Earphone loss post virtual audio experience maintenance method and system based on generative adversarial network

By generating a virtual microphone signal on the lost earphone side using a generative adversarial network and combining it with the signal on the retained earphone side for binaural signal fusion, the problem of auditory discomfort after earphone loss is solved, and a high-fidelity audio experience is restored.

CN122640685APending Publication Date: 2026-08-25CHENGDU RUUSHUI SHISHA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610793598.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies cannot effectively maintain the binaural stereo experience after the headphones are lost, resulting in auditory discomfort and a broken experience. Furthermore, existing solutions fail to reconstruct high-fidelity audio based on the individual acoustic characteristics of the user's ear canal.

Method used

By acquiring the audio recording of the earphone wearing status stored before the earphone was lost, a virtual microphone signal on the side where the earphone was lost is generated using a generative adversarial network. This signal is then combined with the real-time signal collected on the side where the earphone is retained to perform binaural signal fusion and reconstruct the binaural auditory perception state before the earphone was lost.

Benefits of technology

It effectively reconstructs the state of binaural hearing perception after the headphones are lost, overcomes the shortcomings of traditional interpolation methods, and achieves high-fidelity audio experience restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122640685A_ABST
    Figure CN122640685A_ABST
Patent Text Reader

Abstract

The application provides a headphone loss-based virtual audio experience maintenance method and system based on a generative adversarial network, and relates to the technical field of virtual audio experience reconstruction. After a headphone loss event is triggered, a set of wearing state recording segments containing an ear canal audio response segment and an environmental reference audio segment stored before loss is obtained; the ear canal audio response segment is subjected to ear canal impulse response deconvolution processing to obtain a personalized ear canal impulse response; the ear canal impulse response and the environmental reference audio segment are input into a generative adversarial network to generate a virtual microphone pickup signal on the headphone loss side, and are fused with a real-time microphone signal on the reserved side into a binaural virtual sound field synthesizer to obtain a compensation audio signal to be played by a loudspeaker on the loss side, which is emitted by a loudspeaker on the reserved side to reconstruct a binaural hearing perception state in the ear canal of a listener before loss, high-fidelity virtual audio reconstruction is achieved, and the continuity of binaural stereo experience after headphone loss is effectively maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual audio experience reconstruction technology, and more specifically, to a method and system for maintaining a virtual audio experience after headphone loss based on generative adversarial networks. Background Technology

[0002] As the core carrier of personal audio devices, wireless headphones, when lost, not only result in financial loss of the device but also cause significant auditory discomfort and a disruption in the audio experience due to the sudden absence of audio content from one ear, which relies on binaural stereo sound. Current audio compensation solutions for lost headphones primarily focus on device location and retrieval, such as estimating distance using Bluetooth signal strength or reporting location via GPS. However, these solutions do not address maintaining the continuity of the audio experience. Some solutions attempt to switch to mono playback mode after the headphones are lost; however, mono playback cannot reproduce the differences in head-related transfer functions that rely on binaural stereo sound, leading to problems such as sound image localization shift and loss of spatial awareness, resulting in a significant decline in the user's auditory experience. Other solutions utilize a general room impulse response model for simple audio interpolation compensation on the lost side; however, such models are based on statistical averages and do not consider the geometric differences in individual ear canals or the acoustic characteristics under wearing conditions. The compensation signal has a low matching degree with the user's actual ear canal, easily causing timbre distortion and spatial localization errors. There is currently no solution for high-fidelity virtual audio reconstruction based on the user's individual ear canal acoustic characteristics and combined with generative adversarial networks. Summary of the Invention

[0003] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for maintaining a virtual audio experience after headphone loss based on generative adversarial networks, the method comprising:

[0004] After the headphone loss event is triggered, a set of audio recording segments of the headphone wearing status stored before the headphone is lost is obtained. The set of audio recording segments of the headphone wearing status includes the audio response segment in the ear canal recorded by the built-in microphone when both sides of the headphone are worn normally and the environmental reference audio segment recorded by the external microphone just before the headphone is lost.

[0005] The ear canal impulse response deconvolution processing is performed on the ear canal audio response segment in the ear canal recording segment set of the earphone wearing state to obtain the ear canal impulse response under the normal wearing state of both sides of the earphone. The ear canal impulse response includes the ear canal formant frequency set and the ear canal decay time parameter.

[0006] The ear canal impulse response corresponding to the audio response segment in the ear canal and the environmental reference audio segment are input into a pre-constructed generative adversarial network to generate virtual microphone signals on the lost earphone side, thereby obtaining the virtual microphone pickup signal on the lost earphone side. The pre-constructed generative adversarial network includes a generator sub-network and a discriminator sub-network.

[0007] The virtual microphone pickup signal on the lost earphone side and the microphone signal on the retained earphone side, which are collected in real time by the built-in microphone on the retained earphone side, are input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain the compensated audio signal to be played by the speaker on the lost earphone side.

[0008] The compensated audio signal is emitted as a sound wave through the speaker unit on the earphone retaining side, reconstructing the binaural auditory perception state before the earphone was lost in the listener's ear canal.

[0009] Furthermore, embodiments of the present invention also provide a virtual audio experience maintenance system based on generative adversarial networks after headphone loss, comprising:

[0010] A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to perform the above-described method for maintaining a virtual audio experience after headphone loss based on a generative adversarial network by executing the machine-executable instructions.

[0011] Based on the above, after the earphone loss event is triggered, a set of recordings of the wearing state, including audio response segments within the ear canal and environmental reference audio segments stored before the loss, is used. Through deconvolution processing of the ear canal impulse response, a personalized ear canal impulse response, including the set of ear canal formant frequencies and ear canal decay time parameters, is extracted. This captures the acoustic fingerprint characteristics of the user's individual ear canal, avoiding the statistical averaging error of general models. Furthermore, this personalized ear canal impulse response and the environmental reference audio segments are input into a generative adversarial network (GAN) containing generator and discriminator subnetworks. The generator learns the conditional mapping relationship from environmental audio to the signal picked up by the virtual microphone on the lost side, while the discriminator adversarially judges the realism of the generated signal. This game-like training process ensures that the generated virtual microphone pickup signal closely approximates the acoustic signal that should be received by the real lost ear canal in terms of both time-domain waveform details and frequency-domain spatial characteristics, effectively overcoming the inherent defects of traditional interpolation methods in high-fidelity reconstruction. Furthermore, the virtual microphone pick-up signal and the microphone signal collected in real time on the retained side are input into a binaural virtual sound field synthesizer for fusion processing. This allows the compensation audio signal emitted by the retained side speaker to simultaneously reconstruct the sound image localization, spectral characteristics, and spatial auditory cues such as binaural time difference and intensity difference on the lost side within the listener's ear canal. Thus, even without physical equipment, the binaural auditory perception state before the loss of the headphones can be fully restored. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the execution flow of the method for maintaining the virtual audio experience after headphone loss based on generative adversarial networks provided in an embodiment of the present invention.

[0013] Figure 2 This is a schematic diagram of the components of a virtual audio experience maintenance system based on generative adversarial networks provided in an embodiment of the present invention.

[0014] Figure 3 This is a simulation diagram of the verification of the virtual audio experience maintenance method based on generative adversarial networks provided in the embodiments of the present invention after headphone loss.

[0015] Figure 4 This is a schematic diagram of the front-end interface of the pairing terminal during the virtual audio maintenance process provided in an embodiment of the present invention. Detailed Implementation

[0016] Figure 1 This is a flowchart illustrating a method for maintaining a virtual audio experience after headphone loss based on a generative adversarial network, according to an embodiment of the present invention. This embodiment uses a pair of wireless stereo headphones with active noise cancellation and ambient sound pass-through as the target headphones. Both sides of the target headphones have built-in microphones, speakers, and inertial sensing components. The target headphones are paired with a smartphone as a terminal, maintaining data interaction via a Bluetooth Low Energy communication link. In this embodiment, audio data acquisition and acoustic feature extraction are all completed locally on the headphones, avoiding risks of cloud transmission and leakage of personal privacy data, and complying with relevant legal requirements for device-side data processing. The following describes the implementation of the technical solution in detail, using the scenario of the right headphone being lost while the left headphone is retained.

[0017] Step S110: After the headphone loss event is triggered, obtain the set of headphone wearing status recording segments stored before the headphone is lost. The headphone wearing status recording segment set includes the ear canal audio response segment recorded by the built-in microphone when both sides of the headphone are worn normally and the environmental reference audio segment recorded by the external microphone just before the headphone is lost.

[0018] The target earphone continuously runs a wearing status monitoring logic during normal wear. This logic uses an infrared proximity sensor and a capacitive touch sensor to jointly determine whether the earphone is in an in-ear wearing state. While in wear mode, the earphone's built-in microphone continuously records the audio response signal within the ear canal at a preset sampling rate, while an external microphone continuously records the environmental reference audio signal of the external space. Both audio signals are stored in the earphone's built-in flash memory in the form of a circular buffer, with the buffer length covering a preset maximum backtracking time. When the wearing status monitoring logic determines that the earphone has switched from a wearing state to a disengaged state and the disengagement duration exceeds a preset disengagement determination time, an earphone loss event is triggered. After the earphone loss event is triggered, the earphone firmware immediately freezes all the ear canal audio response signal segments and environmental reference audio signal segments stored in the circular buffer and transfers them to a non-volatile storage area, forming a set of earphone wearing status recording segments. The collection of recording clips in the earphone wearing state includes the ear canal audio response clip, which contains the dual-channel ear canal audio response signals recorded by the built-in microphones of the left and right ears when both earphones are worn normally, and the environmental reference audio clip, which contains the mono or stereo environmental audio signals recorded by the external microphones just before the earphones are lost.

[0019] Step S120: Perform ear canal impulse response deconvolution processing on the ear canal audio response segments in the ear canal wearing state recording segment set to obtain the ear canal impulse response under normal wearing state of both sides of the earphone. The ear canal impulse response includes the ear canal resonant frequency set and the ear canal decay time parameter.

[0020] Step S121: Perform excitation signal and response signal separation processing on the audio response segment in the ear canal, and extract the excitation signal component and the ear canal response signal component from the audio response segment in the ear canal.

[0021] The excitation signal in the audio response segment within the ear canal is a known test signal or music signal played by the headphone speaker, which is recorded in the headphone firmware before playback. Using the recorded excitation signal as the reference input and the audio response segment within the ear canal as the observed output, a Wiener deconvolution filter is constructed. The frequency domain transfer function of the Wiener deconvolution filter is calculated by dividing the excitation signal power spectral density by the sum of the excitation signal power spectral density and the noise power spectral density, and then multiplying this by the ratio of the cross-power spectral density of the excitation signal and the observed output to the excitation signal power spectral density. The noise power spectral density is the background noise power spectrum measured by the headphone's built-in microphone in a silent environment.

[0022] Step S122: Perform all-pole linear predictive coding analysis on the separated ear canal response signal components to extract the ear canal formant frequency set.

[0023] Linear predictive coding analysis was performed on the components of the ear canal response signal. The linear prediction order was set based on the audible frequency range and the acoustic characteristics of the ear canal. The linear predictive coding analysis used the Levinson-Dubin recursive algorithm to solve the autocorrelation equations for the linear prediction coefficients. These linear prediction coefficients were then converted into polynomial coefficients, and the conjugate complex roots of the polynomials were solved. Each pair of conjugate complex roots corresponds to an ear canal formant, with the formant frequency taken as f = (theta * fs) / (2 * pi), where theta is the complex root phase angle and fs is the sampling frequency; the formant amplitude is taken as the complex root modulus. The frequencies and amplitudes of all formants constitute the ear canal formant frequency set.

[0024] Step S123: Perform energy attenuation envelope fitting on the separated ear canal response signal components to extract the ear canal attenuation time parameter.

[0025] The Hilbert transform is applied to the ear canal response signal components to extract the signal envelope. The natural logarithm of the signal envelope is taken; the slope of the decrease in the logarithm of the envelope over time is the attenuation rate. The ear canal attenuation time parameter is the time required for the signal energy to decay from its initial value to one millionth of its initial value, calculated using the formula T60 = 3 / (alpha*fs), where alpha is the attenuation rate and fs is the sampling frequency in seconds. The attenuation time is calculated for different frequency bands in the ear canal response signal components to obtain the ear canal attenuation time parameter as a function of frequency.

[0026] Step S124: Combine the set of ear canal resonant frequencies and the ear canal decay time parameters into a parameterized representation of the ear canal impulse response under normal wearing conditions on both sides of the earphone.

[0027] The set of ear canal resonance frequencies extracted in step S122 and the ear canal decay time parameters extracted in step S123 are arranged from low to high according to the resonance frequencies. Each resonance peak corresponds to a frequency value, an amplitude value, and a decay time value. The above parameters are stored in a structured data format to form a parameterized representation of the ear canal impulse response, and stored in the non-volatile storage area of ​​the earphone.

[0028] Step S130: Input the ear canal impulse response corresponding to the audio response segment in the ear canal and the environmental reference audio segment into a pre-constructed generative adversarial network to generate a virtual microphone signal on the lost earphone side, thereby obtaining the virtual microphone pickup signal on the lost earphone side. The pre-constructed generative adversarial network includes a generator sub-network and a discriminator sub-network.

[0029] Step S131: Extract the ear canal formant frequency set from the ear canal impulse response, encode the frequency position and amplitude value of each formant in the ear canal formant frequency set into a formant structure feature sequence, input the formant structure feature sequence into the generator subnetwork of the pre-constructed generative adversarial network, and perform nonlinear mapping processing on the formant structure feature sequence through the ear canal feature embedding layer to obtain the latent encoding vector of ear canal acoustic characteristics.

[0030] The formant structure feature sequence is encoded by concatenating the frequency and amplitude values ​​of each formant into a two-dimensional vector. All formant vectors are arranged in ascending frequency order to form the formant structure feature sequence. The ear canal feature embedding layer consists of alternating fully connected layers and batch normalization layers. The activation function of the fully connected layers is a modified linear unit with leakage. The output of the ear canal feature embedding layer is a fixed-dimensional latent encoding vector of ear canal acoustic characteristics. This latent encoding vector will be injected as conditional information into the subsequent processing layers of the generator subnetwork.

[0031] Step S132: Extract the last complete audio frame unit before the moment the earphone is lost from the environmental reference audio segment, perform multi-resolution time-frequency decomposition processing on the last complete audio frame unit to obtain a set of audio frame multi-scale spectrum layers containing different time granularities and frequency resolutions, input the set of audio frame multi-scale spectrum layers into the environmental context awareness branch of the generator sub-network, and perform causal temporal feature aggregation processing on the set of audio frame multi-scale spectrum layers through a sequence of ordered dilated convolutional layers to obtain an environmental audio context feature map.

[0032] The final complete audio frame unit is taken from the last complete audio frame before the moment the headphones are lost, and the audio frame length is a preset frame length. Multi-resolution time-frequency decomposition uses multiple short-time Fourier transforms with different window lengths: short, medium, and long windows. Each window length corresponds to a spectral layer, and multiple spectral layers are stacked along the depth direction to form a multi-scale spectral layer set for the audio frame. The environmental context-aware branch consists of a sequence of ordered dilated convolutional layers. The dilation rate of each dilated convolution increases in powers of 2, starting from dilation rate 1 and increasing layer by layer to the preset maximum dilation rate. The dilated convolutional layers use causal convolution operations to ensure that the output at the current time depends only on the input at the current time and past time steps. The output of the ordered dilated convolutional layer sequence is the environmental audio context feature map.

[0033] Step S133: The latent encoding vector of the ear canal acoustic characteristics and the environmental audio context feature map are concatenated along the channel dimension in the feature fusion layer of the generator sub-network. The concatenated features are then subjected to gated fusion processing to obtain the conditionally generated feature representation of the acoustic scene on the side where the earphone is lost.

[0034] The latent encoding vector of ear canal acoustic characteristics is first mapped to a two-dimensional feature map with the same spatial dimension as the ambient audio context feature map through a fully connected layer. This feature map is then concatenated with the ambient audio context feature map along the channel dimension. Gated fusion processing uses gated linear units to divide the concatenated feature into two equal parts along the channel dimension. One part serves as a gating signal, mapped to gating weights between 0 and 1 using a sigmoid activation function. The other part serves as the feature to be modulated. The two parts are multiplied element-wise to obtain the fused feature. The conditionally generated feature representation is the fused feature.

[0035] Step S134: Input the conditional generation feature representation of the acoustic scene on the side where the earphone is lost into the stepwise upsampling signal reconstruction module of the generator sub-network, and gradually restore the temporal resolution of the signal through a series of deconvolutional layers to generate candidate virtual microphone pickup signals on the side where the earphone is lost.

[0036] The stepwise upsampling signal reconstruction module consists of cascaded deconvolutional layers. Each deconvolutional layer has an upsampling rate of 2, a preset kernel size, and a leaky modified linear unit (LMU) activation function. The final deconvolutional layer has one output channel, uses a hyperbolic tangent activation function, and has an output value range of -1 to +1. The output of the deconvolutional layer sequence is the time-domain waveform sampling point sequence of the candidate virtual microphone's picked-up signal.

[0037] Step S135: Perform frequency band energy distribution analysis on the candidate virtual microphone pickup signal on the lost earphone side to extract the candidate signal frequency band energy distribution vector; perform frequency band energy distribution analysis on the retained side microphone signal collected in real time by the built-in microphone on the retained earphone side to extract the reference signal frequency band energy distribution vector.

[0038] The frequency band energy distribution analysis process uses a one-third octave band filter bank to decompose the audio signal into multiple frequency bands. The root mean square value of the signal sampling points within each frequency band is calculated as the energy value of that band. The energy values ​​of all frequency bands are arranged into a vector to obtain the frequency band energy distribution vector. The candidate signal frequency band energy distribution vector is denoted as V1, and the reference signal frequency band energy distribution vector is denoted as V2.

[0039] Step S136: Calculate the frequency band energy distribution deviation between the candidate signal frequency band energy distribution vector and the reference signal frequency band energy distribution vector, and add the frequency band energy distribution deviation as a frequency domain consistency constraint term to the loss calculation process of the generator sub-network.

[0040] The frequency band energy distribution deviation is taken as the mean square error of V1 and V2. The mean square error is equal to the sum of the squares of the energy differences of each frequency band divided by the total number of frequency bands. The frequency domain consistency constraint term is taken as this mean square error multiplied by a preset weighting coefficient. This constraint term is added to the adversarial loss term of the generator sub-network to form the total loss function of the generator sub-network.

[0041] Step S137: During the adversarial alternation training of the generator subnetwork and the discriminator subnetwork, the discriminator subnetwork receives the candidate virtual microphone pickup signal and the real lost-side microphone acquisition signal and outputs the discrimination confidence score. The generator subnetwork jointly updates the network weights based on the discrimination confidence score and the frequency domain consistency constraint term. After the adversarial training converges, the generator subnetwork outputs the virtual microphone pickup signal of the lost earphone side.

[0042] The discriminator subnetwork is a convolutional neural network classifier. The input is a mono audio signal waveform, and the output is a discrimination confidence score between 0 and 1. During training, the discriminator and generator subnetworks alternately update their weights. The loss function of the discriminator subnetwork is the binary cross-entropy of the real signal discrimination confidence score and 1, plus the binary cross-entropy of the generated signal discrimination confidence score and 0. The loss function of the generator subnetwork is the binary cross-entropy of the generated signal discrimination confidence score and 1, plus a frequency domain consistency constraint term. An adaptive moment estimation optimizer is used for weight updates. Training converges when the preset number of iterations is reached or the loss value no longer decreases. After convergence, the candidate virtual microphone pickup signal output by the generator subnetwork is the virtual microphone pickup signal on the side where the earphone is lost.

[0043] Step S140: Input the virtual microphone pickup signal on the lost earphone side and the microphone signal on the retained earphone side, which are collected in real time by the built-in microphone on the retained earphone side, into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain the compensated audio signal to be played by the speaker on the lost earphone side.

[0044] Step S141: Perform short-time Fourier transform processing on the virtual microphone pickup signal on the lost side of the earphone to obtain the virtual side time spectrum; perform short-time Fourier transform processing on the reserved side microphone signal collected in real time by the built-in microphone on the reserved side of the earphone to obtain the reserved side time spectrum.

[0045] The window function for the short-time Fourier transform is the Hanning window, with preset window length and step size. Short-time Fourier transforms are performed on the virtual microphone pickup signal and the retained microphone signal respectively, yielding their respective complex time-frequency spectra. The virtual side time-frequency spectrum is denoted as S1(t, f), and the retained side time-frequency spectrum is denoted as S2(t, f), where t is the time frame index and f is the frequency index.

[0046] Step S142: Perform cross-power spectrum phase analysis on the virtual side time spectrum and the reserved side time spectrum in each time-frequency unit, extract the binaural phase difference corresponding to each time-frequency unit, and aggregate the binaural phase differences of all time-frequency units into a binaural phase difference time-frequency diagram.

[0047] The binaural phase difference is taken as the phase difference between S1(t,f) and S2(t,f), and is calculated as angle(S1(t,f)*conj(S2(t,f))). The phase difference is calculated for all time frames and all frequency indices to construct a binaural phase difference time-frequency diagram.

[0048] Step S143: Perform energy ratio analysis on the virtual side time spectrum and the reserved side time spectrum in each time-frequency unit, extract the binaural energy ratio corresponding to each time-frequency unit, and aggregate the binaural energy ratios of all time-frequency units into a binaural energy ratio time-frequency diagram.

[0049] The binaural energy ratio is taken as the ratio of |S1(t,f)|^2 / |S2(t,f)|^2. The energy ratio is calculated for all time frames and all frequency indices to construct a binaural energy ratio time-frequency diagram.

[0050] Step S144: Input the time-frequency diagram of the binaural phase difference and the time-frequency diagram of the binaural energy ratio into the spatial orientation resolution layer of the pre-constructed binaural virtual sound field synthesizer. The spatial orientation resolution layer maps the binaural cues into azimuth and elevation parameters of the virtual sound source relative to the center of the listener's head, thereby obtaining the spatial orientation parameter sequence of the virtual sound source on the side where the headphones are lost.

[0051] The spatial orientation resolution layer consists of a deep neural network. The input is a multi-channel feature map obtained by concatenating the binaural phase difference time-frequency map and the binaural energy ratio time-frequency map along the channel dimension. The network structure includes convolutional layers, pooling layers, and fully connected layers. The network output is a two-dimensional vector containing azimuth and elevation parameters. After frame-by-frame processing, a sequence of spatial orientation parameters is obtained, with each frame containing both azimuth and elevation parameters.

[0052] Step S145: Input the virtual side-time spectrum and the spatial orientation parameter sequence into the sound field rendering layer of the pre-constructed binaural virtual sound field synthesizer. Through the sound field rendering layer, call the head-related transmission characteristic dataset that matches the spatial orientation parameter sequence, and apply the corresponding head-related transmission gain and head-related transmission phase shift to each time-frequency unit in the virtual side-time spectrum to obtain the virtual side-rendered time spectrum carrying spatial orientation clues.

[0053] The head-related transfer characteristic dataset is a pre-stored database of head-related transfer functions, indexed by azimuth and elevation. For the spatial azimuth parameters of each frame, the head-related transfer function corresponding to the nearest azimuth and elevation angles is retrieved from the database. This head-related transfer function is a complex frequency response, with its amplitude response being the head-related transfer gain and its phase response being the head-related transfer phase shift. The complex amplitude of each frequency index in the virtual-side time-spectrum for the corresponding frame is multiplied by the head-related transfer gain, and the head-related transfer phase shift is added to obtain the virtual-side rendering time-spectrum.

[0054] Step S146: Convert the virtual side rendering time spectrum into a time domain signal through inverse short-time Fourier transform to obtain the headphone loss side spatial audio signal carrying spatial orientation clues. Perform perceptual loudness measurement processing on the headphone loss side spatial audio signal to obtain the virtual side perceptual loudness value. Perform perceptual loudness measurement processing on the retained side microphone signal to obtain the retained side perceptual loudness value. Calculate the loudness difference between the virtual side perceptual loudness value and the retained side perceptual loudness value.

[0055] The inverse short-time Fourier transform employs a superposition and addition method. Perceived loudness measurement uses a loudness measurement algorithm based on K-weighted filtering and energy integration, with the calculated loudness value expressed as a square. The loudness difference is calculated by subtracting the retained-side perceived loudness value from the virtual-side perceived loudness value.

[0056] Step S147: Perform gain adjustment processing on the spatial audio signal of the headphone loss side carrying spatial orientation clues according to the loudness difference amount to obtain the loudness-adjusted spatial audio signal of the headphone loss side, and use the loudness-adjusted spatial audio signal of the headphone loss side as the compensation audio signal to be played by the headphone loss side speaker.

[0057] The gain adjustment is set to -delta_L*k, where delta_L is the loudness difference and k is the preset gain coefficient. Each sample point of the spatial audio signal on the headphone loss side carrying spatial orientation clues is multiplied by 10^(gain / 20), where gain is the gain adjustment amount, to obtain the loudness-adjusted spatial audio signal on the headphone loss side.

[0058] Step S150: The compensated audio signal is emitted as a sound wave through the speaker unit on the earphone retaining side to reconstruct the binaural auditory perception state before the earphone was lost in the listener's ear canal.

[0059] After receiving the compensation audio signal, the speaker unit on the retained earphone side converts the digital audio signal into an analog electrical signal via a digital-to-analog converter. This analog signal is then amplified by a power amplifier to drive the speaker diaphragm to vibrate, generating sound waves. These sound waves enter the ear canal through the airflow on the retained ear side, while some are conducted through the skull to the cochlea on the opposite side. The spatial audio information from the lost earphone side carried in the compensation audio signal is fused with the normally transmitted ambient sound information from the retained ear side in the auditory center, reconstructing the binaural auditory perception state before the earphone loss.

[0060] In the above embodiments, based on Figure 2The architecture shown in this embodiment further describes the specific interaction logic between the hardware components. The acquisition of the in-ear canal audio response segment and the environmental reference audio segment relies on the built-in microphones on both sides of the earphone and the external microphone. These microphones operate continuously when the earphone is in normal wearing condition, and their recorded signals are sent to the audio processing unit of the main controller via an analog-to-digital converter. When a loss event is triggered, the main controller reads the pre-stored set of earphone wearing state recording segments from non-volatile storage media and performs deconvolution processing of the ear canal impulse response to extract the set of ear canal formant frequencies and ear canal decay time parameters. Subsequently, the main controller sends the parameterized ear canal impulse response and the contextual features extracted from the environmental reference audio segment together into a pre-built generative adversarial network (GAN). This GAN can be implemented at the hardware level by a dedicated neural network accelerator or digital signal processor, and its generator subnetwork is responsible for synthesizing the virtual microphone pickup signal on the side where the earphone is lost. The signal picked up by the virtual microphone is further combined with the signal from the microphone on the earphone's earpiece side, which is collected in real time by the built-in microphone on the earphone's earpiece side, and then fed into a pre-built binaural virtual sound field synthesizer module. In the binaural virtual sound field synthesizer, the binaural phase difference and energy ratio analysis, the application of the head-related transfer function, and the loudness adjustment are completed. The resulting compensated audio signal is then converted from digital to analog and emitted by the speaker unit on the earphone's earpiece side, thereby reconstructing the binaural auditory perception state before the loss in the listener's ear canal.

[0061] Step S210: Obtain multiple sets of audio response segments in the ear canal with different wearing tightness recorded by the built-in microphone under normal wearing conditions on both sides of the earphone, and the corresponding wearing tightness labels, to construct a wearing tightness condition training dataset. Perform ear canal impulse response deconvolution processing on each set of audio response segments in the wearing tightness condition training dataset to obtain ear canal impulse response samples corresponding to the wearing tightness labels. Extract ear canal formant frequency set samples and ear canal decay time parameter samples from the ear canal impulse response samples.

[0062] The training dataset for wearing tightness conditions was constructed through a controlled wearing experiment. Subjects wore the headphones at preset tightness levels. Tightness was quantified by measuring the contact pressure between the headphones and the ear canal wall using pressure sensors. The pressure values ​​were divided into multiple discrete intervals, each corresponding to a tightness level label. At each tightness level, the headphones' built-in speaker played a logarithmic sweep signal. The sweep signal frequency increased logarithmically from low to high frequencies, and the sweep duration covered the entire audible frequency range. The headphones' built-in microphone simultaneously recorded the sound wave response signal within the ear canal; this recorded signal became the sample audio response segment within the ear canal at that tightness level. Each sample set was accompanied by a corresponding tightness level label. For each sample group, deconvolution processing of the ear canal impulse response is performed. The Wiener deconvolution algorithm is used, and its frequency domain expression is H(f) = Pxy(f) / (Pxx(f) + lambda), where f is the frequency, Pxy is the cross-power spectral density of the excitation signal and the response signal, Pxx is the auto-power spectral density of the excitation signal, and lambda is the noise regularization coefficient. The ear canal impulse response samples obtained by deconvolution are then analyzed by linear predictive coding to extract the ear canal formant frequency set samples, and the ear canal decay time parameter samples are extracted by the Schroeder inverse integral method. The Schroeder inverse integral method is calculated as EDC(t) = integral from t to infinity of h^2(tau) dtau, where h is the impulse response, EDC is the energy decay curve, and the decay time is the time required for EDC to decay from 0dB to -60dB.

[0063] Step S220: Perform resonance peak shift statistical analysis on the ear canal resonance peak frequency set samples corresponding to different wearing tightness labels, and extract the shift direction and shift amplitude law of resonance peak frequency with the wearing tightness; and perform attenuation time statistical analysis on the ear canal attenuation time parameter samples corresponding to different wearing tightness labels, and extract the extension trend and compression trend law of attenuation time parameter with the wearing tightness.

[0064] The statistical analysis of resonant peak offset uses the set of ear canal resonant peak frequencies at the tightest wearing level as a benchmark. The offset of each resonant peak frequency relative to the benchmark value is calculated for each other tightness level. The offset ΔF = Fk - Fref, where Fk is the resonant peak frequency at the k-th tightness level and Fref is the corresponding resonant peak frequency at the tightest level. Linear regression is performed on the offset for each tightness level, with the regression equation ΔF = a × P + b, where P is the wearing pressure value, a is the offset slope parameter, and b is the intercept parameter. The sign of the offset slope parameter a indicates the offset direction, and the absolute value of a indicates the offset amplitude. The above regression analysis is performed on all resonant peaks to obtain the offset direction and amplitude parameters for each resonant peak. The statistical analysis of decay time uses the ear canal decay time parameter at the tightest wearing level as a benchmark. The relative change rate of decay time for each frequency band at each tightness level is calculated as r = (Tk - Tref) / Tref, where Tk is the decay time at the k-th tightness level and Tref is the decay time for the corresponding frequency band at the tightest level. A linear regression was performed on the relative rate of change, with the regression equation being r = c × P + d, where c is the slope parameter of the decay change and d is the intercept parameter. A positive c indicates that the decay time increases as the wear becomes looser, while a negative c indicates that the decay time is compressed as the wear becomes looser.

[0065] Step S230: After the headphone loss event is triggered, the current ear canal response signal in the ear canal on the ear canal is collected in real time by the built-in microphone on the ear canal retention side. The current ear canal response signal is processed by formant detection and decay time estimation to obtain the current ear canal formant frequency and the current ear canal decay time.

[0066] The built-in microphone on the earphone retainer side continuously collects sound wave signals from the ear canal on the retainer side at a preset sampling frequency. Simultaneously, the earphone retainer side speaker plays a short-time logarithmic sweep signal as a probe signal. The parameters of the sweep signal are consistent with those used in step S210 when collecting the audio response segment sample from the ear canal. The built-in microphone on the earphone retainer side records the sound wave response signal from the ear canal during the probe signal playback; this signal is the current ear canal response signal. The current ear canal response signal undergoes the same Wiener deconvolution processing as in step S210 to obtain the current ear canal impulse response. Then, linear predictive coding analysis is performed on the current ear canal impulse response to obtain the current ear canal formant frequency set. Linear predictive coding analysis uses the Levinson-Dubin recursive algorithm to solve the Yule-Walker equation to obtain linear prediction coefficients. After converting the linear prediction coefficients into polynomial coefficients, roots are calculated. Each pair of conjugate complex roots corresponds to a formant, and the formant frequency is f = (theta*fs) / (2*pi), where theta is the complex root phase angle and fs is the sampling frequency. The Schroeder inverse integration method is used to obtain the energy decay curve of the current ear canal impulse response. The time required for the energy decay curve to decay from 0dB to -60dB is read as the current ear canal decay time parameter.

[0067] Step S240: Based on the offset direction and offset amplitude of the resonant frequency as a function of wearing tightness, the estimated current wearing tightness of the earphone retaining side is calculated. Based on the estimated current wearing tightness of the earphone retaining side, a subset of ear canal impulse response samples matching the current wearing tightness is selected from the wearing tightness condition training dataset. The ear canal impulse response sample subset is used to perform online fine-tuning training of the generator subnetwork of the pre-constructed generative adversarial network.

[0068] The inversion calculation process matches the frequency value of each resonant in the current ear canal resonant frequency set with the regression equation of each resonant in step S220. The matching method involves substituting the frequency offset ΔFcur=Fcur-Fref of each resonant into the regression equation ΔF=a×P+b for each tightness level, calculating the corresponding wearing pressure value P=(ΔFcur-b) / a. The weighted average of all resonant pressure values ​​is taken as the estimated tightness of the current wearing condition, with the weight taken as the energy normalized value of each resonant, and higher weights assigned to resonants with higher energy. The estimated tightness of the current wearing condition is compared with the pressure value range corresponding to each tightness level label, and the best-matching tightness level label is selected. All ear canal impulse response samples corresponding to this label are selected from the tightness condition training dataset to form a sample subset. This sample subset is used to fine-tune the generator subnetwork online. The fine-tuning training uses a mini-batch gradient descent method, with the batch size being the number of samples in the sample subset. The loss function is the total loss function of the generator subnetwork defined in step S130. The learning rate is set to a preset proportion of the pre-training learning rate, and the number of iterations is set to a preset number of fine-tuning iterations. After fine-tuning training, the generator subnetwork weights are obtained to adapt to the current tightness of the fit.

[0069] Step S250: After online fine-tuning training is completed, the step of inputting the ear canal impulse response corresponding to the audio response segment in the ear canal and the environmental reference audio segment into the pre-constructed generative adversarial network to generate virtual microphone signals on the lost side of the earphone is executed again to obtain a virtual microphone pickup signal adapted to the listener's current wearing tightness. The virtual microphone pickup signal adapted to the listener's current wearing tightness and the microphone signal of the retained side built-in microphone of the earphone are input into the pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain a compensated audio signal adapted to changes in wearing tightness. The compensated audio signal adapted to changes in wearing tightness is emitted as sound waves through the speaker unit on the retained side of the earphone.

[0070] The generator subnetwork in step S130 is replaced with the finely tuned generator subnetwork, and all processing steps from S130 to S150 are re-executed to finally generate a compensated audio signal that adapts to changes in wearing tightness and is emitted by the speaker on the earphone's retaining side.

[0071] Step S310: Perform auditory filter bank analysis on the virtual microphone pickup signal on the lost earphone side, decomposing the virtual microphone pickup signal on the lost earphone side into multiple auditory sub-band signals through a gamma-tonal filter bank that simulates the frequency selectivity characteristics of the basilar membrane of the human ear; perform auditory filter bank analysis on the microphone signal of the retained earphone side built-in microphone acquired in real time, decomposing the retained earphone side microphone signal into multiple auditory sub-band signals through the gamma-tonal filter bank; perform sub-band cross-correlation analysis on each auditory sub-band signal on the lost earphone side and the corresponding auditory sub-band signal on the retained earphone side, and extract the sub-band binaural time difference parameter and the sub-band binaural cross-correlation coefficient for each auditory sub-band.

[0072] The impulse response expression of the gamma-pass filter bank is g(t) = t^(N-1)*exp(-2*pi*b*t)*cos(2*pi*fc*t+phi), where N is the filter order, b is the equivalent rectangular bandwidth parameter, fc is the center frequency, and phi is the phase parameter. The center frequency fc is uniformly distributed across the audible frequency range according to the equivalent rectangular bandwidth scale. The signal picked up by the virtual microphone and the signal from the microphone on the retained side are passed through each gamma-pass filter, and the output of the filter is the corresponding auditory sub-band signal. Cross-correlation is performed on each pair of sub-band signals with the same center frequency, and the cross-correlation function expression is R(tau) = integrallofx(t)*y(t+tau)dt, where x is the auditory sub-band signal on the lost side of the earphone, and y is the auditory sub-band signal on the retained side of the earphone. The delay time taumax corresponding to the peak value of the cross-correlation function is the binaural time difference parameter of that sub-band. The peak normalized amplitude of the cross-correlation function rho=R(taumax) / sqrt(Rxx(0)*Ryy(0)) is the subband two-ear cross-correlation coefficient of this subband, where Rxx and Ryy are the autocorrelation functions of x and y, respectively.

[0073] Step S320: Evaluate the reliability of the binaural cue for each auditory subband based on the binaural cross-correlation coefficient of the subband. Mark auditory subbands with binaural cross-correlation coefficients lower than a preset correlation threshold as first reliable subbands, and mark auditory subbands with binaural cross-correlation coefficients not lower than a preset correlation threshold as second reliable subbands.

[0074] The binaural cross-correlation coefficient ρ of each subband is compared with a preset correlation threshold ρth. If ρ < ρth, the subband is marked as the first reliable subband, indicating that the correlation of the binaural signals in this subband is weak and the reliability of the binaural cue is low. If ρ ≥ ρth, the subband is marked as the second reliable subband, indicating that the binaural cue is reliable.

[0075] Step S330: For the first reliable sub-band, perform inter-sub-band binaural cue interpolation processing using the binaural time difference parameters of the adjacent second reliable sub-band to obtain the inferred binaural time difference parameters of the first reliable sub-band; and, for the earphone loss side auditory sub-band signal corresponding to the first reliable sub-band, perform time delay compensation and phase alignment processing using the inferred binaural time difference parameters to obtain the binaural cue corrected first reliable sub-band signal.

[0076] For each first reliable sub-band, search for the nearest second reliable sub-band above and below its frequency. Let the center frequency of the lower second reliable sub-band be f1, and the binaural time difference be τ1; let the center frequency of the upper second reliable sub-band be f2, and the binaural time difference be τ2. The center frequency of the first reliable sub-band is f. The binaural time difference parameter τest is calculated by linear interpolation: τest = τ1 + (τ2 - τ1) × (f - f1) / (f2 - f1). Time delay compensation is performed on the auditory sub-band signal on the earphone loss side corresponding to the first reliable sub-band. Time delay compensation is achieved through linear phase rotation in the Fourier transform domain. The compensated frequency domain signal Xcomp(f) = X(f) × exp(-j × 2πf × τest), where X(f) is the original frequency domain signal, and j is the imaginary unit. The compensated frequency domain signal is then subjected to inverse Fourier transform to obtain the time domain signal, which is the first reliable sub-band signal after binaural cue correction.

[0077] Step S340: Perform sub-band synthesis and reconstruction processing on the first reliable sub-band signal after binaural cue correction and the earphone loss side auditory sub-band signal corresponding to the second reliable sub-band to obtain a virtual microphone pickup signal with enhanced binaural cue consistency. Input the virtual microphone pickup signal and the retention side microphone signal collected in real time by the built-in microphone on the earphone retention side into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain a compensated audio signal with enhanced binaural cue consistency. Emitter the compensated audio signal through the speaker unit on the earphone retention side.

[0078] The earphone loss side auditory sub-band signals of all sub-bands are reconstructed through sub-band synthesis. Sub-band synthesis reconstruction involves passing each sub-band signal through a gamma-pass synthesis filter corresponding to its center frequency, and then summing the outputs of all filters. The impulse response of the synthesis filter is the same as that of the decomposition filter. The time-domain signal obtained from the synthesis reconstruction is the virtual microphone pickup signal with enhanced binaural cue consistency. This virtual microphone pickup signal and the retained side microphone signal are then processed according to steps S140 to S150 for binaural signal fusion and sound wave emission.

[0079] Step S410: Acquire a dynamic environmental audio segment recorded by the external microphone when the listener slowly rotates their head while wearing both headphones normally. The dynamic environmental audio segment includes binaural cues indicating the continuous change in the relative azimuth of environmental sound sources during the listener's head rotation. Perform frame-by-frame sound source azimuth estimation processing on the dynamic environmental audio segment, extract the instantaneous sound source azimuth angle sequence and instantaneous sound source elevation angle sequence corresponding to each frame, and use the instantaneous sound source azimuth angle sequence and instantaneous sound source elevation angle sequence as the dynamic azimuth ground truth label.

[0080] Sound source location estimation employs a deep learning-based sound source localization algorithm. The input is the short-time Fourier transform spectrum of the binaural microphone signal, and the output is the azimuth and elevation angles. The sound source localization network structure is a convolutional recurrent neural network. Convolutional layers extract local time-frequency features, recurrent layers model time dependencies, and fully connected output layers output the classification probability distributions of azimuth and elevation angles. During training, a cross-entropy loss function is used, and the training data consists of binaural recordings with precise azimuth annotations collected in an anechoic chamber. Dynamic environmental audio segments are input frame-by-frame into the sound source localization network. Each frame outputs the probability distributions of azimuth and elevation angles, and the azimuth and elevation angles corresponding to the maximum probabilities are taken as the instantaneous sound source azimuth and elevation angles. The instantaneous azimuth angles of all frames constitute the instantaneous sound source azimuth angle sequence, and the instantaneous elevation angles of all frames constitute the instantaneous sound source elevation angle sequence. These two sequences serve as ground truth labels for dynamic location.

[0081] Step S420: Using the microphone signal channel corresponding to the lost earphone side in the dynamic environmental audio segment as the training target signal and the microphone signal channel corresponding to the retained earphone side as the training input signal, a training data pair for generating the lost earphone signal is constructed. Based on the training data pair for generating the lost earphone signal and the dynamic azimuth ground truth label, an azimuth-aware auxiliary training branch is added to the generator subnetwork of the pre-constructed generative adversarial network. An azimuth prediction subnetwork is derived from the intermediate feature layer of the generator subnetwork through the azimuth-aware auxiliary training branch. The azimuth prediction subnetwork outputs a predicted azimuth sequence and a predicted elevation sequence.

[0082] The input to the training data pair for generating the signal from the lost earphone side is the mono audio signal from the microphone channel on the retained earphone side, and the target output is the mono audio signal from the microphone channel on the lost earphone side. The orientation awareness-assisted training branch is derived from the conditional generation feature representation layer of the generator sub-network. This layer is the feature layer resulting from the fusion of the latent encoding vector of the ear canal acoustic characteristics and the environmental audio context feature map in step S133. The orientation prediction sub-network contains a global average pooling layer in the time dimension and two parallel fully connected output layers. One fully connected output layer outputs the predicted azimuth angle, and the other fully connected output layer outputs the predicted elevation angle. The global average pooling layer compresses the conditional generation feature representation into a fixed-dimensional feature vector in the time dimension. The predicted azimuth angle and predicted elevation angle are output by a Softmax classifier, with the number of classification categories taking the total number of categories corresponding to the preset azimuth angle resolution and elevation angle resolution. After frame-by-frame processing, the predicted azimuth angle sequence and predicted elevation angle sequence are obtained.

[0083] Step S430: Calculate the azimuth prediction deviation between the predicted azimuth sequence and the instantaneous sound source azimuth sequence in the dynamic azimuth ground truth label; calculate the elevation prediction deviation between the predicted elevation sequence and the instantaneous sound source elevation sequence in the dynamic azimuth ground truth label; use the azimuth prediction deviation and the elevation prediction deviation as azimuth perception loss terms; perform weighted joint training processing on the azimuth perception loss terms and the adversarial loss terms of the generator sub-network to obtain the weight parameters of the azimuth perception enhanced generator sub-network.

[0084] The azimuth prediction bias is taken as the classification cross-entropy loss of the predicted azimuth, calculated as Laz = -Σqaz × log(paz), where qaz is the one-hot encoding of the azimuth in the dynamic azimuth ground truth label, and paz is the softmax probability of the predicted azimuth. The elevation prediction bias is similarly taken as Lel = -Σqel × log(pel). The azimuth-aware loss term Lap = Laz + Lel. The adversarial loss term Ladv of the generator subnetwork is taken as the generator adversarial loss defined in step S130. The total loss function Ltotal = Ladv + λ × Lap, where λ is the azimuth-aware loss weight coefficient. An adaptive moment estimation optimizer is used to backpropagate and train the total loss function. After training, the weight parameters of the azimuth-aware enhanced generator subnetwork are obtained.

[0085] Step S440: After the headphone loss event is triggered, the generator sub-network for enhanced orientation is invoked to generate virtual microphone signals on the lost headphone side of the environmental reference audio segment, resulting in a virtual microphone pickup signal that retains the orientation information of the sound source. The virtual microphone pickup signal that retains the orientation information of the sound source is then input into a pre-constructed binaural virtual sound field synthesizer along with the microphone signal of the retained headphone side built-in microphone, which is collected in real time. This binaural signal fusion processing is then performed to obtain a compensated audio signal that retains the orientation information of the sound source. Finally, the compensated audio signal that retains the orientation information of the sound source is emitted as a sound wave through the speaker unit on the retained headphone side.

[0086] The generator subnetwork in step S130 is replaced with a generator subnetwork that enhances orientation awareness, and the processing flow from steps S131 to S137 is executed to generate a virtual microphone pickup signal that retains the orientation awareness information of the sound source. The virtual microphone pickup signal and the retained side microphone signal are then processed by binaural signal fusion and sound wave emission according to the processing flow from steps S140 to S150.

[0087] Step S510: Obtain audio response segments from both ear canals recorded by the built-in microphones under normal wearing conditions of both earphones. Perform independent component analysis on the audio response segments from both ear canals, decomposing the left and right ear microphone signals into multiple statistically independent auditory source signal components. Extract spatial features from each auditory source signal component, determining its spatial orientation attribute based on the energy distribution ratio and phase difference between the left and right ear microphone signals. Cluster all auditory source signal components according to their spatial orientation attributes, forming sets of auditory source signal components corresponding to different spatial orientation clusters. Each spatial orientation cluster corresponds to a virtual sound source object.

[0088] Independent Component Analysis (ICA) employs a fast ICA algorithm, with the objective function being the maximization of negative entropy. The approximate expression for negative entropy is J(Y)∝(E[G(Y)]-E[G(V)])², where Y represents the separated independent components, V is a standard normally distributed variable, and G is a nonlinear function. Iterative solution uses a fixed-point iteration method, with the iterative formula W(k+1)=E[X×g(W(k)T×X)]-E[g'(W(k)T×X)]×W(k), where W is the column vector of the unmixing matrix, X is the observed signal matrix, T denotes transpose, g is the derivative of G, and g' is the second derivative of g. For each independent component, the energy distribution ratio r=EL / (EL+ER) is calculated, where EL is the energy of the independent component in the left ear signal, and ER is the energy of the independent component in the right ear signal. The phase difference Δφ is calculated as arg(XL) - arg(XR), where XL is the complex spectrum of the independent component in the left ear signal, and XR is the complex spectrum of the independent component in the right ear signal. Spatial orientation attributes are determined by both the energy distribution ratio and the phase difference. Orientation clustering uses the K-means clustering algorithm, with cluster feature vectors set to [r, cos(Δφ), sin(Δφ)], and the number of clusters K set to a preset value. Each cluster corresponds to a virtual sound source object.

[0089] Step S520: After the headphone loss event is triggered, the virtual microphone pickup signal from the lost headphone side and the microphone signal from the retained headphone side (real-time acquisition by the built-in microphone on the retained headphone side) are input into the pre-built binaural virtual sound field synthesizer. The pre-built binaural virtual sound field synthesizer performs independent spatial rendering processing on the virtual microphone pickup signal according to spatial orientation clusters to obtain the clustered spatial rendering signal corresponding to each spatial orientation cluster. Loudness hierarchy processing is performed on the clustered spatial rendering signals corresponding to different spatial orientation clusters. Differentiated loudness gains are applied to different spatial orientation clusters according to preset spatial orientation priority rules to obtain loudness-hierarchical clustered rendering signals.

[0090] The clustered independent spatial rendering process involves calling the sound field rendering layer of step S145 for each spatial orientation cluster. Using the spatial orientation of the cluster as input, a corresponding head-related transfer function is applied to the virtual sound source signal corresponding to that cluster. The preset spatial orientation priority rule assigns the highest priority to the directly forward orientation cluster, assigning a loudness gain Gmax; the middle priority to the side orientation cluster, assigning a loudness gain Gmid; and the lowest priority to the rear orientation cluster, assigning a loudness gain Gmin, where Gmax > Gmid > Gmin. The clustered spatial rendering signal for each spatial orientation cluster is multiplied by the corresponding gain value to obtain a hierarchical loudness clustered rendering signal.

[0091] Step S530: Perform inter-cluster time alignment processing on the loudness hierarchical cluster rendering signal to compensate for the arrival time offset caused by the difference in virtual sound source distance between clusters in different spatial orientations, and obtain a time-aligned cluster rendering signal. Perform binaural signal fusion processing on the time-aligned cluster rendering signal and the signal of the retaining side microphone collected in real time by the built-in microphone on the retaining side of the earphone to obtain a compensated audio signal of multi-source cluster rendering. The compensated audio signal of multi-source cluster rendering is emitted as a sound wave through the speaker unit on the retaining side of the earphone.

[0092] The sound wave propagation delay τd corresponding to the virtual sound source distance is d / c, where d is the virtual sound source distance parameter and c is the speed of sound. Delay compensation is applied to each cluster of rendered signals by shifting the cluster signal forward by τd in the time domain. After delay compensation, the signals of each cluster are time-aligned. The cluster signals are then added together and processed according to steps S140 to S150 for binaural signal fusion and sound wave emission.

[0093] Step S610: Obtain the environmental reference audio segment recorded by the external microphone within a preset time period before the earphone is lost, perform auditory scene analysis processing on the environmental reference audio segment, and separate the transient sound event sequence and the steady-state background sound sequence from the environmental reference audio segment. The transient sound event sequence contains sudden short-term sounds in the environment, and the steady-state background sound sequence contains background sounds that exist continuously in the environment.

[0094] Auditory scene analysis employs a sound source separation algorithm based on nonnegative matrix factorization (NMF). NMF decomposes the amplitude-time spectrum V of the environmental reference audio segment into the product of a basis matrix W and an activation matrix H, V ≈ W × H, where W is the matrix of frequency multiplied by the number of basis vectors, and H is the matrix of the number of basis vectors multiplied by the number of time frames. The temporal variance is calculated for each row of the decomposed activation matrix H; rows with large variance correspond to transient sound events, and rows with small variance correspond to steady-state background sound. Based on the row classification results, the amplitude-time spectrum Vtrans of transient sound events and the amplitude-time spectrum Vstat of steady-state background sound are reconstructed. Using the phase information of the original time spectra, Vtrans and Vstat are converted into time-domain signals through inverse short-time Fourier transform, yielding the transient sound event sequence and the steady-state background sound sequence.

[0095] Step S620: Perform event boundary detection processing on the transient sound event sequence to determine the start and end timestamps of each transient sound event, and extract the time-domain waveform segment and spectral envelope features of each transient sound event. Perform long-term spectral statistical processing on the steady-state background sound sequence to extract the long-term average spectral distribution and spectral fluctuation range parameters of the steady-state background sound sequence.

[0096] Event boundary detection is achieved by calculating the short-time energy envelope of a transient acoustic event sequence. The short-time energy envelope is E(t) = Σx(n)² × w(n), where w is a window function. An energy threshold Eth is set. The time when E(t) rises from below Eth to above Eth is marked as the start timestamp, and the time when E(t) falls from above Eth to below Eth is marked as the end timestamp. The signal segment between each pair of start and end timestamps constitutes a time-domain waveform segment of a transient acoustic event. The spectral envelope feature is taken as the smoothed amplitude spectrum curve of this segment. The long-term average spectral distribution is taken as the time average of the amplitude spectrum of the steady-state background sound sequence, Savg(f) = Σ|X(t,f)| / T, and the spectral fluctuation range parameter is taken as a preset multiple of the standard deviation of the amplitude time series at each frequency.

[0097] Step S630: After the headphone loss event is triggered, for each transient sound event in the transient sound event sequence, insert the corresponding transient sound event time-domain waveform segment into the virtual microphone pickup signal on the headphone loss side according to the start and end timestamps of the transient sound event, to obtain the virtual microphone pickup signal after the transient sound event is inserted. Perform transient sound event and steady-state background sound fusion transition processing on the virtual microphone pickup signal after the transient sound event is inserted, and apply fade-in and fade-out envelopes to the boundary regions of the transient sound events to obtain a virtual microphone pickup signal that fuses transient and steady-state sounds.

[0098] The insertion method involves superimposing a transient sound event time-domain waveform segment at the corresponding time position of the signal picked up by the virtual microphone using a cross-fading technique. The cross-fading interval is taken from each preset length interval before and after the starting timestamp. Within the cross-fading interval, the superimposed signal is the original signal multiplied by (1-α) plus the transient event signal multiplied by α, with α linearly transitioning from 0 to 1 and then linearly returning to 0. The fade-in and fade-out envelope uses a raised cosine function, expressed as w(n) = 0.5 × (1 - cos(2πn / N)), where n is the sample index and N is the length of the fade interval.

[0099] Step S640: Generate a steady-state background sound compensation signal based on the long-term average spectral distribution and the spectral fluctuation range parameters. Superimpose the steady-state background sound compensation signal onto the virtual microphone pickup signal obtained by fusing transient and steady-state signals to obtain a virtual microphone pickup signal jointly reconstructed from transient background. Input the virtual microphone pickup signal jointly reconstructed from transient background and the microphone signal from the earphone's built-in microphone (real-time acquisition) into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain a compensated audio signal jointly reconstructed from transient background. Emitter the compensated audio signal jointly reconstructed from transient background through the speaker unit on the earphone's ear canal to reconstruct a binaural auditory perception state in which both transient sound events and steady-state background sounds are completely preserved.

[0100] The steady-state background noise compensation signal is generated as follows: white noise is used as the excitation signal, and its amplitude spectrum is adjusted to a long-term average spectral distribution Savg(f), with the adjustment method being Xcomp(f) = Xnoise(f) × Savg(f). Within the amplitude range defined by the spectral fluctuation range parameter, the adjusted signal is subjected to random gain modulation, where the modulation gain g randomly takes values ​​from 1-δ to 1+δ, and δ is the fluctuation range parameter. The superposition method is direct time-domain addition. After superposition, binaural signal fusion and sound wave emission are performed according to the processing flow from steps S140 to S150.

[0101] For example, the method may further include: Step S710: During the training phase of the pre-constructed generative adversarial network, acquire ear canal audio response fragment samples and corresponding environmental reference audio fragment samples from multiple different listener individuals to construct a cross-listener generative adversarial network pre-training dataset. Perform listener individual identity encoding processing on the cross-listener generative adversarial network pre-training dataset, assign a listener individual identity embedding vector to each listener individual, and input the listener individual identity embedding vector as a conditional input along with the ear canal impulse response and environmental reference audio fragments into the generator sub-network. By introducing the listener individual identity embedding vector, perform cross-listener joint adversarial training on the generator sub-network and discriminator sub-network to obtain the generator sub-network base weights and discriminator sub-network base weights that share acoustic prior knowledge across listeners.

[0102] The listener's individual identity embedding vector is a trainable fixed-dimensional vector with a preset dimension. Each listener is assigned a unique embedding vector, which is randomly initialized before training and optimized along with the network weights during training. The embedding vector is mapped through a fully connected layer to a vector with the same dimension as the latent encoding vector of ear canal acoustic characteristics. The mapped embedding vector is added to the latent encoding vector of ear canal acoustic characteristics and used as the joint conditional input to the generator subnetwork. Cross-listener joint adversarial training is performed according to the adversarial training process in step S137, and the training data is randomly sampled from the pre-training dataset of the cross-listener generative adversarial network. After training, the weight parameters of the generator subnetwork and the discriminator subnetwork are saved as base weights.

[0103] Step S720: After the headphone loss event is triggered, obtain the audio response segment in the ear canal stored by the current listener before the headphone was lost. Perform ear canal impulse response deconvolution processing on the current listener's ear canal audio response segment to obtain the current listener's ear canal impulse response. Use the base weights of the generator subnetwork as initial weights, and perform a few iterations of fine-tuning training on the generator subnetwork using the current listener's ear canal impulse response. During the fine-tuning training process, freeze the shallow weights near the input layer in the generator subnetwork, and only update the deep weights near the output layer in the generator subnetwork.

[0104] The ear canal impulse response deconvolution processing is performed according to steps S121 to S124. Shallow weights include the ear canal feature embedding layer weights in step S131 and the pre-set layer weights of the context-aware branch in step S132. Deep weights include the feature fusion layer weights in step S133 and the progressive upsampling signal reconstruction module weights in step S134. Fine-tuning training uses a single-sample training method, and the number of iterations is dynamically adjusted based on the convergence of the loss value.

[0105] Step S730: After the iterative fine-tuning training converges, the current listener's ear canal impulse response and the environmental reference audio segment are input into the fine-tuned generator sub-network to generate virtual microphone signals on the lost side of the earphone, resulting in a personalized virtual microphone pickup signal. This signal is then input into a pre-constructed binaural virtual sound field synthesizer along with the microphone signal from the retained side of the earphone, which is collected in real time. The resulting personalized compensation audio signal is then emitted as a sound wave through the speaker unit on the retained side of the earphone.

[0106] The generator subnetwork in step S130 is replaced with the finely tuned generator subnetwork, and a personalized and adapted compensated audio signal is generated and transmitted according to the processing flow of steps S130 to S150.

[0107] Step S810: After the headphone loss event is triggered, perform fundamental frequency trajectory extraction processing on the virtual microphone pickup signal on the headphone loss side, estimate the voice fundamental frequency value or music fundamental frequency value frame by frame from the virtual microphone pickup signal to form a fundamental frequency trajectory curve that changes with time, perform fundamental frequency continuity analysis processing on the fundamental frequency trajectory curve, detect discontinuous intervals in the fundamental frequency trajectory curve that have fundamental frequency jumps or fundamental frequency interruptions, and mark the detected discontinuous intervals as fundamental frequency discontinuities.

[0108] Fundamental frequency estimation employs the autocorrelation function method, where the autocorrelation function R(τ) = Σx(n) × x(n+τ). The search term identifies the maximum peak value of R(τ) within the corresponding time delay range of the fundamental frequency. The peak position corresponds to the time delay τ0, and the fundamental frequency value f0 = fs / τ0, where fs is the sampling frequency. Fundamental frequency continuity analysis is achieved by calculating the absolute value of the fundamental frequency difference between adjacent frames, Δf0 = |f0(t+1) - f0(t)|. Positions where Δf0 exceeds a preset transition threshold are marked as discontinuous interval boundaries. If the estimated fundamental frequency value of a frame is zero while the fundamental frequency of adjacent frames is non-zero, it is marked as a fundamental frequency interruption. Continuous regions of fundamental frequency transitions and interruptions constitute fundamental frequency discontinuities.

[0109] Step S820: Extract the mean and slope of the fundamental frequency from the adjacent normal fundamental frequency intervals on both sides of the fundamental frequency discontinuity segment. Use the mean and slope of the fundamental frequency to perform fundamental frequency interpolation reconstruction processing on the fundamental frequency discontinuity segment to obtain the fundamental frequency trajectory curve after fundamental frequency continuity restoration. Extract the time-domain waveform segment corresponding to the fundamental frequency discontinuity segment from the signal picked up by the virtual microphone. Perform harmonic structure analysis processing on the time-domain waveform segment to extract the harmonic frequency sequence and harmonic amplitude sequence from the time-domain waveform segment.

[0110] The fundamental frequency interpolation uses linear interpolation. Let the mean fundamental frequency of the normal interval to the left of the fundamental frequency discontinuity be fL, and the mean fundamental frequency of the normal interval to the right be fR. The length of the discontinuity is N frames, and the interpolated fundamental frequency of the nth frame is f(n) = fL + (fR - fL) × n / (N-1), where n ranges from 0 to N-1. Harmonic structure analysis is achieved by performing a short-time Fourier transform on the time-domain waveform segment. The peak values ​​at integer multiples of the fundamental frequency are searched in the short-time Fourier transform amplitude spectrum. The peak frequencies constitute the harmonic frequency sequence fh(k) = k × f0, and the peak amplitudes constitute the harmonic amplitude sequence Ah(k).

[0111] Step S830: Based on the fundamental frequency reconstructed value in the fundamental frequency discontinuity segment interval of the fundamental frequency trajectory curve after fundamental frequency continuity restoration, regenerate the harmonic frequency sequence and harmonic amplitude sequence that match the harmonic structure of the fundamental frequency reconstructed value. Generate the completed waveform segment after fundamental frequency continuity restoration through sine wave superposition and synthesis processing. Replace the original fundamental frequency discontinuity segment time domain waveform segment in the virtual microphone pickup signal with the completed waveform segment to obtain the virtual microphone pickup signal after fundamental frequency continuity restoration.

[0112] The mathematical expression for the superposition and synthesis of sine waves is s(n) = ΣAh(k) × sin(2π × fh(k) × n / fs + φk), which sums over all harmonics. φk is the initial phase of the k-th harmonic, which is the extrapolated value of the harmonic phase of the adjacent normal interval of the original discontinuous segment. When completing the waveform segment to replace the original discontinuous segment, a cross-gradient processing is applied at the replacement boundary. The length of the cross-gradient interval is a preset value, and the gradation function is a raised cosine function.

[0113] Step S840: The virtual microphone pickup signal after fundamental frequency continuity restoration and the microphone signal of the earphone retaining side built-in microphone collected in real time are input into the pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain the compensated audio signal of fundamental frequency continuity restoration. The compensated audio signal of fundamental frequency continuity restoration is emitted as a sound wave through the speaker unit of the earphone retaining side.

[0114] Binaural signal fusion and sound wave emission are performed according to the processing flow of steps S140 to S150.

[0115] Step S910: Extract the interaural time difference parameter between the left and right ear microphone signals when both ears are normally worn from the ear canal audio response segments in the ear canal recording segment set of the earphone wearing state recording segment set. Use the fluctuation sequence of the interaural time difference parameter within a preset time period as the earphone wearing stability index. Input the fluctuation sequence of the interaural time difference parameter into a pre-constructed wearing state prediction network. The wearing state prediction network performs trend modeling processing on the fluctuation sequence of the interaural time difference parameter and outputs the estimated probability distribution of earphone loosening or displacement at future moments. When the estimated probability distribution exceeds a preset probability limit, trigger the earphone's built-in microphone to perform emergency caching processing on the current ear canal audio response segment. Store the urgently cached ear canal audio response segment in the earphone's non-volatile storage area as ear canal acoustic snapshot data at the moment the earphone is lost.

[0116] Interaural time difference (ITD) parameters are extracted using a cross-correlation method, and the time delay corresponding to the peak value of the cross-correlation function is the ITD. The ITD fluctuation sequence is obtained by taking ITD values ​​from multiple consecutive frames, and the fluctuation features are the standard deviation σITD and the linear regression slope sITD of the ITD sequence. The wearing status prediction network adopts a long short-term memory (LSTM) recurrent network structure. The input is the feature vector [σITD, sITD] of the ITD fluctuation sequence. After passing through two LSTM layers, it is followed by a fully connected layer and a Softmax output layer. The output is the probability distribution of three states: stable, loose, and dislodged. When the probability of the dislodged state exceeds a preset probability threshold, an emergency buffer is triggered.

[0117] Step S920: After the earphone loss event is actually triggered, the ear canal acoustic snapshot data is read from the earphone's non-volatile storage area. The ear canal acoustic snapshot data is then subjected to ear canal impulse response deconvolution processing to obtain the instantaneous ear canal impulse response at the moment of earphone loss. The instantaneous ear canal impulse response is used to perform instantaneous conditional injection processing on the generator subnetwork of the pre-constructed generative adversarial network. By encoding the set of ear canal formant frequencies in the instantaneous ear canal impulse response into a conditional vector and modulating it channel-by-channel with the intermediate layer features of the generator subnetwork, instantaneous fine-tuning weights of the generator subnetwork are obtained, conditioned on the ear canal state at the moment of loss.

[0118] Channel-by-channel modulation employs a characteristic linear modulation method. The set of ear canal formant frequencies is encoded into a conditional vector and then mapped to two vectors—a scaling factor γ and an offset factor β—through a fully connected layer. For each channel c of the feature map in the intermediate layer of the generator sub-network, the modulated feature is γc × Fc + βc, where Fc is the c-th channel of the original feature map. The modulated generator sub-network is fine-tuned with a few iterations based on the instantaneous ear canal impulse response to obtain the instantaneous fine-tuning weights.

[0119] Step S930: The generator sub-network, based on the ear canal state at the moment of loss, is invoked to process the environmental reference audio segment for generating a virtual microphone signal on the headphone loss side. This generates a virtual microphone pickup signal based on the ear canal state at the moment of loss. This virtual microphone pickup signal, along with the retention-side microphone signal acquired in real-time by the built-in microphone on the headphone retention side, is input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing, resulting in a compensated audio signal for ear canal state compensation at the moment of loss. The compensated audio signal for ear canal state compensation at the moment of loss is then emitted as sound waves through the speaker unit on the headphone retention side.

[0120] The compensation audio signal for compensating for the loss of ear canal state is generated and transmitted according to the processing flow of steps S130 to S150.

[0121] Step S1010: Acquire an environmental multi-source audio segment containing multiple simultaneously active sound sources, recorded by an external microphone under normal earphone wearing conditions. Perform source separation processing on the environmental multi-source audio segment to obtain multiple single-source audio tracks and spatial orientation labeling information corresponding to each single-source audio track. Input the multiple single-source audio tracks into the pre-constructed binaural virtual sound field synthesizer according to their respective spatial orientation labeling information for independent spatial rendering processing to obtain the single-source virtual microphone rendering signal of each single-source audio track on the earphone loss side. Perform linear superposition processing on all single-source virtual microphone rendering signals to obtain a multi-source synthesized virtual microphone reference signal. At the same time, use the real microphone pickup signal corresponding to the earphone loss side in the environmental multi-source audio segment as the multi-source real microphone reference signal.

[0122] The sound source separation process employs a speech separation algorithm based on deep clustering. The amplitude-time spectrum of the multi-source mixed signal is embedded into a high-dimensional space. Each time-frequency unit in the high-dimensional space is clustered, with each cluster corresponding to a sound source. A time-frequency mask for each sound source is constructed based on the clustering results. The time-frequency spectrum of the mixed signal is multiplied by the time-frequency mask of each sound source to obtain the time-frequency spectrum of each individual sound source. Then, an inverse short-time Fourier transform is used to obtain the audio track of each individual sound source. Spatial orientation information is extracted source-by-source using the sound source orientation estimation algorithm in step S410. Independent spatial rendering is performed according to step S145.

[0123] Step S1020: Input the multi-source synthesized virtual microphone reference signal and the multi-source real microphone reference signal into the discriminator subnetwork of the pre-constructed generative adversarial network for discriminative training in a multi-source scene to obtain the weight parameters of the multi-source scene discriminator subnetwork. During the training process of the generator subnetwork, a multi-source decoupling loss term is introduced. By performing source separation processing on the virtual microphone pickup signal generated by the generator subnetwork and comparing it with the source-by-source signal similarity of the multiple single-source audio tracks, the multi-source decoupling loss value is obtained.

[0124] The multi-source decoupling loss term is calculated as follows: The virtual microphone pickup signal output from the generator sub-network undergoes the same source separation processing as in step S1010 to obtain individual source components of the generated signal. For each separated individual source component, its scale-invariant signal-to-noise ratio (SNR) compared to the corresponding real single-source audio track is calculated. The scale-invariant SNR is SI-SNR = 10 × log10(||s_target||² / ||e_noise||²), where s_target is the orthogonal projection of the real single-source signal onto the generated single-source signal, and e_noise is the residual error signal. The negative value of the average scale-invariant SNR of all sources is taken as the multi-source decoupling loss value.

[0125] Step S1030: The multi-source decoupling loss value is superimposed on the total loss of the generator sub-network for joint backpropagation training to obtain the weight parameters of the multi-source decoupling generator sub-network. After the headphone loss event is triggered, the multi-source decoupling generator sub-network is invoked to generate virtual microphone signals on the headphone loss side of the environmental reference audio segment, resulting in multi-source decoupling virtual microphone pickup signals. The multi-source decoupling virtual microphone pickup signals and the retention side microphone signals collected in real time by the built-in microphone on the headphone retention side are input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain multi-source decoupling compensated audio signals. The multi-source decoupling compensated audio signals are then emitted as sound waves through the speaker unit on the headphone retention side.

[0126] The total loss function is Ltotal = Ladv + λ1 × Lfreq + λ2 × Ldecoup, where Ladv is the adversarial loss, Lfreq is the frequency domain consistency constraint term, Ldecoup is the multi-source decoupling loss value, and λ1 and λ2 are weighting coefficients. After training, the multi-source decoupling compensated audio signal is generated and transmitted according to the processing flow from steps S130 to S150.

[0127] Step S1110: Obtain the bone conduction signal component caused by the listener's own voice and the air conduction signal component caused by the headphone speaker in the audio response segment in the ear canal under normal wearing conditions on both sides of the headphone. Perform signal separation processing on the bone conduction signal component and the air conduction signal component to obtain the pure bone conduction ear canal response and the pure air conduction ear canal response.

[0128] Signal separation processing utilizes the time delay and spectral characteristics differences between bone conduction and air conduction signals. The bone conduction signal is synchronized with the listener's vocalization, while the air conduction signal experiences a system delay from playback through the headphone speaker to microphone pickup. An adaptive filtering algorithm is used, with the headphone speaker signal as the reference input and the microphone pickup signal as the desired output. The adaptive filter output is an estimate of the air conduction signal, which is subtracted from the microphone pickup signal to obtain the bone conduction signal estimate. The adaptive filter employs a normalized least mean square algorithm, with the weight coefficient update formula being W(n+1)=W(n)+μ×e(n)×X(n) / (X(n)T×X(n)+ε), where μ is the step size parameter, e(n) is the error signal, X(n) is the reference input vector, and ε is a small regularization constant.

[0129] Step S1120: Perform spectral feature extraction processing on the pure bone conduction ear canal response to obtain the set of bone conduction resonant frequencies and the bone conduction spectral envelope of the bone conduction sound in the ear canal when the listener speaks. Perform spectral feature extraction processing on the pure air conduction ear canal response to obtain the set of air conduction resonant frequencies and the air conduction spectral envelope of the sound played by the headphone speaker through the air in the ear canal.

[0130] Spectral feature extraction employs linear predictive coding analysis to extract the set of formant frequencies. The spectral envelope is taken from the frequency response amplitude value of the linear predictive coding filter.

[0131] Step S1130: Construct the bone conduction and air conduction ear canal coupling transmission characteristics of an individual listener based on the bone conduction resonant frequency set and the air conduction resonant frequency set. The bone conduction and air conduction ear canal coupling transmission characteristics describe the acoustic superposition mode of bone conduction sound and air conduction sound in the ear canal when the listener produces sound.

[0132] The coupling transfer characteristics are taken from the amplitude ratio and phase difference of the bone conduction spectrum envelope and the air conduction spectrum envelope at each frequency point. The amplitude ratio Hcouple(f) = |Sbone(f)| / |Sair(f)|, and the phase difference Δφcouple(f) = arg(Sbone(f)) - arg(Sair(f)), where Sbone is the bone conduction spectrum and Sair is the air conduction spectrum.

[0133] Step S1140: After the headphone loss event is triggered, determine whether the currently played audio source signal contains the sidetone signal component of the listener's own voice communication. If it contains the sidetone signal component, extract the sidetone signal component and perform bone conduction simulation processing to obtain a simulated bone conduction ear canal response. Input the simulated bone conduction ear canal response into the generator subnetwork of the pre-constructed generative adversarial network. Perform joint conditional generation processing on the simulated bone conduction ear canal response and the environmental reference audio segment through the generator subnetwork to obtain a virtual microphone pickup signal containing bone conduction auditory cues.

[0134] The bone conduction simulation process involves passing the lateral tone signal component through a bone conduction coupling filter, with the filter's frequency response taking the values ​​of Hcouple(f) and Δφcouple(f) from step S1130. The simulated bone conduction canal response is then transformed using an inverse Fourier transform to obtain the time-domain signal. The joint conditional generation process inputs the simulated bone conduction canal response as an additional condition into the generator subnetwork, concatenating it with the canal impulse response condition before inputting it together.

[0135] Step S1150: The virtual microphone pickup signal containing bone conduction auditory cues and the microphone signal collected in real time by the microphone on the earphone's reserved side are input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain a compensated audio signal containing bone conduction auditory cues. The compensated audio signal containing bone conduction auditory cues is then emitted as a sound wave through the speaker unit on the earphone's reserved side.

[0136] Binaural signal fusion and sound wave emission are performed according to the processing flow of steps S140 to S150.

[0137] Step S1210: Obtain the audio response segment inside the ear canal recorded by the built-in microphone under normal wearing conditions on both sides of the earphone. Perform ear canal acoustic impedance estimation processing on the audio response segment inside the ear canal. By analyzing the energy ratio of the reflected wave to the incident wave of different frequency components in the audio response segment inside the ear canal, obtain the ear canal acoustic impedance frequency response curve. Extract the frequency positions of the impedance maxima and impedance minima at the resonant peak frequency from the ear canal acoustic impedance frequency response curve. Use the frequency positions of the impedance maxima and impedance minima as the set of ear canal acoustic impedance feature points.

[0138] The acoustic impedance of the ear canal is estimated using the time-domain reflectometry method. A short pulse is emitted into the ear canal, and the reflected signal is collected. The incident and reflected waves are separated through a time window. The reflection coefficient R(f) = Pr(f) / Pi(f), where Pr is the spectrum of the reflected wave and Pi is the spectrum of the incident wave. The acoustic impedance Z(f) = Z0 × (1 + R(f)) / (1 - R(f)), where Z0 is the characteristic impedance of air. The frequency of impedance maxima is taken as the frequency corresponding to the local maxima of Z(f), and the frequency of impedance minima is taken as the frequency corresponding to the local minima of Z(f).

[0139] Step S1220: After the headphone loss event is triggered, a test sweep signal of preset intensity is emitted into the listener's ear canal through the built-in microphone on the headphone retaining side. At the same time, the built-in microphone on the headphone retaining side synchronously picks up the reflected sound signal in the ear canal. The picked-up reflected sound signal in the ear canal is processed for real-time ear canal acoustic impedance estimation to obtain a real-time ear canal acoustic impedance frequency response curve. The real-time impedance maximum frequency position and the real-time impedance minimum frequency position are extracted from the real-time ear canal acoustic impedance frequency response curve. The real-time impedance maximum frequency position and the real-time impedance minimum frequency position are compared with the ear canal acoustic impedance feature point set for frequency offset processing to obtain the ear canal acoustic impedance feature point frequency offset.

[0140] The frequency offset comparison process calculates the frequency difference ΔfZ = fZ_realtime - fZ_reference for each corresponding feature point, where fZ_realtime is the real-time impedance feature point frequency and fZ_reference is the corresponding feature point frequency in the set of acoustic impedance feature points in the ear canal.

[0141] Step S1230: Based on the frequency offset of the characteristic point of the ear canal acoustic impedance, perform frequency correction processing on the set of ear canal resonant frequencies in the ear canal impulse response to obtain the ear canal resonant frequency correction amount. Use the ear canal resonant frequency correction amount to perform resonant frequency shift processing on the ear canal impulse response to obtain the corrected ear canal impulse response. Input the corrected ear canal impulse response and the environmental reference audio segment into a pre-constructed generative adversarial network to perform virtual microphone signal generation processing on the headphone loss side to obtain a virtual microphone pickup signal based on real-time ear canal acoustic impedance correction.

[0142] The resonant frequency shifting process adds a corresponding frequency offset ΔfZ to each resonant frequency in the ear canal impulse response. After correction, a compensated audio signal is generated and transmitted according to the processing flow from steps S130 to S150.

[0143] Step S1240: The virtual microphone pickup signal based on real-time ear canal acoustic impedance correction and the microphone signal of the earphone retaining side built-in microphone collected in real time are input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain a compensated audio signal based on real-time ear canal acoustic impedance correction. The compensated audio signal based on real-time ear canal acoustic impedance correction is emitted as a sound wave through the speaker unit of the earphone retaining side.

[0144] Binaural signal fusion and sound wave emission are performed according to the processing flow of steps S140 to S150.

[0145] Combination Figure 3 Content, Figure 3 This is a simulation diagram verifying the virtual audio experience maintenance method based on generative adversarial networks after headphone loss provided in an embodiment of the present invention. Specifically, Figure 3Figure A shows a comparison of the time-domain waveforms of the audio response segment inside the ear canal before the earphone is lost and the environmental reference audio segment. The upper waveform is the audio response segment inside the ear canal recorded by the built-in microphone when both earphones are worn normally, and the lower waveform is the environmental reference audio segment recorded by the external microphone just before the earphone is lost. The alignment error between the two is less than 0.1 milliseconds, and the sampling rate of both is 48 kHz, indicating that the two signals are highly synchronized in time. Figure B shows the ear canal impulse response waveform and its parameterized characteristics obtained by deconvolution processing. The figure marks the set of ear canal resonant frequencies (such as the main resonant at 3.2 kHz) and ear canal decay time parameters (such as 8 milliseconds for RT60). These parameters constitute the personalized ear canal acoustic conditions required by the generator sub-network. Figure C shows the frequency domain verification results of the generative adversarial network output signal. The blue solid line represents the power spectral density of the environmental reference audio segment, and the red dashed line represents the power spectral density of the signal picked up by the virtual microphone generated by the generator sub-network. The two highly overlap in spectral shape, with a generator loss of only 0.15, a discriminator accuracy of 51%, and a signal-to-noise ratio of 28 dB. This demonstrates that the generated signal successfully reproduces the true spectral characteristics of the environmental audio in the frequency domain energy distribution. Figure D uses a polar coordinate plot to show the spatial orientation resolution effect of the binaural virtual sound field synthesizer. The solid line represents the true binaural auditory perception state before the headphones are lost (sound source azimuth angle 120°), and the dashed line represents the spatial orientation of the lost-side speaker compensation audio signal reconstructed after fusing the virtual microphone-picked signal with the retained-side signal. The maximum level difference between the two is less than 1.5 dB, the phase difference is less than 5°, and the effective frequency range covers 200 Hz to 8 kHz, indicating that the compensation audio signal successfully preserves the spatial orientation clues of the lost-side sound source. Figure E presents the subjective rating results through a bar chart. The subjective ratings (spatial awareness, speech intelligibility, and comfort) for the uncompensated (mono-ear mode) are all around 2 points, while the ratings for the proposed method (virtual audio maintenance) are all above 4.7 points. The overall MOS score has increased from 1.8 points to 4.7 points, which significantly verifies the effectiveness of the proposed method in maintaining binaural auditory perception.

[0146] Combination Figure 4 Content, Figure 4This is a schematic diagram of the front-end interface of the pairing terminal during the virtual audio maintenance process provided in this embodiment of the invention. It shows the interactive interface when a user performs virtual audio experience maintenance operations through the terminal application interface after the headphones are lost. The top of the interface displays the current status as "Virtual Audio Maintenance", indicating that the system has been successfully activated and entered the lost-side audio compensation mode. A 3D head model is displayed in the center of the interface. The left headphone is marked "In Use" (green lit), indicating that the retained headphone is working normally; the right headphone is marked "Lost" (red mark), and the lost status is indicated by a dotted outline, while also noting "Virtual sound field compensation activated, restoring your binaural hearing". The interface features a status card at the bottom. The left side displays the status information for "Virtual Right Ear (AI Generated)," including details such as the sound source azimuth (120°), formant lock status (normal), and transient events (door closed), demonstrating that the generator sub-network not only reproduced the spectral characteristics but also successfully captured acoustic events in the environment. The right side displays the status information for "Real Left Ear (Real-time)," including wearing tightness (medium) and ear canal impedance (calibrated), indicating that the system is monitoring the acoustic state of the preserved ear canal in real time for dynamic compensation. The bottom function bar provides interactive controls such as "Recalibrate," "Spatial Focus," "Sidetone Mode," and "Audio Diagnostics," allowing users to fine-tune the virtual sound field compensation effect. At the very bottom is a "Listen" button, which allows users to listen to the compensated binaural audio and provide subjective feedback through the prompt "Is the spatial sense accurate?" The overall interface, through a visualized head model and real-time parameter feedback, intuitively demonstrates the generation status and compensation effect of the virtual audio signal on the lost side, achieving effective interaction between the technical solution and user operation.

[0147] In an exemplary embodiment, a virtual audio experience maintenance system based on generative adversarial networks (GANs) after headphone loss is provided. This system can be a terminal, server, etc., and its internal structure includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, near-field communication, or other technologies. When the computer program is executed by the processor, it implements a method for maintaining a virtual audio experience after headphone loss based on generative adversarial networks. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the shell of a virtual audio experience maintenance system after headphone loss based on generative adversarial networks, or external keyboards, touchpads, or mice, etc.

[0148] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A method for maintaining a virtual audio experience after headphone loss based on generative adversarial networks, characterized in that, The method includes: After the headphone loss event is triggered, a set of audio recording segments of the headphone wearing status stored before the headphone is lost is obtained. The set of audio recording segments of the headphone wearing status includes the audio response segment in the ear canal recorded by the built-in microphone when both sides of the headphone are worn normally and the environmental reference audio segment recorded by the external microphone just before the headphone is lost. The ear canal impulse response deconvolution processing is performed on the ear canal audio response segment in the ear canal recording segment set of the earphone wearing state to obtain the ear canal impulse response under the normal wearing state of both sides of the earphone. The ear canal impulse response includes the ear canal formant frequency set and the ear canal decay time parameter. The ear canal impulse response corresponding to the audio response segment in the ear canal and the environmental reference audio segment are input into a pre-constructed generative adversarial network to generate virtual microphone signals on the lost earphone side, thereby obtaining the virtual microphone pickup signal on the lost earphone side. The pre-constructed generative adversarial network includes a generator sub-network and a discriminator sub-network. The virtual microphone pickup signal on the lost earphone side and the microphone signal on the retained earphone side, which are collected in real time by the built-in microphone on the retained earphone side, are input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain the compensated audio signal to be played by the speaker on the lost earphone side. The compensated audio signal is emitted as a sound wave through the speaker unit on the earphone retaining side, reconstructing the binaural auditory perception state before the earphone was lost in the listener's ear canal.

2. The method for maintaining the virtual audio experience after headphone loss based on generative adversarial networks according to claim 1, characterized in that, The step of inputting the ear canal impulse response corresponding to the audio response segment in the ear canal and the environmental reference audio segment into a pre-constructed generative adversarial network for virtual microphone signal generation processing on the lost earphone side to obtain the virtual microphone pickup signal on the lost earphone side includes: The ear canal formant frequency set is extracted from the ear canal impulse response. The frequency position and amplitude value of each formant in the ear canal formant frequency set are encoded as a formant structure feature sequence. The formant structure feature sequence is input into the generator subnetwork of the pre-constructed generative adversarial network. The formant structure feature sequence is nonlinearly mapped through the ear canal feature embedding layer to obtain the latent encoding vector of ear canal acoustic characteristics. The last complete audio frame unit before the moment of headphone loss is extracted from the environmental reference audio segment. The last complete audio frame unit is subjected to multi-resolution time-frequency decomposition processing to obtain a set of audio frame multi-scale spectrum layers containing different time granularities and frequency resolutions. The set of audio frame multi-scale spectrum layers is input into the environmental context awareness branch of the generator sub-network. Causal temporal feature aggregation processing is performed on the set of audio frame multi-scale spectrum layers through a sequence of ordered dilated convolutional layers to obtain an environmental audio context feature map. The latent encoding vector of the acoustic characteristics of the ear canal and the feature map of the environmental audio context are concatenated along the channel dimension in the feature fusion layer of the generator sub-network. The concatenated features are then subjected to gated fusion processing to obtain the conditionally generated feature representation of the acoustic scene on the side where the earphone is lost. The conditional generation feature representation of the acoustic scene on the side where the earphone is lost is input into the stepwise upsampling signal reconstruction module of the generator subnetwork. The temporal resolution of the signal is gradually restored through a series of deconvolutional layers to generate candidate virtual microphone pickup signals on the side where the earphone is lost. The candidate virtual microphone pickup signal on the lost earphone side is processed by frequency band energy distribution analysis to extract the candidate signal frequency band energy distribution vector. The reserved earphone microphone signal collected in real time by the built-in microphone on the reserved earphone side is processed by frequency band energy distribution analysis to extract the reference signal frequency band energy distribution vector. Calculate the frequency band energy distribution deviation between the candidate signal frequency band energy distribution vector and the reference signal frequency band energy distribution vector, and add the frequency band energy distribution deviation as a frequency domain consistency constraint term to the loss calculation process of the generator sub-network; During the adversarial alternation training of the generator subnetwork and the discriminator subnetwork, the discriminator subnetwork receives the candidate virtual microphone pickup signal and the real lost microphone acquisition signal and outputs the discrimination confidence score. The generator subnetwork jointly updates the network weights based on the discrimination confidence score and the frequency domain consistency constraint term. After the adversarial training converges, the generator subnetwork outputs the virtual microphone pickup signal of the lost earphone side.

3. The method for maintaining the virtual audio experience after headphone loss based on generative adversarial networks according to claim 1, characterized in that, The step of inputting the virtual microphone pickup signal from the lost earphone side and the microphone signal from the retained earphone side built-in microphone in real time into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain the compensated audio signal to be played by the speaker on the lost earphone side includes: The virtual microphone signal picked up on the lost side of the earphone is processed by a short-time Fourier transform to obtain the virtual side time spectrum. The microphone signal of the retained side, which is collected in real time by the built-in microphone on the retained side of the earphone, is processed by a short-time Fourier transform to obtain the retained side time spectrum. The cross-power spectrum phase analysis is performed on the virtual side time spectrum and the reserved side time spectrum in each time frequency unit to extract the binaural phase difference corresponding to each time frequency unit, and the binaural phase differences of all time frequency units are aggregated into a binaural phase difference time frequency map; Energy ratio analysis is performed on the virtual side time spectrum and the reserved side time spectrum in each time-frequency unit. The binaural energy ratio corresponding to each time-frequency unit is extracted, and the binaural energy ratios of all time-frequency units are aggregated into a binaural energy ratio time-frequency map. The binaural phase difference time-frequency diagram and the binaural energy ratio time-frequency diagram are input into the spatial orientation resolution layer of the pre-constructed binaural virtual sound field synthesizer. The binaural cues are mapped into azimuth and elevation parameters of the virtual sound source relative to the center of the listener's head through the spatial orientation resolution layer, thereby obtaining the spatial orientation parameter sequence of the virtual sound source on the side where the earphone is lost. The virtual side-time spectrum and the spatial orientation parameter sequence are input into the sound field rendering layer of the pre-constructed binaural virtual sound field synthesizer. The sound field rendering layer calls the head-related transmission characteristic dataset that matches the spatial orientation parameter sequence, and applies the corresponding head-related transmission gain and head-related transmission phase shift to each time-frequency unit in the virtual side-time spectrum to obtain the virtual side-rendered time spectrum carrying spatial orientation clues. The virtual side rendering spectrum is converted into a time-domain signal through inverse short-time Fourier transform to obtain the spatial audio signal of the lost earphone side carrying spatial orientation clues. The spatial audio signal of the lost earphone side is subjected to perceptual loudness measurement processing to obtain the virtual side perceptual loudness value. The microphone signal of the retained side is subjected to perceptual loudness measurement processing to obtain the retained side perceptual loudness value. The loudness difference between the virtual side perceptual loudness value and the retained side perceptual loudness value is calculated. The gain adjustment process is performed on the spatial audio signal of the lost earphone carrying spatial orientation clues according to the loudness difference, to obtain the loudness-adjusted spatial audio signal of the lost earphone, and the loudness-adjusted spatial audio signal of the lost earphone is used as the compensation audio signal to be played by the speaker of the lost earphone.

4. The method for maintaining the virtual audio experience after headphone loss based on generative adversarial networks according to claim 1, characterized in that, The method further includes: Multiple sets of audio response segments in the ear canal with different wearing tightness and corresponding wearing tightness labels are obtained by the built-in microphones under normal wearing conditions on both sides of the earphone. A training dataset for wearing tightness conditions is constructed. Each set of audio response segments in the ear canal in the training dataset for wearing tightness conditions is subjected to deconvolution processing of the ear canal impulse response to obtain the ear canal impulse response sample corresponding to the wearing tightness label. The ear canal formant frequency set sample and ear canal decay time parameter sample are extracted from the ear canal impulse response sample. We performed a statistical analysis of the resonance peak frequency set samples corresponding to different tightness labels in the ear canal to extract the direction and magnitude of the resonance peak frequency shift as the tightness of the ear canal changes; and we performed a statistical analysis of the decay time parameter samples corresponding to different tightness labels in the ear canal to extract the extension and compression trends of the decay time parameter as the tightness of the ear canal changes. After the earphone loss event is triggered, the current ear canal response signal in the ear canal on the earphone retention side is collected in real time by the built-in microphone on the earphone retention side. The current ear canal response signal is processed by formant detection and decay time estimation to obtain the current ear canal formant frequency and the current ear canal decay time. Based on the shift direction and magnitude of the resonant frequency as a function of wearing tightness, the estimated current wearing tightness of the earphone retaining side is calculated. Based on the estimated current wearing tightness of the earphone retaining side, a subset of ear canal impulse response samples matching the current wearing tightness is selected from the wearing tightness condition training dataset. The ear canal impulse response sample subset is then used to perform online fine-tuning training of the generator subnetwork of the pre-constructed generative adversarial network. After online fine-tuning training is completed, the step of inputting the ear canal impulse response corresponding to the audio response segment in the ear canal and the environmental reference audio segment into the pre-constructed generative adversarial network to generate virtual microphone signal on the lost side of the earphone is re-executed to obtain a virtual microphone pickup signal adapted to the current tightness of the listener's current wearing. The virtual microphone signal, which adapts to the listener's current wearing tightness, and the microphone signal from the earphone's retaining side, which is collected in real time by the earphone's retaining side microphone, are input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing. This results in a compensated audio signal that adapts to changes in wearing tightness. The compensated audio signal is then emitted as a sound wave through the earphone's retaining side speaker unit, reconstructing a binaural auditory perception state that matches the listener's current wearing state within the ear canal.

5. The method for maintaining the virtual audio experience after headphone loss based on generative adversarial networks according to claim 1, characterized in that, The method further includes: The virtual microphone pickup signal on the lost earphone side is processed by auditory filter bank analysis. The virtual microphone pickup signal on the lost earphone side is decomposed into multiple auditory sub-band signals by a gamma-tonal filter bank that simulates the frequency selectivity characteristics of the basilar membrane of the human ear. The microphone signal on the retained earphone side, which is collected in real time by the built-in microphone on the retained earphone side, is also processed by auditory filter bank analysis. The retained earphone signal is decomposed into multiple auditory sub-band signals by the gamma-tonal filter bank. Sub-band cross-correlation analysis is performed on each auditory sub-band signal on the lost earphone side and the corresponding auditory sub-band signal on the retained earphone side to extract the sub-band binaural time difference parameter and the sub-band binaural cross-correlation coefficient for each auditory sub-band. The credibility of the binaural cue for each auditory subband is evaluated based on the binaural cross-correlation number of each subband. Auditory subbands with binaural cross-correlation numbers lower than a preset correlation threshold are marked as first credible subbands, and auditory subbands with binaural cross-correlation numbers not lower than a preset correlation threshold are marked as second credible subbands. For the first reliable sub-band, inter-sub-band binaural cue interpolation is performed using the binaural time difference parameters of the adjacent second reliable sub-band to obtain the inferred binaural time difference parameters of the first reliable sub-band; and for the earphone loss side auditory sub-band signal corresponding to the first reliable sub-band, time delay compensation and phase alignment processing are performed using the inferred binaural time difference parameters to obtain the binaural cue corrected first reliable sub-band signal. The first reliable sub-band signal after binaural cue correction and the earphone loss side auditory sub-band signal corresponding to the second reliable sub-band are subjected to sub-band synthesis and reconstruction processing to obtain a virtual microphone pickup signal with enhanced binaural cue consistency. The virtual microphone pickup signal and the retention side microphone signal collected in real time by the built-in microphone on the earphone retention side are input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain a compensated audio signal with enhanced binaural cue consistency. The compensated audio signal is emitted as a sound wave through the speaker unit on the earphone retention side to reconstruct the binaural auditory perception state with consistent binaural cue in the listener's ear canal.

6. The method for maintaining the virtual audio experience after headphone loss based on generative adversarial networks according to claim 1, characterized in that, The method further includes: The system acquires dynamic environmental audio segments recorded by the external microphone when the listener slowly rotates their head while wearing both headphones normally. The dynamic environmental audio segments include binaural cues of the continuous change in the relative position of the environmental sound sources during the listener's head rotation. The dynamic environment audio segment is subjected to frame-by-frame sound source azimuth estimation processing. The instantaneous sound source azimuth angle sequence and instantaneous sound source elevation angle sequence corresponding to each frame are extracted, and the instantaneous sound source azimuth angle sequence and instantaneous sound source elevation angle sequence are used as dynamic azimuth ground truth labels. The microphone signal channel corresponding to the lost earphone side in the dynamic environment audio segment is used as the training target signal, and the microphone signal channel corresponding to the retained earphone side is used as the training input signal to construct a training data pair for generating the lost earphone side signal. Based on the lost earphone signal, a training data pair is generated and the dynamic azimuth ground truth label is generated. An azimuth awareness auxiliary training branch is added to the generator subnetwork of the pre-constructed generative adversarial network. An azimuth prediction subnetwork is derived from the intermediate feature layer of the generator subnetwork through the azimuth awareness auxiliary training branch. The azimuth prediction subnetwork outputs a predicted azimuth angle sequence and a predicted elevation angle sequence. Calculate the azimuth prediction deviation between the predicted azimuth sequence and the instantaneous sound source azimuth sequence in the dynamic azimuth ground truth label, calculate the elevation prediction deviation between the predicted elevation sequence and the instantaneous sound source elevation sequence in the dynamic azimuth ground truth label, use the azimuth prediction deviation and the elevation prediction deviation as azimuth perception loss terms, and perform weighted joint training processing on the azimuth perception loss terms and the adversarial loss terms of the generator subnetwork to obtain the weight parameters of the generator subnetwork with enhanced azimuth perception. After the headphone loss event is triggered, the generator sub-network for enhanced orientation is invoked to process the virtual microphone signal of the lost headphone side on the environmental reference audio segment, resulting in a virtual microphone pickup signal that retains the orientation information of the sound source. The virtual microphone pickup signal that retains the orientation information of the sound source is then input into a pre-constructed binaural virtual sound field synthesizer along with the microphone signal of the retained side that is collected in real time by the built-in microphone on the retained side of the headphone for binaural signal fusion processing, resulting in a compensated audio signal that retains the orientation information of the sound source. The compensated audio signal that retains the orientation information of the sound source is then emitted as a sound wave through the speaker unit on the retained side of the headphone, reconstructing the binaural auditory perception state with correct orientation of the sound source in the listener's ear canal.

7. The method for maintaining the virtual audio experience after headphone loss based on generative adversarial networks according to claim 1, characterized in that, The method further includes: The audio response segments inside the ear canals of both ears, recorded by the built-in microphones under normal wearing conditions of both ears, are obtained. Independent component analysis is performed on the audio response segments inside the ear canals to decompose the left ear microphone signal and the right ear microphone signal into multiple statistically independent auditory source signal components. Spatial feature extraction is performed on each auditory source signal component. The spatial orientation attribute of each auditory source signal component is determined based on the energy distribution ratio and phase difference of each auditory source signal component in the left ear microphone signal and the right ear microphone signal. All auditory source signal components are clustered according to their spatial orientation attributes to form sets of auditory source signal components corresponding to different spatial orientation clusters. Each spatial orientation cluster corresponds to a virtual sound source object. After the headphone loss event is triggered, the virtual microphone pickup signal on the lost headphone side and the microphone signal on the retained headphone side built-in microphone collected in real time are input into the pre-built binaural virtual sound field synthesizer. The virtual microphone pickup signal is processed by the pre-built binaural virtual sound field synthesizer according to the spatial orientation cluster to obtain the clustered spatial rendering signal corresponding to each spatial orientation cluster. Loudness hierarchy processing is performed on the clustered spatial rendering signals corresponding to different spatial orientation clusters. Different loudness gain is applied to different spatial orientation clusters according to the preset spatial orientation priority rules to obtain loudness hierarchy clustered rendering signals. The loudness-hierarchical cluster rendering signal is subjected to inter-cluster time alignment processing to compensate for the arrival time offset caused by the difference in virtual sound source distance between clusters in different spatial orientations, resulting in a time-aligned cluster rendering signal. The time-aligned cluster rendering signal is then fused with the signal from the microphone on the earphone's retaining side, which is collected in real time by the microphone on the retaining side, to obtain a compensated audio signal for multi-source cluster rendering. The compensated audio signal for multi-source cluster rendering is then emitted as a sound wave through the speaker unit on the earphone's retaining side, reconstructing the binaural auditory perception state containing multiple independent and distinguishable virtual sound sources in the listener's ear canal.

8. The method for maintaining the virtual audio experience after headphone loss based on generative adversarial networks according to claim 1, characterized in that, The method further includes: The system acquires an environmental reference audio segment recorded by the external microphone within a preset time period before the earphone is lost. It performs auditory scene analysis processing on the environmental reference audio segment and separates a transient sound event sequence and a steady-state background sound sequence from the environmental reference audio segment. The transient sound event sequence contains sudden short-term sounds in the environment, and the steady-state background sound sequence contains continuous background sounds in the environment. The transient acoustic event sequence is subjected to event boundary detection processing to determine the start and end timestamps of each transient acoustic event, and the time-domain waveform segment and spectral envelope features of each transient acoustic event are extracted. Long-term spectral statistical processing is performed on the steady-state background sound sequence to extract the long-term average spectral distribution and spectral fluctuation range parameters of the steady-state background sound sequence; After the headphone loss event is triggered, for each transient sound event in the transient sound event sequence, the corresponding transient sound event time-domain waveform segment is inserted into the virtual microphone pickup signal on the headphone loss side according to the start timestamp and end timestamp of the transient sound event, so as to obtain the virtual microphone pickup signal after the transient sound event is inserted. The virtual microphone pickup signal after the insertion of transient sound events is subjected to a fusion transition process of transient sound events and steady-state background sound. A fade-in and fade-out envelope is applied to the boundary region of the transient sound events to obtain a virtual microphone pickup signal that fuses transient and steady-state sounds. A steady-state background sound compensation signal is generated based on the long-term average spectral distribution and the spectral fluctuation range parameter. The steady-state background sound compensation signal is then superimposed on the virtual microphone pickup signal obtained by fusing transient and steady-state signals to obtain a virtual microphone pickup signal that is jointly reconstructed from transient background. The virtual microphone pickup signal reconstructed from the transient background and the microphone signal collected in real time from the microphone on the earphone's reserved side are input into a pre-constructed binaural virtual sound field synthesizer for binaural signal fusion processing to obtain the compensated audio signal reconstructed from the transient background. The compensated audio signal, which is jointly reconstructed from the transient background, is emitted as a sound wave through the speaker unit on the earphone retaining side, thereby reconstructing the binaural auditory perception state in the listener's ear canal where both transient sound events and steady-state background sounds are completely preserved.

9. A machine-readable storage medium, characterized in that, Used to store the processor's machine-executable instructions; The processor is configured to execute the virtual audio experience maintenance method based on a generative adversarial network according to any one of claims 1 to 8 by executing the machine-executable instructions.

10. A virtual audio experience maintenance system based on generative adversarial networks after headphone loss, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the virtual audio experience maintenance method based on a generative adversarial network according to any one of claims 1 to 8 by executing the machine-executable instructions.