Speaker sound quality optimization methods, devices, and audio equipment based on reinforcement learning
By employing a reinforcement learning-based speaker sound quality optimization method, which utilizes audio bitstreams and sound field perception records for state representation, separates the acoustic environment and speaker response factors, and adaptively adjusts the parameters of the audio equipment, the problem of unstable sound quality optimization effect of the audio equipment in non-ideal acoustic environments is solved, and the efficiency and robustness of autonomous sound quality optimization are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing audio equipment is susceptible to changes in sound quality optimization in non-ideal acoustic environments due to environmental changes and changes in playback content, making it difficult to adaptively adjust and resulting in unstable sound quality performance.
A speaker sound quality optimization method based on reinforcement learning is adopted. By acquiring the audio bitstream and sound field perception record of the audio device, nonlinear state representation and sound field topology analysis are performed. The implicit causal graph structure of the agent with causal decoupling strategy is used to separate the acoustic environment and the speaker nonlinear response factor, generate a sound quality transfer action sequence, and adaptively adjust the audio signal processing link parameters.
It improves the efficiency and generalization robustness of audio equipment in autonomous sound quality optimization in complex acoustic environments, and can maintain the optimization effect when the environment changes and the content played changes.
Smart Images

Figure CN122496749A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio signal processing technology, and in particular to a speaker sound quality optimization method, apparatus and audio equipment based on reinforcement learning. Background Technology
[0002] With the widespread adoption of smart speakers, home theaters, and in-vehicle audio systems, consumers are demanding higher sound quality from audio equipment in diverse listening environments. The actual listening experience of audio equipment depends not only on the physical characteristics and algorithm configuration of its own electroacoustic components, such as speaker units, amplifier circuits, and digital signal processors, but also significantly on the physical acoustic environment in which it is situated. The characteristics of the physical acoustic environment determine the propagation path, reflection modes, and reverberation attenuation process of sound waves after they radiate from the speaker in space.
[0003] In existing technologies, room acoustic correction techniques are typically used to improve the sound quality of audio equipment in non-ideal acoustic environments. Some adaptive audio systems monitor acoustic feedback in real time through built-in microphones and dynamically adjust filter coefficients using adaptive filtering algorithms to suppress feedback howling or compensate for near-field acoustic effects. Furthermore, in the professional audio and consumer electronics fields, there are also multi-band equalizers or preset scene modes that allow users to manually adjust parameters to a limited extent based on subjective listening experience.
[0004] However, the existing methods mentioned above generally suffer from drawbacks such as sensitivity to the position of the test microphone, difficulty in distinguishing between environmental acoustic defects and the speaker's own nonlinear distortion, and inability to adaptively balance objective acoustic indicators and subjective listening preferences. As a result, the sound quality optimization effect degrades significantly after the listening position changes or the content played is changed. Summary of the Invention
[0005] This application provides a speaker sound quality optimization method, apparatus, and audio equipment based on reinforcement learning, which addresses the technical problem of significant degradation in sound quality optimization effects after changes in listening position or playback content.
[0006] This application provides a reinforcement learning-based speaker sound quality optimization method, applied to a reinforcement learning-based speaker sound quality optimization system. The method includes: Acquire audio bitstream playback records and sound field perception records of audio equipment in the current physical acoustic environment; The audio bitstream playback record is subjected to nonlinear state characterization extraction processing to obtain a nonlinear auditory state characterization, and the sound field perception record is subjected to sound field topology analysis processing to obtain a sound field topology state characterization. Using the nonlinear auditory state representation and the sound field topology state representation as joint state input, the joint state input is decomposed into acoustic environment inherent attribute factor representation and speaker nonlinear response factor representation through the implicit causal graph structure of the preset causal decoupling strategy agent, and a sound quality transfer action sequence is generated based on the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation. The audio quality transfer action sequence is executed to update the operating parameters of the adjustable processing nodes in the audio signal processing link. The subsequent audio bitstream is played according to the updated audio signal processing link, and a set of auditory abstract primitive feedback is collected. A reward signal is generated using the aforementioned auditory abstract primitive feedback set. The joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input are stored as experience transfer tuples in the counterfactual auditory trajectory playback buffer, so that the causal decoupling strategy agent can perform counterfactual policy gradient search and update based on the implicit causal graph structure.
[0007] One embodiment of this application provides a speaker sound quality optimization device based on reinforcement learning, the device comprising: The audio recording acquisition module is used to acquire audio bitstream playback records and sound field perception records of audio equipment in the current physical acoustic environment; The state representation extraction module is used to perform nonlinear state representation extraction processing on the audio bitstream playback record to obtain a nonlinear auditory state representation, and to perform sound field topology analysis processing on the sound field perception record to obtain a sound field topology state representation. The sound quality transfer processing module is used to decompose the joint state input into acoustic environment inherent attribute factor representation and speaker nonlinear response factor representation through the implicit causal graph structure of the preset causal decoupling strategy agent, using the nonlinear auditory state representation and the sound field topology state representation as joint state input, and generate a sound quality transfer action sequence based on the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation. The migration action execution module is used to execute the sound quality migration action sequence to update the operating parameters of the adjustable processing node in the audio signal processing link, and to perform playback processing on the subsequent audio bitstream according to the updated audio signal processing link and collect the listening perception abstract primitive feedback set. The policy gradient update module is used to generate a reward signal using the auditory abstract primitive feedback set, and store the joint state input, the sound quality transfer action sequence, the reward signal and the updated joint state input as experience transfer tuples in the counterfactual auditory trajectory playback buffer, so that the causal decoupling policy agent can perform counterfactual policy gradient search and update based on the implicit causal graph structure.
[0008] One embodiment of this application provides an audio device, which is communicatively connected to a reinforcement learning-based speaker sound quality optimization system. The reinforcement learning-based speaker sound quality optimization system is used to interact with the audio device to implement any of the reinforcement learning-based speaker sound quality optimization methods described above.
[0009] One embodiment of this application provides a speaker sound quality optimization system based on reinforcement learning, including: A processor; a storage device having a computer program stored thereon; a network interface for providing network communication functions; when the computer program is executed by the processor, the processor enables the processor to implement any of the reinforcement learning-based speaker sound quality optimization methods described above.
[0010] One embodiment of this application provides a readable storage medium storing a program or instructions, which, when executed by a processor, implements the steps of the reinforcement learning-based speaker sound quality optimization method.
[0011] Therefore, the embodiments of this application have the following beneficial effects: By acquiring the audio bitstream playback record and sound field perception record of the audio device in the current physical acoustic environment, and performing nonlinear state representation extraction processing and sound field topology analysis processing respectively, a dual state representation system that takes into account both the nonlinear characteristics of human auditory perception and the spatial propagation attributes of the physical sound field is constructed. Based on this, using the nonlinear auditory state representation and the sound field topology state representation as joint state input, the implicit causal graph structure of the intelligent agent using the pre-set causal decoupling strategy is used to separate and decouple the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation mixed in the joint state input at the causal level. This allows the subsequently generated sound quality transfer action sequence to independently respond to changes in the environmental acoustic boundary and changes in the speaker's own distortion characteristics, avoiding the interference of mutual confusion between environmental and device factors on sound quality adjustment decisions. By executing a sequence of sound quality transfer actions, the operating parameters of adjustable processing nodes in the audio signal processing chain are adaptively updated. A reward signal is generated by collecting a set of auditory abstract primitive feedbacks that fuses objective sound field measurement data and subjective auditory abstract primitive scores. Then, an experience transfer tuple containing the joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input is stored in a counterfactual auditory trajectory playback buffer. This supports the agent using a causal decoupling strategy to perform counterfactual policy gradient search and update based on an implicit causal graph structure. This allows the agent to infer potential auditory reward differences between different action sequences under the same environmental conditions from limited interaction experience, improving the efficiency and generalization robustness of audio equipment's autonomous sound quality optimization in unknown and complex acoustic environments. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating a speaker sound quality optimization method based on reinforcement learning, provided in an embodiment of this application.
[0014] Figure 2 This is a schematic diagram of the basic structure of a speaker sound quality optimization system based on reinforcement learning, provided in an embodiment of this application.
[0015] Figure 3 This is a functional block diagram of a speaker sound quality optimization device provided in an embodiment of this application.
[0016] Figure 4 This is a schematic diagram illustrating an application scenario interaction provided in an embodiment of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0018] Please see Figure 1 , Figure 1 This is a flowchart of a speaker sound quality optimization method based on reinforcement learning provided in an embodiment of this application. The method can be executed by a speaker sound quality optimization system based on reinforcement learning, or it can be jointly executed by the speaker sound quality optimization system based on reinforcement learning and a server. The method may include steps 110-150.
[0019] This application provides a speaker sound quality optimization method based on reinforcement learning. This method can be applied to electronic devices such as smart speakers, home theater systems, and car audio systems that require adaptive adjustment of sound reproduction quality. The method separates the inherent physical properties of the acoustic environment from the nonlinear distortion characteristics of the speaker itself at the causal level through a causal decoupling strategy agent. Based on the separated factors, a sound quality transfer action sequence is generated that acts on multiple adjustable points in the audio signal processing link. By collecting auditory feedback and combining it with a counterfactual experience playback mechanism, the strategy agent is iteratively updated, enabling the speaker to autonomously explore and converge to the optimal sound quality parameter configuration state that is adapted to the current physical acoustic environment and the current playback content.
[0020] For ease of explanation, the following section uses an audio device as a unified implementation carrier to elaborate on the specific implementation method of the technical solution of this application. This audio device is equipped with a sound field topology sensing unit array, which consists of multiple spatially distributed sound pressure sensing units capable of capturing the acoustic response characteristics of the physical environment surrounding the speaker in real time. Simultaneously, the audio device integrates a flexibly configurable audio signal processing link, which includes multiple adjustable processing nodes such as crossover network nodes, dynamic gain control nodes, and spatial virtualization nodes. The operating parameters of each node can be dynamically modified by receiving external commands.
[0021] Step 110: Obtain audio bitstream playback records and sound field perception records of the audio device in the current physical acoustic environment.
[0022] In this embodiment, the audio equipment is in normal working condition and continuously plays the audio bitstream from the audio source device, which can be a network streaming media server, local storage medium, or a mobile terminal connected wirelessly. The audio bitstream playback record consists of metadata and frame-level feature parameters extracted in real time by the audio decoder during the decoding process.
[0023] Specifically, during the process of decoding the input compressed audio bitstream into a pulse-code modulation (PMC) format digital audio data stream, the audio decoder simultaneously extracts the modified cosine transform domain coefficient sequence carried by each frame of audio data. This coefficient sequence is organized frame by frame, with each frame containing a set of coefficient values corresponding to the number of frequency domain subbands. Simultaneously, the audio decoder also extracts transient identifiers for each audio frame from the frame header information and auxiliary data areas of the audio bitstream. These transient identifiers indicate whether the current frame contains a sudden, non-steady-state acoustic event, such as the instant a percussion instrument is struck or the beginning of a plosive sound in speech. The audio bitstream playback record is composed of the modified cosine transform domain coefficient sequence arranged in timestamp order and the transient identifier sequence.
[0024] Meanwhile, an array of sound field topology sensing units, positioned at different locations within the speaker enclosure, continuously collects sound pressure level (SPL) values from the surrounding environment at a preset sampling rate. Each SPL unit outputs a SPL value at each sampling moment. The SPL values collected by all SPL units at the same time are arranged according to their positions in the spatial coordinate system, forming an instantaneous SPL gradient change vector. The sound field sensing record consists of multiple instantaneous SPL gradient change vectors continuously collected within a time window. Additionally, the sound field sensing record includes envelope curves of SPL decay over time, independently recorded by each SPL unit. These envelope curves reflect the reverberation decay process as sound waves gradually dissipate energy after excitation in the physical environment. The audio bitstream playback record and the sound field sensing record are synchronized on the time axis and stored in real-time in a temporary state representation buffer within the speaker equipment.
[0025] Step 120: Perform nonlinear state characterization extraction processing on the audio bitstream playback record to obtain a nonlinear auditory state characterization, and perform sound field topology analysis processing on the sound field perception record to obtain a sound field topology state characterization.
[0026] After obtaining the audio bitstream playback record and the sound field perception record, the representation extraction module inside the audio equipment performs differentiated feature engineering processes on the two types of records to generate standardized state representations that can respectively characterize the auditory perception characteristics of the human ear and the physical sound field spatial structure. The processing of the audio bitstream playback record aims to simulate the nonlinear perception mechanism of the human ear to sound signals, thereby extracting auditory state representations that are highly correlated with subjective listening experience. The processing of the sound field perception record aims to analyze the propagation path, reflection mode, and energy attenuation law of sound waves in physical space from multi-channel sound pressure data, thereby extracting topological state representations that are highly correlated with the objective properties of the acoustic environment.
[0027] The generation process of the two types of representations will be explained in detail below: Step 121: Extract the modified cosine transform domain coefficient sequence from the audio bitstream playback record, and rearrange it according to the human ear critical frequency band scale to generate a critical frequency band energy sequence; perform loudness perception compression processing on each frequency band energy value in the critical frequency band energy sequence to obtain an auditory excitation spectrum representation.
[0028] In this embodiment, the characterization extraction module first reads the modified cosine transform domain coefficient sequence from the audio bitstream playback record in the state characterization temporary buffer. This modified cosine transform domain coefficient sequence has been weighted and superimposed using a window function during the audio decoding stage, and each coefficient corresponds to the spectral energy component at a specific frequency index. Based on the equivalent rectangular bandwidth scale widely used in research on the masking effect of human hearing, the characterization extraction module maps the modified cosine transform domain coefficients on the linear frequency axis to multiple non-uniformly divided critical frequency bands.
[0029] The specific mapping method is as follows: For each critical frequency band, determine the index range of all modified cosine transform domain coefficients contained within the frequency interval defined by the lower and upper limits. Square all modified cosine transform domain coefficients within this index range and sum them to obtain the original energy value of the critical frequency band. After performing the above operation on all critical frequency bands, a critical frequency band energy sequence is obtained, the length of which is equal to the total number of critical frequency bands used.
[0030] Subsequently, the characterization extraction module performs loudness-perceived compression processing on the energy values of each frequency band in the critical frequency band energy sequence. The loudness-perceived compression processing employs a power-law compression rule, using each frequency band energy value as the base and a preset perceived compression factor as the exponent for power-law operation, resulting in the compressed auditory excitation value. The perceived compression factor is less than 1, used to simulate the nonlinear compression characteristics of the human auditory system's response to sound intensity. After power-law compression processing, the critical frequency band energy sequence is converted into an auditory excitation spectrum representation. This representation is a real-valued vector with the same number of dimensions as the critical frequency band, where each element represents the subjective perceived intensity of sound energy within that critical frequency band.
[0031] Step 122: Extract the transient identifier sequence of each audio frame from the audio bitstream playback record, count the frequency of transient identifier occurrence within the sliding time window to generate an envelope mutation density curve, perform peak-preserving smoothing on the envelope mutation density curve, extract local peak regions exceeding the perception threshold, record the start time position, peak amplitude, and duration of each local peak region to obtain a mutation event set; map each event in the mutation event set into a multi-dimensional feature vector containing a normalized start time parameter, a normalized peak amplitude parameter, and a normalized duration parameter, and combine all multi-dimensional feature vectors to determine the temporal mutation representation.
[0032] Simultaneously, the characterization extraction module extracts temporal abrupt change features from the transient identifier sequence in the audio bitstream playback record. The transient identifier sequence is a one-dimensional sequence of Boolean values, the same length as the audio frame sequence, where frames marked as true indicate the presence of a transient event. The characterization extraction module defines a sliding time window with a preset window length, which slides across the transient identifier sequence in single-frame increments. At each position the window stops at, the number of frames marked as true within the window is counted, and this count is used as the amplitude of the envelope abrupt change density curve at the center of the current window. After the window traverses the entire transient identifier sequence, an envelope abrupt change density curve corresponding to the audio time axis is generated.
[0033] Next, peak-holding smoothing is performed on the envelope mutation density curve. The specific process is as follows: a hold time parameter is set. When the curve amplitude rises to a new local maximum, the smoothed curve amplitude remains unchanged at that local maximum for the subsequent hold time, unless a higher local maximum appears within the hold time. If the curve amplitude decays to a lower level after the hold time, the smoothed curve amplitude decays along the falling edge of the envelope. After obtaining the smoothed envelope mutation density curve, the characterization extraction module sets a perception threshold to identify all continuous intervals on the curve where the amplitude exceeds the perception threshold, and treats each continuous interval as a local peak region. For each local peak region, its starting position, the maximum amplitude within the region as the peak amplitude, and the duration of the region are recorded. All identified local peak regions together constitute the mutation event set.
[0034] Then, the characterization extraction module normalizes and vectorizes each event in the mutation event set. The ratio of the start time position to the total duration of the audio segment is converted into a normalized start time parameter, the relative intensity of the peak amplitude to the perceptual threshold is converted into a normalized peak amplitude parameter, and the ratio of the duration to the sliding window length is converted into a normalized duration parameter. For each mutation event, the above three normalized parameters are combined into a three-dimensional feature vector. The three-dimensional feature vectors corresponding to all events in the mutation event set are arranged in the order of their occurrence, together forming a temporal mutation representation.
[0035] Step 123: Represent the auditory excitation spectrum as a first multidimensional vector set, and the temporal abrupt change as a second multidimensional vector set. Align and concatenate all multidimensional feature vectors in the first and second multidimensional vector sets in the time dimension to generate a joint multidimensional vector as the nonlinear auditory state representation.
[0036] After obtaining the auditory excitation spectrum representation and the temporal abrupt change representation respectively, the representation extraction module performs a representation fusion operation to form a unified nonlinear auditory state representation. Although the auditory excitation spectrum representation is formally represented by a vector corresponding to each audio frame, it represents the quasi-static perceptual characteristics of the frequency domain energy distribution, while the temporal abrupt change representation is sparsely distributed, with effective vectors existing only near the time point where the abrupt event is detected. The representation extraction module first expands the auditory excitation spectrum representation along the time axis into a first multidimensional vector set of the same length as the audio frame sequence, where each element is the auditory excitation spectrum vector of the corresponding audio frame. For the temporal abrupt change representation, the representation extraction module generates a second multidimensional vector set of the same length as the audio frame sequence and with the same vector dimension. This set fills the local peak region corresponding to the abrupt event with the corresponding three-dimensional feature vector, and fills the remaining non-event regions with zero vectors.
[0037] Then, the first and second multidimensional vector sets are aligned along the time dimension to ensure that the two vectors at the same audio frame index position are strictly corresponding in time. After alignment, for each audio frame index position, the auditory excitation spectrum vector and the temporal change vector at that position are concatenated at the end of the vectors, i.e., the two vectors are connected end-to-end along the feature dimension to generate a higher-dimensional joint multidimensional vector. The joint multidimensional vectors at all audio frame index positions are combined in temporal order to obtain the nonlinear auditory state representation. This nonlinear auditory state representation integrates the nonlinear perception information of the human ear on the energy distribution of the sound spectrum and the temporal sensitivity information of the human ear to transient changes in sound.
[0038] Step 124: Extract the sound pressure level gradient change sequence from the sound field sensing record. The sound pressure level gradient change sequence is composed of sound pressure level values collected by multiple sound pressure sensing units in the sound field topology sensing unit array at the same time, arranged according to their spatial positions. Perform spatial gradient field decomposition processing on the sound pressure level gradient change sequence. Based on the Helmholtz decomposition principle, the sound pressure level gradient change sequence is separated into an irrotational gradient field part and a divergence-free gradient field part. The irrotational gradient field part corresponds to the direct sound propagation path contribution of the sound wave from the sound source directly to each sound pressure sensing unit, and the divergence-free gradient field part corresponds to the reflected sound propagation path contribution of the sound wave after boundary reflection to each sound pressure sensing unit.
[0039] For the extraction of the sound field topology state representation, the representation extraction module reads the multi-channel sound pressure data collected by the sound field topology sensing unit array from the sound field sensing record. At each sampling moment, all M sound pressure sensing units in the sound field topology sensing unit array synchronously output a set containing M sound pressure level values. The representation extraction module arranges the M sound pressure level values at each sampling moment according to the preset coordinate positions of each sound pressure sensing unit in physical space, forming the sound pressure level gradient change vector at that sampling moment.
[0040] Within an observation time window, the sound pressure level gradient change vectors at all sampling times are arranged in chronological order, forming a sound pressure level gradient change sequence. To separate the different propagation path components of the sound wave from this sound pressure level gradient change sequence, the characterization extraction module performs spatial gradient field decomposition processing on the sound pressure level gradient change vector at each sampling time. The theoretical basis of this decomposition processing is the Helmholtz decomposition principle, that is, any sufficiently smooth vector field can be uniquely decomposed into a superposition of an irrotational field component and a divergence-free field component.
[0041] In this application scenario, the sound pressure gradient field can be considered as a vector field. The characterization extraction module first calculates the gradient tensor of the sound pressure level gradient change vector in the spatial domain, and then reconstructs the scalar potential function and the vector potential function by solving the Poisson equation numerically. The negative gradient of the scalar potential function corresponds to the irrotational gradient field portion, where the sound energy flow direction is from the sound source outwards. This represents the sound pressure distribution contributed by the direct sound propagation path from the speaker unit of the audio equipment to each sound pressure sensing unit without any boundary reflection. The curl of the vector potential function corresponds to the divergence-free gradient field portion, where the sound energy flow exhibits a closed vortex shape. This represents the sound pressure distribution jointly contributed by the reflected sound propagation paths after one or more reflections from walls, ceilings, floors, and indoor object surfaces to each sound pressure sensing unit. Through spatial gradient field decomposition, the original sound pressure level gradient change vector is separated into two mutually orthogonal components at each sampling time.
[0042] Step 125: The irrotational gradient field portion is converted into a first spatial distribution matrix composed of the sound pressure amplitude and phase at the location of each sound pressure sensing unit and used as a representation of the direct sound field component. The divergence-free gradient field portion is converted into a second spatial distribution matrix composed of the sound pressure amplitude and phase at the location of each sound pressure sensing unit and used as a representation of the reflected sound field component.
[0043] After completing the spatial gradient field decomposition, the characterization extraction module constructs matrix representations for both the irrotational and divergent gradient field components. For the irrotational gradient field component, the characterization extraction module uses the spatial distribution of the scalar potential function to calculate the sound pressure amplitude and phase at the spatial location of each sound pressure sensing unit. The sound pressure amplitude is determined by the magnitude of the gradient of the potential function at that location, while the sound pressure phase is calculated by combining the product of the direction angle of the gradient of the potential function at that location and the wave number of the sound wave with a reference phase. The sound pressure amplitudes and phases at all M sound pressure sensing unit locations are organized into an M×2 first spatial distribution matrix. The element in the m-th row and 1-th column of the matrix stores the sound pressure amplitude of the m-th sound pressure sensing unit, and the element in the m-th row and 2-th column stores the sound pressure phase of the m-th sound pressure sensing unit. This first spatial distribution matrix is the representation of the direct sound field component.
[0044] Similarly, for the divergence-free gradient field portion, the characterization extraction module uses the curl distribution of the vector potential function to calculate the sound pressure amplitude and phase contributed by the reflected sound field at the spatial location of each sound pressure sensing unit. Using the same matrix organization method as above, the calculated reflected sound pressure amplitude and phase at each sound pressure sensing unit location are organized into an M×2 second spatial distribution matrix, which is the characterization of the reflected sound field components. The characterization of the direct sound field components and the characterization of the reflected sound field components describe, in the spatial dimension, the relative distribution pattern of the direct energy from the sound source and the boundary reflected energy in the current physical environment.
[0045] Step 126: Extract the reverberation attenuation envelope sequence from the sound field sensing record. The reverberation attenuation envelope sequence is composed of the envelope curves of sound pressure level attenuation over time recorded by each sound pressure sensing unit in the sound field topology sensing unit array. Perform initial attenuation segment separation processing on each envelope curve in the reverberation attenuation envelope sequence. Extract an initial segment from the arrival time of the direct sound to the preset early reflection time boundary from each envelope curve. Perform arrival direction estimation processing on the initial segment to generate arrival direction angle parameters and relative energy parameters corresponding to each early reflection sound event. The arrival direction angle parameters and relative energy parameters corresponding to all early reflection sound events constitute the early reflection sound direction distribution characterization.
[0046] In addition to spatial gradient field decomposition, the characterization extraction module also needs to extract parameters reflecting the reverberation characteristics in the time domain from the sound field perception record. The sound field perception record contains the envelope curve of the sound pressure level decaying over time, independently recorded by each sound pressure sensing unit in the sound field topology sensing unit array. These envelope curves constitute a reverberation decay envelope sequence. Each envelope curve starts after the audio equipment emits a test pulse or stops playing a steady-state noise segment, recording the process of the sound energy at the location of that sensing unit gradually decaying from the initial state to the background noise level.
[0047] The characterization extraction module performs initial attenuation segment separation processing on each envelope curve in the reverberation attenuation envelope sequence. First, the arrival time of the direct sound is located on each envelope curve using a peak detection algorithm. This moment typically corresponds to the instant when the amplitude of the envelope curve rises sharply to its maximum value. Then, a preset early reflection time boundary is set, which is estimated based on the round-trip time of the sound wave between the audio equipment and the nearest reflecting wall. From the arrival time of the direct sound to the preset early reflection time boundary, an initial segment of the envelope curve is extracted. This initial segment contains several discrete early reflection events immediately following the direct sound. The arrival direction estimation process is performed on the initial segment of each envelope curve. The arrival direction estimation process calculates the incident azimuth and elevation angles of the reflected sound based on the arrival time difference or phase difference before different sensing units in the sound field topology sensing unit array receive the same early reflected sound wave, generating arrival direction angle parameters.
[0048] Simultaneously, by analyzing the ratio of the local amplitude jump caused by the reflected sound on the envelope curve to the amplitude of the direct sound, a relative energy parameter is generated. The arrival direction angle parameters and relative energy parameters of all early reflected sound events detected from the initial segments of the envelope curves of all sound pressure sensing units are collected together to form an early reflected sound direction distribution characterization, which describes the density of early reflected sound in three-dimensional space and the energy intensity in each direction.
[0049] Step 127: Perform late decay segment analysis on each envelope curve in the reverberation decay envelope sequence, extract the portion of each envelope curve after the preset early reflection time boundary, and perform spatial averaging on the late decay segments of each sound pressure sensing unit to obtain an average late decay curve. Extract the late reverberation decay rate parameter based on the logarithmic slope of the average late decay curve, and use the late reverberation decay rate parameter as a characterization of late diffusion field energy decay.
[0050] The characterization extraction module further analyzes the late portion of the reverberation attenuation envelope sequence. For each envelope curve, the portion from the preset early reflection time boundary until the envelope amplitude decays to the background noise level is extracted; this portion is called the late attenuation segment. Within the late attenuation segment, the sound wave undergoes sufficient multiple reflections in space, the sound field tends to diffuse, the energy is uniformly distributed in all directions, and the envelope curve exhibits a relatively smooth exponential decay shape. To eliminate the potential local standing wave or interference effects at the location of a single sensing unit, the characterization extraction module performs spatial averaging on the late attenuation segments from all M sound pressure sensing units.
[0051] Specifically, the amplitudes of the M late-decay curves at the same time offset are arithmetically averaged to generate an average late-decay curve. Then, a logarithmic coordinate transformation is performed on the average late-decay curve, keeping the time axis linear while converting the amplitude axis to a logarithmic scale. In the transformed logarithmic coordinate system, the average late-decay curve will approximate as a straight line with a negative slope within the stable decay range. The characterization extraction module calculates the absolute value of the slope of this line using a linear fitting algorithm and uses this absolute value as the late-decay reverberation decay rate parameter. This parameter reflects the degree of sound energy decay per unit time in the diffuse sound field and is directly related to the equivalent sound-absorbing area of the physical environment and the room volume. The extracted late-decay reverberation decay rate parameter is used as a characterization of the energy decay of the late-decay diffuse field.
[0052] Step 128: The direct sound field component representation, the reflected sound field component representation, the early reflected sound direction distribution representation, and the late diffused field energy attenuation representation are spliced together in a preset combination order to generate a composite multidimensional array containing spatial sound field structure and time reverberation attenuation properties as the sound field topology state representation.
[0053] After extracting the above sub-representations, the representation extraction module integrates the direct sound field component representation, reflected sound field component representation, early reflected sound direction distribution representation, and late diffused field energy attenuation representation into a unified sound field topological state representation. Since the data structure dimensions of each sub-representation are different, the first spatial distribution matrix and the second spatial distribution matrix are both M×2 two-dimensional arrays; the early reflected sound direction distribution representation can be represented as a list containing multiple triplets of data, each triplet containing azimuth, elevation, and relative energy; the late diffused field energy attenuation representation is a scalar value.
[0054] To facilitate unified processing by the subsequent causal decoupling strategy agent, the representation extraction module expands and concatenates the aforementioned heteroproton representations into a one-dimensional composite multidimensional array. The concatenation operation is performed according to a preset combination order, which may be: first, arranging all elements of the first spatial distribution matrix sequentially; then, arranging all elements of the second spatial distribution matrix sequentially; next, expanding and arranging the triplet data from the early reflected sound direction distribution representation in the order of detection time; and finally, appending the late reverberation decay rate parameter value to the end of the array. The composite multidimensional array generated through this multidimensional array concatenation operation is the sound field topological state representation. The sound field topological state representation completely encodes the spatial sound field structure and temporal reverberation decay properties of the current physical acoustic environment.
[0055] Step 130: Using the nonlinear auditory state representation and the sound field topology state representation as joint state input, the joint state input is decomposed into acoustic environment inherent attribute factor representation and speaker nonlinear response factor representation through the implicit causal graph structure of the preset causal decoupling strategy agent, and a sound quality transfer action sequence is generated based on the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation.
[0056] After generating nonlinear auditory state representations and sound field topological state representations, these two are passed as joint state inputs to a pre-defined causal decoupling strategy agent. This causal decoupling strategy agent is a reinforcement learning policy network with an embedded implicit causal graph structure. Its core function is to decouple and separate the confounding factors affecting the final listening quality at the causal level, thereby independently evaluating the impact of the audio equipment's own nonlinear distortion on sound quality without interference from environmental factors, and generating accurate sound quality transfer actions accordingly. Step 131: Read the nonlinear auditory state representation and the sound field topology state representation from the state representation temporary buffer, and perform a head-to-tail concatenation operation on the vector dimension of the nonlinear auditory state representation and the sound field topology state representation to generate a joint state input vector.
[0057] The input interface of the causal decoupling strategy agent reads the generated nonlinear auditory state representation vector and sound field topology state representation vector for the current time step from the temporary state representation buffer. Let the dimension of the nonlinear auditory state representation vector be Daud and the dimension of the sound field topology state representation vector be Dgeo. The input interface performs a vector concatenation operation, directly appending the element sequence of the sound field topology state representation vector to the end of the element sequence of the nonlinear auditory state representation vector, generating a joint state input vector with dimensions Daud + Dgeo. This joint state input vector contains all available state information from the auditory perception path and the physical sound field perception path at the current time step, and is the sole information input for the causal decoupling strategy agent to perform subsequent causal inference and decision-making.
[0058] Step 132: Input the joint state input vector into the implicit causal graph structure input layer of the causal decoupling strategy agent. The implicit causal graph structure consists of multiple causal nodes and directed causal edges between nodes. Each causal node corresponds to a potential physical factor that affects the sound quality.
[0059] The joint state input vector is first fed into the implicit causal graph structure input layer of the causal decoupling strategy agent. This implicit causal graph structure is a predefined directed acyclic graph computational model containing multiple causal nodes. Each causal node is represented in the model as a learnable embedding vector, with its initial value randomly set and updated via backpropagation during subsequent training. Directed causal edges between nodes represent a hypothetical causal relationship, i.e., the state of a parent node causally influences the state of its child nodes. The connections between directed causal edges are pre-constructed based on prior knowledge of acoustic physics; for example, a room volume causal node points to a reverberation time causal node, a loudspeaker unit diaphragm displacement causal node points to a harmonic distortion causal node, and so on. Each directed causal edge is associated with a causal edge weight parameter, which represents the strength of the causal influence of the parent node on its child nodes. During the forward propagation of the implicit causal graph structure, information from the joint state input vector flows and transforms between the embedding vectors of the causal nodes through a series of message passing and node update operations in the graph neural network.
[0060] Step 133: In the first sub-network of the implicit causal graph structure, the joint state input vector is mapped to the acoustic environment intrinsic attribute factor node through the first set of causal edge weights, and the acoustic environment intrinsic attribute factor node outputs the acoustic environment intrinsic attribute factor representation.
[0061] The implicit causal graph structure is internally divided into a first sub-network and a second sub-network, which are used to extract environmental intrinsic attribute factors and speaker response factors, respectively. In the first sub-network, after the input layer receives the joint state input vector, it selectively passes information to causal nodes related to physical environmental attributes through the first set of causal edge weights, such as causal nodes related to room size, wall sound absorption coefficient, and listening position.
[0062] During information transmission, the network structure forcibly blocks the influence of causal nodes related to the speaker's nonlinear response, ensuring that the first sub-network is only sensitive to environmental factors. The end of the first sub-network converges to a special acoustic environment intrinsic property factor node. The output embedding vector of this node is the representation of the acoustic environment intrinsic property factor. This representation vector extracts the potential invariant description related to the current physical acoustic environment spatial structure, boundary sound absorption characteristics, and the relative position of the listener from the joint state input vector, and is independent of the content being played, the volume, and the amplifier distortion state at this moment.
[0063] Step 134: In the second sub-network of the implicit causal graph structure, the joint state input vector is mapped to the speaker nonlinear response factor node through the second set of causal edge weights. During the mapping process, a conditional independence constraint is introduced from the acoustic environment intrinsic attribute factor node, so that the speaker nonlinear response factor node outputs a speaker nonlinear response factor representation that is decoupled from the acoustic environment intrinsic attribute factor representation.
[0064] In parallel with the first subnetwork, the second subnetwork transmits information from the joint state input vector to causal nodes related to the speaker's nonlinear response through a second set of causal edge weights, such as power amplifier gain causal nodes, speaker unit suspension compliance causal nodes, magnetic circuit symmetry causal nodes, and crossover phase offset causal nodes.
[0065] To ensure that the extracted speaker nonlinear response factor representation and the acoustic environment inherent attribute factor representation are statistically independent, conditional independence constraints are introduced during network training and inference.
[0066] Specifically, a mutual information minimization regularization term is added to the loss function. This term measures the dependence of the conditional probability distributions of the acoustic environment's inherent attribute factors and the conditional probability distributions of the speaker's nonlinear response factor factors on the given joint state input. By optimizing this loss function through gradient descent, the network is prompted to learn a mapping relationship that minimizes the amount of information that can be linearly predicted by environmental factors in the embedding vectors output by the speaker's nonlinear response factor nodes.
[0067] Finally, the output of the second sub-network is a speaker nonlinear response factor representation. This representation vector characterizes the nonlinear distortion characteristics introduced by the audio equipment's own electroacoustic conversion link under the current playback conditions and signal dynamics, such as the energy distribution of harmonic distortion components, intermodulation distortion products, and compression dynamics.
[0068] Step 135: Input the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation into the action generation module of the causal decoupling strategy agent. The action generation module includes action parameter generation branches that correspond one-to-one with the adjustable processing nodes in the audio signal processing link.
[0069] After decoupling the causal factors, the representation vectors of the acoustic environment's inherent attribute factors and the speaker's nonlinear response factor factors are input together into the action generation module of the causal decoupling strategy agent. The action generation module adopts a multi-branch parallel output architecture, which internally sets up multiple parallel action parameter generation branches, each corresponding to an adjustable processing node in the audio signal processing chain.
[0070] Specifically, the audio signal processing link in this embodiment includes a frequency division network node, a dynamic gain control node, and a spatial virtualization node. Therefore, the action generation module is correspondingly equipped with three action parameter generation branches, namely the first action parameter generation branch, the second action parameter generation branch, and the third action parameter generation branch. Each branch is composed of a stack of fully connected neural network layers and nonlinear activation function layers, which are responsible for mapping the input factor representation to the specific parameter adjustment amount of the corresponding processing node.
[0071] Step 136: Generate branch generation frequency band energy redistribution action, gain curve shape adjustment action and sound image trajectory reshaping action using different action parameters in the action generation module, and determine the sound quality transfer action sequence.
[0072] In this step, the three action parameter generation branches each output their respective adjustment commands. The output of the first action parameter generation branch is interpreted as a frequency band energy redistribution action, which includes a set of control words for resetting the frequency division points and slope parameters of each frequency band in the crossover network nodes. The output of the second action parameter generation branch is interpreted as a gain curve shape adjustment action, which includes a set of control words for modifying the compression characteristic curve threshold, ratio, and time constant of the dynamic gain control node. The output of the third action parameter generation branch is interpreted as a sound image trajectory reshaping action, which includes a set of control words for adjusting the azimuth coordinates of the virtual sound source and the spatial reflection gain of the spatial virtualization node.
[0073] The action generation module merges and packages the three parallel generated action instructions according to the preset instruction encapsulation format to generate a complete audio quality transfer action sequence. This audio quality transfer action sequence is a single decision output made by the causal decoupling strategy agent in response to the current joint state input.
[0074] Step 1361: In the first action parameter generation branch corresponding to the frequency division network node, a frequency band energy redistribution action is generated based on the reflected sound field direction distribution information in the acoustic environment inherent attribute factor representation and the frequency response non-flatness characteristics in the speaker nonlinear response factor representation.
[0075] Specifically, in the first action parameter generation branch, its input layer receives the spliced representations of the acoustic environment's inherent attribute factors and the speaker's nonlinear response factor representations. The neural network layer within this branch first performs attention-weighted processing on the acoustic environment's inherent attribute factor representations, focusing on extracting components related to the directional distribution of the reflected sound field. This directional distribution information reveals which direction's early reflected sound energy is too strong, potentially leading to a widened sound image or decreased clarity.
[0076] Simultaneously, the network layer also extracts components reflecting the non-flat characteristics of the current speaker's frequency response from the speaker's nonlinear response factor characterization, such as the inherent peaks and valleys in a certain frequency band. The first action parameter generation branch calculates based on the above two types of information and outputs a set of parameter increment values for adjusting the crossover network nodes. For example, when significant early reflection energy is detected from the sidewall direction and the speaker itself has a response peak in that frequency band, the frequency band energy redistribution action output by the branch will instruct the crossover network nodes to appropriately increase the crossover point frequency or increase the attenuation slope of the filter in that frequency band to suppress the energy output of that frequency band and balance the overall listening experience.
[0077] Step 1362: In the second action parameter generation branch corresponding to the dynamic gain control node, a gain curve shape adjustment action is generated based on the late reverberation decay rate information in the acoustic environment inherent attribute factor representation and the large signal compression characteristics in the speaker nonlinear response factor representation.
[0078] In the second action parameter generation branch, the network layer focuses on resolving the late reverberation decay rate parameter from the inherent property factors of the acoustic environment. A slow late reverberation decay rate indicates a long room reverberation time and significant sound trailing. Applying excessive dynamic compression to the audio signal in this case will further exacerbate the energy accumulation of reverberant sound, leading to a muddy listening experience. Simultaneously, the network layer analyzes the large-signal compression characteristics in the speaker's nonlinear response factor representation. This characteristic describes the soft inflection point compression phenomenon in the gain curve when the input signal amplitude approaches the speaker's maximum power handling. The second action parameter generation branch integrates the late reverberation decay rate and the large-signal compression characteristics to output gain curve shape adjustment actions. For example, in acoustic environments with long reverberation times, the branch may output parameter adjustment commands to increase the compression threshold, decrease the compression ratio, or shorten the release time to maintain the transient impact of the sound and avoid the adverse effects of the compressor exacerbating environmental reverberation.
[0079] Step 1363: In the third action parameter generation branch corresponding to the spatial virtualization node, based on the early reflected sound direction distribution information in the acoustic environment inherent attribute factor representation and the directivity deviation characteristics in the speaker nonlinear response factor representation, a sound image trajectory reshaping action is generated.
[0080] The third branch, which generates motion parameters, focuses on spatial sound image processing. The network layer extracts early reflected sound direction distribution information from the inherent properties of the acoustic environment, analyzing whether there is significant asymmetry in early reflected sound in the current physical environment, such as stronger reflections from the left wall than the right wall. Simultaneously, it extracts directivity deviation characteristics from the speaker's nonlinear response factor characterization. This characteristic describes the frequency response variation of the speaker at different off-axis angles, particularly the off-axis frequency response attenuation caused by directivity narrowing in the high-frequency range.
[0081] The third action parameter generation branch, combining the early reflected sound direction distribution and directivity deviation characteristics, outputs a sound image trajectory reshaping action. This action repositions the perceived location of the virtual sound source by adjusting the head-related transfer function parameters or crosstalk cancellation filter coefficients used to generate the virtual sound source within the spatial virtualization node. For example, if excessive left-side reflections are detected, causing the overall sound image to shift to the left, and the off-axis attenuation of the speaker in the high-frequency band exacerbates this shift, the sound image trajectory reshaping action output by the branch may instruct the spatial virtualization node to slightly adjust the virtual playback location of the central sound image to the right to compensate for the sound image shift caused by the physical environment and the speaker's directivity.
[0082] Step 140: Execute the audio quality transfer action sequence to update the operating parameters of the adjustable processing nodes in the audio signal processing link, and perform playback processing on the subsequent audio bitstream according to the updated audio signal processing link and collect the listening perception abstract primitive feedback set.
[0083] After generating the audio quality transfer action sequence, the agent of the causal decoupling strategy immediately parses and deploys the action sequence, transforms the action instructions into actual modifications to the hardware registers or software variable values in the audio signal processing link, drives the updated signal processing link to process the subsequent audio stream, and simultaneously initiates the feedback acquisition process to obtain a quantitative evaluation of the current action execution effect.
[0084] Step 141: Analyze the frequency band energy redistribution action in the sound quality migration action sequence, extract the frequency division network parameter adjustment instruction string contained in the frequency band energy redistribution action, and update the frequency division point register value and slope register value of the frequency division network node in the audio signal processing link according to the frequency division network parameter adjustment instruction string.
[0085] In this step, the frequency band energy redistribution action field in the sound quality transfer action sequence is first parsed. This field consists of a series of instruction strings encoded according to a preset protocol. Each instruction string contains a register address, a value to be written, and an optional parity bit. The instruction string whose address corresponds to the crossover point register of the crossover network node is identified, and the crossover point parameter value to be written is extracted from it. For example, the instruction string might instruct that the crossover point between the woofer and midrange driver be adjusted from the default value to a new value. Subsequently, the new crossover point parameter value is written to the crossover point register of the crossover network node via the bus interface or direct register write operation. Similarly, the slope register values corresponding to the slope of each frequency band filter in the crossover network node are parsed and updated. The slope register values determine the steepness of the attenuation at the cutoff band edge of the filter. After completing the write operation of all relevant registers of the crossover network node, the frequency response segmentation characteristics of the node are changed accordingly.
[0086] Step 142: Analyze the gain curve shape adjustment action in the sound quality transfer action sequence, extract the gain curve parameter adjustment instruction string contained in the gain curve shape adjustment action, and update the threshold inflection point register value, compression ratio curve slope register value, and start-up release time constant register value of the dynamic gain control node in the audio signal processing link according to the gain curve parameter adjustment instruction string.
[0087] Next, the gain curve shape adjustment action field is analyzed, which contains a series of parameter adjustment instructions for the dynamic gain control node. The instruction string corresponding to the threshold inflection point register is extracted; the threshold inflection point defines the input level boundary at which the dynamic gain control node begins to attenuate or boost the signal. Simultaneously, the instruction string corresponding to the compression ratio curve slope register is extracted; the compression ratio curve slope determines the ratio of the output level change relative to the input level after the input level exceeds the threshold. Furthermore, instruction strings corresponding to the start-up time constant register and the release time constant register are extracted; the start-up time constant determines the speed of the gain attenuation process after the signal exceeds the threshold, and the release time constant determines the speed of the gain recovery process after the signal falls back below the threshold. The extracted new parameter values are written into the corresponding registers of the dynamic gain control node, thereby instantly reshaping the input-output level transmission curve shape and dynamic response behavior of the node.
[0088] Step 143: Analyze the sound image trajectory reshaping action in the sound quality transfer action sequence, extract the spatial parameter adjustment instruction string contained in the sound image trajectory reshaping action, and update the virtual sound source orientation register value and spatial reflection intensity register value of the spatial virtualization node in the audio signal processing link according to the spatial parameter adjustment instruction string.
[0089] Then, the sound image trajectory reshaping action field is parsed to extract the spatial parameter adjustment instruction string for the spatial virtualization node. The spatial virtualization node internally maintains a set of virtual sound source azimuth registers, each storing the azimuth, elevation, and distance parameters of a virtual sound source in a three-dimensional polar coordinate system. Based on the instruction string, the value of the specified virtual sound source azimuth register is updated, thereby changing the perceived location of the virtual sound source.
[0090] Simultaneously, the spatial reflection intensity register value of the spatial virtualization node is parsed and updated. This register value is used to control the energy ratio of early reflected sound to late diffused sound or the overall gain level in the artificial reverberation algorithm. By modifying the spatial reflection intensity register, the strength of the sound field immersion and spatial sense simulated by the spatially virtualized sound signal can be adjusted.
[0091] Step 144: After completing the register value update operation of all adjustable processing nodes, it is determined that the audio signal processing link has switched to the updated parameter configuration state, and receives the subsequent audio bit stream from the audio source device. The subsequent audio bit stream is input into the audio decoder for decoding processing to obtain a pulse code modulation format digital audio data stream. The pulse code modulation format digital audio data stream is sequentially fed into the frequency division network node, dynamic gain control node and spatial virtualization node in the audio signal processing link in the updated parameter configuration state for signal processing.
[0092] Once all target registers of the frequency division network node, dynamic gain control node, and spatial virtualization node have been successfully written with the new values, a state transition confirmation signal is issued, indicating that the audio signal processing link has been migrated to the updated parameter configuration state. Afterward, subsequent audio bitstreams from the audio source device no longer undergo the previous old parameter configuration processing but instead enter the new processing flow. The subsequent audio bitstream first enters the audio decoder, which executes decompression and reconstruction algorithms to convert the compressed domain bitstream into a time-continuous, amplitude-quantized pulse-code modulation (PCM) digital audio data stream. This PCM data stream is then sequentially fed into the audio signal processing link.
[0093] The data stream first flows through a frequency division network node configured with new parameters. This node divides the full-band audio data stream into multiple independent sub-data streams, including low-frequency, mid-frequency, and high-frequency bands, based on the newly set division points and slopes. Each sub-data stream then enters a dynamic gain control node in parallel. This node independently applies dynamic range compression or expansion to each sub-data stream according to newly set thresholds, ratios, and time constants. Finally, the dynamically processed sub-data streams rejoin in a spatial virtualization node. Based on newly set virtual sound source locations and spatial reflection intensities, the spatial virtualization node performs spatial effects processing on the audio data stream, including convolution reverberation, crosstalk cancellation, and binaural synthesis.
[0094] Step 145: Obtain the processed digital audio data stream at the output end of the audio signal processing link, and convert the processed digital audio data stream into an analog audio signal to drive the speaker unit of the audio device to reproduce sound, forming an updated sound field in the current physical acoustic environment. In the updated sound field, the updated sound pressure level gradient change sequence and the updated reverberation attenuation envelope sequence are collected by the sound field topology sensing unit array.
[0095] The digital audio data stream, processed by the audio signal processing link, is converted into an analog audio signal by a digital-to-analog converter at its output. This analog audio signal is amplified by a power amplifier and then drives the speaker unit of the audio equipment to perform electroacoustic conversion, radiating sound waves into the current physical acoustic environment and forming an updated sound field. The sound waves propagate in the updated sound field and interact with the environmental boundaries.
[0096] To quantitatively evaluate the changes in objective acoustic effects resulting from the execution of this sound quality transfer sequence, the sound field topology sensing unit array was triggered again to perform a complete sound field sensing sampling. The acquisition process was the same as that in step 110 for obtaining the sound field sensing record. Each sound pressure level sensing unit in the sound field topology sensing unit array synchronously acquired sound pressure level values, generating an updated sound pressure level gradient change sequence. Simultaneously, each sound pressure level sensing unit also re-recorded the sound pressure level attenuation process after the test signal stopped, generating an updated reverberation attenuation envelope sequence. These two types of objective acoustic measurement data constitute the objective feedback portion of the auditory abstract primitive feedback set.
[0097] Step 146: Receive the set of subjective listening abstract primitive scores input by the user through the listening evaluation interface. The set of subjective listening abstract primitive scores includes the timbre transparency abstract primitive score, the sound image positioning clarity abstract primitive score, and the listening comfort abstract primitive score.
[0098] While the audio equipment is playing according to the updated parameter configuration, or during playback intervals, the system can obtain subjective listening feedback through user interaction. The audio equipment can be associated with a listening evaluation interface running on a mobile terminal or a remote control with a touchscreen. This interface guides the user to evaluate the just-played audio quality segment from several abstract dimensions, in the form of a slider or a level selector.
[0099] Specifically, the evaluation interface requires users to independently rate the abstract primitives of timbre transparency, sound image localization clarity, and listening comfort. The timbre transparency rating reflects the user's subjective judgment of whether the sound is clear, layered, and free of muffled or harsh sounds. The sound image localization clarity rating reflects the user's subjective judgment of whether the virtual sound source's location is clear and whether the sound field width and depth feel natural. The listening comfort rating reflects the user's subjective feeling of fatigue after prolonged listening and whether the sound is harsh or oppressive. The user-submitted ratings for these three abstract primitives are wirelessly transmitted back to the audio equipment and used as the subjective feedback component of the listening comfort abstract primitive feedback set.
[0100] Step 147: Combine the updated sound pressure level gradient change sequence, the updated reverberation attenuation envelope sequence, and the subjective listening abstract primitive score set to form the listening abstract primitive feedback set.
[0101] The feedback integration module inside the audio equipment receives updated sound pressure level gradient change sequences and updated reverberation attenuation envelope sequences from the sound field topology sensing unit array, and simultaneously receives a set of subjective listening abstract primitive scores from the listening evaluation interface. The feedback integration module timestamps these data and packages them into a composite data object, which is the listening abstract primitive feedback set. This set comprehensively records the dual changes caused at both the objective physical sound field level and the subjective human ear perception level after the execution of the sound quality transfer action sequence.
[0102] Step 150: Generate a reward signal using the auditory abstract primitive feedback set, and store the joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input as experience transfer tuples in the counterfactual auditory trajectory playback buffer, so that the causal decoupling strategy agent can perform counterfactual policy gradient search and update based on the implicit causal graph structure.
[0103] After acquiring the set of auditory abstract primitive feedback, the reinforcement learning process enters the reward signal calculation and policy update stage. The system needs to quantify the effect of this action execution into a scalar reward signal to guide the parameter optimization direction of the causal decoupled policy agent. At the same time, this interaction experience is stored in an experience replay buffer with a counterfactual reasoning mechanism. The agent will subsequently sample from this buffer and update the policy gradient based on the implicit causal graph structure.
[0104] Step 151: Extract the updated sound pressure level gradient change sequence from the auditory abstract primitive feedback set, and calculate the spatial gradient smoothness index of the updated sound pressure level gradient change sequence. The spatial gradient smoothness index is determined by the variance statistic of the sound pressure level difference between adjacent spatial sampling points in the updated sound pressure level gradient change sequence.
[0105] First, the objective acoustic feedback data is processed: the updated sound pressure level gradient change sequence is read from the auditory abstract primitive feedback set. For the sound pressure level gradient change vector at each sampling time in this sequence, the absolute value of the difference between the sound pressure level values measured by two adjacent sound pressure sensing units in physical space is calculated. By traversing all adjacent sampling point pairs in the vector, a set of sound pressure level difference values is obtained.
[0106] Then, the variance statistic is calculated for this set of sound pressure level differences. The calculation process for the variance statistic is as follows: first, calculate the arithmetic mean of this set of sound pressure level differences; then, calculate the square of the difference between each difference and the mean; finally, calculate the arithmetic mean of these squared values. The resulting variance statistic is the spatial gradient smoothness index for that sampling time. The spatial gradient smoothness index is averaged over time for all sampling times to obtain the overall spatial gradient smoothness index, which represents the smoothness of the sound field spatial distribution throughout the entire observation period. The smaller the value of this index, the smoother the change in sound pressure level in space and the more uniform the sound field distribution.
[0107] Step 152: Extract the updated reverberation attenuation envelope sequence from the auditory abstract primitive feedback set, and calculate the ratio of early reflection sound to late diffusion field energy of the updated reverberation attenuation envelope sequence. The ratio of early reflection sound to late diffusion field energy is determined by the ratio of the integrated energy within the time window of the early reflection sound to the integrated energy within the time window of the late diffusion field.
[0108] Next, the reverberation decay envelope sequence is processed: the updated reverberation decay envelope sequence is extracted from the auditory abstract primitive feedback set and spatially averaged to obtain an average reverberation decay envelope curve. An early reflection time window is defined, starting at the arrival time of the direct sound and lasting for a preset duration, typically tens of milliseconds after the arrival of the direct sound. Simultaneously, a late diffusion field time window is defined, starting at the end of the early reflection time window and lasting for a preset duration or until the envelope decays to the background noise level. For the average reverberation decay envelope curve, the square of the envelope amplitude is integrated over time within the early reflection time window to obtain the integrated energy of the early reflection.
[0109] Similarly, the square of the envelope amplitude is integrated over time within the late diffusion field time window to obtain the late diffusion field integrated energy. The ratio of the early reflection integrated energy to the late diffusion field integrated energy is the early reflection to late diffusion field energy ratio index. The larger the index value, the stronger the perceived clarity and intimacy of the sound, while the smaller the index value, the more prominent the spatial immersion and reverberation of the sound.
[0110] Step 153: Extract the timbre transparency abstract primitive score, sound image localization clarity abstract primitive score, and listening comfort abstract primitive score from the subjective listening abstract primitive score set from the listening abstract primitive feedback set. Then, perform a linear weighted summation of the timbre transparency abstract primitive score, the sound image localization clarity abstract primitive score, and the listening comfort abstract primitive score with preset weights to generate a comprehensive subjective listening score.
[0111] Furthermore, the system processes subjective feedback data: it reads the scores for timbre transparency, sound image localization clarity, and auditory comfort from the auditory abstract primitive feedback set. Internally, it stores a set of preset weighting coefficients: timbre transparency weighting coefficient, sound image localization clarity weighting coefficient, and auditory comfort weighting coefficient, the sum of which equals 1. The timbre transparency score is multiplied by the timbre transparency weighting coefficient to obtain the first weighting term; the sound image localization clarity score is multiplied by the sound image localization clarity weighting coefficient to obtain the second weighting term; and the auditory comfort score is multiplied by the auditory comfort weighting coefficient to obtain the third weighting term. Then, the three weighting terms are arithmetically summed, and the result is the comprehensive subjective auditory score, which is a comprehensive quantitative assessment of the user's multidimensional subjective experience.
[0112] Step 154: Based on the spatial gradient smoothness index, the ratio of early reflected sound to late diffused field energy index, and the preset objective index weighting coefficient, generate an objective acoustic quality comprehensive score.
[0113] Optionally, objective indicators can be processed in a manner similar to weighted summation. Specifically, the calculated spatial gradient smoothness index is multiplied by a preset first objective indicator weighting coefficient to obtain a first objective weighting term; the calculated ratio of early reflected sound energy to late diffused field energy is multiplied by a preset second objective indicator weighting coefficient to obtain a second objective weighting term. The sum of these two objective indicator weighting coefficients can be equal to 1. The first and second objective weighting terms are added together to obtain the overall objective acoustic quality score. In some implementations, the spatial gradient smoothness index and the ratio of early reflected sound energy to late diffused field energy index can also be normalized to their maximum and minimum values to make their numerical ranges match the range of subjective scores before weighted summation.
[0114] Step 155: Input the subjective listening perception comprehensive score and the objective acoustic quality comprehensive score into a pre-trained nonlinear fusion mapping network. Through the multilayer perceptron structure of the nonlinear fusion mapping network, perform feature interaction and nonlinear transformation processing on the subjective listening perception comprehensive score and the objective acoustic quality comprehensive score, and output the comprehensive listening perception improvement.
[0115] To integrate subjective and objective evaluations, the system utilizes a pre-trained nonlinear fusion mapping network, a compact multilayer perceptron model. Its input layer has two neurons, receiving the subjective auditory perception score and the objective acoustic quality score, respectively. The network contains several hidden layers, each consisting of multiple neurons connected by fully connected weight matrices. A nonlinear activation function, such as a hyperbolic tangent activation function or a rectified linear unit activation function, is applied after each hidden layer. The output layer contains only one neuron. After being fed into the input layer, the subjective and objective auditory perception scores undergo multiple linear transformations and nonlinear activations in the hidden layers, resulting in deep interaction and fusion of their features. This ultimately generates a comprehensive auditory improvement at the output layer. The weight parameters of this nonlinear fusion mapping network have been pre-trained to minimize the error between the network output and the overall evaluation of sound quality improvement by human experts.
[0116] Step 156: Read the historical comprehensive listening value from the historical data storage area under the previous parameter configuration state. The historical comprehensive listening value is calculated using the same nonlinear fusion mapping network after the previous sound quality migration action sequence was executed.
[0117] To calculate the change in perceived sound quality resulting from this action, a historical comprehensive listening value is retrieved from the audio equipment's historical data storage area. This historical comprehensive listening value was generated and stored after the last execution of the sound quality transfer action sequence, using a set of listening abstract primitives, and processed through the same objective index calculation, subjective scoring weighting, and nonlinear fusion mapping network process. The historical comprehensive listening value represents the baseline level of perceived sound quality achievable with the audio equipment's parameter configuration before this action was executed.
[0118] Step 157: Subtract the overall listening quality improvement from the historical overall listening quality value to obtain the listening quality change. Apply a piecewise linear reward shaping function to the listening quality change to map positive changes to a positive reward interval and negative changes to a negative reward interval to obtain the reward signal.
[0119] Subtract the overall listening quality improvement output in step 155 from the historical overall listening quality value read in step 156 to obtain the change in listening quality. If the change is positive, it means that the current sound quality transfer action has improved the listening quality compared to the previous state; if the change is negative, it means that the listening quality has decreased.
[0120] Subsequently, a pre-defined piecewise linear reward shaping function is applied to map the changes in perceived sound quality. The piecewise linear reward shaping function is defined as follows: when the input change is greater than zero, the function output is proportional to the input change, and the proportionality coefficient is positive; when the input change is less than zero, the function output is also proportional to the input change, but the proportionality coefficient is negative, and the absolute value of this coefficient can be set to be greater than the absolute value of the positive proportionality coefficient to impose a more severe penalty on deterioration in sound quality; when the input change is equal to zero, the function output is zero. The scalar output value obtained after mapping by the piecewise linear reward shaping function is the final reward signal used for reinforcement learning updates.
[0121] Step 158: The joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input are encapsulated as an experience transfer tuple to obtain a data structure as the experience transfer tuple and stored in the counterfactual auditory trajectory playback buffer. The counterfactual policy gradient search is then performed in conjunction with the monitoring of the numerical change trend of the loss function of the causal decoupling strategy agent.
[0122] After generating the reward signal for this interaction, the system encapsulates the entire process data of this decision and feedback into an experience transfer tuple. Specifically, the joint state input generated in step 131 is marked as the state before the action is executed, the sound quality transfer action sequence determined in step 136 is marked as the action to be executed, the reward signal calculated in step 157 is marked as the immediate reward, and the updated joint state input regenerated based on the auditory abstract primitive feedback set collected in step 147 is marked as the state after the action is executed.
[0123] The four data objects mentioned above are packaged into a single data structure, which is an experience transfer tuple. This experience transfer tuple is stored at the end of the storage queue of the counterfactual auditory trajectory playback buffer. The counterfactual auditory trajectory playback buffer is a circular queue structure with a fixed maximum capacity. When the storage space is full, the newly stored experience transfer tuple will overwrite the oldest stored tuple.
[0124] Step 1581: Mark the joint state input as the state before the action is executed, mark the sound quality transfer action sequence as the action to be executed, mark the reward signal as the immediate reward, mark the updated joint state input regenerated based on the auditory abstract primitive feedback set as the state after the action is executed, and encapsulate the state before the action, the action to be executed, the immediate reward, and the state after the action as a data structure as the experience transfer tuple, and store it at the end of the storage queue of the counterfactual auditory trajectory playback buffer.
[0125] This step details the specific encapsulation and storage process of the experience transfer tuple. The state labeling module adds a semantic label of "state before action" to the joint state input. The action labeling module adds a semantic label of "action performed" to the sound quality transfer action sequence. The reward labeling module adds a semantic label of "instant reward" to the reward signal. After the auditory abstract primitive feedback set is collected, the state update module immediately triggers a new representation extraction process to generate an updated joint state input and adds a semantic label of "state after action". The system serializes these four semantically labeled data items according to the standard reinforcement learning tuple format of "state, action, reward, next state" to generate a compact data structure, and pushes it to the tail of the counterfactual auditory trajectory playback buffer queue.
[0126] Step 1582: When the number of experience transfer tuples stored in the counterfactual auditory trajectory playback buffer reaches a preset batch sampling threshold, a batch of experience transfer tuples is randomly selected from the counterfactual auditory trajectory playback buffer. The gradient of the loss function of the causal decoupling strategy agent is calculated using the selected experience transfer tuples. The first set of causal edge weights, the second set of causal edge weights, and the weight parameters of each action parameter generation branch inside the causal decoupling strategy agent are updated using the gradient descent optimization algorithm.
[0127] The system continuously monitors the number of experience transition tuples in the counterfactual auditory trajectory playback buffer. When the cumulative number reaches a preset batch sampling threshold, an offline policy learning and update process is initiated. A batch of experience transition tuples is extracted from the buffer using a uniform random sampling strategy. For each experience transition tuple in this batch, the loss function value of the causal decoupling policy agent is calculated using the pre-action state, the action performed, the immediate reward, and the post-action state data contained therein. In this embodiment, the loss function used is a composite loss function that integrates policy gradient loss and value function estimation error.
[0128] Automatic differentiation is used to calculate the partial derivatives of the loss function with respect to all trainable parameters within the causal decoupling agent, i.e., the gradient. These trainable parameters include the weights of the first set of causal edges associated with the first sub-network in the implicit causal graph structure, the weights of the second set of causal edges associated with the second sub-network, and the weights of the fully connected layers in each action parameter generation branch of the action generation module. Subsequently, using a gradient descent optimization algorithm, such as the Adam optimization algorithm, all the above weight parameters are fine-tuned based on the calculated gradient values, with the update magnitude controlled by a preset learning rate parameter.
[0129] Step 1583: Monitor the trend of loss function value change of the causal decoupling strategy agent in multiple consecutive iteration update cycles. When the fluctuation amplitude of the loss function value change trend is less than the preset convergence judgment threshold, it is determined that the preset update stop condition is met, and the current snapshot of the causal decoupling strategy agent parameters is saved for the real-time sound quality optimization process of the audio equipment.
[0130] After each parameter update in step 1582, the updated loss function value is recorded. In practical applications, a sliding window can be set up to store a sequence of loss function values from the most recent consecutive update cycles. The variance or standard deviation of the loss function values in this sequence is calculated as a measure of fluctuation. This fluctuation is then compared to a preset convergence threshold. When the fluctuation of the loss function values for several consecutive cycles is less than this convergence threshold, the training process of the causal decoupling strategy agent is considered to have met the preset update stopping condition, meaning the strategy network has converged to near a local optimum. At this point, all parameter values of the current causal decoupling strategy agent are saved as a parameter snapshot file and stored in the non-volatile memory of the audio equipment. This parameter snapshot will be loaded for subsequent real-time sound quality optimization inference processes.
[0131] After the causal decoupling policy agent completes one or more rounds of updates based on counterfactual policy gradient search, in order to further improve training efficiency and policy robustness, the method of this embodiment may also include dynamic adjustment of the sampling probability of the experience replay buffer (step 210) and adaptive constraint on the learning step size of the policy network (step 310).
[0132] Step 210: Obtain multiple experience transition tuples from the counterfactual auditory trajectory playback buffer, each experience transition tuple containing a joint state input and a reward signal; using the joint state input in the multiple experience transition tuples as a first set of analysis objects, perform state representation analysis on each joint state input in the first set of analysis objects to separate the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation, obtaining an environmental factor sequence and a response factor sequence; using the reward signal in the multiple experience transition tuples as a second set of analysis objects, perform reward value distribution fitting on each reward signal in the second set of analysis objects to determine the long-tail offset trend representation of the reward distribution; re-evaluate the causal effect intensity of the environmental factor sequence and the response factor sequence through the implicit causal graph structure to generate an intervention tendency weight sequence, perform weighted recombination processing on the environmental factor sequence and the response factor sequence based on the intervention tendency weight sequence to generate a priority experience sampling distribution representation, and adjust the sampling probability weight of each experience transition tuple in the counterfactual auditory trajectory playback buffer based on the priority experience sampling distribution representation.
[0133] Once a sufficient number of experience transition tuples have accumulated in the counterfactual auditory trajectory playback buffer, the system can initiate a background analysis task. This task first retrieves a batch of experience transition tuples from the buffer. For each experience transition tuple in this batch, the pre-action state vector is re-inputted into the implicit causal graph structure of the causal decoupling strategy agent. A forward propagation is performed, but the gradient is not calculated, to separate and record the acoustic environment intrinsic attribute factor representation vector and the speaker nonlinear response factor representation vector in that state. The environmental factor representations of all tuples constitute an environmental factor sequence, and the response factor representations constitute a response factor sequence.
[0134] Simultaneously, the reward signal scalar value is extracted from each tuple to form a second set of analytical objects. This set of reward values is then fitted with a distribution, for example, using kernel density estimation to fit its probability density function. The extent of tail extension in the density function is analyzed to determine the long-tailed shift trend of the reward distribution. This long-tailed shift trend reflects whether certain specific combinations of environmental and response factors can lead to abnormally high or low rewards.
[0135] Then, the system uses an implicit causal graph structure to score the intervention tendency for each set of environmental factors and response factors. The intervention tendency score is obtained by calculating the ratio of the conditional probability to the marginal probability of a specific response factor under given environmental factors, reflecting the novelty or counterfactual importance of the state combination. Using the intervention tendency score as a weight, the environmental factor sequence and the response factor sequence are weighted and recombined to generate a preferred empirical sampling distribution representation, which is specifically represented as a probability weight array of the same length as all current empirical transition tuples in the buffer.
[0136] Thus, when random sampling is performed in step 1582, uniform sampling is no longer used. Instead, importance sampling is performed based on the probability weight array, so that experience transfer tuples located in the long tail region of the reward distribution or with a high tendency to intervene are selected for policy updates with a higher probability.
[0137] Step 310: Obtain the first set of causal edge weights and the second set of causal edge weights of the agent before and after the update; using the first set of causal edge weights before and after the update as the first comparison object, perform edge weight change magnitude analysis on the first comparison object to generate a first structural sensitivity thermal representation; using the second set of causal edge weights before and after the update as the second comparison object, perform edge weight change direction consistency analysis on the second comparison object to generate a second structural sensitivity thermal representation; perform topology-preserving fusion processing based on the implicit causal graph structure on the first and second structural sensitivity thermal representations to generate a graph structure evolution trend representation, which is used to reflect the learning dynamics of the agent on different causal decoupling paths; based on the graph structure evolution trend representation, constrain the parameter update step size of the action generation module of the agent to adjust the weight adjustment magnitude of different action parameter generation branches in the subsequent counterfactual policy gradient search update process.
[0138] After multiple updates and iterations of the policy network, the system can analyze the learning dynamics of the causal graph structure within the network and adjust the learning step size accordingly. The system obtains the first and second sets of causal edge weight matrices of the causal decoupling policy agent before and after the update from parameter snapshots. For the first set of causal edge weights, the absolute value of the change in weight value for each edge before and after the update is calculated. All the absolute values of the changes in weight value are arranged according to their positions in the implicit causal graph structure to generate a first structural sensitivity thermal representation. The numerical values in this thermal representation reflect the sensitivity of each edge in the first sub-network to recent reward signals. For the second set of causal edge weights, not only the magnitude of the change but also the sign of the change is calculated. If the weight adjustment direction of the same edge is opposite in two adjacent updates, the edge is recorded as an oscillating edge. Combining the consistency of the change magnitude and direction, a second structural sensitivity thermal representation is generated.
[0139] Next, the system performs topology-preserving fusion of the first and second structure-sensitivity thermal representations. Topology-preserving fusion means that during weighted combination, the hierarchical order of the implicit causal graph structure itself is preserved. The fused representation generates a graph structure evolution trend representation, which describes the convergence state and learning activity of each part of the network in vector or matrix form.
[0140] Finally, the system adjusts the parameter update step size of the three action parameter generation branches in the action generation module during subsequent updates based on the graph structure evolution trend representation. For example, if the graph structure evolution trend representation shows that the causal path related to the first action parameter generation branch is still undergoing drastic adjustments, the learning rate of that branch can be appropriately reduced to stabilize training; conversely, if the causal path related to a branch has stabilized, its learning rate can be maintained or slightly increased to accelerate convergence.
[0141] In the methods described in the various embodiments of this application, the execution entity for all steps, including record acquisition, representation extraction, factor decomposition, action sequence generation, parameter update instruction issuance, feedback set acquisition and processing, reward signal calculation, and policy gradient search and update, is a speaker sound quality optimization system (system) based on reinforcement learning. The speaker equipment and its internal configuration are the data source for the speaker sound quality optimization system to perceive the external physical acoustic environment state or to receive the sound quality transfer action sequence instructions generated by the speaker sound quality optimization system and adjust its own acoustic response characteristics and electroacoustic conversion behavior accordingly. The speaker sound quality optimization system completes the state perception, parameter configuration, and iterative optimization of the above-mentioned controlled object by calling the internally preset representation extraction function unit, causal decoupling strategy agent, reward calculation function unit, and policy update function unit, without performing any decision-making calculation steps within the above-mentioned controlled object.
[0142] Please see Figure 2 The figure is a schematic diagram of the basic structure of a speaker sound quality optimization system 200 based on reinforcement learning provided in an embodiment of this application. The speaker sound quality optimization system 200 based on reinforcement learning includes: a processor 201; a storage device 202 on which a computer program 2020 is stored; and a network interface 203 for providing network communication functions. When the computer program 2020 is executed by the processor 201, the processor 201 implements any of the speaker sound quality optimization methods based on reinforcement learning.
[0143] Please see Figure 3 This application provides a functional block diagram of a speaker sound quality optimization device, which includes: The audio recording acquisition module is used to acquire audio bitstream playback records and sound field perception records of audio equipment in the current physical acoustic environment; The state representation extraction module is used to perform nonlinear state representation extraction processing on the audio bitstream playback record to obtain a nonlinear auditory state representation, and to perform sound field topology analysis processing on the sound field perception record to obtain a sound field topology state representation. The sound quality transfer processing module is used to decompose the joint state input into acoustic environment inherent attribute factor representation and speaker nonlinear response factor representation through the implicit causal graph structure of the preset causal decoupling strategy agent, using the nonlinear auditory state representation and the sound field topology state representation as joint state input, and generate a sound quality transfer action sequence based on the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation. The migration action execution module is used to execute the sound quality migration action sequence to update the operating parameters of the adjustable processing node in the audio signal processing link, and to perform playback processing on the subsequent audio bitstream according to the updated audio signal processing link and collect the listening perception abstract primitive feedback set. The policy gradient update module is used to generate a reward signal using the auditory abstract primitive feedback set, and store the joint state input, the sound quality transfer action sequence, the reward signal and the updated joint state input as experience transfer tuples in the counterfactual auditory trajectory playback buffer, so that the causal decoupling policy agent can perform counterfactual policy gradient search and update based on the implicit causal graph structure.
[0144] Based on the same inventive concept, an audio device is also provided, which is communicatively connected to a reinforcement learning-based speaker sound quality optimization system. The reinforcement learning-based speaker sound quality optimization system is used to interact with the audio device to implement any of the reinforcement learning-based speaker sound quality optimization methods described above.
[0145] Based on the above, a readable storage medium is provided, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the above method are implemented.
[0146] Please refer to the following: Figure 4This application embodiment acquires audio bitstream playback records and sound field perception records of audio equipment in the current physical acoustic environment, and performs nonlinear state representation extraction processing and sound field topology analysis processing respectively, constructing a dual state representation system that takes into account both the nonlinear characteristics of human auditory perception and the spatial propagation properties of the physical sound field. Based on this, using the nonlinear auditory state representation and the sound field topology state representation as joint state input, the implicit causal graph structure of the intelligent agent using the pre-set causal decoupling strategy separates and decouples the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation mixed in the joint state input at the causal level. This allows the subsequently generated sound quality transfer action sequence to independently respond to changes in the environmental acoustic boundary and changes in the speaker's own distortion characteristics, avoiding interference from the mutual confusion of environmental and equipment factors on sound quality adjustment decisions. By executing a sequence of sound quality transfer actions, the operating parameters of adjustable processing nodes in the audio signal processing chain are adaptively updated. A reward signal is generated by collecting a set of auditory abstract primitive feedback that integrates objective sound field measurement data and subjective auditory abstract primitive scores. Then, the experience transfer tuple containing the joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input is stored in the counterfactual auditory trajectory playback buffer. This supports the agent of the causal decoupling strategy to carry out counterfactual policy gradient search and update based on the implicit causal graph structure. This enables the agent to infer the potential auditory reward differences of different action sequences under the same environmental conditions from limited interaction experience, thereby improving the efficiency and generalization robustness of the sound quality autonomous optimization of audio equipment in unknown and complex acoustic environments.
[0147] Furthermore, it should be noted that this application also provides a computer program product, which may include a computer program that can be stored in a computer-readable storage medium. The processor of the reinforcement learning-based speaker sound quality optimization system reads the computer program from the computer-readable storage medium, and the processor can execute the computer program, causing the reinforcement learning-based speaker sound quality optimization system to perform the aforementioned... Figure 1 The methods described in the corresponding embodiments are already known, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program product embodiments related to this application, please refer to the description of the method embodiments of this application.
[0148] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
Claims
1. A reinforcement learning-based sound quality optimization method for a sound bar, characterized in that, The method includes: Acquire audio bitstream playback records and sound field perception records of audio equipment in the current physical acoustic environment; The audio bitstream playback record is subjected to nonlinear state characterization extraction processing to obtain a nonlinear auditory state characterization, and the sound field perception record is subjected to sound field topology analysis processing to obtain a sound field topology state characterization. Using the nonlinear auditory state representation and the sound field topology state representation as joint state input, the joint state input is decomposed into acoustic environment inherent attribute factor representation and speaker nonlinear response factor representation through the implicit causal graph structure of the preset causal decoupling strategy agent, and a sound quality transfer action sequence is generated based on the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation. The audio quality transfer action sequence is executed to update the operating parameters of the adjustable processing nodes in the audio signal processing link. The subsequent audio bitstream is played according to the updated audio signal processing link, and a set of auditory abstract primitive feedback is collected. A reward signal is generated using the aforementioned auditory abstract primitive feedback set. The joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input are stored as experience transfer tuples in the counterfactual auditory trajectory playback buffer, so that the causal decoupling strategy agent can perform counterfactual policy gradient search and update based on the implicit causal graph structure.
2. The method according to claim 1, characterized in that, The nonlinear state representation extraction process for the audio bitstream playback record to obtain a nonlinear auditory state representation includes: Extract the modified cosine transform domain coefficient sequence from the audio bitstream playback record, and rearrange it according to the human ear critical frequency band scale to generate the critical frequency band energy sequence. The loudness perception compression processing is performed on the energy value of each frequency band in the critical frequency band energy sequence to obtain the auditory excitation spectrum characterization; The transient identifier sequence of each audio frame is extracted from the audio bitstream playback record. The frequency of transient identifier occurrence is counted within the sliding time window to generate an envelope mutation density curve. Peak preservation smoothing is performed on the envelope mutation density curve. Local peak regions exceeding the perception threshold are extracted. The start time position, peak amplitude and duration of each local peak region are recorded to obtain a set of mutation events. Each event in the set of mutation events is mapped to a multidimensional feature vector containing a normalized start time parameter, a normalized peak amplitude parameter, and a normalized duration parameter. The temporal mutation characterization is determined by combining all the multidimensional feature vectors. The auditory excitation spectrum is represented as a first multidimensional vector set, and the temporal abrupt change is represented as a second multidimensional vector set. All multidimensional feature vectors in the first and second multidimensional vector sets are aligned and concatenated in the time dimension to generate a joint multidimensional vector as the nonlinear auditory state representation.
3. The method according to claim 1, characterized in that, The step of performing sound field topology analysis on the sound field sensing record to obtain a sound field topology state representation includes: The sound pressure level gradient change sequence is extracted from the sound field sensing record. The sound pressure level gradient change sequence is composed of the sound pressure level values collected by multiple sound pressure sensing units in the sound field topology sensing unit array at the same time, arranged according to their spatial positions. The sound pressure level gradient change sequence is subjected to spatial gradient field decomposition. Based on the Helmholtz decomposition principle, the sound pressure level gradient change sequence is separated into an irrotational gradient field part and a divergence-free gradient field part. The irrotational gradient field part corresponds to the direct sound propagation path contribution of the sound wave from the sound source to each sound pressure sensing unit, and the divergence-free gradient field part corresponds to the reflected sound propagation path contribution of the sound wave after boundary reflection to each sound pressure sensing unit. The irrotational gradient field is converted into a first spatial distribution matrix composed of the sound pressure amplitude and phase at the location of each sound pressure sensing unit and used as a representation of the direct sound field component. The divergence-free gradient field is converted into a second spatial distribution matrix composed of the sound pressure amplitude and phase at the location of each sound pressure sensing unit and used as a representation of the reflected sound field component. The reverberation attenuation envelope sequence is extracted from the sound field sensing record. The reverberation attenuation envelope sequence is composed of the envelope curve of the sound pressure level attenuating over time recorded by each sound pressure sensing unit in the sound field topology sensing unit array. The initial attenuation segment separation process is performed on each envelope curve in the reverberation attenuation envelope sequence. An initial segment from the arrival time of the direct sound to the preset early reflection time boundary is extracted from each envelope curve. The arrival direction estimation process is performed on the initial segment to generate the arrival direction angle parameter and relative energy parameter corresponding to each early reflection sound event. The arrival direction angle parameter and relative energy parameter corresponding to all early reflection sound events constitute the early reflection sound direction distribution characterization. Late decay segment analysis is performed on each envelope curve in the reverberation decay envelope sequence. The portion of each envelope curve after the preset early reflection time boundary is extracted. The late decay segment of each sound pressure sensing unit is spatially averaged to obtain the average late decay curve. The late reverberation decay rate parameter is extracted based on the logarithmic slope of the average late decay curve. The late reverberation decay rate parameter is used as a characterization of the energy decay of the late diffusion field. The direct sound field component representation, the reflected sound field component representation, the early reflected sound direction distribution representation, and the late diffuse field energy attenuation representation are spliced together in a preset combination order to generate a composite multidimensional array containing spatial sound field structure and time reverberation attenuation attributes as the sound field topology state representation.
4. The method according to any one of claims 1-3, characterized in that, The method uses the nonlinear auditory state representation and the sound field topological state representation as joint state inputs. Through a preset causal decoupling strategy, the intelligent agent decomposes the joint state inputs into acoustic environment intrinsic attribute factor representations and speaker nonlinear response factor representations using an implicit causal graph structure. Based on these acoustic environment intrinsic attribute factor representations and speaker nonlinear response factor representations, a sound quality transfer action sequence is generated, including: Read the nonlinear auditory state representation and the sound field topology state representation from the state representation temporary buffer, and perform a head-to-tail concatenation operation on the vector dimension of the nonlinear auditory state representation and the sound field topology state representation to generate a joint state input vector. The joint state input vector is input to the implicit causal graph structure input layer of the causal decoupling strategy agent. The implicit causal graph structure consists of multiple causal nodes and directed causal edges between nodes. Each causal node corresponds to a potential physical factor that affects the sound quality. In the first sub-network of the implicit causal graph structure, the joint state input vector is mapped to the acoustic environment intrinsic attribute factor node through the first set of causal edge weights, and the acoustic environment intrinsic attribute factor node outputs the acoustic environment intrinsic attribute factor representation. In the second sub-network of the implicit causal graph structure, the joint state input vector is mapped to the speaker nonlinear response factor node through the second set of causal edge weights. During the mapping process, a conditional independence constraint is introduced from the acoustic environment intrinsic attribute factor node, so that the speaker nonlinear response factor node outputs a speaker nonlinear response factor representation that is decoupled from the acoustic environment intrinsic attribute factor representation. The acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation are input to the action generation module of the causal decoupling strategy agent. The action generation module includes action parameter generation branches that correspond one-to-one with the adjustable processing nodes in the audio signal processing link. The action generation module generates branch generation frequency band energy redistribution actions, gain curve shape adjustment actions, and sound image trajectory reshaping actions using different action parameters, and determines the sound quality transfer action sequence.
5. The method according to claim 4, characterized in that, The process of generating branch frequency band energy redistribution actions, gain curve shape adjustment actions, and sound image trajectory reshaping actions using different action parameters in the action generation module, and determining the sound quality transfer action sequence, includes: In the first action parameter generation branch corresponding to the frequency division network node, the frequency band energy redistribution action is generated based on the reflected sound field direction distribution information in the acoustic environment inherent attribute factor characterization and the frequency response non-flatness characteristics in the speaker nonlinear response factor characterization. In the second action parameter generation branch corresponding to the dynamic gain control node, based on the late reverberation decay rate information in the acoustic environment inherent attribute factor characterization and the large signal compression characteristics in the speaker nonlinear response factor characterization, a gain curve shape adjustment action is generated. In the third action parameter generation branch corresponding to the spatial virtualization node, the sound image trajectory reshaping action is generated based on the early reflected sound direction distribution information in the acoustic environment inherent attribute factor representation and the directivity deviation characteristics in the speaker nonlinear response factor representation. The frequency band energy redistribution action, the gain curve shape adjustment action, and the sound image trajectory reshaping action are merged and packaged according to a preset instruction encapsulation format to form the sound quality transfer action sequence.
6. The method according to claim 1, characterized in that, The process of executing the audio quality transfer action sequence to update the operating parameters of the adjustable processing nodes in the audio signal processing link, and performing playback processing on subsequent audio bitstreams and collecting a set of auditory perception abstract primitive feedback based on the updated audio signal processing link includes: The frequency band energy redistribution action in the sound quality migration action sequence is analyzed, the frequency division network parameter adjustment instruction string contained in the frequency band energy redistribution action is extracted, and the frequency division point register value and slope register value of the frequency division network node in the audio signal processing link are updated according to the frequency division network parameter adjustment instruction string. The gain curve shape adjustment action in the sound quality transfer action sequence is analyzed, the gain curve parameter adjustment instruction string contained in the gain curve shape adjustment action is extracted, and the threshold inflection point register value, compression ratio curve slope register value and start-up release time constant register value of the dynamic gain control node in the audio signal processing link are updated according to the gain curve parameter adjustment instruction string. The sound image trajectory reshaping action in the sound quality transfer action sequence is analyzed, the spatial parameter adjustment instruction string contained in the sound image trajectory reshaping action is extracted, and the virtual sound source orientation register value and spatial reflection intensity register value of the spatial virtualization node in the audio signal processing link are updated according to the spatial parameter adjustment instruction string. After completing the register value update operation of all adjustable processing nodes, it is determined that the audio signal processing link has switched to the updated parameter configuration state and receives the subsequent audio bit stream from the audio source device. The subsequent audio bit stream is input into the audio decoder for decoding processing to obtain a pulse code modulation format digital audio data stream. The pulse code modulation format digital audio data stream is sequentially fed into the frequency division network node, dynamic gain control node and spatial virtualization node in the audio signal processing link in the updated parameter configuration state for signal processing. The processed digital audio data stream is acquired at the output end of the audio signal processing link, and the processed digital audio data stream is converted into an analog audio signal to drive the speaker unit of the audio equipment to reproduce sound, forming an updated sound field in the current physical acoustic environment. In the updated sound field, the updated sound pressure level gradient change sequence and the updated reverberation attenuation envelope sequence are collected by the sound field topology sensing unit array. The system receives a set of subjective listening abstract primitive scores input by the user through the listening evaluation interface. The set of subjective listening abstract primitive scores includes timbre transparency abstract primitive scores, sound image localization clarity abstract primitive scores, and listening comfort abstract primitive scores. The updated sound pressure level gradient change sequence, the updated reverberation attenuation envelope sequence, and the subjective listening perception abstract primitive score set are combined to form the listening perception abstract primitive feedback set.
7. The method according to claim 1 or 6, characterized in that, The process of generating a reward signal using the auditory abstract primitive feedback set, and storing the joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input as experience transfer tuples in the counterfactual auditory trajectory playback buffer, allows the causal decoupling strategy agent to perform counterfactual policy gradient search and update based on the implicit causal graph structure, including: The updated sound pressure level gradient change sequence is extracted from the auditory abstract primitive feedback set, and the spatial gradient smoothness index of the updated sound pressure level gradient change sequence is calculated. The spatial gradient smoothness index is determined by the variance statistic of the sound pressure level difference between each adjacent spatial sampling point in the updated sound pressure level gradient change sequence. The updated reverberation attenuation envelope sequence is extracted from the auditory abstract primitive feedback set, and the ratio of early reflection energy to late diffusion field energy of the updated reverberation attenuation envelope sequence is calculated. The ratio of early reflection energy to late diffusion field energy is determined by the ratio of the integral energy within the time window of the early reflection to the integral energy within the time window of the late diffusion field. Extract the timbre transparency abstract primitive score, sound image localization clarity abstract primitive score, and listening comfort abstract primitive score from the subjective listening abstract primitive score set from the listening abstract primitive feedback set. Then, perform a linear weighted summation of the timbre transparency abstract primitive score, the sound image localization clarity abstract primitive score, and the listening comfort abstract primitive score with preset weights to generate a comprehensive subjective listening score. Based on the spatial gradient smoothness index, the ratio of early reflected sound to late diffused field energy index, and the preset objective index weighting coefficient, an objective acoustic quality comprehensive score is generated. The subjective listening perception score and the objective acoustic quality score are input into a pre-trained nonlinear fusion mapping network. The multilayer perceptron structure of the nonlinear fusion mapping network is used to perform feature interaction and nonlinear transformation processing on the subjective listening perception score and the objective acoustic quality score, and output the overall listening perception improvement. The historical comprehensive listening value under the previous parameter configuration state is read from the historical data storage area. The historical comprehensive listening value is calculated using the same nonlinear fusion mapping network after the previous sound quality migration action sequence was executed. The improvement in overall listening quality is subtracted from the historical overall listening quality value to obtain the change in listening quality. A piecewise linear reward shaping function is then applied to the change in listening quality to map positive changes to a positive reward interval and negative changes to a negative reward interval, thus obtaining the reward signal. The joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input are encapsulated as an experience transfer tuple to obtain a data structure as the experience transfer tuple, which is then stored in the counterfactual auditory trajectory playback buffer. The counterfactual policy gradient search is then performed in conjunction with the monitoring of the numerical change trend of the loss function of the causal decoupling strategy agent.
8. The method according to claim 1, characterized in that, The step of encapsulating the joint state input, the sound quality transfer action sequence, the reward signal, and the updated joint state input as an experience transition tuple to obtain a data structure as the experience transition tuple and storing it in the counterfactual auditory trajectory playback buffer, and combining this with monitoring the trend of the loss function value change of the causal decoupling strategy agent to perform counterfactual policy gradient search and update, includes: The joint state input is marked as the state before the action is executed, the sound quality transfer action sequence is marked as the action to be executed, the reward signal is marked as the immediate reward, and the updated joint state input regenerated based on the auditory abstract primitive feedback set is marked as the state after the action is executed. The state before the action, the action to be executed, the immediate reward, and the state after the action are encapsulated into a data structure as the experience transfer tuple and stored at the end of the storage queue of the counterfactual auditory trajectory playback buffer. When the number of experience transfer tuples stored in the counterfactual auditory trajectory playback buffer reaches the preset batch sampling threshold, a batch of experience transfer tuples is randomly selected from the counterfactual auditory trajectory playback buffer. The gradient of the loss function of the causal decoupling strategy agent is calculated using the selected experience transfer tuples, and the first set of causal edge weights, the second set of causal edge weights, and the weight parameters of each action parameter generation branch inside the causal decoupling strategy agent are updated through the gradient descent optimization algorithm. The system monitors the trend of the loss function value of the causal decoupling strategy agent in multiple consecutive iteration update cycles. When the fluctuation amplitude of the loss function value trend is less than the preset convergence judgment threshold, it is determined that the preset update stop condition is met, and the current snapshot of the causal decoupling strategy agent parameters is saved for the real-time sound quality optimization process of the audio equipment.
9. A speaker sound quality optimization device based on reinforcement learning, characterized in that, The device includes: The audio recording acquisition module is used to acquire audio bitstream playback records and sound field perception records of audio equipment in the current physical acoustic environment; The state representation extraction module is used to perform nonlinear state representation extraction processing on the audio bitstream playback record to obtain a nonlinear auditory state representation, and to perform sound field topology analysis processing on the sound field perception record to obtain a sound field topology state representation. The sound quality transfer processing module is used to decompose the joint state input into acoustic environment inherent attribute factor representation and speaker nonlinear response factor representation through the implicit causal graph structure of the preset causal decoupling strategy agent, using the nonlinear auditory state representation and the sound field topology state representation as joint state input, and generate a sound quality transfer action sequence based on the acoustic environment inherent attribute factor representation and the speaker nonlinear response factor representation. The migration action execution module is used to execute the sound quality migration action sequence to update the operating parameters of the adjustable processing node in the audio signal processing link, and to perform playback processing on the subsequent audio bitstream according to the updated audio signal processing link and collect the listening perception abstract primitive feedback set. The policy gradient update module is used to generate a reward signal using the auditory abstract primitive feedback set, and store the joint state input, the sound quality transfer action sequence, the reward signal and the updated joint state input as experience transfer tuples in the counterfactual auditory trajectory playback buffer, so that the causal decoupling policy agent can perform counterfactual policy gradient search and update based on the implicit causal graph structure.
10. An audio device, characterized in that, The audio device is communicatively connected to a reinforcement learning-based speaker sound quality optimization system, which interacts with the audio device to implement the method described in any one of claims 1-8.