Intelligent energy-saving sound equipment playing control method and system

By preprocessing the original audio signal and compressing it into a low-dimensional semantic vector, combined with a neural sound field model and a distributed clock synchronization protocol, the delay time of the speaker array is dynamically adjusted, solving the problem of the lack of adaptive ability of smart audio equipment in dynamic environments, and realizing intelligent and efficient energy-saving audio playback.

CN120692504AInactive Publication Date: 2025-09-23GANZHOU DEHUIDA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510900862.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing technology, smart audio devices lack an audio playback control method based on the coordinated optimization of semantic compression and energy efficiency feedback in a dynamic acoustic environment. In the existing technology, smart audio devices lack an audio playback control method based on the coordinated optimization of semantic compression and energy efficiency feedback in a dynamic acoustic environment. In the existing technology, smart audio devices lack the ability to adapt to dynamic changes in the environment, especially how to dynamically adjust playback parameters to reduce energy consumption without reducing sound quality.

Method used

By preprocessing the original audio input signal, using a variational autoencoder to perform low-dimensional semantic vector compression, and combining the sound wave reflection time difference feedback from the terminal device, a neural sound field model is generated, the phase distribution of the speaker array is calculated, and the delay time is coordinated through a distributed clock synchronization protocol. The energy efficiency data of the speaker drive circuit and the ambient noise spectrum characteristics are dynamically adjusted to generate a dynamic control parameter set.

Benefits of technology

It achieves the goal of improving data processing efficiency and system intelligent perception capabilities while ensuring the quality of audio transmission, effectively supporting the intelligence and efficiency of energy-saving audio playback control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692504A_ABST
    Figure CN120692504A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent energy-saving sound equipment playing control method and system, and relates to the technical field of playing control, and the method comprises the steps: carrying out the preprocessing of an original audio input signal, inputting the signal into a variational auto-encoder, compressing the signal into a low-dimensional semantic vector, and according to the sound wave reflection time difference fed back by terminal equipment, obtaining a low-dimensional semantic vector; obtaining the distance between the terminal equipment and the audio signal source; according to the low-dimensional semantic vector, an audio waveform is reconstructed through a generative adversarial network, an initial neural sound field model is constructed, real-time correction is carried out through a physical sound wave propagation equation, and a neural sound field model is generated; based on the neural sound field model, calculating the phase distribution of the multi-loudspeaker array, and coordinating the delay time of the loudspeaker array through a distributed clock synchronization protocol to obtain the sound field focusing precision. An original audio input signal is subjected to multi-level preprocessing, and a processed high-dimensional audio signal is input into a variational auto-encoder for compression of a low-dimensional semantic vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of playback control, and in particular to an intelligent energy-saving audio playback control method and system thereof. Background Art

[0002] With the deep integration of multidisciplinary technologies such as artificial intelligence, the Internet of Things, and acoustic modeling, smart audio devices have gradually evolved from traditional sound playback terminals into comprehensive platforms capable of semantic recognition, spatial perception, and interactive decision-making. In audio processing, the application of deep neural networks, particularly in compression and reconstruction tasks, has significantly improved the transmission and storage efficiency of audio data. Simultaneously, the development of neural rendering and 3D modeling technologies has also driven the initial exploration of Neural Radiant Fields (NeRF) in sound field modeling, enabling a more refined representation of spatial information during sound wave propagation.

[0003] In the area of ​​spatial sound control, focusing the three-dimensional sound field with distributed speaker arrays has become an important means of enhancing the immersive audio experience. However, due to the complex spatial interference and ambient noise in the actual sound field environment, the real-time and stability of sound field modeling still face technical challenges. In addition, how to dynamically adjust playback parameters to reduce energy consumption without reducing sound quality has gradually become a core issue in intelligent audio control. Current technical solutions mostly use static modeling or empirical parameter adjustment to control playback, which lacks the ability to adapt to dynamic changes in the environment. There is still much room for improvement in strategy updates and playback energy efficiency management, especially in the linkage mechanism between integrated spatial modeling and energy consumption control. A unified and efficient technical path has not yet been formed. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an intelligent energy-saving audio playback control method to solve the problem of lack of audio playback control based on collaborative optimization of semantic compression and energy efficiency feedback in dynamic acoustic environments.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides an intelligent energy-saving audio playback control method, which includes: The original audio input signal is preprocessed and input into a variational autoencoder to compress it into a low-dimensional semantic vector. At the same time, the distance between the terminal device and the audio signal source is obtained based on the sound wave reflection time difference fed back by the terminal device. Based on the low-dimensional semantic vector, the audio waveform is reconstructed through a generative adversarial network to build an initial neural sound field model. This model is then corrected in real time using the physical sound wave propagation equation to generate a neural sound field model. Based on the neural sound field model, the phase distribution of the multi-speaker array is calculated, and the delay time of the speaker array is coordinated through the distributed clock synchronization protocol to obtain the sound field focusing accuracy; According to the sound field focusing accuracy, the variational autoencoder is dynamically adjusted through the policy gradient algorithm, and the energy efficiency data of the speaker driving circuit and the ambient noise spectrum characteristics are multi-dimensionally fused through the multimodal feature fusion method to generate a dynamic control parameter set.

[0007] As a preferred solution of the intelligent energy-saving audio playback control method of the present invention, wherein: the input is compressed into a low-dimensional semantic vector into a variational autoencoder, the steps are as follows: Input the preprocessed audio input signal into the variational autoencoder to obtain the mean and variance of the latent space; The variational autoencoder is trained using historical audio data, and the mean and variance of the latent space are input into the trained variational autoencoder for reparameter sampling to generate a low-dimensional semantic vector.

[0008] As a preferred solution of the intelligent energy-saving audio playback control method described in the present invention, the reconstructing the audio waveform through the generative adversarial network refers to inputting the low-dimensional semantic vector into the generative adversarial network, generating an audio waveform and comparing it with the real audio waveform, optimizing the generated audio waveform through the loss function, and generating a reconstructed audio waveform.

[0009] As a preferred solution of the intelligent energy-saving audio playback control method of the present invention, the steps of constructing the initial neural sound field model are as follows: Collect audio signal sample data with spatial coordinate annotations, and use multimodal data fusion and training methods to construct neural radiation fields; Short-time Fourier transform is used to divide the reconstructed audio waveform into short-time frame signals weighted by window functions. The short-time frame signals weighted by window functions and 3D spatial coordinates are integrated into the neural radiation field through the multimodal sound field implicit method to generate an initial neural sound field model.

[0010] As a preferred solution of the intelligent energy-saving audio playback control method of the present invention, wherein: the real-time correction is performed through the physical sound wave propagation equation to generate a neural sound field model, the steps are as follows: Substitute the distance between the terminal device and the audio source into the free space attenuation formula, extract the attenuation coefficient based on the inverse square law, and substitute the acoustic impedance values ​​of the incident medium and the reflecting medium in the neural sound field model into the impedance ratio formula to determine the reflection parameter through the acoustic impedance ratio of the incident medium to the reflecting medium. According to the attenuation coefficient and the reflection parameter, the propagation characteristics in the initial neural sound field model are adjusted to obtain the neural sound field model.

[0011] As a preferred solution of the intelligent energy-saving audio playback control method of the present invention, wherein: the delay time of the speaker array is coordinated by the distributed clock synchronization protocol to obtain the sound field focusing accuracy, the steps are as follows: The three-dimensional spatial position of each speaker unit in the neural sound field model is extracted through the back propagation positioning method, the propagation distance between the speaker and the sound output unit is calculated, and the propagation delay compensation method is used to identify the compensation delay time of each speaker with the speaker closest to the sound output unit as a reference; The clock deviation of each speaker node is obtained by a two-way delay measurement method, and then the local control timestamp of each speaker is obtained from the clock deviation and the compensation delay time by a synchronization timestamp derivation method; The reconstructed audio waveform is delayed according to the local control timestamp of each speaker, the sound field energy distribution data is collected, and the sound field focusing accuracy is obtained through the sound energy density analysis method.

[0012] As a preferred solution of the intelligent energy-saving audio playback control method of the present invention, wherein: the steps of generating a dynamic control parameter set are as follows: Based on the sound field focusing accuracy, the compression rate of the variational autoencoder is dynamically corrected through the policy gradient algorithm; Collect energy efficiency data of the speaker drive circuit and environmental noise spectrum characteristics; Through the multimodal feature fusion method, the energy efficiency data of the speaker driving circuit and the spectrum characteristics of the ambient noise are fused to generate a multi-dimensional fusion feature vector; The multi-dimensional fusion feature vector is encoded by the modified variational autoencoder to generate a set of dynamic control parameters.

[0013] In a second aspect, the present invention provides an intelligent energy-saving audio playback control system, comprising: The semantic compression module preprocesses the original audio input signal and inputs it into the variational autoencoder to compress it into a low-dimensional semantic vector. At the same time, it obtains the distance between the terminal device and the audio signal source based on the sound wave reflection time difference fed back by the terminal device; The sound field modeling module reconstructs the audio waveform based on the low-dimensional semantic vector through a generative adversarial network to build an initial neural sound field model. It then makes real-time corrections using the physical sound wave propagation equation to generate a neural sound field model. The focus control module calculates the phase distribution of the multi-speaker array based on the neural sound field model and coordinates the delay time of the speaker array through the distributed clock synchronization protocol to obtain the sound field focusing accuracy; The dynamic control module dynamically adjusts the variational autoencoder through the policy gradient algorithm according to the sound field focusing accuracy, and uses the multimodal feature fusion method to perform multi-dimensional feature fusion on the energy efficiency data of the speaker drive circuit and the ambient noise spectrum characteristics to generate a set of dynamic control parameters.

[0014] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the intelligent energy-saving audio playback control method as described in the first aspect of the present invention is implemented.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the intelligent energy-saving audio playback control method as described in the first aspect of the present invention is implemented.

[0016] The beneficial effects of the present invention are as follows: by performing multi-level preprocessing on the original audio input signal and inputting the processed high-dimensional audio signal into a variational autoencoder for low-dimensional semantic vector compression, and combining the sound wave reflection time difference feedback from the terminal device to accurately calculate the physical distance between the device and the audio signal source, efficient semantic compression of audio data and dynamic perception of the spatial environment are achieved. At the same time, spatial position information is integrated into the compression process, providing a reliable foundation for the adaptive selection of subsequent wireless communication protocols and optimization of the audio playback environment. Ultimately, the goal is to improve data processing efficiency and system intelligent perception capabilities while ensuring audio transmission quality, effectively supporting the intelligent and efficient control of energy-saving audio playback. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 The figure is a flow chart of an intelligent energy-saving audio playback control method.

[0019] Figure 2 This is a flow chart of the dynamic control module in the intelligent energy-saving audio playback control method.

[0020] Figure 3 This is a flowchart of the sound field modeling module in the intelligent energy-saving audio playback control method.

[0021] Figure 4 This is a flow chart of the focus control module in the intelligent energy-saving audio playback control method. DETAILED DESCRIPTION

[0022] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0025] Reference Figures 1 to 4 , is an embodiment of the present invention, which provides an intelligent energy-saving audio playback control method, comprising the following steps: S1: The original audio input signal is preprocessed and input into the variational autoencoder to be compressed into a low-dimensional semantic vector. At the same time, the distance between the terminal device and the audio signal source is obtained based on the sound wave reflection time difference fed back by the terminal device.

[0026] Specifically, the following steps are included: S1.1: Feed the preprocessed audio input signal into the variational autoencoder to obtain the mean and variance of the latent space.

[0027] Specifically, the preprocessed audio input signal is input into the encoder part of the variational autoencoder to obtain the mean vector and variance logarithm vector of the latent space corresponding to the Gaussian distribution parameters. The encoder consists of multiple stacked convolutional layers and fully connected layers. The convolutional layer extracts time-frequency features layer by layer through local perception and weight sharing, and the fully connected layer maps the features to the statistical parameters of the latent space.

[0028] For example, when the latent space dimension is 64, the mean vector and the variance log vector must both be 64-dimensional to ensure a one-to-one correspondence between the mean and variance log at each latent dimension position.

[0029] S1.2: Use historical audio data to train the variational autoencoder, and input the mean and variance of the latent space into the trained variational autoencoder for reparameter sampling to generate a low-dimensional semantic vector.

[0030] Specifically, historical audio data is input into the variational autoencoder for training. During the training process, parameter learning is performed through the encoder and decoder structures. The training goal is to simultaneously minimize the reconstruction error between the input audio data and the reconstruction result, as well as the KL divergence between the latent space distribution and the standard normal distribution. After the training is completed, the mean vector and variance vector in the pre-obtained latent space are input into the trained variational autoencoder, and the re-parameter sampling method is used to generate a low-dimensional semantic vector.

[0031] S2: Based on the low-dimensional semantic vector, the audio waveform is reconstructed through a generative adversarial network to build an initial neural sound field model, and the physical sound wave propagation equation is used to perform real-time corrections to generate a neural sound field model.

[0032] Specifically, the following steps are included: S2.1: Input the low-dimensional semantic vector into the generative adversarial network to generate an audio waveform and compare it with the real audio waveform. The generated audio waveform is optimized through the loss function to generate the reconstructed audio waveform.

[0033] Specifically, the low-dimensional semantic vector transmitted to the terminal device is input into the generative adversarial network, and an audio waveform corresponding to the low-dimensional semantic vector is generated through the conditional generative modeling method. The audio waveform corresponding to the low-dimensional semantic vector is compared point by point with the real audio collected by a high-quality microphone and preprocessed. The discriminator in the generative adversarial network is used to distinguish the audio waveform corresponding to the low-dimensional semantic vector from the real audio, and a discrimination result of the audio waveform and the real audio waveform is generated. Subsequently, the perceptual loss method is used to construct a loss function and update the generator parameters to finally generate a reconstructed audio waveform.

[0034] The perceptual loss method is used to construct the loss function and update the generator parameters. It is based on the pre-trained deep neural network obtained by perceptual feature comparison training, extracts the high-level features of the generated audio waveform and the real audio waveform, and calculates the difference between the feature representations of the two as the perceptual loss value. Subsequently, the perceptual loss value is used as the optimization target, and the gradient of the loss function with respect to the generator parameters is calculated through the back-propagation algorithm. The generator parameters are updated through the gradient information to improve the quality of the generated audio waveform and finally generate the reconstructed audio waveform.

[0035] S2.2: Collect audio signal sample data with spatial coordinate annotations, and use multimodal data fusion and training methods to construct the neural radiation field.

[0036] Specifically, audio signal sample data with spatial coordinate annotations is collected, and visual information and spatial position information of the corresponding scene are obtained synchronously, and integrated to generate a multimodal dataset containing audio signals, image information and spatial coordinates. During data preprocessing, the spectral features of the audio signal are extracted by short-time Fourier transform, the image information is encoded by convolutional neural network encoding method, and the spatial coordinate data is normalized by numerical normalization method. The multimodal dataset is used as input, and a multimodal data fusion and training method is adopted. By jointly optimizing the correlation between audio features, visual features and spatial information, a neural radiation field that includes spatial perception ability is trained.

[0037] During the training process, the mapping relationship between audio signal sample data and spatial coordinates is modeled through the neural radiation field, so that the neural radiation field has the ability to generate audio features corresponding to spatial positions based on spatial positions.

[0038] The steps for synchronously acquiring visual information and spatial position information for the corresponding scene are as follows: while collecting audio signal sample data, use an acquisition device equipped with a visual sensor and a positioning device (for example, a mobile device with an RGB camera and an IMU, GPS, or lidar) to synchronously record the scene image and spatial coordinates.

[0039] Specifically, for example: using a portable terminal device equipped with a depth camera and an inertial measurement unit in an indoor environment, while the collector records the audio signal, the camera continuously obtains image frames corresponding to the time of the audio sampling point, and the IMU records the spatial position and posture coordinates of the terminal device at that moment.

[0040] The aforementioned joint optimization of the correlation between audio features, visual features, and spatial information refers to the semantic alignment and spatiotemporal consistency of modal data (audio features, visual features, and spatial information) in a common scene. This mainly includes the following three aspects: Semantic consistency relevance refers to whether the audio and visual features semantically represent the same event or entity. For example, if the image information contains a person speaking, the audio should contain the corresponding speech signal, and the "opening mouth" action in the visual feature should correspond to the semantic meaning of "sound" in the audio.

[0041] Spatial position correspondence refers to the relationship between the spatial source of an audio signal and its spatial coordinates. For example, the characteristics of an audio signal received at a certain coordinate position (e.g., x, y, z) should be spatially aligned with the image information appearing at that position in the visual scene. This is then modeled through neural radiation fields to achieve an optimized mapping from spatial coordinates to audio distribution.

[0042] Temporal synchronization correlation refers to the synchronization of audio signals and visual images in the temporal dimension. For example, at a spatial coordinate location (e.g., x, y, z), when the visual information reflects a local environmental feature such as an obstacle, opening structure, or wall material, the corresponding reflection, attenuation, or interference characteristics in the audio signal should also appear within the same time window.

[0043] S2.3: Use short-time Fourier transform to divide the reconstructed audio waveform into short-time frame signals weighted by window functions; use the multimodal sound field implicit method to integrate the short-time frame signals weighted by window functions and 3D spatial coordinates into the neural radiation field to generate an initial neural sound field model.

[0044] Specifically, the reconstructed audio waveform is processed by short-time Fourier transform, the complete audio waveform is divided into multiple continuous short-time frame signals weighted by window functions, and time-frequency features are extracted in each short-time frame signal weighted by window functions; according to the timestamp of each short-time frame signal weighted by window functions, the extracted time-frequency features are aligned with the three-dimensional spatial coordinate data corresponding to the audio time-frequency features, and a pair consisting of an audio time-frequency feature vector and a three-dimensional spatial coordinate vector is constructed, and multiple pairs are arranged in order of events to construct a joint input sample containing time window, frequency distribution and spatial position information; the short-time frame signal weighted by window function and the three-dimensional spatial coordinate data are fused using a multimodal sound field implicit method, a neural network is obtained by constructing a multi-layer fully connected neural network structure, and the neural radiation field is trained using a supervised learning method based on the joint input sample, and according to the mapping relationship between the time window, frequency distribution and spatial position information in the joint input sample, an initial neural sound field model that can express the audio propagation characteristics in space is established.

[0045] S2.4: The physical sound wave propagation equations include the free space attenuation formula and the impedance ratio formula.

[0046] Specifically, the physical sound wave propagation equation describes the energy attenuation of sound waves in space and the reflection behavior of medium boundaries (for example, when sound waves encounter the boundaries of different media such as air and walls, ground or other solid surfaces, part of the sound wave energy is reflected back to the propagation medium, and part of the sound wave energy penetrates the interface and continues to propagate). It characterizes the spatial physical mechanism of audio signal propagation; the free space attenuation formula is used to express the decreasing relationship between the energy density of sound waves as the propagation distance increases, reflecting the energy loss process in the propagation path; the impedance ratio formula is used to describe the reflection and transmission behavior of sound waves at the boundaries of different media, and calculates the reflection coefficient and transmission coefficient based on the ratio of the acoustic impedance between the two media, thereby quantifying the waveform changes and energy distribution caused by medium changes during sound wave propagation.

[0047] S2.5: Substitute the distance between the terminal device and the audio source into the free space attenuation formula, extract the attenuation coefficient based on the inverse square law, and substitute the acoustic impedance values ​​of the incident medium and the reflecting medium in the neural sound field model into the impedance ratio formula. The reflection parameters are determined by the acoustic impedance ratio of the incident medium to the reflecting medium.

[0048] Specifically, according to the inverse square law, the distance between the terminal device and the audio source Substituting into the free space attenuation formula, we can get the sound pressure level corresponding to the actual space distance. , and by and The difference reflects the attenuation coefficient of the distance attenuation effect. For example: =80dB, =1m, =4m, the attenuation coefficient is 12.04dB.

[0049] The free space attenuation formula is as follows, ; in, Indicates the distance between the terminal device and the audio source is The sound pressure level, Indicates reference distance The sound pressure level at Indicates the distance between the terminal device and the audio source. Indicates the reference distance.

[0050] The reference distance is a standard distance used to define the starting point of sound pressure in free-space acoustic measurements. Furthermore, it must be set in the far field of the sound source to avoid near-field effects on the sound pressure level. For example, standards such as IEC60268 and ISO 9613-1 use a default reference distance of 1 meter when testing sound source output in free-field conditions.

[0051] The acoustic impedance values ​​of the incident medium and the reflecting medium in the neural sound field model are substituted into the impedance ratio formula as input parameters. Based on the ratio of the acoustic impedances of the two media, the reflection coefficient when the sound wave is reflected at the medium boundary is calculated.

[0052] The impedance ratio formula is as follows, ; in, Expressed as the reflection coefficient, is expressed as the acoustic impedance of the incident medium, Expressed as the acoustic impedance of the reflecting medium.

[0053] The reflection coefficient is the absolute value of the energy reflection ratio. The theoretical value range of the reflection coefficient is [0,1], where =0 means full transmission, no reflection, =1 indicates complete reflection and no transmission.

[0054] The sound pressure sensor and the electric velocity sensor are used to jointly collect the sound pressure data and particle vibration velocity of the sound wave. The sound pressure data of the sound wave is multiplied by the particle vibration velocity to obtain the incident sound energy intensity. After that, the incident sound energy intensity is multiplied by the reflection coefficient to obtain the reflected sound energy intensity.

[0055] S2.6: Adjust the propagation characteristics in the initial neural sound field model according to the attenuation coefficient and the reflection parameter to obtain the neural sound field model.

[0056] Specifically, the attenuation coefficient extracted from the inverse square law is used as the energy loss parameter in the propagation process. At the same time, the reflection coefficient calculated based on the acoustic impedance ratio formula is introduced into the neural sound field model to adjust the reflection intensity of the sound wave at the boundary of the medium. By comprehensively applying the attenuation coefficient and the reflection parameter, the propagation path and energy distribution of the sound wave are corrected, and the neural sound field model's ability to express sound wave transmission in the actual sound field environment is improved. Finally, a neural sound field model that takes into account spatial attenuation and interface reflection effects is obtained.

[0057] For example, the attenuation coefficient can be expressed as -10dB, and the reflection coefficient ranges from 0 to 1, both of which effectively adjust the spatial distribution of acoustic energy in the model.

[0058] S3: Based on the neural sound field model, the phase distribution of the multi-speaker array is calculated, and the delay time of the speaker array is coordinated through the distributed clock synchronization protocol to obtain the sound field focusing accuracy.

[0059] Specifically, the following steps are included: S3.1: Extract the three-dimensional spatial position of each speaker unit in the neural sound field model through the backpropagation localization method, calculate the propagation distance between the speaker and the sound output unit, and use the propagation delay compensation method to identify the compensation delay time of each speaker with the speaker closest to the sound output unit as a reference.

[0060] Specifically, the three-dimensional spatial position of each speaker unit in the neural sound field model is extracted through the back propagation positioning method, and the position of each speaker unit is spatially paired with the position of the sound output unit. The Euclidean distance formula between two points in three-dimensional space is used to calculate the propagation path length between each speaker unit and the sound output unit, and the actual propagation distance between the speaker unit and the sound output unit in the physical space is obtained. The speed of sound conversion method is used to convert the actual propagation distance between the speaker unit and the sound output unit in the physical space into a propagation delay corresponding to the actual propagation distance. Subsequently, a propagation delay compensation method is adopted, with the speaker unit closest to the sound output unit as a reference benchmark, and the propagation path delay difference of each speaker unit relative to the reference speaker unit is identified, thereby determining the compensation delay time of each speaker unit.

[0061] For example, if the reference speaker is 0.5 meters away from the sound output unit, and the distance to the other speaker is 0.7 meters, and the speed of sound is taken as an example value of 340 meters per second, then the propagation delay difference between the two is approximately (0.7-0.5) / 340≈0.00059 seconds, and the calculated compensation delay time is 0.59 milliseconds.

[0062] The expression for compensating the delay time is as follows: ; in, Indicates the Compensation delay time of each speaker unit, Indicates the The propagation path length of each speaker unit, Indicates the shortest propagation path length among all speaker units, It represents the speed of sound waves in the propagation medium.

[0063] S3.2: Obtain the clock deviation of each speaker node through a two-way delay measurement method, and then obtain the local control timestamp of each speaker from the clock deviation and the compensated delay time through a synchronization timestamp derivation method.

[0064] Specifically, the time information of sending and receiving control signals between the sound output unit and each speaker node is recorded respectively through the two-way delay measurement method, and the propagation time of the control signal in the round trip process is calculated. The clock deviation of each speaker node is deduced from this. Then, based on the synchronization timestamp derivation method, the clock deviation of each speaker node and the obtained compensation delay time are combined to derive the local control timestamp of each speaker node.

[0065] For example, when the propagation delay of a speaker node is 0.002 seconds, the clock deviation is 0.0005 seconds, and the local control timestamp of the reference speaker is 2.000 seconds, it can be deduced that the local control timestamp of the speaker is 2.0025 seconds.

[0066] S3.3: Delay the reconstructed audio waveform according to the local control timestamp of each speaker, collect the sound field energy distribution data, and obtain the sound field focusing accuracy through the sound energy density analysis method.

[0067] Specifically, the reconstructed audio waveform is delayed according to the local control timestamp of each speaker, and the audio waveforms output by each speaker are aligned in the time domain within the target position, where the target position is the three-dimensional focusing coordinate point preset in the sound field design. Subsequently, with the target position as the center, the target space range is defined according to the focusing tolerance radius set in the application scenario, the sound field energy distribution data is collected within the target space range, and the sound energy density distribution within the unit volume is calculated by the sound energy density analysis method to obtain the sound field focusing accuracy.

[0068] For example, if the sound energy density collected within a radius of 0.05 meters from the focusing point reaches more than 90% of the total sound energy, it can be judged that the sound field focusing accuracy is high.

[0069] S4: Based on the sound field focusing accuracy, the variational autoencoder is dynamically adjusted through the policy gradient algorithm. The energy efficiency data of the speaker drive circuit and the ambient noise spectrum characteristics are fused in multiple dimensions through the multimodal feature fusion method to generate a set of dynamic control parameters.

[0070] Specifically, the following steps are included: S4.1: Based on the sound field focusing accuracy, the compression rate of the variational autoencoder is dynamically corrected through the policy gradient algorithm.

[0071] Specifically, the sound field focusing accuracy is used as the core evaluation criterion of the reward function. The reward function is constructed by mapping the sound field focusing accuracy to numerical rewards. The policy gradient algorithm is used to continuously sample the reconstructed audio waveform corresponding to the low-dimensional semantic vector generated by the variational autoencoder at different compression rates, and the propagation delay processing is performed on the reconstructed audio waveform corresponding to the low-dimensional semantic vector to obtain the sound field energy distribution data corresponding to the reconstructed audio waveform; the sound field focusing accuracy is used as the reward signal of the policy gradient algorithm, and the policy parameters related to the compression rate are updated in each round of training; according to the iterative adjustment results of the policy parameters, the compression rate setting of the variational autoencoder is corrected, so that the variational autoencoder can optimize the focusing effect of the audio restored by the low-dimensional semantic vector in space while maintaining the semantic expression ability, and finally realize the dynamic correction of the compression rate of the variational autoencoder.

[0072] For example, in an audio reconstruction task for a conference room scenario, the initial compression rate of the variational autoencoder is set to 0.25, indicating that the input audio signal is compressed to 25% of its original dimension during the encoding process. The low-dimensional semantic vector generated at this compression rate is input into a generative adversarial network to reconstruct the corresponding audio waveform and play it through a multi-speaker array.

[0073] S4.2: Collect energy efficiency data of the speaker driver circuit and the environmental noise spectrum characteristics.

[0074] Specifically, by measuring the input electric power and the corresponding output sound power of the speaker driving circuit, and adopting the sound energy conversion rate analysis method to obtain the sound energy conversion efficiency of the speaker driving circuit, the sound energy conversion efficiency is used as the energy efficiency ratio data of the speaker driving circuit; at the same time, the ambient noise signal is collected by using a microphone array, and the energy distribution of the ambient noise signal in each frequency range is extracted by short-time Fourier transform, and the extracted spectral energy value is normalized to obtain the spectral characteristics of the ambient noise.

[0075] For example, by collecting the average sound pressure level of 85 dB generated by a speaker under 1 watt input power, the energy efficiency ratio can be calculated as 85 dB / W. In the environmental noise spectrum analysis, if there is a significant energy peak at a frequency of 500 Hz, this frequency band is recorded as the main interference frequency band.

[0076] S4.3: The energy efficiency data of the speaker driving circuit and the spectrum characteristics of the ambient noise are fused through a multimodal feature fusion method to generate a multi-dimensional fusion feature vector.

[0077] Specifically, the energy efficiency data of the speaker driving circuit and the spectrum characteristics of the ambient noise are normalized and feature extracted to obtain numerical normalized energy efficiency data and a multidimensional spectrum feature vector describing the ambient noise. The normalized energy efficiency data and the multidimensional spectrum feature vector are synthesized into a multidimensional fusion feature vector through a feature-level fusion strategy.

[0078] For example, the normalized energy efficiency ratio value of 0.75 and the energy ratios of the ambient noise in the low-frequency (100 Hz–300 Hz) and high-frequency (2000 Hz–4000 Hz) frequency bands of 0.45 and 0.20 are sequentially concatenated to generate a three-dimensional fusion feature vector [0.75, 0.45, 0.20].

[0079] S4.4: Encode the multi-dimensional fusion feature vector through the modified variational autoencoder to generate a set of dynamic control parameters.

[0080] Specifically, the multi-dimensional fused feature vector is input into the encoder part of the modified variational autoencoder, and the high-dimensional input features are mapped to the latent space through a multi-layer neural network structure. The mean vector and variance vector of the latent space are extracted, and then the latent variables are sampled from the latent space distribution through the reparameterization technique. Finally, the latent variables are mapped through the decoder to generate a set of dynamic control parameters.

[0081] This embodiment also provides an intelligent energy-saving audio playback control system, including: The semantic compression module preprocesses the original audio input signal and inputs it into the variational autoencoder to compress it into a low-dimensional semantic vector. At the same time, it obtains the distance between the terminal device and the audio signal source based on the sound wave reflection time difference fed back by the terminal device; The sound field modeling module reconstructs the audio waveform based on the low-dimensional semantic vector through a generative adversarial network to build an initial neural sound field model. It then makes real-time corrections using the physical sound wave propagation equation to generate a neural sound field model. The focus control module calculates the phase distribution of the multi-speaker array based on the neural sound field model and coordinates the delay time of the speaker array through the distributed clock synchronization protocol to obtain the sound field focusing accuracy; The dynamic control module dynamically adjusts the variational autoencoder through the policy gradient algorithm according to the sound field focusing accuracy, and uses the multimodal feature fusion method to perform multi-dimensional feature fusion on the energy efficiency data of the speaker drive circuit and the ambient noise spectrum characteristics to generate a set of dynamic control parameters.

[0082] This embodiment also provides a computer device suitable for the intelligent energy-saving audio playback control method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the intelligent energy-saving audio playback control method proposed in the above embodiment.

[0083] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.

[0084] This embodiment also provides a storage medium having a computer program stored thereon. When the program is executed by a processor, the program implements the intelligent energy-saving audio playback control method proposed in the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0085] In summary, the present invention achieves efficient semantic compression of audio data and dynamic perception of the spatial environment by performing multi-level preprocessing on the original audio input signal, inputting the processed high-dimensional audio signal into a variational autoencoder for low-dimensional semantic vector compression, and accurately calculating the physical distance between the device and the audio signal source in combination with the sound wave reflection time difference fed back by the terminal device. At the same time, the spatial position information is integrated into the compression process, providing a reliable foundation for the adaptive selection of subsequent wireless communication protocols and the optimization of the audio playback environment. Ultimately, the data processing efficiency and the system's intelligent perception capabilities are improved while ensuring the quality of audio transmission, effectively supporting the intelligent and efficient control of energy-saving audio playback.

[0086] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An intelligent energy-saving audio playback control method, characterized by: include, The original audio input signal is preprocessed and input into a variational autoencoder to compress it into a low-dimensional semantic vector. At the same time, the distance between the terminal device and the audio signal source is obtained based on the sound wave reflection time difference fed back by the terminal device. Based on the low-dimensional semantic vector, the audio waveform is reconstructed through a generative adversarial network to build an initial neural sound field model. This model is then corrected in real time using the physical sound wave propagation equation to generate a neural sound field model. Based on the neural sound field model, the phase distribution of the multi-speaker array is calculated, and the delay time of the speaker array is coordinated through the distributed clock synchronization protocol to obtain the sound field focusing accuracy; According to the sound field focusing accuracy, the variational autoencoder is dynamically adjusted through the policy gradient algorithm, and the energy efficiency data of the speaker driving circuit and the ambient noise spectrum characteristics are multi-dimensionally fused through the multimodal feature fusion method to generate a dynamic control parameter set.

2. The intelligent energy-saving audio playback control method according to claim 1, characterized in that: The input is compressed into a low-dimensional semantic vector in the variational autoencoder. The steps are as follows: Input the preprocessed audio input signal into the variational autoencoder to obtain the mean and variance of the latent space; The variational autoencoder is trained using historical audio data, and the mean and variance of the latent space are input into the trained variational autoencoder for reparameter sampling to generate a low-dimensional semantic vector.

3. The intelligent energy-saving audio playback control method according to claim 1, characterized in that: The reconstructing of the audio waveform by generating an adversarial network refers to inputting a low-dimensional semantic vector into a generative adversarial network, generating an audio waveform and comparing it with a real audio waveform, optimizing the generated audio waveform through a loss function, and generating a reconstructed audio waveform.

4. The intelligent energy-saving audio playback control method according to claim 1, characterized in that: The steps of constructing the initial neural sound field model are as follows: Collect audio signal sample data with spatial coordinate annotations, and use multimodal data fusion and training methods to construct neural radiation fields; Short-time Fourier transform is used to divide the reconstructed audio waveform into short-time frame signals weighted by window functions. The short-time frame signals weighted by window functions and 3D spatial coordinates are integrated into the neural radiation field through the multimodal sound field implicit method to generate an initial neural sound field model.

5. The intelligent energy-saving audio playback control method according to claim 1, characterized in that: The steps of performing real-time correction through the physical sound wave propagation equation to generate the neural sound field model are as follows: Substitute the distance between the terminal device and the audio source into the free space attenuation formula, extract the attenuation coefficient based on the inverse square law, and substitute the acoustic impedance values ​​of the incident medium and the reflecting medium in the neural sound field model into the impedance ratio formula to determine the reflection parameter through the acoustic impedance ratio of the incident medium to the reflecting medium. According to the attenuation coefficient and the reflection parameter, the propagation characteristics in the initial neural sound field model are adjusted to obtain the neural sound field model.

6. The intelligent energy-saving audio playback control method according to claim 1, characterized in that: The distributed clock synchronization protocol is used to coordinate the delay time of the speaker array and obtain the sound field focusing accuracy. The steps are as follows: The three-dimensional spatial position of each speaker unit in the neural sound field model is extracted through the back propagation positioning method, the propagation distance between the speaker and the sound output unit is calculated, and the propagation delay compensation method is used to identify the compensation delay time of each speaker with the speaker closest to the sound output unit as a reference; The clock deviation of each speaker node is obtained by a two-way delay measurement method, and then the local control timestamp of each speaker is obtained from the clock deviation and the compensation delay time by a synchronization timestamp derivation method; The reconstructed audio waveform is delayed according to the local control timestamp of each speaker, the sound field energy distribution data is collected, and the sound field focusing accuracy is obtained through the sound energy density analysis method.

7. The intelligent energy-saving audio playback control method according to claim 1, characterized in that: The steps of generating a dynamic control parameter set are as follows: Based on the sound field focusing accuracy, the compression rate of the variational autoencoder is dynamically corrected through the policy gradient algorithm; Collect energy efficiency data of the speaker drive circuit and environmental noise spectrum characteristics; Through the multimodal feature fusion method, the energy efficiency ratio data of the speaker driving circuit and the spectrum characteristics of the ambient noise are fused to generate a multi-dimensional fusion feature vector; The multi-dimensional fusion feature vector is encoded by the modified variational autoencoder to generate a set of dynamic control parameters.

8. An intelligent energy-saving audio playback control system, based on the intelligent energy-saving audio playback control method according to any one of claims 1 to 7, characterized in that: include, The semantic compression module preprocesses the original audio input signal and inputs it into the variational autoencoder to compress it into a low-dimensional semantic vector. At the same time, it obtains the distance between the terminal device and the audio signal source based on the sound wave reflection time difference fed back by the terminal device; The sound field modeling module reconstructs the audio waveform based on the low-dimensional semantic vector through a generative adversarial network to build an initial neural sound field model. It then makes real-time corrections using the physical sound wave propagation equation to generate a neural sound field model. The focus control module calculates the phase distribution of the multi-speaker array based on the neural sound field model and coordinates the delay time of the speaker array through the distributed clock synchronization protocol to obtain the sound field focusing accuracy; The dynamic control module dynamically adjusts the variational autoencoder through the policy gradient algorithm according to the sound field focusing accuracy, and uses the multimodal feature fusion method to perform multi-dimensional feature fusion on the energy efficiency data of the speaker drive circuit and the ambient noise spectrum characteristics to generate a set of dynamic control parameters.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the intelligent energy-saving audio playback control method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent energy-saving audio playback control method according to any one of claims 1 to 7 are implemented.