Wireless audio terminal super-dense networking and multi-source data fusion method

By collecting temporal features of the sound field and wireless channel state information in wireless audio terminals, unsupervised clustering is performed using Bayesian nonparametric clustering and graph convolutional networks. Combined with multi-agent deep reinforcement learning and complex network theory, node density and transmission power are dynamically adjusted to solve the problems of co-channel interference and redundant transmission in high-density deployment of wireless audio terminals, thus achieving efficient audio data transmission and adaptive networking.

CN122294256APending Publication Date: 2026-06-26SHENZHEN ZUNTE DIGITAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN ZUNTE DIGITAL CO LTD
Filing Date
2026-05-19
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In scenarios with high-density deployment of wireless audio terminals, existing technologies cannot identify the spatial correlation of sound fields among terminals within the same sound field coverage area, leading to co-channel interference and redundant transmission, and a disconnect between resource allocation and the physical characteristics of the sound field of audio services.

Method used

By collecting temporal features of the sound field and wireless channel state information, unsupervised clustering is performed using a Bayesian nonparametric clustering model and a graph convolutional network. Combined with multi-agent deep reinforcement learning and complex network theory, node density and transmission power are dynamically adjusted to optimize time slot resource allocation and multi-hop relay transmission paths. Quantum heuristic optimization and conditional generative adversarial networks are introduced to recover audio frequency bands.

Benefits of technology

It effectively avoids co-channel interference, reduces redundant transmission, improves throughput and resilience, and enables adaptive anomaly recovery and efficient transmission of audio data in ultra-dense networking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122294256A_ABST
    Figure CN122294256A_ABST
Patent Text Reader

Abstract

This invention discloses a method for ultra-dense wireless audio terminal networking and multi-source data fusion, belonging to the field of ultra-dense networking and audio processing technology. The method involves the terminal collecting sound field temporal characteristics, wireless channel status, and location information; performing distributed time-slot collaborative scheduling based on sound field spatial correlation, allocating orthogonal short time slots to suppress co-channel interference; implementing unsupervised clustering of multi-source data and generating fusion weights through a Dirichlet process variational autoencoder; dynamically adjusting point density and transmission power using the Navier-Stokes equation, and optimizing relay links by combining small-world networks and quantum heuristics; compensating for high-frequency audio gaps using a conditional generative adversarial network, and uploading the data after joint encoding of the source and channel; and employing a clonal selection algorithm to achieve self-healing of network anomalies. This invention can reduce interference and redundant transmission, improving audio transmission quality and system robustness in ultra-dense scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention discloses a method for ultra-dense networking and multi-source data fusion of wireless audio terminals, belonging to the field of ultra-dense networking and audio processing technology. Background Technology

[0002] In existing high-density deployment scenarios for wireless audio terminals, conventional networking schemes typically allocate time slots based on physical layer wireless channel state information. Each terminal acts as a single data transmission node, independently reporting its audio stream to the central access point. In this conventional mode, the resource scheduling module only monitors physical channel indicators such as signal-to-noise ratio and received signal strength indication. When these channel indicators meet preset thresholds, the terminal is allowed to occupy transmission time slots for data transmission.

[0003] The core problem with the aforementioned conventional technical solutions lies in the disconnect between time slot resource allocation and the physical characteristics of the audio service's sound field, resulting in the inability to eliminate co-channel interference and redundant transmission. Specifically, multiple terminals within the same sound field coverage area collect direct and reflected sound from the same sound source, and their audio data are highly correlated in physical space. However, existing technologies allocate time slots solely based on the wireless channel status, failing to identify this acoustic spatial correlation. This causes terminals with strong spatial correlation to transmit highly repetitive sound field data in the same or adjacent time slots, leading to co-channel interference and transmission collisions, resulting in the ineffective use of channel resources. Summary of the Invention

[0004] The purpose of this invention is to provide a solution to the problems described in the background section.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for ultra-dense networking of wireless audio terminals and fusion of multi-source data includes: the wireless audio terminal, as a micro access node and terminal node, collects the temporal characteristics of the sound field of its own pickup channel and the wireless channel state information; Distributed time-slot cooperative scheduling allocates orthogonal short time-slot transmission resources to wireless audio terminals within the same sound field coverage area based on the spatial correlation of sound field characteristics. The Bayesian nonparametric clustering model performs unsupervised clustering on sound field data, wireless channel state data and terminal location data uploaded by multiple wireless audio terminals, and generates multi-source data fusion weights within the same cluster. The fused audio data is accessed through the cluster head node to the upper-layer wireless LAN. The cluster head node dynamically adjusts the node density and transmission power of the network within the cluster based on the weight of the multi-source data fusion.

[0006] Preferably, the spatial correlation calculation method for sound field features in distributed time-slot collaborative scheduling is as follows: wireless audio terminals within the same sound field coverage area are constructed as a dynamic topology graph, with the terminal location as the node coordinate and the cross-correlation coefficient of the sound field time-domain features as the edge weight. The dynamic topology graph is input into a graph convolutional network to extract high-dimensional spatial correlation features that include the geometric distribution of sound sources and acoustic attenuation characteristics. Based on these high-dimensional spatial correlation features, orthogonal short time slot transmission resources are pre-allocated.

[0007] Preferably, based on the pre-allocation of orthogonal short time slot transmission resources, multi-agent deep reinforcement learning is introduced to eliminate time slot conflicts: each wireless audio terminal is used as an independent agent, the current high-dimensional spatial related features and wireless channel state information are used as the state space, the time slot offset and transmission power fine-tuning are used as the action space, and the minimum co-channel interference penalty and the maximum throughput are used as the joint reward function. The optimal time slot scheduling strategy is output through the multi-agent near-end policy optimization algorithm.

[0008] Preferably, the Bayesian nonparametric clustering model is implemented using a Dirichlet process variational autoencoder: the sound field data, wireless channel state data and terminal location data are concatenated into a multi-source heterogeneous tensor; The encoder network maps multi-source heterogeneous tensors to the latent manifold space. Nonparametric clustering is performed in the latent manifold space based on the Dirichlet process prior. The lower bound of evidence is maximized through variational inference, the number of clusters is automatically determined, and the fusion weights of multi-source data within each cluster are output.

[0009] Preferably, in the step of dynamically adjusting the node density and transmission power of the cluster network based on the multi-source data fusion weights, the Navier-Stokes equations from computational fluid dynamics are introduced for modeling: The wireless audio terminals within the cluster are mapped as fluid micro-elements, and the multi-source data fusion weights are mapped as fluid density and pressure fields. The velocity field and pressure gradient of the node distribution are calculated by solving the discretized Navier-Stokes equations. The velocity field is converted into dynamic addition and deletion instructions for node density, and the pressure gradient is converted into dynamic adjustment step size of transmission power.

[0010] Preferably, after the cluster head node adjusts the cluster topology according to the dynamic addition and deletion instructions of node density, a small-world network reconstruction mechanism from complex network theory is introduced to optimize the intra-cluster communication links: Calculate the average path length and clustering coefficient of the current cluster topology. If the average path length is greater than a preset threshold, introduce long-range connection edges through a reconnection strategy. The establishment of long-range connection edges is based on the Wasserstein distance between different clusters in the potential manifold space to improve the robustness and synchronous convergence speed of multi-hop relay transmission within the cluster.

[0011] Preferably, when performing multi-hop relay transmission via long-range connection edges, a quantum heuristic optimization algorithm is introduced to optimize the transmission path: The relay node sequence is encoded into a quantum bit state, and the quantum bit probability amplitude is updated through a quantum rotation gate. The weighted sum of end-to-end delay and transmission power adjustment step size is used as the fitness function to perform a collapse measurement on the quantum state. The optimal multi-hop relay path from the cluster to the cluster head node is output to reduce the queuing delay of the fused audio data accessing the upper-layer wireless LAN.

[0012] Preferably, when the cluster head node performs deep fusion of the audio data aggregated within the cluster, a conditional generative adversarial network is introduced to compensate for missing audio frequency bands: The multi-source data fusion weights and wireless channel state data are used as the input of the conditional vector generator. The generator reconstructs the high-frequency acoustic details lost due to co-channel interference based on the conditional vector. The discriminator uses terminal location data as a priori constraint, and performs adversarial training until Nash equalization. It then performs frequency domain splicing of the high-frequency acoustic details output by the generator with the fused audio data.

[0013] Preferably, before the frequency-domain spliced ​​audio data is accessed by the upper-layer wireless LAN, rate-distortion theory and mutual information maximization criteria are introduced for joint source-channel coding: A joint cost function is constructed with audio data distortion and wireless channel transmission bit error rate as independent variables. Under the constraint of the total bandwidth of the upper-layer wireless LAN, Pareto optimal solutions for the source coding rate and channel coding rate that minimize the joint cost function are found. Based on the Pareto optimal solution, the frequency-domain spliced ​​audio data is adaptively modulated and hierarchically coded for transmission.

[0014] Preferably, during the hierarchical coding and transmission process based on the Pareto optimal solution, a clonal selection algorithm from the biological immune system is introduced to recover from network anomalies. Real-time monitoring of gradient changes in the joint cost function; triggering an abnormal immune response when the gradient mutation exceeds a preset threshold. The current optimal multi-hop relay path and orthogonal short-slot scheduling strategy are mapped to antibodies. Diverse scheduling schemes are generated through cloning, high-frequency mutation and receptor editing operations. The scheme that fastest inhibits the increase of the joint cost function is used as the memory cell and stored in the cluster head node to achieve the adaptive evolution of ultra-dense networking.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention addresses the core problem of the disconnect between time slot allocation and the physical characteristics of the sound field by introducing spatial correlation features of the sound field before time slot allocation. By using the cross-correlation coefficient of the sound field as the edge weight of the dynamic topology graph and combining it with a graph convolutional network to extract high-dimensional features containing the geometric distribution characteristics of the sound sources, the pre-allocation of time slots is directly controlled by acoustic physical laws. Based on the pre-allocation, a multi-agent near-end policy optimization algorithm combined with wireless channel state information is used to eliminate time slot conflicts, ensuring that terminals within the same sound field coverage area are allocated orthogonal short time slots, thus avoiding co-channel interference between spatially correlated terminals from the physical source. Simultaneously, a Dirichlet process variational autoencoder is used to perform unsupervised clustering and variational inference on the multi-source heterogeneous tensor composed of sound field data, channel state, and location, automatically determining clusters and generating fusion weights. Combining the Navier-Stokes equations, the fusion weights are mapped to a fluid physical field, and the solved velocity field and pressure gradient are directly converted into node density increase / decrease commands and transmission power adjustment step sizes, reducing the repeated transmission of highly correlated redundant sound field data in the network.

[0016] 2. After dynamic adjustment of the intra-cluster topology, this invention utilizes a small-world network reconstruction mechanism, introducing long-range connection edges based on the probability distribution distance in the potential manifold space. This shortens the average path length of multi-hop relays and improves the resilience and synchronization convergence speed of the intra-cluster topology. For multi-hop transmission path planning, a heuristic optimization method using quantum bit state encoding and quantum rotating gate updates is employed. End-to-end delay and power adjustment step size are used as fitness functions, reducing the queuing delay for fused audio data accessing the upper-layer network. To address audio loss caused by interference, a conditional generative adversarial network is used. High-frequency acoustic details are reconstructed using multi-source data fusion weights and channel data as conditional vectors. A discriminator is combined to introduce prior position constraints, recovering audio frequency band information. During the data uplink phase, Pareto optimal solutions for source and channel coding rates are solved based on rate-distortion theory and the mutual information maximization criterion, achieving adaptive hierarchical coding transmission. During network operation, the gradient change of the joint cost function is monitored, triggering a cloning selection algorithm to generate diverse scheduling schemes and retain the optimal solution, achieving adaptive anomaly recovery in high-density networks. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the overall process of ultra-dense networking and multi-source data fusion of the wireless audio terminal according to the present invention. Figure 2 This is a flowchart of the distributed time-slot collaborative scheduling process of the present invention; Figure 3 This is a flowchart of the Bayesian nonparametric clustering and weight calculation process of the present invention; Figure 4 This is a flowchart illustrating the dynamic adjustment of node density and transmission power according to the present invention. Figure 5This is a flowchart of the intra-cluster topology optimization and relay path optimization of the present invention; Figure 6 This is a flowchart of the audio deep fusion and network anomaly recovery process of the present invention. Detailed Implementation

[0018] As a preferred embodiment, please refer to the appendix. Figure 1 The wireless audio terminal acquires the time-domain audio signal of its own pickup channel through a local pickup module, denoted as... ,in A unique identifier for the terminal. For discrete-time indexing, the time-domain audio signal is processed by frame segmentation, with a frame length of 20ms, a frame shift of 10ms, and a Hanning window applied. The short-time energy, zero-crossing rate, and temporal envelope of each frame are extracted as temporal features of the sound field. The wireless audio terminal measures the channel frequency response of the current wireless channel through the wireless communication module to obtain the wireless channel state information, denoted as... ,in This is an OFDM subcarrier frequency index, containing amplitude and phase information for each subcarrier; the three-dimensional spatial coordinates of the terminal are obtained through the positioning module, denoted as... The unit is meters. The wireless audio terminal broadcasts its own sound field temporal characteristics, wireless channel state information, and terminal location data to other terminals within the same sound field coverage area via a side link. The same sound field coverage area is a spatial region where the straight-line distance between terminals is less than a preset sound field-related threshold.

[0019] Distributed time-slot collaborative scheduling allocates orthogonal short time-slot transmission resources to wireless audio terminals within the same sound field coverage area based on the spatial correlation of the sound field temporal characteristics of each terminal. The spatial correlation of the sound field temporal characteristics is determined by two terminals... and The cross-correlation coefficient of the sound field time-domain characteristics is characterized by the following formula: in, for The mean, for standard deviation For mathematical expectation calculation, The range of values ​​is The closer its absolute value is to 1, the stronger the spatial correlation of the sound field characteristics of the two terminals.

[0020] The formula is based on the acoustic spatial propagation law. The cross-correlation coefficient of the audio signal collected by the pickup terminals within the same sound field coverage area is strongly coupled with the spatial distance between the terminals and the incident angle of the sound source. This is consistent with the acoustic attenuation characteristics of spherical wave propagation in a free field. The absolute value of the cross-correlation coefficient decreases exponentially with the increase of the distance between the terminals. It can characterize the spatial correlation of the sound field and provide the core judgment basis of the physical layer for subsequent time slot allocation.

[0021] Distributed time slot collaborative scheduling uses a distributed consensus algorithm to allocate non-overlapping orthogonal short time slots to spatially correlated terminals. The duration of each short time slot is 125μs, corresponding to the MU-OFDMA time slot granularity of the Wi-Fi 7 protocol. Each short time slot corresponds to a set of orthogonal subcarrier resources, thus avoiding co-channel interference caused by spatially correlated terminals transmitting data in the same time slot.

[0022] The cluster head node collects sound field data, wireless channel status data, and terminal location data uploaded by all wireless audio terminals within the same sound field coverage area. It then performs unsupervised clustering of this multi-source data using a Bayesian nonparametric clustering model, generating multi-source data fusion weights within the same cluster. First, it processes the multi-source data vectors of each terminal... After performing min-max normalization, the calculation formula is as follows: in, and These represent the minimum and maximum values ​​of the corresponding data dimensions for all terminals within the same sound field coverage area. The Bayesian nonparametric clustering model employs a Dirichlet process mixture model, with the base distribution set as a normal-inverse Wissaud distribution, and the concentration parameter... When set to 1.0, the number of clusters is automatically determined by posterior inference through Gibbs sampling, without the need to preset the number of clusters.

[0023] This concentration parameter The setting is based on the concentration parameter in the Dirichlet process. Directly control the number of clusters generated. The larger the value, the more clusters are generated, and vice versa. This embodiment focuses on typical wireless audio terminal deployment scenarios such as conference rooms, concert halls, and stadiums, completing 500 sets of scenario simulation tests, covering an ultra-dense deployment range of 8 to 128 terminals. The test results show that when… At that time, the model can automatically adapt to the number of sound sources in different scenarios, and the matching degree between the number of clusters and the actual number of independent sound sources reaches 96.2%, which can accurately achieve unsupervised clustering of the same sound source terminal; if This can easily lead to under-clustering problems where multiple sound sources merge into the same cluster, resulting in a matching degree decrease of ≥31%; if This can easily lead to over-clustering, where the same sound source splits into multiple clusters, resulting in a matching accuracy decrease of ≥24%. Meanwhile, The settings conform to the conventional prior settings of unsupervised clustering in Bayesian nonparametric models, taking into account both the convergence speed and clustering accuracy of the model. The model's single-frame inference time is ≤5ms, which meets the latency requirements of real-time audio processing.

[0024] For each generated cluster Calculate intra-cluster terminals The multi-source data fusion weight is calculated using the following formula: in, For clusters The center vector, For clusters The mean of the trace of the covariance matrix is ​​used for summation operations across clusters. All terminals within The fusion weights satisfy the normalization constraint, and the sum of the fusion weights of all terminals within the same cluster is 1.

[0025] The cluster head node performs weighted fusion of the temporal audio data of all terminals within the cluster based on the generated multi-source data fusion weights to obtain the fused audio data. The calculation formula is as follows: The fused audio data is accessed through the uplink communication link of the cluster head node to the upper-layer wireless LAN. The cluster head node dynamically adjusts the node density and transmission power within the cluster based on the multi-source data fusion weight: for terminals with a fusion weight lower than a preset weight threshold, the cluster head node sends a sleep command to control the corresponding terminal into a low-power sleep state, reducing the node density within the cluster and minimizing redundant data transmission; for active terminals within the cluster, their transmission power is adjusted according to the fusion weight, calculated using the following formula: in, The maximum permissible transmission power for wireless audio terminals to comply with radio management regulations is determined by the fusion weight. Terminals with higher fusion weights are allocated lower transmission power, thereby reducing co-channel interference while ensuring transmission reliability.

[0026] As a preferred embodiment, refer to the appendix. Figure 2 In distributed time-slot collaborative scheduling, wireless audio terminals within the same sound field coverage area are first constructed into a dynamic topology map. , where the set of nodes Within the domain Each of the following is a list of wireless audio terminals and their initial attribute vectors. It includes terminal location information and temporal characteristics of the sound field; edge set In the middle, node and Edge weights between That is, the absolute value of the cross-correlation coefficient of the time-domain characteristics of the sound field, when When the edge weight is below a preset threshold, the corresponding edge is not established, and the adjacency matrix of the dynamic topology graph is... The elements satisfy: when ,otherwise .

[0027] The dynamic topology graph is input into a two-layer graph convolutional network to extract high-dimensional spatial correlation features containing the geometric distribution of sound sources and acoustic attenuation characteristics. The first layer convolution operation formula of the graph convolutional network is as follows: in, For the normalized adjacency matrix, It is the identity matrix. It is a degree matrix and ; This is the first-layer trainable weight matrix, with dimensions (input feature dimension, 64). It is the ReLU activation function; This is a matrix consisting of the initial attribute vectors of all nodes.

[0028] The formula for the second layer convolution operation is: in, The second-layer trainable weight matrix has dimensions (64, 32), and the output is... Each row vector in the graph represents the 32-dimensional high-dimensional spatial correlation feature of the corresponding node. The graph convolutional network is trained offline using contrastive loss, and the weights after training are embedded in the NPU of the wireless audio terminal, outputting the high-dimensional spatial correlation features in real time during online inference. Based on the output high-dimensional spatial correlation features, K-means clustering is performed to group nodes with high feature similarity into the same category. Different orthogonal short time slots are assigned to different categories, and orthogonal subcarriers within the same time slot are assigned to nodes within the same category, thus completing the pre-allocation of orthogonal short time slot transmission resources.

[0029] Based on the pre-allocation of orthogonal short time-slot transmission resources, multi-agent deep reinforcement learning is introduced to eliminate time-slot conflicts. A multi-agent near-end policy optimization algorithm is adopted, implemented with a framework of centralized training and distributed execution. Each wireless audio terminal is treated as an independent agent, and the state space of each agent is as follows: The local observation vector includes its own high-dimensional spatial correlation features, the amplitude mean and phase variance of its own wireless channel state information, the one-hot encoded time slot occupancy status of neighboring terminals, and the measured interference power of the current time slot, for a total of 51 dimensions. The action space of each agent... The continuous action space contains two dimensions: time slot offset. The range of values ​​is Used to adjust the start time of the transmission time slot; transmission power fine-tuning amount. The range of values ​​is It is used to adjust the transmission power to reduce interference.

[0030] A global joint reward function is set, which is shared by all agents. The calculation formula is as follows: in, Normalized system throughput is the ratio of the number of data packets successfully transmitted in the current time slot to the total number of data packets, with a value ranging from [value missing]. ; The co-channel interference penalty is equal to the ratio of the average interference power measured by all terminals in the current time slot to the maximum permissible interference power, and its value range is [value range missing]. ; This is the weighting coefficient, preset to 1.2, used to balance the optimization objectives of throughput improvement and interference suppression.

[0031] The design principle of this reward function is as follows: a centralized shared reward mechanism is adopted to solve the credit allocation problem in multi-agent reinforcement learning. The optimization objectives of all agents are strongly bound to the global network performance, avoiding the deterioration of global interference caused by the selfish optimization of a single agent, and ensuring the global optimality of distributed time slot scheduling.

[0032] This weighting coefficient The design is based on the following: the core optimization objective of this invention is to suppress co-channel interference in ultra-dense network scenarios, while ensuring that the system throughput meets the real-time audio transmission requirements. This embodiment uses 1000 Monte Carlo simulation experiments in ultra-dense scenarios with terminal densities ranging from 0.5 to 5 terminals / m² to... Gradient traversal tests were performed in the range of 0.5 to 2.0. The test results show that when At that time, the system's co-channel interference suppression rate reached 89.7%, while the throughput remained at a peak level of 92.3%, achieving an optimal balance between interference suppression and throughput; if The interference suppression effect decreases by ≥22%, failing to meet the anti-interference requirements of ultra-dense scenarios; if The system throughput decreased by ≥15%, failing to meet the bandwidth requirements for real-time audio transmission.

[0033] Each agent's policy network and value network are both three-layer fully connected networks. The policy network takes a state space vector as input and outputs the Gaussian distribution mean and variance of the actions; the value network takes a state space vector as input and outputs the value estimate for the corresponding state. The cluster head node acts as the central node, collecting the state, action, and reward data of all agents for centralized training. After updating the network parameters, it distributes the updated data to all wireless audio terminals. When the terminals are running online, they perform policy network inference only through their local NPU, outputting the optimal time slot offset and transmit power fine-tuning to eliminate time slot conflicts and achieve optimal time slot scheduling.

[0034] As a preferred embodiment, refer to the appendix. Figure 3 The Bayesian nonparametric clustering model is implemented using a Dirichlet process variational autoencoder. First, the sound field data, wireless channel state data, and terminal location data of each wireless audio terminal are concatenated into a multi-source heterogeneous tensor: the sound field data consists of 10 frames of time-domain audio signals, flattened into a dimension... The vector; the wireless channel state data consists of the amplitude and phase information of 320 subcarriers, flattened into a dimension. The vector; the terminal position data is 3D coordinates, flattened to a dimension. The vector; concatenating the above vectors yields the dimension. Multi-source heterogeneous vectors After normalization, it is used as the input of the Dirichlet process variational autoencoder.

[0035] The Dirichlet process variational autoencoder comprises an encoder network and a decoder network. The encoder network consists of a 5-layer one-dimensional convolutional network and a fully connected layer. The input is a normalized multi-source heterogeneous vector. The convolutional kernel sizes are 16, 8, 4, 2, and 1, respectively, with strides of 2, 2, 2, 2, and 1, and channel numbers of 64, 128, 256, 128, and 64, respectively. After global average pooling, the input is fed into the fully connected layer, mapping to 32-dimensional latent variables in the latent manifold space. It also outputs the variational distribution parameters of the mixing coefficients of the Dirichlet process.

[0036] In the latent manifold space, nonparametric clustering is performed based on the Dirichlet process prior. The Dirichlet process is constructed using a break-down construction method, and the mixing ratio is calculated using the following formula: in, , For concentration parameters, For the first The mixing ratio of each cluster, , The preset cutoff value is 20, which covers the maximum number of clusters in the actual scenario.

[0037] Model training is achieved by maximizing the lower bound of evidence through variational inference. The formula for calculating the lower bound of evidence is as follows: in, It is a variational distribution. For the generation distribution of the decoder, For KL divergence calculation, The latent variables are a mixture of Gaussian priors, with each cluster corresponding to a Gaussian distribution. , The prior distribution is the Dirichlet process. During training, gradient backpropagation is achieved through reparameterization techniques to maximize the lower bound of evidence. After training, the model weights are fixed in the NPU of the cluster head node.

[0038] During online inference, the cluster head node inputs the collected multi-source heterogeneous vectors into the encoder network to obtain the latent variables of the corresponding terminal. and terminals Belongs to the Posterior probability of each cluster The system automatically determines the number of clusters and discards invalid clusters whose sum of posterior probabilities is below a preset threshold. The fusion weights of multi-source data within the same cluster are obtained by normalizing the posterior probabilities, calculated using the following formula: As an effective embodiment, refer to the appendix. Figure 4 When the cluster head node dynamically adjusts the node density and transmission power of the cluster network based on the weights of multi-source data fusion, the Navier-Stokes equations from computational fluid dynamics are introduced for modeling. First, the minimum bounding box of the three-dimensional position coordinates of all wireless audio terminals within the cluster is used as the computational domain. ,Will Discretize into a uniform 3D mesh, mesh step size Each grid cell corresponds to a fluid micro-element.

[0039] Establish the mapping relationship between the fluid physics field and the network parameters: If a wireless audio terminal exists within a mesh cell, the density field of that micro-element... This is equal to the multi-source data fusion weight of the corresponding terminal. If there are multiple terminals within a grid cell, the density field is the sum of the fusion weights; if there are no terminals within a grid cell, the density field value is 0. Pressure field It is positively correlated with the density field, and the calculation formula is: in, The environmental pressure baseline value is set to 1.0; This is the pressure coefficient, preset to 5.0.

[0040] The pressure coefficient is set based on the following: the linear mapping relationship between the pressure field and the density field needs to be adapted to the dynamic adjustment range of the transmission power of the wireless audio terminal. This embodiment is based on the general radio management specifications of a maximum transmission power of 20dBm and a minimum transmission power of -10dBm for the terminal. It uses a combination of fluid dynamics simulation and network-based testing to... Traversal tests were performed in the range of 1.0 to 10.0. The test results show that when At this time, the dynamic range of the pressure field is 1.0 to 6.0, which can be linearly mapped to the full range adjustment of the terminal transmission power. The power adjustment step size has a resolution of 0.5 dBm, perfectly matching the gradient change of the audio fusion weights; if The pressure field dynamic range is insufficient, and the resolution of power adjustment decreases by ≥60%, making fine power control impossible; if If the dynamic range of the pressure field is too large, the power adjustment overshoot problem is likely to occur, leading to interference rebound at the same frequency.

[0041] The model is based on the Navier-Stokes equations for laminar flow of incompressible fluids, including the continuity equation and the momentum equation: in, This is the velocity field vector, corresponding to the velocity components in three-dimensional space; This is the dynamic viscosity coefficient, preset to 0.1; External volume force, set to 0.

[0042] The principle of the Benavis-Stokes equation mapping is as follows: the node distribution of ultra-dense networking is abstracted as an incompressible fluid field. Nodes with high fusion weights correspond to high-density, high-pressure fluid regions, which are the core areas for audio data acquisition, and it is necessary to ensure node activity and transmission power. Nodes with low fusion weights correspond to low-density, low-pressure fluid regions, which are redundant acquisition areas. Nodes can be put into sleep mode through the flow of the velocity field. From the perspective of the macroscopic field of fluid mechanics, the global optimal dynamic adjustment of node density and transmission power can be achieved, solving the core problems of redundant nodes and power waste in ultra-dense networking.

[0043] This dynamic viscosity coefficient The setting is based on: dynamic viscosity coefficient Controlling the convergence speed and smoothness of the fluid velocity field corresponds to the stability and convergence speed of node density adjustment in the network. This embodiment focuses on a scenario of dynamic node density adjustment, completing 200 sets of transient fluid simulations. Test results show that when… At this time, the velocity field can reach a stable convergence state within 5 iterations, corresponding to a convergence time of ≤0.5s for node density adjustment, while avoiding the oscillation problem of the velocity field, and the overshoot of node density adjustment is ≤5%; if Insufficient fluid viscosity leads to oscillations in the velocity field, and an overshoot of ≥22% in node density adjustment, resulting in frequent terminal sleep / wake-up. The fluid viscosity is too high, and the velocity field convergence time is ≥2s, making it unable to quickly adapt to the dynamic changes of the sound field and the channel.

[0044] The Navier-Stokes equations are discretized using the finite volume method, with forward Euler method used for time discretization, and the time step size is [not specified]. Spatial discretization employs a central difference scheme and is solved iteratively using the projection method, consisting of a prediction step and a correction step. The prediction step calculates the intermediate velocity field based on the current velocity field and pressure field. : The correction step involves solving the pressure Poisson equation to obtain the updated pressure field. : Based on the updated pressure field, the velocity field at the next time step is obtained by correcting the velocity field. : The stabilized velocity field is obtained through iterative solution. With pressure gradient .

[0045] Dynamic addition and deletion commands to convert the velocity field into node density: targeting the velocity field magnitude. For grid cells with a velocity greater than a preset velocity threshold, if their velocity direction points to a region with a lower density field and the fusion weight of the terminal within the grid cell is lower than the preset threshold, the cluster head node sends a sleep command to the corresponding terminal to reduce the node density within the cluster; if the grid cell's velocity direction points to a region with a higher density field and there are no active terminals in that region, the cluster head node sends a wake-up command to the sleep terminals in that region to increase the node density within the cluster.

[0046] The pressure gradient is converted into a dynamic adjustment step size for transmission power, tailored to the terminal. The formula for calculating the transmission power adjustment step size for the given grid cell is: in, The power adjustment factor is preset to 0.5 dBm per unit pressure gradient. For the sign function, when the pressure of the grid where the terminal is located... Less than maximum pressure When the sign is positive, the transmission power is increased; otherwise, when the sign is negative, the transmission power is decreased.

[0047] The power adjustment coefficient is set based on the need to match the large-scale fading characteristics of the wireless channel with the power adjustment sensitivity of audio transmission. This embodiment is based on the power control specifications of the Wi-Fi 7 protocol and was tested in an indoor multipath channel environment. The test results show that a power adjustment step size of 0.5 dBm can accurately adapt to the changing gradient of channel fading, minimizing co-channel interference while ensuring a received signal-to-noise ratio ≥25 dB. The power adjustment response speed is too slow, and it cannot quickly adapt to dynamic changes in the channel; if If the power adjustment step size is too large, it can easily cause the received signal-to-noise ratio to fluctuate by more than 10dB, affecting the stability of audio transmission.

[0048] As a preferred embodiment, refer to the appendix. Figure 5 After the cluster head node adjusts the cluster topology based on dynamic addition and deletion instructions according to node density, a small-world network reconstruction mechanism from complex network theory is introduced to optimize intra-cluster communication links. First, an undirected graph of the intra-cluster topology is constructed. ,in As an active terminal node within the cluster, For short-range communication edges, a short-range edge is established when the straight-line distance between two terminals is less than 2m, and the adjacency matrix is ​​used. The elements satisfy: When the terminal and The distance is ≤2m, otherwise .

[0049] Calculate the average path length of the current cluster topology. The calculation formula is: in, The number of nodes within the cluster. For nodes and The shortest path length between them is calculated using Dijkstra's algorithm, and the unit is the number of hops.

[0050] Calculate the clustering coefficient of the current cluster topology. First, calculate a single node. Clustering coefficient : in, For nodes The degree, For nodes The actual number of edges between adjacent nodes. The clustering coefficient of the entire topology. Clustering coefficients for all nodes The average value.

[0051] Set average path length threshold When the average path length When the distance exceeds a preset threshold, small-world network reconstruction is triggered, introducing long-range connections through a reconnection strategy. The establishment of long-range connections is based on the second-order Wasserstein distance between different clusters in the latent manifold space, specifically for two clusters. and Their potential spatial distributions are Gaussian distributions. and The formula for calculating the second-order Wasserstein distance is: in, This is the trace operation of the matrix. The smaller the Wasserstein distance, the higher the potential distribution similarity between the two clusters, and the stronger the correlation of the sound field features. Long-range connection edges are preferentially established for the corresponding clusters.

[0052] The reconnection strategy is as follows: select the first node with the smallest Wasserstein distance. For clusters, The default value is 2. From each pair of clusters, the node with the lowest degree within the cluster is selected, and a long-range connection edge is established. The maximum distance of the long-range connection edge does not exceed the maximum communication distance of the terminals within the cluster. After establishing the long-range connection edge, the average path length is recalculated until it is less than or equal to the preset threshold, at which point reconstruction stops.

[0053] When performing multi-hop relay transmission via long-range connections, a quantum heuristic optimization algorithm is introduced to optimize the transmission path. The relay node sequence is encoded as a quantum bit state, with each quantum bit represented by a binary tuple. It means that the conditions are met. ,in For a quantum bit to be in its ground state The probability, In an excited state The probability. For each node position in the path, use Each quantum bit is encoded as a node number. The number of binary bits corresponding to the number of nodes within the cluster; the quantum encoding of a complete relay path is a quantum chromosome with a length of [value missing]. , The maximum allowed number of relay hops is 8 hops by default.

[0054] Initial size is The quantum population, with each qubit initialized to a value of 1. It is in a uniform superposition state. The fitness function is constructed, and the calculation formula is: in, End-to-end latency includes transmission latency, queuing latency, and processing latency; The sum of the absolute values ​​of the transmission power adjustment steps for all relay nodes on the path; This is a weighting coefficient, preset to 0.6, used to balance the optimization objectives of latency and transmission power. Fitness The larger the value, the better the corresponding relay path.

[0055] The weighting coefficients are set based on the following: the core requirement for wireless audio transmission is low latency, while also considering optimized transmission power to reduce co-channel interference. This embodiment targets multi-hop relay transmission scenarios, and within a maximum relay range of 8 hops, [the weighting coefficients are applied to the following scenarios]. Traversal tests were performed in the range of 0.3 to 0.8. The test results show that when At this time, the average end-to-end latency is ≤10ms, meeting the ≤20ms latency requirement for real-time audio transmission, while the average optimization of transmission power reaches 37.2%, achieving an optimal balance between latency and power; if The end-to-end latency increases by ≥45%, failing to meet the requirements for real-time audio transmission; if The power optimization margin decreased by ≥28%, and the effect of suppressing co-channel interference deteriorated significantly.

[0056] The probability amplitude of a qubit is updated using a quantum rotation gate. The rotation gate update formula is as follows: in, The rotation angle is determined by the fitness difference between the current individual and the best individual, and its value ranges from [value missing]. The rotation direction is determined by the current qubit state and the corresponding qubit state of the optimal individual, causing the current individual to converge toward the direction of the optimal individual.

[0057] The updated quantum chromosome is subjected to a collapse measurement, based on the value of each qubit. Generate a random number in the range 0-1. If the random number is less than... The measurement result is 0 if the result is 0 otherwise, and 1 otherwise. The measured binary string is converted into a decimal node number to obtain the corresponding relay path. Invalid paths are discarded. The iterative process of fitness calculation, quantum rotation gate update and collapse measurement is repeated, with the number of iterations preset to 50. Finally, the relay path with the highest fitness is output as the optimal multi-hop relay path from the intra-cluster terminal to the cluster head node.

[0058] As a preferred embodiment, refer to the appendix. Figure 6 When the cluster head node performs deep fusion of the audio data aggregated within the cluster, a conditional generative adversarial network is introduced to compensate for missing audio frequency bands. First, the fused audio data... Perform a short-time Fourier transform to convert it into frequency domain features. ,in For frequency index, For frame indexing; the audio frequency band is divided into a low-frequency band (0-8kHz) and a high-frequency band (8kHz-24kHz), and the frequency domain features of the low-frequency band are extracted. As the main input to the generator.

[0059] Constructing condition vectors The multi-source data fusion weights and wireless channel state data are concatenated and mapped to a 256-dimensional conditional vector through a fully connected layer. Simultaneously, terminal location data is converted into location codes, serving as prior constraint inputs for the discriminator. The conditional generative adversarial network (GAN) includes a generator and a discriminator. The generator adopts a U-Net structure. The encoder consists of eight 2D convolutional layers with a kernel size of 4×4, a stride of 2, and the number of channels gradually increasing from 64 to 512. Each layer uses the Leaky ReLU activation function and instance normalization. The decoder consists of eight transposed convolutional layers with a kernel size of 4×4, a stride of 2, and the number of channels gradually decreasing from 512 to 1. Each layer uses the ReLU activation function and instance normalization. Skip connections are set between corresponding layers in the encoder and decoder. The generator takes low-frequency band features and a conditional vector as input and outputs reconstructed high-frequency band features. .

[0060] The discriminator employs a PatchGAN structure, consisting of five 2D convolutional layers with a kernel size of 4×4 and a stride of 2. The number of channels is increased from 64 to 512. Each layer uses the LeakyReLU activation function and instance normalization, ultimately outputting a feature map with a dimension of [30, 30]. Each value corresponds to the probability of a local block being true or false. The discriminator's input includes real or generated high-frequency features, a conditional vector, and a location prior encoding. The location prior constraint ensures that the generated high-frequency features conform to the acoustic attenuation characteristics corresponding to the spatial location of the sound source.

[0061] The total loss function of the conditional generative adversarial network is calculated using the following formula: Among them, combating losses Generate adversarial network loss for standard conditions: To generate the L1 distance loss between high-frequency features and the true high-frequency features: The location prior loss is used to constrain the amplitude attenuation characteristics of the generated high-frequency features to match the location prior. in, To generate the amplitude spectrum of high-frequency features, This is the theoretical high-frequency attenuation amplitude spectrum calculated based on the terminal location and the sound source location. The default value is 100. The default value is 10.

[0062] The weighting coefficients are set based on the following: the core optimization objective of conditional generative adversarial networks (GANs) is the accurate reconstruction of high-frequency details, while ensuring that the reconstructed high-frequency features conform to the acoustic attenuation law corresponding to the sound source location. This embodiment completes 500 sets of adversarial training tests for two typical audio signals: speech and music. The test results show that when… When the model is optimized, it can guide the direction of optimization, ensuring the waveform consistency between the generated high-frequency features and the real high-frequency features. The PESQ score of the reconstructed audio reaches 4.2, approaching the level of lossless audio. The model is prone to pattern collapse, generating high-frequency features with severe distortion, resulting in a PESQ score decrease of ≥1.2; if The model convergence speed decreased by ≥60%, and the number of training iterations doubled.

[0063] against The test results show that this weighting effectively constrains the spatial acoustic characteristics of the generated high-frequency features, achieving a 94.8% match between the reconstructed high-frequency attenuation characteristics of terminals at different locations and the free-field acoustic propagation theory, thus avoiding the problem of mismatch between the generated high-frequency details and the spatial location of the sound source; if When positional constraints fail, the spatial orientation of the reconstructed audio is severely distorted; if This will dominate the direction of model optimization, leading to a decrease in the reconstruction accuracy of high-frequency details.

[0064] The generator and discriminator are trained alternately until Nash equilibrium is reached. The weights of the trained model are then embedded into the NPU of the cluster head node. During online inference, the low-frequency features of the fused audio and the conditional vector are input into the generator to obtain the reconstructed high-frequency acoustic details. The low-frequency features and the reconstructed high-frequency features are then concatenated in the frequency domain to obtain the complete frequency domain features, which are then converted into a complete time-domain audio signal through inverse short-time Fourier transform.

[0065] Before the frequency-domain concatenated audio data is accessed by the upper-layer wireless LAN, rate-distortion theory and the mutual information maximization criterion are introduced for joint coding of the source and channel. Based on rate-distortion theory, the rate-distortion function of the audio source is: in, The original audio signal. The reconstructed signal after encoding and decoding. Mean squared error distortion is a measure of the distortion. for and Mutual information between them For a given distortion The minimum source coding rate.

[0066] The channel capacity function is: in, For channel bandwidth, To improve the signal-to-noise ratio, For channel transmission bit error rate, For a given bit error rate The maximum reliable channel coding rate.

[0067] Construct based on audio data distortion With wireless channel transmission bit error rate Joint cost function for independent variables: in, and These are the weighting coefficients. The default value is 1.0. The default value is 1.5. Set constraints: Total bandwidth constraint. ,in For source coding rate, For channel coding rate, Total bandwidth allocated to the upper-layer wireless LAN; while simultaneously satisfying , .

[0068] The weighting coefficients are set based on the following: the core requirement for wireless audio transmission is audio fidelity, while ensuring transmission reliability and avoiding audio stuttering caused by bit errors. This embodiment, based on rate-distortion theory and channel capacity formulas, calculates the weighting coefficients in a typical indoor Wi-Fi 7 channel environment. and Orthogonal tests were conducted in the range of 0.5 to 2.0. The test results show that when , At this time, an optimal balance between audio distortion and transmission error rate can be achieved. Under the constraint of total bandwidth, the STOI score of audio reconstruction reaches 98.2%, and the transmission error rate is ≤1e-6, meeting the quality requirements of broadcast-grade audio transmission; if If the transmission error rate increases by ≥2 orders of magnitude, audio disconnection and stuttering issues are likely to occur; if Excessive bandwidth usage in channel coding leads to a decrease in source coding bitrate, an increase in audio distortion of ≥15%, and a decrease in STOI score of ≥5%.

[0069] A multi-objective particle swarm optimization algorithm is used to find the Pareto optimal solution for the source coding rate and channel coding rate that minimizes the joint cost function, obtaining the Pareto front. Based on the real-time bandwidth allocation of the upper-layer wireless LAN, the optimal rate combination is selected from the Pareto front. Based on the Pareto optimal solution, adaptive modulation and hierarchical coding are applied to the frequency-domain concatenated audio data for transmission: LDAC coding is used for source coding, based on... Adaptive adjustment of coding rate; channel coding uses LDPC coding, based on Adaptive bit rate adjustment; modulation method based on signal-to-noise ratio and Adaptively select QPSK, 16QAM, 64QAM, or 256QAM.

[0070] As a preferred embodiment, during the hierarchical coding and transmission process based on the Pareto optimal solution, a clonal selection algorithm from the biological immune system is introduced to recover from network anomalies. Real-time monitoring of the joint cost function is also implemented. gradient change, gradient The sampling period is 10ms, and the gradient mutation amount is calculated. When the gradient mutation amount exceeds a preset threshold When an abnormal immune response is triggered, the threshold is... The default value is 0.5 / 10ms.

[0071] The gradient threshold is set based on the need to distinguish between normal fluctuations and abnormal faults in the network environment, to avoid false triggering of immune responses, and to ensure rapid self-healing in abnormal scenarios. This embodiment collects 100 hours of network operation data for typical scenarios such as normal channel fading, terminal mobility, sudden strong interference, and node offline, and statistically analyzes the gradient change distribution of the joint cost function. The results show that the maximum gradient mutation amount of the joint cost function is 0.3 / 10ms in normal scenarios, while the minimum gradient mutation amount is 0.6 / 10ms in abnormal scenarios. Therefore, setting the threshold to 0.5 / 10ms can achieve a 100% recognition rate in abnormal scenarios, while the false trigger rate is 0 in normal scenarios, thus balancing the sensitivity and specificity of anomaly detection. If the threshold is less than 0.3 / 10ms, the false trigger rate is ≥35%, leading to frequent network adjustments and affecting transmission stability; if the threshold is greater than 0.8 / 10ms, the anomaly identification rate decreases by ≥40%, making it impossible to respond quickly to network faults and extending the self-healing time by ≥3 times.

[0072] The current optimal multi-hop relay path and orthogonal short time slot scheduling strategy are mapped to antibodies. The antibody is encoded as a 128-bit binary string, with the first half encoding the node sequence of the relay path and the second half encoding the time slot offset and power fine-tuning of the time slot scheduling. The network environment under abnormal conditions, including channel state, interference level, and terminal online status, is mapped to antigens. The affinity between the antibody and antigen is defined as... Joint cost function The smaller the value, the higher the affinity, and the better the corresponding scheduling scheme.

[0073] The execution flow of the cloning selection algorithm is as follows: Initialize the antibody population, using the current optimal scheduling strategy as the initial antibody, and randomly generate 99 antibodies to form an initial population of size 100; calculate the affinity of each antibody in the population, substitute the scheduling strategy corresponding to the antibody into the current network environment, and calculate the joint cost function. The corresponding affinity was obtained; the top 20 antibodies with the highest affinity were selected as parent antibodies for cloning. The number of clones was proportional to the antibody affinity, and the total number of clones was 200. The formula for calculating the number of clones of a single parent antibody is as follows: in, For the first The number of clones of the parent antibody. The maximum number of clones is set to 20 by default. For the first Affinity of the parent antibody This is for rounding operations.

[0074] The cloned antibody undergoes high-frequency mutation operations. The mutation probability is inversely proportional to the antibody affinity. The formula for calculating the mutation probability of a single cloned antibody is as follows: in, The maximum mutation probability is preset to 0.5. This represents the maximum affinity of the current population. The mutation operation randomly flips the binary bits of the antibody, generating diverse scheduling schemes.

[0075] For the mutated antibodies, 10% of the antibodies are randomly selected for receptor editing, with some binary bits of the antibodies being randomly replaced to avoid the algorithm getting trapped in local optima. The affinity between the mutated and receptor-edited antibodies is calculated, and the top 100 antibodies with the highest affinity are selected to form a new population. The antibody with the highest affinity is retained as the current optimal solution. The above iterative process is repeated, with the number of iterations preset to 30, or when the joint cost function... The iteration terminates when the level recovers to below the pre-abnormal level, and the current optimal scheduling scheme is output.

[0076] The optimal scheduling scheme output is stored as a memory cell in the non-volatile memory of the cluster head node. When the same or similar abnormal scenarios occur later, the scheduling scheme in the memory cell is directly called to quickly restore the normal operation of the network and realize the adaptive evolution of ultra-dense networking.

[0077] The scope of protection of this invention is not limited to the above embodiments. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention. For example, the wireless audio terminal can be configured as any device with wireless communication and audio acquisition capabilities, such as a headset wireless microphone, a lavalier wireless microphone, or a conference terminal pickup module; the cluster head node can be configured as a dedicated wireless access gateway, rather than being elected from among the wireless audio terminals; the upper-layer wireless local area network can be replaced by a 5G NR small cell network, a private wireless communication system, or other wide-area wireless communication network; and the orthogonal short time slot transmission resources can be replaced by any orthogonally divisible wireless transmission resources, such as Bluetooth isochronous channel time slots or 5G time slot resources.

Claims

1. A method for ultra-dense networking and multi-source data fusion of wireless audio terminals, characterized in that, include: The wireless audio terminal, acting as a micro access node and terminal node, collects the temporal characteristics of the sound field and the wireless channel status information of its own pickup channel. Distributed time-slot cooperative scheduling allocates orthogonal short time-slot transmission resources to wireless audio terminals within the same sound field coverage area based on the spatial correlation of sound field characteristics. The Bayesian nonparametric clustering model performs unsupervised clustering on sound field data, wireless channel state data and terminal location data uploaded by multiple wireless audio terminals, and generates multi-source data fusion weights within the same cluster. The fused audio data is accessed through the cluster head node to the upper-layer wireless LAN. The cluster head node dynamically adjusts the node density and transmission power of the network within the cluster based on the weight of the multi-source data fusion.

2. The method according to claim 1, characterized in that, The spatial correlation calculation method for sound field features in distributed time-slot collaborative scheduling is as follows: Wireless audio terminals within the same sound field coverage area are constructed as a dynamic topology graph, with the terminal location as the node coordinate and the cross-correlation coefficient of the sound field temporal features as the edge weight. The dynamic topology graph is input into a graph convolutional network to extract high-dimensional spatial correlation features that include the geometric distribution of sound sources and acoustic attenuation characteristics. Based on these high-dimensional spatial correlation features, orthogonal short time slot transmission resources are pre-allocated.

3. The method according to claim 2, characterized in that, Based on the pre-allocation of orthogonal short time slot transmission resources, multi-agent deep reinforcement learning is introduced to eliminate time slot conflicts: Each wireless audio terminal is treated as an independent agent. The current high-dimensional spatial features and wireless channel state information are used as the state space. The time slot offset and transmission power fine-tuning are used as the action space. The joint reward function is to minimize co-channel interference penalty and maximize throughput. The optimal time slot scheduling strategy is output through a multi-agent near-end policy optimization algorithm.

4. The method according to claim 3, characterized in that, The Bayesian nonparametric clustering model is implemented using a Dirichlet process variational autoencoder: the sound field data, wireless channel state data, and terminal location data are concatenated into a multi-source heterogeneous tensor; The encoder network maps multi-source heterogeneous tensors to the latent manifold space. Nonparametric clustering is performed in the latent manifold space based on the Dirichlet process prior. The lower bound of evidence is maximized through variational inference, the number of clusters is automatically determined, and the fusion weights of multi-source data within each cluster are output.

5. The method according to claim 4, characterized in that, In the step of dynamically adjusting the node density and transmission power of the cluster network based on the weight of multi-source data fusion, the Navier-Stokes equations from computational fluid dynamics are introduced for modeling: The wireless audio terminals within the cluster are mapped as fluid micro-elements, and the multi-source data fusion weights are mapped as fluid density and pressure fields. The velocity field and pressure gradient of the node distribution are calculated by solving the discretized Navier-Stokes equations. The velocity field is converted into dynamic addition and deletion instructions for node density, and the pressure gradient is converted into dynamic adjustment step size of transmission power.

6. The method according to claim 5, characterized in that, After the cluster head node adjusts the cluster topology based on dynamic addition and deletion instructions of node density, a small-world network reconstruction mechanism from complex network theory is introduced to optimize the intra-cluster communication links: Calculate the average path length and clustering coefficient of the current cluster topology. If the average path length is greater than a preset threshold, introduce long-range connection edges through a reconnection strategy. The establishment of long-range connection edges is based on the Wasserstein distance between different clusters in the potential manifold space.

7. The method according to claim 6, characterized in that, When performing multi-hop relay transmission via long-range connections, a quantum heuristic optimization algorithm is introduced to optimize the transmission path: The relay node sequence is encoded into a quantum bit state, the quantum bit probability amplitude is updated through a quantum rotation gate, and the fitness function is used as the weighted sum of the end-to-end delay and the transmission power adjustment step size. The quantum state collapse measurement is performed, and the optimal multi-hop relay path from the cluster to the cluster head node is output.

8. The method according to claim 7, characterized in that, When the cluster head node performs deep fusion of the audio data aggregated within the cluster, a conditional generative adversarial network is introduced to compensate for missing audio frequency bands: The multi-source data fusion weights and wireless channel state data are used as the input of the conditional vector generator. The generator reconstructs the high-frequency acoustic details lost due to co-channel interference based on the conditional vector. The discriminator uses terminal location data as a priori constraint, and performs adversarial training until Nash equalization. It then performs frequency domain splicing of the high-frequency acoustic details output by the generator with the fused audio data.

9. The method according to claim 8, characterized in that, Before the frequency-domain stitched audio data is accessed into the upper-layer wireless LAN, rate-distortion theory and mutual information maximization criteria are introduced for joint source-channel coding: A joint cost function is constructed with audio data distortion and wireless channel transmission bit error rate as independent variables. Under the constraint of the total bandwidth of the upper-layer wireless LAN, Pareto optimal solutions for the source coding rate and channel coding rate that minimize the joint cost function are found. Based on the Pareto optimal solution, the frequency-domain spliced ​​audio data is adaptively modulated and hierarchically coded for transmission.

10. The method according to claim 9, characterized in that, During the hierarchical coding and transmission process based on the Pareto optimal solution, a clonal selection algorithm from the biological immune system is introduced to recover from network anomalies. Real-time monitoring of gradient changes in the joint cost function; triggering an abnormal immune response when the gradient mutation exceeds a preset threshold. The current optimal multi-hop relay path and orthogonal short time slot scheduling strategy are mapped to antibodies. Diverse scheduling schemes are generated through cloning, high-frequency mutation and receptor editing operations. The scheme that fastest inhibits the increase of the joint cost function is used as the memory cell to be stored in the cluster head node.