Multi-channel Audio Synchronization Method, Device, Equipment and Storage Medium for Conference Terminal Equipment

By collecting audio data streams of multiple terminal devices and encoding ambient audio information, combining local sensitive hashing algorithms and multi-agent reinforcement learning model, the problem of multi-channel audio synchronization is solved, and high-quality audio synchronization and user experience improvement is achieved.

CN119276843BActive Publication Date: 2025-05-27SHENZHEN YIYANG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411459365.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-05-27
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

In the conference scenario where multiple terminal devices participate, it is difficult for the prior art to achieve accurate synchronization of multiple audio, and problems such as audio loss, echo and crosstalk often occur, affecting the user's meeting experience and communication efficiency.

Method used

By collecting and encoding the audio data stream of each conference terminal device, embedding it into the audio data stream, time-frequency domain conversion and peak detection are performed, audio feature vectors are generated using a local sensitive hash algorithm, cross-correlation functions between terminal devices, determining time delay, building a two-layer cache structure and using Kalman filter to predict network state, optimizing communication resource allocation through multi-agent reinforcement learning model, and finally achieving global synchronous optimization through gradient descent algorithm.

Benefits of technology

It significantly improves the audio quality and user experience of multi-party remote meetings, enhances the robustness of audio synchronization, reduces the impact of network jitter and burst noise, and realizes accurate timeline adjustment and content alignment of multiple audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119276843B_ABST
    Figure CN119276843B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of multi-channel audio synchronization, and discloses a multi-channel audio synchronization method, device, equipment and storage medium for a conference terminal device. The method includes: collecting the audio data streams of each conference terminal device to obtain embedded audio data streams; performing time-frequency domain conversion and peak detection to obtain audio feature vectors; calculating the cross-correlation function between terminal devices to obtain a global time relationship graph; constructing a two-layer cache structure to adjust the residence time of audio data in the cache to obtain a buffering strategy; using the buffering strategy as a state to input into a multi-agent reinforcement learning model for communication resource allocation to obtain a resource allocation strategy; constructing a global synchronization optimization problem based on the resource allocation strategy, and solving for the optimal synchronization parameters through a gradient descent algorithm to perform time-axis adjustment and content alignment on the embedded audio data streams to obtain synchronized multi-channel audio, thereby improving the audio quality and user experience of multi-party remote conferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of multi-channel audio synchronization, and particularly to a multi-channel audio synchronization method, device, equipment, and storage medium for conference terminal devices. Background Art

[0002] In a conference scenario involving multiple terminal devices, due to the influence of network latency, device performance differences, and environmental factors, achieving precise synchronization of multi-channel audio still faces huge challenges. Traditional audio synchronization methods often rely on simple timestamp comparison or fixed buffering strategies, and it is difficult to adapt to complex and changeable network environments and dynamic conference scenarios.

[0003] These methods often have problems such as audio desynchronization, echo, and crosstalk when dealing with multi-terminal and long-duration conferences, seriously affecting the user's conference experience and communication efficiency. In addition, existing technologies also seem powerless in dealing with problems such as environmental noise, network jitter, and device heterogeneity, and it is difficult to ensure the audio synchronization quality in various complex situations. Summary of the Invention

[0004] This application provides a multi-channel audio synchronization method, device, equipment, and storage medium for conference terminal devices, thereby improving the audio quality and user experience of multi-party remote conferences.

[0005] In the first aspect of this application, a multi-channel audio synchronization method for conference terminal devices is provided. The multi-channel audio synchronization method for conference terminal devices includes:

[0006] Collect the audio data stream of each conference terminal device, and encode the environmental audio information into frequency domain features and embed them into the audio data stream to obtain an embedded audio data stream;

[0007] Perform time-frequency domain conversion and peak detection on the embedded audio data stream, and map the peak points to binary strings through the locality-sensitive hashing algorithm to obtain an audio feature vector;

[0008] Calculate the cross-correlation function between terminal devices based on the audio feature vector, determine the time delay through multi-scale analysis and median filtering, and obtain a global time relationship graph;

[0009] Construct a two-layer cache structure according to the global time relationship graph, predict the network state through a Kalman filter, and adjust the residence time of audio data in the cache to obtain a buffering strategy;

[0010] Use the buffering strategy as the state to input into a multi-agent reinforcement learning model, and optimize the communication resource allocation through a double deep Q network to obtain a resource allocation strategy;

[0011] Construct a global synchronization optimization problem based on the resource allocation strategy, solve for the optimal synchronization parameters through the gradient descent algorithm, perform time-axis adjustment and content alignment on the embedded audio data stream, and obtain the synchronized multi-channel audio.

[0012] The second aspect of this application provides a multi-channel audio synchronization device for a conference terminal device. The multi-channel audio synchronization device for a conference terminal device includes:

[0013] An acquisition module, configured to acquire the audio data stream of each conference terminal device, and encode the environmental audio information as frequency-domain features and embed them into the audio data stream to obtain an embedded audio data stream;

[0014] A mapping module, configured to perform time-frequency domain conversion and peak detection on the embedded audio data stream, and map the peak points to binary strings through the locality-sensitive hashing algorithm to obtain audio feature vectors;

[0015] A calculation module, configured to calculate the cross-correlation function between terminal devices based on the audio feature vectors, determine the time delay through multi-scale analysis and median filtering, and obtain a global time relationship graph;

[0016] A construction module, configured to construct a two-layer cache structure according to the global time relationship graph, predict the network state through a Kalman filter, and adjust the residence time of audio data in the cache to obtain a buffering strategy;

[0017] An allocation module, configured to use the buffering strategy as the state to input into a multi-agent reinforcement learning model, optimize the communication resource allocation through a double deep Q-network, and obtain a resource allocation strategy;

[0018] A synchronization module, configured to construct a global synchronization optimization problem based on the resource allocation strategy, solve for the optimal synchronization parameters through the gradient descent algorithm, perform time-axis adjustment and content alignment on the embedded audio data stream, and obtain the synchronized multi-channel audio.

[0019] The third aspect of this application provides an electronic device, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory so that the electronic device executes the above-mentioned multi-channel audio synchronization method for a conference terminal device.

[0020] The fourth aspect of this application provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the above-mentioned multi-channel audio synchronization method for a conference terminal device.

[0021] Compared with the prior art, the present application has the following beneficial effects: By encoding environmental audio information into frequency-domain features and embedding them into the audio data stream, the present invention enhances the robustness of audio synchronization and can better cope with background noise interference in different conference environments. The local sensitive hashing algorithm is used to map peak points into binary strings, greatly reducing the storage and transmission overhead of audio features while maintaining the distinctiveness of the features. The introduction of multi-scale analysis and median filtering techniques improves the accuracy and stability of time delay estimation and effectively overcomes the influence of network jitter and burst noise. A two-layer cache structure is constructed and combined with a Kalman filter to predict the network state, realizing dynamic adjustment of audio data caching and effectively balancing synchronization accuracy and system delay. The multi-agent reinforcement learning model and double deep Q-network are used to optimize communication resource allocation, enabling the system to adaptively cope with complex and changing network environments and improving resource utilization efficiency. By constructing a global synchronization optimization problem and using the gradient descent algorithm to solve it, precise time axis adjustment and content alignment of multiple channels of audio are achieved, ensuring the audio synchronization to the greatest extent. The present invention significantly improves the audio quality and user experience of multi-party remote conferences. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0023] The structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have technical essence. Any modification of the structure, change of the proportional relationship, or adjustment of the size should still fall within the scope that can be covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.

[0024] Figure 1 It is a schematic flowchart of the method for synchronizing multiple channels of audio of the conference terminal device provided by the embodiment of the present invention;

[0025] Figure 2 It is a schematic block diagram of the structure of the device for synchronizing multiple channels of audio of the conference terminal device provided by the embodiment of the present invention;

[0026] Figure 3 It is a schematic block diagram of the structure of the electronic device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.

[0029] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0030] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Please refer to Figure 1 , an embodiment of the method for multi-channel audio synchronization of a conference terminal device in an embodiment of this application includes:

[0031] Step 100: Collect the audio data streams of each conference terminal device, and encode the environmental audio information into frequency-domain features and embed them into the audio data streams to obtain embedded audio data streams;

[0032] It can be understood that the execution subject of this application can be a multi-channel audio synchronization device for a conference terminal device, or a terminal or a server, and specifically, it is not limited here. In the embodiments of this application, the server is used as the execution subject for illustration.

[0033] Specifically, the environment around each conference terminal device is acoustically sampled to obtain real environmental background sound information. The collected sound signal is preprocessed by a high-pass filter to filter out low-frequency interference and retain the high-frequency components in the environment, resulting in environmental background audio. The processed environmental background audio is spectrally analyzed, and by analyzing the energy aggregation of different frequency components, frequency components that are continuously stable over time and have distinct characteristics are selected. These frequency components represent the significant features in the environmental sound and can describe the characteristics of the environment, thus obtaining characteristic environmental audio. The characteristic environmental audio is subjected to discrete cosine transform to convert the characteristic environmental audio from a time-domain signal to a frequency-domain space, obtaining frequency-domain environmental audio features. At the same time, the audio input signal of the conference terminal device is sampled and quantized, and the collected analog audio signal is converted into a digital signal by an analog-to-digital converter to obtain an audio data stream for processing. The audio data stream is frame-divided, dividing the continuous audio data into audio frames of a fixed length to obtain frame-divided audio data. The frame-divided audio data is subjected to fast Fourier transform to convert the frame-divided audio data in the time domain to the frequency-domain space, obtaining frequency-domain audio data. In the frequency-domain space, based on the frequency-domain environmental audio features, the mid- and high-frequency bands of the frequency-domain audio data are modulated to embed the environmental audio information into the audio data stream, obtaining embedded frequency-domain audio data. The embedded frequency-domain audio data is subjected to inverse fast Fourier transform to convert the embedded frequency-domain audio data back to the time-domain space, obtaining an embedded audio data stream.

[0034] Step 200: Perform time-frequency domain conversion and peak detection on the embedded audio data stream, and map the peak points to binary strings through the locality-sensitive hashing algorithm to obtain an audio feature vector;

[0035] Specifically, the embedded audio data stream is segmented, and the continuous embedded audio data stream is divided into time segments of a fixed length to obtain an audio segment sequence, which facilitates feature extraction and synchronization operations for each segment. Each audio segment in the audio segment sequence is transformed to convert the time-domain signal into a time-frequency representation, obtaining a time-frequency spectrogram that reflects the energy distribution of each frequency component of the audio signal at different time points, helping to identify the characteristics of the signal at different times and frequencies. The energy of the time-frequency spectrogram is calculated to quantify the energy distribution of each time-frequency unit. The energy value of the time-frequency unit is calculated by an integration method to obtain an energy distribution matrix. The energy distribution matrix is used to represent the signal intensity at each time and frequency combination in the time-frequency spectrogram. Based on the energy distribution matrix, peak detection is performed on each time-frequency unit to obtain a set of peak points in the time-frequency spectrogram, identifying the important frequency components with representativeness in the time-frequency spectrogram. Coordinate quantization is performed on the set of peak points to map the time-frequency coordinates to a discrete grid, obtaining quantized peak points. Through coordinate quantization, the representation of the time-frequency coordinates is simplified, making it easier to compare and match the peak points between different audio segments. A feature descriptor is constructed based on the quantized peak points. By calculating the time-frequency difference between adjacent peak points, a difference vector is obtained. The difference vector is used to describe the relationship between adjacent peak points. The difference vector is input into a locality-sensitive hashing function, and a hash value is generated through random projection and threshold comparison to obtain a binary feature string. Locality-sensitive hashing can maintain the distance relationship between similar data in a high-dimensional space, mapping similar features to similar hash values and simplifying the matching process between audio segments. Grouping and merging operations are performed on the binary feature string. Through bit operations, multiple feature strings are combined into a vector of a fixed length to obtain an audio feature vector.

[0036] Step 300: Calculate the cross-correlation function between terminal devices based on the audio feature vector, determine the time delay through multi-scale analysis and median filtering, and obtain a global time relationship graph;

[0037] It should be noted that the audio feature vectors are regarded as time series data and decomposed by time series. By resampling the audio feature vectors at different time scales, multiple subsequences with different resolutions are generated. These subsequences represent the information of the audio feature vectors at multiple preset time scales, forming a multi-scale feature sequence. Based on the multi-scale feature sequence, cross-correlation calculations are performed on the feature sequences of each pair of conference terminal devices. A fixed-size time window is selected and slid simultaneously on the multi-scale feature sequences of the two conference terminal devices. At each position, the Pearson correlation coefficient of the data within the time window is calculated, and the correlation coefficient values at different time offsets are recorded to form a cross-correlation function curve. The cross-correlation function curve represents the similarity of the feature sequences between the two devices at different time offsets, and the peak points on the curve represent the optimal synchronization time delay between the two devices. Peak detection is performed on the cross-correlation function curve to search for local maximum points on the cross-correlation function curve. For each local maximum point, two points on its left and right are taken, a total of five points of data, and a quadratic curve is fitted using these points. Through quadratic curve fitting, the vertex position of the curve is calculated, and the vertex position is taken as the exact position of the peak to obtain candidate values of the initial time delay. A multi-scale time delay matrix is constructed based on the candidate values of the initial time delay. Each row of the multi-scale time delay matrix represents a time scale, and each column represents the delay between a pair of devices. By analyzing different time scales, various time-varying information contained in the audio feature vectors is captured. On this basis, weighted averaging is performed on each column in the multi-scale time delay matrix to obtain a fused time delay value, which represents the comprehensive delay estimate at multiple time scales. Median filtering is performed on the fused time delay value. An odd-sized sliding window is set and slid on the time delay sequence. Each time it slides, all the values within the sliding window are taken, sorted by size, and the middle value is selected as the output at the current position to obtain a smooth time delay sequence. Median filtering effectively suppresses the noise in the sequence and removes abnormal delays caused by network jitter or sudden interference, making the final time delay sequence smoother and more stable. Based on the smooth time delay sequence, a time delay graph is constructed. Each conference terminal device is regarded as a node in the time delay graph, and the time delay between the devices is used as the weight of the edge. Then, the Kruskal algorithm is used to construct a minimum spanning tree. The goal of the Kruskal algorithm is to select a set of edges with the minimum total delay to obtain the topological structure between the devices, which represents the mutual dependence of each terminal device in time synchronization to ensure the global minimum delay. When globally optimizing the constructed device topological structure, a relative time coordinate system with a reference device as the origin is established. The time delay between each pair of adjacent devices is regarded as a constraint condition, and a least squares optimization problem is constructed.By solving the least squares optimization problem, the position of each conference terminal device in the unified time coordinate system is obtained, and the consistent time coordinate is obtained. Based on the consistent time coordinate, the positions of each conference terminal device in the unified time coordinate system are mapped onto a two-dimensional plane, where the distance between devices reflects their time synchronization relationship. The force-directed algorithm is used to adjust the positions of these nodes to obtain the global time relationship graph. The force-directed algorithm simulates the mutual forces between nodes, making the nodes present a more balanced layout that conforms to the time synchronization relationship on the plane. The finally obtained global time relationship graph intuitively shows the synchronization situation among various conference terminal devices.

[0038] Step 400: Construct a two-layer cache structure according to the global time relationship graph, predict the network state through the Kalman filter, and adjust the residence time of audio data in the cache to obtain the buffering strategy;

[0039] Specifically, based on the global time relationship graph, an appropriate amount of cache space is allocated to each conference terminal device, and this cache space is divided into two layers: a fast cache and a deep cache. The fast cache is used to store the recently received audio data, mainly aiming at low latency, so that the latest audio data can be quickly responded to and enter the processing flow, while the deep cache is used to store the processed historical audio data, providing buffering for audio data with a longer time span. Through the two-layer cache structure, the real-time performance and stability of the audio stream are more effectively balanced, and the synchronization effect of the audio data is improved. Timestamp analysis is performed on the data in the two-layer cache structure to obtain the time pattern of the arrival of the audio data. By calculating the time interval between adjacent audio frames, the time series of the arrival of the audio data is obtained, and the state space model of the Kalman filter is constructed. In this model, the arrival time of the audio data is regarded as the observed value, while the network transmission delay is used as the state variable, and the time-varying characteristics and delay fluctuations in the network transmission are described through this model. The Kalman filter is a tool for state estimation of time-varying systems. By dynamically adjusting the observed values, it makes the best estimate of the system state and is suitable for predicting network delays. In order to enable the Kalman filter to effectively predict the network state, the parameters of the network state prediction model are initialized. Through the analysis of historical audio data, the initial state estimate and the error covariance matrix are calculated to obtain the initial parameters of the Kalman filter. On this basis, based on the initial parameters of the Kalman filter, the state prediction and update of the newly arrived audio data are performed. Through the time update equation, the network state at the next moment is predicted, and then the measurement update equation is used to correct the prediction result to obtain the optimal estimate of the current network state. Through the cycle process of prediction and correction, the Kalman filter makes the estimation of the network state have good accuracy and real-time performance. Based on the optimal estimate of the current network state, the ideal residence time of the audio data in the cache is calculated. By comparing the predicted network transmission delay with the target playback delay, the target value for cache adjustment is obtained. If the network delay is large, the residence time of the audio data in the cache needs to be extended to avoid data incoherence during playback; if the network delay is small, the cache time is reduced to reduce the impact of the delay. Through the comparison between the network state predicted by the Kalman filter and the target playback requirements, the target value for cache adjustment is obtained. Based on the cache adjustment target value, the data in the two-layer cache structure is reorganized to achieve the optimal cache distribution. By adjusting the data allocation ratio between the fast cache and the deep cache, the buffering time of the audio data is finely controlled. For the audio data that arrives ahead of time, its processing is appropriately delayed to match the playback rhythm; for the audio data that arrives late, the processing is accelerated to make it as synchronous as possible with the audio data streams of other devices. Through the flexible allocation of the fast cache and the deep cache, the synchronization of the audio data in time is achieved, and the playback delay difference between different devices is reduced, thereby improving the overall synchronization quality of the conference audio.Evaluate the performance of the caching strategy to verify its effectiveness and adaptability under the current network conditions. By calculating metrics such as buffer utilization, packet loss rate, and audio quality, obtain the performance metrics of the caching strategy. Buffer utilization reflects the usage of caching resources. If the utilization is too high, it will lead to excessive latency, while if the utilization is too low, it cannot fully cope with network fluctuations; the packet loss rate is an important metric to measure the reliability of network transmission. If the packet loss rate is high, it indicates that the redundancy of the cache needs to be increased to reduce the impact of packet loss; the audio quality metric is an important standard directly reflecting the user experience, including audio continuity, clarity, etc. Based on these performance metrics, make necessary adjustments to the current caching strategy to ensure that the caching strategy can adapt to the current network conditions and finally obtain the optimal cache management solution.

[0040] Step 500: Take the caching strategy as the state and input it into the multi-agent reinforcement learning model. Optimize the communication resource allocation through the double deep Q-network to obtain the resource allocation strategy;

[0041] Specifically, feature extraction is performed on the caching strategy to extract a buffer state feature vector, including information such as cache utilization rate, latency status of the audio stream, and fluctuations in the network status. Based on the buffer state feature vector, a reinforcement learning environment is constructed, where each conference terminal device is regarded as an independent agent, and the environmental state of each agent is represented by its buffer state feature vector. At the same time, the action space is set to the selectable communication frequency bands and transmission powers, thus constituting a multi-agent learning environment. Each agent selects an appropriate action based on its own state to optimize its communication and resource utilization. In the multi-agent learning environment, the design of the reward function is carried out. The role of the reward function is to guide each agent to make a choice that is most beneficial to the overall goal of the system. The design of the reward function comprehensively considers three factors: audio synchronization degree, transmission latency, and channel utilization rate. The audio synchronization degree is used to measure the quality of audio data synchronization among various conference terminal devices, the transmission latency reflects the time delay of data transmission from one device to another, and the channel utilization rate is a measure of the efficiency of communication resource usage. By combining these factors in the form of a weighted sum, an optimization objective for reinforcement learning is constructed. Based on the optimization objective of reinforcement learning, a double deep Q-network model is constructed. The double deep Q-network includes an online network and a target network, and these two networks have the same structure, consisting of multiple fully connected layers, with the input being the current state vector and the output being the Q-values of each possible action. In the initialization stage, all parameters of the double deep Q-network are randomly initialized, and the structure of the network enables it to effectively map complex environmental states, providing an optimal action selection basis for each agent. To enable the double deep Q-network to effectively learn, an experience replay buffer is generated through random sampling. This buffer is used to store the experience data of each agent's interaction with the environment, including information such as state, action, reward, and next state. The existence of the experience pool enables the network to be trained from past experiences, avoiding the problem of unstable training caused by data correlation. Based on the experience pool, the double deep Q-network is trained. By randomly extracting a batch of experience data from the experience pool, the mean squared error between the target Q-value and the current Q-value is calculated, and this error is used as the loss function. Then, the parameters of the online network are updated using the backpropagation algorithm. Through multiple iterations and training, the online network gradually learns how to make the best action selection from the current environmental state, obtaining a trained double deep Q-network. Based on the trained double deep Q-network, this network is used to select actions for each agent, achieving a balance between exploration and exploitation through the ε-greedy strategy. The core of the ε-greedy strategy is to select a random action with a certain probability (exploration), and select the action with the largest current Q-value with the remaining probability (exploitation). This strategy enables new optimal strategies to be found during the training process and the learned knowledge to be fully utilized, thereby selecting the optimal communication frequency band and transmission power for each agent to form an initial resource allocation strategy.Globally coordinate the initial allocation strategy to ensure that there are no conflicts in resource allocation among all agents within the system. For example, multiple agents may select the same communication frequency band or too high a transmission power, resulting in channel conflicts or power overload problems. To solve the resource conflicts, a centralized resource allocator is constructed to collect the state information and action selections of all agents, and based on this information, coordinate and adjust the allocation strategies of each agent. By reasonably adjusting the initial allocation strategies of each agent, resource conflicts are eliminated, and a globally optimal resource allocation strategy is obtained.

[0042] Step 600: Construct a global synchronization optimization problem based on the resource allocation strategy, solve for the optimal synchronization parameters through the gradient descent algorithm, perform time-axis adjustment and content alignment on the embedded audio data stream, and obtain the synchronized multi-channel audio.

[0043] Specifically, based on the resource allocation strategy, the communication parameters of each conference terminal device are quantified, and communication parameters such as frequency band selection and transmission power are converted into numerical representations to construct a communication resource allocation matrix. The communication resource allocation matrix can intuitively reflect the communication resources allocated to each device. Feature extraction is performed on the communication resource allocation matrix. By calculating the channel capacity and transmission delay between conference terminal devices, a network performance feature vector is obtained, which is used to characterize the performance of the network state, including transmission bandwidth, delay, and communication stability, etc. Based on the network performance feature vector and the global time relationship graph, a global synchronization optimization objective function is constructed. The global synchronization optimization objective function takes the time synchronization error and network resource utilization rate as optimization indicators, where the time synchronization error measures the synchronization accuracy of audio data between devices, and the network resource utilization rate reflects the usage efficiency of communication resources. To ensure that the synchronization effect meets the established requirements, constraint conditions are set for the global synchronization optimization objective function, and a constrained optimization problem is constructed by limiting the maximum allowable synchronization error and the minimum necessary transmission rate. The maximum allowable synchronization error ensures that the synchronization quality of audio between terminals does not exceed the acceptable range, while the minimum necessary transmission rate ensures that the communication bandwidth is sufficient to support the stable transmission of audio data. Based on the constrained optimization problem, a Lagrangian function is constructed. By introducing Lagrange multipliers, the constraint conditions are incorporated into the global synchronization optimization objective function, and the original constrained problem is transformed into an unconstrained optimization problem. The advantage of the unconstrained optimization problem is that it can be solved by classical optimization algorithms, such as the gradient descent method, to obtain the optimal synchronization parameters. By iteratively calculating the gradient of the Lagrangian function and gradually updating the optimization variables, the optimal solution of the global synchronization optimization objective function is converged to obtain the optimal synchronization parameters. The gradient descent method adjusts the error of each iteration, making the synchronization error gradually decrease and finally reaching the optimal synchronization state. Based on the obtained optimal synchronization parameters, the time axis of the embedded audio data stream is adjusted. By performing interpolation operations on audio frames, the audio time axes of each conference terminal device are aligned to a unified reference time axis to obtain a time-aligned audio stream. The interpolation operation is achieved by inserting appropriate sampling points between adjacent audio frames. The interpolated audio stream eliminates the time differences between different devices, making the audio data of each terminal device synchronized on the unified time axis. Content alignment is performed on the time-aligned audio stream to ensure that the content of the audio is consistent among all devices. By analyzing the similarity of audio feature vectors, duplicate or missing audio segments in the audio stream are identified. Through the similarity analysis of audio feature vectors, the duplicate or missing segments are accurately located and corresponding measures are taken for processing. For duplicate audio segments, the redundant parts are deleted to simplify the audio stream; for missing audio segments, they are filled in by interpolation or using the audio data of adjacent devices to ensure the integrity and consistency of the audio stream. Through the dual alignment of the time axis and content, finally, synchronized multi-channel audio is obtained.

[0044] In the embodiments of the present application, by encoding environmental audio information into frequency-domain features and embedding them into the audio data stream, the present invention enhances the robustness of audio synchronization and can better cope with background noise interference in different conference environments. The local sensitive hashing algorithm is used to map peak points into binary strings, greatly reducing the storage and transmission overhead of audio features while maintaining the distinctiveness of the features. The introduction of multi-scale analysis and median filtering techniques improves the accuracy and stability of time delay estimation and effectively overcomes the influence of network jitter and burst noise. A two-layer cache structure is constructed and combined with a Kalman filter to predict the network state, realizing the dynamic adjustment of audio data caching and effectively balancing the synchronization accuracy and system delay. The multi-agent reinforcement learning model and double deep Q-network are used to optimize the communication resource allocation, enabling the system to adaptively cope with complex and changing network environments and improving the resource utilization efficiency. By constructing a global synchronization optimization problem and using the gradient descent algorithm to solve it, the precise time axis adjustment and content alignment of multiple channels of audio are realized, ensuring the audio synchronization to the greatest extent. The present invention significantly improves the audio quality and user experience of multi-party remote conferences.

[0045] In a specific embodiment, the process of executing step 100 may specifically include the following steps:

[0046] Collect sound in the environment of each conference terminal device, preprocess the collected sound signal through a high-pass filter to obtain environmental background audio, and perform spectrum analysis on the environmental background audio. By calculating and selecting the frequency components that are continuously stable and have characteristics through energy aggregation, characteristic environmental audio is obtained;

[0047] Perform discrete cosine transform on the characteristic environmental audio to convert the characteristic environmental audio into the frequency-domain space to obtain frequency-domain environmental audio features, and sample and quantize the audio input of the conference terminal device. Convert the analog signal into a digital signal through an analog-to-digital converter to obtain an audio data stream;

[0048] Perform frame splitting on the audio data stream, divide the continuous audio data into audio frames of a fixed length to obtain frame-split audio data, and perform fast Fourier transform on the frame-split audio data to convert the frame-split audio data into the frequency-domain space to obtain frequency-domain audio data;

[0049] Based on the frequency-domain environmental audio features, modulate the middle and high frequency bands of the frequency-domain audio data, embed the environmental audio information into the audio data stream to obtain embedded frequency-domain audio data, and perform inverse fast Fourier transform on the embedded frequency-domain audio data to convert the embedded frequency-domain audio data back into the time-domain space to obtain an embedded audio data stream.

[0050] Specifically, the environment of each conference terminal device is acoustically sampled, and the built-in microphone of the device or other external audio sensors are used to capture the sound signals in the environment. The environmental sound signals contain audio information from many different sources, including the ambient noise of the device, background music, speech, etc. To extract useful environmental features from them, the collected sound signals are preprocessed, and a high-pass filter is used to eliminate low-frequency noise. The main function of the high-pass filter is to remove low-frequency interference (such as power hum, wind noise, etc.) and retain the part with higher energy characteristics in the mid- and high-frequency bands, which can be specifically expressed as:

[0051] ;

[0052] Among them, is the output signal after being processed by the high-pass filter, is the original environmental sound signal, is the attenuation coefficient of the filter, which is used to control the degree of removal of low-frequency components. After preprocessing by the high-pass filter, the environmental background audio is obtained. The spectrum analysis is performed on the environmental background audio after being processed by the high-pass filter to extract the main frequency components in the signal. By performing spectrum analysis on the audio signal, the energy distribution at each frequency is quantified. The energy aggregation degree calculation is used to determine which frequency components are continuously stable in time and have significant characteristics. Calculate the energy

[0053] ;

[0054] Among them, represents the average energy at frequency , is the analysis time window, is the frequency at time the spectral amplitude on. In this way, the frequency components with significant characteristics in energy and stable in time are selected as the characteristic environmental audio. The characteristic environmental audio contains important sound characteristics in the device environment, such as the frequency components of background noise. The characteristic environmental audio is further processed, and the discrete cosine transform is used to convert it into the frequency domain space to obtain the frequency domain environmental audio characteristics. The role of the discrete cosine transform is to convert the time domain signal into a frequency domain representation, and retain the core information of the signal by removing high-frequency noise and compressing energy. The formula of the discrete cosine transform is:

[0055] ;

[0056] Among them, is the transformed frequency domain coefficient, is the input signal in the time domain, is the total number of samples. Through discrete cosine transform, the time-domain information of the characteristic ambient audio is converted into frequency-domain coefficients , and these frequency-domain coefficients can compactly represent the ambient characteristics. Meanwhile, sample and quantize the audio input of the conference terminal device, convert the analog signal into a digital signal through an analog-to-digital converter, and obtain an audio data stream. The audio data stream is the result of digitizing the input analog audio and is represented as a series of discrete audio samples. Perform frame segmentation on the audio data stream, divide the continuous audio data into multiple audio frames of a fixed length, and each frame contains several consecutive audio sampling points. Through frame segmentation, local analysis and processing of the audio data are carried out to improve the real-time performance of processing. Perform a fast Fourier transform on the framed audio data to convert the audio data from the time domain to the frequency-domain space and obtain frequency-domain audio data. The frequency-domain audio data contains the amplitude and phase information of the audio signal at each frequency. Based on the frequency-domain ambient audio characteristics, modulate the frequency-domain representation of the audio data stream to embed the ambient audio information into the audio data stream. Select the mid-high frequency bands in the frequency-domain audio data for the modulation operation. These bands are usually the parts where the energy of the audio signal is relatively concentrated and are also the parts where the human ear is most sensitive to the audio quality. By modulating the characteristic information of the ambient audio onto these bands, the embedding of audio characteristics is effectively achieved while ensuring the naturalness and coherence of the original audio signal. For example, for the characteristic frequency components in the frequency domain , perform amplitude modulation with the ambient characteristic coefficient to obtain the embedded frequency-domain audio data:

[0057] ;

[0058] where is the amplitude of the original frequency-domain audio data at frequency , is a factor controlling the embedding intensity, is the frequency-domain coefficient of the characteristic ambient audio. Through modulation, the ambient audio characteristics are embedded into the original audio signal, enabling the audio stream to carry ambient information. Perform an inverse fast Fourier transform on the embedded frequency-domain audio data to convert it back from the frequency domain to the time domain and obtain the embedded audio data stream.

[0059] In a specific embodiment, the process of executing step 200 may specifically include the following steps:

[0060] Perform segmentation processing on the embedded audio data stream, divide the embedded audio data stream into time segments of a fixed length to obtain a sequence of audio segments, and perform a transformation on each audio segment in the sequence of audio segments to convert the time-domain signal into a time-frequency representation and obtain a spectrogram;

[0061] Calculate the energy of the time-frequency spectrogram, calculate the energy value of each time-frequency unit through the integral method, obtain the energy distribution matrix, and based on the energy distribution matrix, perform peak detection on each time-frequency unit to obtain the peak point set;

[0062] Quantify the coordinates of the peak point set, map the time-frequency coordinates to a discrete grid to obtain the quantified peak points, and based on the quantified peak points, construct a feature descriptor. By calculating the time-frequency difference between adjacent peak points, obtain the difference vector;

[0063] Input the difference vector into the locality-sensitive hashing function, generate a hash value through random projection and threshold comparison to obtain a binary feature string, and group and merge the binary feature strings. Combine multiple feature strings into a fixed-length vector through bit operations to obtain the audio feature vector.

[0064] Specifically, divide the embedded audio data stream into time segments of a fixed length to obtain an audio segment sequence. Let the audio data stream be , segment it to obtain time segments of length :

[0065] ;

[0066] Among them, represents the th audio segment, is the length of each segment, is the time variable. Through the segmentation method, the entire audio data stream is divided into segments, and each segment contains a part of the audio information. Perform a transformation on each audio segment to convert it from a time-domain signal to a time-frequency representation to obtain a time-frequency spectrogram. Through the short-time Fourier transform, provide the local features of the signal simultaneously in time and frequency:

[0067] ;

[0068] Among them, represents the time-frequency representation of the th audio segment at frequency and time , is the input signal, is a window function, usually a Hanning window or a Gaussian window, used to limit the local range of the signal in time, represents the imaginary unit. This conversion reveals the time characteristics and frequency characteristics of the time-domain signal simultaneously to form a time-frequency spectrogram. Calculate the energy of the time-frequency spectrogram, calculate the energy value of each time-frequency unit through the integral method to obtain the energy distribution matrix. The purpose of energy calculation is to determine the intensity of the audio signal at different times and frequencies. The specific formula is:

[0069] ;

[0070] wherein, represents the energy at time and frequency , and represents the squared amplitude of the spectrogram at this position. In this way, the energy distribution matrix is obtained, which describes the energy situation of the signal on each small unit in the time-frequency space. Based on the energy distribution matrix, peak detection is performed to determine at which times and frequencies the audio signal has significant features. By comparing each time-frequency unit, local energy peaks are found, and a set of peak points is obtained, which represents the important frequency components and events in the signal. Coordinate quantization is performed on the set of peak points, and its time-frequency coordinates are mapped onto a discrete grid to obtain quantized peak points. Suppose the coordinates of a peak point are , by quantizing the frequency and time into discrete grid coordinates respectively:

[0071] ;

[0072] wherein, and are the quantized frequency and time coordinates, and are the quantization steps of frequency and time respectively, and round represents the rounding operation. The purpose of quantization is to map the continuous time-frequency coordinates onto discrete grid points, reduce the data volume and improve the feature matching ability. Based on the quantized peak points, a feature descriptor is constructed. By calculating the time-frequency difference between adjacent peak points, a difference vector is obtained. The difference vector is used to describe the relationship between adjacent important events in the signal, so as to capture the global structural features of the signal. Suppose the coordinates of two adjacent peak points are and , then their difference vector is expressed as:

[0073] ;

[0074] The difference vector describes the change pattern of the audio signal in the time-frequency space and is important information for feature matching and synchronization. To improve the operability and compactness of audio features, the difference vector is input into a locality-sensitive hashing function, and a hash value is generated through random projection and threshold comparison to obtain a binary feature string. The role of locality-sensitive hashing is to reduce the dimension of high-dimensional data while maintaining the similarity between similar data, and accelerate similarity search. By taking the inner product of the difference vector with a set of random vectors and comparing with a threshold to generate a binary hash value:

[0075] ;

[0076] Wherein, represents the generated binary hash value, is a random vector, and · represents the inner product operation. In this way, the difference vector is mapped to a binary string of a fixed length. The obtained binary feature strings are grouped and combined, and multiple feature strings are combined into a vector of a fixed length through bit operations to obtain an audio feature vector.

[0077] In a specific embodiment, the process of executing step 300 may specifically include the following steps:

[0078] Perform time series decomposition on the audio feature vector. By resampling the audio feature vector according to different time scales, multiple subsequences with different resolutions are generated. Each subsequence contains the information of the audio feature vector at a preset time scale, and a multi-scale feature sequence is obtained;

[0079] Based on the multi-scale feature sequence, perform cross-correlation calculation on the feature sequences of each pair of conference terminal devices. Specifically: select a time window of a fixed size, slide the time window simultaneously on the multi-scale feature sequences of the two conference terminal devices, calculate the Pearson correlation coefficient for the data within the time window, record the correlation coefficient values at different time offsets, and form a cross-correlation function curve;

[0080] Perform peak detection on the cross-correlation function curve. By searching for local maximum points on the cross-correlation function curve, and taking two points on the left and right of each local maximum point, use the data of these five points to fit a quadratic curve, and calculate the vertex position of the quadratic curve as the exact peak position to obtain an initial time delay candidate value;

[0081] Construct a multi-scale time delay matrix based on the initial time delay candidate values. Organize the time delay candidate values obtained under different time scales into a matrix form. Each row of the multi-scale time delay matrix represents a time scale, and each column represents the delay between a pair of devices. Then, perform weighted averaging on each column in the multi-scale time delay matrix to obtain a fused time delay value;

[0082] Perform median filtering on the fused time delay value. Set a sliding window of an odd size, slide the sliding window on the time delay sequence, and each time take all the values within the sliding window, sort them by size, and select the middle value as the output at the current position to obtain a smoothed time delay sequence;

[0083] Construct a time-delay graph based on the smoothed time-delay sequence. Consider each conference terminal device as a node in the graph, and the time delay between conference terminal devices as the weight of the edge. Then use the Kruskal algorithm to construct a minimum spanning tree, select the edge set with the minimum total delay, and obtain the topological structure between devices;

[0084] Globally optimize the topological structure between devices. Establish a relative time coordinate system with a reference device as the origin. Consider the time delay between each pair of adjacent devices as a constraint condition, construct a least squares optimization problem, solve the least squares optimization problem to obtain the position of each conference terminal device in the unified time coordinate system, and obtain the consistent time coordinates;

[0085] Based on the consistent time coordinates, map the position of each conference terminal device in the unified time coordinate system to a two-dimensional plane. The distance between conference terminal devices reflects their time synchronization relationship. Then use the force-directed algorithm to adjust the node positions to obtain the global time relationship graph.

[0086] Specifically, perform time series decomposition on the audio feature vectors. By resampling the audio feature vectors at different time scales, generate multiple subsequences with different resolutions. The multi-scale analysis can effectively capture the features of the signal at different time scales, thereby improving the robustness and accuracy of audio synchronization. Let the original audio feature vector be , and by resampling it, generate the multi-scale feature sequence , where represents different time scales. For example, the time scales from short to long are . Through decomposition, each subsequence contains the information of the audio feature vector at a preset time scale, and the obtained multi-scale feature sequence can characterize the dynamic characteristics of the audio signal at each time level. Based on the multi-scale feature sequence, perform cross-correlation calculation on the feature sequences of each pair of conference terminal devices to determine the time synchronization deviation between each device. Select a fixed-size time window for each pair of devices, slide this time window simultaneously on the multi-scale feature sequences of two conference terminal devices, and calculate the Pearson correlation coefficient within each time window. The Pearson correlation coefficient is calculated as follows:

[0087] ;

[0088] where, and are the feature vector values of the two devices within the time window respectively, and are the means of the corresponding feature vectors, is the number of samples within the time window. By this method, the correlation coefficients of two feature sequences are calculated at different time offsets to obtain the cross-correlation function curve. The cross-correlation function curve describes the similarity of signals between two devices at different time offsets, and the peak in it represents the optimal time synchronization point. To accurately determine the peak position, peak detection is performed on the cross-correlation function curve. Local maximum points are searched on the cross-correlation function curve, and two points are taken on each side of each local maximum point, a total of five points of data. These points are used to fit a quadratic curve. The fitted quadratic curve is expressed as:

[0089] ;

[0090] where, 、 and are the coefficients of the quadratic curve, is the time offset, is the cross-correlation value. The exact position of the peak is obtained by solving the vertex position of the quadratic curve, and the formula is:

[0091] ;

[0092] Through the formula, the exact time offset position of each local peak is obtained, thereby determining the candidate values of the initial time delay. Based on the candidate values of the initial time delay, a multi-scale time delay matrix is constructed. The candidate values of the time delay obtained at different time scales are organized in matrix form, where each row of the matrix represents a time scale and each column represents the delay between a pair of devices. Let the time delay matrix be , and its element represents the time delay between the th pair of devices at the th time scale. To obtain a fused time delay value, a weighted average is performed on each column in the time delay matrix to obtain the fused time delay value:

[0093] ;

[0094] where, is the fused time delay of the th pair of devices, is the weight of the th time scale. By appropriate weight selection, the contributions of different time scales in the synchronization calculation are better balanced. Median filtering is performed on the fused time delay value to eliminate abnormal delays introduced by network fluctuations or other interferences. A sliding window of odd size , slide the sliding window over the time delay sequence. Each time, take all the values within the sliding window, sort them by size, and select the middle value as the output at the current position to obtain a smoothed time delay sequence. The median filtering method can effectively remove sudden noises, making the time delay sequence smoother and more stable. Based on the smoothed time delay sequence, construct a time delay graph. Consider each conference terminal device as a node in the graph, and the time delay between devices as the weight of the edge between nodes in the graph. Then use the Kruskal algorithm to construct a minimum spanning tree, and select the set of edges with the minimum total delay from it to obtain the topological structure between devices. The construction of the minimum spanning tree ensures the establishment of the lowest-delay connection between various devices to achieve global time synchronization. Perform global optimization on the topological structure between devices, and establish a relative time coordinate system with a reference device as the origin. Consider the time delay between each pair of adjacent devices as a constraint condition, and construct a least squares optimization problem. Let the time positions of each device be , and the goal is to minimize the following objective function:

[0095] ;

[0096] where, is the optimization objective function, represents that there is an edge between device and device , and is the time delay between the two devices. By solving the least squares optimization problem, obtain the relative positions of each device in the unified time coordinate system to get the consistent time coordinates. Based on the consistent time coordinates, map the positions of each conference terminal device in the unified time coordinate system to a two-dimensional plane, and this mapping reflects the time synchronization relationship between various devices. To optimize the distribution of nodes on the plane, use the force-directed algorithm to adjust the positions of the nodes. The force-directed algorithm makes the node distribution more balanced by simulating the attractive and repulsive forces between nodes, while maintaining the relative distances between each node, to obtain the global time relationship graph.

[0097] In a specific embodiment, the process of executing step 400 may specifically include the following steps:

[0098] Based on the global time relationship graph, allocate cache space for each conference terminal device, divide the cache space into two layers: a fast cache and a deep cache. The fast cache is used to store the recently received audio data, and the deep cache is used to store the processed historical audio data, to obtain a two-layer cache structure;

[0099] ​Perform timestamp analysis on the data in the two - layer cache structure. By calculating the time intervals between adjacent audio frames, obtain the time series of the arrival of audio data. Based on the time series of the arrival of audio data, construct the state - space model of the Kalman filter, taking the arrival time of audio data as the observation value and the network transmission delay as the state variable, to obtain the network state prediction model;

[0100] Initialize the parameters of the network state prediction model. Calculate the initial state estimate and error covariance matrix through historical data to obtain the initial parameters of the Kalman filter. Based on the initial parameters of the Kalman filter, perform state prediction and update on the newly arrived audio data. Predict the network state at the next moment through the time - update equation, and then use the measurement - update equation to correct the prediction result to obtain the optimal estimate of the current network state;

[0101] According to the optimal estimate of the network state, calculate the ideal residence time of audio data in the cache. By subtracting the predicted network transmission delay from the target playback delay, obtain the cache adjustment target value;

[0102] Based on the cache adjustment target value, reorganize the data in the two - layer cache structure. By adjusting the data distribution ratio between the fast cache and the deep cache, delay the advanced audio data and accelerate the lagged audio data to obtain the cache data distribution;

[0103] Evaluate the performance of the cache data distribution. By calculating the buffer utilization rate, packet loss rate, and audio quality metrics, obtain the performance metrics of the buffering strategy, and adjust the buffering strategy according to the performance metrics to finally obtain the buffering strategy adapted to the current network state.

[0104] Specifically, allocate cache space for each conference terminal device and divide it into a two - layer cache structure, namely a fast cache and a deep cache. The fast cache is used to store the recently received audio data, which are the latest audio packets that have not been deeply processed. The role of the fast cache is to achieve low - latency instant response to ensure the real - time playback of audio data. Relatively, the deep cache is used to store the historical audio data that has been further processed to guarantee the continuity and overall quality of the entire audio stream. Through the two - layer cache structure, a balance can be found between real - time performance and stability. Perform timestamp analysis on the data in the two - layer cache structure. By recording the arrival time of each audio frame, calculate the time interval between adjacent audio frames to obtain the time series of the arrival of audio data. Let the audio frame arrival time series be , where represents the arrival time of the -th audio frame, and the time interval is expressed as:

[0105] ;

[0106] By calculating the time interval, the delay and jitter characteristics of audio data in network transmission are understood. Based on the time series of the arrival of audio data, a state space model of the Kalman filter is constructed, taking the arrival time of audio data as the observation value and the network transmission delay as the state variable, resulting in a network state prediction model. The state space model of the Kalman filter is described as:

[0107] ;

[0108] ;

[0109] Among them, represents the state of the system, that is, the network transmission delay, is the state transition matrix, which describes the evolution of the delay over time, is the process noise, is the observation value, that is, the arrival time of the audio data, is the observation matrix, is the observation noise. Through the state space model, the transmission state of the network is effectively modeled and predicted. To enable the Kalman filter to accurately predict the network state, its parameters are initialized, and the initial state estimate and error covariance matrix are calculated. The initial state estimate is estimated based on historical data, while the error covariance matrix is used to describe the uncertainty of the initial state. Let the historical arrival time of the audio data be , and the initial state estimate is expressed as:

[0110] ;

[0111] Among them, is the number of historical data samples, represents the arrival time of the th historical sample, is the mean of the historical samples. Through the initialization method, the Kalman filter better reflects the initial conditions of the network state. After initializing the Kalman filter, it enters the prediction and update stages to perform state prediction and update on newly arrived audio data. The network state at the next moment is predicted through the time update equation:

[0112] ;

[0113] Then, the measurement update equation is used to correct the prediction result to obtain the optimal estimate of the current network state:

[0114] ;

[0115] Among them, is the Kalman gain, which determines the influence degree of the observed value on the state estimation correction. Through the iterative prediction and update process, the Kalman filter gradually obtains an accurate estimate of the network transmission delay. Based on the obtained optimal estimate of the network state, the ideal residence time of the audio data in the buffer is calculated. The predicted network transmission delay is compared with the target playback delay. Let the target playback delay be , and the network delay predicted by the Kalman filter be , then the cache adjustment target value is expressed as:

[0116] ;

[0117] The cache adjustment target value represents the adjustment amount that should be made to the audio data in the cache to ensure the smoothness and synchronization of audio playback. Based on this adjustment target value, the data in the two-layer cache structure is reorganized by adjusting the data allocation ratio between the fast cache and the deep cache. For example, for the audio data that arrives in advance, it is stored in the deep cache and processed with appropriate delay, while for the audio data that arrives late, its processing process is accelerated and it is directly stored in the fast cache to minimize the impact of delay on audio quality as much as possible. By reasonably adjusting the allocation between the fast cache and the deep cache, an optimized distribution of the cache data is achieved, thereby ensuring the synchronization and stability of the audio under different network states. The performance of the cache data distribution is evaluated by calculating the buffer utilization rate, packet loss rate, and audio quality metrics to measure the effect of the current buffering strategy. The buffer utilization rate is defined as:

[0118] ;

[0119] Among them, represents the used cache space, is the total cache capacity. By evaluating the buffer utilization rate, the usage of the cache resources is understood to determine whether the buffering strategy needs to be adjusted. The packet loss rate is used to reflect the situation of audio data loss during network transmission due to insufficient cache or unstable transmission. Audio quality metrics such as signal-to-noise ratio directly measure the final audio output quality. In the performance evaluation, the current buffering strategy is adaptively adjusted according to the obtained performance metrics. When the buffer utilization rate is high and the packet loss rate increases, the capacity of the deep cache is increased to store more historical audio data, thereby reducing the packet loss phenomenon. And when the audio quality metric drops, the ratio of the fast cache to the deep cache is adjusted to ensure that the key audio frames can be processed and played in a timely manner. Through dynamic adjustment, a buffering strategy suitable for the current network state is finally obtained.

[0120] In a specific embodiment, the process of executing step 500 may specifically include the following steps:

[0121] Extract features from the buffering policy to obtain a buffering state feature vector, and construct a reinforcement learning environment based on the buffering state feature vector. Treat each conference terminal device as an agent, use the buffering state feature vector as the environmental state, set the action space as the selectable communication frequency bands and transmission powers, and obtain a multi-agent learning environment;

[0122] Design a reward function for the multi-agent learning environment. By comprehensively considering the audio synchronization degree, transmission delay, and channel utilization rate, construct a reward function in the form of a weighted sum to obtain the optimization objective of reinforcement learning;

[0123] Construct a double deep Q-network model based on the optimization objective of reinforcement learning, including an online network and a target network. The two networks have the same structure and are composed of multiple fully connected layers. The input is the state vector, and the output is the Q value of each action, to obtain an initialized double deep Q-network;

[0124] Initialize the parameters of the initialized double deep Q-network, generate an experience replay buffer through random sampling method, and store the experience data of the interaction between the agent and the environment to obtain an experience pool;

[0125] Train the double deep Q-network based on the experience pool. By randomly extracting a batch of data from the experience pool, calculate the mean square error between the target Q value and the current Q value as the loss function, and use the backpropagation algorithm to update the parameters of the online network to obtain a trained double deep Q-network;

[0126] Use the trained double deep Q-network for action selection. Achieve a balance between exploration and exploitation through the ε-greedy policy, select the optimal communication frequency band and transmission power for each agent to obtain an initial allocation policy;

[0127] Conduct global coordination on the initial allocation policy. By constructing a centralized resource allocator, collect the state information and action selections of all agents, and solve possible resource conflicts to obtain a resource allocation policy.

[0128] Specifically, extract features from the buffering policy to obtain a buffering state feature vector. The buffering state feature vector describes the current state of the system, including information such as the buffer utilization rate of each device, the delay and arrival of audio data, etc. Let the buffering state feature of each device be , where represents the device number, and the feature vector contains the following information: buffer utilization rate , delay , channel condition etc. By extracting these features, a complete feature vector describing the current state of each conference terminal device is obtained. Based on the buffer state feature vector, a reinforcement learning environment is constructed. In the reinforcement learning environment, each conference terminal device is regarded as an agent, and the buffer state feature vector is used as the state of the environment. The action space of the agent is set to the selectable communication frequency bands and transmission powers. Let the communication frequency band be and the transmission power be , then the action is expressed as:

[0129] ;

[0130] By defining the action space, each agent selects a communication frequency band and transmission power suitable for the current network condition to achieve the best communication effect, improve the stability and quality of audio synchronization, and obtain a multi-agent learning environment. All devices interact as independent agents in this environment. To guide each agent to select appropriate actions, a reward function is designed for the multi-agent learning environment. The reward function comprehensively considers several important factors such as audio synchronization degree, transmission delay, and channel utilization rate. Let the synchronization degree be , the transmission delay be , and the channel utilization rate be , then the reward function is expressed in the form of a weighted sum:

[0131] ;

[0132] where , and are the weight coefficients of each factor, respectively, used to balance the influence among the synchronization degree, delay, and channel utilization rate. Through the design of the reward function, each agent is prompted to ensure the synchronization quality of the audio, minimize the delay as much as possible, and make full use of the network channel resources when selecting actions. Based on the above reward function and reinforcement learning objective, a double deep Q-network (DQN) model is constructed. The double deep Q-network includes two networks: one is the online network (OnlineNetwork), and the other is the target network (Target Network). The two networks have the same structure, consisting of multiple fully connected layers. The input is the state vector, and the output is the Q value of each action. The Q value is used to measure the expected return of taking a certain action in the current state. Let the input state be , and the output of the Q-network is the Q value of each possible action :

[0133] ;

[0134] where represents the state The lower action 's Q value is a parameter of the network. With this structure, the double deep Q network learns how to select the best communication strategy for each device in different states. After the network is constructed, the parameters of the initialized double deep Q network are initialized. To conduct effective training, an experience replay buffer is generated by random sampling method, which is used to store the experience data of the interaction between the agent and the environment. The experience data includes the current state , action , reward and the next state . By accumulating these experience data, an experience pool is obtained for subsequent Q network training. The double deep Q network is trained based on the experience pool. By randomly extracting a batch of experience data from the experience pool, the mean square error between the target Q value and the current Q value is calculated and used as the loss function. The calculation formula of the target Q value is as follows:

[0135] ;

[0136] where is the target Q value, is the reward for the current action, is the discount factor, represents the parameters of the target network. The mean square error loss function is:

[0137] ;

[0138] By minimizing the loss function , the parameters of the online network are updated using the backpropagation algorithm, thereby gradually improving the performance of the network to obtain the trained double deep Q network. When the network training is completed, the trained double deep Q network is used for action selection. To balance exploration and exploitation, the -greedy strategy is adopted. With probability , an action is randomly selected (exploration), and with probability , the action with the largest current Q value is selected (exploitation). Through this strategy, each agent selects the optimal communication frequency band and transmission power in most cases, but can also maintain a certain degree of exploration to discover new possible optimal strategies, obtaining the initial communication resource allocation strategy. The initial allocation strategies of each agent are globally coordinated to solve resource conflicts. A centralized resource allocator is constructed to collect the state information and action selections of all agents, and adjust the resource allocations of each device based on this information. The goal of the centralized resource allocator is to solve problems such as frequency band conflicts and power allocation overloads that occur to ensure that each device can communicate without interfering with other devices. Through global coordination, an optimized resource allocation strategy is obtained.

[0139] In a specific embodiment, the process of executing step 600 may specifically include the following steps:

[0140] Based on the resource allocation strategy, quantify the communication parameters of each conference terminal device. By converting the frequency band selection and transmission power into numerical representations, obtain a communication resource allocation matrix, and perform feature extraction on the communication resource allocation matrix. By calculating the channel capacity and transmission delay between conference terminal devices, obtain a network performance feature vector;

[0141] Based on the network performance feature vector and the global time relationship graph, construct a global synchronization optimization objective function. By taking the time synchronization error and network resource utilization rate as optimization metrics, and setting constraint conditions for the global synchronization optimization objective function. By limiting the maximum allowable synchronization error and the minimum necessary transmission rate, obtain an optimization problem with constraints;

[0142] Based on the optimization problem with constraints, construct a Lagrangian function. By introducing Lagrange multipliers to incorporate the constraint conditions into the global synchronization optimization objective function, obtain an unconstrained optimization problem. And according to the unconstrained optimization problem, by iteratively calculating the gradient of the global synchronization optimization objective function and updating the optimization variables, obtain the optimal synchronization parameters;

[0143] Based on the optimal synchronization parameters, perform time-axis adjustment on the embedded audio data stream. By performing interpolation operations on audio frames, align the audio time axes of each conference terminal device to a unified reference time axis to obtain a time-aligned audio stream. And perform content alignment on the time-aligned audio stream. By analyzing the similarity of audio feature vectors, identify and eliminate duplicate or missing audio segments, and finally obtain synchronized multi-channel audio.

[0144] Specifically, based on the resource allocation strategy, quantify the communication parameters of each conference terminal device. Convert the communication frequency band selection and transmission power of each conference terminal device into numerical forms. Let the communication frequency band be , and the transmission power be , then the quantified communication parameters are represented as a numerical vector:

[0145] ;

[0146] where represents the quantified representation of the communication parameters of the th terminal device. By quantifying the communication parameters of all devices, construct a communication resource allocation matrix , where the th row represents the Communication resource allocation for devices. Feature extraction is performed on the communication resource allocation matrix, and the channel capacity and transmission delay between devices are calculated. The channel capacity is calculated by the Shannon formula, which is used to measure the maximum communicable data rate between two devices. The formula is as follows:

[0147] ;

[0148] where represents the channel bandwidth, is the transmission power, is the channel gain, which reflects the signal propagation effect between device and device , is the noise power density. In this way, the performance metrics for communication resource allocation are obtained, and then the transmission delay between conference terminal devices is calculated , which describes the time delay for audio data to be transmitted from one device to another. Based on the channel capacity and transmission delay, the network performance eigenvector is obtained, which describes the communication status of the entire network. Based on the network performance eigenvector and the global time relationship graph, a global synchronization optimization objective function is constructed. The global synchronization optimization objective function takes the time synchronization error and network resource utilization rate as optimization metrics. Let the time synchronization error be , and the network resource utilization rate be , then the optimization objective function is expressed as:

[0149] ;

[0150] where is the global synchronization optimization objective function, and are weight coefficients, which respectively control the relative importance of the synchronization error and resource utilization rate in the optimization. In this form, while ensuring the time synchronization accuracy, the utilization efficiency of network resources is improved. To limit the optimization problem, constraint conditions are set for the global synchronization optimization objective function, including limiting the maximum allowable synchronization error and the minimum necessary transmission rate. Let the maximum allowable synchronization error be , and the minimum necessary transmission rate be , then the constraint conditions are expressed as:

[0151] ;

[0152] The constraint conditions ensure that the optimization process meets the basic requirements in terms of synchronization error and transmission rate, thus achieving high-quality audio synchronization. Based on the constrained optimization problem, a Lagrangian function is constructed. By introducing Lagrange multipliers, the constraint conditions are incorporated into the global synchronization optimization objective function to obtain an unconstrained optimization problem. Let the Lagrange multipliers be and , the Lagrangian function is expressed as:

[0153] ;

[0154] where and are non - negative Lagrange multipliers used to penalize violations of the constraints. By solving the unconstrained Lagrangian function, the originally complex constrained optimization problem is transformed into an equivalent unconstrained problem, thus enabling the use of numerical optimization methods such as gradient descent for solution. By performing iterative calculations on the Lagrangian function, solving its gradient and updating the optimization variables, the optimal synchronization parameters are gradually found. Let the synchronization parameter at the current iteration step be , then the parameter update for the next step is expressed as:

[0155] ;

[0156] where is the learning rate, which controls the update amplitude at each step, and represents the gradient of the synchronization parameter. Through multiple iterations, the optimal synchronization parameters that minimize the Lagrangian function are found. Based on the obtained optimal synchronization parameters, the time axis of the embedded audio data stream is adjusted. To keep the audio of all conference terminal devices synchronized on a unified reference time axis, interpolation operations are performed on the audio frames to align the audio data in time. By means of interpolation, the audio time axis of each device is aligned to the reference time axis to obtain a time - aligned audio stream. Let the audio frame of device be , and the aligned audio frame obtained by interpolation is:

[0157] ;

[0158] where is the interpolation coefficient, and is the time synchronization error. The purpose of the interpolation operation is to enable the audio streams of each device to be synchronized at the same time point, thus avoiding audio misalignment caused by asynchronous issues. After completing the time alignment, content alignment is performed on the time - aligned audio stream to ensure the consistency of the audio content. By analyzing the similarity of audio feature vectors, duplicate or missing segments in the audio stream are identified. Let the feature vectors of two audio segments be and respectively, then their similarity is calculated by cosine similarity:

[0159] ;

[0160] where represents the similarity between audio segments, · represents the vector dot product, represents the norm of the vector. By calculating the similarity, duplicate segments are identified and eliminated, or missing segments are completed, so as to finally obtain the synchronized multi-channel audio stream.

[0161] The method for synchronizing multi-channel audio of the conference terminal device in the embodiment of the present application is described above. Next, the multi-channel audio synchronization device 10 of the conference terminal device in the embodiment of the present application will be described. Please refer to Figure 2 One embodiment of the multi-channel audio synchronization device 10 of the conference terminal device in the embodiment of the present application includes:

[0162] An acquisition module 11, configured to acquire the audio data stream of each conference terminal device, and encode the environmental audio information into frequency domain features and embed them into the audio data stream to obtain an embedded audio data stream;

[0163] A mapping module 12, configured to perform time-frequency domain conversion and peak detection on the embedded audio data stream, and map the peak points to binary strings through the locality-sensitive hashing algorithm to obtain audio feature vectors;

[0164] A calculation module 13, configured to calculate the cross-correlation function between terminal devices based on the audio feature vectors, determine the time delay through multi-scale analysis and median filtering, and obtain a global time relationship graph;

[0165] A construction module 14, configured to construct a two-layer cache structure according to the global time relationship graph, predict the network state through a Kalman filter, and adjust the residence time of audio data in the cache to obtain a buffering strategy;

[0166] An allocation module 15, configured to use the buffering strategy as the state to input into a multi-agent reinforcement learning model, optimize the communication resource allocation through a double deep Q network, and obtain a resource allocation strategy;

[0167] A synchronization module 16, configured to construct a global synchronization optimization problem based on the resource allocation strategy, solve the optimal synchronization parameters through a gradient descent algorithm, and perform time axis adjustment and content alignment on the embedded audio data stream to obtain synchronized multi-channel audio.

[0168] Through the collaborative cooperation of the above-mentioned various components, by encoding environmental audio information into frequency-domain features and embedding them into the audio data stream, the present invention enhances the robustness of audio synchronization and can better cope with background noise interference in different conference environments. The local sensitive hashing algorithm is used to map peak points into binary strings, greatly reducing the storage and transmission overhead of audio features while maintaining the distinctiveness of the features. The introduction of multi-scale analysis and median filtering techniques improves the accuracy and stability of time delay estimation and effectively overcomes the impact of network jitter and burst noise. A two-layer cache structure is constructed and combined with a Kalman filter to predict the network state, realizing dynamic adjustment of audio data caching and effectively balancing synchronization accuracy and system delay. The multi-agent reinforcement learning model and double deep Q-network are used to optimize communication resource allocation, enabling the system to adaptively cope with complex and changing network environments and improving resource utilization efficiency. By constructing a global synchronization optimization problem and using the gradient descent algorithm to solve it, precise time axis adjustment and content alignment of multiple channels of audio are achieved, ensuring audio synchronization to the greatest extent. The present invention significantly improves the audio quality and user experience of multi-party remote conferences.

[0169] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of the structure of the electronic device 300 provided by an embodiment of the present application. The electronic device 300 includes a processor 301 and a memory 302, and the processor 301 and the memory 302 are connected through a device bus 203. Among them, the memory 302 may include a non-volatile storage medium and an internal memory.

[0170] The non-volatile storage medium can store a computer program. The computer program includes program instructions, and when the program instructions are executed by the processor 301, the processor 301 can be enabled to execute any of the above multi-channel audio synchronization methods for conference terminal devices.

[0171] The processor 301 is used to provide computing and control capabilities to support the operation of the entire electronic device 300.

[0172] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 301, the processor 301 can be enabled to execute any of the above multi-channel audio synchronization methods for conference terminal devices.

[0173] Those skilled in the art can understand that Figure 3 the structure shown in

[0174] It should be understood that the processor 301 may be a central processing unit (CPU), and the processor 301 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0175] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described electronic device 300 can refer to the corresponding process of the foregoing multi-channel audio synchronization method for conference terminal devices, and will not be elaborated herein.

[0176] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the one or more processors are caused to implement the multi-channel audio synchronization method for conference terminal devices provided by the embodiment of the present application.

[0177] Among them, the computer-readable storage medium may be an internal storage unit of the foregoing embodiment of the electronic device 300, such as the hard disk or memory of the electronic device 300. The computer-readable storage medium may also be an external storage device of the electronic device 300, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped with the electronic device 300.

[0178] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0179] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0180] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A multi-channel audio synchronization method for conference terminal equipment, characterized in that: The method comprises: The audio data stream of each conference terminal device is collected, and the environmental audio information is encoded into a frequency domain feature and embedded into the audio data stream to obtain an embedded audio data stream; Performing time-frequency domain conversion and peak detection on the embedded audio data stream, mapping the peak point into a binary string by using a local sensitive hashing algorithm, and obtaining an audio feature vector; Calculate the cross-correlation function between terminal devices based on the audio feature vector, determine the time delay through multi-scale analysis and median filtering, and obtain a global time relationship diagram; A two-layer cache structure is constructed according to the global time relationship graph, a network state is predicted by a Kalman filter, and the residence time of the audio data in the cache is adjusted to obtain a buffering strategy; The buffer strategy is input as a state into a multi-agent reinforcement learning model, and the communication resource allocation is optimized through a dual deep Q network to obtain a resource allocation strategy; Based on the resource allocation strategy, a global synchronization optimization problem is constructed, the optimal synchronization parameters are solved by the gradient descent algorithm, the time axis of the embedded audio data stream is adjusted and the content is aligned to obtain synchronized multi-channel audio; specifically, the communication parameters of each conference terminal device are quantified based on the resource allocation strategy, the communication resource allocation matrix is ​​obtained by converting the frequency band selection and the transmission power into numerical representation, and the communication resource allocation matrix is ​​feature extracted, and the network performance feature vector is obtained by calculating the channel capacity and transmission delay between the conference terminal devices; based on the network performance feature vector and the global time relationship graph, a global synchronization optimization objective function is constructed, the time synchronization error and the network resource utilization are used as optimization indicators, and the global synchronization optimization objective function is constrained by limiting the maximum Allowing synchronization error and minimum necessary transmission rate, a constrained optimization problem is obtained; based on the constrained optimization problem, a Lagrangian function is constructed, and the constraints are integrated into the global synchronization optimization objective function by introducing Lagrangian multipliers to obtain an unconstrained optimization problem, and according to the unconstrained optimization problem, the gradient of the global synchronization optimization objective function is iteratively calculated and the optimization variables are updated to obtain the optimal synchronization parameters; based on the optimal synchronization parameters, the time axis of the embedded audio data stream is adjusted, and the audio time axis of each conference terminal device is aligned to a unified reference time axis by interpolating the audio frames to obtain a time-aligned audio stream, and the content of the time-aligned audio stream is aligned, and the repeated or missing audio segments are identified and eliminated by analyzing the similarity of the audio feature vectors, so as to finally obtain synchronized multi-channel audio.

2. The multi-channel audio synchronization method of conference terminal equipment according to claim 1, characterized in that: The method collects the audio data stream of each conference terminal device, and encodes the ambient audio information into frequency domain features and embeds them into the audio data stream to obtain an embedded audio data stream, including: The sound of the environment of each conference terminal device is collected, and the collected sound signal is preprocessed through a high-pass filter to obtain the environmental background audio, and the spectrum analysis of the environmental background audio is performed to obtain the characteristic environmental audio by calculating the energy concentration and selecting the continuous, stable and characteristic frequency components; Performing discrete cosine transform on the characteristic environment audio, converting the characteristic environment audio into frequency domain space to obtain frequency domain environment audio features, sampling and quantizing the audio input of the conference terminal device, converting the analog signal into a digital signal through an analog-to-digital converter to obtain an audio data stream; Performing frame processing on the audio data stream, dividing the continuous audio data into audio frames of fixed length to obtain framed audio data, and performing fast Fourier transform on the framed audio data to convert the framed audio data into frequency domain space to obtain frequency domain audio data; Based on the frequency domain environmental audio characteristics, the mid-high frequency bands of the frequency domain audio data are modulated, the environmental audio information is embedded into the audio data stream to obtain embedded frequency domain audio data, and the embedded frequency domain audio data is inverse fast Fourier transformed to convert the embedded frequency domain audio data back to the time domain space to obtain an embedded audio data stream.

3. The multi-channel audio synchronization method of conference terminal equipment according to claim 2, characterized in that: The embedded audio data stream is subjected to time-frequency domain conversion and peak detection, and the peak point is mapped into a binary string by a local sensitive hashing algorithm to obtain an audio feature vector, including: Segmenting the embedded audio data stream, dividing the embedded audio data stream into time segments of fixed length to obtain an audio segment sequence, and transforming each audio segment in the audio segment sequence to convert the time domain signal into a time-frequency representation to obtain a time-frequency spectrum; Performing energy calculation on the time-frequency spectrum, calculating the energy value of each time-frequency unit by an integration method to obtain an energy distribution matrix, and performing peak detection on each time-frequency unit based on the energy distribution matrix to obtain a peak point set; Quantizing the coordinates of the peak point set, mapping the time-frequency coordinates to a discrete grid to obtain quantized peak points, constructing feature descriptors based on the quantized peak points, and obtaining a difference vector by calculating the time-frequency difference between adjacent peak points; The difference vector is input into a local sensitive hash function, a hash value is generated through random projection and threshold comparison to obtain a binary feature string, and the binary feature strings are grouped and merged, and multiple feature strings are combined into a vector of fixed length through bit operations to obtain an audio feature vector.

4. The multi-channel audio synchronization method of conference terminal equipment according to claim 3, characterized in that: The method of calculating the cross-correlation function between the terminal devices based on the audio feature vector, determining the time delay through multi-scale analysis and median filtering, and obtaining a global time relationship diagram includes: Performing time series decomposition on the audio feature vector, generating a plurality of subsequences with different resolutions by resampling the audio feature vector according to different time scales, each subsequence containing information of the audio feature vector at a preset time scale, and obtaining a multi-scale feature sequence; Based on the multi-scale feature sequence, a cross-correlation calculation is performed on the feature sequence of each pair of conference terminal devices, specifically: a time window of a fixed size is selected, the time window is simultaneously slid on the multi-scale feature sequences of the two conference terminal devices, the Pearson correlation coefficient is calculated for the data in the time window, and the correlation coefficient values ​​under different time offsets are recorded to form a cross-correlation function curve; Perform peak detection on the cross-correlation function curve, search for local maximum points on the cross-correlation function curve, select two points on the left and right of each local maximum point, use the data of these five points to fit a quadratic curve, calculate the vertex position of the quadratic curve as the precise peak position, and obtain an initial time delay candidate value; Constructing a multi-scale time delay matrix based on the initial time delay candidate value, organizing the time delay candidate values ​​obtained at different time scales into a matrix form, wherein each row of the multi-scale time delay matrix represents a time scale, and each column represents a delay between a pair of devices, and then performing weighted averaging on each column in the multi-scale time delay matrix to obtain a fused time delay value; Perform median filtering on the fused time delay value, set a sliding window of an odd size, slide the sliding window on the time delay sequence, take all the values ​​in the sliding window each time and sort them by size, then select the middle value as the output of the current position, to obtain a smoothed time delay sequence; A time delay graph is constructed based on the smoothed time delay sequence, each conference terminal device is regarded as a node in the graph, the time delay between conference terminal devices is used as the weight of the edge, and then a minimum spanning tree is constructed using the Kruskal algorithm, and the edge set with the minimum total delay is selected to obtain the topological structure between the devices; The topological structure between the devices is globally optimized, a relative time coordinate system with the reference device as the origin is established, the time delay between each pair of adjacent devices is regarded as a constraint condition, a least squares optimization problem is constructed, and the least squares optimization problem is solved to obtain the position of each conference terminal device in the unified time coordinate system to obtain a consistent time coordinate; Based on the consistent time coordinate, the position of each conference terminal device in the unified time coordinate system is mapped to a two-dimensional plane. The distance between the conference terminal devices reflects their time synchronization relationship. Then, the node position is adjusted using a force-directed algorithm to obtain a global time relationship diagram.

5. The multi-channel audio synchronization method of conference terminal equipment according to claim 4, characterized in that: The two-layer cache structure is constructed according to the global time relationship graph, the network state is predicted by the Kalman filter, the residence time of the audio data in the cache is adjusted, and the buffering strategy is obtained, including: Based on the global time relationship graph, cache space is allocated to each conference terminal device, and the cache space is divided into two layers: a fast cache and a deep cache, wherein the fast cache is used to store recently received audio data, and the deep cache is used to store processed historical audio data, thereby obtaining a two-layer cache structure; Performing timestamp analysis on the data in the two-layer cache structure, obtaining a time series of audio data arrival by calculating the time interval between adjacent audio frames, and constructing a state space model of a Kalman filter based on the time series of audio data arrival, taking the arrival time of the audio data as an observation value and the network transmission delay as a state variable, to obtain a network state prediction model; Initializing the parameters of the network state prediction model, calculating the initial state estimate and the error covariance matrix through historical data, obtaining the initial parameters of the Kalman filter, and based on the initial parameters of the Kalman filter, predicting and updating the state of the newly arrived audio data, predicting the network state at the next moment through the time update equation, and then correcting the prediction result using the measurement update equation to obtain the optimal estimate of the current network state; Calculating an ideal residence time of the audio data in the cache based on the optimal estimate of the network state, and obtaining a cache adjustment target value by subtracting the predicted network transmission delay from the target playback delay; Based on the cache adjustment target value, the data in the two-layer cache structure is reorganized, and the data allocation ratio between the fast cache and the deep cache is adjusted to delay the processing of the advanced audio data and accelerate the processing of the lagging audio data to obtain the cache data distribution; A performance evaluation is performed on the cache data distribution, and a performance indicator of the buffer strategy is obtained by calculating the buffer utilization, packet loss rate and audio quality indicator. The buffer strategy is adjusted according to the performance indicator to finally obtain a buffer strategy that adapts to the current network status.

6. The multi-channel audio synchronization method of conference terminal equipment according to claim 5, characterized in that: The buffer strategy is used as a state input into a multi-agent reinforcement learning model, and communication resource allocation is optimized through a dual-depth Q network to obtain a resource allocation strategy, including: Extract features of the buffer strategy to obtain a buffer state feature vector, and build a reinforcement learning environment based on the buffer state feature vector, regard each conference terminal device as an intelligent agent, use the buffer state feature vector as the environment state, set the action space as an optional communication frequency band and transmission power, and obtain a multi-agent learning environment; Designing a reward function for the multi-agent learning environment, constructing a reward function in the form of a weighted sum by comprehensively considering audio synchronization, transmission delay, and channel utilization, and obtaining an optimization target for reinforcement learning; A dual-depth Q network model is constructed based on the optimization objective of the reinforcement learning, including an online network and a target network, the two networks have the same structure, are composed of multiple fully connected layers, the input is a state vector, and the output is a Q value of each action, to obtain an initialized dual-depth Q network; Initializing parameters of the initialized dual-depth Q network, generating an experience replay buffer by a random sampling method, storing experience data of the interaction between the agent and the environment, and obtaining an experience pool; The dual-depth Q network is trained based on the experience pool, batch data is randomly extracted from the experience pool, a mean square error between a target Q value and a current Q value is calculated as a loss function, and an online network parameter is updated using a back propagation algorithm to obtain a trained dual-depth Q network; The trained dual-depth Q network is used to select actions, and an ε-greedy strategy is used to strike a balance between exploration and utilization, and an optimal communication frequency band and transmission power are selected for each intelligent agent to obtain an initial allocation strategy; The initial allocation strategy is globally coordinated, and a centralized resource allocator is constructed to collect the state information and action selections of all agents, resolve possible resource conflicts, and obtain a resource allocation strategy.

7. A multi-channel audio synchronization device for conference terminal equipment, characterized in that: The multi-channel audio synchronization device of the conference terminal equipment comprises: A collection module, used to collect the audio data stream of each conference terminal device, and encode the ambient audio information into frequency domain features and embed them into the audio data stream to obtain an embedded audio data stream; A mapping module, used to perform time-frequency domain conversion and peak detection on the embedded audio data stream, and map the peak point into a binary string through a local sensitive hashing algorithm to obtain an audio feature vector; A calculation module, used to calculate the cross-correlation function between terminal devices based on the audio feature vector, determine the time delay through multi-scale analysis and median filtering, and obtain a global time relationship diagram; A construction module is used to construct a two-layer cache structure according to the global time relationship graph, predict the network state through a Kalman filter, adjust the residence time of the audio data in the cache, and obtain a buffering strategy; An allocation module, configured to input the buffer strategy as a state into a multi-agent reinforcement learning model, optimize the communication resource allocation through a dual deep Q network, and obtain a resource allocation strategy; The synchronization module is used to construct a global synchronization optimization problem based on the resource allocation strategy, solve the optimal synchronization parameters through the gradient descent algorithm, adjust the time axis and align the content of the embedded audio data stream, and obtain synchronized multi-channel audio; specifically includes: based on the resource allocation strategy, quantify the communication parameters of each conference terminal device, convert the frequency band selection and the transmission power into numerical representation, obtain the communication resource allocation matrix, and extract features from the communication resource allocation matrix, and obtain the network performance feature vector by calculating the channel capacity and transmission delay between the conference terminal devices; based on the network performance feature vector and the global time relationship diagram, construct a global synchronization optimization objective function, use the time synchronization error and the network resource utilization as optimization indicators, and set constraints on the global synchronization optimization objective function, through The maximum allowable synchronization error and the minimum necessary transmission rate are limited to obtain a constrained optimization problem; based on the constrained optimization problem, a Lagrangian function is constructed, and the constraints are integrated into the global synchronization optimization objective function by introducing Lagrangian multipliers to obtain an unconstrained optimization problem; and according to the unconstrained optimization problem, the gradient of the global synchronization optimization objective function is iteratively calculated and the optimization variables are updated to obtain the optimal synchronization parameters; based on the optimal synchronization parameters, the time axis of the embedded audio data stream is adjusted, and the audio time axis of each conference terminal device is aligned to a unified reference time axis by interpolating the audio frames to obtain a time-aligned audio stream, and the content of the time-aligned audio stream is aligned, and the repeated or missing audio segments are identified and eliminated by analyzing the similarity of the audio feature vectors, so as to finally obtain synchronized multi-channel audio.

8. An electronic device, characterized in that: The electronic device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instruction in the memory so that the electronic device executes the multi-channel audio synchronization method of a conference terminal device according to any one of claims 1 to 6.

9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the multi-channel audio synchronization method for a conference terminal device as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Multi-path embedded video coding method and system based on generative artificial intelligence

    CN119583816A