Multi-modal code stream transmission bandwidth dynamic allocation optimization method and device
By performing feature extraction and code stream demand prediction on image and voiceprint data, combined with the deep actor-critician algorithm, the bandwidth is dynamically allocated to maximize transmission efficiency, and the problem of limited multimodal data transmission bandwidth in the distribution network is solved, achieving efficient and reliable data transmission and accurate fault detection.
Patent Information
- Application Number
- CN202510226024.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-13
AI Technical Summary
In the existing distribution network monitoring technology, the large amount of information transmission of image and voiceprint data leads to bandwidth limitations. The traditional static bandwidth allocation strategy cannot adapt to the dynamic changes of multimodal data flow, resulting in low bandwidth utilization, low data transmission efficiency, increased delay, and affecting the accuracy of fault detection.
A multimodal code stream transmission bandwidth dynamic allocation optimization method is proposed. By extracting the image data and voiceprint data, calculating the code stream demand prediction value using a timing neural network, building a bandwidth allocation optimization model, and using a deep actor-critician algorithm to solve the decision process, dynamically allocating bandwidth to maximize transmission efficiency.
Real-time and reliable transmission of multimodal code stream data is realized, the utilization rate of bandwidth resources is improved, the delay is reduced, the accuracy of fault detection is enhanced, and different monitoring scenarios are adapted.
Smart Images

Figure CN119996207A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to communication network technology, and in particular to a method and device for dynamically allocating and optimizing multi-modal code stream transmission bandwidth. Background Art
[0002] In the operation of the power system, the intelligent development of the distribution network plays a vital role in improving resource utilization, data transmission efficiency, reducing latency and improving detection accuracy. With the access of new services, the amount of multi-modal real-time information data that the distribution network needs to transmit has increased significantly, which has put forward higher requirements on the monitoring performance of the distribution network and equipment.
[0003] In the existing distribution network monitoring technology, images and voiceprints are two important multimodal data sources that are widely used in the monitoring and fault diagnosis of power equipment. Image data can provide intuitive visual information about the equipment, such as damaged shells, signs of corrosion, and overheated areas, while voiceprint data can capture audio signals when the equipment is running and identify audio features such as abnormal noise, arc discharge sound, and vibration sound. However, the large amount of information transmission of image data and voiceprint data leads to bandwidth limitations. Under high resolution and high sampling rate, the amount of data increases dramatically, which poses a huge challenge to the communication bandwidth of the distribution network.
[0004] Traditional bandwidth allocation methods often adopt static allocation strategies, which cannot adapt to the dynamic changes of multimodal data flows in distribution networks, resulting in low utilization of bandwidth resources, low data transmission efficiency, increased latency, and to some extent, affecting the accuracy of fault detection. In addition, when processing multimodal data, existing technologies fail to fully utilize the inherent connection and complementarity between image and voiceprint data, lack effective data fusion mechanisms, and have problems such as not fully considering the interweaving of multimodal data, large prediction deviations, and poor adaptability to different monitoring scenarios.
[0005] Therefore, there is an urgent need to reasonably and dynamically optimize the bandwidth allocation strategy to ensure the real-time and reliable transmission of multimodal code stream data and reduce the redundant waste of distribution network resources. Summary of the invention
[0006] Based on this, the present invention aims to propose a method and device for optimizing dynamic allocation of multimodal code stream transmission bandwidth, extract features of image data and voiceprint data, use the extracted features to predict code stream transmission demand, and dynamically allocate bandwidth based on the demand prediction value.
[0007] In a first aspect, the present invention provides a method for dynamically allocating and optimizing bandwidth for multi-modal code stream transmission, comprising:
[0008] Acquire image data and voiceprint data to be transmitted;
[0009] Extract features from the image data and the voiceprint data respectively, and calculate the image feature quantity and the voiceprint feature quantity;
[0010] Use time series neural network to calculate the bitrate demand prediction value;
[0011] A bandwidth allocation optimization model is constructed based on image features and voiceprint features with the optimization goal of maximizing bandwidth transmission efficiency. The state set of the bandwidth allocation optimization model includes the bitrate demand prediction value. The bandwidth allocation optimization model is solved to obtain the bandwidth allocation plan.
[0012] Furthermore, the bandwidth allocation optimization model is solved to obtain a bandwidth allocation solution including:
[0013] The bandwidth allocation optimization model is modeled as a Markov decision process. The state set includes available bandwidth, bitrate demand prediction value and bitrate demand prediction deviation. The action set includes voiceprint dimension, image semantic compression factor and bitrate bandwidth allocation decision. The bandwidth allocation optimization model is defined as the reward function of the decision. The deep actor-critic algorithm is used to solve the Markov decision process to obtain the bandwidth allocation plan. The playback experience pool of the deep actor-critic algorithm is constructed according to the prediction deviation.
[0014] Furthermore, feature extraction is performed on the image data and the voiceprint data respectively, and the image feature quantity and the voiceprint feature quantity are calculated to include:
[0015] Constructing the first convolutional neural network and the second convolutional neural network;
[0016] Using the first convolutional neural network to extract features from the image data to obtain image features, and using the second convolutional neural network to extract features from the voiceprint data to obtain voiceprint features;
[0017] The image features and the voiceprint features are statistically analyzed to obtain image feature quantities and voiceprint feature quantities respectively;
[0018] Considering the similarity between the image features and the voiceprint features, the total amount of state data to be transmitted is calculated based on the image feature quantities and the voiceprint feature quantities.
[0019] Furthermore, the total amount of state data to be transmitted is calculated as follows:
[0020]
[0021] in, represents the total amount of state data to be transmitted at time t, Indicates At the moment, the image code stream channel c Unique feature quantity, Indicates The moment voiceprint code stream channel d unique feature quantities; a represents the index of similar features between image features and voiceprint features, Indicates The image code stream channel c at the moment The amount of data with similar features, Indicates The voiceprint code stream channel d at the moment The amount of data with similar features, and They represent the bandwidth allocated to the image code stream channel c and the voiceprint code stream d at time t, respectively. represents the channel gain of the image code stream and the voiceprint code stream at time t, is the transmission power of the image and voiceprint code stream, is the unilateral power spectral density of additive white Gaussian noise, represents the electromagnetic interference power of the image and voiceprint channels at time t, is the Gompertz function, which is used to accurately reflect the marginal benefit impact of the number of features and feature similarity on the amount of transmitted data.
[0022] Furthermore, the image features are statistically analyzed to obtain image feature quantities including:
[0023] Original image feature quantity of statistical image features;
[0024] The image features are semantically compressed and the final image feature quantity is calculated as follows:
[0025]
[0026] in, represents the semantic compression factor of the image feature at the tth moment, represents the original image feature at the tth moment, Represents the image features actually transmitted after semantic compression.
[0027] Furthermore, the voiceprint features are statistically analyzed to obtain the voiceprint feature quantities including:
[0028] The voiceprint features are compressed in the frequency dimension and then the voiceprint feature quantities are counted.
[0029] Furthermore, the bandwidth allocation optimization model is constructed as follows:
[0030]
[0031] in, represents the multi-modal code stream transmission efficiency at time t, represents the voiceprint feature dimension, represents the semantic compression factor of image features, and They represent the bandwidth allocated to the image code stream channel c and the voiceprint code stream channel d at time t respectively.
[0032] Furthermore, the bitstream demand prediction value is calculated using the time series neural network, including:
[0033] Constructing a first temporal neural network and a second temporal neural network;
[0034] The first time series neural network and the second time series neural network are used to respectively predict the image code stream transmission demand and the voiceprint code stream transmission demand;
[0035] The interleaving coefficients of the image code stream channel and the voiceprint code stream channel are calculated using the correlation interleaving network. The image code stream transmission demand and the voiceprint code stream transmission demand are interleaved and corrected according to the interleaving coefficients to obtain the image code stream demand prediction value and the voiceprint code stream demand prediction value.
[0036] In a second aspect, the present invention proposes a device for dynamically allocating and optimizing bandwidth for multi-modal code stream transmission, comprising:
[0037] A data acquisition module, used to acquire image data and voiceprint data to be transmitted;
[0038] The feature calculation module extracts features from the image data and the voiceprint data, and calculates the image feature quantity and the voiceprint feature quantity;
[0039] A demand prediction module is used to calculate the bitstream demand prediction value using a time series neural network;
[0040] The bandwidth allocation module is used to construct a bandwidth allocation optimization model with the maximization of bandwidth transmission efficiency as the optimization goal based on the image feature quantity and the voiceprint feature quantity. The state set of the bandwidth allocation optimization model includes the bit rate demand prediction value. The bandwidth allocation optimization model is solved to obtain the bandwidth allocation plan.
[0041] In a third aspect, the present invention provides an electronic device comprising a memory storing computer executable instructions and a processor, wherein when the computer executable instructions are executed by the processor, the device executes the various steps of the multi-modal code stream transmission bandwidth dynamic allocation optimization method provided in the first aspect.
[0042] In a fourth aspect, the present invention provides a readable storage medium storing a computer executable program, which, when executed, can implement the various steps of the multi-modal code stream transmission bandwidth dynamic allocation optimization method provided in the first aspect.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] The present invention proposes a method for optimizing the dynamic allocation of bandwidth for multimodal code stream transmission. First, feature extraction is performed on multimodal data to be transmitted, including images and voiceprints, and the number of features is calculated. A further embodiment considers the influence of feature similarity between multimodal data on marginal benefits of transmission, so as to improve the adaptability of analysis results to the actual environment. A time series neural network is used to calculate the code stream demand prediction value, and a bandwidth allocation optimization model with maximizing bandwidth transmission efficiency as the optimization goal is constructed according to image feature quantities and voiceprint feature quantities. The optimization model is further modeled as a decision-making process, and a deep actor-critic algorithm is used to solve the decision-making process. The prediction deviation of the time series neural network is used to construct a playback experience pool in the decision-making process, so that the algorithm learns the optimal strategy of each parameter under different prediction accuracies, enhances the adaptability of the bandwidth allocation method to different monitoring scenarios, and overcomes the problem of sparse training samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0046] Figure 1 This is a flowchart of a method for dynamically allocating and optimizing bandwidth for multi-modal code stream transmission provided by an embodiment of the present invention;
[0047] Figure 2 It is a structural diagram of a multi-modal code stream transmission bandwidth dynamic allocation optimization device provided by an embodiment of the present invention;
[0048] Figure 3 This is a diagram of the electronic device architecture provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] See also Figure 1 An embodiment of the present invention provides a method for dynamically allocating and optimizing bandwidth for multi-modal code stream transmission, comprising the following steps:
[0051] Step S110: Obtain image data and voiceprint data to be transmitted.
[0052] In the scenario of multimodal data transmission optimization, it is first necessary to obtain the image data and voiceprint data to be transmitted. Specifically in the distribution equipment monitoring scenario, the image data usually comes from the camera at the equipment site, and the voiceprint data is generally the audio signal when the equipment is running, such as equipment noise, voice conversations, machine vibration sounds, etc.
[0053] Furthermore, the collected image data and voiceprint data can be preprocessed before feature extraction, including denoising, normalization, time-frequency transformation, etc.
[0054] Step S120: Extract features from the image data and voiceprint data respectively, and calculate image feature quantities and voiceprint feature quantities.
[0055] This step constructs independent code stream channels for image data and voiceprint data. In order to obtain key features in complex image information, the first convolutional neural network is used to extract features from the image code stream, and the second convolutional neural network is used to extract features from the voiceprint code stream. Image features and voiceprint features are obtained respectively, so as to calculate the feature quantity of multimodal data.
[0056] In a specific embodiment, step S120 includes the following steps:
[0057] Step S121. Constructing a first convolutional neural network and a second convolutional neural network;
[0058] Step S122. Using the first convolutional neural network to extract features from the image data to obtain image features, and using the second convolutional neural network to extract features from the voiceprint data to obtain voiceprint features;
[0059] Step S123. Count the image features and the voiceprint features to obtain image feature quantities and voiceprint feature quantities respectively;
[0060] Step S124. Considering the similarity between the image features and the voiceprint features, the total amount of state data to be transmitted is calculated based on the image feature quantities and the voiceprint feature quantities.
[0061] Optionally, due to the limited communication resources in the actual monitoring process, it is necessary to perform semantic compression on the image features to improve the efficiency of multimodal code stream transmission. In general, if the image semantic compression factor is given, the amount of feature data after compression can be determined, and the image feature amount actually transmitted is expressed as:
[0062]
[0063] in, represents the semantic compression factor of the image feature at the tth moment, represents the original image feature at the tth moment, Represents the image features actually transmitted after semantic compression.
[0064] Optionally, the voiceprint feature extraction process includes:
[0065] The voiceprint data is pre-emphasized, framed and windowed, then fast Fourier transformed and the power spectrum density is calculated. The power spectrum of the voiceprint code stream is spliced to extract the voiceprint features.
[0066] As a preferred implementation, considering that the extracted voiceprint feature code stream is still a large amount of data, the voiceprint features can be compressed in the frequency dimension before the voiceprint feature quantity is counted. Specifically, the frequency dimension reflects the energy intensity of the sound signal at different frequencies, and peaks will appear in certain frequency bands. The key information of the voiceprint features can be retained through filter bank processing, and the original high-resolution spectrum data can be converted into a low-dimensional feature representation that still has recognition capabilities through methods such as downsampling and principal component analysis.
[0067] Optionally, considering that image and voiceprint data have unique features within the modality and similar features between modalities, for unique features within the modality, as long as the data of the modality is successfully transmitted, the complete amount of information can be obtained; for similar features between modalities, the more information is transmitted, the greater the amount of information, but there is a decreasing marginal effect, and the decreasing speed is related to the similarity between modalities and the number of features. Therefore, considering the decreasing marginal benefit of similar features, the marginal benefit effect of the number of features and feature similarity on the amount of data transmitted is introduced, and then the total amount of state data to be transmitted is calculated based on the image feature quantity and voiceprint feature quantity. The total amount of state data to be transmitted can be further expressed as follows:
[0068]
[0069] in, represents the total amount of state data to be transmitted at time t, Indicates At the moment, the image code stream channel c Unique feature quantity, Indicates The moment voiceprint code stream channel d unique feature quantities; a represents the index of similar features between image features and voiceprint features, Indicates The image code stream channel c at the moment The amount of data with similar features, Indicates The voiceprint code stream channel d at the moment The amount of data with similar features, and They represent the bandwidth allocated to the image code stream channel c and the voiceprint code stream d at time t, respectively. represents the channel gain of the image code stream and the voiceprint code stream at time t, is the transmission power of the image and voiceprint code stream, is the unilateral power spectral density of additive white Gaussian noise, represents the electromagnetic interference power of the image and voiceprint channels at time t, It is a Gompertz function, which is used to accurately reflect the marginal benefit effect of the number of features and feature similarity on the amount of data transmitted. The higher the feature similarity between image features and voiceprint features, the stronger the marginal effect of data transmission will be.
[0070] Step S130: Calculate the bit rate demand prediction value using the time series neural network.
[0071] The traditional bitstream bandwidth allocation method does not take into account the data transmission requirements of the bitstream. In order to maximize the amount of transmitted data, this step uses a time series neural network to predict the transmission requirements of the image bitstream and voiceprint bitstream, and outputs the bitstream requirement values for several future time steps through regression prediction.
[0072] As an optional embodiment, step S130 includes the following steps:
[0073] Step S131. Constructing a first time series neural network and a second time series neural network;
[0074] Step S132. Use the first time series neural network and the second time series neural network to respectively predict the image code stream transmission demand and the voiceprint code stream transmission demand;
[0075] Step S133. Calculate the interleaving coefficients of the image code stream channel and the voiceprint code stream channel using the correlation interleaving network, perform interleaving correction on the image code stream transmission demand and the voiceprint code stream transmission demand according to the interleaving coefficients, and obtain the image code stream demand prediction value and the voiceprint code stream demand prediction value.
[0076] This embodiment introduces a correlation interleaving network and introduces historical correlation when predicting demand, so that network training focuses on the characteristics of frequently transmitted bitstream data, improves the adaptability of multi-modal bitstream prediction results to the operating environment, and the bitstream transmission demand predicted by the interleaved and corrected timing neural network can be expressed as the sum of the transmission demand of data-unique features and the transmission demand of similar features between modalities.
[0077] Step S140. Construct a bandwidth allocation optimization model based on the image feature and the voiceprint feature with the optimization goal of maximizing bandwidth transmission efficiency. The state set of the bandwidth allocation optimization model includes a bitrate demand prediction value. Solve the bandwidth allocation optimization model to obtain a bandwidth allocation solution.
[0078] This step takes the bandwidth allocation optimization problem as the goal of maximizing bandwidth transmission efficiency. By dynamically allocating bandwidth to meet the needs of different modal code streams, in order to better capture the system state and decision impact, the bandwidth allocation problem is modeled as a decision process. The state set of the decision process includes the code stream demand prediction value calculated in the previous step. The state set is the set of all possible states of the environment. Each state represents a possible configuration or situation of the system. Through the selection of continuous state-action pairs, the bandwidth utilization is gradually optimized, thereby achieving the global goal of maximizing transmission efficiency.
[0079] More specifically, the bandwidth allocation optimization model is modeled as a Markov decision process. The state set includes available bandwidth, bitrate demand prediction value and bitrate demand prediction deviation. The action set includes voiceprint dimension, image semantic compression factor and bitrate bandwidth allocation decision. The bandwidth allocation optimization model is defined as the reward function of the decision. The deep actor-critic algorithm is used to solve the Markov decision process to obtain the bandwidth allocation plan, where the playback experience pool of the deep actor-critic algorithm is constructed according to the prediction deviation.
[0080] The Markov decision process is used to describe the process of making decisions in a given environment. In the distribution equipment monitoring scenario applied in the embodiment of the present invention, the state set of the decision process is defined as available bandwidth, bitstream demand prediction value and bitstream demand prediction deviation, where available bandwidth represents the total bandwidth resources that can be allocated in the current system, bitstream demand prediction value is the prediction result of future bitstream demand using a time series neural network (such as LSTM), which includes both image bitstream demand and voiceprint bitstream demand, and bitstream demand prediction deviation represents the error between the prediction value and the actual demand, which is used to measure the uncertainty of the prediction model, thereby providing additional information for decision making; the action set is defined as voiceprint dimension, image semantic compression factor and bitstream bandwidth allocation decision, where voiceprint dimension represents the impact of the dimension of voiceprint features (for example, through dimensionality reduction or feature selection) on voiceprint data transmission demand, image semantic compression factor represents the impact of image features after compression on image data transmission demand, and bitstream bandwidth allocation decision represents the current transmission bandwidth allocated to image data and voiceprint data respectively. On this basis, the optimization objective is defined as the reward function of the decision process, that is, maximizing bandwidth transmission efficiency.
[0081] The embodiment of the present invention proposes to use a deep actor-critic algorithm to solve the decision-making process. The output of the algorithm is the mean and variance of each optimization variable. Then, the inverse distribution function method is used to generate specific values for each optimization variable to obtain a bandwidth allocation scheme. In the learning phase of the algorithm, samples are placed in different levels of playback experience pools according to the prediction deviation. During playback learning, sparse sample learning is added. Part of the samples are extracted from each playback experience pool with the same probability to form a sample set. Based on this, a loss function for algorithm training is established, so that the algorithm learns the optimal strategy for each optimization variable under different prediction accuracies.
[0082] The method of the present invention is further described below through a specific embodiment.
[0083] An embodiment of the present invention provides a method for dynamically allocating and optimizing bandwidth for multi-modal code stream transmission, comprising the following steps:
[0084] Step S210: construct an image feature extraction network and a voiceprint feature extraction network respectively.
[0085] This step uses a convolutional neural network as the architecture of the feature extraction network, assuming that the training time set is , T is the total number of training moments.
[0086] The extraction of image features can be expressed as:
[0087]
[0088] in, is the input image of channel c at time t, is the convolution kernel parameter of channel c at time t, It is the feature map output by channel c at the tth moment.
[0089] Due to the limited communication resources in the actual monitoring process, it is necessary to perform semantic compression on the image features. The number of image features actually transmitted after compression is expressed as:
[0090]
[0091] in, is the image semantic compression factor at time t, reflecting the degree of semantic compression of image data equipment. is the original feature data volume of channel c at time t, is the amount of feature data actually transmitted at time t after compression, and the image feature set at time t after compression is expressed as ,in is the feature map output by channel c at time t after compression, and C is the number of image code stream channels.
[0092] The extraction of voiceprint features is to pre-emphasize, frame and window the voiceprint data, then perform fast Fourier transform on the voiceprint code stream and calculate the power spectral density. Finally, the power spectrum of the voiceprint data is spliced to extract the voiceprint features.
[0093] Since the voiceprint feature code flow after extraction is still large, the voiceprint code stream needs to be further compressed. The current analysis of voiceprint data shows that the voiceprint data components are mainly concentrated on the 50Hz frequency component in the range of 0 to 4kHz, and the frequency is evenly distributed in the range of 0 to 8kHz. Therefore, the voiceprint data can be compressed in the frequency dimension based on the triangular filter group. The voiceprint feature extraction and compression process can be expressed as:
[0094]
[0095] in, is the voiceprint feature extracted from channel d at time t, is the voiceprint feature after compression of channel d at time t, represents the triangular filter bank compression function, for The center frequency vector at time , where for Moment The center frequency of the triangular filter, for The total number of triangular filters at time, Determines the key frequency of the filter bank, Determines the dimension of the compressed voiceprint, that is, ,in is the dimension of the voiceprint after compression at time t, express The number of features extracted by each triangular filter at the moment, the feature set of the compressed voiceprint code stream can be expressed as , D is the number of voiceprint code stream channels.
[0096] Step S220: Use the image feature extraction network and the voiceprint feature extraction network to extract image features and voiceprint features respectively.
[0097] Step S230: Count the image feature quantities and voiceprint feature quantities, and calculate the multimodal feature data quantity based on the marginal benefit of similar features.
[0098] Image and voiceprint data have unique features within the modality and similar features between modalities. For unique features within the modality, as long as the data of the modality is successfully transmitted, the complete amount of information can be obtained; for similar features between modalities, the more information is transmitted, the greater the amount of information, but there is a decreasing marginal effect, and the decreasing speed is related to the similarity between modalities and the number of features. Therefore, considering the decreasing marginal benefits of similar features, the marginal benefit effect of the number of features and feature similarity on the amount of data transmitted is introduced, and then the total amount of state data to be transmitted is calculated based on the image feature quantity and voiceprint feature quantity. The total amount of state data to be transmitted can be further expressed as follows:
[0099]
[0100] in, represents the total amount of state data to be transmitted at time t, Indicates At the moment, the image code stream channel c Unique feature quantity, Indicates The moment voiceprint code stream channel d unique feature quantities; a represents the index of similar features between image features and voiceprint features, Indicates The image code stream channel c at the moment The amount of data with similar features, Indicates The voiceprint code stream channel d at the moment The amount of data with similar features, and They represent the bandwidth allocated to the image code stream channel c and the voiceprint code stream d at time t, respectively. represents the channel gain of the image code stream and the voiceprint code stream at time t, is the transmission power of the image and voiceprint code stream, is the unilateral power spectral density of additive white Gaussian noise, represents the electromagnetic interference power of the image and voiceprint channels at time t, is the Gompertz function, which is used to accurately reflect the marginal benefit impact of the number of features and feature similarity on the amount of transmitted data.
[0101] The optimization goal of the embodiment of the present invention is to maximize the transmission efficiency, which is specifically expressed as the ratio of the total state data volume to the allocated bandwidth, as shown below:
[0102]
[0103] in, represents the multi-modal code stream transmission efficiency at time t, represents the voiceprint feature dimension, represents the semantic compression factor of image features, and They represent the bandwidth allocated to the image code stream channel c and the voiceprint code stream channel d at time t respectively.
[0104] S240. Calculate the code stream demand prediction value based on correlation interleaving using a time series neural network.
[0105] Since the traditional multimodal code stream bandwidth allocation method does not take into account the data transmission requirements of the code stream, in order to maximize the amount of data transmitted by the distribution network equipment status, this step predicts the multimodal data transmission requirements based on the interleaving of multimodal data.
[0106] The image data code stream demand prediction network and the voiceprint data code stream demand prediction network are constructed respectively. This embodiment adopts the LSTM network as the architecture of the prediction network. The image code stream demand prediction process can be expressed as:
[0107]
[0108] in, , , They are the status information of the forget gate, input gate and output gate of the LSTM network of the image code stream channel c respectively; , , , They are the weight indices of the forget gate, input gate, output gate and candidate unit of the LSTM network of the image code stream channel c respectively; , , , are the offsets of the LSTM network of the image code stream channel c; and They are The input and hidden layer output of the LSTM network of the image code stream channel c at time instant; for The hidden layer output of the LSTM network of the image code stream channel c at time instant; is the state information of the candidate unit of the LSTM network of the image code stream channel c; and is the activation function, which is used to update the state information for the forget gate, input gate and output gate of the LSTM network, and to create new state information for the candidate unit.
[0109] The voiceprint code stream demand prediction process can be expressed as:
[0110]
[0111] in, , , They are the state information of the forget gate, input gate, and output gate of the LSTM network of the voiceprint code stream channel d; , , , They are the weight indices of the forget gate, input gate, output gate and candidate unit of the LSTM network of the voiceprint code stream channel d; , , , These are the offsets of the LSTM network of the voiceprint code stream channel d; , They are -1 The input and hidden layer output of the LSTM network of the voiceprint code stream channel d at time instant; For the Hidden layer output of the LSTM network of voiceprint code stream channel d at time instant. It is the state information of the candidate unit of the LSTM network of the voiceprint code stream channel d.
[0112] This embodiment also introduces a correlation interleaving network to calculate the interleaving degree of image data and voiceprint data, and introduces the influence of historical correlation and other modal information in the demand prediction process. The code stream demand prediction based on this can be expressed as:
[0113]
[0114] In the formula, is the interleaving coefficient of the image code stream channel c and the voiceprint code stream channel d at the tth moment, which is used to realize the interaction of different code stream information in the demand prediction process and improve the prediction accuracy. , , are neural network parameters. is the set of probability densities of the LSTM network hidden layer outputs of the image code stream channel c and the voiceprint code stream channel d at the tth moment, where , They are the marginal probability density functions of the hidden layer outputs of the LSTM network for the image code stream channel c and the voiceprint code stream channel d respectively. is the set of feature historical correlations at different time scales, where is the historical correlation degree of the features of channel c of the image code stream at time t, is the historical correlation degree of voiceprint code stream channel d at time t.
[0115] The introduction of feature history correlation can make network training focus on the features of frequently transmitted bitstream data and improve the adaptability of multimodal bitstream prediction results to the operating environment. is the linear rectification activation function, are normalization functions, both of which can improve the stability of the demand forecasting training process.
[0116] After the interleaving correction of the correlation interleaving network, the final output of the LSTM network of the image code stream channel c and the voiceprint code stream channel d at the tth moment can be expressed as and ,in, , They are the image unique feature transmission requirements and similar feature transmission requirements predicted by the LSTM network of the image code stream channel c, , They are the voiceprint unique feature transmission requirements and similar feature transmission requirements predicted by the LSTM network of the voiceprint code stream channel d.
[0117] Step S250. Construct a bandwidth allocation optimization model with maximizing bandwidth transmission efficiency as the optimization goal based on the image feature quantity and the voiceprint feature quantity, and use the deep actor-critic algorithm to solve the bandwidth allocation optimization model to obtain a bandwidth allocation plan.
[0118] This embodiment models the optimization objective as a Markov decision process, which includes:
[0119] (1) State set
[0120] The state set includes available bandwidth, bitrate demand prediction value and prediction deviation, which can be expressed as:
[0121]
[0122] in, is the available bandwidth at time t, is the LSTM empirical bias at time t, that is, the set of prediction biases of the training rounds before time t.
[0123] (2) Action set
[0124] The action set is defined as the joint decision of voiceprint dimension, image semantic compression factor, and image and voiceprint bandwidth allocation, expressed as:
[0125]
[0126] (3) Reward Function
[0127] It is defined as the optimization goal of the bandwidth allocation optimization model, that is, maximizing the bandwidth transmission efficiency.
[0128] The value of each optimization variable solved by the deep actor-critic algorithm (DAC) satisfies the normal distribution. The output of DAC is the mean and variance of each optimization variable, that is, , , , , , , , , the probability density function is expressed as:
[0129]
[0130] According to the above probability density function, the inverse distribution function method is used to generate the specific value of each optimization variable. Among them, the voiceprint dimension needs to be rounded, and the image semantic compression factor needs to be normalized. For the image bandwidth and voiceprint bandwidth, if the sum of the generated image bandwidth and voiceprint bandwidth is not greater than the available bandwidth, that is , the generated image bandwidth and voiceprint bandwidth are the final allocated bandwidth, otherwise, the available bandwidth is allocated in equal proportion according to the generated results.
[0131] This embodiment uses the prediction deviation to construct the playback experience pool of the DAC algorithm. The prediction deviation is expressed as:
[0132]
[0133] in, and are the true values of the unique feature requirements and similar feature requirements of the image code stream channel c at time t, and are the true values of the unique feature requirements and similar feature requirements of the voiceprint code stream channel d at the tth moment, and is the weight coefficient.
[0134] Four sub-experience pools are constructed using the prediction deviation, and the levels are divided into excellent, good, medium, and poor, which are represented as , , , . According to the prediction deviation, the samples are placed in the corresponding experience pool, where if , then put it into the excellent experience pool; if , then put it into the good experience pool, if , then put it into the middle experience pool; if , then put it into the difference experience pool, where , and It is the experience pool classification threshold.
[0135] When replaying learning, add sparse sample learning, extract some samples from the four sub-experience pools with the same probability Constructing the loss function , which allows the DAC algorithm to learn the optimal strategies for voiceprint dimension, image semantic compression ratio, image data bandwidth, and voiceprint data bandwidth allocation under different prediction accuracies of the time series network, and enhance the adaptability of the bandwidth allocation method.
[0136] The network parameters of the DAC algorithm are updated using the gradient ascent method according to the loss function, and the demand prediction network is updated according to the prediction deviation. The above process can be expressed as:
[0137]
[0138] In the formula, for The samples, for The number of samples, , , , are samples from excellent, good, medium, and poor experience pools, respectively. , , , for , , , The corresponding sample size is to ensure equal sampling probability. . is the value function of the DAC algorithm, are the network parameters of the critic network in the DAC algorithm, It represents the maximum value function value obtained based on the state space and action space at time t+1. Refers to the value function value calculated based on the state space and action space at time t.
[0139] The above embodiments propose a method for optimizing the dynamic allocation of bandwidth for multimodal code stream transmission. First, feature extraction is performed on the multimodal data to be transmitted, including images and voiceprints, and the number of features is calculated. A further embodiment considers the impact of feature similarity between multimodal data on the marginal benefit of transmission, thereby improving the adaptability of the analysis result to the actual environment. A time series neural network is used to calculate the code stream demand prediction value, and a bandwidth allocation optimization model with maximizing bandwidth transmission efficiency as the optimization goal is constructed based on image feature quantities and voiceprint feature quantities. The optimization model is further modeled as a decision-making process, and the decision-making process is solved using a deep actor-critic algorithm. The prediction deviation of the time series neural network is used to construct a playback experience pool in the decision-making process, so that the algorithm learns the optimal strategy for each parameter under different prediction accuracies, thereby enhancing the adaptability of the bandwidth allocation method to different monitoring scenarios and overcoming the problem of sparse training samples.
[0140] The above-disclosed method can be implemented by using various forms of equipment, so the present invention also discloses a multi-modal code stream transmission bandwidth dynamic allocation optimization device corresponding to the above-mentioned method, and a specific embodiment is given below for detailed description.
[0141] like Figure 2 As shown, an embodiment of the present invention provides a device for dynamically allocating and optimizing bandwidth for multi-modal code stream transmission, comprising:
[0142] The data acquisition module 202 is used to acquire the image data and voiceprint data to be transmitted;
[0143] The feature calculation module 204 performs feature extraction on the image data and the voiceprint data, and calculates the image feature quantity and the voiceprint feature quantity;
[0144] The demand prediction module 206 is used to calculate the bit stream demand prediction value using a time series neural network;
[0145] The bandwidth allocation module 208 is used to construct a bandwidth allocation optimization model with the optimization goal of maximizing bandwidth transmission efficiency based on the image feature quantity and the voiceprint feature quantity. The state set of the bandwidth allocation optimization model includes the bitrate demand prediction value. The bandwidth allocation optimization model is solved to obtain a bandwidth allocation plan.
[0146] The implementation principle and technical effects of the device provided in the embodiments of the present application are the same as those of the aforementioned method embodiments. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned method embodiments.
[0147] The methods and related devices mentioned in the above embodiments are described with reference to the method flow charts and / or structural diagrams provided in the embodiments of the present application. Specifically, each process and / or block in the method flow charts and / or structural diagrams, as well as the combination of processes and / or blocks in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 Schematic diagram of one or more processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A flow or multiple flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.
[0148] The following embodiments are described by taking the method applied to a computer device as an example. It can be understood that the computer device can be any device with computing and processing functions, and can be but not limited to a server or a personal laptop computer, etc. In one embodiment, the computer device can be an application server, and the application server can be a server for running an application to be tested.
[0149] See also Figure 3 , which shows a hardware block diagram of an electronic device, the electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0150] like Figure 3 As shown, the electronic device includes: at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4;
[0151] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 communicate with each other through the communication bus 4;
[0152] The processor 1 may be a central processing unit CPU, or an application-specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0153] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), etc., such as at least one disk memory;
[0154] The memory stores a program, and the processor can call the program stored in the memory, wherein the program is used to implement various processing flows of the aforementioned multi-modal code stream transmission bandwidth dynamic allocation optimization.
[0155] An embodiment of the present invention further provides a readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processing flow of the multi-modal code stream transmission bandwidth dynamic allocation optimization solution provided in the above embodiment and / or any possible implementation method in combination with the embodiment is implemented.
[0156] The above-mentioned embodiments have described the present invention in particular detail with respect to possible scenarios, and those skilled in the art will recognize that the present invention can be practiced through other embodiments. The specific naming of components, the capitalization of terms, attributes, data structures, or any other programming or structural aspects are not mandatory or important, and the mechanisms or features of the present invention may have different names, forms, or procedures. The system may be implemented by a combination of hardware and software (as described), entirely by hardware elements, or entirely by software elements. The specific division of functions between the various system components described herein is exemplary only and not mandatory; on the contrary, the functions performed by a single system component may be performed by multiple components, or the functions performed by multiple components may be performed by a single component.
[0157] Those skilled in the art should understand that the various steps of the above disclosed method can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, optionally, they can be implemented with program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the embodiments of the present invention are not limited to any specific combination of hardware and software.
[0158] These computing device executable programs (also referred to as programs, software, software applications, or code) include machine instructions for programmable processors, and these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0159] Certain aspects of the present invention include process steps and instructions described herein in the form of algorithms. It should be noted that the process steps and instructions of the present invention can be implemented in software, firmware and / or hardware, and when implemented by software, it can be downloaded, stored on different platforms used by various operating systems and operated from the platforms.
[0160] Those skilled in the art will understand that the structures shown in the accompanying drawings are merely block diagrams of partial structures related to the scheme of the present application, and do not constitute a limitation on the terminal device to which the scheme of the present application is applied. The specific terminal device may include more or fewer components than those shown in the figures, or combine certain components, or have a different arrangement of components.
[0161] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "possible design" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0162] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0163] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for dynamically allocating and optimizing bandwidth for multi-modal code stream transmission, characterized in that: include: Acquire image data and voiceprint data to be transmitted; Extract features from the image data and the voiceprint data respectively, and calculate the image feature quantity and the voiceprint feature quantity; Use time series neural network to calculate the bitrate demand prediction value; A bandwidth allocation optimization model is constructed based on image features and voiceprint features with the optimization goal of maximizing bandwidth transmission efficiency. The state set of the bandwidth allocation optimization model includes the bitrate demand prediction value. The bandwidth allocation optimization model is solved to obtain the bandwidth allocation plan.
2. The method according to claim 1, characterized in that: Solving the bandwidth allocation optimization model to obtain a bandwidth allocation solution includes: The bandwidth allocation optimization model is modeled as a Markov decision process. The state set includes available bandwidth, bitrate demand prediction value and bitrate demand prediction deviation. The action set includes voiceprint dimension, image semantic compression factor and bitrate bandwidth allocation decision. The bandwidth allocation optimization model is defined as the reward function of the decision. The deep actor-critic algorithm is used to solve the Markov decision process to obtain the bandwidth allocation plan. The playback experience pool of the deep actor-critic algorithm is constructed according to the prediction deviation.
3. The method according to claim 1, characterized in that The feature extraction of the image data and the voiceprint data is performed respectively, and the image feature quantity and the voiceprint feature quantity are calculated and obtained, including: Constructing the first convolutional neural network and the second convolutional neural network; Using the first convolutional neural network to extract features from the image data to obtain image features, and using the second convolutional neural network to extract features from the voiceprint data to obtain voiceprint features; The image features and the voiceprint features are statistically analyzed to obtain image feature quantities and voiceprint feature quantities respectively; Considering the similarity between the image features and the voiceprint features, the total amount of state data to be transmitted is calculated based on the image feature quantities and the voiceprint feature quantities.
4. The method according to claim 3, characterized in that: The total amount of state data to be transmitted is calculated as follows: in, represents the total amount of state data to be transmitted at time t, Indicates At the moment, the image code stream channel c Unique feature quantity, Indicates The moment voiceprint code stream channel d unique feature quantities; a represents the index of similar features between image features and voiceprint features, Indicates The image code stream channel c at the moment The amount of data with similar features, Indicates The voiceprint code stream channel d at the moment The amount of data with similar features, and They represent the bandwidth allocated to the image code stream channel c and the voiceprint code stream d at time t, respectively. represents the channel gain of the image code stream and the voiceprint code stream at time t, is the transmission power of the image and voiceprint code stream, is the unilateral power spectral density of additive white Gaussian noise, represents the electromagnetic interference power of the image and voiceprint channels at time t, is the Gompertz function, which is used to accurately reflect the marginal benefit impact of the number of features and feature similarity on the amount of transmitted data.
5. The method according to claim 3, characterized in that: The image feature quantity obtained by performing statistics on the image feature includes: Original image feature quantity of statistical image features; The image features are semantically compressed and the final image feature quantity is calculated as follows: in, represents the semantic compression factor of the image feature at the tth moment, represents the original image feature at the tth moment, Represents the image features actually transmitted after semantic compression.
6. The method according to claim 1, characterized in that The bandwidth allocation optimization model is constructed as follows: in, represents the multi-modal code stream transmission efficiency at time t, represents the voiceprint feature dimension, represents the semantic compression factor of image features, and They represent the bandwidth allocated to the image code stream channel c and the voiceprint code stream channel d at time t respectively.
7. The method according to claim 1, characterized in that The method of calculating the bit stream demand prediction value by using a time series neural network includes: Constructing a first temporal neural network and a second temporal neural network; The first time series neural network and the second time series neural network are used to respectively predict the image code stream transmission demand and the voiceprint code stream transmission demand; The interleaving coefficients of the image code stream channel and the voiceprint code stream channel are calculated using the correlation interleaving network. The image code stream transmission demand and the voiceprint code stream transmission demand are interleaved and corrected according to the interleaving coefficients to obtain the image code stream demand prediction value and the voiceprint code stream demand prediction value.
8. A device for dynamically allocating and optimizing bandwidth for multi-modal code stream transmission, characterized in that: include: A data acquisition module, used to acquire image data and voiceprint data to be transmitted; The feature calculation module extracts features from the image data and the voiceprint data, and calculates the image feature quantity and the voiceprint feature quantity; A demand prediction module is used to calculate the bitstream demand prediction value using a time series neural network; The bandwidth allocation module is used to construct a bandwidth allocation optimization model with the maximization of bandwidth transmission efficiency as the optimization goal based on the image feature quantity and the voiceprint feature quantity. The state set of the bandwidth allocation optimization model includes the bitrate demand prediction value. The bandwidth allocation optimization model is solved to obtain the bandwidth allocation plan.
9. An electronic device, characterized in that: The device comprises a memory storing computer executable instructions and a processor, and when the computer executable instructions are executed by the processor, the device executes the multi-modal code stream transmission bandwidth dynamic allocation optimization method as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that: A computer executable program is stored, and when the program is executed, the multi-modal code stream transmission bandwidth dynamic allocation optimization method as described in any one of claims 1 to 7 can be implemented.
Citation Information
Cited By
Equipment fault voiceprint recognition monitoring system and method based on cloud edge cooperation
CN120602518A
GIS image voiceprint joint perception method and system based on multi-modal interaction and edge intelligent optimization
CN121659195A