Audio and video wireless optimization transmission system and method based on content perception
By building a wireless transmission platform for the client and server, performing frame rate division, feature extraction and key frame marking, and formulating a transmission strategy based on the network channel status, the problems of idle resources and high bit error rate in traditional audio and video wireless transmission methods are solved, and efficient and stable audio and video transmission is achieved.
Patent Information
- Application Number
- CN202411817552.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-12-11
AI Technical Summary
Traditional wireless audio and video transmission methods cannot adapt to real-time changes in bandwidth, resulting in idle resources or transmission congestion, and cannot accurately meet dynamic transmission needs. In addition, the high bit error rate and unstable latency of wireless channels affect the integrity and accuracy of audio and video, resulting in a poor user experience.
The content-aware audio and video wireless optimization transmission system builds a wireless transmission platform for the client and server, performs frame rate division, feature extraction, and key frame marking, formulates a transmission strategy based on the network channel status, and uses adaptive signal conversion and coding compression technology to optimize the transmission strategy.
It achieves the rational allocation and optimal utilization of network resources, improves transmission efficiency and system stability, reduces data transmission volume, ensures the integrity and accuracy of audio and video, and enhances user experience.
Smart Images

Figure CN119697426B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless transmission technology, and more specifically, to a content-aware based audio and video wireless optimization transmission system and method. Background Art
[0002] In the current digital age, wireless transmission of audio and video content has become an indispensable part of people's daily lives and work. With the popularity of smartphones, tablets, and other mobile devices, users' demand for accessing and sharing audio and video content anytime, anywhere is growing. However, due to the complexity and resource limitations of wireless network environments, traditional audio and video wireless transmission methods face a series of challenges.
[0003] Compared with existing technologies, traditional audio and video wireless transmission methods mostly use fixed encoding and unified bandwidth allocation strategies, which are difficult to adapt to real-time bandwidth changes and cannot accurately meet the dynamic transmission needs of audio and video, resulting in idle resources or transmission congestion.
[0004] The inherent characteristics of audio and video data also exacerbate transmission difficulties. Large data volumes and a high level of redundant information, when transmitted directly without optimized processing, will significantly increase network load. For example, the massive pixel data of high-definition video and the dense sampling points of high-fidelity audio will quickly exhaust limited bandwidth without effective compression and filtering. Furthermore, audio and video with different frame rates, resolutions, and content complexity have different transmission requirements. The traditional "one-size-fits-all" approach cannot account for this diversity, resulting in the obstruction of critical information transmission or the excessive use of resources by secondary content, reducing transmission efficiency and quality.
[0005] Furthermore, the inherent characteristics of wireless channels bring challenges such as high bit error rates and unstable latency. Signals are susceptible to multipath fading and interference noise, resulting in data loss or erroneous reception, affecting the integrity and accuracy of audio and video. Delay jitter disrupts audio and video synchronization, causing issues such as image and sound out of sync, frame freezes, and frame skipping. This is particularly prominent in real-time interactive scenarios, significantly reducing user experience satisfaction.
[0006] In view of this, the present invention proposes a content-aware based audio and video wireless optimization transmission system and method to solve the above problems. Summary of the Invention
[0007] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solutions:
[0008] The content-aware audio and video wireless optimized transmission method includes:
[0009] Step 1: Build a wireless transmission platform based on the client and server architecture, where the client includes a sending end and a receiving end;
[0010] Step 2: Based on the transmitting end, the audio and video data to be transmitted is obtained, and the frame rate of the data is divided to obtain the frame rate audio and video of several single frames;
[0011] Step 3: Extract features of the obtained frame rate audio and video to obtain corresponding audio and video feature information, and based on the feature information, mark key frames of the frame rate audio and video in the corresponding audio and video data;
[0012] Step 4: Upload the key-frame marked audio and video data to the server, and the server constructs and executes a corresponding wireless transmission strategy based on the received audio and video data;
[0013] Step 5: The receiving end receives the transmitted audio and video data, verifies its integrity, and displays it to the user after passing the verification.
[0014] Furthermore, the client is composed of a receiving end and a sending end, and each client is provided with a unique device end number, wherein the client corresponding to each user can serve as a receiving end and a sending end; the sending end is used to obtain the audio and video data required to be transmitted by the user and perform data processing on it; the receiving end is used to receive the audio and video data transmitted by the sending end; the server end is used to obtain the network channel status between the sending end and the receiving end, and formulate and execute corresponding transmission strategies in combination with the audio and video data after data processing.
[0015] Furthermore, based on the transmitting end, the audio and video data to be transmitted is obtained, and the frame rate of the data is divided to obtain the frame rate audio and video of several single frames. The process includes:
[0016] The transmitting end is provided with a collection unit, and when the collection unit identifies that an audio and video transmission task is generated in the corresponding transmitting end, the collection unit collects data of the audio and video stream required to be transmitted in the corresponding audio and video transmission task to obtain corresponding audio and video data;
[0017] The collected audio and video data is divided into data according to the video frame rate to obtain several single-frame audio and video frame rates.
[0018] Furthermore, the process of extracting features from the obtained frame rate audio and video to obtain corresponding audio and video feature information includes:
[0019] Extract data from the acquired frame rate audio and video to obtain corresponding frame rate audio and frame rate images; at the same time, assign corresponding timestamps to the corresponding frame rate audio and frame rate images based on the acquisition time of the audio and video;
[0020] Then, preprocessing the obtained frame rate audio to obtain a corresponding frame rate spectrogram;
[0021] According to the time sequence corresponding to the timestamps, the corresponding frame rate spectrogram and frame rate image are input into the pre-built audio and video feature model to obtain the corresponding audio and video feature information.
[0022] Further, data sampling is performed on the obtained frame rate audio to obtain a plurality of data sampling points, and then N1 consecutive data sampling points are merged to obtain a plurality of sampled audios (i.e., the corresponding frame rate audio is unequally divided to obtain a plurality of consecutive sampled audios); wherein N1 is a constant, and the total number of data sampling points N1 corresponding to different sampled audios is not exactly the same;
[0023] The obtained sampled audio is subjected to adaptive signal transformation to obtain the corresponding sampled spectrum. The corresponding adaptive signal transformation formula is as follows: Where ni represents the i-th sampled audio corresponding to the n-th frame rate audio; FFT() represents the fast Fourier transform operation; δ n represents the windowing function, which is adaptively selected by the signal frequency and length of the corresponding frame rate audio; P ni Indicates the sampling spectrum of the i-th sampled audio corresponding to the n-th frame rate audio; X ni Indicates the total number of data sampling points in the i-th sampled audio corresponding to the n-th frame rate audio; * indicates convolution operation;
[0024] Obtain the spectrum frequency within the corresponding sampling spectrum and perform logarithmic operation on it to obtain the corresponding logarithmic spectrum line; the formula for performing logarithmic operation is: Where, F(t) represents the corresponding logarithmic spectrum line; DCT() represents the discrete cosine transform operation, and pf represents the spectrum frequency in the corresponding sampled spectrum; wherein, one spectrum frequency corresponds to one logarithmic spectrum line;
[0025] Furthermore, based on the obtained logarithmic spectral lines, the spectral coefficients corresponding to the corresponding collected audio are constructed. The spectral coefficients can be used to capture the spectral envelope information of the speech signal, reduce the redundant information in the original speech signal, and improve the efficiency and accuracy of the speech recognition system;
[0026] The formula for obtaining the corresponding spectral coefficient is:
[0027] Where i is the index of the sampled audio, a is the logarithmic spectrum corresponding to the corresponding sampled audio, m represents the index of the number of pre-selected filters, and M is the total number of pre-selected filters; SP (i,a) represents the spectral coefficient corresponding to the i-th sampled audio; S(i, m) represents the i-th sampled audio after filtering by the m-th filter;
[0028] Obtain the spectral coefficients corresponding to all sampled audio corresponding to the corresponding frame rate audio, and construct the corresponding frame rate spectrogram based on them.
[0029] Furthermore, the construction process of the audio and video feature model includes:
[0030] The backbone network of the audio and video feature model is defined as an improved convolutional neural network; the basic architecture of the improved convolutional neural network is an input layer, a convolution layer, a pooling layer, a feature fusion layer, and an output layer;
[0031] The input layer is used to receive input vectors and perform image normalization on them to meet model requirements; the input vectors include frame rate spectrograms and frame rate images. Image normalization refers to scaling the image size of the input vectors to meet model input requirements, and performing image dimensionality upscaling on the scaled input vectors. Image dimensionality upscaling refers to mapping the image from the original pixel space to a higher-dimensional feature space to better represent the image features.
[0032] The convolution layer uses multi-scale convolution kernels to perform convolution operations on the input vector processed by the input layer to obtain the corresponding audio features and image features;
[0033] The pooling layer is used to perform an adaptive pooling operation on the output vector of the convolutional layer;
[0034] The feature fusion layer is used to perform feature fusion on the audio features and image features after the pooling operation to obtain corresponding audio and video fusion features;
[0035] The output layer is equipped with a fully connected layer, which uses the flatten algorithm to expand the audio and video fusion features to meet the dimensionality requirements of the fully connected layer. The fully connected layer is then used to reduce the dimensionality of the data; after the dimensionality reduction process is completed, the result is output;
[0036] Acquire several sets of historical audio and video data and audio and video feature information corresponding to the corresponding audio and video data; and construct a corresponding training data set based on the data; the training data set consists of several training samples;
[0037] The obtained training data set is divided into a training set and a test set, and the corresponding audio and video feature model is iteratively trained based on the training set until the loss function tends to converge, the model parameters are saved, and the corresponding audio and video feature model is verified based on the test set. If the verification passes, the corresponding audio and video feature model is output and put into use. If the verification fails, the iterative training continues until it meets the requirements.
[0038] Furthermore, the operation formula of the convolution operation is: Where Z l1+1 and Z l1Represent the output vector and input vector of the l1+1th convolution layer respectively; (i, j) represents the dynamic area position during the convolution operation; ω l1+1 Refers to the connection weight between the l1 layer and the l1+1 layer, b l1+1 represents the bias term, Represents the convolution operation;
[0039] The formula for the pooling operation is: Where s0 represents the step size, which is used to determine the pooling speed during the pooling operation; p∈(1,∞) is used to represent the degree of pooling, and when p takes the maximum value, it indicates the maximum pooling operation; Indicates the position of the dynamic area after pooling in the feature map corresponding to the k-th input vector in the l2-th pooling layer, where k is a natural number; x and y refer to the horizontal and vertical local offsets in the corresponding feature map; B1 and B2 are natural numbers, respectively used for the size of the pooling window in the horizontal and vertical directions during the pooling operation; l2 represents the index of the pooling layer;
[0040] The corresponding feature fusion process includes:
[0041] Obtain the audio features and image features completed by the pooling operation, and perform two-dimensional convolution processing on them respectively to obtain the corresponding convolution audio Ka and Va and convolution image Qv; perform dot product operation on the obtained convolution audio Ka and Qv to obtain the corresponding dot product weight, and assign it to the convolution audio Va to obtain the corresponding initial fusion feature Fav; the acquisition formula of the corresponding initial fusion feature is: Where Fv represents the image feature, β represents the learning parameter, softmax is the pre-selected activation function; d represents the feature dimension of the corresponding convolution audio and convolution image; ⊙ represents the dot product operation; T represents the matrix transpose;
[0042] Perform two-dimensional convolution processing on the obtained initial fusion features, and repeat the above initial fusion feature acquisition process to obtain the corresponding audio and video fusion features;
[0043] Among them, the formula for two-dimensional convolution processing is: Ka = Fa*Wa; Qv = Fv*Wv; Va = Fa*Wb; Wa, Wv and Wb respectively represent the network weight matrices in the network model training process; Fa represents the audio feature.
[0044] Furthermore, the process of marking the frame rate audio and video in the corresponding audio and video data as key frames based on the obtained audio and video feature information, and constructing and executing the corresponding wireless transmission strategy based on the received audio and video data includes:
[0045] Obtain audio and video feature information corresponding to each frame rate audio and video in the corresponding audio and video data; randomly select N2 audio and video feature information corresponding to the frame rate audio and video as the corresponding initial cluster center; N2 is a natural number;
[0046] Obtain the Euclidean distances between other audio and video feature information and the corresponding initial cluster centers respectively; obtain the audio and video feature information with the smallest Euclidean distance from the corresponding initial cluster center, and construct corresponding feature cluster pairs based on the audio and video feature information;
[0047] Obtain the Euclidean distance between different feature cluster pairs, and based on the above cluster pair acquisition process, perform secondary clustering on the corresponding feature cluster pairs, and so on, until several feature cluster sets are obtained;
[0048] Arbitrarily obtain an audio and video feature information in the feature cluster set, and obtain the average value of the Euclidean distance between it and other frame rate audio and video; obtain the frame rate audio and video image corresponding to the audio and video feature information with the minimum average value, and perform frame difference operation on it with the adjacent frame rate audio and video images to obtain the corresponding frame difference factor, and determine whether it is a key frame based on the frame difference factor. If so, mark the corresponding frame rate audio and video as a key frame;
[0049] After the key frame marking is completed, the audio and video data after the key frame marking is encoded and compressed based on the pre-set second encoding scheme to obtain the corresponding second encoded audio; and the second encoded audio and the corresponding audio and video transmission task are uploaded to the server;
[0050] When the server receives the corresponding second encoded audio, the server obtains transmission data during the transmission of the corresponding second encoded audio;
[0051] The server obtains the device terminal numbers of the corresponding sending end and transmitting end based on the received audio and video transmission tasks, and constructs a corresponding temporary wireless transmission channel based on the device terminal numbers; and collects the status data of each network channel in the corresponding temporary wireless transmission channel to obtain corresponding channel network data;
[0052] Pre-allocating audio and video network bandwidth for each frame rate within the corresponding audio and video data based on the obtained transmission data;
[0053] Based on the obtained transmission data corresponding to the audio at each frame rate and the pre-allocated network bandwidth, and in combination with the corresponding channel network data, network channel allocation is performed to obtain the corresponding initial channel transmission strategy;
[0054] After the allocation is completed, if the idle network bandwidth of the network channel allocated to the corresponding frame rate audio and video is not greater than the pre-allocated network bandwidth, the encoding compression scheme of the corresponding frame rate audio and video is changed to the pre-set third encoding scheme; if the idle network bandwidth of the network channel allocated to the corresponding frame rate audio and video is greater than the pre-allocated network bandwidth; the encoding compression scheme of the corresponding frame rate audio and video is changed to the pre-set first encoding scheme;
[0055] After the change is completed, the corresponding initial channel transmission strategy is updated and returned to the sender for execution.
[0056] Furthermore, the receiving end is used to receive the transmitted audio and video data, perform integrity verification on it, and present it to the user after passing the verification. The process includes:
[0057] When the receiving end recognizes that the corresponding temporary wireless transmission channel is established, the receiving end constructs a cache node; a cache slot is set in the cache node; when the receiving end receives the corresponding coded audio, it decodes it, obtains the timestamp in the decoded coded audio, and verifies whether there is an identical timestamp in an existing cache slot in the corresponding cache node; if not, it stores the timestamp in the corresponding cache slot and constructs a new cache slot; if it exists, it ignores the corresponding coded audio, wherein the coded audio is one of the first coded audio, the second coded audio, and the third coded audio;
[0058] After the transmission is completed, the cache node will merge the decoded encoded audio based on the timestamp and key frame mark to obtain the corresponding first audio and video, and perform frame rate sequence verification and frame rate error verification on it; if there is no corresponding frame rate sequence error and frame rate error verification, the cache node will play the frame rate audio and video in the received first audio and video before and after the timestamp. If there is an error, the frame rate position of the missing frame rate audio and video is obtained and fed back to the server, and the server will retransmit the corresponding frame rate audio and video, and so on.
[0059] The content-aware audio and video wireless optimization transmission system includes:
[0060] The data acquisition module is used to capture the audio and video data to be transmitted and divide it into frame rates to obtain audio and video at several frame rates;
[0061] A data processing module is used to process the obtained frame rate audio and video data and obtain corresponding audio and video feature information based on the data processing results;
[0062] A transmission strategy module is used to mark the key frames of audio and video at each frame rate in the corresponding audio and video data based on the audio and video feature information, and to construct a corresponding channel transmission strategy based on the key frame marking;
[0063] The user display module is used to receive the transmitted audio and video data, perform integrity verification on it, and display it to the user after the verification is passed.
[0064] The technical effects and advantages of the content-aware audio and video wireless optimization transmission system and method of the present invention are as follows:
[0065] 1. By obtaining the network channel status between the sender and receiver and formulating corresponding transmission strategies based on the processed audio and video data, it is possible to achieve reasonable allocation and optimal utilization of network resources and avoid resource waste;
[0066] 2. By preprocessing, extracting features, and encoding and compressing audio and video data, as well as building temporary wireless transmission channels and allocating network channels on the server side, it can effectively address network fluctuations and packet loss, improving the robustness and stability of the system.
[0067] 3. By building a wireless transmission platform based on the client and server, and adopting technologies such as frame rate division, feature extraction and key frame marking, it is possible to effectively compress audio and video data, reduce the amount of data transmission, and thus improve the efficiency of wireless transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 Schematic diagram of the content-aware audio and video wireless optimization transmission method of the present invention;
[0069] Figure 2 Schematic diagram of the content-aware audio and video wireless optimization transmission system of the present invention. DETAILED DESCRIPTION
[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0071] Example 1
[0072] See also Figure 1 As shown, the content-aware audio and video wireless optimized transmission method of this embodiment includes:
[0073] Step 1: Build a wireless transmission platform based on the client and server architecture. The client includes the sender and the receiver.
[0074] Step 2: Based on the transmitting end, the audio and video data to be transmitted is obtained, and the frame rate of the data is divided to obtain the frame rate audio and video of several single frames;
[0075] Step 3: Extract features of the obtained frame rate audio and video to obtain corresponding audio and video feature information, and based on the feature information, mark key frames of the frame rate audio and video in the corresponding audio and video data;
[0076] Step 4: Upload the keyframe-marked audio and video data to the server, which constructs and executes the corresponding wireless transmission strategy based on the received audio and video data;
[0077] Step 5: The receiving end receives the transmitted audio and video data, verifies its integrity, and displays it to the user after passing the verification;
[0078] It should be further explained that, in the specific implementation process, the client consists of a receiving end and a sending end, and each client is provided with a unique device terminal number, wherein the client corresponding to each user can serve as both a receiving end and a sending end; the sending end is used to obtain the audio and video data required to be transmitted by the user, and perform data processing on it; the receiving end is used to receive the audio and video data transmitted by the sending end; the server end is used to obtain the network channel status between the sending end and the receiving end, and formulate and execute corresponding transmission strategies based on the audio and video data after data processing.
[0079] It should be further explained that, in a specific implementation process, the process of obtaining the audio and video data to be transmitted based on the transmitting end and dividing the frame rate thereof to obtain the frame rate audio and video of several single frames includes:
[0080] The transmitting end is provided with a collection unit. When the collection unit identifies that an audio and video transmission task is generated in the corresponding transmitting end, the collection unit collects data of the audio and video stream required to be transmitted in the corresponding audio and video transmission task to obtain the corresponding audio and video data;
[0081] Then, the collected audio and video data is divided according to the video frame rate to obtain several single-frame audio and video frames.
[0082] It should be further explained that, in a specific implementation process, the process of extracting features from the obtained frame rate audio and video to obtain corresponding audio and video feature information includes:
[0083] Extract data from the acquired frame rate audio and video to obtain corresponding frame rate audio and frame rate images; at the same time, assign corresponding timestamps to the corresponding frame rate audio and frame rate images based on the acquisition time of the audio and video;
[0084] Then, preprocessing the obtained frame rate audio to obtain a corresponding frame rate spectrogram;
[0085] According to the time sequence corresponding to the timestamps, the corresponding frame rate spectrogram and frame rate image are input into the pre-built audio and video feature model to obtain the corresponding audio and video feature information.
[0086] It should be further explained that, in a specific implementation process, the process of preprocessing the obtained frame rate audio and obtaining the corresponding frame rate spectrogram includes:
[0087] Data sampling is performed on the obtained frame rate audio to obtain a plurality of data sampling points, and then N1 consecutive data sampling points are merged to obtain a plurality of sampled audios (i.e., the corresponding frame rate audio is divided into unequal parts to obtain a plurality of consecutive sampled audios); wherein N1 is a constant, and the total number of data sampling points N1 corresponding to different sampled audios is not exactly the same;
[0088] The obtained sampled audio is subjected to adaptive signal transformation to obtain the corresponding sampled spectrum. The corresponding adaptive signal transformation formula is as follows: Where ni represents the i-th sampled audio corresponding to the n-th frame rate audio; FFT() represents the fast Fourier transform operation; δ n represents the windowing function, which is adaptively selected by the signal frequency and length of the corresponding frame rate audio; P ni Indicates the sampling spectrum of the i-th sampled audio corresponding to the n-th frame rate audio; X ni Indicates the total number of data sampling points in the i-th sampled audio corresponding to the n-th frame rate audio; * indicates convolution operation;
[0089] Obtain the spectrum frequency within the corresponding sampling spectrum and perform logarithmic operation on it to obtain the corresponding logarithmic spectrum line; the formula for performing logarithmic operation is: Where, F(t) represents the corresponding logarithmic spectrum line; DCT() represents the discrete cosine transform operation, and pf represents the spectrum frequency in the corresponding sampled spectrum; wherein, one spectrum frequency corresponds to one logarithmic spectrum line;
[0090] Furthermore, based on the obtained logarithmic spectrum lines, the spectral coefficients corresponding to the corresponding collected audio are constructed. The spectral coefficients can be used to capture the spectral envelope information of the speech signal, reducing the redundant information in the original speech signal and improving the efficiency and accuracy of the speech recognition system.
[0091] The formula for obtaining the corresponding spectral coefficient is:
[0092] Where i is the index of the sampled audio, a is the logarithmic spectrum corresponding to the corresponding sampled audio, m represents the index of the number of pre-selected filters, and M is the total number of pre-selected filters; SP (i,a)represents the spectral coefficient corresponding to the i-th sampled audio; S(i, m) represents the i-th sampled audio after filtering by the m-th filter;
[0093] Obtaining the spectral coefficients corresponding to all sampled audios corresponding to the corresponding frame rate audio, and constructing the corresponding frame rate spectrogram based on them;
[0094] It should be further explained that, in the specific implementation process, the construction process of the audio and video feature model includes:
[0095] The backbone network of the audio and video feature model is defined as an improved convolutional neural network. The basic architecture of the improved convolutional neural network is composed of input layer, convolution layer, pooling layer, feature fusion layer and output layer.
[0096] The input layer receives input vectors and normalizes them to meet the model requirements. The input vectors include frame rate spectrograms and frame rate images. Image normalization involves scaling the input vectors to meet the model input requirements and performing image upscaling on the scaled input vectors. Image upscaling involves mapping the image from its original pixel space to a higher-dimensional feature space to better represent the image's features.
[0097] The convolution layer uses multi-scale convolution kernels to perform convolution operations on the input vector processed by the input layer to obtain the corresponding audio features and image features;
[0098] The operation formula of the convolution operation is: Where Z l1+1 and Z l1 Represent the output vector and input vector of the l1+1th convolution layer respectively; (i, j) represents the dynamic area position during the convolution operation; ω l1+1 Refers to the connection weight between the l1 layer and the l1+1 layer, b l1+1 represents the bias term, Represents the convolution operation;
[0099] The pooling layer is used to perform adaptive pooling operations on the output vectors of the convolutional layer. The formula for the pooling operation is: Where s0 represents the step size, which is used to determine the pooling speed during the pooling operation; p∈(1,∞) is used to indicate the degree of pooling (i.e., pooling type), and when p takes the maximum value, it indicates the maximum pooling operation; Indicates the position of the dynamic area after pooling in the feature map corresponding to the k-th input vector in the l2-th pooling layer, where k is a natural number; x and y refer to the horizontal and vertical local offsets in the corresponding feature map; B1 and B2 are natural numbers, respectively used for the size of the pooling window in the horizontal and vertical directions during the pooling operation; l2 represents the index of the pooling layer;
[0100] The feature fusion layer is used to fuse the audio features and image features after the pooling operation; the corresponding fusion process is:
[0101] Obtain the audio features and image features completed by the pooling operation, and perform two-dimensional convolution processing on them respectively to obtain the corresponding convolution audio Ka and Va and convolution image Qv; perform dot product operation on the obtained convolution audio Ka and Qv to obtain the corresponding dot product weight, and assign it to the convolution audio Va to obtain the corresponding initial fusion feature Fav; the acquisition formula of the corresponding initial fusion feature is: Wherein, Fv represents the image feature, β represents the learning parameter, softmax is the pre-selected excitation function; d represents the feature dimension of the corresponding convolution audio and convolution image; ⊙ represents the dot product operation; T represents the matrix transpose, and the convolution audio and convolution image in the present invention are the feature matrices obtained after two-dimensional convolution processing of the audio features and the image features;
[0102] The formula for two-dimensional convolution processing is: Ka = Fa*Wa; Qv = Fv*Wv; Va = Fa*Wb; Wa, Wv, and Wb represent the network weight matrices in the network model training process; Fa represents the audio feature;
[0103] Then, the obtained initial fusion features are subjected to two-dimensional convolution processing, and the above initial fusion feature acquisition process is repeated to obtain the corresponding audio and video fusion features;
[0104] The output layer is equipped with a fully connected layer, which uses the flatten algorithm to expand the audio and video fusion features to meet the dimensionality requirements of the fully connected layer. The fully connected layer is then used to reduce the dimensionality of the data. After the dimensionality reduction process is completed, the results are output.
[0105] The loss function of the audio and video feature model is defined as Where N3 is the total number of training samples input to the audio and video feature model, h is an index variable used to traverse all training samples, T`(h) is the output of the audio and video feature model corresponding to the h-th training sample, and T(h) is the label of the h-th training sample;
[0106] Acquire several sets of historical audio and video data and audio and video feature information corresponding to the corresponding audio and video data; and construct a corresponding training data set based on the data; the training data set consists of several training samples;
[0107] The obtained training data set is divided into a training set and a test set, and the corresponding audio and video feature model is iteratively trained based on the training set until the loss function tends to converge, the model parameters are saved, and the corresponding audio and video feature model is verified based on the test set. If the verification passes, the corresponding audio and video feature model is output and put into use. If the verification fails, the iterative training continues until it meets the requirements.
[0108] It should be further explained that, in a specific implementation process, the process of marking the key frames of the frame rate audio and video in the corresponding audio and video data based on the obtained audio and video feature information includes:
[0109] Obtain audio and video feature information corresponding to each frame rate audio and video in the corresponding audio and video data; randomly select audio and video feature information corresponding to N2 frame rates of audio and video as the corresponding initial cluster center, where N2 is a natural number;
[0110] Obtain the Euclidean distances between other audio and video feature information and the corresponding initial cluster centers respectively; obtain the audio and video feature information with the smallest Euclidean distance to the corresponding initial cluster center, and construct the corresponding feature cluster pair based on it; obtain other audio and video feature information excluding the feature cluster pair, and repeat the corresponding feature cluster pair acquisition process until all audio and video feature information has a corresponding feature cluster pair; or the Euclidean distance between the audio and video feature information and other audio and video feature information is not less than a pre-set distance threshold;
[0111] Then, the Euclidean distances between different feature cluster pairs are obtained, and based on the above cluster pair acquisition process, the corresponding feature cluster pairs are clustered twice, and so on, until several feature cluster sets are obtained;
[0112] Arbitrarily obtain an audio and video feature information in the feature cluster set, and obtain the average value of the Euclidean distance between it and other frame rate audio and video; obtain the frame rate audio and video image corresponding to the audio and video feature information with the minimum average value, and perform frame difference operation on it with the adjacent frame rate audio and video image to obtain a corresponding frame difference factor; if the corresponding frame difference factor is not less than a preset frame difference threshold, it indicates that the corresponding frame rate audio and video is a key frame, and then the corresponding frame rate audio and video is marked as a key frame in the corresponding audio and video data; if the corresponding frame difference factor is less than the preset frame difference threshold, it indicates that the corresponding frame rate audio and video is not a key frame; then obtain the frame rate audio and video corresponding to the audio and video with the second minimum average value, and so on;
[0113] After the key frame marking is completed, the audio and video data after the key frame marking is encoded and compressed based on the pre-set second encoding scheme to obtain the corresponding second encoded audio; and the second encoded audio and the corresponding audio and video transmission task are uploaded to the server.
[0114] It should be further explained that, in a specific implementation process, the process of constructing and executing a corresponding wireless transmission strategy based on the received audio and video data includes:
[0115] When the server receives the corresponding second encoded audio, the server obtains the transmission data of the corresponding second encoded audio during transmission; the transmission data includes parameters such as transmission space, transmission rate, transmission delay, and transmission packet loss rate;
[0116] Furthermore, based on the received audio and video transmission tasks, the server obtains the device terminal numbers of the corresponding sending end and transmitting end, and builds a corresponding temporary wireless transmission channel based on them; and collects the status data of each network channel in the corresponding temporary wireless transmission channel to obtain the corresponding channel network data; the channel network data includes parameters such as channel bandwidth and channel delay;
[0117] Based on the obtained transmission data, the audio and video network bandwidth of each frame rate in the corresponding audio and video data is pre-allocated; the formula for pre-allocation is: Where C g C represents the pre-allocated network bandwidth for the g-th frame rate audio and video; total is the total idle network bandwidth in the corresponding temporary wireless transmission channel; R g Y represents the transmission space required for the g-th frame rate audio and video; g The audio weight of the g-th frame rate audio or video is determined by whether the corresponding frame rate audio is a key frame and the Euclidean distance between the corresponding frame rate audio and the nearest key frame; g represents the index of the frame rate audio or video; wherein, the constraint condition for network bandwidth allocation within the above frame rate audio or video is to use the second coding scheme for coding compression;
[0118] Then, based on the transmission data corresponding to the obtained audio at each frame rate and the pre-allocated network bandwidth, and in combination with the corresponding channel network data, network channel allocation is performed to obtain the corresponding initial channel transmission strategy;
[0119] After the allocation is completed, if the idle network bandwidth of the network channel allocated to the corresponding frame rate audio and video is not greater than the pre-allocated network bandwidth, the encoding compression scheme of the corresponding frame rate audio and video is changed to the pre-set third encoding scheme; if the idle network bandwidth of the network channel allocated to the corresponding frame rate audio and video is greater than the pre-allocated network bandwidth; the encoding compression scheme of the corresponding frame rate audio and video is changed to the pre-set first encoding scheme;
[0120] After the change is completed, the corresponding initial channel transmission strategy is updated and returned to the sender for execution. The updated initial channel transmission strategy is not the final transmission strategy. The corresponding transmission strategy excludes the frame rate audio and video being transmitted. The transmission strategies corresponding to the remaining frame rates audio and video will be updated in real time based on the channel network information.
[0121] It should be further explained that, in the specific implementation process, when changing the corresponding encoding scheme, it is necessary to consider both the idle bandwidth of the network channel and the idle total network bandwidth in the audio and video wireless transmission channel. If the total network bandwidth is sufficient, the frame rate of audio and video with key frame markers will be increased first.
[0122] Among them, the pre-set first coding scheme, second coding scheme and third coding scheme refer to coding processing processes using different modulation modes and coding methods, among which the size of the coding processing effect is: first coding scheme>second coding scheme>third coding scheme.
[0123] It should be further explained that, in a specific implementation, the receiving end receives the transmitted audio and video data, verifies its integrity, and displays it to the user after verification. The process includes:
[0124] When the receiving end recognizes that the corresponding temporary wireless transmission channel is established, the receiving end will establish a cache node; a cache slot is set in the cache node; when the receiving end receives the corresponding encoded audio, it decodes it, obtains the timestamp in the decoded encoded audio, and verifies whether the same timestamp exists in the existing cache slot in the corresponding cache node; if not, it stores the same timestamp in the corresponding cache slot and establishes a new cache slot; if it exists, it ignores the corresponding encoded audio, wherein the encoded audio is one of the first encoded audio, the second encoded audio, and the third encoded audio;
[0125] After the transmission is completed, the cache node will merge the decoded encoded audio based on the timestamp and key frame mark to obtain the corresponding first audio and video, and perform frame rate sequence verification and frame rate error verification on it; if there is no corresponding frame rate sequence error and frame rate error verification, the cache node will play the frame rate audio and video in the received first audio and video according to the timestamp before and after. Among them, for problems such as buffering and freezing during audio and video playback, the cache center will adaptively select appropriate bit rate stream and resolution parameters for playback based on the current network status of the receiving end; users can also set bit rate stream, resolution and other parameters based on the receiving end for local cache playback;
[0126] If there is an error, the frame rate position of the missing frame rate audio and video is obtained and fed back to the server, and the server retransmits the corresponding frame rate audio and video, and so on;
[0127] The frame rate error verification refers to error verification based on a pre-set automatic repeat request (ARQ) strategy; the automatic repeat request (ARQ) strategy is a prior art and the inventors will not elaborate on it in detail.
[0128] The present invention constructs a wireless transmission platform; in terms of transmission efficiency, through frame rate division and key frame marking, key information is accurately screened, the transmission volume is reduced, and the coding scheme and channel bandwidth are flexibly allocated according to network conditions to ensure smoothness, smooth peaks and fill valleys, alleviate congestion, and improve overall efficiency and stability; in terms of quality assurance, the feature extraction model deeply explores the audio spectrum and image features and integrates them, key frame screening ensures integrity and coherence, the receiving end verifies and corrects errors, and multiple mechanisms work together to provide users with an orderly, high-quality audio-visual experience, so that wirelessly transmitted audio and video have both efficiency and quality.
[0129] Example 2
[0130] See also Figure 2 As shown, for the parts not described in detail in this embodiment, please refer to the description of Example 1, which provides a content-aware audio and video wireless optimization transmission system, including:
[0131] The data acquisition module is used to obtain the audio and video data to be transmitted and divide it into frame rates to obtain several frame rate audio and video;
[0132] A data processing module is used to process the obtained frame rate audio and video data and obtain corresponding audio and video feature information based on the data processing results;
[0133] A transmission strategy module is used to mark the key frames of audio and video at each frame rate in the corresponding audio and video data based on the audio and video feature information, and to construct a corresponding channel transmission strategy based on the key frame marking;
[0134] The user display module is used to receive the transmitted audio and video data, verify its integrity, and display it to the user after passing the verification;
[0135] The modules are connected via wired and / or wireless means to achieve data transmission between modules.
[0136] Example 3
[0137] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the operation mode of the above-mentioned content-aware-based audio and video wireless optimization transmission system and method is implemented.
[0138] Since the electronic device introduced in this embodiment is an electronic device used to implement the content-aware audio and video wireless optimization transmission system and method in the embodiment of this application, based on the content-aware audio and video wireless optimization transmission system and method introduced in the embodiment of this application, technical personnel in this field can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application will not be introduced in detail here. As long as technical personnel in this field implement the electronic device used by the content-aware audio and video wireless optimization transmission system and method in the embodiment of this application, it falls within the scope of protection of this application.
[0139] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.
[0140] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for users of ordinary skill in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A content-aware audio and video wireless optimization transmission method, characterized in that: include: Step 1: Build a wireless transmission platform based on the client and server architecture, where the client includes a sending end and a receiving end; Step 2: Based on the transmitting end, the audio and video data to be transmitted is obtained, and the frame rate of the data is divided to obtain the frame rate audio and video of several single frames; Step 3: Extract features of the obtained frame rate audio and video to obtain corresponding audio and video feature information, and based on the feature information, mark key frames of the frame rate audio and video in the corresponding audio and video data; The process of key-frame marking of the frame rate audio and video in the corresponding audio and video data includes: Obtain audio and video feature information corresponding to each frame rate audio and video in the corresponding audio and video data; randomly select N2 audio and video feature information corresponding to the frame rate audio and video as the corresponding initial cluster center; N2 is a natural number; Obtain the Euclidean distances between other audio and video feature information and the corresponding initial cluster centers respectively; obtain the audio and video feature information with the smallest Euclidean distance from the corresponding initial cluster center, and construct corresponding feature cluster pairs based on the audio and video feature information; Obtain the Euclidean distance between different feature cluster pairs, and based on the above cluster pair acquisition process, perform secondary clustering on the corresponding feature cluster pairs, and so on, until several feature cluster sets are obtained; Arbitrarily obtain an audio and video feature information in the feature cluster set, and obtain the average value of the Euclidean distance between it and other frame rate audio and video; obtain the frame rate audio and video image corresponding to the audio and video feature information with the minimum average value, and perform frame difference operation on it with the adjacent frame rate audio and video images to obtain the corresponding frame difference factor, and determine whether it is a key frame based on the frame difference factor. If so, mark the corresponding frame rate audio and video as a key frame; Step 4: Upload the key-frame marked audio and video data to the server, and the server constructs and executes a corresponding wireless transmission strategy based on the received audio and video data; Step 5: The receiving end receives the transmitted audio and video data, verifies its integrity, and displays it to the user after passing the verification.
2. The method for wireless optimized transmission of audio and video based on content perception according to claim 1, characterized in that: The client is composed of a receiving end and a sending end, and each client is provided with a unique device end number; The sending end is used to obtain the audio and video data that the user needs to transmit and perform data processing on it; the receiving end is used to receive the audio and video data transmitted by the sending end; the service end is used to obtain the network channel status between the sending end and the receiving end, and formulate and execute corresponding transmission strategies based on the audio and video data after data processing.
3. The method for wireless optimized transmission of audio and video based on content perception according to claim 2, characterized in that: The process of obtaining the audio and video data to be transmitted based on the sending end and dividing the frame rate thereof to obtain the frame rate audio and video of several single frames includes: The transmitting end is provided with a collection unit, and when the collection unit identifies that an audio and video transmission task is generated in the corresponding transmitting end, the collection unit collects data of the audio and video stream required to be transmitted in the corresponding audio and video transmission task to obtain corresponding audio and video data; The collected audio and video data is divided into data according to the video frame rate to obtain several single-frame audio and video frame rates.
4. The content-aware audio and video wireless optimization transmission method according to claim 3, characterized in that: The process of extracting features from the obtained frame rate audio and video to obtain corresponding audio and video feature information includes: Extract data from the acquired frame rate audio and video to obtain corresponding frame rate audio and frame rate images; at the same time, assign corresponding timestamps to the corresponding frame rate audio and frame rate images based on the acquisition time of the audio and video; Preprocessing the obtained frame rate audio to obtain a corresponding frame rate spectrogram; According to the time sequence corresponding to the timestamps, the corresponding frame rate spectrogram and frame rate image are input into the pre-built audio and video feature model to obtain the corresponding audio and video feature information.
5. The method for wireless optimized transmission of audio and video based on content perception according to claim 4, characterized in that: The process of preprocessing the obtained frame rate audio includes: Performing unequal audio division on the corresponding frame rate audio to obtain a plurality of continuous sampled audios, where the sampled audios are composed of a plurality of continuous data sampling points; The obtained sampled audio is subjected to adaptive signal transformation to obtain the corresponding sampled spectrum. The corresponding adaptive signal transformation formula is as follows: Where ni represents the i-th sampled audio corresponding to the n-th frame rate audio; FFT() represents the fast Fourier transform operation; δ n represents the windowing function; P ni Indicates the sampling spectrum of the i-th sampled audio corresponding to the n-th frame rate audio; X ni Indicates the total number of data sampling points in the i-th sampled audio corresponding to the n-th frame rate audio; * indicates convolution operation; Obtain the spectrum frequency within the corresponding sampling spectrum and perform logarithmic operation on it to obtain the corresponding logarithmic spectrum line; the formula for performing logarithmic operation is: Where, F(t) represents the corresponding logarithmic spectrum line; DCT() represents the discrete cosine transform operation, and pf represents the spectrum frequency in the corresponding sampled spectrum; The spectral coefficients corresponding to the collected audio are constructed based on the obtained logarithmic spectral lines. The formula for obtaining the corresponding spectral coefficients is: Where i is the index of the sampled audio, a is the logarithmic spectrum corresponding to the corresponding sampled audio, m represents the index of the number of pre-selected filters, and M is the total number of pre-selected filters; SP (i,a) represents the spectral coefficient corresponding to the i-th sampled audio; S(i, m) represents the i-th sampled audio after filtering by the m-th filter; Obtain the spectral coefficients corresponding to all sampled audio corresponding to the corresponding frame rate audio, and construct the corresponding frame rate spectrogram based on them.
6. The method for wireless optimized transmission of audio and video based on content perception according to claim 4, characterized in that: The construction process of the audio and video feature model includes: The backbone network of the audio and video feature model is defined as an improved convolutional neural network; the basic architecture of the improved convolutional neural network is an input layer, a convolution layer, a pooling layer, a feature fusion layer, and an output layer; The input layer is used to receive the input vector and perform image normalization on it; The convolution layer uses multi-scale convolution kernels to perform convolution operations on the input vector processed by the input layer to obtain the corresponding audio features and image features; The pooling layer is used to perform an adaptive pooling operation on the output vector of the convolutional layer; The feature fusion layer is used to perform feature fusion on the audio features and image features after the pooling operation to obtain corresponding audio and video fusion features; The output layer is provided with a fully connected layer, which uses the flatten algorithm to expand the audio and video fusion features, and then uses the fully connected layer to perform dimensionality reduction processing on the data; after the dimensionality reduction processing is completed, the result is output; Acquire several sets of historical audio and video data and audio and video feature information corresponding to the corresponding audio and video data; and construct a corresponding training data set based on the data; the training data set consists of several training samples; The obtained training data set is divided into a training set and a test set, and the corresponding audio and video feature model is iteratively trained based on the training set until the loss function tends to converge, the model parameters are saved, and the corresponding audio and video feature model is verified based on the test set. If the verification passes, the corresponding audio and video feature model is output and put into use. If the verification fails, the iterative training continues until it meets the requirements.
7. The method for wireless optimized transmission of audio and video based on content perception according to claim 6, characterized in that: The operation formula of the convolution operation is: Where Z l1+1 and Z l1 Represent the output vector and input vector of the l1+1th convolution layer respectively; (i, j) represents the dynamic area position during the convolution operation; ω l1+1 Refers to the connection weight between the l1 layer and the l1+1 layer, b l1+1 represents the bias term, Represents the convolution operation; The formula for the pooling operation is: Where s0 represents the step length; represents the position of the dynamic area after the pooling operation in the feature map corresponding to the k-th input vector in the l2-th pooling layer, where k is a natural number; x and y refer to the horizontal and vertical local offsets in the corresponding feature map; B1 and B2 are used for the horizontal and vertical sizes of the pooling window during the pooling operation, respectively; l2 represents the index of the pooling layer; p∈(1,∞), used to indicate the degree of pooling; The corresponding feature fusion process includes: Obtain the audio features and image features completed by the pooling operation, and perform two-dimensional convolution processing on them respectively to obtain the corresponding convolution audio Ka and Va and convolution image Qv; perform dot product operation on the obtained convolution audio Ka and Qv to obtain the corresponding dot product weight, and assign it to the convolution audio Va to obtain the corresponding initial fusion feature Fav; the acquisition formula of the corresponding initial fusion feature is: Where Fv represents the image feature, β represents the learning parameter, softmax is the pre-selected activation function; d represents the feature dimension of the corresponding convolution audio and convolution image; ⊙ represents the dot product operation; T represents the matrix transpose; Perform two-dimensional convolution processing on the obtained initial fusion features, and repeat the above initial fusion feature acquisition process to obtain the corresponding audio and video fusion features; Among them, the formula for two-dimensional convolution processing is: Ka = Fa*Wa; Qv = Fv*Wv; Va = Fa*Wb; Wa, Wv and Wb respectively represent the network weight matrices in the network model training process; Fa is the audio feature.
8. The method for wireless optimized transmission of audio and video based on content perception according to claim 1, characterized in that: The process of constructing and executing a corresponding wireless transmission strategy based on the received audio and video data includes: After the key frame marking is completed, the audio and video data after the key frame marking is encoded and compressed based on the pre-set second encoding scheme to obtain the corresponding second encoded audio; and the second encoded audio and the corresponding audio and video transmission task are uploaded to the server; When the server receives the corresponding second encoded audio, the server obtains transmission data during the transmission of the corresponding second encoded audio; The server obtains the device terminal numbers of the corresponding sending end and transmitting end based on the received audio and video transmission tasks, and constructs a corresponding temporary wireless transmission channel based on the device terminal numbers; and collects the status data of each network channel in the corresponding temporary wireless transmission channel to obtain corresponding channel network data; Pre-allocating audio and video network bandwidth for each frame rate within the corresponding audio and video data based on the obtained transmission data; Based on the obtained transmission data corresponding to the audio at each frame rate and the pre-allocated network bandwidth, and in combination with the corresponding channel network data, network channel allocation is performed to obtain the corresponding initial channel transmission strategy; After the allocation is completed, if the idle network bandwidth of the network channel allocated to the corresponding frame rate audio and video is not greater than the pre-allocated network bandwidth, the encoding compression scheme of the corresponding frame rate audio and video is changed to the pre-set third encoding scheme; if the idle network bandwidth of the network channel allocated to the corresponding frame rate audio and video is greater than the pre-allocated network bandwidth; the encoding compression scheme of the corresponding frame rate audio and video is changed to the pre-set first encoding scheme; After the change is completed, the corresponding initial channel transmission strategy is updated and returned to the sender for execution.
9. The method for wireless optimized transmission of audio and video based on content perception according to claim 8, characterized in that: The receiving end receives the transmitted audio and video data, verifies its integrity, and displays it to the user after verification. The process includes: When the receiving end recognizes that the corresponding temporary wireless transmission channel is established, the receiving end constructs a cache node; a cache slot is set in the cache node; when the receiving end receives the corresponding coded audio, it decodes it, obtains the timestamp in the decoded coded audio, and verifies whether there is an identical timestamp in an existing cache slot in the corresponding cache node; if not, it stores the timestamp in the corresponding cache slot and constructs a new cache slot; if it exists, it ignores the corresponding coded audio, wherein the coded audio is one of the first coded audio, the second coded audio, and the third coded audio; After the transmission is completed, the cache node will merge the decoded encoded audio based on the timestamp and key frame mark to obtain the corresponding first audio and video, and perform frame rate sequence verification and frame rate error verification on it; if there is no corresponding frame rate sequence error and frame rate error verification, the cache node will play the frame rate audio and video in the received first audio and video before and after the timestamp. If there is an error, the frame rate position of the missing frame rate audio and video is obtained and fed back to the server, and the server will retransmit the corresponding frame rate audio and video, and so on.
10. A content-aware audio and video wireless optimization transmission system, which is used to implement the content-aware audio and video wireless optimization transmission method according to any one of claims 1 to 9, characterized in that: include: The data acquisition module is used to obtain the audio and video data to be transmitted and divide it into frame rates to obtain several frame rate audio and video; A data processing module is used to process the obtained frame rate audio and video data and obtain corresponding audio and video feature information based on the data processing results; A transmission strategy module is used to mark the key frames of audio and video at each frame rate in the corresponding audio and video data based on the audio and video feature information, and to construct a corresponding channel transmission strategy based on the key frame marking; The user display module is used to receive the transmitted audio and video data, perform integrity verification on it, and display it to the user after the verification is passed.
Citation Information
Patent Citations
Real-time video transmission method based on WebRTC
CN117834984A