A video transmission method and system based on streaming processing
Through multi-level feature extraction and keyframe selection of video frames, combined with dynamic path optimization technology and error hiding technology, the shortcomings of existing video transmission methods in path optimization, traffic allocation and picture quality optimization are solved, and efficient and stable transmission of high-definition video streams are achieved.
Patent Information
- Application Number
- CN202510134163.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The existing video transmission methods have limited performance in efficient path optimization, dynamic traffic allocation and picture quality optimization, and are less adaptable to complex video content and non-ideal network environments.
Through multi-level feature extraction and keyframe selection of video frames, combined with dynamic path optimization technology, the Q learning algorithm and luminescent firefly algorithm are used to dynamically adjust path selection and traffic allocation, improve transmission efficiency, and ensure high-fidelity presentation of video streams through error hiding and image quality enhancement technologies.
It realizes efficient and stable transmission of video data, improves transmission efficiency and image quality optimization effects, adapts to complex video content and non-ideal network environments, and provides a brand new video transmission solution.
Smart Images

Figure CN119562138B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video transmission, and particularly to a video transmission method and system based on streaming processing. Background Art
[0002] With the rapid development of Internet and multimedia technologies, video transmission has become an important field in data communication. In traditional methods of video transmission, video stream protocols based on fixed bitrate (such as RTMP, HLS) are widely used for the transmission of real-time video streams. However, the fixed bitrate method is difficult to adapt to complex network environments. When the network bandwidth fluctuates, it is easy to cause problems such as stuttering and degraded video quality. In recent years, the introduction of streaming processing technology has provided new solutions for the real-time performance and stability of video transmission. Streaming processing can fragment data streams and, combined with network dynamic adjustment algorithms, make video transmission more flexible and efficient. To further improve the transmission quality, various optimization strategies have been proposed in the academic and industrial communities. For example, transmission methods based on adaptive streaming technology can dynamically adjust video quality and bitrate by real-time monitoring of network status, significantly improving network utilization efficiency. At the same time, the application of deep learning technology makes the preprocessing and feature extraction of video frames more refined, providing support for complex tasks such as semantic recognition and scene detection. However, the performance of existing solutions is still limited in aspects such as efficient path optimization, dynamic traffic allocation, and video quality optimization, and their adaptability to complex video content and non-ideal network environments is weak. Summary of the Invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides a video transmission method and system based on streaming processing, which solves the problems that the performance of existing solutions is still limited in aspects such as efficient path optimization, dynamic traffic allocation, and video quality optimization, and their adaptability to complex video content and non-ideal network environments is weak.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] In the first aspect, the present invention provides a video transmission method based on streaming processing, which includes:
[0007] Collect video streams, extract video frames after preprocessing, encapsulate the video frames into picture streams, perform data fragmentation on the picture streams, allocate initial paths for the fragmented data, and calculate the initial traffic ratio;
[0008] Optimize path selection according to network feedback and adjust path traffic allocation to obtain a path optimization result, and transmit the fragmented data according to the path optimization result;
[0009] After receiving the sharded data, the sharded data is recombined into a video stream and the video quality is optimized, and the transmitted video data is stored.
[0010] As a preferred solution of the video transmission method based on streaming processing according to the present invention, wherein: the step of collecting a video stream, preprocessing it, and extracting video frames means using a high-resolution imaging device to capture a high-definition video stream in real time at a frame rate of 30fps, storing the collected video stream in a circular buffer in units of frames, marking the timestamps of the video frames and sorting them according to the timestamps to form a frame sequence, and performing preprocessing on the frame sequence;
[0011] For each frame image A, pixel values are extracted from the RGB channels, and color moments are calculated, including extracting the mean value C1, variance C2, and skewness C3 of the color, and combining C1, C2, and C3 to form a color feature vector;
[0012] Each frame image is converted into a grayscale image, the co-occurrence matrix of each pair of pixels is calculated based on the neighborhood relationship, and the texture features of each frame image A are calculated, including contrast B1, homogeneity B2, and correlation B3, and combining B1, B2, and B3 to form a texture feature vector;
[0013] Based on the grayscale image of texture analysis, the occurrence frequency of each grayscale value is statistically analyzed, normalized to a probability distribution p(f), and the Shannon entropy E of each frame image is calculated based on p(f) f , measuring the information complexity of the image;
[0014] Load the pre-trained VGG16 model, freeze the weights of the convolutional layers and select the Conv1 to Conv5 layers as the multi-layer feature output layers. Each layer extracts low-level structural features and high-level semantic features respectively, and inputs the frame sequence into the VGG16 model to extract features layer by layer;
[0015] Calculate the gradient magnitude of the frame sequence and statistically analyze the frequency distribution of the gradient magnitude to generate a multi-channel histogram H M ;
[0016] Combine the multi-layer convolutional features of VGG16 with the HGMF-MC features to obtain a fused deep feature vector F, fuse the deep feature vector F with the color, texture, and information complexity features to obtain the final deep feature vector, and store the feature vectors of all frames as a feature vector sequence;
[0017] Use a convolutional neural network as a classification model, input the comprehensive feature vector of each frame into the classification model, calculate the classification probability, set a classification threshold D. If the classification probability is greater than or equal to the classification threshold D, then mark the current frame as a semantic scene frame, otherwise mark the current frame as a static scene frame, and mark the frame sequence as a semantic scene frame set and a static scene frame set;
[0018] In the set of computed semantic scene frames, calculate the gradient components and gradient magnitudes of each frame of the image. Divide the gradient magnitude of each frame into multiple channels, and perform bucket statistics on the gradient magnitude of each channel to construct the histogram HGMF;
[0019] Calculate the difference values of the histograms HGMF for each pair of adjacent frames one by one, calculate the mean and standard deviation of all inter-frame differences, and set the threshold s according to the sum of the mean and standard deviation;
[0020] Traverse the set of semantic scene frames, select the frames with inter-frame difference values greater than the threshold s as key frames, and add timestamps to each key frame;
[0021] Perform 2D-DWT on each frame of the image in the set of static scene frames to extract the low-frequency feature S;
[0022] Use the perceptual hashing method to generate hash values for each frame of low-frequency features, calculate the differences between adjacent frame hash values through the Hamming distance, calculate the mean and standard deviation of all hash value differences, and set the threshold g according to the sum of the mean and standard deviation;
[0023] Traverse the hash value difference sequence, select the frames with local extrema greater than g as key frames, and add timestamps to each key frame;
[0024] Dynamically calculate the frame extraction frequency F according to the actual network bandwidth, number of windows, and video content;
[0025] Sort the set of key frames according to the timestamps, select a frame sequence for key frame extraction according to the dynamically calculated frame extraction frequency, and combine the extracted key frames into a set.
[0026] As a preferred solution of the video transmission method based on streaming processing according to the present invention, wherein: after encapsulating the video frames into a picture stream, perform data sharding, allocate an initial path for the sharded data and calculate the initial traffic ratio means encapsulating the video frames into the JPG picture format through FFmpeg, integrating the picture frames to form a picture stream, setting a time window to segment the picture stream, dividing the picture frames within the time window into one shard, and each shard is marked as F k where k represents the current shard number;
[0027] Add timestamps to each shard, set a priority identifier for each shard according to the data transmission requirements, generate a CRC check code for the complete data content of each shard, append the check code to the end of the shard to form a complete shard structure, and obtain a sequence of sharded data with completion marks;
[0028] Collect the network status information of all current paths and normalize the data, use a network performance probe to update the path status in real time, and store the monitoring results as a path attribute table;
[0029] Calculate the comprehensive score of each path by the weighted fusion method, sort the paths in descending order according to the comprehensive score, and select the first m paths as the effective path set;
[0030] For the filtered path set, use the Pareto distribution model to calculate the initial traffic allocation ratio of each path in the set;
[0031] According to the shard priority identifier, adjust the path allocation ratio, and generate a path allocation table by combining the path network status information and the adjusted allocation ratio.
[0032] As a preferred scheme of the video transmission method based on streaming processing described in the present invention, wherein: the obtaining the path optimization result by optimizing the path selection according to the network feedback and adjusting the path traffic allocation means calculating the reward value R of each path according to the network status;
[0033] Use the Q-learning algorithm to dynamically update the path value, select the optimal path according to the path value and generate the optimal path set;
[0034] Calculate the delay difference between paths i and j and calculate the path attractiveness according to the delay difference ;
[0035] According to the luminous firefly algorithm, dynamically adjust the path traffic ratio ;
[0036] Normalize all path traffic ratios, iteratively update the path traffic ratios, set the maximum number of iterations, and stop iterating until the maximum number of iterations is reached to obtain the final path traffic ratio. According to the final traffic ratio and reward value of the path, generate the final optimization weight G for each path;
[0037] Divide the shard data into a high-priority data set and a secondary-priority data set according to the priority identifier. Sort the path set in descending order according to the optimization weight, set the thresholds U and O, and U>O. If the path weight is greater than or equal to the threshold U, it is a high-priority path group and transmits the high-priority data set. If the path weight is greater than or equal to the threshold O and less than the threshold U, it is a medium-priority path group and transmits the secondary-priority data set. If the path weight is less than the threshold O, it is a low-priority path group and serves as a backup path to be used when congestion occurs in the high-priority path group and the medium-priority path group;
[0038] Integrate all path allocation results to generate the final optimized path allocation table.
[0039] As a preferred solution of the video transmission method based on streaming processing according to the present invention, wherein: transmitting the shard data according to the path optimization result means traversing each data shard, selecting the target path group according to the priority, initializing the path scheduler, transmitting the shards in the high-priority dataset to the target receiver through a low-latency transmission protocol, transmitting the shards in the secondary-priority dataset to the target receiver through a bandwidth optimization protocol, recording the transmission results of each path to generate a transmission log and storing it in the database.
[0040] As a preferred solution of the video transmission method based on streaming processing according to the present invention, wherein: after receiving the shard data, reorganizing the shard data into a video stream and performing video quality optimization means receiving the shard data transmitted by each path through the transmission protocol and using cyclic redundancy check to detect the integrity of each shard, recording the shard numbers with failed checks and requesting retransmission from the sender. If the retransmission fails, interpolation compensation is performed using adjacent shards;
[0041] Sorting the shard data according to the shard number and timestamp, merging the shard data of all paths by number into a complete time sequence to obtain a complete data frame sequence, initializing the video decoder at the receiver, decoding the data frame sequence frame by frame to generate a video stream, and using error concealment technology to correct frame errors that occur during the decoding process;
[0042] Using bilateral filtering to remove noise and retain edge details for the decoded image frame sequence, using histogram equalization to enhance the image contrast, and performing detail enhancement through inverse wavelet transform, and re-encoding the optimized image frame sequence into a complete video stream.
[0043] As a preferred solution of the video transmission method based on streaming processing according to the present invention, wherein: storing the transmitted video data means compressing the video stream using context-adaptive binary arithmetic coding, storing the compressed video stream in layers, and uploading the data to the cloud for backup after storage.
[0044] In a second aspect, the present invention provides a video transmission system based on streaming processing, including,
[0045] A data collection module for collecting video streams and performing preprocessing;
[0046] A data sharding and path initialization module for encapsulating the preprocessed video frames into a picture stream and performing data sharding, allocating initialization paths for the shards and calculating the initial traffic ratio;
[0047] A path optimization module for dynamically optimizing path selection based on network feedback and adjusting path traffic allocation;
[0048] A data transmission module for efficiently transmitting the shard data to the receiver according to the path optimization result;
[0049] The video quality optimization module is used to receive fragmented data, reorganize it into a video stream, and optimize the video quality at the same time.
[0050] The storage backup module is used to compress and store the transmitted video stream and upload it to the cloud for backup.
[0051] Thirdly, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the video transmission method based on streaming processing described in the first aspect of the present invention is implemented.
[0052] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the video transmission method based on streaming processing described in the first aspect of the present invention is implemented.
[0053] The beneficial effects of the present invention are as follows: through multi-level feature extraction and key frame selection of video frames, the present invention realizes the fragmentation encapsulation of video data. Combining with the dynamic path optimization technology, the path selection and traffic allocation are dynamically adjusted based on the Q-learning algorithm and the glowworm swarm optimization algorithm, improving the transmission efficiency. The video image processing and dynamic network optimization technologies are fully combined, providing a new solution for the efficient and stable transmission of high-definition video streams. Description of the Drawings
[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 It is a flowchart of the video transmission method based on streaming processing in Embodiment 1;
[0056] Figure 2 It is a schematic diagram of the video transmission system based on streaming processing in Embodiment 1. Detailed Embodiments
[0057] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed embodiments of the present invention in conjunction with the drawings of the specification.
[0058] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways than those specifically described herein. Those skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0059] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other from other embodiments.
[0060] Example 1, referring to Figure 1 and Figure 2 , which is the first embodiment of the present invention. This embodiment provides a video transmission method based on streaming processing, including the following steps:
[0061] S1. Collect the video stream, extract video frames after preprocessing, encapsulate the video frames into a picture stream, perform data sharding on the picture stream, allocate an initial path for the sharded data, and calculate the initial traffic ratio;
[0062] Specifically, collecting the video stream and extracting video frames after preprocessing means using a high-resolution camera device to capture a high-definition video stream in real time at a frame rate of 30fps, storing the captured video stream in a circular buffer in units of frames, marking the timestamps of the video frames, sorting them according to the timestamps to form a frame sequence, and performing preprocessing on the frame sequence;
[0063] The preprocessing refers to using the local difference detection method to detect salt-and-pepper noise, and marking the video frames with detected salt-and-pepper noise as type 1, and other noises as type 2;
[0064] Apply median filtering to the frames marked as type 1 to remove salt-and-pepper noise, and apply bilateral filtering to the frames marked as type 2 to retain edge details;
[0065] Through salt-and-pepper noise detection and various filtering methods (such as median filtering and bilateral filtering), effectively remove image noise, retain the edge details of the frames, improve the quality of the frames, and ensure the temporal consistency of the frame sequence through timestamp marking and sorting, supporting subsequent dynamic optimization strategies based on time information;
[0066] For each frame image A, extract pixel values from the RGB channels, calculate the color moments, including extracting the mean C1, variance C2, and skewness C3 of the color, and combine C1, C2, and C3 to form a color feature vector;
[0067] Convert each frame of the image into a grayscale image, calculate the co-occurrence matrix of each pair of pixels based on the neighborhood relationship, and calculate the texture features of each frame of image A, including contrast B1, homogeneity B2, and correlation B3. Combine B1, B2, and B3 to form a texture feature vector. The calculation formulas for B1, B2, and B3 are as follows:
[0068] ,
[0069] ,
[0070] ,
[0071] In the formula, ij is the grayscale level index, and P(i, j) is the joint probability at position (i, j) in the grayscale co-occurrence matrix. and are the normalized means of the grayscale levels, is the normalized standard deviation of the grayscale levels;
[0072] The RGB color moment and the grayscale co-occurrence matrix provide multi-dimensional frame description information, enhancing the discrimination ability of video frames in scene changes. High-quality frames achieve in-depth analysis of information through multi-level feature extraction, and color and texture features provide strong supplements for depth features, improving the classification accuracy. The preprocessed high-quality frames provide accurate basic data such as color, texture, and timestamps for key frame extraction, ensuring the precise selection of key frames;
[0073] Based on the grayscale image of texture analysis, count the occurrence frequency of each grayscale value, normalize it to the probability distribution p(f), and calculate the Shannon entropy E of each frame of the image based on p(f) f , measuring the information complexity of the image:
[0074]
[0075] Take the Shannon entropy as the information complexity feature;
[0076] Load the pre-trained VGG16 model, freeze the weights of the convolutional layers, and select the Conv1 to Conv5 layers as the multi-level feature output layers. Extract low-level structural features and high-level semantic features respectively for each layer. Input the frame sequence into the VGG16 model and extract features layer by layer:
[0077] Conv1 layer: Capture the edge and contour information of the image;
[0078] Conv2 layer: Extract local texture features;
[0079] Conv3 layer: Generate middle-level shape and region features;
[0080] Conv4 layer: Capture complex image structure information;
[0081] Conv5 layer: aggregates high-level semantic features;
[0082] Calculate the gradient amplitude of the frame sequence and count the frequency distribution of the gradient amplitude to generate a multi-channel histogram H M ;
[0083] From Conv1 to Conv5 layers, low, medium, and high-level features are extracted respectively to ensure that information of different granularities is fully utilized. By freezing the weights of the convolutional layer, the capabilities of the pre-trained model are fully utilized to capture the semantic hierarchical information of the image and improve the recognition ability of the classification model. The multi-channel histogram HGMF and color and texture features are combined to generate a fused deep feature vector F, which improves the comprehensiveness and expressiveness of the features.
[0084] Combine the multi-layer convolutional features of VGG16 with the HGMF-MC features to obtain the fused deep feature vector F, fuse the deep feature vector F with the color, texture, and information complexity features to obtain the final deep feature vector, and store the feature vectors of all frames as a feature vector sequence;
[0085] The extraction of deep features depends on high-quality frames. Frames that have undergone noise removal and feature extraction can better reflect their low-level and high-level semantic features. Multi-level deep features provide rich semantic information support for frame classification and key frame extraction, enhance the accurate evaluation of inter-frame differences, and promote the accuracy of key frame selection. The fusion of deep features, color features, texture features, and information complexity features makes the stored feature sequences more distinguishable and applicable.
[0086] Use convolutional neural network as the classification model, input the comprehensive feature vector of each frame into the classification model, calculate the classification probability, set the classification threshold D based on the ROC curve, and if the classification probability is greater than or equal to the classification threshold D, mark the current frame as a semantic scene frame, otherwise mark the current frame as a static scene frame, and mark the frame sequence as a semantic scene frame set and a static scene frame set;
[0087] By classifying semantic scene frames and static scene frames, unnecessary static data is reduced from participating in subsequent transmission and optimization, thus improving efficiency. Based on the comprehensive feature vector input, the classification performance of the convolutional neural network is better than that of the traditional classifier, and the classification accuracy is higher. By classifying scenes, the amount of data that needs to be transmitted and stored is reduced, and the efficiency of system resource utilization is improved. The classification model directly uses the generated deep feature vectors, and the classification performance depends on the comprehensiveness and accuracy of the feature expression. The classified semantic scene frame set and static scene frame set provide basic data for key frame extraction, ensuring that the key frame selection is concentrated in scenes with significant semantic changes, thus reducing the interference of meaningless data.
[0088] In the set of computed semantic scene frames, calculate the gradient components and gradient magnitudes of each frame image. Divide the gradient magnitude of each frame into multiple channels, and perform bucket statistics on the gradient magnitude of each channel to construct the histogram HGMF;
[0089] Calculate the difference values of the histograms HGMF for each pair of adjacent frames pairwise, calculate the mean and standard deviation of all inter-frame differences, and set the threshold s according to the sum of the mean and standard deviation;
[0090] Traverse the set of semantic scene frames, select the frames with inter-frame difference values greater than the threshold s as key frames, and add timestamps to each key frame;
[0091] Perform 2D-DWT on each frame image in the set of static scene frames to extract the low-frequency feature S:
[0092]
[0093] where H and W are the height and width of the image, f(x, y) is the gray value of the image, is the Haar function, is the currently used wavelet scale, h is the translation parameter of the wavelet function in the vertical direction, w is the translation parameter of the wavelet function in the horizontal direction, and x and y are the position indices of the pixels in the horizontal and vertical directions;
[0094] Use the perceptual hashing method to generate hash values for each frame of low-frequency features, calculate the differences between adjacent frame hash values through the Hamming distance, calculate the mean and standard deviation of all hash value differences, and set the threshold a according to the sum of the mean and standard deviation;
[0095] Traverse the hash value difference sequence, select the frames with local extrema greater than a as key frames, and add timestamps to each key frame;
[0096] The key frame selection strategy based on the inter-frame gradient magnitude difference and hash value difference ensures that the key frames reflect the main information of semantic changes;
[0097] Dynamically calculate the frame extraction frequency F according to the actual network bandwidth, number of windows, and video content:
[0098]
[0099] where L is the total frame processing capacity, determined by the maximum processing capacity supported by the system hardware (CPU, GPU), q is the number of parallel windows, and Y is the priority factor of the key frames;
[0100] Sort the set of key frames according to the time stamp, select a frame sequence for key frame extraction based on the dynamically calculated frame extraction frequency, and combine the extracted key frames into a set. The dynamic frame extraction frequency method adjusts the number of key frames dynamically according to the network bandwidth, improves the flexibility of transmission, and only transmits the set of key frames, significantly reducing data redundancy and optimizing the bandwidth utilization rate. The key frame extraction depends on the classification results of the semantic scene frame set and the static scene frame set. The key frames extracted from the semantic scene frames can effectively reflect the main changes in the video content. The extracted set of key frames provides a high-value data source for subsequent feature fusion, transmission, and storage.
[0101] Combining a deep learning model and a dynamic optimization algorithm not only comprehensively extracts multi-dimensional features of images (such as color, texture, information complexity, etc.), but also significantly reduces data redundancy through semantic scene frame classification and key frame extraction. The dynamic frame extraction frequency adjustment and path optimization strategy accurately adapt to the volatility of the network bandwidth, effectively improving the reliability and efficiency of transmission. At the same time, through error concealment and image quality enhancement technologies, high-fidelity presentation after video recombination is ensured. The present invention not only solves the problems of low transmission efficiency and insufficient image quality optimization in the prior art, but also realizes the multi-scenario adaptability of video data transmission and processing, providing an innovative and systematic solution for streaming media transmission and intelligent video analysis.
[0102] Further, after encapsulating the video frames into a picture stream, perform data sharding, allocate an initial path for the sharded data and calculate the initial traffic ratio. Encapsulate the video frames into the JPG picture format through FFmpeg, integrate the picture frames to form a picture stream. FFmpeg is a widely used open-source multimedia processing framework that supports functions such as audio and video recording, conversion, streaming, and playback. Use FFmpeg to extract and encapsulate the frames in the video stream into the JPG picture format to form a standardized picture stream, which is convenient for subsequent sharding operations. Set a time window to split the picture stream, divide the picture frames within the time window into a shard, and each shard is marked as F k , where k represents the current shard number. After data sharding, identify through the number F k to track the status of each shard, which is convenient for precise retransmission in case of data loss or delay;
[0103] Add a time stamp to each shard and set a priority identifier for each shard according to the data transmission requirements.
[0104] Generate a CRC checksum for the complete data content of each shard, append the checksum to the end of the shard to form a complete shard structure, and obtain a shard data sequence with completion markers. Timestamps ensure the order of shard data, and priority identifiers support dynamically allocating transmission resources according to the importance of the content. For example, allocate a high-priority path for key-frame shards to ensure the timely transmission of key frames. By appending the CRC checksum to the end of the shard, the receiving end can quickly verify the integrity of the data shard. If an error is found, it can request retransmission or correct it through redundant data, improving the reliability of data transmission;
[0105] Collect the network status information of all current paths and normalize the data. Use network performance probes to update the path status in real time and store the monitoring results as a path attribute table;
[0106] The network status information includes bandwidth, latency, and packet loss rate;
[0107] Calculate the comprehensive score of each path by using a weighted fusion method for the network status information, sort the paths in descending order according to the comprehensive score, and select the top m paths as the effective path set;
[0108] For the filtered path set, use the Pareto distribution model to calculate the initial traffic allocation ratio of each path in the set. The Pareto distribution is a probability distribution model commonly used to describe the imbalance of path traffic allocation in network traffic. Using the Pareto distribution model to calculate the initial traffic ratio can effectively balance the preferential utilization of high-efficiency paths and the backup allocation of redundant paths, preventing the occurrence of transmission bottlenecks;
[0109] According to the shard priority identifier, adjust the path allocation ratio, generate a path allocation table by combining the path network status information and the adjusted allocation ratio. By combining the path network status information and the priority identifier, dynamically adjusting the path allocation ratio can adapt to real-time network fluctuations and ensure the high-priority transmission requirements of important data.
[0110] By encapsulating video frames into a picture stream and performing shard transmission, the present invention effectively overcomes the deficiencies of traditional video transmission technologies in a network environment with a high packet loss rate and large bandwidth fluctuations. The addition of shard numbers and timestamps ensures the traceability and order of shards; the combination of priority identifiers and dynamic path optimization algorithms significantly improves the stability and efficiency of data transmission. At the same time, the integrity detection of CRC checks and the traffic allocation strategy of the Pareto distribution make the transmission process more reliable and efficient. The present invention is applicable to scenarios with strong multi-path transmission optimization requirements and has a wide range of application prospects.
[0111] S2. Optimize path selection according to network feedback and adjust path traffic allocation to obtain a path optimization result, and transmit shard data according to the path optimization result;
[0112] Specifically, optimizing the path selection according to network feedback and adjusting the path traffic distribution to obtain the path optimization result means calculating the reward value R of each path based on the network state:
[0113]
[0114] In the formula, is the delay of path a, is the packet loss rate of path a, is the bandwidth utilization rate of path a, and are the weight factors of the packet loss rate and the bandwidth utilization rate, which are adjusted and set according to application requirements;
[0115] Use the Q-learning algorithm to dynamically update the path value, select the optimal path according to the path value, and generate the optimal path set;
[0116] Calculating the reward value through the weighted fusion of delay, packet loss rate, and bandwidth utilization rate can accurately measure the path performance in multiple dimensions;
[0117] Calculate the delay difference between paths i and j and calculate the path attractiveness according to the delay difference :
[0118]
[0119] In the formula, is the initial attractiveness, is the light intensity attenuation coefficient, which is set through historical experience to control the attractiveness attenuation speed, is the delay difference between path i and path j, reflecting the difference degree of the performance of the two paths;
[0120] According to the luminous firefly algorithm, dynamically adjust the path traffic ratio :
[0121]
[0122] In the formula, and are the current traffic ratios of paths i and j, is the path attractiveness, is the random perturbation factor, is the generated normal distribution random number with a mean of 0, as the variance;
[0123] Dynamically adjusting the traffic ratio through the path attractiveness can achieve load balancing in multi-path transmission, avoiding resource waste or path congestion;
[0124] Normalize all path traffic ratios, iteratively update the path traffic ratios, set the maximum number of iterations, and stop iterating until the maximum number of iterations is reached to obtain the final path traffic ratios. Based on the final traffic ratios and reward values of the paths, generate the final optimized weight G for each path:
[0125] ,
[0126] Classify the sharded data into high-priority datasets and secondary-priority datasets according to the priority identifiers. Sort the path set in descending order according to the optimized weights. Set thresholds U and O through statistical analysis of historical data, and U > O. If the path weight is greater than or equal to threshold U, it is a high-priority path group for transmitting high-priority datasets. If the path weight is greater than or equal to threshold O and less than threshold U, it is a medium-priority path group for transmitting secondary-priority datasets. If the path weight is less than threshold O, it is a low-priority path group as a backup path to be used when congestion occurs in the high-priority path group and the medium-priority path group;
[0127] Dividing the high-priority path group, medium-priority path group, and backup path group through the path weight thresholds can achieve differentiated scheduling of resources;
[0128] Integrate all path allocation results to generate the final optimized path allocation table;
[0129] The path allocation table includes path numbers, traffic ratios, reward values, comprehensive weights, priority groupings, and path traffic.
[0130] Through core steps such as reward value calculation, path attractiveness modeling, and optimization using the glowworm swarm optimization algorithm, the problems of path selection, traffic allocation, and priority scheduling in multi-path transmission scenarios are systematically solved. Combining the path allocation table and the dynamic priority grouping mechanism, this method can not only achieve efficient video data transmission in complex network environments but also significantly improve the stability and robustness of the system.
[0131] Further, transmitting sharded data according to the path optimization result means traversing each data shard, selecting the target path group according to the priority, improving the transmission success rate of high-priority shards, ensuring that critical data is preferentially guaranteed, dynamically adapting to changes in the network environment, effectively avoiding path congestion, improving the transmission efficiency, initializing the path scheduler, providing dynamic path scheduling capabilities, flexibly adapting to real-time changes in the network state, reducing the residence time of data shards in the low-priority path group, improving the overall system throughput, transmitting the shards in the high-priority data set to the target receiver through a low-latency transmission protocol (UDP) to ensure that the data transmission of critical tasks meets real-time requirements, such as the key frames in a video call or the transmission of high-priority service data, transmitting the shards in the secondary-priority data set to the target receiver through a bandwidth optimization protocol (Link Aggregation Control Protocol LACP) to enhance the bandwidth utilization of the system and reduce resource waste caused by link idleness, and recording the transmission results of each path to generate a transmission log and storing it in the database;
[0132] The transmission log includes the number of successfully transmitted shards, the packet loss rate, and the transmission delay.
[0133] The dynamic sharded transmission method based on the path optimization result effectively solves the problems of static path selection, insufficient priority scheduling, and lack of transmission performance monitoring in the prior art, provides low-latency guarantee for high-priority data transmission, and fully utilizes network bandwidth resources to complete secondary-priority data transmission. At the same time, the system establishes a performance feedback mechanism through the transmission log, further enhancing the intelligence and optimization capabilities of the transmission system, and providing strong technical support for the video transmission of streaming processing.
[0134] S3. After receiving the sharded data, reconstruct the sharded data into a video stream and perform picture quality optimization, and store the transmitted video data;
[0135] Specifically, reconstructing the sharded data into a video stream and performing picture quality optimization after receiving the sharded data means receiving the sharded data transmitted by each path through the transmission protocol and using cyclic redundancy check (CRC) to detect the integrity of each shard, ensuring the integrity of data transmission, quickly identifying transmission errors, improving transmission reliability, and laying a foundation for video stream reconstruction, recording the shard numbers with verification failures and requesting retransmission from the sender. If the retransmission fails, interpolation compensation is performed using adjacent shards. The retransmission mechanism provides the first line of security defense to ensure that repairable errors are minimized. When the retransmission fails, the video data is restored to the greatest extent through the interpolation compensation mechanism to avoid the interruption of the overall picture caused by the loss of a single shard;
[0136] Sort the shard data according to the shard number and timestamp, merge the shard data of all paths into a complete time series according to the number to obtain a complete data frame series, reorganize the transmitted shards into a complete time series according to the shard number and timestamp to ensure the order consistency of the frame data, initialize the video decoder at the receiving end, decode the data frame series frame by frame to generate a video stream, initialize the video decoder and decode frame by frame to ensure the efficient restoration of the video stream, provide accurate frame-level error detection capabilities, provide input for error concealment techniques, and use error concealment techniques to correct frame errors during the decoding process;
[0137] The error correction includes spatial domain correction and temporal domain correction;
[0138] The spatial domain correction refers to filling the error area with the pixel information of adjacent frames, and the temporal domain correction refers to interpolation repair through adjacent frames on the time axis. The spatial domain correction directly fills the pixel area of the error frame to reduce obvious visual defects. The temporal domain correction uses the time series characteristics to restore the dynamic picture, improving the naturalness and coherence after repair. The comprehensive correction method improves the video quality and is applicable to different scenarios and error types;
[0139] Use bilateral filtering to remove noise and retain edge details for the decoded image frame series, use histogram equalization to enhance the image contrast, and perform detail enhancement through inverse wavelet transform. Re-encode the optimized image frame series into a complete video stream.
[0140] Through the CRC checksum and retransmission mechanism, ensure the high integrity and accuracy of data transmission. The interpolation compensation mechanism effectively copes with irrecoverable data loss and improves the availability of the video stream. From image preprocessing to error correction, combine multiple optimization techniques to improve the image quality. Through filtering and enhancement techniques, improve the clarity and contrast of the decoded video. Dynamically adjust the retransmission and correction strategies to adapt to different network conditions and ensure the continuity and stability of video stream transmission. Cover the entire process from data reception to reorganization, decoding and then optimization, build an efficient video transmission method, provide a targeted solution, and meet the needs of multiple scenarios and multiple terminals.
[0141] Furthermore, storing the transmitted video data means compressing the video stream using context-adaptive binary arithmetic coding, storing the compressed video stream in layers, and uploading the data to the cloud for backup after storage.
[0142] An efficient video data storage solution is constructed through three core steps: CABAC compression, hierarchical storage, and cloud backup. The CABAC technology significantly improves the compression efficiency of video data, reduces transmission costs. The hierarchical storage strategy optimizes the utilization rate of storage resources and enhances the access efficiency of key data. Cloud backup enhances data security and scalability, meeting the diverse needs of modern video data management.
[0143] This embodiment also provides a video transmission system based on streaming processing, including:
[0144] A data collection module for collecting video streams and performing preprocessing;
[0145] A data sharding and path initialization module for encapsulating preprocessed video frames into picture streams, performing data sharding, allocating initialization paths for the shards, and calculating the initial traffic ratio;
[0146] A path optimization module for dynamically optimizing path selection based on network feedback and adjusting path traffic allocation;
[0147] A data transmission module for efficiently transmitting the sharded data to the receiving end according to the path optimization result;
[0148] A video quality optimization module for receiving the sharded data, recombining it into a video stream, and optimizing the video quality at the same time;
[0149] A storage and backup module for compressing and storing the transmitted video stream and uploading it to the cloud for backup.
[0150] This embodiment also provides a computer device applicable to the case of a video transmission method based on streaming processing, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the video transmission method based on streaming processing as proposed in the above embodiment.
[0151] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0152] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the video transmission method based on streaming processing as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0153] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A video transmission method based on streaming processing, characterized in that: include, After collecting the video stream for preprocessing, the color feature vector and texture feature vector are constructed, the Shannon entropy of each frame is calculated and used as the information complexity feature, the multi-layer convolution features of the image are extracted through the VGG16 model, the extracted and calculated features are fused to obtain the final deep feature vector, the video frames are divided into semantic scene frames and static scene frames through the convolutional neural network, the key frames of the two types of video frames are extracted respectively, and the frame extraction frequency is calculated according to the real-time data for key frame extraction; After encapsulating the video frames into image streams, data is segmented, initialization paths are assigned to the segmented data, and high-priority paths are assigned to key frame segments before calculating the initial traffic ratio; Optimize path selection and adjust path traffic distribution according to network feedback to obtain path optimization results, and transmit shard data according to the path optimization results; After receiving the fragmented data, the fragmented data is reassembled into a video stream and the image quality is optimized, and the transmitted video data is stored; The optimizing path selection and adjusting the path flow distribution according to the network feedback to obtain the path optimization result refers to calculating the reward value R of each path according to the network status; Use Q learning algorithm to dynamically update path value, select the optimal path according to the path value and generate the optimal path set; Calculate the delay difference between path pair i and j and calculate the path attractiveness based on the delay difference ; Dynamically adjust the path flow ratio according to the glowing firefly algorithm ; Normalize the flow ratios of all paths, iterate and update the flow ratios of the paths, set the maximum number of iterations, and stop the iteration until the maximum number of iterations is reached to obtain the final flow ratio of the paths. Generate the final optimization weight G for each path based on the final flow ratio and reward value of the path. The shard data is divided into a high priority data set and a second priority data set according to the priority identifier, and the path set is sorted in descending order according to the optimization weight. Thresholds U and O are set, and U>O. If the path weight is greater than or equal to the threshold U, it is a high priority path group, which transmits the high priority data set. If the path weight is greater than or equal to the threshold O and less than the threshold U, it is a medium priority path group, which transmits the second priority data set. If the path weight is less than the threshold O, it is a low priority path group, which is used as a backup path when congestion occurs in the high priority path group and the medium priority path group. Integrate all path allocation results to generate the final optimized path allocation table.
2. The video transmission method based on streaming processing as claimed in claim 1, characterized in that: The collecting of video streams and preprocessing and then extracting video frames refers to using a high-resolution camera device to capture a high-definition video stream in real time at a frame rate of 30fps, storing the collected video stream in a ring buffer in units of frames, marking the timestamps of the video frames and sorting them by the timestamps to form a frame sequence, and preprocessing the frame sequence; For each frame of image A, pixel values are extracted from the RGB channels and color moments are calculated, including extracting the color mean C1, variance C2 and skewness C3, and combining C1, C2 and C3 to form a color feature vector; Convert each frame of image into a grayscale image, calculate the co-occurrence matrix of each pair of pixels based on the neighborhood relationship, calculate the texture features of each frame of image A, including contrast B1, homogeneity B2 and correlation B3, and combine B1, B2 and B3 to form a texture feature vector; Based on the grayscale image of texture analysis, the frequency of occurrence of each grayscale value is counted and normalized to the probability distribution p(f). Based on p(f), the Shannon entropy E of each frame image is calculated. f , measures the information complexity of the image; Load the pre-trained VGG16 model, freeze the convolutional layer weights and select Conv1 to Conv5 as the multi-layer feature output layer. Each layer extracts low-level structural features and high-level semantic features respectively. Input the frame sequence into the VGG16 model and extract features layer by layer. Calculate the gradient amplitude of the frame sequence and count the frequency distribution of the gradient amplitude to generate a multi-channel histogram H M ; Combine the multi-layer convolutional features of VGG16 with the HGMF-MC features to obtain the fused deep feature vector F, fuse the deep feature vector F with the color, texture, and information complexity features to obtain the final deep feature vector, and store the feature vectors of all frames as a feature vector sequence; Use convolutional neural network as the classification model, input the comprehensive feature vector of each frame into the classification model, calculate the classification probability, set the classification threshold D, if the classification probability is greater than or equal to the classification threshold D, then mark the current frame as a semantic scene frame, otherwise mark the current frame as a static scene frame, and mark the frame sequence as a semantic scene frame set and a static scene frame set; Calculate the gradient component and gradient amplitude of each frame in the semantic scene frame set, divide the gradient amplitude of each frame into multiple channels, and perform bucket statistics on the gradient amplitude of each channel to construct the histogram HGMF; Calculate the difference value of the histogram HGMF of each pair of adjacent frames one by one, calculate the mean and standard deviation of the differences between all frames, and set the threshold s according to the sum of the mean and standard deviation; Traverse the semantic scene frame set, select frames whose inter-frame difference values are greater than the threshold s as key frames, and add a timestamp to each key frame; Perform 2D-DWT on each frame in the static scene frame set to extract low-frequency features S; Use the perceptual hashing method to generate a hash value for the low-frequency features of each frame, calculate the difference in hash values of adjacent frames through the Hamming distance, calculate the mean and standard deviation of all hash value differences, and set the threshold g according to the sum of the mean and standard deviation; Traverse the hash value difference sequence, select the frames whose local extremum is greater than g as key frames, and add a timestamp to each key frame; Dynamically calculate the frame rate F based on the actual network bandwidth, number of windows and video content; The key frame set is sorted by timestamp, a frame sequence is selected for key frame extraction according to a dynamically calculated frame extraction frequency, and the extracted key frames are combined into a set.
3. The video transmission method based on streaming processing as claimed in claim 2, characterized in that: The method of encapsulating the video frame into a picture stream and then performing data slicing, allocating an initialization path for the sliced data and calculating the initial flow ratio refers to encapsulating the video frame into a JPG picture format through FFmpeg, integrating the picture frames to form a picture stream, setting a time window to segment the picture stream, dividing the picture frame in the time window into a slice, and each slice is marked as F k , where k represents the current shard number; Add a timestamp to each shard, set a priority identifier for each shard according to data transmission requirements, generate a CRC checksum for the complete data content of each shard, append the checksum to the end of the shard to form a complete shard structure, and obtain a shard data sequence that has been marked. Collect network status information of all current paths and normalize the data, use network performance probes to update path status in real time, and store the monitoring results as a path attribute table; The network status information is calculated by weighted fusion method to obtain the comprehensive score of each path, the paths are sorted in descending order according to the comprehensive score, and the first m paths are selected as the valid path set; For the selected path set, the Pareto distribution model is used to calculate the initial flow allocation ratio of each path in the set; According to the fragment priority identifier, the path allocation ratio is adjusted, and the path allocation table is generated by combining the path network status information and the adjusted allocation ratio.
4. The video transmission method based on streaming processing as claimed in claim 3, characterized in that: The transmitting of the sliced data according to the path optimization result refers to traversing each data slice, selecting the target path group according to the priority, initializing the path scheduler, transmitting the slices in the high priority data set to the target receiving end through the low-latency transmission protocol, transmitting the slices in the second priority data set to the target receiving end through the bandwidth optimization protocol, recording the transmission results of each path to generate a transmission log and storing it in the database.
5. The video transmission method based on streaming processing as claimed in claim 4, characterized in that: After receiving the fragmented data, reorganizing the fragmented data into a video stream and optimizing the image quality means receiving the fragmented data transmitted by each path through the transmission protocol and using a cyclic redundancy check to detect the integrity of each fragment, recording the fragment number that failed the check and requesting retransmission from the sending end, and if the retransmission fails, using adjacent fragments for interpolation compensation; Sorting the slice data according to the slice number and timestamp, merging the slice data of all paths into a complete time sequence according to the number to obtain a complete data frame sequence, initializing the video decoder at the receiving end, decoding the data frame sequence frame by frame, generating a video stream, and correcting the frame errors that occur during the decoding process using error concealment technology; Bilateral filtering is used to remove noise and retain edge details from the decoded image frame sequence, histogram equalization is used to enhance image contrast, and inverse wavelet transform is used to enhance details. The optimized image frame sequence is re-encoded into a complete video stream.
6. The video transmission method based on streaming processing as claimed in claim 5, characterized in that: The video data for storage and transmission refers to compressing the video stream using context-adaptive binary arithmetic coding, storing the compressed video stream in layers, and uploading the data to the cloud for backup after storage.
7. A video transmission system based on streaming processing, based on the video transmission method based on streaming processing according to any one of claims 1 to 6, characterized in that: include, Data collection module, used to collect video streams and perform preprocessing; The data slicing and path initialization module is used to encapsulate the preprocessed video frames into image streams and perform data slicing, allocate initialization paths for the slicing, and calculate the initial traffic ratio; Path optimization module, used to dynamically optimize path selection and adjust path traffic distribution based on network feedback; Data transmission module, used to efficiently transmit the shard data to the receiving end according to the path optimization result; The image quality optimization module is used to receive the fragmented data and reassemble it into a video stream while optimizing the image quality; The storage backup module is used to compress and store the transmitted video stream and upload it to the cloud for backup.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the video transmission method based on streaming processing described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the video transmission method based on streaming processing described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Video material marking method and device, equipment and medium
CN114547375A
Frame extraction method and system for road test vehicle image information
CN116580370A
Live video data optimized recording and storing method
CN117979050A