Method for compressing video data in Internet of Vehicles scene
By adopting a hybrid architecture image compression method in the field of Internet of Vehicles, using dual-branch structure and lightweight fine-tuning technology, the problem of low compression efficiency in the field of Internet of Vehicles is solved, and high-quality and efficient video data compression is achieved.
Patent Information
- Application Number
- CN202510318842.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing video compression technology has shown limitations in the field of Internet of Vehicles, and it is difficult to effectively reduce storage and transmission costs in scenarios with high resolution and large data volumes.
Using a hybrid architecture image compression method, feature extraction, data distribution adaptive quantization and entropy encoding are performed through the dual-branch structure and lightweight fine-tuning technology in the image encoder and decoder, combined with the CNN and Transformer modules.
The compressed video data quality is significantly improved, the average bit rate of pixels is reduced, the storage efficiency of Internet of Vehicles video data is improved, and stability is maintained in different environments.
Smart Images

Figure CN120128733A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image compression, and specifically to a method for compressing video data in the context of the Internet of Vehicles (IoV). Background Art
[0002] With the rapid development of intelligent connected vehicle technology, modern vehicles are equipped with thousands of sensors that generate a large amount of signal data in real time, covering multiple aspects such as vehicle status, driving behavior, and environmental perception. The types of sensor data are complex and diverse, and the data acquisition frequency has gradually increased from the second level to the millisecond level, resulting in a significant increase in the data generation rate. The large-scale generation of this data poses a severe challenge to the data processing and storage capabilities of enterprises. Especially in the Internet of Vehicles (IoV) system, the communication and data storage systems of vehicles need to be able to handle petabytes of massive data. The sources of this data are extensive, including in-vehicle systems, vehicle hardware, user behavior data, and third-party partners, etc. The data volume is huge and has high real-time requirements.
[0003] Taking the in-vehicle six-camera as an example, this device captures 30 high-resolution images per second, with each image having a resolution of 1920x1080 and a data volume of approximately 1.12GB / second. In IoV applications, the real-time transmission and storage of such high-resolution video data require very high bandwidth and storage space. Especially in application scenarios such as autonomous driving, traffic flow analysis, and road monitoring, a large amount of image data needs to be stored for a long time to support intelligent analysis and decision-making. Therefore, how to effectively compress this image data to reduce storage and transmission costs has become an important challenge faced by database systems.
[0004] Currently, traditional video compression standards such as the H.26x series are relatively mature in terms of compression efficiency. It adopts a hybrid video coding framework and significantly reduces the redundancy of video data through techniques such as intra-frame and inter-frame prediction, integer discrete cosine transform (DCT), adaptive entropy coding (CABAC and CAVLC), and deblocking filters. While ensuring relatively high visual quality, the H.26x series has significantly reduced bitrate requirements in video storage and transmission and is widely used in fields such as streaming media, high-definition video, and video conferencing. In recent years, image compression technologies based on neural networks have shown great potential. Learning-based image compression methods represented by Deep ImagePrior and Neural Image Compression model the statistical distribution of image data through deep learning models, enabling the compression algorithm to adapt to the characteristics of the image and retain more detailed information under the premise of a higher compression ratio. These methods can achieve higher-quality image restoration with limited data transmission bandwidth and effectively reduce the storage space requirements.
[0005] However, the above two methods are video compression methods in the general field, and the video compression effect in the Internet of Vehicles (IoV) field is not obvious. Therefore, it is of great significance to develop a compression scheme suitable for high-resolution video data in the IoV to effectively reduce the storage and transmission costs of in-vehicle camera image data.
[0006] Traditional general compression algorithms (such as H.264 / H.265) have been widely used in the field of video compression and have high compression efficiency, but they show some obvious limitations in the Internet of Vehicles (IoV) scenario. These traditional algorithms mainly achieve compression through techniques such as inter-frame prediction, intra-frame prediction, discrete cosine transform (DCT), and entropy coding. However, in scenarios with high resolution and large data volume, especially in the continuous high-frame-rate video data generated by in-vehicle cameras, it is often difficult for traditional compression techniques to further improve the compression ratio while ensuring video quality. This results in excessive transmission bandwidth and storage requirements for data in real-time transmission and storage, making it difficult to meet the requirements of practical applications.
[0007] Neural network methods use deep learning models to model the statistical characteristics of image data and can maintain high image quality at a higher compression ratio. However, most current neural network compression models are trained based on general image datasets (such as ImageNet). These datasets contain a wide range of scenes and contents, but there are significant differences from the characteristics of image data in the IoV scenario. Because the data distribution in the IoV is different from traditional data, and the entropy coding is different, directly using neural network compression models trained on general datasets such as ImageNet often cannot achieve the best compression effect in IoV video data. These models lack adaptability to the specific data distribution in the IoV environment in their design, resulting in their inability to fully exploit the characteristics of in-vehicle camera images and making it difficult to strike a balance between efficient compression and high-quality reconstruction. Summary of the Invention
[0008] Aiming at the deficiencies of the prior art, the purpose of the present invention is to propose a compression method for video data in the IoV scenario, including:
[0009] Step 1: Encode the original image data through an image encoder to obtain a feature representation of the original image data;
[0010] Step 2: Add a uniformly distributed noise to the feature representation of the original image data to obtain a feature representation after adding noise. By introducing the cumulative distribution function (CDF), quantize the feature representation after adding noise to obtain a target quantized feature representation, and then calculate the quantized probability distribution based on the target quantized feature representation;
[0011] Step 3: Perform entropy encoding on each quantized feature representation in the target quantized feature representation to obtain an encoded feature representation, and then calculate the probability distribution after entropy encoding;
[0012] Step 4: Perform entropy decoding on the encoded feature representation to obtain the restored quantized feature representation, and perform inverse quantization on the restored quantized feature representation to obtain the restored continuous feature representation;
[0013] Step 5: Through the CNN deconvolution and attention mechanism in the image decoder, decode the restored continuous feature representation to obtain the reconstructed image data, and then calculate the probability that the original image data is successfully reconstructed given the target quantized feature representation based on the reconstructed image data;
[0014] Step 6: Calculate the optimization objective function according to the quantized probability distribution, the probability that the original image data is successfully reconstructed given the target quantized feature representation, and the probability distribution after entropy encoding;
[0015] Step 7: Update the parameters of the image encoder and the parameters of the image decoder according to the optimization objective function;
[0016] Step 8: Obtain multiple pieces of original image data, and repeat Steps 1 to 7 until the number of repeated executions of Steps 1 to 7 reaches the preset number of rounds, to obtain the trained image encoder and image decoder. Use the trained image encoder to implement image compression through the processes of image encoding, quantization, and entropy encoding.
[0017] Optionally, Step 1 specifically includes:
[0018] The original image data is encoded by multiple basic blocks in the image encoder. Specifically, in the first basic block, the original image data is used as the input of the first basic block. The original image data is processed by the general feature extraction model to obtain the general image feature information. At the same time, the original image data is processed by the vehicle networking feature extraction model to obtain the vehicle networking image feature information. The general image feature information and the vehicle networking image feature information are fused to obtain the combined feature representation of the first basic block. The combined feature representation of the first basic block is used as the input of the next basic module. Similarly, the input combined feature representation is processed by the general feature extraction model and the vehicle networking feature extraction model to obtain the general image feature information of the combined feature representation and the vehicle networking image feature information of the combined feature representation. The general image feature information of the combined feature representation and the vehicle networking image feature information of the combined feature representation are fused to obtain the combined feature representation of this basic block. The combined feature representation of this basic block is used as the input of the next basic block of this basic block. Similarly, the operations of processing and feature fusion are repeated through the general feature extraction model and the vehicle networking feature extraction model until the last basic block. The combined feature output by the last basic block is the feature representation of the original image data.
[0019] Optionally, the general feature extraction model is obtained in the following way:
[0020] The image data in the general dataset ImageNet is input into the hybrid Transformer-CNN architecture model to obtain the predicted general image feature information. Based on the predicted general image feature information and the true general image feature information, the loss value is calculated. Based on the loss value, the model parameters of the hybrid Transformer-CNN architecture model are updated until the loss value is less than the preset threshold, and the general feature extraction model is obtained;
[0021] The vehicle networking feature extraction model is obtained by fine-tuning the general feature extraction model based on the existing public vehicle networking dataset through lightweight fine-tuning technology on the basis of the general feature extraction model.
[0022] Optionally, in step 2, based on the target quantization feature representation, the quantized probability distribution is calculated, which is specifically implemented by the following formula:
[0023]
[0024] where x represents the original image data, represents the target quantization feature representation, represents the quantized probability distribution, represents the parameters of the image encoder, U represents the uniform distribution, y iDenote the \(i\)-th feature representation in the feature representation after adding noise. Denote \(y\). i The quantized feature representation.
[0025] Optionally, the probability distribution after entropy coding in step 3 is calculated specifically through the following formula:
[0026]
[0027] Where, Denote the target quantized feature representation. Denote the probability distribution after entropy coding. Denote the parameters of the entropy encoder. Is The \(i\)-th component of Denote the probability distribution of \(y\) under the condition of given i The probability distribution. Denote the quantized feature representation of the \(i\)-th feature representation in the feature representation after adding noise.
[0028] Optionally, in step 5, based on the reconstructed image data, calculate the probability that the original image data is successfully reconstructed given the target quantized feature representation, which is specifically implemented through the following formula:
[0029]
[0030] Where, \(x\) represents the original image data. Denote the target quantized feature representation. Denote the probability that the original image data is successfully reconstructed given the target quantized feature representation, \(\theta\) g Denote the parameters of the image decoder. Denote the reconstructed image data, \(\lambda\) is a hyperparameter used to measure the importance of the mean square error between the original image data and the reconstructed image data. Denote the mean square error between the original image data and the reconstructed image data.
[0031] Optionally, step 6 is specifically implemented through the following formula:
[0032]
[0033] Where, \(x\) represents the original image data. Denote the target quantized feature representation, \(q\) represents the vehicle networking feature extraction model, \(p\) x Denote the distribution of the original image data. Denote the true probability distribution after quantization. Is used to measure The KL divergence of the difference between \(q\) and Represents the quantized probability distribution, Represents the probability distribution after entropy coding, Represents the probability that the original image data is successfully reconstructed given the target quantized feature representation, and const represents the preset residual.
[0034] The beneficial effects produced by adopting the above technical solution are as follows:
[0035] The present invention designs a dual-branch structure in the image encoder, namely, the branch of the general feature extraction model and the branch of the vehicle networking feature extraction model, which solves the problem that the model trained by the existing learning-based image compression algorithm does not match the characteristics of video data in the vehicle networking field, and significantly improves the quality of the decompressed video data. By using the lightweight fine-tuning technology and fine-tuning the parameters of the general feature extraction model based on the existing publicly available vehicle networking dataset, the present invention can quickly train a model suitable for the characteristics of vehicle networking video data, greatly reducing the number of parameters to be learned and improving the model training efficiency. Especially when dealing with a large amount of vehicle networking video data, the improvement in efficiency is particularly obvious. The hybrid architecture designed by the present invention can capture the local texture features and global context dependencies of the image, significantly reducing the average bit rate of pixels and improving the storage efficiency of vehicle networking video data. The present invention is not affected by the complex and variable content of vehicle networking videos and can still maintain high stability whether in the countryside, city or other environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic flowchart of a method for compressing video data in a vehicle networking scenario according to an embodiment of the present invention;
[0037] Figure 2 It is a schematic flowchart of another method for compressing video data in a vehicle networking scenario according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The following will further describe in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0039] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide an efficient compression method specifically for video data in the Internet of Vehicles (IoV) scenario to meet the strict requirements for video data transmission and storage in the IoV scenario. The present invention utilizes the advantages of neural networks in feature extraction and data distribution adaptation, enabling the compression algorithm to be optimized according to the specific distribution of in-vehicle camera data. Specifically, the present invention introduces lightweight fine-tuning techniques such as LoRA (Low-Rank Adaptation) and Adapter Layer, enabling the neural network model to adaptively compress video data in the IoV scenario without significantly increasing the computational cost. This lightweight fine-tuning method is particularly effective when the computing resources of vehicles are limited, and can optimize the model performance without significantly increasing the hardware burden, thereby compressing in-vehicle camera data.
[0040] Specifically, the present invention provides a compression method for video data in the IoV scenario, combined with Figure 1 and Figure 2 , which may include the following steps:
[0041] Step 1: Encode the original image data through an image encoder to obtain a feature representation of the original image data;
[0042] Encode the original image data through multiple basic blocks in the image encoder. Specifically, in the first basic block, take the original image data as the input of the first basic block, process the original image data through a general feature extraction model. Specifically, process the input features through the convolutional weight W and the ReLU activation function to obtain the general image feature information. At the same time, process the original image data through the IoV feature extraction model to obtain the IoV image feature information, and perform feature fusion Hout on the general image feature information and the IoV image feature information to obtain the combined feature representation of the first basic block, that is, the hidden representation Hin.
[0043] Take the combined feature representation of the first basic block as the input of the next basic module. Similarly, process the input combined feature representation through the general feature extraction model and the IoV feature extraction model to obtain the general image feature information of the combined feature representation and the IoV image feature information of the combined feature representation, perform feature fusion on the general image feature information of the combined feature representation and the IoV image feature information of the combined feature representation to obtain the combined feature representation of this basic block, and take the combined feature representation of this basic block as the input of the next basic block of this basic block. Similarly, repeat the operations of processing and feature fusion through the general feature extraction model and the IoV feature extraction model until the last basic block. The combined feature output by the last basic block is the feature representation of the original image data.
[0044] Among them, the general feature extraction model is obtained through the following method:
[0045] Input the image data in the general dataset ImageNet into the hybrid Transformer-CNN architecture model to obtain the predicted general feature information of the image. Based on the predicted general feature information of the image and the true general feature information of the image, calculate the loss value. Update the model parameters of the hybrid Transformer-CNN architecture model based on the loss value until the loss value is less than the preset threshold to obtain the general feature extraction model;
[0046] Among them, the core of the hybrid Transformer-CNN architecture model involves combining a convolutional neural network (CNN) and a Transformer module. In the CNN part, a multi-level residual convolutional block is adopted to extract the local texture features of the image (such as edges, color distribution, etc.). In the Transformer part, a multi-head self-attention mechanism is used to model the global context dependence of the image. In addition, the output features of the CNN and the Transformer are fused through skip connections to generate the final latent representation.
[0047] The vehicle networking feature extraction model is obtained by fine-tuning the general feature extraction model based on the existing public vehicle networking dataset, such as KITTI, through lightweight fine-tuning techniques. Specifically, lightweight fine-tuning techniques such as LoRA (Low-Rank Adaptation) and Adapter Layer can be used to achieve this. These techniques fine-tune the convolutional weights through additional low-rank matrices WA and WB to adapt to the specific data distribution at the lowest computational cost. This branch outputs the dedicated features in the field of vehicle networking.
[0048] In image compression, the main purpose of quantization is to convert continuous numerical values into discrete numerical values, which is convenient for entropy coding to process. Since continuous numerical values cannot be directly used for finite symbol coding, they must be discretized first. In the quantization process, generally, rounding or floor operation is directly performed. However, in this rounding step, because the continuous numerical values become discrete numerical values, it makes it impossible to optimize the quantization process through standard backpropagation in neural network training. In this solution, by using the method of adding uniform noise, a uniform distribution noise is added to the continuous values before quantization to avoid the gradient blockage caused by non-continuous values. The specific quantization process is shown in Step 2.
[0049] Step 2: Add a uniform distribution noise to the feature representation of the original image data to obtain the feature representation after adding noise. Quantize the feature representation after adding noise by introducing the cumulative distribution function CDF to obtain the target quantized feature representation;
[0050] Furthermore, based on the target quantization feature representation, the quantized probability distribution is calculated, which is specifically implemented through the following formula:
[0051]
[0052] Among them, x represents the original image data, represents the target quantization feature representation, represents the quantized probability distribution, represents the parameters of the image encoder, U represents the uniform distribution, and y i represents the i-th feature representation in the feature representation after adding noise, represents y i the quantized feature representation.
[0053] Step 3: Perform entropy coding on each quantized feature representation in the target quantization feature representation to obtain the encoded feature representation;
[0054] Furthermore, calculate the probability distribution after entropy coding, which is specifically implemented through the following formula:
[0055]
[0056] Among them, represents the target quantization feature representation, represents the probability distribution after entropy coding, represents the parameters of the entropy encoder, is the i-th component of, represents the probability distribution of y under the condition of given i and, represents the quantized feature representation of the i-th feature representation in the feature representation after adding noise.
[0057] The purpose of entropy coding is to minimize the average coding length of each symbol. The prerequisite for the optimal solution is to obtain the true probability distribution of the symbols. If the wrong actual distribution is used, the average coding length will increase. After entropy coding, the quantized features are encoded into a compact binary representation (such as "0110...1110"), further compressing the data volume. This part of the processing ensures that the image data is transmitted with a smaller bandwidth occupancy during the transmission process.
[0058] Step 4: Perform entropy decoding on the encoded feature representation to obtain the restored quantized feature representation. These features are still discrete quantization values. Perform inverse quantization on the restored quantized feature representation to obtain the restored continuous feature representation, so as to restore the distribution of the original features as much as possible. The final output after inverse quantization is the restored continuous feature representation, which is used as the input to the image decoder for further image reconstruction.
[0059] Step 5: Decode the restored continuous feature representation through CNN deconvolution and attention mechanism in the image decoder to obtain the reconstructed image data;
[0060] Combined with Figure 2 , the image decoding image decoder consists of multiple basic blocks (symmetric to the encoder structure), restores the feature representation layer by layer, and reconstructs it back into a form close to the original input image. The output reconstructed image data should retain as much key information in the original image as possible to meet the requirements for image quality in the vehicle networking scenario.
[0061] Furthermore, based on the reconstructed image data, calculate the probability that the original image data is successfully reconstructed given the target quantization feature representation, which is specifically achieved through the following formula:
[0062]
[0063] where \(x\) represents the original image data, represents the target quantization feature representation, represents the probability that the original image data is successfully reconstructed given the target quantization feature representation, \(\theta\) g represents the parameters of the image decoder, represents the reconstructed image data, \(\lambda\) is a hyperparameter used to measure the importance of the mean square error between the original image data and the reconstructed image data, represents the mean square error between the original image data and the reconstructed image data.
[0064] Step 6: Calculate the optimization objective function according to the quantized probability distribution, the probability that the original image data is successfully reconstructed given the target quantization feature representation, and the entropy-coded probability distribution, which is specifically achieved through the following formula:
[0065]
[0066] where \(x\) represents the original image data, represents the target quantization feature representation, \(q\) represents the vehicle networking feature extraction model, \(p\) x represents the distribution of the original image data, represents the quantized true probability distribution, is the KL divergence used to measure the difference between \(q\) and , represents the quantized probability distribution, represents the entropy-coded probability distribution, represents the probability that the original image data is successfully reconstructed given the target quantization feature representation, and const represents the preset residual.
[0067] The core of the objective optimization function is to minimize a KL divergence, which measures and the true posterior The distance between them is minimized to approximate the true posterior distribution as closely as possible because the true posterior distribution is difficult to compute as it involves integrating or summing over the entire data space. So we introduce a substitute, a computable approximate distribution
[0068] The optimization objective consists of two parts: the distortion term (weighted distortion) is used to measure the loss between the reconstructed image and the original image x, and the closer it is to 0, the better. The bitrate term (rate) is used to measure the number of bits required for encoding The smaller the better, which describes the length of the compressed data.
[0069] During the training process, the posterior distribution is related to the actual task and is used to describe the distribution of latent variables. Design an approximate distribution with simplified calculations to approximate the true posterior distribution, thereby optimizing the computational efficiency of the model. By balancing the coding error and the bitrate, an optimizable objective function is constructed, and posterior distribution estimation is used to improve the training effect of the model.
[0070] The optimization objective function is one of the core inventive points in the present invention because it end-to-end guides the joint training of various modules such as the encoder, quantization, entropy coding, and decoder. By balancing the coding error (distortion term) and the coding cost (bitrate term), the overall optimization of the video data compression effect in the vehicle networking scenario is achieved. In the following image coding steps, the original image x is processed by the encoder to generate a feature representation y; in the quantization inference and entropy coding steps, y is quantized and probability modeled to obtain a symbol sequence; throughout the process, the optimization objective function is used as a loss function to guide the joint optimization of the parameters of each step.
[0071] Step 7: Update the parameters of the image encoder and the parameters of the image decoder according to the optimization objective function;
[0072] Among them, updating the parameters of the image encoder means updating the parameters of the general feature extraction model and the parameters of the vehicle networking feature extraction model.
[0073] Step 8: Obtain multiple original image data, repeat Steps 1 to 7 until the number of repetitions of Steps 1 to 7 reaches a preset number of rounds, to obtain a trained image encoder and image decoder. Use the trained image encoder to achieve image compression through the processes of image coding, quantization, and entropy coding.
[0074] It can also be understood that the present invention obtains a plurality of original image data and the true probability distribution after quantization corresponding to the original image data. Based on this, taking the original image data as an input sample, from step 1 to step 7, the parameters of the image encoder and the parameters of the image decoder are updated once. Then, taking a new original image data as an input sample, execute step 1 to step 7, and update the parameters of the image encoder and the parameters of the image decoder again. Keep repeating until the number of executions of step 1 to step 7 reaches the preset number of rounds, that is, the number of times the parameters of the image encoder and the parameters of the image decoder are updated reaches the preset number of rounds, and the updated image encoder and image decoder are obtained.
[0075] Thus, perform image encoding on the image data to be compressed to obtain the feature representation of the image data to be compressed. Add a noise with a uniform distribution to the feature representation of the image data to be compressed, perform quantization and entropy encoding on the feature representation of the image data to be compressed after adding the noise, and obtain the encoded feature representation of the image to be compressed, thereby realizing the compression process of the image.
[0076] The present invention cleverly combines the existing learning-based image / video compression algorithms and the model fine-tuning technology LoRA, and proposes a comprehensive solution aiming at the deficiencies of the existing methods in video compression in the vehicle networking field, aiming to improve the video compression efficiency and compression quality in the vehicle networking field. The following are the core highlights of the present invention:
[0077] 1. Training paradigm of the hybrid architecture: The entire training process adopts a Transformer-CNN hybrid architecture, combining the local feature extraction ability of CNN and the global modeling ability of Transformer to learn the optimal latent representation. Compared with the pure CNN or pure Transformer structure, the hybrid architecture can better balance the computational efficiency and compression quality, and adapt to the diverse video data in the vehicle networking scenario.
[0078] 2. Dual-branch structure to enhance adaptability: Introduce a dual-branch structure to learn general image features and vehicle networking-specific image features respectively. This enables the model to not only retain general image information but also optimize for the data distribution in the vehicle networking scenario, thereby improving the compression performance.
[0079] 3. Efficient fine-tuning mechanism: Use LoRA for lightweight fine-tuning to reduce the computational cost and optimization difficulty, enabling the model to efficiently adapt to the data characteristics in the vehicle networking field. LoRA can significantly improve the performance on specific vehicle networking datasets by only updating a small number of parameters without affecting the original generalization ability of the model.
[0080] By comprehensively utilizing existing learning-based image / video compression algorithms and LoRA fine-tuning technology, the present invention achieves efficient compression of video data in the field of vehicle networking. Compared with existing traditional video compression algorithms and neural network image compression strategies, it demonstrates significant advantages and excellent effects. The following are the main effects and advantages of the present invention:
[0081] 1. Improved compression quality of video data: The present invention designs a dual-branch structure, which solves the problem of the mismatch between the model trained by existing learning-based image compression algorithms and the characteristics of video data in the field of vehicle networking, and significantly improves the quality of decompressed video data.
[0082] 2. Reduced model training time: By using LoRA fine-tuning technology, the present invention can quickly train a model suitable for the characteristics of vehicle networking video data, greatly reducing the number of parameters to be learned and improving the model training efficiency. Especially when dealing with a large amount of vehicle networking video data, the improvement in efficiency is particularly obvious.
[0083] 3. Reduced average bit rate of pixels: The hybrid architecture designed by the present invention can capture local texture features and global context dependencies of images, significantly reducing the average bit rate of pixels and improving the storage efficiency of vehicle networking video data.
[0084] 4. Adapt to complex and variable video data: The present invention is not affected by the complex and variable content of vehicle networking videos and can still maintain high stability whether in rural areas, cities or other environments.
[0085] In the experiment, we compared the compression effects of the present invention with existing learning-based image compression algorithms / video compression algorithms on multiple vehicle networking data sets. The results show that under the same hardware environment, the present invention improves the PSNR (Peak Signal-to-Noise Ratio) for evaluating the quality of decompressed images and the SSIM (Structural Similarity Index) for measuring the similarity between the decoded image and the original image by about 15%, reduces the average number of bits per pixel (Bit-rate) by 17%, and also maintains high reliability and stability. The present invention demonstrates significant technical advantages and application value in the video compression of the vehicle networking field.
[0086] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for compressing video data in a vehicle networking scenario, characterized in that: include: Step 1: Encode the original image data through the image encoder to obtain the feature representation of the original image data; Step 2: Add a uniformly distributed noise to the feature representation of the original image data to obtain the feature representation after adding the noise. By introducing the cumulative distribution function CDF, the feature representation after adding the noise is quantized to obtain the target quantized feature representation. Then, based on the target quantized feature representation, the probability distribution after quantization is calculated. Step 3: Perform entropy coding on each quantized feature representation in the target quantized feature representation to obtain a coded feature representation, and then calculate the probability distribution after entropy coding; Step 4: entropy decoding is performed on the encoded feature representation to obtain a restored quantized feature representation, and dequantization is performed on the restored quantized feature representation to obtain a restored continuous feature representation; Step 5: Decode the recovered continuous feature representation through the CNN deconvolution and attention mechanism in the image decoder to obtain the reconstructed image data, and then calculate the probability of the original image data being successfully reconstructed given the target quantitative feature representation based on the reconstructed image data; Step 6: Calculate the optimization objective function based on the probability distribution after quantization, the probability of the original image data being successfully reconstructed given the target quantization feature representation, and the probability distribution after entropy coding; Step 7: Update the parameters of the image encoder and the parameters of the image decoder according to the optimization objective function; Step 8: Obtain multiple original image data, repeat steps 1 to 7 until the number of repetitions of steps 1 to 7 reaches a preset number of rounds, and obtain the trained image encoder and image decoder. Use the trained image encoder to compress the image through the processes of image encoding, quantization, and entropy encoding.
2. The method for compressing video data in a vehicle networking scenario according to claim 1, characterized in that: Step 1 specifically includes: The original image data is encoded by multiple basic blocks in the image encoder. Specifically, in the first basic block, the original image data is used as the input of the first basic block, and the original image data is processed by the general feature extraction model to obtain the general feature information of the image. At the same time, the original image data is processed by the Internet of Vehicles feature extraction model to obtain the Internet of Vehicles image feature information. The general feature information of the image and the Internet of Vehicles image feature information are feature fused to obtain the combined feature representation of the first basic block, and the combined feature representation of the first basic block is used as the input of the next basic module. Similarly, the input combined feature representation is processed by the general feature extraction model and the Internet of Vehicles feature extraction model to obtain the general feature information of the image represented by the combined feature and the Internet of Vehicles image feature information represented by the combined feature. The general feature information of the image represented by the combined feature and the Internet of Vehicles image feature information represented by the combined feature are feature fused to obtain the combined feature representation of the basic block, and the combined feature representation of the basic block is used as the input of the next basic block of the basic block. Similarly, the processing and feature fusion operations by the general feature extraction model and the Internet of Vehicles feature extraction model are repeated until the last basic block, and the combined feature output by the last basic block is the feature representation of the original image data.
3. The method for compressing video data in a vehicle networking scenario according to claim 2, characterized in that: The general feature extraction model is obtained in the following way: Input the image data in the general dataset ImageNet into the hybrid Transformer-CNN architecture model to obtain the general feature information of the predicted image, calculate the loss value based on the general feature information of the predicted image and the general feature information of the real image, and update the model parameters of the hybrid Transformer-CNN architecture model based on the loss value until the loss value is less than the preset threshold, thereby obtaining a general feature extraction model; The Internet of Vehicles feature extraction model is obtained by fine-tuning the general feature extraction model based on the existing Internet of Vehicles public dataset through lightweight fine-tuning technology.
4. The method for compressing video data in a vehicle networking scenario according to claim 1, characterized in that: In step 2, based on the target quantitative feature representation, the quantized probability distribution is calculated, which is specifically implemented by the following formula: Among them, x represents the original image data, represents the target quantitative feature representation, represents the quantized probability distribution, represents the parameters of the image encoder, U represents uniform distribution, y i represents the i-th feature representation in the feature representation after adding noise, Represents y i Quantized feature representation.
5. The method for compressing video data in a vehicle networking scenario according to claim 1, characterized in that: The probability distribution after entropy coding is calculated in step 3, which is specifically implemented by the following formula: in, represents the target quantitative feature representation, represents the probability distribution after entropy coding, represents the parameters of the entropy encoder, for The i-th component of Indicates that in a given Under the condition of y i The probability distribution of It represents the quantized feature representation of the i-th feature representation in the feature representation after adding noise.
6. The method for compressing video data in a vehicle networking scenario according to claim 1, characterized in that: In step 5, based on the reconstructed image data, the probability of the original image data being successfully reconstructed given the target quantitative feature representation is calculated, which is specifically implemented by the following formula: Among them, x represents the original image data, represents the target quantitative feature representation, represents the probability that the original image data is successfully reconstructed given the target quantitative feature representation, θ g Represents the parameters of the image decoder, represents the reconstructed image data, λ is a hyperparameter used to measure the importance of the mean square error between the original image data and the reconstructed image data, Represents the mean square error between the original image data and the reconstructed image data.
7. The method for compressing video data in a vehicle networking scenario according to claim 1, characterized in that: Step 6 is specifically implemented by the following formula: Among them, x represents the original image data, represents the target quantitative feature representation, q represents the Internet of Vehicles feature extraction model, and p x represents the distribution of the original image data, represents the true probability distribution after quantization, To measure q and The KL divergence of the difference between represents the quantized probability distribution, represents the probability distribution after entropy coding, It represents the probability that the original image data is successfully reconstructed given the target quantitative feature representation, and const represents the preset residual.
Citation Information
Patent Citations
Character-object interaction action detection method based on scene graph information refining and compression
CN117789294A
Electrical impedance image reconstruction method
CN117911713A
Complex traffic scene analysis method and system based on high-order video technology
CN118247741A
Fast VVC intra-frame coding method for machine video coding
CN119211557A
Image semantic segmentation method based on TransDeep model
CN119360028A