Novel digital hybrid coding method based on large model

Through a new digital hybrid encoding method based on large models, combined with multi-layer feature learning, neural network encoding, adaptive entropy coding and A3C optimization strategy, the problems of insufficient compression efficiency and poor adaptability in processing complex data are solved, and efficient, flexible and adaptable data compression effects are achieved.

CN120124682AActive Publication Date: 2025-06-10广州企通云网络科技有限公司

Patent Information

Application Number
CN202510208727.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-10
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Traditional data encoding technology is insufficient compression efficiency when processing complex and variable data, poor adaptability, difficult to optimize, and limited performance when facing high-dimensional and diverse data.

Method used

A new digital hybrid encoding method based on large models is adopted, through multi-layer feature learning, neural network encoding, adaptive entropy coding and A3C optimization strategies, the encoding parameters are dynamically adjusted to adapt to different channel environments and data characteristics, and the lightweight encoding model is extracted using diffusion model and knowledge distillation technology.

Benefits of technology

It significantly improves data compression efficiency and encoding quality, achieves high efficiency, flexibility and adaptability in complex application scenarios, and brings significant performance improvements to data transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124682A_ABST
    Figure CN120124682A_ABST
Patent Text Reader

Abstract

The invention discloses a novel digital hybrid coding method based on a large model. The method comprises the following steps: S1, preprocessing to-be-coded data; s2, performing feature coding on the preprocessed data by adopting a diffusion model driven feature learning method; s3, performing neural network coding on the feature coding data, and dynamically adjusting the coding bit number based on a variable bit rate quantization method; s4, performing adaptive entropy coding on the neural network coding data, and performing probability density transformation based on a normalized flow entropy coding method; s5, discrete transform coding is carried out on entropy coding data, and hierarchical quantization is carried out on different frequency components by adopting a hierarchical sub-block quantization method; and S6, based on the code length distribution and entropy coding parameters of the A3C optimization transformation coding data, extracting a lightweight coding model from the diffusion model by using a knowledge distillation method, and generating final optimization coding data. Through the hybrid coding and dynamic optimization strategy, the data compression efficiency and the coding quality are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital encoding technologies, and particularly to a novel digital hybrid encoding method based on large models. Background Art

[0002] With the advent of the digital age, data transmission and storage are facing increasing demands. Especially in the processing of high-quality audio and video, images, text, and other multimedia data, how to efficiently compress and encode has become a key issue. Traditional data compression techniques such as Huffman coding, arithmetic coding, discrete cosine transform (DCT), discrete wavelet transform (DWT), etc. are widely used in image, audio, and video coding. These methods reduce the storage and transmission costs of data by effectively reducing redundant information. However, with the increasing diversity of data types and the complexity of the transmission environment, traditional coding techniques often face the following problems when dealing with large-scale data: insufficient compression efficiency, poor adaptability to different channel environments, difficulty in optimizing the complexity during the encoding process, and limited performance of the algorithm when facing data with higher dimensions and more complex features.

[0003] Firstly, traditional techniques based on transform coding (such as DCT, DWT) can achieve data compression to a certain extent, but their effectiveness depends on pre-set feature extraction rules and fixed transform models. These transform methods often perform well under fixed data types and feature structures, but are inadequate for complex and variable data (such as multi-modal data or data from different sensors). In particular, for some high-dimensional data with complex spatio-temporal dependencies, traditional feature extraction methods and compression strategies are difficult to obtain ideal compression effects. In addition, traditional coding techniques usually ignore the context information of the data and have poor adaptability to data in different channel environments. Therefore, in practical applications, complex manual adjustments often need to be made according to specific channel and data characteristics.

[0004] Secondly, coding methods based on neural networks have shown excellent performance in various data compression tasks in recent years. Deep learning models can handle complex data distributions and feature structures by automatically learning feature representations. However, when applied to data coding, these models often face problems such as a large scale of training data, high consumption of computing resources, and poor adaptability to different application scenarios. Although some deep learning-based coding methods have achieved certain success, most of them still rely on fixed coding architectures and are difficult to be optimized dynamically according to different channel environments, data types, and actual application scenarios.

[0005] Furthermore, although existing adaptive entropy coding methods have played a certain role in improving coding efficiency, there are still certain deficiencies in the process of dealing with complex data. Traditional entropy coding methods (such as Huffman coding, arithmetic coding, etc.) have certain advantages in compression efficiency, but their effects are often limited when facing high-dimensional and various types of data. Adaptive entropy coding can optimize the compression ratio by adjusting the probability distribution, but such methods usually require a large amount of computation and memory overhead, and it is often difficult to adapt to real-time changes under different channel conditions in practical applications. Especially for complex data streams or applications in uncertain environments, it is very difficult for these traditional entropy coding methods to achieve efficient, real-time, and adaptive adjustment.

[0006] In addition, although some existing transform coding methods (such as DCT, DWT, etc.) effectively decompose signals, the flexibility and adaptability of their methods are poor, and it is difficult to perform dynamic optimization according to different channel environments. Traditional transform coding often processes in a fixed domain and cannot adaptively select the most suitable coding strategy according to real-time channel environments and data characteristics. Transform coding methods usually cannot well handle the ever-changing channel conditions in modern communication networks, resulting in limited compression efficiency and transmission effects.

[0007] Therefore, how to provide a new digital hybrid coding method based on large models is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0008] An object of the present invention is to propose a new digital hybrid coding method based on large models. Through multi-layer feature learning, neural network coding, adaptive entropy coding, and A3C optimization strategy, the present invention can efficiently process complex data, significantly improve data compression efficiency and coding quality. By dynamically adjusting coding parameters, it adapts to different channel environments and data characteristics, providing a more optimized compression scheme. At the same time, a lightweight coding model is extracted by using diffusion models and knowledge distillation technology, realizing high efficiency, flexibility, and adaptability in complex application scenarios, and bringing significant performance improvements to data transmission and storage.

[0009] A new digital hybrid coding method based on large models according to an embodiment of the present invention includes the following steps:

[0010] S1. Preprocess the data to be encoded, perform normalization and feature extraction according to the data type, and generate preprocessed data;

[0011] S2. Perform feature encoding on the preprocessed data by using a feature learning method driven by a diffusion model, and extract global relationships and optimal representations by using latent variable representations to generate feature-encoded data;

[0012] S3. Perform neural network encoding on the feature encoding data, use an end-to-end neural network encoding structure for optimal compression representation, and dynamically adjust the encoding bit number based on the variable bit rate quantization method to generate neural network encoding data;

[0013] S4. Perform adaptive entropy encoding on the neural network encoding data, use the context-aware probability modeling method to predict the prior probability distribution of data blocks, and perform probability density transformation based on the normalizing flow entropy encoding method to generate entropy encoding data;

[0014] S5. Perform discrete transform encoding on the entropy encoding data. Based on the dynamic domain conversion method, select a suitable transformation method according to the data characteristics, including discrete cosine transform, discrete wavelet transform or fast Fourier transform, and perform hierarchical sub-block quantization on different frequency components to generate transform encoding data;

[0015] S6. Optimize the code length allocation and entropy encoding parameters of the transform encoding data based on A3C, dynamically adjust the entropy encoding parameters according to different channel environments, and at the same time use the knowledge distillation method to extract a lightweight encoding model from the diffusion model to generate the final optimized encoding data.

[0016] Optionally, the S2 specifically includes:

[0017] S21. Perform feature decomposition on the preprocessed data, use the multi-scale feature extraction method to perform multi-level partitioning on the data, construct feature representations of different scales, and generate a set of data blocks {X i}, where X i represents the i-th data block;

[0018] S22. Perform diffusion modeling on the set of data blocks {X i}, use the forward diffusion process to construct the probability evolution path of the data blocks, and gradually evolve the data into a standard Gaussian distribution:

[0019]

[0020] Among them, represents the state of the i-th data block at the t-th time step in the diffusion process, represents the state of the i-th data block at the (t - 1)-th time step in the diffusion process, α t represents the attenuation coefficient at the t-th time step, ∈ t represents random Gaussian noise, and N(0, I) represents the standard normal distribution with zero mean and unit variance;

[0021] S23. Perform the reverse diffusion process on the data after diffusion modeling, use the denoising network to estimate the noise term, and perform reverse derivation at multiple time steps to obtain the denoised data representation:

[0022]

[0023] Among them, represents the noise term estimated by the denoising network;

[0024] S24. Perform latent variable mapping based on the denoised data representation, and use a variational autoencoder to learn the latent variable representation of the data:

[0025] Z i = μ i + σ i · ∈, ∈ ~ N(0, I);

[0026] Among them, Z i represents the latent variable representation of the i-th data block, μ i represents the mean corresponding to the i-th data block, σ i represents the standard deviation corresponding to the i-th data block, and ∈ represents a random variable following the standard normal distribution N(0, I);

[0027] S25. Perform global relationship modeling on the latent variable representation, use a Transformer structure to calculate the global relationship between data blocks, and construct a self-attention weight matrix:

[0028]

[0029] Among them, A ij represents the attention weight between the i-th data block and the j-th data block, softmax represents the normalization function, Q i represents the query vector of the i-th data block, K j represents the key vector of the j-th data block, and d represents the feature dimension;

[0030] S26. Extract the optimal representation based on the global relationship modeling result, use a feed-forward neural network to fuse self-attention features, and generate feature-encoded data.

[0031] Optionally, the specific steps of S3 include:

[0032] S31. Perform feature transformation on the feature-encoded data, use a high-dimensional non-linear projection method to map it to a multi-scale feature space, and construct an input data set {F i}, where F i represents the i-th feature-encoded data block, and use an adaptive feature enhancement technology to generate a high-dimensional feature representation {H i}, where H i represents the high-dimensional feature representation of the i-th feature-encoded data block;

[0033] S32. Use an end-to-end neural network encoding structure to perform deep compression representation on the high-dimensional feature representation {H i}, dynamically adjust the feature weights using a variable-channel depth convolution network, and calculate the compression representation {C i} through a high-order non-linear mapping function:

[0034] C i = σ(W 3 ·Swish(W 2 ·ReLU(W 1 H i + b 1 )+ b 2 )+ b 3 );

[0035] Among them, C i represents the compression representation of the i-th feature encoding data block, σ represents the activation function, W 1 , W 2 and W 3 represent the mapping weight matrices, b 1 , b 2 and b 3 represent the bias terms, ReLU represents the rectified linear unit activation function, and Swish represents the differentiable activation function;

[0036] S33. Perform variable bitrate quantization processing on the compression representation {C i}, adopt a non-uniform logarithmic scale quantization method, and generate the quantization encoded data {Q i}, where Q i represents the i-th feature quantization encoded data:

[0037]

[0038] Among them, Q(C i ) represents the quantization value of the i-th feature encoding data block, β represents the non-uniform quantization scale parameter, max(|C i |) represents the maximum value of the i-th feature encoding data block; sign(·) represents the sign function: when C i is positive, sign(C i ) = 1; when C i is negative, sign(C i ) = -1; when C i is zero, sign(C i ) = 0;

[0039] S34. Calculate the information redundancy between different data blocks based on the cross-modal contrast learning method, optimize the data compression expression using a hierarchical contrast loss function, and generate the neural network encoded data:

[0040]

[0041] Among them, L contrastive represents the cross-modal feature contrast loss, N represents the total number of feature encoding data blocks, exp represents the exponential function, and S(Q i , Q j ) represents the similarity measure between the i-th feature quantization encoded data Q i and the j-th feature quantization encoded data Q j , and S(Q i , Q k ) represents the similarity measure between the i-th feature quantization encoded data Q i and the k-th feature quantization encoded data Q k . τ represents the temperature parameter; I(y i = y j ) represents the indicator function. When y i = y j , I(y i = y j ) = 1. When y i ≠ y j , I(y i = y j ) = 0.

[0042] Optionally, the S4 specifically includes:

[0043] S41. Divide the neural network encoded data into several neural network encoded data blocks and assign indexes, and record the context information for each neural network encoded data block, including the encoding results of adjacent neural network encoded data blocks and the global statistical features;

[0044] S42. Based on the context-aware probability modeling method, predict the prior probability of each neural network encoded data block, and use the self-attention mechanism to comprehensively consider the local and global dependencies to generate the probability distribution model of the neural network encoded data block;

[0045] S43. Perform normalized flow entropy coding conversion on the probability distribution model, select a reversible and differentiable flow model to map the data distribution, so that the data tends to be uniformly distributed;

[0046] S44. Based on the normalized flow conversion result, combined with the context prediction information of the neural network encoded data block, adaptively allocate the code length by means of bit rate regulation or entropy threshold adjustment, so that the encoding length used by the high-frequency neural network encoded data block is greater than the encoding length used by the low-frequency neural network encoded data block;

[0047] S45. Perform entropy encoding on the neural network encoded data block after adaptive code length allocation, generate a compressed bitstream using arithmetic coding, and perform correlation verification;

[0048] S46. Output the compressed bitstream with correct correlation verification as entropy encoded data.

[0049] Optionally, the S5 specifically includes:

[0050] S51. Perform feature analysis on the entropy encoded data, calculate the time-domain variation characteristics and spectral energy distribution of the entropy encoded data block;

[0051] S52. Based on the feature analysis result of the entropy encoded data block, use the dynamic domain transformation method, select discrete cosine transform, discrete wavelet transform or fast Fourier transform according to the characteristics of the entropy encoded data, and perform transformation on the entropy encoded data block to obtain the transformed entropy encoded data block;

[0052] S53. Perform frequency decomposition on the transformed entropy encoded data block, divide several sub-bands based on the spectral energy density, calculate the signal energy of each sub-band, and determine the optimal decomposition scale based on the local gradient change analysis;

[0053] S54. Perform non-uniform quantization on the transform coefficients of different sub-bands through transform domain gradient constraint:

[0054]

[0055] where, E i,j represents the quantization value of the j-th frequency component in the i-th entropy encoded data block, C i,j represents the transform coefficient, Δ i,j represents the adaptive quantization step size, round(·) represents the rounding operation, A i,j represents the attention weight of the j-th frequency component in the i-th entropy encoded data block, γ represents the gradient regulation coefficient, represents the local gradient of the transform coefficient along the frequency axis;

[0056] S55. Perform transform domain recombination on the non-uniformly quantized entropy encoded data block to generate transform encoded data.

[0057] Optionally, the S6 specifically includes:

[0058] S61. Optimize the code length allocation of the transform encoded data using the A3C algorithm, construct the state space and action space, and calculate the optimal code length allocation strategy through the policy network and value function:

[0059]

[0060] Among them, θ represents the set of parameters of the policy network, ξ represents the learning rate, and N 1 represents the total number of transformed coding data blocks, and s i represents the state of the i-th transformed coding data block, and π θ (s i ) represents the probability distribution of the policy network π θ selecting an action in the state s i . represents the parameter gradient, A represents the parameter gradient, and A i represents the advantage estimate value, and V θ (s i ) represents the expected total reward obtained by adopting the current policy in the state s i . ← represents the update operation;

[0061] S62. According to the code length allocation optimized by the A3C algorithm, dynamically adjust the entropy coding parameters of the transformed coding data blocks, adaptively configure the entropy coding algorithm based on different channel environments, and the optimization goal is to maximize the coding efficiency;

[0062] S63. Use the knowledge distillation method to extract a lightweight coding model from the diffusion model, and transfer the knowledge of the complex model to the student model through the teacher-student model training strategy:

[0063]

[0064] Among them, L KD represents the knowledge distillation loss function, G i represents the i-th transformed coding data block, P teacher (G i ) represents the probability distribution of the teacher model, P student (G i ) represents the probability distribution of the student model, D KL (·||·) represents the Kullback-Leibler divergence, which is used to measure the difference between the teacher model and the student model;

[0065] S64. Based on the lightweight coding model, combined with the changes in the channel environment, dynamically adjust the code length and entropy coding parameters of the transformed coding data blocks to generate the final optimized coding data.

[0066] The beneficial effects of the present invention are:

[0067] First, by combining diffusion models, neural network encoding, entropy encoding, and transform encoding, the present invention can achieve more efficient compression during large-scale data transmission. The diffusion model can effectively represent and reduce the dimensionality of data by learning the latent structure of the data, thereby reducing redundant information and increasing the compression ratio. At the same time, neural network encoding can automatically extract the features of complex data, avoiding the limitations of the manually set feature extraction rules in traditional encoding methods, making the encoding process more intelligent and adaptive. Compared with traditional transform- and entropy-coding-based techniques, the hybrid encoding method of the present invention can better adapt to different types of input data, improving the adaptability and accuracy of data compression.

[0068] Secondly, the present invention introduces a reinforcement learning optimization strategy during the encoding process, which can dynamically adjust the encoding strategy under changing channel environments and data characteristics. This innovative design enables the present invention to perform adaptive optimization according to the actual application scenario, greatly improving the flexibility and real-time performance of encoding. Existing traditional encoding methods often rely on fixed rules and models and lack the ability to adapt to environmental changes. However, through reinforcement learning optimization, the present invention can not only evaluate the effect of data transmission in real time but also adjust the compression algorithm and encoding process under different channel conditions to ensure the best encoding effect in various network environments. This adaptive optimization mechanism not only improves the stability of data transmission but also effectively addresses issues such as network fluctuations and transmission delays, ensuring efficient compression and transmission of data.

[0069] Furthermore, the hybrid encoding method of the present invention has high computational efficiency. Compared with traditional deep learning encoding methods, although deep learning techniques perform well in feature extraction and compression effects, their high computational complexity and high demand for computing resources pose certain bottlenecks in practical applications. However, by introducing techniques such as knowledge distillation, the present invention not only ensures the efficient training of the model but also reduces the consumption of computing resources. By adopting distributed computing and hierarchical optimization strategies, the invention enables the model to achieve real-time encoding and decoding at a relatively low computational cost while maintaining an efficient compression effect, especially suitable for scenarios of large-scale data transmission and real-time communication.

[0070] In addition, the digital hybrid coding method of the present invention has strong multi-modal adaptation ability. Traditional coding techniques are usually optimized for a single data type. However, with the rise of the Internet of Things, intelligent sensors, and big data applications, the emergence of multi-modal data has brought greater challenges to coding. Based on the multi-modal data compression method of the present invention, it can simultaneously process inputs from different sensors and different data sources. Whether it is audio-visual data, image data, or sensor data from different devices, it can be efficiently compressed and transmitted through a unified coding framework. This multi-modal adaptation ability not only improves the flexibility of the system but also reduces the transmission compatibility issues between different data types, making the coding method applicable to a wider range of application fields, including smart homes, vehicle networking, medical health, and other fields.

[0071] Finally, thanks to the advanced algorithms and optimization strategies adopted by the present invention, the compression efficiency of the data has been significantly improved, and a higher compression ratio has been achieved in practical applications. By combining technologies such as neural networks, transform coding, and entropy coding, the present invention can greatly reduce the requirements for data storage and transmission while ensuring the quality and integrity of the data, reducing the storage cost and the consumption of transmission bandwidth. This advantage is particularly important for fields such as big data transmission and real-time communication, which can effectively improve the data processing speed, reduce the transmission delay, and enhance the performance and response ability of the overall system. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0073] Figure 1 is the overall flowchart of a novel digital hybrid coding method based on a large model proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0075] Refer to Figure 1 , a novel digital hybrid coding method based on a large model, includes the following steps:

[0076] S1. Preprocess the data to be encoded, perform normalization and feature extraction according to the data type, and generate preprocessed data;

[0077] S2. Use a feature learning method driven by a diffusion model to perform feature encoding on the preprocessed data, and use latent variable representation to extract global relationships and optimal representations, generating feature-encoded data;

[0078] S3. Perform neural network encoding on the feature encoding data, use an end-to-end neural network encoding structure for optimal compression representation, and dynamically adjust the encoding bit number based on the variable bit rate quantization method to generate neural network encoded data;

[0079] S4. Perform adaptive entropy encoding on the neural network encoded data, use the context-aware probability modeling method to predict the prior probability distribution of data blocks, and perform probability density transformation based on the normalizing flow entropy encoding method to generate entropy encoded data;

[0080] S5. Perform discrete transform encoding on the entropy encoded data, based on the dynamic domain conversion method, select a suitable transformation method according to data characteristics, including discrete cosine transform, discrete wavelet transform or fast Fourier transform, and perform hierarchical sub-block quantization for different frequency components to generate transform encoded data;

[0081] S6. Optimize the code length allocation and entropy encoding parameters of the transform encoded data based on A3C, dynamically adjust the entropy encoding parameters according to different channel environments, and at the same time use the knowledge distillation method to extract a lightweight encoding model from the diffusion model to generate the final optimized encoded data.

[0082] In this embodiment, the S2 specifically includes:

[0083] S21. Perform feature decomposition on the preprocessed data, use the multi-scale feature extraction method to perform multi-level partitioning on the data, construct feature representations of different scales, and generate a data block set {X i}, where X i represents the i-th data block;

[0084] S22. Perform diffusion modeling on the data block set {X i}, use the forward diffusion process to construct the probability evolution path of data blocks, and gradually evolve the data into a standard Gaussian distribution:

[0085]

[0086] where, represents the state of the i-th data block at the t-th time step in the diffusion process, represents the state of the i-th data block at the (t - 1)-th time step in the diffusion process, α t represents the attenuation coefficient at the t-th time step, ∈ t represents random Gaussian noise, and N(0, I) represents the standard normal distribution with zero mean and unit variance;

[0087] S23. Perform a reverse diffusion process on the data after diffusion modeling, estimate the noise term using a denoising network, and perform reverse derivation at multiple time steps to obtain a denoised data representation:

[0088]

[0089] wherein, represents the noise term estimated by the denoising network;

[0090] S24. Perform latent variable mapping based on the denoised data representation, and use a variational autoencoder to learn the latent variable representation of the data:

[0091] Z i = μ i + σ i · ∈, ∈ ~ N(0, I);

[0092] wherein, Z i represents the latent variable representation of the i-th data block, μ i represents the mean corresponding to the i-th data block, σ i represents the standard deviation corresponding to the i-th data block, and ∈ represents a random variable following the standard normal distribution N(0, I);

[0093] S25. Perform global relationship modeling on the latent variable representation, use a Transformer structure to calculate the global relationship between data blocks, and construct a self-attention weight matrix:

[0094]

[0095] wherein, A ij represents the attention weight between the i-th data block and the j-th data block, softmax represents a normalization function, Q i represents the query vector of the i-th data block, K j represents the key vector of the j-th data block, and d represents the feature dimension;

[0096] S26. Extract the optimal representation based on the global relationship modeling result, and use a feed-forward neural network to fuse the self-attention features to generate feature-encoded data.

[0097] In this embodiment, the specific steps of S3 are as follows:

[0098] S31. Perform feature transformation on the feature-encoded data, map it to a multi-scale feature space using a high-dimensional non-linear projection method, construct an input data set {F i}, where F i represents the i-th feature-encoded data block, and use an adaptive feature enhancement technology to generate a high-dimensional feature representation {H i}, where Hi Represents the high-dimensional feature representation of the i-th feature encoding data block;

[0099] S32. Use an end-to-end neural network encoding structure to perform deep compression representation on the high-dimensional feature representation {H i}, dynamically adjust the feature weights using a variable-channel depth convolution network, and calculate the compression representation {C i}:

[0100] C i = σ(W 3 ·Swish(W 2 ·ReLU(W 1 H i + b 1 ) + b 2 ) + b 3 );

[0101] Among them, C i represents the compression representation of the i-th feature encoding data block, σ represents the activation function, W 1 , W 2 and W 3 represent the mapping weight matrices, b 1 , b 2 and b 3 represent the bias terms, ReLU represents the rectified linear unit activation function, and Swish represents the differentiable activation function;

[0102] S33. Perform variable bitrate quantization processing on the compression representation {C i}, adopt the non-uniform logarithmic scale quantization method, and generate the quantization encoded data {Q i}, where Q i represents the i-th feature quantization encoded data:

[0103]

[0104] Among them, Q(C i ) represents the quantization value of the i-th feature encoding data block, β represents the non-uniform quantization scale parameter, max(|C i |) represents the maximum value of the i-th feature encoding data block; sign(·) represents the sign function: when C i is positive, sign(C i ) = 1; when C i is negative, sign(C i ) = -1; when C i is zero, sign(C i ) = 0;

[0105] S34. Calculate the information redundancy between different data blocks based on the cross-modal contrast learning method, optimize the data compression representation using a hierarchical contrast loss function, and generate neural network encoded data:

[0106]

[0107] Among them, L contrastive represents the cross-modal feature contrast loss, N represents the total number of feature encoded data blocks, exp represents the exponential function, S(Q i , Q j ) represents the similarity measure between the i-th feature quantization encoded data Q i and the j-th feature quantization encoded data Q j , S(Q i , Q k ) represents the similarity measure between the i-th feature quantization encoded data Q i and the k-th feature quantization encoded data Q k , τ represents the temperature parameter; I(y i = y j ) represents the indicator function. When y i = y j , I(y i = y j ) = 1. When y i ≠ y j , I(y i = y j ) = 0.

[0108] In this embodiment, the S4 specifically includes:

[0109] S41. Divide the neural network encoded data into several neural network encoded data blocks and assign indexes, and record the context information for each neural network encoded data block, including the encoding results of adjacent neural network encoded data blocks and global statistical features;

[0110] S42. Based on the context-aware probability modeling method, predict the prior probability of each neural network encoded data block, and use the self-attention mechanism to comprehensively consider local and global dependencies to generate a probability distribution model for the neural network encoded data block;

[0111] S43. Perform a normalized flow entropy coding transformation on the probability distribution model, select a reversible and differentiable flow model to map the data distribution, so that the data tends to be uniformly distributed;

[0112] S44. Based on the normalizing flow conversion result, combine the context prediction information of the neural network encoded data block, and adaptively allocate the code length by means of bit rate regulation or entropy threshold adjustment, so that the encoding length used by the high-frequency neural network encoded data block is greater than the encoding length used by the low-frequency neural network encoded data block;

[0113] S45. Perform entropy encoding processing on the neural network encoded data block after adaptively allocating the code length, generate a compressed bitstream using arithmetic coding, and perform correlation verification;

[0114] S46. Output the compressed bitstream after successful correlation verification as the entropy encoded data.

[0115] In this embodiment, the specific steps of S5 are as follows:

[0116] S51. Analyze the characteristics of the entropy encoded data, and calculate the time-domain variation characteristics and spectral energy distribution of the entropy encoded data block;

[0117] S52. Based on the feature analysis result of the entropy encoded data block, adopt a dynamic domain conversion method, select discrete cosine transform, discrete wavelet transform or fast Fourier transform according to the characteristics of the entropy encoded data, and perform the transform on the entropy encoded data block to obtain the transformed entropy encoded data block;

[0118] S53. Perform frequency decomposition on the transformed entropy encoded data block, divide several sub-bands based on the spectral energy density, calculate the signal energy of each sub-band, and determine the optimal decomposition scale based on the local gradient change analysis;

[0119] S54. Perform non-uniform quantization on the transform coefficients of different sub-bands through transform domain gradient constraint:

[0120]

[0121] where, E i,j represents the quantization value of the j-th frequency component in the i-th entropy encoded data block, C i,j represents the transform coefficient, Δ i,j represents the adaptive quantization step size, round(·) represents the rounding operation, A i,j represents the attention weight of the j-th frequency component in the i-th entropy encoded data block, γ represents the gradient regulation coefficient, represents the local gradient of the transform coefficient along the frequency axis;

[0122] S55. Perform transform domain recombination on the non-uniformly quantized entropy encoded data block to generate transform encoded data.

[0123] In this embodiment, the specific steps of S6 are as follows:

[0124] S61. Optimize the code length allocation of the transformed encoded data using the A3C algorithm, construct the state space and action space, and calculate the optimal code length allocation strategy through the policy network and value function:

[0125]

[0126] Among them, θ represents the parameter set of the policy network, ξ represents the learning rate, N 1 represents the total number of transformed encoded data blocks, s i represents the state of the i-th transformed encoded data block, π θ (s i ) represents the probability distribution of the policy network π θ selecting an action in the state s i , represents the parameter gradient, A i represents the advantage estimate value, V θ (s i ) represents the expected total reward obtained by taking the current policy in the state s i ; ← represents the update operation;

[0127] S62. Dynamically adjust the entropy coding parameters of the transformed encoded data blocks according to the code length allocation optimized by the A3C algorithm, adaptively configure the entropy coding algorithm based on different channel environments, and the optimization goal is to maximize the coding efficiency;

[0128] S63. Use the knowledge distillation method to extract a lightweight coding model from the diffusion model, and transfer the knowledge of the complex model to the student model through the teacher-student model training strategy:

[0129]

[0130] Among them, L KD represents the knowledge distillation loss function, G i represents the i-th transformed encoded data block, P teacher (G i ) represents the probability distribution of the teacher model, P student (G i ) represents the probability distribution of the student model, D KL (·||·) represents the Kullback-Leibler divergence, which is used to measure the difference between the teacher model and the student model;

[0131] S64. Based on the lightweight coding model, combined with the changes in the channel environment, dynamically adjust the code length and entropy coding parameters of the transformed encoded data blocks to generate the final optimized encoded data.

[0132] Example 1:

[0133] To verify the feasibility of the present invention in implementation, the present invention is applied to an intelligent city data transmission system in an Internet of Things (IoT) environment. This system involves a large number of different types of sensors, collecting multi-modal data including temperature, humidity, air pressure, video surveillance, audio monitoring, etc. The data is transmitted to the data center through a wireless network for processing. Due to the complexity and enormity of these data transmissions, traditional coding and compression methods often face problems such as low compression ratio, large data transmission delay, and unstable compression efficiency when the channel environment changes in practical applications. Therefore, an efficient and adaptive coding scheme is needed to optimize data transmission.

[0134] In this embodiment, a new digital hybrid coding method based on large models is adopted to code the sensor data in the intelligent city. This method makes full use of the advantages of diffusion models, neural network coding, entropy coding, and transform coding, and combines the optimization strategy of reinforcement learning to dynamically adjust the coding parameters according to the real-time changing channel environment, so as to improve the data compression efficiency and ensure the integrity and recoverability of the data under unstable network conditions.

[0135] Specifically, the data collected by the sensor nodes, including signals from temperature sensors, humidity sensors, air quality monitors, and video surveillance cameras, will first undergo preprocessing. The preprocessing steps include data normalization, noise removal, and feature extraction. For example, video data first undergoes color space conversion, audio data undergoes signal normalization, and sensor data undergoes normalization and denoising. The preprocessed data is input into the diffusion model for feature learning and coding. Through the diffusion process, the data gradually evolves into a standard Gaussian distribution, which is convenient for subsequent compression and coding.

[0136] Next, neural network coding is used to perform deep compression representation on the preprocessed data, and the coding bit number is dynamically adjusted based on the variable bit rate quantization method. In this way, the compression efficiency is significantly improved. During the coding process, the neural network adaptively adjusts the coding strategy according to different data types and channel conditions to ensure that each data block can be coded at the optimal bit rate during transmission.

[0137] In addition, in the adaptive entropy coding stage, the context-aware probability modeling method is used to predict the prior probability distribution of the data block, and the probability density of the data block is transformed through normalizing flow entropy coding. This process effectively removes the redundant data in the coding process and improves the compression efficiency. At the same time, multiple entropy coding methods, such as Huffman coding and arithmetic coding, are combined in the entropy coding process to ensure the optimal compression effect under different data types.

[0138] The transform coding stage uses a dynamic domain conversion method based on data characteristics. According to the frequency characteristics of different sensor data, an appropriate transform method is selected (such as discrete cosine transform, discrete wavelet transform, or fast Fourier transform). This transform can better separate the low-frequency and high-frequency components of the signal, thus providing favorable conditions for subsequent quantization and compression. In the transformed data, a hierarchical sub-block quantization method is adopted, with a larger quantization for high-frequency components and a precise quantization for low-frequency components. This quantization method effectively balances the compression ratio and data quality.

[0139] At the end of the entire coding process, the code length allocation and entropy coding parameters optimized based on the A3C algorithm will further optimize the transform-coded data. The coding parameters are dynamically adjusted according to different channel environments, and a lightweight coding model is extracted from the diffusion model through knowledge distillation to reduce computational resource consumption and ensure that the coding can operate efficiently on resource-constrained devices.

[0140] Through the above method, the finally generated coded data not only has a high compression ratio but also can be automatically optimized according to changes in network conditions and transmission environments, providing an efficient, stable, and adaptive coding scheme for data transmission in smart cities.

[0141] To verify the effectiveness of the present invention, we conducted experimental data collection and comparison. The data collection location was a data center in a certain smart city, and the performance of the method of the present invention and traditional methods was compared in various scenarios during the experiment.

[0142] The experimental results show that the method of the present invention is significantly superior to traditional coding methods in terms of data compression ratio, transmission delay, and robustness in changing channel environments. The specific experimental data are shown in the following table:

[0143] Table 1 Performance comparison table of different coding methods in the smart city data transmission system

[0144]

[0145] By analyzing the data in Table 1 above, first, in terms of the data compression ratio, the compression ratio of the method of the present invention reaches 58.7%, which is 16.2 percentage points higher than that of traditional compression methods. The compression ratio of traditional compression methods is 42.5%, while the compression ratios of the adaptive entropy coding method and the neural network-based compression method are 50.3% and 53.2% respectively. This result shows that the method of the present invention has obvious advantages in improving the compression ratio and can effectively reduce the bandwidth required for data transmission, saving network resources.

[0146] Secondly, the encoding time and decoding time are another key factors affecting data transmission efficiency. According to the experimental data, the method of the present invention has an encoding time of only 0.9 seconds and a decoding time of 0.8 seconds, both of which are the shortest. In contrast, the encoding time of the traditional compression method is 1.2 seconds and the decoding time is 1.0 seconds. The encoding and decoding times of the adaptive entropy coding method and the neural network-based compression method are also relatively long respectively. Especially when dealing with complex data, the computational overhead of the traditional method and other methods is relatively large. The time optimization of the method of the present invention benefits from its efficient encoding structure and reinforcement learning optimization strategy, thus reducing the time delay of data processing and improving the real-time response ability of the system.

[0147] In terms of transmission delay, the method of the present invention performs excellently with a transmission delay of 150 milliseconds, a 25% reduction compared to the 200 milliseconds of the traditional compression method. This shows that the method of the present invention can reduce the transmission delay of data in the network while ensuring a high compression rate, thereby improving the performance of the overall system. The delays of the adaptive entropy coding method and the neural network-based compression method are 180 milliseconds and 170 milliseconds respectively. Although they are better than the traditional method, they still do not reach the excellent performance of the method of the present invention.

[0148] In terms of the bit error rate, the bit error rate of the method of the present invention is 0.8%, significantly lower than the 5.8% of the traditional compression method, and much lower than the other two methods (3.1% for the adaptive entropy coding method and 2.3% for the neural network-based compression method). This shows that the method of the present invention can better maintain the accuracy and reliability of data during signal transmission. Especially under complex channel conditions, it can effectively reduce data loss and errors and ensure the integrity of data.

[0149] Finally, from the perspective of channel adaptability, the method of the present invention performs the best among all the tested methods and is rated as excellent. It can dynamically adjust the encoding strategy according to different channel environments, ensuring stable performance under various network conditions. Relatively speaking, the channel adaptability of the traditional compression method is poor and it is easily affected by network fluctuations. Although the adaptive entropy coding method and the neural network-based compression method perform better, their adaptive capabilities still cannot be compared with that of the method of the present invention.

[0150] Generally speaking, the method of the present invention performs excellently in all key performance indicators. Especially in terms of compression ratio, encoding efficiency, transmission delay, bit error rate and channel adaptability, it is superior to the prior art. Through this method, the data transmission system of the smart city can achieve efficient, low-latency and robust compression and transmission, suitable for various data transmission requirements in practical applications.

[0151] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.

Claims

1. A new digital hybrid coding method based on a large model, characterized in that: The steps include: S1. Preprocess the data to be encoded, perform normalization and feature extraction according to the data type, and generate preprocessed data; S2, using a diffusion model driven feature learning method to feature encode the preprocessed data, and using latent variable representation to extract global relationships and optimal representations to generate feature encoded data; S3, performing neural network encoding on the feature encoded data, using an end-to-end neural network encoding structure for optimal compression representation, and dynamically adjusting the number of encoding bits based on a variable bit rate quantization method to generate neural network encoded data; S4, performing adaptive entropy coding on the neural network coded data, using a context-aware probability modeling method to predict a priori probability distribution of a data block, performing probability density transformation based on a normalized stream entropy coding method, and generating entropy coded data; S5, performing discrete transform coding on the entropy coded data, selecting a suitable transform method according to data characteristics based on a dynamic domain conversion method, including discrete cosine transform, discrete wavelet transform or fast Fourier transform, and performing hierarchical quantization for different frequency components using a hierarchical sub-block quantization method to generate transform coded data; S6. Optimize the code length allocation and entropy coding parameters of the transform coded data based on A3C, dynamically adjust the entropy coding parameters according to different channel environments, and use the knowledge distillation method to extract a lightweight coding model from the diffusion model to generate the final optimized coded data.

2. According to the novel digital hybrid coding method based on large model in claim 1, it is characterized in that: The S2 specifically includes: S21, feature decomposition is performed on the preprocessed data, a multi-scale feature extraction method is used to perform multi-level division on the data, feature representations of different scales are constructed, and a data block set {X i }, where X i Represents the i-th data block; S22, for the data block set {X i }Diffusion modeling is performed, and the probability evolution path of the data block is constructed using the forward diffusion process, so that the data gradually evolves into a standard Gaussian distribution: in, represents the state of the i-th data block at the t-th time step of the diffusion process, represents the state of the i-th data block at the t-1th time step of the diffusion process, α t represents the attenuation coefficient of the tth time step, ∈ t represents random Gaussian noise, N(0,I) represents the standard normal distribution with zero mean and unit variance; S23, performing a reverse diffusion process on the data after the diffusion modeling, using a denoising network to estimate the noise term, and performing reverse deduction on multiple time steps to obtain a denoised data representation: in, represents the noise term estimated by the denoising network; S24, performing latent variable mapping based on the denoised data representation, and using a variational autoencoder to learn the latent variable representation of the data: Z i =μ i +s i ·∈,∈~N(0,I); Among them, Z i represents the latent variable representation of the i-th data block, μ i represents the mean corresponding to the i-th data block, σ i represents the standard deviation corresponding to the i-th data block, ∈ represents a random variable that obeys the standard normal distribution N(0,I); S25, perform global relationship modeling on the latent variable representation, use the Transformer structure to calculate the global relationship between data blocks, and construct a self-attention weight matrix: Among them, A ij represents the attention weight between the i-th data block and the j-th data block, softmax represents the normalization function, Q i represents the query vector of the i-th data block, K j represents the key vector of the jth data block, and d represents the feature dimension; S26. Extract the optimal representation based on the global relationship modeling result, use a feedforward neural network to fuse the self-attention features, and generate feature encoding data.

3. According to claim 1, a novel digital hybrid coding method based on a large model is characterized in that: The S3 specifically includes: S31, performing feature transformation on the feature coded data, mapping it to a multi-scale feature space using a high-dimensional nonlinear projection method, and constructing an input data set {F i }, where F i represents the i-th feature encoding data block, and uses adaptive feature enhancement technology to generate a high-dimensional feature representation {H i }, where H i Represents the high-dimensional feature representation of the i-th feature encoding data block; S32, using an end-to-end neural network encoding structure to represent the high-dimensional feature {H i } to perform deep compression representation, dynamically adjust feature weights using a variable channel deep convolutional network, and calculate the compressed representation through a high-order nonlinear mapping function {C i }: <h2 style=";text-align:left;direction:ltr">C<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> =σ(W3·Swish(W2·ReLU(W1H<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> +b1)+b2)+b3); Among them, C i represents the compressed representation of the i-th feature encoding data block, σ represents the activation function, W1, W2 and W3 represent the mapping weight matrices, b1, b2 and b3 represent the bias terms, ReLU represents the rectified linear unit activation function, and Swish represents the differentiable activation function; S33, the compression expression {C i }Perform variable bit rate quantization processing and use non-uniform logarithmic scale quantization method to generate quantized coded data {Q i }, where Q i Represents the quantized encoding data of the i-th feature: Among them, Q(C i ) represents the quantization value of the i-th feature coding data block, β represents the non-uniform quantization scale parameter, max(|C i |) represents the maximum value of the i-th feature encoding data block; sign(·) represents the sign function: when C i When it is a positive number, sign(C i )=1; when C i When it is a negative number, sign(C i )=-1; when C i When it is zero, sign(C i )=0; S34. Based on the cross-modal contrastive learning method, the information redundancy between different data blocks is calculated, and the hierarchical contrast loss function is used to optimize the data compression expression to generate neural network encoding data: Among them, L contrastive represents the cross-modal feature contrast loss, N represents the total number of feature encoding data blocks, exp represents the exponential function, S(Q i ,Q j ) represents the i-th feature quantization coded data Q i and the jth feature quantized coded data Q j The similarity measure between i ,Q k ) represents the i-th feature quantization coded data Q i and the kth feature quantized coded data Q k The similarity measure between them, τ represents the temperature parameter; I(y i =y j ) represents the indicator function, when y i =y j When I(y i =y j )=1, when y i ≠y j When I(y i =y j )=0.

4. A novel digital hybrid coding method based on a large model according to claim 1, characterized in that: The S4 specifically includes: S41, dividing the neural network encoded data into a number of neural network encoded data blocks and assigning indexes, and recording context information for each neural network encoded data block, including encoding results and global statistical features of adjacent neural network encoded data blocks; S42, based on the context-aware probability modeling method, predict the prior probability of each neural network encoding data block, use the self-attention mechanism to comprehensively consider the local and global dependencies, and generate a probability distribution model of the neural network encoding data block; S43, performing normalized stream entropy coding conversion on the probability distribution model, selecting a reversible and differentiable stream model to map the data distribution, so that the data tends to be evenly distributed; S44, based on the normalized stream conversion result and in combination with the context prediction information of the neural network coded data block, adaptively allocating the code length by bit rate control or entropy threshold adjustment, so that the code length used by the high-frequency neural network coded data block is greater than the code length used by the low-frequency neural network coded data block; S45, performing entropy coding processing on the neural network coded data block after adaptively allocating the code length, using arithmetic coding to generate a compressed bit stream, and performing association verification; S46, outputting the compressed bit stream after the association check as entropy coded data.

5. A novel digital hybrid coding method based on a large model according to claim 1, characterized in that: The S5 specifically includes: S51, performing feature analysis on the entropy coded data, and calculating the time domain variation characteristics and spectrum energy distribution of the entropy coded data block; S52, based on the characteristic analysis result of the entropy coded data block, adopt a dynamic domain conversion method, select discrete cosine transform, discrete wavelet transform or fast Fourier transform according to the characteristics of the entropy coded data, and perform transformation on the entropy coded data block to obtain a transformed entropy coded data block; S53, performing frequency decomposition on the transformed entropy coded data block, dividing it into a number of sub-frequency bands based on spectrum energy density, calculating the signal energy of each sub-frequency band, and determining an optimal decomposition scale based on local gradient change analysis; S54, performing non-uniform quantization on transform coefficients of different sub-bands through transform domain gradient constraints: Among them, E i,j represents the quantized value of the jth frequency component in the i-th entropy coded data block, C i,j represents the transformation coefficient, Δ i,j represents the adaptive quantization step size, round(·) represents the rounding operation, A i,j represents the attention weight of the jth frequency component in the i-th entropy coded data block, γ represents the gradient control coefficient, Represents the local gradient of the transform coefficient along the frequency axis; S55 , performing transform domain reorganization on the non-uniformly quantized entropy coded data block to generate transform coded data.

6. A novel digital hybrid coding method based on a large model according to claim 1, characterized in that: The S6 specifically includes: S61, using the A3C algorithm to optimize the code length allocation of the transform coded data, constructing a state space and an action space, and calculating the optimal code length allocation strategy through a strategy network and a value function: Among them, θ represents the parameter set of the policy network, ξ represents the learning rate, N1 represents the total number of transform coding data blocks, and s i represents the state of the i-th transform coded data block, π θ (s i ) represents the policy network π θ In status i The probability distribution of the selected action is represents the parameter gradient, A i represents the advantage estimate, V θ (s i ) means in state s i The expected total reward obtained by taking the current strategy under the current policy, ← represents the update operation; S62, dynamically adjusting the entropy coding parameters of the transform coded data block according to the code length allocation optimized by the A3C algorithm, adaptively configuring the entropy coding algorithm based on different channel environments, and optimizing the goal to maximize coding efficiency; S63. Use the knowledge distillation method to extract a lightweight encoding model from the diffusion model, and transfer the knowledge of the complex model to the student model through the teacher-student model training strategy: Among them, L KD represents the knowledge distillation loss function, G i represents the i-th transform coded data block, P teacher (G i ) represents the probability distribution of the teacher model, P student (G i ) represents the probability distribution of the student model, D KL (·||·) represents the Kullback-Leibler divergence, which is used to measure the difference between the teacher model and the student model; S64. Based on the lightweight coding model and in combination with changes in the channel environment, the code length and entropy coding parameters of the transform coding data block are dynamically adjusted to generate the final optimized coding data.

Citation Information

Patent Citations

  • Video coding intra-frame code rate control method based on deep reinforcement learning

    CN111294595A

  • Probability entropy modeling image coding, decoding and compression method based on conditional diffusion

    CN117119204A

  • Parallel context modeling using information shared between tiles

    CN117501696A

  • Video compression method and system based on variational auto-encoder improved entropy model

    CN119011851A

  • Universal image fusion method and system based on diffusion model

    CN119151813A

Cited By

  • Fast coding and decoding method and system for real-time audio

    CN120766691A

  • A fast codec method and system for real-time audio

    CN120766691B