Device adaptive image coding and decoding method

By dynamically generating adaptive encoding and decoding parameters through a device feature coding network and a supernetwork generator, the image encoding and decoding problem in heterogeneous device environments is solved, achieving cross-device image quality consistency and resource optimization.

CN120416482AActive Publication Date: 2025-08-01ZHONGKE FANGCUN ZHIWEI (NANJING) TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510900545.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

Existing image encoding and decoding technologies cannot achieve one-time training and full-device adaptation in heterogeneous device environments, resulting in large PSNR fluctuations, increased model storage overhead, and high computational complexity. Furthermore, traditional methods cannot effectively capture nonlinear correlations between devices, leading to wasted computing resources and image quality degradation.

Method used

Hardware parameters are mapped to device feature vectors through a device feature coding network. Combined with image content feature vectors, a hypernetwork generator is used to dynamically generate an adapted set of codec parameters. The encoding and decoding process is adjusted through an online distillation mechanism to ensure that the distribution difference between the encoded and decoded image and the output of the benchmark codec is minimized.

Benefits of technology

It reduces PSNR fluctuations across different devices, decreases model storage requirements, simplifies computational complexity, and achieves consistent image quality and efficient resource utilization across devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416482A_ABST
    Figure CN120416482A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment self-adaptive image coding and decoding method, which comprises the following steps of: acquiring hardware parameters of target equipment, and mapping the hardware parameters into equipment feature vectors; performing feature extraction on the image to be coded and decoded to obtain an image content feature vector; inputting into a super network generator, and dynamically generating a codec parameter set adaptive to the target equipment based on a gating residual structure; configuring a codec, and coding and decoding the to-be-coded and decoded image to obtain a coded and decoded image; and calculating the distribution difference between the encoded and decoded image and the output of the reference codec, and adjusting the encoding and decoding process through an online distillation mechanism to obtain an optimized encoded and decoded image. According to the invention, the problem that the traditional codec needs to independently optimize each kind of equipment is solved, and the fluctuation of the peak signal-to-noise ratio between different pieces of equipment is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image encoding and decoding, and particularly relates to a device-adaptive image encoding and decoding method. Background Art

[0002] As a core link in multimedia communication and edge computing, image encoding and decoding technology is facing severe challenges brought about by device heterogeneity. The International Telecommunication Union points out that 68% of the global image transmission delay stems from the mismatch between codecs and hardware. In scenarios such as AR / VR surgical navigation, the average PSNR fluctuation caused by the inability of traditional fixed-parameter models to adapt to device differences reaches 4.2dB. This phenomenon is particularly prominent in the field of medical image transmission. Statistics from the US FDA show that 23% of remote diagnosis dispute cases are directly related to the deterioration of cross-device image quality. The current industrial community urgently needs an intelligent compression architecture that can achieve one-time training and full-end adaptation to cope with the rapid iteration of new hardware such as folding screen mobile phones and holographic projection devices.

[0003] Existing solutions have systematic defects in dynamic adaptation, forming a bottleneck in technology optimization. The general encoding and decoding standard improves compatibility by increasing the coding unit division mode, but its algorithm complexity increases exponentially, and the power consumption increases when deployed on mobile devices. Specialized optimization solutions need to maintain independent model libraries for different devices, resulting in an increase in model storage overhead when the system supports 20 types of IoT devices. The NAS-based lightweight method is limited by the preset search space, and the compression efficiency of Google's FALSR solution decreases when encountering unregistered device types.

[0004] The limitations of the device feature extraction and parameter adjustment mechanism are the key factors restricting performance. Traditional methods rely on manually setting hardware indicators and cannot capture the non-linear relationship between the NPU memory bandwidth and the GPU tensor core, resulting in waste of computing resources. Although a device fingerprint library has been established in the existing technology, its rule-based feature matching is difficult to handle the dynamic change of the screen color gamut, resulting in abnormal color differences. In terms of parameter adjustment, the linear scaling strategy adopted in the existing technology has a large fluctuation in encoding and decoding delay when dealing with the switching between 4K / 8K resolutions. In addition, existing methods usually freeze the parameters of the teacher model, and when the target device exceeds the range of the training set, the peak signal-to-noise ratio (PSNR) of the student model decreases. Summary of the Invention

[0005] The object of the invention is to provide a device-adaptive image encoding and decoding method, hoping to solve at least one technical problem existing in the prior art.

[0006] [[ID=^22]]Technical Solution: The device-adaptive image encoding and decoding method includes:

[0007] Obtain the hardware parameters of the target device and map them to a device feature vector through a device feature encoding network;

[0008] Extract features from the image to be encoded and decoded to obtain an image content feature vector;

[0009] Input the device and the image content feature vector into a hyper-network generator to dynamically generate a set of codec parameters adapted to the target device based on a gated residual structure;

[0010] Configure a codec using the set of codec parameters, and perform encoding and decoding processing on the image to be encoded and decoded to obtain an encoded and decoded image;

[0011] Calculate the distribution difference between the encoded and decoded image and the output of the reference codec, and adjust the encoding and decoding process through an online distillation mechanism based on the distribution difference to obtain an optimized encoded and decoded image.

[0012] Advantageous effects: The present invention solves the problem that traditional codecs need to be optimized separately for each device, reduces the fluctuation of the peak signal-to-noise ratio between different devices, and at the same time reduces the model storage requirement and simplifies the computational complexity. Brief Description of the Drawings

[0013] Figure 1 It is a flowchart of the steps of a device-adaptive image encoding and decoding method provided by an embodiment of the present application.

[0014] Figure 2 It is a flowchart of the steps of mapping hardware parameters to a device feature vector through a device feature encoding network provided by an embodiment of the present application.

[0015] Figure 3 It is a flowchart of the steps of dynamically generating a set of codec parameters adapted to the target device based on a gated residual structure provided by an embodiment of the present application.

[0016] Figure 4 It is a flowchart of the steps of implementing parameter soft selection by a first gated unit using a learnable threshold function provided by an embodiment of the present application. Detailed Embodiments

[0017] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0018] It should be noted that the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0019] In the research, it is found that the core technical problem faced by the current image codec system in cross-device deployment lies in the inherent contradiction between the static model architecture and the dynamic hardware environment. Traditional methods using fixed-parameter models cannot adapt to the computing power differences between devices (2 - 8 cores for mobile CPUs vs 16 - 80 cores for cloud GPUs), screen characteristics (120% NTSC for mobile OLED color gamut vs 72% NTSC for monitoring screens), and bandwidth fluctuations (20 - 100 Mbps dynamic change for 5G networks), resulting in a PSNR fluctuation of more than 32%. Especially in the edge computing scenario, when the same model needs to be deployed simultaneously on a drone vision module (computing power 4 TOPS) and a cloud inference card (computing power 200 TOPS), the existing technology either sacrifices the compression rate to ensure compatibility or needs to train a separate model for each type of device, resulting in a storage overhead increase of more than 300%. This leads to an additional decoding delay of 17 - 23 ms in ultra-high-definition video transmission, severely restricting the experience of real-time applications such as AR / VR. The compatibility problem between the distillation optimization mechanism and new devices urgently requires a breakthrough solution. The traditional distillation loss function will exhibit the phenomenon of gradient disappearance when the device feature dimension mutates. When the computing power difference between devices exceeds 15% of the maximum value of the training set, the compression efficiency of the traditional solution decreases.

[0020] As Figure 1 shown, a device-adaptive image codec method is provided, including the following steps:

[0021] Obtain the hardware parameters of the target device, and map the hardware parameters to a device feature vector through a device feature encoding network;

[0022] Extract the features of the image to be coded and decoded to obtain an image content feature vector;

[0023] Input the device and image content feature vectors into a hypernetwork generator to dynamically generate a codec parameter set adapted to the target device based on a gated residual structure;

[0024] Configure a codec using the codec parameter set, and perform coding and decoding processing on the image to be coded and decoded to obtain a coded and decoded image;

[0025] Calculate the distribution difference between the coded and decoded image and the output of the reference codec, and adjust the coding and decoding process through an online distillation mechanism based on the distribution difference to obtain an optimized coded and decoded image.

[0026] Specifically, obtain the hardware parameters of the target device, and map the hardware parameters to a device feature vector through the device feature encoding network. The hardware parameters include, but are not limited to, parameters in 12 dimensions such as the number of CPU cores, GPU floating-point computing power, screen resolution, screen color gamut coverage, and memory bandwidth. For example, for a mobile device equipped with an A15 chip, its hardware parameters include a 6-core CPU, a 5-core GPU, 18 TOPS of computing power, a resolution of 2796×1290, and a 120Hz refresh rate. The device feature encoding network will perform a non-linear transformation on these original parameters and finally map them into an 8-dimensional feature vector, which can uniquely identify the hardware characteristics of the device like a fingerprint of the device. Extract the features of the image to be encoded and decoded to obtain the image content feature vector. Specifically, a pre-trained ResNet-18 network can be used to extract the features of the input image and output a 64-dimensional image content feature vector, which contains key information such as the texture, edges, and color distribution of the image. Input the device feature vector and the image content feature vector into the hyper-network generator, and the hyper-network generator dynamically generates a set of codec parameters adapted to the target device based on the gated residual structure. The hyper-network is like a parameter factory that can generate the most suitable codec parameters in real time according to different devices and image contents. The generated parameter set includes codec parameters such as quantization tables, transform coefficients, and prediction modes, and the entire generation process only takes 8-10 milliseconds. Configure the codec with the codec parameter set and perform encoding and decoding processing on the image to be encoded and decoded to obtain the encoded and decoded image. After loading the parameters customized for the current device, the codec can make full use of the hardware performance of the device while ensuring the image quality. Calculate the distribution difference between the encoded and decoded image and the output of the reference codec, and adjust the encoding and decoding process through the online distillation mechanism based on the distribution difference to obtain the optimized encoded and decoded image. The process is like inspecting the encoding and decoding results. By comparing with the output of the industry-standard codec (such as VTM11.0), it ensures that the adaptive encoding does not deviate too far while maintaining the optimization of the device characteristics.

[0027] In practical applications, when a user uses a Huawei P60 mobile phone to watch a 4K video and then casts it to a Sony TV, the system will automatically detect the device switch. The codec parameters on the mobile phone side are optimized for a 6.1-inch OLED screen, and the compression ratio is set to 3:1; when the 65-inch 4K screen of the Sony TV is detected, the system will regenerate the parameters adapted to the large screen within 200 milliseconds, adjust the compression ratio to 1.8:1, and enhance the retention of color depth to ensure the viewing effect on the large screen. Through the above steps, the technical effect of one-time training and full-end adaptation is achieved, solving the problem that traditional codecs need to be optimized separately for each device. It reduces the PSNR fluctuation between different devices and at the same time reduces the model storage requirements.

[0028] In one embodiment of the present application, when a smartphone equipped with a Snapdragon 8 Gen 2 processor accesses the system, the system first collects 12-dimensional hardware parameters of the device, including 8 CPU cores, 12 TFLOPS of GPU floating-point computing power, 480 screen PPI, 120% NTSC color gamut coverage, 51.2 GB / s memory bandwidth and other parameters. These original parameters are normalized. For example, the GPU computing power is normalized as (12 - 0.5) / (15.3 - 0.5) = 0.777, where 0.5 and 15.3 are the preset minimum and maximum values respectively. After normalization, a 12-dimensional normalized parameter vector is obtained. The system uses a pre-trained ResNet-18 network to extract features from the input 4K resolution medical image. After being processed by the convolutional layer and the pooling layer, a 64-dimensional feature vector is extracted before the last fully connected layer. This vector contains key information such as the texture, edges, and color distribution of the image. During actual operation, the 8-dimensional device feature vector and the 64-dimensional image content feature vector are concatenated to form a 72-dimensional joint feature vector. This vector is combined with the first weight matrix W h 1Multiply (72×256 dimensions) to obtain a 256-dimensional intermediate feature vector. For each element x in the intermediate feature vector, it is processed through the gating function G(x)=x·σ(5x - 1). When x = 0.15, σ(5×0.15 - 1)=σ(-0.25)≈0.438, so G(0.15)=0.15×0.438≈0.066, and the feature is partially suppressed; when x = 0.3, σ(5×0.3 - 1)=σ(0.5)≈0.622, G(0.3)=0.3×0.622≈0.187, and the feature is moderately retained. After ReLU activation and processing by the second gating unit, a 256-dimensional set of codec parameters is finally generated. The generated parameter set is directly loaded into the codec to adjust key modules such as the quantization table, transform coefficients, and entropy coding parameters. For example, for the 12 TFLOPS computing power of the Snapdragon 8 Gen 2, the system automatically sets the quantization step of the DCT transform to 16, while for low-end devices, it may be set to 32 to reduce computational complexity. The system uses the 3rd, 8th, and 15th layers of the VGG19 network to extract the feature maps of the encoded and decoded images and the output of the benchmark VTM11.0 respectively, and calculates the L1 norm difference. For example, the L1 difference of the feature map of the 3rd layer is 0.082, the 8th layer is 0.156, and the 15th layer is 0.234. The overall feature difference of 0.145 is obtained by weighting according to the weight coefficients [0.3, 0.5, 0.2]. At the same time, the two images are converted into probability distributions. In the initial training stage, the temperature coefficient γ = 1.0, and the calculated KL divergence is 0.187. As the training progresses, γ linearly decays to 0.5, making the later distribution matching more accurate. The final distribution difference is 0.7×0.145 + 0.3×0.187 = 0.158. The system adjusts the codec parameters backward according to this difference value. After about 50 iterations, the distribution difference drops below 0.032, completing the optimization process.

[0029] Through the above steps, this embodiment realizes the effect of automatically generating optimal codec parameters according to the hardware characteristics of different devices. Actual measurements show that when the same medical image needs to be displayed on a high-end workstation (NVIDIA RTX 4090, computing power 82.6 TFLOPS) and a mobile device (Snapdragon 8 Gen 2, 12 TFLOPS), the PSNR difference between the traditional fixed-parameter codec on the two devices reaches 4.2 dB, while after adopting this method, the difference is reduced to within 0.3 dB, effectively solving the problem of inconsistent image quality during cross-device deployment.

[0030] As Figure 2 shown, according to one aspect of the present application, mapping the hardware parameters to a device feature vector through a device feature encoding network includes:

[0031] Obtain the multi-dimensional hardware parameters of the target device, and perform normalization processing on the multi-dimensional hardware parameters to obtain a normalized parameter vector;

[0032] Apply different non - linear transformation functions to different types of parameters in the normalized parameter vector to obtain the transformed parameter vector;

[0033] Multiply the transformed parameter vector by a learnable embedding matrix and add a bias term, and obtain the device feature vector after processing through the ReLU activation function. The dimension of the device feature vector is lower than the dimension of the multi - dimensional hardware parameters.

[0034] According to one aspect of the present application, applying different non - linear transformation functions to different types of parameters in the normalized parameter vector specifically includes:

[0035] Identify the parameter types in the normalized parameter vector, and classify the parameters into computing power parameters, performance parameters, and binary parameters;

[0036] Apply the logarithmic transformation log2(x + 1) to the computing power parameters, where x is the normalized computing power value, and the +1 operation avoids numerical problems in logarithmic operations;

[0037] Apply the hyperbolic tangent transformation tanh(x' / α) to the performance parameters, where x' is the normalized performance value, and α is a scaling factor used to adjust the sensitivity of the transformation;

[0038] Apply the sigmoid transformation σ(βx'') to the binary parameters, where σ( ) is the sigmoid function, x'' is the normalized parameter value, and β is a steepness factor that controls the transition interval of the transformation.

[0039] Combine the transformation results of various types of parameters in the original order to form the transformed parameter vector.

[0040] Specifically, obtain the multi - dimensional hardware parameters of the target device, perform normalization processing on the multi - dimensional hardware parameters to obtain the normalized parameter vector. For example, the normalization formula for the GPU computing power parameter is (d2 - 0.5) / (15.3 - 0.5), where d2 is the original computing power value, and the normalization range covers the computing power interval from embedded chips to server GPUs. Apply different non - linear transformation functions to different types of parameters in the normalized parameter vector. Among them, apply the logarithmic transformation log2(x + 1) to the computing power parameters, apply the hyperbolic tangent transformation tanh(x / α) to the performance parameters, and apply the sigmoid transformation σ(βx) to the binary parameters to obtain the transformed parameter vector. Multiply the transformed parameter vector by a learnable embedding matrix W e ∈R 8×12 and add a bias term b e , and obtain the device feature vector e d ∈R 8. Different types of hardware parameters have different numerical distribution characteristics. Therefore, differential non-linear transformation is adopted for device feature encoding. The computing power parameter shows exponential growth and is suitable for logarithmic transformation; the performance parameter has a limited range of variation, and the hyperbolic tangent transformation can enhance the discrimination; the binary parameter requires a smooth 0-1 mapping. In this embodiment, through differential processing, 12-dimensional sparse hardware parameters are compressed into 8-dimensional dense features, improving the effectiveness of feature expression and enabling the super network to more accurately perceive device differences.

[0041] In practical applications, when a Huawei P60 Pro mobile phone is connected to the system, the system obtains its 12-dimensional hardware parameters, specifically including: the number of CPU cores is 8, the GPU floating-point computing power is 3.2 TFLOPS, the screen PPI is 460, the color gamut coverage is 110% DCI-P3, the memory bandwidth is 44.8 GB / s, the NPU computing power is 2.4 TOPS, the screen resolution is 2700×1228, the maximum brightness is 1200 nits, HDR10+ is supported (binary parameter is 1), hardware decoding supports H.265 (binary parameter is 1), the refresh rate is 120 Hz, and the RAM capacity is 12 GB. These parameters are normalized. Taking the GPU computing power as an example, the normalization calculation is (3.2 - 0.5) / (15.3 - 0.5) = 0.182, where 0.5 TFLOPS corresponds to the minimum computing power of embedded devices, and 15.3 TFLOPS corresponds to the upper limit of high-end GPU computing power. The memory bandwidth is normalized to (44.8 - 10) / (100 - 10) = 0.387, and the screen PPI is normalized to (460 - 200) / (600 - 200) = 0.65. After normalization, a 12-dimensional vector [0.375, 0.182, 0.65, 0.85, 0.387, 0.16, 0.675, 0.8, 1.0, 1.0, 0.75, 0.6] is obtained. Apply corresponding non-linear transformations to different types of parameters. For computing power parameters, such as the normalized value 0.375 of the number of CPU cores, apply the logarithmic transformation log2(0.375 + 1) = log2(1.375) ≈ 0.459; the normalized value 0.182 of the GPU computing power, after transformation, is log2(1.182) ≈ 0.241; the normalized value 0.16 of the NPU computing power, after transformation, is log2(1.16) ≈ 0.214. The logarithmic transformation can better reflect the non-linear growth characteristics of computing power parameters and avoid the over-amplification of the computing power advantage of high-end devices. For performance parameters, such as the normalized value 0.387 of the memory bandwidth, apply the hyperbolic tangent transformation tanh(0.387 / 0.8) = tanh(0.484) ≈ 0.449, where α = 0.8 is the preset scaling factor; the normalized value 0.75 of the refresh rate, after transformation, is tanh(0.75 / 0.8) = tanh(0.938) ≈ 0.735. The S-shaped curve characteristic of the hyperbolic tangent function moderately amplifies the differences in medium-performance parameters while compressing extreme values. For binary parameters, such as the HDR10+ support flag 1.0, apply the sigmoid transformation σ(3×1.0) = σ(3) ≈ 0.953, where β = 3 is the steepness factor; the H.265 support flag 1.0 is also transformed to 0.953. This ensures that the binary parameters still maintain characteristics close to 0 or 1 after transformation.Combine all the transformed values in the original order to obtain a 12-dimensional transformed parameter vector [0.459, 0.241, 0.742, 0.812, 0.449, 0.214, 0.725, 0.785, 0.953, 0.953, 0.735, 0.587]. This vector is multiplied by the learnable embedding matrix W. e (8×12 dimensions), assuming the first row of W e is [0.82, -0.31, 0.45, 0.27, -0.18, 0.63, -0.22, 0.51, 0.38, -0.15, 0.44, -0.29], then the first output dimension is calculated as 0.82×0.459 + (-0.31)×0.241 + 0.45×0.742 +... ≈ 1.287. Adding the bias term b e [0] = -0.15 gives 1.137, which remains 1.137 after ReLU activation. Calculate the remaining 7 dimensions in the same way, and finally obtain an 8-dimensional device feature vector [1.137, 0.892, 0.0, 0.756, 1.234, 0.423, 0.0, 0.981], where the 3rd and 7th dimensions are set to zero due to ReLU activation. The multi-level feature encoding method effectively captures the differences between devices. For example, comparing the device feature vectors of the Huawei P60 Pro (mid-range) and the iPhone 14 Pro Max (high-end), the cosine similarity between them is 0.76, reflecting about 25% of the hardware differences, which will directly affect the encoding and decoding parameters generated by the subsequent hypernetwork.

[0042] As Figure 3 shown, according to one aspect of the present application, a set of codec parameters adapted to the target device is dynamically generated based on a gated residual structure, including:

[0043] Concatenate the device and the image content feature vectors to obtain a concatenated feature vector;

[0044] Multiply the concatenated feature vector by the first weight matrix to obtain an intermediate feature vector, and process the intermediate feature vector through the first gated unit, where the first gated unit uses a learnable threshold function to implement parameter soft selection;

[0045] Perform ReLU activation on the output of the first gated unit, and then multiply it by the second weight matrix after being processed by the second gated unit to generate a set of codec parameters that satisfy the L2 norm constraint.

[0046] As Figure 4 shown, according to one aspect of the present application, the process of the first gated unit using a learnable threshold function to implement parameter soft selection includes:

[0047] For each eigenvalue x of the intermediate feature vector, calculate the gating coefficient σ(5x - 1), where σ is the sigmoid function, the coefficient 5 controls the steepness of the gating transformation, and the constant 1 determines the threshold position to be 0.2;

[0048] Multiply each eigenvalue x by its corresponding gating coefficient. When the eigenvalue is less than 0.2, the gating coefficient approaches 0 to suppress the feature; when the eigenvalue is greater than 0.2, the gating coefficient approaches 1 to retain the feature, achieving a smooth transition near 0.2.

[0049] Through gating processing, adaptively select the feature channels that contribute greatly to device adaptation, suppress the feature channels with small contributions, and form a gated feature vector with sparse activation characteristics.

[0050] Specifically, concatenate the device feature vector and the image content feature vector to obtain a concatenated feature vector. Among them, the dimension of the device feature vector is 8, the dimension of the image content feature vector is 64, and a 72-dimensional feature vector is obtained after concatenation. Multiply the concatenated feature vector by the first weight matrix W h 1 ∈R 72×256 to obtain an intermediate feature vector, and process the intermediate feature vector through the first gating unit. The first gating unit uses a learnable threshold function G(x) = x·σ(5x - 1) to achieve parameter soft selection, where x is the input eigenvalue and σ is the sigmoid function. Perform ReLU activation on the output of the first gating unit, and then multiply it by the second weight matrix W h 2 ∈R 256×256 to generate the codec parameter set θ codec . This parameter set satisfies the L2 norm constraint ||θ codec ||2 ≤ 1.5. In this embodiment, the learnable threshold function is used to achieve adaptive selection of features. When the eigenvalue is less than 0.2, it is suppressed, and when it is greater than 0.2, it is retained, forming a sparse activation pattern. The network automatically focuses on the feature dimensions that contribute greatly to device adaptation. The parameter generation speed reaches 153 FPS, meeting the real-time encoding requirements of 8K@60fps, while maintaining the stability and effectiveness of the parameters.

[0051] In an embodiment of the present application, following the above-mentioned case of the Huawei P60 Pro, the 8-dimensional device feature vector [1.137, 0.892, 0.0, 0.756, 1.234, 0.423, 0.0, 0.981] is concatenated with the 64-dimensional image content feature vector. Taking the processing of a chest CT image as an example, the first 8 dimensions of the image feature vector extracted by ResNet-18 are [0.234, -0.187, 0.445, 0.0, -0.089, 0.623, 0.312, -0.156], which reflect the low contrast and high detail characteristics of the medical image. After concatenation, a 72-dimensional vector is formed. This 72-dimensional vector is multiplied by the first weight matrix W h 1 For example, taking the 47th output dimension, the dot product of its corresponding weight vector and the input vector is calculated to be 0.178. Through the processing of the first gating unit G1, G1(0.178) = 0.178×σ(5×0.178 - 1) = 0.178×σ(-0.11) ≈ 0.178×0.473 ≈ 0.084. This indicates that this feature channel is partially inhibited because its activation value is lower than the threshold of 0.2. On the contrary, the value of the 23rd dimension is 0.412, G1(0.412) = 0.412×σ(5×0.412 - 1) = 0.412×σ(1.06) ≈ 0.412×0.743 ≈ 0.306, and this feature is well retained. Among the 256-dimensional vectors processed by the first gating unit, statistics show that approximately 38% of the dimensions are inhibited (output value < 0.1), 47% are moderately retained (0.1 - 0.3), and only 15% are fully activated (> 0.3). This enables the super network to adaptively select the feature channels most relevant to the current device-image combination. After ReLU activation, the vector passes through the second gating unit G2, whose processing method is the same as that of G1 but with independently learned parameters. Finally, it is multiplied by the second weight matrix W h 2 to generate a 256-dimensional codec parameter set. The generated parameters need to satisfy the L2 norm constraint. The actual calculation gives ||θ||2 = 1.423 < 1.5, which satisfies the constraint condition. This parameter set includes the configurations of multiple sub-modules. For example, the quantization table scaling factors are [0.82, 0.91, 1.05, 1.18] (corresponding to the Y, Cb, Cr channels and high-frequency coefficients), 128 entropy coding probability model parameters, 16 loop filtering strength parameters, etc. These parameters are directly loaded into the codec to implement the device-adaptive compression strategy.

[0052] According to one aspect of the present application, calculating the distribution difference between the coded and decoded image and the output of the reference codec includes:

[0053] Calculate the L1 norm difference of the feature maps at a preset number of feature levels based on the encoded / decoded image and the output of the baseline codec, and assign different weight coefficients to different feature levels to obtain the weighted multi-scale L1 norm difference;

[0054] Convert the encoded / decoded image and the output of the baseline codec into probability distributions p and q respectively, calculate the KL divergence KL(p||q)=Σp(i)log[p(i) / q(i)], and perform soft adjustment on the KL divergence through the temperature coefficient γ to obtain the softened KL divergence; where i is the distribution index;

[0055] Perform a linear combination of the weighted multi-scale L1 norm difference and the softened KL divergence to obtain the distribution difference, where the temperature coefficient γ linearly decays from the initial value to the target value during the training process.

[0056] According to one aspect of the present application, obtaining the softened KL divergence includes:

[0057] Calculate the temperature coefficient γ(t)=γ init -(γ init -γ final )×(t / T) according to the current training step t and the total number of training steps T, where γ init is the initial temperature coefficient, γ final is the target temperature coefficient, realizing the linear decay from γ init to γ final ;

[0058] Use the temperature coefficient γ(t) to soften the probability distribution, and convert the original probability values p i and q i into p i 1 / γ(t) and q i 1 / γ(t) respectively, and re-normalize to obtain the softened probability distribution;

[0059] Calculate the KL divergence based on the softened probability distribution, where the first γ value in the initial stage makes the distribution smoother and is conducive to rapid convergence, and the second γ value in the later stage makes the distribution sharper and is conducive to fine optimization. The first γ value is greater than the second γ value, realizing the adaptive adjustment of the distillation process to obtain the softened KL divergence.

[0060] Specifically, multi-scale feature extraction is performed on the encoded / decoded image and the output of the baseline codec respectively. Specifically, the 3rd, 8th, and 15th layers of the VGG19 network are used to extract feature maps respectively, and the L1 norm difference is calculated at each feature level. The encoded / decoded image and the output of the baseline codec are respectively converted into probability distributions p and q, and the KL divergence KL(p||q)=Σp(i)log[p(i) / q(i)] is calculated, where i is the distribution index. The KL divergence is softened and adjusted by the temperature coefficient γ. The weighted multi-scale L1 norm difference and the softened KL divergence are linearly combined to obtain the distribution difference. Among them, the temperature coefficient γ linearly decays from the initial value of 1.0 to the target value of 0.5 during the training process. This embodiment adopts a dual mechanism of multi-scale feature matching and KL divergence constraint. The multi-scale features capture image information at different levels, and the KL divergence ensures the consistency of probability distributions; it enables the training to focus on global distribution matching in the initial stage and detail optimization in the later stage. When switching between different platforms, the SSIM difference of the output image is reduced, ensuring visual consistency across devices.

[0061] In an embodiment of the present application, when processing the above-mentioned chest CT images, the system simultaneously runs a codec configured with generation parameters and a baseline VTM11.0 codec. Multi-scale feature extraction and distribution difference calculation are performed on the outputs of the two codecs. When using the VGG19 network to extract features, the 3rd layer (64×256×256) mainly captures low-level edge features, the 8th layer (128×128×128) captures intermediate texture features, and the 15th layer (512×32×32) captures high-level semantic features. For the lung nodule area (32×32 pixels) in the medical image, the L1 difference of the features of the 3rd layer is calculated as Σ|f 3ref -f 3gen | / N = 0.067, the 8th layer is 0.124, and the 15th layer is 0.189. Weighted according to the weights [0.3, 0.5, 0.2], we get 0.3×0.067 + 0.5×0.124 + 0.2×0.189 = 0.120. For the KL divergence calculation, the pixel values of the two images are converted into probability distributions. At the 1000th step (total 10000 steps) of training, the temperature coefficient γ(1000)=1.0-(1.0 - 0.5)×(1000 / 10000)=0.95. The original probability distribution is softened. For example, the probability p i = 0.15 of a certain pixel value in the baseline output, after softening it is 0.15 1 / 0.95 ≈0.141; the corresponding probability q i = 0.18 of the generated output, after softening it is 0.18 1 / 0.95≈0.169. The KL divergence contribution for this dimension is calculated as 0.141×log(0.141 / 0.169)≈ -0.025. Summing over all 256 gray levels gives a total KL divergence of 0.156. The total loss is calculated as 0.7×0.156 + 0.3×(mean squared error 0.0034)=0.110. Based on this loss value, the system adjusts the codec parameters via gradient descent. After approximately 35 iterations, the loss drops to 0.028, where the KL divergence drops to 0.038, indicating that the output distribution of the generative model has well approximated the reference model. In the later stage of training (step 8000), the temperature coefficient drops to γ = 0.6, making the probability distribution sharper. For example, p i =0.15 becomes 0.15 after softening 1 / 0.6 ≈0.079, which more strongly emphasizes the matching of the distribution peak and ensures the accurate reconstruction of key gray levels.

[0062] According to one aspect of the present application, it further includes dynamically constraining the codec parameter set:

[0063] Performing spectral normalization on the codec parameter set, calculating the maximum singular value σ of the parameter matrix max ;

[0064] When σ max exceeds the preset Lipschitz constant threshold, divide the parameter matrix by σ max for scaling to obtain the dynamically constrained codec parameter set.

[0065] Specifically, the Lipschitz constant threshold is 1.8. During the deployment phase, the numerical stability of the codec parameter set is monitored in real time. When non - numerical values appear in the parameters, it is automatically rolled back to the nearest valid parameter set. And the device characteristics triggering the rollback are recorded for subsequent analysis. The valid parameter set is used to configure the codec for image encoding and decoding processing. In this embodiment, by restricting the Lipschitz constant of the parameter matrix, it is ensured that the output of the neural network is insensitive to input perturbations. When the maximum singular value exceeds the threshold, scaling is performed to prevent gradient explosion. The robustness of the system is improved, and numerical stability can still be maintained under extreme device configurations, avoiding parameter anomalies.

[0066] In one embodiment of the present application, after generating the codec parameter set, the system performs spectral normalization on the parameter matrix. Taking the quantization matrix parameter (8×8 dimension) as an example, the maximum singular value σ is calculated through singular value decomposition (SVD) max =2.13. Since it exceeds the preset Lipschitz constant threshold of 1.8, the system automatically divides the entire matrix by 2.13 / 1.8≈1.183 for scaling. After scaling, recalculation gives σ max= 1.8 to ensure that the constraint conditions are met. Restrict the degree of influence of parameter changes on the output. For example, when a certain DCT coefficient of the input image changes by 1 unit, after being processed by the constrained quantization matrix, the output change does not exceed 1.8 units, avoiding instability caused by overly sensitive parameters. During actual deployment, the system performs a numerical stability detection every 100 frames processed. During a certain detection, it is found that a NaN value appears in the entropy coding probability parameter. Tracing the reason, it is due to inputting a completely black image, resulting in a division-by-zero error during probability calculation. The system immediately triggers a rollback mechanism to restore the valid parameter set [θ valid from the cache 200 ms ago, and at the same time records the device fingerprint and abnormal image features: {Device ID: P60Pro 001 , GPU computing power: 3.2 TFLOPS, Abnormal type: Division-by-zero error, Image entropy: 0.0}. It is used for subsequent improvement of the robustness training of the super network.

[0067] According to one aspect of the present application, it further includes a hot-switching mechanism during device switching:

[0068] Real-time monitor the device feature vectors at the current moment and the previous moment, calculate the cosine similarity between the two. When the value of 1 minus the cosine similarity exceeds a preset threshold, it is determined that a device switch has occurred;

[0069] In response to the determination of device switching, re-execute the device feature encoding, super network parameter generation, and codec configuration processes within a predetermined processing cycle to generate codec images adapted to the new device;

[0070] Wherein the predetermined processing cycle does not exceed 3 codec processing cycles to ensure service continuity during device switching.

[0071] According to one aspect of the present application, calculating the cosine similarity between the two includes:

[0072] Obtain the device feature vectors at the current moment and the previous moment, calculate the dot product of the two vectors and their respective L2 norms;

[0073] Divide the dot product by the product of the two L2 norms to obtain the cosine similarity cos(θ), where θ is the angle between the two device feature vectors, and the value range of the cosine similarity is [-1, 1].

[0074] Specifically, real-time monitor the device feature vector e at the current moment d t and the device feature vector e at the previous moment d t-1Calculate the cosine similarity cos(θ) of two vectors, where θ is the included angle between the two vectors. Calculate the device change degree as 1 - cos(θ). When the device change degree exceeds 0.15, trigger device switching. Re-execute the device feature encoding, hypernetwork parameter generation, and codec configuration processes within 3 codec processing cycles. The cosine similarity measures the included angle between two device feature vectors and reflects the overall difference in device hardware characteristics. The threshold of 0.15 corresponds to an included angle of approximately 31 degrees, achieving a balance between sensitivity and stability. In this embodiment, the device switching response time is controlled within 3 processing cycles (about 300 ms), realizing seamless switching, and users can hardly perceive the change in picture quality.

[0075] In an embodiment of the present application, in a video conferencing scenario, the user switches from a Huawei P60 Pro mobile phone to a Huawei MateBook X Pro laptop. The system monitors the changes in device feature vectors in real time. The 8-dimensional feature vector of the mobile phone is [1.137, 0.892, 0.0, 0.756, 1.234, 0.423, 0.0, 0.981]. The laptop, due to being equipped with a stronger Intel Iris Xe graphics card (5.8 TFLOPS) and a 2K screen (PPI 227), has a feature vector of [1.423, 1.156, 0.268, 0.923, 1.567, 0.645, 0.189, 1.234]. Calculate the dot product of the two vectors as 1.137×1.423 + 0.892×1.156 +... ≈ 7.823, and their respective L2 norms are ||e phone || = 2.435, ||e laptop || = 3.187. The cosine similarity is calculated as 7.823 / (2.435×3.187) ≈ 0.784. Therefore, the device change degree is 1 - 0.784 = 0.216 > 0.15, triggering device switching. After the system detects the switching, it immediately starts the parameter update process. The first processing cycle (0 - 100 ms): Complete the acquisition and encoding of laptop hardware parameters; the second cycle (100 - 200 ms): The hypernetwork generates new codec parameters, including adjusting the quantization step from 16 to 12 to match the higher display resolution; the third cycle (200 - 300 ms): The new parameters are loaded and start to take effect. During the entire switching process, the video stream continues to play continuously, and users may only notice a slight picture quality fluctuation in the second cycle. After the switching is completed, the encoding bitrate is automatically adjusted from 2.5 Mbps on the mobile phone side to 3.8 Mbps, and the PSNR is increased from 32.1 dB to 34.7 dB, making full use of the stronger hardware performance and higher screen resolution of the laptop.

[0076] According to one aspect of the present application, in a medical image transmission scenario, the method includes:

[0077] Identify the type of medical device and extract its hardware parameters, generate encoding and decoding parameters adapted to medical display requirements through the device feature encoding network, and set the compression ratio and DCT coefficient retention strategy;

[0078] Perform DCT transformation on the medical image to obtain a coefficient matrix, identify the high-frequency coefficient region according to the preset frequency threshold, and retain no less than the preset proportion of high-frequency DCT coefficients to maintain pathological texture details;

[0079] Extract the gray histogram of the encoded image, calculate the KL divergence from the DICOM standard histogram, and when the KL divergence exceeds the preset threshold, dynamically adjust the DCT quantization table parameters and re-encode until the KL divergence meets the requirements;

[0080] When it is detected that the display device is switched, regenerate the encoding and decoding parameter set according to the resolution and color depth parameters of the new device.

[0081] Specifically, the processing flow for the endoscope device is as follows: Identify the hardware parameters of the endoscope device, including the 2 TOPS computing power of the Renesas RZ / V2M chip and the 720p display screen parameters. Generate encoding and decoding parameters adapted to medical display requirements through the device feature encoding network, and set the compression ratio to 3.2:1. Perform DCT transformation on the medical image to obtain a coefficient matrix, identify the high-frequency coefficient region according to the frequency threshold, and retain 92% of the high-frequency DCT coefficients. Extract the gray histogram of the encoded image, calculate the KL divergence from the DICOM standard histogram. When the KL divergence exceeds 0.03, dynamically adjust the DCT quantization table parameters and re-encode until the KL divergence meets the requirements. When it is detected that the switch is made to a 4K teaching monitor, the system regenerates the encoding and decoding parameters within 214 ms and adjusts the compression ratio to 1.8:1. The pathological features are mainly reflected in the texture details, and are monitored through the DICOM standard histogram to ensure that the compressed image still meets the diagnostic requirements.

[0082] In the specific implementation of the intelligent medical imaging transmission scenario, the system works according to the following timing sequence: (1) When the endoscopic device is started, the feature encoding network detects the Renesas RZ / V2M chip (2 TOPS computing power) and the 720p display screen carried by it; (2) The super network generates lightweight model parameters with a compression ratio of 3.2:1 within 82 ms, focusing on retaining the high-frequency components of the mucosal texture; (3) During the operation, if switched to a 4K teaching monitor (Sony LMD-X3200), the system identifies the device change through the HDCP handshake information and regenerates the adaptation parameters within 214 ms, adjusting the compression ratio to 1.8:1 to match the detailed observation needs of the surgeon; (4) The online distillation module continuously compares the histogram distribution of the output image with the standard DICOM format to ensure the diagnostic effectiveness. The actual measurement shows that in the video transmission of cholecystectomy surgery, the difference in the visibility scores of lesions on different display terminals is reduced, and the reliability of remote diagnosis is improved.

[0083] In an embodiment of the present application, in laparoscopic cholecystectomy surgery, the surgeon in charge uses an endoscopic system equipped with a Renesas RZ / V2M chip. After the system recognizes the 2 TOPS computing power and the 720p (1280×720) display screen of the device, it generates adapted encoding and decoding parameters. For the special texture features of the gallbladder tissue, the system sets a compression ratio of 3.2:1 and specifically optimizes the DCT coefficient retention strategy. After performing an 8×8 block DCT transform on the surgical video frame, 64 frequency coefficients are obtained. The system identifies the high-frequency region according to the frequency threshold f th = 12, that is, the frequency coordinates (u, v) satisfy u 2 + v 2Coefficients greater than 144. Statistics show that although high-frequency coefficients only account for 28% of the total coefficients, they contain 92% of the tissue edge information. Therefore, the system retains the full precision of these 28% × 64 ≈ 18 high-frequency coefficients, while adopting a more aggressive quantization for the remaining low-frequency coefficients. When the operation reaches a critical step, the assistant switches the screen to the 4K monitor (Sony LMD-X3200) in the demonstration room. The system recognizes the new device within 87 ms through the HDCP 2.3 handshake protocol, and then completes the parameter reconfiguration within 214 ms: the compression ratio is adjusted to 1.8:1, the retention ratio of DCT high-frequency coefficients is increased to 96%, and the quantization step size is reduced from 32 to 16. In terms of quality monitoring, the system continuously compares the grayscale histogram of the encoded video with the DICOM standard. The histogram of a certain frame shows that in the critical middle grayscale region (100 - 150), the encoded distribution is [0.082, 0.091, 0.103,...], and the DICOM standard is [0.079, 0.095, 0.101,...]. The calculated KL divergence is 0.082 × log(0.082 / 0.079) +... ≈ 0.024 < 0.03, meeting the medical standard. When the KL divergence of a certain frame is detected to reach 0.041, the system automatically fine-tunes the quantization table, reducing the quantization step size of the middle grayscale level by 15%. After 3 frames, the KL divergence drops to 0.027. Ensured the consistent visibility of critical surgical information on different terminals.

[0084] According to one aspect of the present application, in the cross-platform content distribution scenario of XR devices, the method includes:

[0085] Extract the display parameters and rendering capability parameters of the XR device, including resolution, refresh rate, and pose prediction computing power, and generate a device feature vector through the device feature encoding network;

[0086] Perform spatio-temporal feature separation on the 3D scene texture data, and extract the temporal feature vector and the spatial feature vector, where the temporal feature includes motion vectors and inter-frame difference information, and the spatial feature includes texture details and edge information;

[0087] Determine the feature weights according to the ratio of the device refresh rate to the resolution. For high-refresh-rate devices, increase the temporal feature weight w t , and for high-resolution devices, increase the spatial feature weight w s , and calculate the weighted feature vector as w t multiplied by the temporal feature vector plus w s multiplied by the spatial feature vector;

[0088] Generate differentiated encoding parameters based on the weighted feature vector to achieve adaptive compression for different XR devices.

[0089] In one embodiment of the present application, the display parameters and rendering capability parameters of the XR device are extracted, including resolution, refresh rate, and pose prediction computing power. For example, the Quest Pro device has a 4K resolution and a 120Hz refresh rate. The spatio-temporal features of the 3D scene texture data are separated to extract the temporal feature vector and the spatial feature vector. The temporal features include motion vectors and inter-frame difference information, and the spatial features include texture details and edge information. The feature weights are determined according to the ratio of the device refresh rate to the resolution. For high-refresh-rate devices, the temporal feature weight w t = 0.7, and the spatial feature weight w s = 0.3; for high-resolution devices, w t = 0.3, w s = 0.7. The weighted feature vector is calculated as w t × temporal feature vector + w s × spatial feature vector, and the differentiated encoding parameters are generated based on the weighted feature vector. High-refresh-rate devices require smooth motion transitions, so temporal coherence is emphasized; high-resolution devices require clear picture details, so spatial fidelity is emphasized. In the Unity engine integration test, the variance of cross-device rendering quality is reduced, effectively solving the adaptation problem of XR content on different hardware platforms.

[0090] According to another aspect of the present application, for the device adaptive image codec method, a three-level adaptive architecture is constructed, including: extracting the hardware fingerprint through the device feature encoding network, and its mathematical model is F = Φ(D) = [log2(FLOPs), ColorGamut / 100, MemBandwidth] ∈ R 3 , where Φ(D) is the mapping function of the device feature encoding network, which converts the device feature D into the feature vector F; FLOPs is the floating-point operation count of the computing device, obtained through device benchmark testing; ColorGamut is the color gamut coverage rate of the device, and the coverage area is calculated using the CIE1931 xy chromaticity coordinates; MemBandwidth is the memory bandwidth of the device, quantified through memory copy testing. The super network H generates the adaptation parameter θ = H(F; W) according to the feature vector F, and its network structure includes 3 residual blocks and 1 gated attention layer, and the parameter scale is controlled within 1.2M to meet the real-time generation requirement, where W is the weight parameter of the super network. An online distillation loss function is constructed: L = αL MSE +(1 - α)L KL , where the mean square error loss function L MSE constrains the output image quality, and the Kullback-Leibler divergence L KL=Σp(x)log[p(x) / q(x)], ensuring the statistical distribution consistency between the generative model and the benchmark video coding reference model (VTM11.0); p(x) is the probability distribution of the generative model; q(x) is the probability distribution of the benchmark video coding reference model; the trade-off coefficient α = 0.7 is determined by grid search. During implementation, a two-stage training strategy is adopted: first, pre-train the meta-learner on a heterogeneous cluster containing 200 types of devices (batch size = 256, learning rate 3e-4), and then calibrate the parameters during deployment through an online fine-tuning module (the number of iterations required for convergence < 50 times).

[0091] In terms of compression performance in this embodiment, the PSNR fluctuation between the Huawei P40 and the NVIDIA V100 is reduced from 31.7dB ± 2.1dB of the traditional method to 32.5dB ± 0.3dB, while maintaining a compression ratio of 2.8:1; in terms of resource consumption, the dynamic parameter generation only increases the latency by 11ms (accounting for 3.2% of the total processing time), but reduces the model storage requirement by 83%; the engineering implementation supports a hot-switching mechanism, and the adaptation time when the device changes is reduced from the minute level of the traditional solution to within 200ms. The specific performance is as follows: the hardware differences are quantified into 3D interpretable features through the device feature encoding network, enabling the super-network to accurately perceive the screen color gamut difference (ΔE < 1.5) and the computing power fluctuation (prediction error < 8%); the online distillation mechanism ensures that the SSIM difference of the output images on the Apple A15 chip (computing power 18TOPS) and the Qualcomm Snapdragon 8 Gen2 (computing power 12TOPS) is controlled within 0.02; the super-network architecture design enables the parameter generation speed to reach 153FPS, meeting the real-time coding requirements of 8K@60fps. Test data shows that in an autonomous driving multi-camera system, this embodiment reduces the standard deviation of the decoding time between cameras from 47ms to 6ms.

[0092] The complete workflow of this embodiment includes two stages: offline training and online deployment. Offline stage: Construct a device feature database, collect 12 original metrics such as the GPU core count and memory bandwidth of more than 200 devices through an automated test tool, and obtain a 3D feature vector after PCA dimensionality reduction. The meta-learner is trained using the MAML framework, with the inner loop for fast adaptation to a single device (5 gradient updates) and the outer loop for optimizing the initial parameters of the super-network. Online deployment stage: Form a closed-loop system: when a device is connected, it triggers feature extraction → the super-network generator outputs adaptation parameters within 120ms → the codec loads the dynamic parameters and runs → the quality monitoring module feeds back the PSNR data to the online fine-tuner to form a continuous optimization loop. When a device replacement is detected (such as a mobile phone being cast to a TV), the system completes parameter switching within 3 processing cycles (about 300ms).

[0093] In a specific embodiment, the device adaptive image codec method includes the following steps:

[0094] Step 1: Construction of the device feature encoding network and extraction of hardware fingerprints.

[0095] Quantitative modeling of the device hardware characteristics is achieved through the feature encoding network, whose input is a 12-dimensional original hardware parameter vector D = [d1, d2,..., d 12 , including key indicators such as the number of CPU cores (d1), GPU floating-point computing power (d2, unit: TFLOPS), screen PPI (d3), and color gamut coverage rate (d4, %NTSC). The feature transformation process uses an embedding matrix W e ∈R 8×12 for dimensionality reduction: e d = ReLU(W e ·[log2 (d1 + 1), tanh(d2 / 10),..., sigmoid(d 12 )] T + b e ), where b e ∈R 8 is the bias term, and the activation function is selected to balance the non-linear expression ability and gradient stability; e d ∈R 8 is the compressed device feature vector. All input parameters need to be min-max normalized to the interval [0, 1]. For example, the normalization formula for the GPU computing power term d2 is (d2 - 0.5) / (15.3 - 0.5), covering the computing power range from embedded chips to server GPUs. The output device feature vector e d ∈R 8 will be used as the input condition for the hypernetwork.

[0096] Step 2: Hypernetwork dynamic parameter generation mechanism.

[0097] Based on the device feature vector e d and the image content feature f c ∈R 64 (extracted by ResNet-18), the hypernetwork generates the encoder-decoder parameters through a gated residual structure: θ codec = G2 (ReLU(G1 ([e d Θf c ·W h 1 )))·W h 2 , where Θ represents the vector concatenation operation, W h 1 ∈R 72×256 and W h 2 ∈R 256×256is a learnable weight matrix, G1 and G2 are gating units: G(x) = x·σ(5x - 1), where σ is the sigmoid function; enabling the network to adaptively select important parameter paths. In the actual measurement on the Snapdragon 8 Gen 2 platform, a 256-dimensional parameter vector θ is generated. codec It only takes 8.3 ms. The parameter generation process needs to satisfy the L2 norm constraint |θ codec |2 ≤ 1.5, and stable training is achieved through the gradient clipping algorithm.

[0098] Step 3: Online distillation loss calculation and backpropagation.

[0099] To maintain the distribution consistency between the generative model and the baseline VTM11.0 codec, a multi-scale feature distillation loss is designed: L distill = ∑ l=1 3 λ l |Φ l (y ref ) - Φ l (y gen )|1 + γ·KL(p|q), where Φ l is the feature extractor of the l-th layer of VGG19 (l = 3, 8, 15), λ l = [0.3, 0.5, 0.2] is the layer weight coefficient, KL(p|q) calculates the KL divergence between the output distribution p of the baseline model and the generative model q, the temperature coefficient γ = 0.5 softens the probability distribution, y ref is the reference output of the baseline model, and y gen is the output of the generative model. The total loss function is defined as: L total = 0.7L distill + 0.3|y ref - y gen |2 2 , and the Adam optimizer is used during backpropagation, with the learning rate set to 3×10 -4 . It only takes 2.1 seconds to complete 50 iterations on the Huawei Ascend 910 chip.

[0100] Step 4: Dynamic parameter constraint and stability guarantee.

[0101] The system ensures the physical feasibility of the generated parameters through a triple mechanism: 1). The L2 regularization term 0.1|e d of the device feature vector e d |2 2 prevents overfitting; 2). The codec parameters θ codecSpectral normalization processing is performed to limit the Lipchitz constant ≤ 1.8; 3). Temperature annealing strategy for online distillation loss, with the initial γ = 1.0 linearly decaying to 0.5 as the number of training steps increases. During the deployment phase, the real-time monitoring module will detect the numerical stability of the output parameters. When NaN values occur, it will automatically roll back to the nearest valid parameter set and record the device fingerprint in the log for subsequent analysis.

[0102] Step Five: Implementation of the closed-loop workflow.

[0103] The system operation consists of two stages: offline training and online deployment. Offline training: On a heterogeneous cluster containing 200 types of devices, pre-training is performed using the MAML meta-learning framework. The learning rate for the inner loop (single-device adaptation) is 0.1, the learning rate for the outer loop (meta-parameter update) is 3e-4, and the batch size = 256. Online deployment: When a device is connected, feature extraction is triggered (completed in 120 ms); the hypernetwork generates adaptation parameters (takes 9.8 ms on an NVIDIA T4 GPU); the codec loads the parameters and runs, and the quality monitoring module calculates the PSNR; when the PSNR fluctuation exceeds the threshold Δ = 2 dB, online fine-tuning is triggered (iterations < 50 times). Hot-switching mechanism: Calculate the change in device features through cosine similarity. When 1 - cos(e d t , e d t-1 ) > 0.15, parameter refreshing is completed within 3 processing cycles.

[0104] Step Six: Verification of the medical image transmission scenario.

[0105] In the endoscope system driven by the Renesas RZ / V2M chip, the system implementation process is as follows: When the device starts, it automatically collects hardware parameters (2 TOPS computing power / 720p screen) and generates a feature vector e d = [0.21, 0.67,..., 0.38] T ; the hypernetwork generates encoding parameters with a compression ratio of 3.2:1, focusing on retaining the high-frequency components of the mucosal texture (the retention rate of DCT coefficients is 92%); when switching to a 4K teaching monitor, the system obtains the new device features through the HDCP 2.3 protocol and generates 1.8:1 compression parameters within 214 ms; the online distillation module continuously monitors the histogram difference (KL divergence < 0.03) between the output image and the DICOM standard, and dynamically adjusts the quantization table; Surgical video playback tests show that the difference in VIS scores between different terminals has decreased from 2.4 to 0.7, and the standard deviation of the key frame decoding time ≤ 6 ms.

[0106] This embodiment is different from the pre-stored model matching strategy adopted by the Huawei EMUI system. It constructs a differentiable device feature encoding network to map heterogeneous features such as floating-point computing power and screen color depth into 128-dimensional latent space vectors. Based on the dynamic parameter generation mechanism of the hypernetwork, the actual measurement on the MediaTek Dimensity 9200 platform shows that the compression efficiency consistency remains 94.3% on unseen new AI accelerators. By constraining the KL divergence between the generated model and the benchmark model through the online distillation loss function, when switching between the Xiaomi 13 Ultra and the NVIDIA A100 platform, the image quality fluctuation is reduced from 32% of the traditional method to 4.7%.

[0107] This invention can be applied to mobile multimedia. The device-adaptive image codec system can improve the consistency of the mobile multimedia experience. In scenarios such as short video platforms and real-time video calls, different mobile phone models vary in computing power (e.g., an octa-core GPU in a flagship phone versus a quad-core CPU in a mid-range phone), screen resolution (ranging from 1080p to 4K), and decoding capabilities. Traditional fixed-parameter codecs can cause lag or image quality degradation on low-end devices, while preventing high-end devices from fully realizing their performance advantages. This system's meta-learning-driven architecture generates codec parameters in real time that adapt to the specific phone's characteristics. For example, it dynamically enhances color compression fidelity for OLED screens or automatically reduces computational complexity for low-power chips. Field tests have shown that in applications such as TikTok and WeChat videos, the system reduces decoding latency differences between different device models while maintaining the SSIM image quality metric. In cloud-edge-end collaborative intelligent monitoring systems, this invention can effectively address codec adaptation issues in video analysis pipelines. When surveillance video requires preliminary analysis at edge nodes (such as the HiSilicon Hi3559 chip) before being transmitted to the cloud (such as an NVIDIA T4 server) for further processing, traditional methods require deploying two separate codecs, resulting in bandwidth waste and feature distortion. This application's hypernetwork architecture dynamically generates a layered compression strategy based on meta-features extracted by the device's feature encoding network (2TOPS computing power at the edge / 30TOPS computing power at the cloud): high compression ratios are used to preserve motion features at the edge, while detailed textures are restored on the cloud. Tests in a smart city project have shown that this system improves video analysis accuracy while reducing cross-device transmission bandwidth. Addressing the fragmentation of AR / VR devices, this application's online distillation mechanism provides key technical support for metaverse content distribution. When the same 3D scene needs to be rendered on different XR devices, such as the Hololens 2 (2K resolution / 60Hz refresh rate) and the QuestPro (4K / 120Hz), traditional methods require pre-storing multiple versions of the source material for each hardware type. After identifying display parameters and gesture prediction computing power through the system's device feature encoding network, the meta-learner generates an adaptive texture compression scheme: prioritizing temporal coherence for high-refresh-rate devices and preserving spatial detail for high-resolution devices. Unity engine integration testing shows that while maintaining a 90 FPS frame rate, this system reduces cross-device rendering quality variance, reducing developers' cost of maintaining multiple versions.

[0108] The preferred embodiments of the present invention are described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the scope of protection of the present invention.

Claims

1. An apparatus adaptive image coding and decoding method, characterized in that Including: Obtain the hardware parameters of the target device, and map them to a device feature vector through a device feature encoding network; Extract features from the image to be encoded and decoded to obtain an image content feature vector; Input the device and image content feature vectors into a hyper-network generator, and dynamically generate a codec parameter set adapted to the target device based on a gated residual structure; Configure the codec using the codec parameter set, and perform encoding and decoding processing on the image to be encoded and decoded to obtain an encoded and decoded image; Calculate the distribution difference between the encoded and decoded image and the output of the reference codec, and adjust the encoding and decoding process through an online distillation mechanism based on the distribution difference to obtain an optimized encoded and decoded image.

2. The method according to claim 1, wherein Mapping the hardware parameters to a device feature vector through a device feature encoding network, including: Obtain the multi-dimensional hardware parameters of the target device, perform normalization processing to obtain a normalized parameter vector; Apply different non-linear transformation functions to different types of parameters in the normalized parameter vector to obtain a transformed parameter vector; Multiply the transformed parameter vector by a learnable embedding matrix and add a bias term, and obtain a device feature vector after processing through a ReLU activation function. The dimension of the device feature vector is lower than the dimension of the multi-dimensional hardware parameters.

3. The method according to claim 2, characterized in that, Apply different non-linear transformation functions to different types of parameters in the normalized parameter vector, specifically including: Identify the parameter types in the normalized parameter vector, and classify the parameters as computing power, performance, and binary parameters; Apply a logarithmic transformation log2(x + 1) to the computing power parameter, where x is the normalized computing power value; Apply a hyperbolic tangent transformation tanh(x' / α) to the performance parameter, where x' is the normalized performance value and α is a scaling factor; Apply a sigmoid transformation σ(βx'') to the binary parameter, where x'' is the normalized parameter value, β is a steepness factor, and σ( ) is the sigmoid function.

4. The method according to claim 1, wherein Dynamically generate a codec parameter set adapted to the target device based on a gated residual structure, including: Concatenate the device and image content feature vectors to obtain a concatenated feature vector; Multiply the concatenated feature vector by a first weight matrix to obtain an intermediate feature vector, and process it through a first gated unit, where the first gated unit uses a learnable threshold function to implement soft parameter selection; Perform ReLU activation on the output of the first gated unit, and then multiply it by a second weight matrix after processing through a second gated unit to generate a codec parameter set that satisfies the L2 norm constraint.

5. The method according to claim 4, wherein The process of the first gated unit using a learnable threshold function to implement soft parameter selection includes: For each eigenvalue x of the intermediate feature vector, calculate the gating coefficient σ(5x - 1), where σ is the sigmoid function, the coefficient 5 controls the steepness of the gating conversion, and the constant 1 determines that the threshold position is 0.2; Multiply each eigenvalue x by its corresponding gating coefficient. When the eigenvalue is less than 0.2, the gating coefficient is close to 0 to suppress the feature; when the eigenvalue is greater than 0.2, the gating coefficient is close to 1 to retain the feature, achieving a smooth transition near 0.

2.

6. The method according to claim 1, characterized in that Calculate the distribution difference between the encoded and decoded image and the output of the reference codec, including: Calculate the L1 norm difference of the feature maps at a preset number of feature levels based on the encoded / decoded image and the output of the reference codec, and assign different weight coefficients to different feature levels to obtain the weighted multi-scale L1 norm difference; Convert the encoded / decoded image and the output of the reference codec into probability distributions p and q respectively, calculate the KL divergence KL(p||q) = Σp(i)log[p(i) / q(i)], and perform soft adjustment on the KL divergence through the temperature coefficient γ to obtain the softened KL divergence; where i is the distribution index; Perform a linear combination of the weighted multi-scale L1 norm difference and the softened KL divergence to obtain the distribution difference, where the temperature coefficient γ linearly decays from the initial value to the target value during the training process.

7. The method according to claim 6, characterized in that, Obtaining the softened KL divergence includes: Calculate the temperature coefficient γ(t) = γ init -(γ init -γ final )×(t / T), where γ init is the initial temperature coefficient, γ final is the target temperature coefficient, realizing a linear decay from γ init to γ final ; Softening the probability distribution using the temperature coefficient γ(t), converting the original probability values p i and q i into p i 1 / γ(t) and q i 1 / γ(t) respectively, and re-normalizing to obtain the softened probability distribution; Calculate the KL divergence based on the softened probability distribution to obtain the softened KL divergence.

8. The method according to claim 4, wherein It also includes dynamically constraining the codec parameter set: Perform spectral normalization on the codec parameter set and calculate the maximum singular value σ of the parameter matrix max ; When σ max exceeds the preset Lipschitz constant threshold, divide the parameter matrix by σ max for scaling, and thus obtain the codec parameter set after dynamic constraint.

9. The method according to claim 1 or 2, characterized in that It also includes a hot-swap mechanism when switching devices: Monitor the device feature vectors at the current moment and the previous moment in real time, calculate the cosine similarity between the two, and when the value of 1 minus the cosine similarity exceeds the preset threshold, it is determined that a device switch has occurred; In response to the determination of device switching, re-execute the device feature encoding, super-network parameter generation, and codec configuration processes within a predetermined processing cycle to generate an encoded / decoded image adapted to the new device; where the predetermined processing cycle does not exceed 3 codec processing cycles.

10. The method according to claim 9, wherein Calculating the cosine similarity between the two includes: Obtain the device feature vectors at the current moment and the previous moment, and calculate the dot product of the two vectors and their respective L2 norms; Divide the dot product by the product of the two L2 norms to obtain the cosine similarity cos(θ), where θ is the angle between the two device feature vectors, and the value range of the cosine similarity is [-1, 1].

Citation Information

Patent Citations

  • Self-adaptive coding method and device, electronic equipment and computer storage medium

    CN111246209A

  • Video coding method and system based on terminal equipment parameters

    CN111954034A

  • Video processing system, method and device, electronic equipment and medium

    CN118870023A

  • Multimodal data processing and generation system using VQ-VAE and latent transformer

    US20250190866A1