AI-based image, video, audio, and PDF lossless compression method and system

CN122845792APending Publication Date: 2026-09-29BEIJING ZHIXINGHUI INFORMATION SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611081709.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

传统方法对所有区域一视同仁地分配相同的量化参数和码率,导致关键区域的细节信息在压缩过程中被过度丢失,而背景区域却占用了不必要的码率资源,造成了编码效率的浪费

Benefits of technology

本发明实施例提供的技术方案带来的有益效果至少包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845792A_ABST
    Figure CN122845792A_ABST
Patent Text Reader

Abstract

This invention discloses an AI-based lossless compression method and system for images, videos, audio, and PDFs, belonging to the field of lossless compression technology. The method includes the following steps: semantic region segmentation, multi-scale semantic feature enhancement, region-differentiated quantization entropy encoding, and generative detail compensation decoding reconstruction. This invention reads multimedia data from images and video frames, distinguishing regions of interest from background regions through semantic segmentation and saliency detection; it uses a pre-trained deep network to extract multi-scale spatial features, combining them with adaptive weighting based on region distribution to generate semantically enhanced feature maps; then, it allocates differentiated quantization parameters according to different region fidelity requirements, quantizing to obtain discrete feature tensors and entropy encoding to generate compressed bitstreams; finally, it decodes and dequantizes to obtain reconstructed feature tensors, which are then processed by a high-frequency detail compensation subnetwork of a generative decoding network to repair textures and sharpen edges, outputting visually lossless, high-quality multimedia files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of lossless compression technology, and in particular to an AI-based lossless compression method and system for images, videos, audio, and PDFs. Background Technology

[0002] PDF documents are widely used in cloud storage, electronic archiving and other scenarios. The volume of multimedia and structured document data is growing exponentially. Traditional standardized compression algorithms (such as JPEG) based on discrete cosine transform and general archiving compression tools cannot simultaneously achieve high compression ratios and lossless, unified compression of multiple file types.

[0003] Traditional compression algorithms often struggle to balance compression efficiency and visual quality, posing significant challenges to storage space and network bandwidth as image resolution continues to increase. Most existing compression methods employ a uniform compression strategy for the entire image or video frame, failing to differentiate based on the importance of different regions within the image content. In visual perception, salient regions (such as faces and foreground objects) are more important than other regions and should be given higher fidelity during compression. Traditional methods allocate the same quantization parameters and bitrate to all regions equally, leading to excessive loss of detail in critical areas during compression, while background areas consume unnecessary bitrate resources, resulting in wasted coding efficiency. In bandwidth-sensitive and ultra-low latency applications, traditional coding schemes based on standards such as H.264 / AVC and HEVC often struggle to simultaneously achieve coding efficiency and perceptual quality of human-like regions (such as faces and gestures), easily leading to problems such as bitrate fluctuations, detail loss, and inconsistent image quality. While deep learning-based image / video compression systems have made progress in rate-distortion performance, most existing methods are limited by insufficient utilization of multi-scale spatial information during feature extraction, ignoring the differences between features at different levels. This leads to problems such as incomplete predicted structure and loss of detail in salient object detection. In the quantization stage, existing techniques typically use static, uniform quantization bit widths, which cannot adapt to the highly diverse data distributions and sensitive features in compression models, nor can they dynamically assign differentiated quantization parameters to different regions based on the semantic information of the image content. In the decoding and reconstruction stage, images or videos reconstructed after lossy compression often suffer from compression artifacts such as texture blurring and reduced edge sharpness. Existing methods have limited ability to recover high-frequency details and texture information lost during compression, making it difficult to achieve truly visually lossless reconstruction results. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide an AI-based lossless compression method and system for images, videos, audio, and PDFs. This method can read multimedia data from images and video frames, distinguish regions of interest from background regions through semantic segmentation and saliency detection; extract multi-scale spatial features using a pre-trained deep network, and generate semantically enhanced feature maps by adaptive weighting based on region distribution; allocate differentiated quantization parameters according to different region fidelity requirements, quantize to obtain discrete feature tensors, and entropy encode to generate compressed bitstreams; finally, decode and dequantize to obtain reconstructed feature tensors, which are then processed by a high-frequency detail compensation subnetwork of a generative decoding network to repair textures and sharpen edges, outputting visually lossless, high-quality multimedia files.

[0005] On the one hand, an AI-based lossless compression method for images, videos, audio, and PDFs is provided. This method includes: Step 1, acquiring the multimedia data to be compressed, which includes video frame sequences and static images; performing semantic segmentation and saliency detection on the multimedia data to identify regions of interest (ROI) and background regions; Step 2, inputting the multimedia data into a pre-trained deep feature extraction network to extract multi-scale spatial features; adaptively weighting the multi-scale spatial features according to the distribution of ROI and background regions to obtain a semantically enhanced feature map; Step 3, inputting the semantically enhanced feature map into an adaptive quantization module; assigning differentiated quantization parameters to different regions of the semantically enhanced feature map according to the preset fidelity requirements of ROI and background regions, and performing quantization processing to obtain a discrete feature tensor; entropy encoding of the discrete feature tensor to generate a compressed bitstream; Step 4, entropy decoding and inverse quantization processing of the compressed bitstream to obtain a reconstructed feature tensor; inputting the reconstructed feature tensor into a pre-trained generative decoding network; performing texture restoration and edge sharpening on the reconstructed feature tensor through a high-frequency detail compensation subnetwork in the generative decoding network to output visually lossless, high-quality multimedia data.

[0006] On the other hand, an AI-based lossless compression system for images, videos, audio, and PDF is provided. This system includes: a semantic region segmentation module, a multi-scale semantic feature enhancement module, a region-differential quantization entropy encoding module, and a generative detail compensation decoding and reconstruction module. The semantic region segmentation module acquires the multimedia data to be compressed, including video frame sequences and still images, and performs semantic segmentation and saliency detection on the multimedia data to identify regions of interest (ROIs) and background regions. The multi-scale semantic feature enhancement module inputs the multimedia data into a pre-trained deep feature extraction network, extracts multi-scale spatial features through the deep feature extraction network, and performs multi-scale spatial feature processing based on the distribution of ROIs and background regions. Adaptive weighting yields a semantically enhanced feature map. A region-differentiated quantization entropy encoding module inputs the semantically enhanced feature map into the adaptive quantization module. Based on the preset fidelity requirements of the region of interest and the background region, it assigns differentiated quantization parameters to different regions of the semantically enhanced feature map and performs quantization processing to obtain a discrete feature tensor. Entropy encoding is then applied to the discrete feature tensor to generate a compressed bitstream. A generative detail compensation decoding and reconstruction module performs entropy decoding and dequantization processing on the compressed bitstream to obtain a reconstructed feature tensor. The reconstructed feature tensor is input into a pre-trained generative decoding network. The high-frequency detail compensation subnetwork within the generative decoding network performs texture restoration and edge sharpening on the reconstructed feature tensor, outputting visually lossless, high-quality multimedia data.

[0007] Beneficial effects The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: By combining semantic segmentation with saliency detection, regions of interest (ROIs) and background regions in an image are accurately distinguished. Multi-scale deep feature extraction and spatial attention adaptive weighting mechanisms are used to strengthen effective features of the subject and suppress redundant background features. Differentiated quantization step sizes are then configured for the two types of regions to achieve region-specific fidelity control. Entropy coding is used to achieve efficient bitstream compression. At the decoding end, a generative network and a multi-scale high-frequency detail compensation sub-network are added to repair texture and edge information lost during quantization. Compared to traditional globally unified coding schemes, this approach, on the one hand, can improve the overall compression ratio while ensuring that core ROIs such as faces, text, and products are free from block artifacts, color banding, blurry text, and blurred edges, achieving a lossless subjective visual effect. On the other hand, it can fully compress low-attention backgrounds such as skies and solid-color distant scenes, minimizing storage and bandwidth consumption. This solution employs lightweight AI network structures such as residual encoders, dilated convolutional multi-scale high-frequency extraction, and adaptive arithmetic / Huffman entropy coding. While ensuring the quality of compression and reconstruction, it controls the computational overhead of encoding, decoding, and inference, avoiding the high complexity and latency of traditional high-definition encoding algorithms. The entire architecture is uniformly adapted to static images and video frame sequences, and can be extended to be compatible with integrated compression processing of audio and embedded image PDF files. It eliminates the need to deploy multiple independent encoding and decoding tools, reducing hardware and software deployment and maintenance costs. It solves the industry pain points of traditional compression technologies, such as the inability to balance high compression ratios with the fidelity of the main image, the inability to distinguish semantic regions for differentiated encoding, the permanent loss of details after decoding, the fragmentation of compression systems for multiple file types, and the excessive consumption of computing resources. It can be widely used in various multimedia processing scenarios such as cloud storage, live streaming, electronic archives, and high-speed file transmission. Attached Figure Description

[0008] Figure 1 A flowchart of an AI-based lossless compression method for images, videos, audio, and PDF provided in this application embodiment; Figure 2 A flowchart illustrating the semantic enhancement feature map generation process of the AI-based lossless compression method for images, videos, and audio PDFs provided in this application embodiment; Figure 3 A schematic diagram of the structure of an AI-based lossless compression system for images, videos, audio, and PDF provided in this application embodiment. Detailed Implementation

[0009] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0010] like Figure 1 The diagram shown is a flowchart of an AI-based lossless compression method for images, videos, audio, and PDF provided in this application. The method includes the following steps: Step 1: Obtain the multimedia data to be compressed. The multimedia data includes video frame sequences and still images. Perform semantic segmentation and saliency detection on the multimedia data to identify the regions of interest and background regions in the multimedia data.

[0011] It should be noted that the specific process for semantic segmentation and saliency detection of multimedia data is as follows: Multimedia data is preprocessed, including at least one of resolution normalization, color space conversion and noise reduction, to obtain preprocessed multimedia data. The preprocessed multimedia data is input into a pre-trained semantic segmentation network. The semantic segmentation network extracts multi-scale semantic features from the preprocessed multimedia data and performs category prediction on each pixel in the multimedia data to generate a semantic segmentation map. Semantic segmentation maps are used to label the semantic categories of different objects in multimedia data and their corresponding pixel regions; The preprocessed multimedia data is input into a pre-trained saliency detection network. The saliency detection network extracts global context features and local detail features of the multimedia data, and calculates the saliency response value of each pixel in the multimedia data based on the global context features and local detail features to generate a saliency map. Saliency maps are used to characterize the degree to which different regions in multimedia data attract visual attention.

[0012] Semantic segmentation and saliency detection of multimedia data also include: The semantic segmentation map and the saliency map are fused together. The saliency response values ​​in the saliency map are weighted and corrected according to the semantic category information marked on the semantic segmentation map to obtain the corrected saliency map. The revised saliency map takes into account both the semantic attributes and visual saliency of the objects. Based on the corrected saliency map and the preset saliency threshold, the pixel regions in the multimedia data are divided into regions of interest and background regions. The region of interest is the area formed by pixels whose saliency response value in the corrected saliency map is greater than or equal to a preset saliency threshold, and the background region is the area formed by pixels whose saliency response value in the corrected saliency map is less than a preset saliency threshold.

[0013] In this embodiment, a 4K product display static image with a resolution of 3840×2160 is selected as the multimedia data to be compressed. A complete semantic segmentation and saliency detection process is executed to accurately divide the region of interest (ROI) and background region. The specific implementation process is as follows: First, multimedia data preprocessing is performed, simultaneously executing resolution normalization, color space conversion, and image denoising. In the resolution normalization stage, the original 3840×2160 image is uniformly adjusted to the standard size of 1024×1024 for network input through proportional scaling and edge zero-padding, eliminating network inference bias caused by differences in the original resolution of different video frames and images. Color space conversion... The initial RGB image is converted to the YCbCr color space, separating the luminance and chrominance channels to reduce the interference of color shift on semantic and saliency feature extraction. Noise reduction is achieved using a 5×5 Gaussian convolution kernel for smoothing filtering, removing grainy salt-and-pepper noise from the camera and Gaussian noise from the sensor, resulting in a clean and uniform 1024×1024 preprocessed multimedia data with consistent pixel values. Subsequently, dual-branch parallel inference is initiated, simultaneously feeding the same preprocessed multimedia data into a semantic segmentation network and a saliency detection network that have undergone offline supervised training with massive amounts of image, text, portrait, and product samples. The semantic segmentation network employs a U-Net encoding / decoding architecture, extracting shallow textures and mid-layer objects layer by layer. The system utilizes multi-scale semantic features of the image structure and high-level global scene. Based on a Softmax classifier, it performs multi-class prediction on each pixel of the image, outputting a pixel-level semantic segmentation map. This semantic segmentation map is labeled with different grayscale values ​​to represent all pixel regions corresponding to different semantic objects such as faces, product subjects, text icons, solid-color walls, distant sky, and blurred ground, clearly defining the spatial distribution boundaries of various entities within the image. The saliency detection network adopts a multi-scale contrast perception architecture, extracting global contextual contour features and local texture detail features in parallel. It combines four dimensions—brightness difference, color contrast, edge gradient, and human eye center visual prior—to calculate the saliency of each pixel in the 0-1 interval within the image. The response value generates a saliency map with the exact same size as the input image. In the saliency map, the higher the pixel grayscale value, the stronger the visual attention attraction of the corresponding area. Low attention areas such as solid-color distant scenery and blurred ground have pixel response values ​​close to 0, while high attention areas such as the main body of the product, face, and text have pixel response values ​​close to 1. After the semantic segmentation map and saliency map are output by the dual-branch network inference, the dual-map fusion weighted correction process is performed. Fixed semantic weight coefficients are pre-configured for different categories of objects in the semantic segmentation map. The semantic category weight of core main body such as products, faces, and printed text is set to 1.5, the weight of secondary foreground objects is set to 1.0, and the semantic weight of sky, wall, and distant background is set to 0.3. Read the category weights of corresponding positions in the semantic segmentation map pixel by pixel, multiply them by the original saliency response values ​​of the same positions in the saliency map, and obtain a modified saliency map after semantic attribute correction. This modified saliency map retains the focusing patterns of human vision while relying on semantic categories to suppress false saliency responses from meaningless backgrounds and strengthen the saliency weights of core entity regions, avoiding the defect of misjudging messy textured backgrounds as highly saliency regions simply by relying on visual contrast. Finally, set a fixed saliency threshold of 0.4, and traverse the modified saliency map pixel by pixel. All consecutive pixel sets with a saliency response value greater than or equal to 0.4 are defined as regions of interest, fully covering core visual objects such as the main product, human faces, and promotional text. Pixel sets with a saliency response value less than 0.4 are uniformly defined as background regions, including low-visual-value parts such as solid-color walls, blurred distant scenery, and blank spaces. This completes the precise pixel-level division of the two types of regions, and outputs a binary mask carrying complete region boundary information.

[0014] Step 2: Input the multimedia data into a pre-trained deep feature extraction network. Extract multi-scale spatial features through the deep feature extraction network. Adaptively weight the multi-scale spatial features according to the distribution of the region of interest and the background region to obtain a semantically enhanced feature map.

[0015] It should be noted that, as Figure 2 The flowchart shown is for the semantic enhancement feature map generation process of the AI-based image, video, audio, and PDF lossless compression method provided in this application embodiment. The process is divided into two parallel preprocessing branches. The first branch inputs multimedia data into a deep feature network, outputs low-level and high-level feature maps, and integrates them to form a multi-scale feature set. The second branch generates a region mask map of the original size based on the divided region of interest and background region, and then downsamples the mask to obtain a multi-scale mask adapted to each feature size. After the results of the two branches are combined, attention weight maps with high foreground weight and low background weight and normalization are generated scale by scale. The weight map is multiplied element by element with the feature map of the same scale to obtain a weighted feature map. All weighted features are uniformly upsampled to the original image resolution and then channel-separated to generate a preliminary fusion feature map. Then, feature integration and channel dimensionality reduction are completed through convolutional layers, and finally, a semantic enhancement feature map is output.

[0016] It should be understood that the specific process for obtaining the semantically enhanced feature map is as follows: Multimedia data is input into a deep feature extraction network, which outputs feature maps with different spatial resolutions. The shallow stage outputs high-resolution low-level feature maps rich in local details, while the deep stage outputs low-resolution high-level feature maps containing global semantic information. The low-level feature maps and high-level feature maps together constitute a multi-scale spatial feature set. The deep feature extraction network adopts an encoder based on a residual structure and contains multiple cascaded feature extraction stages. Each stage consists of a convolutional layer, a batch normalization layer, an activation function layer, and a downsampling layer.

[0017] Based on the divided regions of interest and background regions, a region mask map with the same initial resolution as the multimedia data is generated. The pixel positions corresponding to the regions of interest in the region mask map are marked with the first identifier value, and the pixel positions corresponding to the background regions are marked with the second identifier value. The region mask is downsampled so that its spatial size is consistent with the spatial size of each feature map in the multi-scale spatial feature set, thus obtaining a multi-scale region mask. The multi-scale region mask is used to indicate the region category to which each spatial location belongs in the feature map at each scale. Based on the feature map at each scale in the multi-scale spatial feature set, and using a spatial attention mechanism, a spatial attention weight map at that scale is generated according to the corresponding multi-scale region mask map. Specifically: Set the spatial attention weight corresponding to the position marked as the first identifier value in the multi-scale region mask image as the first weight value.

[0018] Generate a spatial attention weight map of that scale based on the multi-scale region mask map of the corresponding scale, and also include: Set the spatial attention weight corresponding to the position marked as the second identifier value as the second weight value, and the first weight value is greater than the second weight value; The weights of the spatial attention weight map are normalized so that the sum of the weights at all locations equals the total number of pixels in the spatial attention weight map. The spatial attention weight map is multiplied element-wise with the feature map at the corresponding scale to obtain the weighted feature map at each scale; the element-wise multiplication enhances the feature response of the region of interest and suppresses the feature response of the background region.

[0019] The weighted feature maps at each scale are upsampled to unify the spatial size of the feature maps at all scales to the initial resolution of the multimedia data. The feature maps at each scale after unification are then stitched together by channel dimension to obtain a preliminary fused feature map that integrates multi-scale information. The initial fused feature map is input into one or more convolutional layers for feature integration and channel dimensionality reduction, and the output is a semantically enhanced feature map.

[0020] In this embodiment, preprocessed multimedia data of commodity images with an original resolution of 1024×1024, after semantic segmentation and saliency detection to divide the region, is used as input. Multi-scale feature extraction, spatial attention adaptive weighting, and multi-scale feature fusion are then performed to generate a semantically enhanced feature map. The details of the entire process are as follows: First, the 1024×1024 multimedia data is fed into an encoder-type deep feature extraction network that has undergone supervised training using massive amounts of high-definition images and video frame samples and is built based on a residual structure. This network consists of four cascaded feature extraction stages. Each stage sequentially stacks a 3×3 standard convolutional layer, a batch normalization layer, a ReLU activation function layer, and a maximum downsampling layer with a stride of 2. Multimedia data propagates layer by layer from the network input. The first shallow layer outputs a low-level feature map with a spatial resolution of 512×512 and 64 channels, fully containing fine local details such as the outline of the product, text strokes, surface texture, and object edges. The second and third intermediate layers output 256×256 and 128×128 transition feature maps, respectively. The fourth deep layer outputs a high-level feature map with a spatial resolution of 64×64 and 512 channels. While the high-level feature map loses subtle pixel textures, it fully stores global semantic information such as the overall scene, the global distribution of the product, and the background layout. The 512×512 low-level feature map, the 256×256 transition feature map, and the 128×128 transition feature map are combined... The feature maps and 64×64 high-level feature maps are uniformly collected to form a complete multi-scale spatial feature set. Next, based on the pixel boundaries of the regions of interest and background regions obtained in the previous step, a binary region mask map that perfectly matches the original 1024×1024 resolution of the multimedia is generated. A first identifier value of 1 and a second identifier value of 0 are pre-set. In the mask map, pixels corresponding to all product main bodies, promotional text, product labels, and other regions of interest are filled with a value of 1, while pixels corresponding to solid-color walls, blurred distant environments, and blank spaces are filled with a value of 0, resulting in a complete original-size region mask map that distinguishes between the two types of regions. Subsequently, mask downsampling adaptation operations are performed on the four sets of feature maps with different resolutions within the multi-scale spatial feature set, using... The step-matching bilinear downsampling method sequentially downsamples the original 1024×1024 region mask to four sizes: 512×512, 256×256, 128×128, and 64×64. These sizes correspond one-to-one with the spatial dimensions of the four feature maps, generating four independent multi-scale region masks. Each multi-scale region mask can accurately indicate whether the spatial pixel position within the corresponding scale feature map belongs to the region of interest or the background region. Then, a spatial attention weight map generation operation is performed scale by scale. For each set of scale feature maps (512×512, 256×256, 128×128, and 64×64), a multi-scale region mask of the same size is retrieved, with a preset first weight value of 1.8 and a second weight value of 0.4. The first weight value is significantly greater than the second weight value. Traverse all pixels in the mask image, assigning a value of 1.8 to the pixel positions of the region of interest (ROI) marked with the first identifier value of 1, and assigning a value of 0.4 to the pixel positions of the background region marked with the second identifier value of 0, generating an initial weight matrix. Then, perform global normalization on this weight matrix, adjusting all weight values ​​based on the total number of pixels in the weight image. After normalization, the sum of the weight values ​​of all pixel positions in the weight image is exactly equal to the total number of pixels in the weight image, resulting in a standardized single-scale spatial attention weight image. Perform element-wise multiplication between each standardized spatial attention weight image and the corresponding scale's original feature image, amplifying the feature response amplitude of the ROI at the pixel level and suppressing the invalid feature amplitude of the background region, resulting in four layers of weighted and corrected feature images at each scale. This effectively filters meaningless textures and redundant color features in the background region, enhancing the detailed feature expression of core areas such as products and text. After weighting, perform bilinear interpolation upsampling on the four layers of weighted feature images at different resolutions. Similarly, the four sets of feature maps (512×512, 256×256, 128×128, and 64×64) are all uniformly upsampled to restore the original 1024×1024 resolution of the multimedia, eliminating the spatial size misalignment problem between multi-scale features. Then, the four weighted feature maps of uniform size are directly stitched and fused along the channel dimension, stacking all levels of detailed and global semantic information to generate a preliminary fused feature map with channel-dimensional superposition and complete multi-scale information fusion. Finally, the preliminary fused feature map with redundant channel dimensions is fed into two consecutively stacked 3×3 convolutional layers. The first convolutional layer is responsible for cross-channel feature integration of multi-channel stacked features, associating similar semantic features extracted from different scales. The second convolutional layer performs channel dimensionality reduction, removing duplicate and redundant feature channels and retaining high-value semantic feature channels. After feature integration and channel compression by the convolutional layers, the final output is a semantically enhanced feature map with uniform size, significantly enhanced main subject details, and greatly suppressed background redundancy features, providing high-quality feature input for the subsequent adaptive differential quantization module.

[0021] Step 3: Input the semantic enhancement feature map into the adaptive quantization module. Based on the preset fidelity requirements of the region of interest and the background region, assign differentiated quantization parameters to different regions of the semantic enhancement feature map and perform quantization processing to obtain discrete feature tensors. Perform entropy encoding on the discrete feature tensors to generate compressed bitstreams.

[0022] Furthermore, the specific process for quantization to obtain discrete feature tensors is as follows: Obtain the preset fidelity requirement parameters corresponding to the region of interest and the background region respectively. The preset fidelity requirement parameters include at least one of the target peak signal-to-noise ratio threshold, the target structural similarity threshold, or the target bit rate allocation ratio. Based on the preset fidelity requirement parameters, the first quantization step size applicable to the region of interest and the second quantization step size applicable to the background region are determined respectively. The first quantization step size is smaller than the second quantization step size so that the region of interest can obtain a higher fidelity than the background region. Based on the divided region of interest and background region, a quantization parameter map with the same spatial size as the semantic enhancement feature map is generated. The first quantization step size is assigned to each pixel position corresponding to the region of interest in the quantization parameter map, and the second quantization step size is assigned to each pixel position corresponding to the background region. The semantic enhancement feature map and the quantization parameter mapping map are spatially mapped. For each feature value in the semantic enhancement feature map, a uniform quantization method with the quantization step size corresponding to that position is used for quantization, mapping the floating-point feature value to a discrete integer value, thus obtaining a discrete feature tensor. The quantization formula is: quantized value = round(feature value / quantization step size), where round is the rounding function.

[0023] The specific steps for entropy encoding discrete feature tensors to generate compressed bitstreams are as follows: The discrete feature tensor is subjected to decorrelation preprocessing, which includes run-length encoding or differential pulse code modulation to eliminate data redundancy between adjacent positions. The probability distribution of symbols after decorrelation preprocessing is statistically analyzed, and the symbols are losslessly encoded using adaptive arithmetic coding or Huffman coding to generate a compressed bitstream. The encoding information of the quantization parameter mapping map and the context model parameters of the entropy encoding are embedded in the header information of the compressed bitstream so that the decoding end can recover the quantization parameters.

[0024] In this embodiment, the object to be processed is a floating-point semantic enhancement feature map that carries enhanced features of the product subject while suppressing background redundancy features. First, the system reads two sets of pre-configured independent preset fidelity requirement parameters for the region of interest (ROI) and the background region. This embodiment simultaneously uses two types of parameters as constraints: a target peak signal-to-noise ratio (PSNR) threshold and a target bitrate allocation ratio. The preset target PSNR threshold for the ROI is 42dB, and the bitrate allocation ratio is 75%. The preset target PSNR threshold for the background region is 30dB, and the bitrate allocation ratio is 25%. The quantization step size for adapting to the two types of regions is derived by reverse reasoning based on the two sets of fidelity indicators. The first quantization step size for adapting to the ROI such as faces, products, and text is calculated through the preset mapping relationship of the encoding. The first quantization step size is set to 0.5, while the second quantization step size, adapted for solid-color walls and blurred distant areas, is set to 2.0. The first quantization step size is much smaller than the second quantization step size. By reducing the quantization error of the region of interest through a smaller quantization interval, the fidelity of the main image details is ensured. The larger quantization interval amplifies the compression of background features and reduces the overall data volume. Subsequently, the region division results output from the previous semantic segmentation and saliency detection are retrieved to generate a quantization parameter mapping map with the same spatial size of 1024×1024 as the semantic enhancement feature map. Region assignment is performed pixel by pixel. All pixels corresponding to the region of interest in the mapping map are uniformly assigned the first quantization step size of 0.5, and all pixels corresponding to the background region are uniformly assigned the second quantization step size of 2.0. Completely construct the parameter mapping relationship that binds spatial location to quantization step size; after completing the mapping map construction, traverse the floating-point semantic enhancement feature map and the quantization parameter mapping map pixel by pixel in spatial alignment. For each floating-point feature value in the feature map, retrieve the dedicated quantization step size in the same coordinate mapping map, and complete the numerical mapping using uniform quantization. The quantization calculation strictly follows the formula quantization value = round(feature value / quantization step size). Use the rounding function to convert the continuously changing floating-point feature amplitude into discrete integer numbers within a finite interval. After traversing all channels and all pixels of the feature map to complete the quantization operation, integrate all discrete integer values ​​to form a discrete feature tensor with the same dimension and size as the original semantic enhancement feature map. After entering the entropy encoding stage, first perform decorrelation preprocessing on the discrete feature tensor. In this embodiment, the differential pulse code modulation method is used to eliminate the numerical correlation redundancy between adjacent pixels of the tensor. Calculate the difference between adjacent signs row by row and channel by channel of the tensor, and use the difference sequence to replace the original Discrete numerical sequences significantly reduce data redundancy. After preprocessing, the frequency of all discrete symbols within the difference sequence is statistically analyzed to construct a dynamic symbol probability distribution model adapted to the current multimedia content. Unlike fixed static probability models, this model optimizes coding efficiency by conforming to the image feature distribution of the current frame. Subsequently, an adaptive arithmetic coding algorithm is used to perform lossless compression coding on all difference symbols, generating a basic binary bitstream. Finally, the complete compressed bitstream is encapsulated. A dedicated metadata storage area is reserved in the bitstream header, synchronously writing the compressed coding data of the quantization parameter mapping map and the context model parameters used in the entropy coding process into the header information. This achieves the binding and storage of quantization region division parameters, coding model parameters, and image feature bitstream. The final encapsulated result is a complete compressed bitstream that can be independently transmitted, stored, and parsed. The decoding end can fully restore the quantization rules and coding context by reading the accompanying parameters embedded in the bitstream header, ensuring that subsequent dequantization and reconstruction processes can accurately reproduce the region-specific quantization logic and avoid decoding distortion.

[0025] Step four: Perform entropy decoding and dequantization on the compressed bitstream to obtain the reconstructed feature tensor; input the reconstructed feature tensor into a pre-trained generative decoding network, and perform texture restoration and edge sharpening on the reconstructed feature tensor through the high-frequency detail compensation subnetwork in the generative decoding network, outputting visually lossless high-quality multimedia data.

[0026] It needs to be clarified that the specific steps for texture restoration and edge sharpening of the reconstructed feature tensor are as follows: Entropy decoding of the compressed bitstream includes: parsing the header information of the compressed bitstream, extracting the encoded data of the quantization parameter mapping map and the context model parameters of the entropy coding from the header information; and using an adaptive arithmetic decoding or Huffman decoding algorithm to perform lossless decoding of the payload of the compressed bitstream based on the context model parameters, to obtain a discrete symbol sequence after decorrelation preprocessing. Reverse run-length decoding or inverse differential pulse code modulation is performed on discrete symbol sequences to recover discrete feature tensors with the same spatial size and number of channels. The discrete feature tensor is dequantized, including: parsing the quantization parameter map from the header information of the compressed bitstream, the quantization parameter map recording the quantization step size corresponding to each spatial position on the semantic enhancement feature map; based on the quantization parameter map, the feature values ​​at each position in the discrete feature tensor are dequantized using the inverse uniform quantization formula to obtain a floating-point reconstructed feature tensor; wherein, the inverse uniform quantization formula is: reconstructed feature value = quantized value × quantization step size, where the quantization step size is taken from the value at the corresponding position in the quantization parameter map, resulting in a floating-point reconstructed feature tensor.

[0027] The reconstructed feature tensor is input into a generative decoding network, which includes a backbone decoding subnetwork and a high-frequency detail compensation subnetwork. The backbone decoding subnetwork adopts an encoder-decoder structure and contains several cascaded upsampling layers, convolutional layers, batch normalization layers, and activation function layers. It is used to gradually restore the reconstructed feature tensor to the original spatial resolution of the multimedia data and output a preliminary reconstructed feature map. The high-frequency detail compensation subnetwork is set in parallel or in series with the backbone decoding subnetwork and is used to enhance the high-frequency details of the reconstructed feature tensor or the preliminary reconstructed feature map.

[0028] Texture restoration and edge sharpening of the reconstructed feature tensor also include: The high-frequency detail compensation subnetwork performs multi-scale high-frequency component extraction on the input feature tensor. The multi-scale high-frequency component extraction includes extracting high-frequency residual features under different receptive fields in parallel through multiple dilated convolutional layers with different dilation rates, and extracting the edge response map of the input feature tensor using the Laplacian operator or the Sobel operator. High-frequency residual features and edge response maps at different scales are concatenated along the channel dimension, and a high-frequency compensation weight map is generated through the attention fusion module. The high-frequency compensation weight map is multiplied element-wise with the input feature tensor, and then superimposed onto the input feature tensor through residual connection to obtain the high-frequency enhanced feature tensor. The preliminary reconstructed feature map output by the backbone decoding subnetwork is fused with the high-frequency enhanced feature tensor output by the high-frequency detail compensation subnetwork. Feature fusion includes element-wise addition or channel concatenation followed by integration through a convolutional layer to obtain a fused enhanced feature map. The fused and enhanced feature map is input into the final image reconstruction convolutional layer, which outputs a reconstructed image or video frame sequence with the same resolution as the original multimedia data. The reconstructed image or video frame sequence achieves a lossless level in both subjective visual quality and objective evaluation metrics. The image reconstruction convolutional layer includes one or more convolutional layers with kernel sizes of 3×3 or 5×5, mapping the fused and enhanced feature map from the feature domain to the pixel domain.

[0029] In this embodiment, the entropy decoding process is initiated first. The decoder reads the compressed bitstream and splits it into a header metadata segment and a payload bitstream segment. The header information is parsed and extracted, separating two key metadata types: compressed encoded data of the quantization parameter mapping map and context model parameters corresponding to adaptive arithmetic coding. The decoding operation environment is initialized based on the extracted context model parameters. The adaptive arithmetic coding logic used by the encoder is matched to perform lossless arithmetic decoding on the payload bitstream. After the decoding operation is completed, the discrete symbol sequence after differential pulse code modulation (DCM) processing in the encoding stage is output. Subsequently, inverse differential pulse code modulation (DCM) operation is performed on the discrete symbol sequence to restore the difference relationship between adjacent values ​​channel by channel and row by row, canceling the data transformation caused by the decorrelation preprocessing at the encoder, and restoring the original discrete feature tensor that is completely consistent with the size and number of channels in the encoding stage. Next, dequantization processing is performed to parse and decode the complete 1024×1024 quantization parameter mapping map from the bitstream header. This mapping map completely retains the quantization step size corresponding to each pixel position in the image, with a step size of 0.5 for the region of interest and a step size of 2 for the background region.0. Traverse all spatial locations and channels of the discrete feature tensor, retrieve the quantization step size within the quantization parameter mapping map of the same coordinates, and strictly reconstruct the feature value according to the inverse uniform quantization formula = quantization value × quantization step size to complete the pixel-by-pixel floating-point restoration, restoring the integer discrete quantization value to a continuous floating-point amplitude, generating a floating-point reconstructed feature tensor without quantization truncation constraints; feed this reconstructed feature tensor into a pre-trained generative decoding network, which consists of a backbone decoding subnetwork and a cascaded high-frequency detail compensation subnetwork. The backbone decoding subnetwork adopts a symmetric encoder-decoder architecture, internally cascading four upsampling layers, 3×3 convolutional layers, batch normalization layers, and ReLU activation function layers, progressively enlarging the size and basic features of the low-resolution reconstructed feature tensor. The image is restored by progressively supplementing its global structure and basic color information, ultimately outputting a preliminary reconstructed feature map of size 1024×1024. This preliminary reconstructed feature map can restore the main outline of the image, but due to quantization compression, it loses a large amount of texture and high-frequency edge details, resulting in slight blurring and blurred edges. A concurrently running high-frequency detail compensation subnetwork uses the original reconstructed feature tensor as input to perform multi-scale high-frequency component extraction. Three sets of dilated convolutional layers with dilation rates of 1, 3, and 5 are set in parallel to capture high-frequency residual features of subtle textures, medium structures, and large contours under small, medium, and large receptive fields, respectively. Simultaneously, the Sobel operator is introduced to perform lateral and vertical gradient calculations on the input tensor to extract clear pixel-level edge response maps. The three sets of dilated layers are then combined... The high-frequency residual features output by the convolution are directly concatenated and fused with the edge response map at the channel dimension. The resulting data is then fed into the built-in channel attention fusion module to calculate the contribution weight of high-frequency information in each channel, generating a globally adapted high-frequency compensation weight map. This weight map is then multiplied element-wise with the input reconstructed feature tensor to amplify the effective high-frequency components. The weighted high-frequency features are then superimposed onto the original reconstructed feature tensor through residual connections to eliminate the high-frequency information attenuation caused by quantization, resulting in a high-frequency enhanced feature tensor. Subsequently, the preliminary reconstructed feature map output by the backbone decoding subnetwork and the high-frequency enhanced feature tensor output by the high-frequency detail compensation subnetwork are fused element-wise to superimpose global basic information and repaired texture edge details, generating a fused enhanced feature map with complete detail information. Finally, the fused and enhanced feature map is fed into a two-stage cascaded 3×3 image reconstruction convolutional layer. The two convolutional layers work together to complete the mapping transformation from the feature domain to the pixel domain, outputting a reconstructed product image with a resolution of 1024×1024, completely consistent with the size of the original multimedia data to be compressed. Objective metrics show that the reconstructed image's region of interest PSNR reaches 42dB and its SSIM value is close to 1. Subjectively, the product texture is delicate and clear, the text strokes are sharp and without adhesion, and the object edges are free of jaggedness and blurring. The compressed background area shows no obvious noise or color banding. It achieves visual losslessness standards in both subjective visual perception and objective image quality evaluation metrics, completely solving the technical defects of traditional compression and decoding, such as permanent loss of subject details and decreased image clarity.

[0030] like Figure 3 The diagram shown is a structural schematic of an AI-based lossless PDF compression system for images, videos, and audio provided in this application embodiment. It includes: a semantic region segmentation module, a multi-scale semantic feature enhancement module, a region-differential quantization entropy encoding module, and a generative detail compensation decoding and reconstruction module. The semantic region segmentation module acquires the multimedia data to be compressed, including video frame sequences and still images. It performs semantic segmentation and saliency detection on the multimedia data to identify the regions of interest (ROI) and background regions. The multi-scale semantic feature enhancement module inputs the multimedia data into a pre-trained deep feature extraction network, extracts multi-scale spatial features through the deep feature extraction network, and enhances the multi-scale spatial features according to the distribution of the ROI and background regions. The features are adaptively weighted to obtain a semantically enhanced feature map. A region-differentiated quantization entropy encoding module inputs the semantically enhanced feature map into an adaptive quantization module. Based on the preset fidelity requirements of the region of interest and the background region, it assigns differentiated quantization parameters to different regions of the semantically enhanced feature map and performs quantization processing to obtain a discrete feature tensor. Entropy encoding is then performed on the discrete feature tensor to generate a compressed bitstream. A generative detail compensation decoding and reconstruction module performs entropy decoding and dequantization processing on the compressed bitstream to obtain a reconstructed feature tensor. The reconstructed feature tensor is input into a pre-trained generative decoding network. The high-frequency detail compensation subnetwork in the generative decoding network performs texture restoration and edge sharpening on the reconstructed feature tensor, outputting visually lossless, high-quality multimedia data.

[0031] In this embodiment, the semantic region segmentation module is first activated to complete the preliminary region recognition. This module reads the original product image multimedia data stored locally and simultaneously performs resolution normalization, color space conversion, and Gaussian denoising preprocessing operations on the image. Then, the preprocessed image is fed in parallel into the offline-trained semantic segmentation network and saliency detection network. The semantic segmentation network outputs a semantic segmentation map carrying pixel category labels, distinguishing different semantic objects such as the product subject, descriptive text, solid-color walls, and blurred backgrounds. The saliency detection network calculates the saliency response value of each pixel using global context and local detail features to generate a saliency response value. The saliency map module performs pixel-level weighted fusion on the two images, corrects the original saliency response values ​​based on semantic category weights to obtain a corrected saliency map, and then performs binary classification of pixel regions based on a preset saliency threshold of 0.4. Pixels with saliency response values ​​greater than or equal to the threshold, such as product and text pixels, are designated as regions of interest, while pixels with saliency response values ​​less than the threshold, such as blank walls and distant scenes, are designated as background regions. A binary region mask matching the size of the original image is output to provide accurate regional spatial identification for subsequent modules. After the region mask is output to the multi-scale semantic feature enhancement module, this module feeds the original multimedia image into a deep feature extraction module built on a residual encoder. The network, in its four-level feature extraction stage, sequentially outputs a 512×512 low-level detail feature map, 256×256 and 128×128 mid-level transition feature maps, and a 64×64 high-level global feature map, integrating them to form a complete multi-scale spatial feature set. The module downsamples the original-size region mask output by the semantic region segmentation module according to the resolution of each layer's feature map, generating corresponding multi-scale region mask maps. Based on a spatial attention mechanism, pixels in the region of interest within the mask are assigned a high weight of 1.8, and background regions are assigned a low weight of 0.4. After normalization, a dedicated attention weight map for each scale is generated. The weight map and the corresponding scale feature map are then compared step-by-step. Element-wise multiplication enhances the main features and suppresses redundant background features. All weighted feature maps are upsampled to the original image resolution of 1024×1024 and then stitched together along the channel dimension to form a preliminary fused feature map. Then, multi-layer convolution completes cross-channel feature integration and channel dimensionality reduction, and finally outputs a semantically enhanced feature map with enhanced details and suppressed background redundancy, which is then passed to the downstream module. After receiving the semantically enhanced feature map, the regional differential quantization entropy encoding module first retrieves the system's preset fidelity parameters and sets the PSNR of the target in the region of interest to 42dB, corresponding to a quantization step size of 0.5, and the PSNR of the target in the background region to 30dB, corresponding to a quantization step size of 2.0. Combining the region mask generation and semantic enhancement feature map with a quantization parameter mapping map of the same size, the module matches the quantization step size of the two types of regions pixel by pixel. A uniform quantization formula is used to complete the quantization mapping of all floating-point feature values ​​within the feature map, converting them into discrete integer values ​​to obtain a discrete feature tensor. The module further performs differential pulse code modulation on the discrete feature tensor to achieve decorrelation preprocessing. The probability distribution of the preprocessed symbols is statistically analyzed to construct a dynamic probability model. Adaptive arithmetic coding is used to complete lossless compression to generate a binary bitstream. Simultaneously, the quantization parameter mapping map encoded data and entropy coding context model parameters are written into the bitstream header metadata, encapsulated to obtain... The system generates a complete, storable, and transmittable compressed bitstream file. When high-definition image restoration is required, the system runs a generative detail compensation decoding and reconstruction module to read the compressed bitstream. The module first parses the bitstream header to extract quantization parameter information and encoding context parameters, performs adaptive arithmetic decoding on the bitstream payload to recover the discrete symbol sequence, restores the discrete feature tensor through inverse differential pulse code modulation, and then uses the inverse uniform quantization formula to dequantize pixel by pixel based on the quantization parameter mapping map stored in the header to restore the continuous floating-point reconstructed feature tensor. The reconstructed feature tensor is then fed into a pre-trained generative decoding network consisting of a backbone decoding subnetwork and a cascaded high-frequency detail compensation subnetwork. The backbone decoding subnetwork gradually restores the feature tensor to its original 1024×1024 resolution using multi-level upsampling, convolution, and normalization layers, outputting a preliminary reconstructed feature map with complete basic contours but missing high-frequency details. The high-frequency detail compensation subnetwork uses parallel dilated convolutions with dilation rates of 1, 3, and 5 to extract multi-scale high-frequency residual features, combined with the Sobel operator to extract image edge response maps. After concatenating the multi-scale high-frequency features with the edge features, the attention fusion module generates a high-frequency compensation weight map, which is then weighted and superimposed onto the original reconstructed feature tensor to obtain the high-frequency enhanced feature tensor. The module then gradually merges the preliminary reconstructed feature map with the high-frequency enhanced feature tensor. Elements are added and fused, then fed into two stages of 3×3 image reconstruction convolutional layers to complete the transformation from the feature domain to the pixel domain, ultimately outputting a 1024×1024 reconstructed product image. This image exhibits clear product textures, sharp and unblurred text edges, and seamless color transitions. Objective PSNR and SSIM metrics both meet visual lossless standards. The entire system relies on the collaborative efforts of four modules. It significantly improves compression ratios and reduces storage and transmission bandwidth consumption through regional differential coding, while leveraging an AI-generated high-frequency compensation mechanism to address the permanent loss of detail inherent in traditional compression. This achieves the dual requirements of high compression ratios and visual losslessness for multimedia data such as images and video frames.

[0032] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope and intent of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. An AI-based lossless compression method for images, videos, audio, and PDFs, characterized in that: Includes the following steps: Step 1: Obtain the multimedia data to be compressed. The multimedia data includes video frame sequences and still images. Perform semantic segmentation and saliency detection on the multimedia data to identify the region of interest and background region in the multimedia data. Step 2: Input the multimedia data into a pre-trained deep feature extraction network, extract multi-scale spatial features through the deep feature extraction network, and adaptively weight the multi-scale spatial features according to the distribution of the region of interest and the background region to obtain a semantically enhanced feature map; Step 3: Input the semantic enhancement feature map into the adaptive quantization module. Based on the preset fidelity requirements of the region of interest and the background region, assign differentiated quantization parameters to different regions of the semantic enhancement feature map and perform quantization processing to obtain discrete feature tensors. The discrete feature tensor is entropy encoded to generate a compressed bitstream; Step four: Perform entropy decoding and dequantization on the compressed bitstream to obtain a reconstructed feature tensor; input the reconstructed feature tensor into a pre-trained generative decoding network, and perform texture restoration and edge sharpening on the reconstructed feature tensor through the high-frequency detail compensation subnetwork in the generative decoding network to output visually lossless high-quality multimedia data.

2. The AI-based lossless compression method for images, videos, audio, and PDFs as described in claim 1, characterized in that: The specific process for semantic segmentation and saliency detection of multimedia data is as follows: The multimedia data is preprocessed, and the preprocessing includes at least one of resolution normalization, color space conversion and noise reduction, to obtain preprocessed multimedia data. The preprocessed multimedia data is input into a pre-trained semantic segmentation network. The semantic segmentation network extracts multi-scale semantic features from the preprocessed multimedia data and performs category prediction on each pixel in the multimedia data to generate a semantic segmentation map. The semantic segmentation map is used to label the semantic categories to which different objects in multimedia data belong and their corresponding pixel regions; The preprocessed multimedia data is input into a pre-trained saliency detection network. The saliency detection network extracts global context features and local detail features of the multimedia data, and calculates the saliency response value of each pixel in the multimedia data based on the global context features and local detail features to generate a saliency map. The saliency map is used to characterize the degree to which each region in the multimedia data attracts visual attention.

3. The AI-based lossless compression method for images, videos, audio, and PDFs as described in claim 2, characterized in that: The semantic segmentation and saliency detection of multimedia data also includes: The semantic segmentation map and the saliency map are fused together, and the saliency response values ​​in the saliency map are weighted and corrected according to the semantic category information marked on the semantic segmentation map to obtain the corrected saliency map. The revised saliency map takes into account both the semantic attributes and visual saliency of the object. Based on the corrected saliency map and the preset saliency threshold, the pixel regions in the multimedia data are divided into regions of interest and background regions. The region of interest is the region formed by pixels whose saliency response value in the corrected saliency map is greater than or equal to the preset saliency threshold, and the background region is the region formed by pixels whose saliency response value in the corrected saliency map is less than the preset saliency threshold.

4. The AI-based lossless compression method for images, videos, audio, and PDFs as described in claim 1, characterized in that: The specific process for obtaining the semantically enhanced feature map is as follows: Multimedia data is input into a deep feature extraction network, which outputs feature maps with different spatial resolutions. The shallow stage outputs low-level feature maps, and the deep stage outputs high-level feature maps. The low-level feature maps and the high-level feature maps together constitute a multi-scale spatial feature set. Based on the divided region of interest and background region, a region mask map with the same initial resolution as the multimedia data is generated. The pixel position corresponding to the region of interest in the region mask map is marked as a first identifier value, and the pixel position corresponding to the background region is marked as a second identifier value. The region mask image is downsampled so that its spatial size is consistent with the spatial size of each feature image in the multi-scale spatial feature set, thus obtaining a multi-scale region mask image. The multi-scale region mask image is used to indicate the region category to which each spatial location belongs in each scale feature image. Based on the feature map at each scale in the multi-scale spatial feature set, and using a spatial attention mechanism, a spatial attention weight map at that scale is generated according to the multi-scale region mask map at the corresponding scale. Specifically: Set the spatial attention weight corresponding to the position marked with the first identifier value in the multi-scale region mask image as the first weight value.

5. The AI-based lossless compression method for images, videos, audio, and PDFs as described in claim 4, characterized in that: The step of generating a spatial attention weight map of a given scale based on the multi-scale region mask map of the corresponding scale further includes: Set the spatial attention weight corresponding to the position marked as the second identifier value as the second weight value, and the first weight value is greater than the second weight value; The weight values ​​of the spatial attention weight map are normalized so that the sum of the weights at all positions is equal to the total number of pixels in the spatial attention weight map. The spatial attention weight map is multiplied element-wise with the feature map at the corresponding scale to obtain the weighted feature maps at each scale. The weighted feature maps at each scale are upsampled to unify the spatial size of all feature maps to the initial resolution of the multimedia data. The feature maps at each scale after unification are then stitched together by channel dimension to obtain a preliminary fused feature map that integrates multi-scale information. The preliminary fused feature map is input into one or more convolutional layers for feature integration and channel dimensionality reduction, and a semantically enhanced feature map is output.

6. The AI-based lossless compression method for images, videos, audio, and PDFs as described in claim 1, characterized in that: The specific process for obtaining the discrete feature tensor through quantization is as follows: Obtain the preset fidelity requirement parameters corresponding to the region of interest and the background region respectively. The preset fidelity requirement parameters include at least one of the target peak signal-to-noise ratio threshold, the target structural similarity threshold, or the target bit rate allocation ratio. Based on the preset fidelity requirement parameters, a first quantization step size applicable to the region of interest and a second quantization step size applicable to the background region are determined respectively. The first quantization step size is smaller than the second quantization step size, so that the region of interest can obtain a higher fidelity than the background region. Based on the divided region of interest and the background region, a quantization parameter mapping map with the same spatial size as the semantic enhancement feature map is generated. The first quantization step size is assigned to each pixel position corresponding to the region of interest in the quantization parameter mapping map, and the second quantization step size is assigned to each pixel position corresponding to the background region. The semantic enhancement feature map and the quantization parameter mapping map are spatially mapped. For each feature value in the semantic enhancement feature map, a uniform quantization method with the quantization step size corresponding to the position as the quantization interval is used to quantize the floating-point feature value and map it to a discrete integer value to obtain the discrete feature tensor.

7. The AI-based lossless compression method for images, videos, audio, and PDFs as described in claim 1, characterized in that: The specific steps for entropy encoding the discrete feature tensor to generate a compressed bitstream are as follows: The discrete feature tensor is subjected to decorrelation preprocessing, which includes running-length encoding or differential pulse code modulation of the discrete feature tensor to eliminate data redundancy between adjacent positions. The probability distribution of the symbols after decorrelation preprocessing is statistically analyzed, and the symbols are losslessly encoded using adaptive arithmetic coding or Huffman coding to generate the compressed bitstream. The encoding information of the quantization parameter mapping map and the context model parameters of the entropy encoding are embedded in the header information of the compressed bitstream so that the decoding end can recover the quantization parameters.

8. The AI-based lossless compression method for images, videos, audio, and PDFs as described in claim 1, characterized in that: The specific steps for texture restoration and edge sharpening of the reconstructed feature tensor are as follows: Entropy decoding of the compressed bitstream includes: parsing the header information of the compressed bitstream, extracting the encoded data of the quantization parameter mapping map and the context model parameters of the entropy encoding from the header information; and using an adaptive arithmetic decoding or Huffman decoding algorithm to perform lossless decoding of the payload of the compressed bitstream based on the context model parameters, to obtain a discrete symbol sequence after decorrelation preprocessing. The discrete symbol sequence is subjected to inverse run-length decoding or inverse differential pulse code modulation to recover the discrete feature tensor with the same spatial size and number of channels; The discrete feature tensor is dequantized, including: parsing the quantization parameter mapping map from the header information of the compressed bitstream, wherein the quantization parameter mapping map records the quantization step size corresponding to each spatial position on the semantic enhancement feature map; and dequantizing the feature values ​​at each position in the discrete feature tensor using the inverse uniform quantization formula according to the quantization parameter mapping map to obtain a floating-point reconstructed feature tensor. The reconstructed feature tensor is input into a generative decoding network, which includes a backbone decoding subnetwork and the high-frequency detail compensation subnetwork.

9. The AI-based lossless compression method for images, videos, audio, and PDFs as described in claim 1, characterized in that: The process of performing texture restoration and edge sharpening on the reconstructed feature tensor also includes: The high-frequency detail compensation subnetwork performs multi-scale high-frequency component extraction on the input feature tensor. The multi-scale high-frequency component extraction includes extracting high-frequency residual features under different receptive fields in parallel through multiple dilated convolutional layers with different dilation rates, and extracting the edge response map of the input feature tensor using the Laplacian operator or the Sobel operator. The high-frequency residual features at different scales and the edge response map are concatenated along the channel dimension, and a high-frequency compensation weight map is generated through the attention fusion module. The high-frequency compensation weight map is multiplied element-wise with the input feature tensor, and then superimposed onto the input feature tensor through residual connection to obtain the high-frequency enhanced feature tensor. The preliminary reconstructed feature map output by the backbone decoding subnetwork is fused with the high-frequency enhanced feature tensor output by the high-frequency detail compensation subnetwork. The feature fusion includes element-wise addition or channel concatenation followed by integration through a convolutional layer to obtain a fused enhanced feature map. The fused and enhanced feature map is input into the final image reconstruction convolutional layer, and the output is a reconstructed image or video frame sequence with the same resolution as the original multimedia data. The reconstructed image or video frame sequence achieves a lossless level in both subjective visual quality and objective evaluation indicators.

10. A system applying the AI-based lossless compression method for images, videos, audio, and PDF as described in any one of claims 1-9, characterized in that, include: The module includes a semantic region segmentation module, a multi-scale semantic feature enhancement module, a regional differential quantization entropy encoding module, and a generative detail compensation decoding and reconstruction module. The semantic region segmentation module is used to acquire multimedia data to be compressed, which includes video frame sequences and still images. It performs semantic segmentation and saliency detection on the multimedia data to identify the region of interest and background region in the multimedia data. The multi-scale semantic feature enhancement module is used to input the multimedia data into a pre-trained deep feature extraction network, extract multi-scale spatial features through the deep feature extraction network, and adaptively weight the multi-scale spatial features according to the distribution of the region of interest and the background region to obtain a semantically enhanced feature map. The region-differentiated quantization entropy encoding module is used to input the semantic enhancement feature map into the adaptive quantization module, allocate differentiated quantization parameters to different regions of the semantic enhancement feature map according to the preset fidelity requirements of the region of interest and the background region, and perform quantization processing to obtain a discrete feature tensor; and perform entropy encoding on the discrete feature tensor to generate a compressed bitstream. The generative detail compensation decoding and reconstruction module is used to perform entropy decoding and dequantization on the compressed bitstream to obtain a reconstructed feature tensor. The reconstructed feature tensor is then input into a pre-trained generative decoding network, and the high-frequency detail compensation subnetwork in the generative decoding network performs texture restoration and edge sharpening on the reconstructed feature tensor to output visually lossless high-quality multimedia data.