Multi-dimensional optimized audio and video lossless compression method and system, computer equipment and storage medium
Through the multi-dimensional optimization of audio and video lossless compression method, RGB-HSL-YUV conversion, block processing and deep learning models are used to solve the problem of insufficient fidelity and computing resources in traditional audio and video compression technology, and efficient and low-redundant audio and video compression is achieved, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN202510884858.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Traditional audio and video compression technology is difficult to maintain high fidelity in high dynamic scenes or complex image content, redundant data processing is rough, and computing resource requirements are high, resulting in a contradiction between compression rate and image quality, especially in real-time applications, where there are problems of latency and insufficient processing speed.
The multi-dimensional optimization of audio and video lossless compression method is adopted, and the non-keyframe feature extraction and reconstruction is performed by RGB-HSL-YUV conversion, blocking processing, DCT transformation, quantization and entropy coding combined with the deep learning model to optimize the data representation of keyframes and non-keyframes.
Significantly reduce data redundancy, maintain high fidelity, reduce computing complexity, adapt to different types of audio and video data, and is suitable for streaming media transmission, video storage and real-time communication fields.
Smart Images

Figure CN120390093A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio - video compression, and particularly relates to a multi - dimensional optimized lossless audio - video compression method, system, computer device, and storage medium. Background Art
[0002] Audio - video data compression technology is an important research direction in the field of digital media processing. With the rapid development of the Internet, the generation and dissemination of audio - video content have shown explosive growth. In order to effectively store and transmit these large - scale data, it is particularly important to develop efficient compression algorithms. Compression technology not only needs to reduce data redundancy as much as possible while ensuring video quality, but also needs to support real - time transmission and playback. In recent years, the application of emerging technologies such as deep learning has provided new ideas for audio - video compression, promoting the progress and development of related technologies.
[0003] Traditional audio - video compression technologies mainly rely on methods such as discrete cosine transform (DCT), motion compensation, and predictive coding. Common coding standards such as H.264 and H.265 analyze the similarity between video frames, utilize the differences between key frames (I - frames) and non - key frames (P - frames, B - frames) to reduce redundant information. Specifically, traditional methods first divide video frames into blocks, then perform DCT transformation on each block, extract frequency - domain information, and perform quantization processing on it. Although these technologies reduce the amount of data to a certain extent, when dealing with complex scenes, they still face the contradiction between compression ratio and image quality.
[0004] Although traditional compression technologies have achieved certain success in video data processing, there are still some obvious defects. First, in high - dynamic scenes or complex image content, traditional methods often have difficulty maintaining the high - fidelity of audio - video, resulting in a decline in image quality after compression. Second, the existing technologies' processing methods for redundant data are relatively rough, unable to fully utilize the multi - dimensional features of images, resulting in ineffective reduction of data redundancy. In addition, traditional methods have high requirements for computing resources, especially in real - time applications, which may lead to delays and insufficient processing speed. Summary of the Invention
[0005] The purpose of the present invention is to provide a lossless audio - video compression method based on multi - dimensional optimization for the problems existing in the above - mentioned prior art.
[0006] The technical solution for achieving the purpose of the present invention is: A multi - dimensional optimized lossless audio - video compression method, the method comprising:
[0007] Step 1, collect an audio - video sequence;
[0008] Step 2, divide the audio - video sequence into key frames and non - key frames;
[0009] Step 3: Perform multi-dimensional information optimization for the key frames;
[0010] Step 4: Compress the key frames;
[0011] Step 5: Optimize the non-key frame data based on the multi-dimensional optimization information obtained in Step 3;
[0012] Step 6: Compress the optimized non-key frame data.
[0013] Further, in Step 3, for the key frames, performing multi-dimensional information optimization specifically includes:
[0014] Step 3-1: Convert the key frame image from RGB to HSL and then to YUV;
[0015] Step 3-2: Divide the image converted in Step 3-1 into multiple pixel blocks, and characterize the luminance and chrominance information of each pixel block to obtain the luminance prediction value and chrominance prediction value of each pixel block;
[0016] Step 3-3: Perform DCT transformation on the image processed in Step 3-2.
[0017] Further, during the conversion process of Step 3-1, it also includes:
[0018] Set a first quantization threshold;
[0019] If any component of a pixel is less than the first quantization threshold, set it to 0.
[0020] Further, in Step 3-2, for each pixel block, construct a prediction model based on the surrounding pixel information, and use this prediction model to obtain the luminance prediction value and chrominance prediction value; the prediction model is specifically:
[0021] ; In the formula, represents the luminance / chrominance prediction value of the pixel block (i, j), N represents the total number of pixels in the pixel block (i, j), represents the luminance value of the k-th pixel in this pixel block, is the corresponding weight, and is adaptively adjusted according to the local gradient information of the pixel.
[0022] Further, Step 5 specifically includes:
[0023] Step 5-1: Convert the non-key frame image from RGB to HSL and then to YUV, and extract the luminance information, chrominance information, and chrominance-to-luminance residual;
[0024] Step 5-2: Divide the image converted in Step 5-1 into multiple pixel blocks. For each pixel block, extract the pixel means in the horizontal and vertical directions;
[0025] Step 5-3: In the luminance domain, for each pixel block, calculate the error values between the actual pixel values and the luminance prediction value and chrominance prediction value of the key frame respectively. The calculation formula is as follows:
[0026]
[0027] In the formula, E represents the error value, I represents the actual pixel value, that is, the pixel mean extracted in Step 5-2, and P represents the luminance prediction value or chrominance prediction value of the key frame;
[0028] Step 5-4: Use the deep model to reconstruct the luminance information, chrominance information and chrominance-to-luminance residual of the non-key frame. The specific process includes:
[0029] Use the multi-dimensional optimization information of the key frame obtained in Step 2 and the error value obtained in Step 5-3 to train the deep learning model; among them, the multi-dimensional optimization information of the key frame is used for supervised learning to construct a non-linear mapping relationship of luminance information, chrominance information and chrominance-to-luminance residual;
[0030] Input the luminance information, chrominance information and chrominance-to-luminance residual of the non-key frame into the trained deep learning model to achieve reconstruction;
[0031] Step 5-5: Perform DCT transformation on the image reconstructed in Step 5-4.
[0032] Further, in Step 5-1, the calculation formula for the chrominance-to-luminance residual is as follows:
[0033]
[0034] In the formula, R represents the chrominance-to-luminance residual, C represents the extracted chrominance information, Y represents the luminance information of the key frame optimized in Step 3, represents the influence function of luminance information on chrominance, and adopts linear or non-linear mapping.
[0035] Further, in Step 4 and Step 6, compress the key frame according to the video standard to ensure that the compressed data conforms to the standard format.
[0036] Further, in Step 4 or Step 6, for the data after DCT transformation, adopt an entropy coding method based on the video standard for compression. The specific process includes:
[0037] Step 4-1: Quantize the DCT transformation coefficient matrix. The quantization process controls the precision of different frequency components through the quantization matrix. The formula is as follows:
[0038]
[0039] Among them, represents the quantized DCT transform coefficient, is the original DCT transform coefficient, is the corresponding element of the standard quantization matrix;
[0040] Step 4-2: Arrange the quantized DCT transform coefficients in a "zigzag" scan order to form a one-dimensional data stream;
[0041] Step 4-3: Use Huffman coding to encode the quantized DCT transform coefficients, construct a variable-length coding table based on the probability of data occurrence, assign relatively short codes in the variable-length coding table to high-frequency data, and assign relatively long codes in the variable-length coding table to low-frequency data.
[0042] On the other hand, a multi-dimensional optimized audio and video lossless compression system is provided. The system includes:
[0043] The first module is used to collect audio and video sequences;
[0044] The second module is used to divide the audio and video sequences into key frames and non-key frames;
[0045] The third module is used to perform multi-dimensional information optimization on the key frames;
[0046] The fourth module is used to compress the key frames;
[0047] The fifth module is used to optimize the non-key frame data based on the multi-dimensional optimization information obtained by the third module;
[0048] The sixth module is used to compress the optimized non-key frame data.
[0049] On the other hand, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the multi-dimensional optimized audio and video lossless compression method is implemented.
[0050] Compared with the prior art, the present invention has the following remarkable advantages:
[0051] (1)Improve compression efficiency: By adopting multi-dimensional information optimization and block processing techniques, the generation of redundant data is significantly reduced. During the processing of key frames, the conversion from RGB to HSL and then from HSL to YUV is utilized to optimize the extraction of luminance and chrominance information, minimizing data redundancy. For non-key frames, data representation is further optimized by extracting residual and depth information. This multi-level compression strategy greatly reduces the volume of the final compressed file, saving a large amount of storage space and bandwidth and improving data transmission efficiency.
[0052] (2)Maintain high fidelity: During the compression process, luminance and chrominance block optimization techniques are adopted to ensure the restoration of image details and colors. By setting reasonable quantization thresholds, pixel values in flat areas are effectively set to zero, while areas with rich details retain more information, ensuring image clarity and realism. In this way, even at a high compression ratio, users can still experience a visual effect similar to the original audio-visual content, meeting the requirements of high fidelity.
[0053] (3)Flexibility and adaptability: By introducing a deep learning model for feature extraction and reconstruction of images, it has strong flexibility and adaptability. The model can adaptively learn the features of different contents during training and can better capture image details when dealing with complex scenes. This enables the solution to be applicable not only to static images but also to dynamic videos, widely adapting to various types of audio-visual data.
[0054] (4)Reduce computational complexity: This solution effectively simplifies the computational requirements during compression through the combination of block optimization and high-density hybrid prediction. The discrete cosine transform (DCT) and its optimized form are used to convert the time-domain signal into a frequency-domain representation, effectively suppressing high-frequency noise. During secondary prediction, the information of the previous frame can be used for efficient prediction, reducing the calculation of redundant data and improving the overall processing speed, making it suitable for real-time application scenarios.
[0055] (5)Wide application prospects: The innovative and efficient nature of the solution of the present invention gives it wide application prospects, especially in the fields of streaming media transmission, video storage, and real-time communication. In streaming media platforms, users' requirements for video quality and smoothness are constantly increasing, and the lossless compression technology provided by this solution can meet such requirements, ensuring no distortion loss during transmission. In addition, in high-definition video storage, the high compression ratio of this solution will help users save storage costs and achieve more efficient content management.
[0056] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings
[0057] Figure 1 It is a schematic diagram of the principle of the multi-dimensional optimized audio-visual lossless compression method in an embodiment.
[0058] Figure 2 Schematic diagram of multi-dimensional information optimization for key frames in an embodiment.
[0059] Figure 3 Schematic diagram of optimizing non-key frame data based on multi-dimensional optimized information in an embodiment. Specific implementation manners
[0060] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0061] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0062] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0063] In one embodiment, in combination with Figure 1 , a multi-dimensional optimized audio and video lossless compression method is provided, and the method includes:
[0064] Step 1, collecting an audio and video sequence;
[0065] Step 2, dividing the audio and video sequence into key frames and non-key frames;
[0066] Step 3, performing multi-dimensional information optimization on the key frames;
[0067] Step 4, compressing the key frames;
[0068] Step 5, optimizing the non-key frame data based on the multi-dimensional optimized information obtained in Step 3;
[0069] Here, the non-key frame data is optimized based on the multi-dimensional optimization information of the key frames to improve the efficiency of subsequent compression and the visual quality after decoding.
[0070] Step 6, compress the optimized non-key frame data.
[0071] Furthermore, in one embodiment, in combination with Figure 2 , for the key frames in step 3, multi-dimensional information optimization is performed, specifically including:
[0072] Step 3-1, perform the conversion of RGB-HSL-YUV on the key frame image;
[0073] Here, the image is converted from the RGB space to the HSL and YUV color spaces, and this conversion helps to separate the luminance and chrominance information and improve the processing ability for different-dimensional data.
[0074] Preferably, during the conversion process, it further includes: setting a first quantization threshold; if any component of a pixel is less than the first quantization threshold, set it to 0. This can ensure that low-intensity noise does not affect subsequent processing.
[0075] Here, the first quantization threshold is set based on the image statistical distribution, and the optimized color representation can reduce unnecessary information redundancy and at the same time reduce the influence of high-frequency noise. After conversion, the luminance information Y and chrominance information U, V of the image data in the YUV space will be optimized respectively, where the luminance information Y is used for subsequent structure analysis, and the chrominance information U, V is used for further residual information extraction.
[0076] Step 3-2, divide the image converted in step 3-1 into multiple pixel blocks, and characterize the luminance and chrominance information of each pixel block to obtain the luminance prediction value and chrominance prediction value of each pixel block;
[0077] Here, the pixel block division method is determined based on the adaptive block size, which can maximize the retention of local features while ensuring the calculation efficiency.
[0078] Step 3-3, perform DCT transformation on the image processed in step 3-2.
[0079] Here, the DCT transformation can convert the pixel data from the time domain to the frequency domain, making the energy of the image more concentrated in the low-frequency components, thereby improving the data compression ratio. The formula for the DCT transformation is as follows:
[0080] .
[0081] Among them, represents the frequency domain coefficient after DCT transformation, is the pixel value in the spatial domain, MMM and N respectively represent the number of rows and columns of the image block, and are normalization factors to ensure that the energy of the transformation remains consistent. The core goal of the DCT transformation is to concentrate the energy in the low-frequency part, making the subsequent coding compression more efficient, reducing unnecessary information at the same time, and improving the overall compression ratio. Combining the aforementioned brightness and chroma optimizations, this step can ensure that the visually important features of the key frames are retained as much as possible during the compression process, laying a foundation for the subsequent key frame compression.
[0082] Preferably, in some embodiments, in step 3-2, for each pixel block, a prediction model is constructed based on the surrounding pixel information, and the brightness prediction value and chroma prediction value are obtained using this prediction model to reduce data redundancy; the prediction model is specifically:
[0083]
[0084] In the formula, represents the brightness / chroma prediction value of the pixel block (i,j), N represents the total number of pixels in the pixel block (i,j), represents the brightness value of the kth pixel in this pixel block, is the corresponding weight, and it is adaptively adjusted according to the local gradient information of the pixels.
[0085] Adopting the solution of this embodiment can effectively reduce the brightness error of the key frames, improve the prediction accuracy, and at the same time provide better basic data for the subsequent DCT transformation.
[0086] Furthermore, in one of the embodiments, in combination with Figure 3 , step 5 specifically includes:
[0087] Step 5-1, perform the conversion of RGB-HSL-YUV on the non-key frame image, and extract the brightness information, chroma information, and chroma-to-brightness residual;
[0088] Here, by calculating the chroma-to-brightness residual, the expression of the chroma information is further refined, and the accuracy of color restoration is improved.
[0089] Here, the formula for the chroma-to-brightness residual is:
[0090]
[0091] In the formula, R represents the chroma-to-brightness residual, C represents the extracted chroma information, Y represents the brightness information of the key frame optimized by step 3, Represents the influence function of luminance information on chrominance, usually adopting linear or non-linear mapping to make the chrominance residual information more representative. This residual information is input into the deep learning model as supplementary information during the subsequent reconstruction process to enhance the color restoration degree of non-key frames.
[0092] Step 5-2: Divide the image converted in Step 5-1 into multiple pixel blocks. For each pixel block, extract the pixel means in the horizontal and vertical directions to enhance the perception ability of the local patterns of the image.
[0093] Step 5-3: In the luminance domain, for each pixel block, calculate the error values between the actual pixel value and the luminance prediction value and chrominance prediction value of the key frame respectively. The calculation formula is:
[0094]
[0095] In the formula, E represents the error value, I represents the actual pixel value, that is, the pixel mean extracted in Step 5-2, and P represents the luminance prediction value or chrominance prediction value of the key frame.
[0096] Step 5-4: Use the deep model to reconstruct the luminance information, chrominance information, and chrominance-to-luminance residual of the non-key frame. The specific process includes:
[0097] Use the multi-dimensional optimization information of the key frame obtained in Step 2 and the error values obtained in Step 5-3 to train the deep learning model; among them, the multi-dimensional optimization information of the key frame is used for supervised learning to construct the non-linear mapping relationship of the luminance information, chrominance information, and chrominance-to-luminance residual.
[0098] Input the luminance information, chrominance information, and chrominance-to-luminance residual of the non-key frame into the trained deep learning model to achieve reconstruction.
[0099] Here, the error values are used to train the deep learning model to enable it to more accurately predict the luminance and chrominance information of the non-key frame, reduce the reconstruction error, and improve the video quality.
[0100] Here, reconstruction is performed based on the learned key frame features to make the data quality of the non-key frame as close as possible to the original frame.
[0101] Step 5-5: Perform DCT transformation on the image reconstructed in Step 5-4.
[0102] Here, through DCT transformation, the spatial domain information is converted into frequency domain information for subsequent compression coding to ensure that the data remains efficient and conforms to the standard format during storage and transmission.
[0103] Further, in one of the embodiments, steps 4 and 6 compress the key frames according to the video standard to ensure that the compressed data conforms to the standard format. At the same time, in the compression process, the optimized multi-dimensional information in step 3 is combined to improve the compression efficiency and reconstruction quality.
[0104] Preferably, in some embodiments, after the DCT transformation, the key frame data has been transformed into the frequency domain representation, where the low-frequency components contain the main image information, and the high-frequency components store the details and edge information.
[0105] Therefore, in step 4 or step 6, for the data after the DCT transformation, an entropy coding method based on the video standard is used for compression. Ensure that the compressed data conforms to the standard format, while minimizing data redundancy and improving storage and transmission efficiency. Common entropy coding methods include run-length coding and variable-length coding. Among them, run-length coding reduces the data storage amount by recording the number of consecutive identical values, and variable-length coding allocates different lengths of codes according to the probability of data occurrence to reduce the storage overhead of high-frequency data.
[0106] The specific process of the entropy coding method includes:
[0107] Step 4-1, quantize the DCT transformation coefficient matrix to reduce the data precision of the high-frequency components and improve the compression ratio. The quantization process controls the precision of different frequency components through the quantization matrix, and the formula is as follows:
[0108]
[0109] Among them, represents the quantized DCT transformation coefficient, is the original DCT transformation coefficient, is the corresponding element of the standard quantization matrix;
[0110] Here, the quantization matrix is optimized according to the human visual characteristics. A smaller quantization step is used for the low-frequency components to retain more image details, while a larger quantization step is used for the high-frequency components to reduce the data storage amount and ensure that the compressed image still has a high perceptual quality.
[0111] Step 4-2, arrange the quantized DCT transformation coefficients in the "zigzag" scan order to form a one-dimensional data stream for subsequent entropy coding;
[0112] Step 4-3, use Huffman coding to encode the quantized DCT transformation coefficients. Construct a variable-length coding table according to the probability of data occurrence, allocate relatively short codes in the variable-length coding table for high-frequency data, and allocate relatively long codes in the variable-length coding table for low-frequency data;
[0113] Here, the construction process of Huffman coding is based on statistical information, calculates the probability distribution of different DCT coefficients, constructs a binary tree accordingly, and the coding rule follows the prefix-free principle to ensure the uniqueness of data decoding. The finally obtained coded data conforms to the video standard, ensuring its correct parsing at the video decoding end. At the same time, combined with the aforementioned optimization process, the key frames still retain sufficient visual information after compression, providing a basis for the optimization of subsequent non-key frames.
[0114] In one embodiment, there is provided a multi-dimensional optimized lossless audio and video compression system, the system comprising:
[0115] A first module, configured to collect an audio and video sequence;
[0116] A second module, configured to divide the audio and video sequence into key frames and non-key frames;
[0117] A third module, configured to perform multi-dimensional information optimization on the key frames;
[0118] A fourth module, configured to compress the key frames;
[0119] A fifth module, configured to optimize non-key frame data based on the multi-dimensional optimization information obtained by the third module;
[0120] A sixth module, configured to compress the optimized non-key frame data.
[0121] For the specific limitations of the multi-dimensional optimized lossless audio and video compression system, reference can be made to the limitations of the multi-dimensional optimized lossless audio and video compression method in the above text, which will not be elaborated here. Each module in the above multi-dimensional optimized lossless audio and video compression system can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0122] In one embodiment, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the following is implemented:
[0123] Step 1, collect an audio and video sequence;
[0124] Step 2, divide the audio and video sequence into key frames and non-key frames;
[0125] Step 3, perform multi-dimensional information optimization on the key frames;
[0126] Step 4, compress the key frames;
[0127] Step 5, optimize the non-key frame data based on the multi-dimensional optimization information obtained in Step 3;
[0128] Step 6, compress the optimized non-key frame data.
[0129] For the specific limitations of each step, reference can be made to the limitations of the audio and video lossless compression method with multi-dimensional optimization in the above text, which will not be elaborated here.
[0130] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it realizes:
[0131] Step 1, collect an audio and video sequence;
[0132] Step 2, divide the audio and video sequence into key frames and non-key frames;
[0133] Step 3, perform multi-dimensional information optimization for the key frames;
[0134] Step 4, compress the key frames;
[0135] Step 5, optimize the non-key frame data based on the multi-dimensional optimization information obtained in Step 3;
[0136] Step 6, compress the optimized non-key frame data.
[0137] For the specific limitations of each step, reference can be made to the limitations of the audio and video lossless compression method with multi-dimensional optimization in the above text, which will not be elaborated here.
[0138] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-dimensional optimized lossless audio and video compression method, characterized in that The method includes: Step 1, collecting an audio-visual sequence; Step 2, dividing the audio-visual sequence into key frames and non-key frames; Step 3, performing multi-dimensional information optimization on the key frames; Step 4, compressing the key frames; Step 5, optimizing the non-key frame data based on the multi-dimensional optimization information obtained in Step 3; Step 6, compressing the optimized non-key frame data.
2. The multi-dimensional optimized audio and video lossless compression method according to claim 1, characterized in that, In Step 3, for the key frames, performing multi-dimensional information optimization specifically includes: Step 3-1, converting the key frame image from RGB-HSL to YUV; Step 3-2, dividing the image converted in Step 3-1 into multiple pixel blocks, and characterizing the luminance and chrominance information of each pixel block to obtain the luminance prediction value and chrominance prediction value of each pixel block; Step 3-3, performing DCT transformation on the image processed in Step 3-2.
3. The multi-dimensional optimized lossless audio and video compression method according to claim 2, wherein, During the conversion process of Step 3-1, it also includes: Setting a first quantization threshold; If any component of a pixel is less than the first quantization threshold, setting it to 0.
4. The multi-dimensional optimized lossless audio and video compression method according to claim 2, wherein In Step 3-2, for each pixel block, constructing a prediction model based on the surrounding pixel information and using this prediction model to obtain the luminance prediction value and chrominance prediction value; the specific prediction model is: ; wherein, represents the luminance / chrominance prediction value of the pixel block (i, j), N represents the total number of pixels within the pixel block (i, j), represents the luminance value of the k-th pixel within this pixel block, is the corresponding weight, and is adaptively adjusted according to the local gradient information of the pixels.
5. The multi-dimensional optimized audio and video lossless compression method according to claim 1, wherein Step 5 specifically includes: Step 5-1, converting the non-key frame image from RGB-HSL to YUV, and extracting the luminance information, chrominance information, and chrominance-to-luminance residual; Step 5-2, dividing the image converted in Step 5-1 into multiple pixel blocks, and for each pixel block, extracting the pixel means in the horizontal and vertical directions; Step 5-3, in the luminance domain, for each pixel block, calculating the error values between the actual pixel value and the luminance prediction value and chrominance prediction value of the key frame respectively, and the calculation formula is: ; In the formula, E represents the error value, I represents the actual pixel value, that is, the pixel mean extracted in Step 5-2, and P represents the luminance prediction value or chrominance prediction value of the key frame; Step 5-4, using a deep model to reconstruct the luminance information, chrominance information, and chrominance-to-luminance residual of the non-key frame, and the specific process includes: Using the multi-dimensional optimization information of the key frame obtained in Step 2 and the error values obtained in Step 5-3 to train a deep learning model; among them, the multi-dimensional optimization information of the key frame is used for supervised learning to construct a non-linear mapping relationship of the luminance information, chrominance information, and chrominance-to-luminance residual; Inputting the luminance information, chrominance information, and chrominance-to-luminance residual of the non-key frame into the trained deep learning model to achieve reconstruction; Step 5-5, performing DCT transformation on the image reconstructed in Step 5-4.
6. The multi-dimensional optimized lossless audio and video compression method according to claim 5, wherein In Step 5-1, the calculation formula of the chrominance-to-luminance residual is: ; Wherein, R represents the residual from chrominance to luminance, C represents the extracted chrominance information, and Y represents the luminance information of the key frame optimized in step 3. represents the influence function of luminance information on chrominance, and linear or non-linear mapping is adopted.
7. The multi-dimensional optimized lossless audio and video compression method according to claim 1, characterized in that Steps 4 and 6 compress the key frames according to the video standard to ensure that the compressed data conforms to the standard format.
8. The multi-dimensional optimized lossless audio and video compression method according to claim 2, characterized in that, In Step 4 or Step 6, for the data after DCT transformation, an entropy coding method based on the video standard is used for compression, and the specific process includes: Step 4-1, quantizing the DCT transformation coefficient matrix, and the quantization process controls the precision of different frequency components through a quantization matrix, and the formula is as follows: ; Among them, represents the quantized DCT transform coefficient, is the original DCT transform coefficient, is the corresponding element of the standard quantization matrix; Step 4-2: Arrange the quantized DCT transform coefficients in the "Zigzag" scan order to form a one-dimensional data stream; Step 4-3: Use Huffman coding to encode the quantized DCT transform coefficients. Construct a variable-length coding table based on the probability of data occurrence. Assign relatively shorter codes in the variable-length coding table to high-frequency data and relatively longer codes in the variable-length coding table to low-frequency data.
9. A multi-dimensional optimized lossless audio and video compression system based on the method according to any one of claims 1 to 8, characterized in that The system includes: A first module for collecting audio-visual sequences; A second module for dividing the audio-visual sequences into key frames and non-key frames; A third module for performing multi-dimensional information optimization on the key frames; A fourth module for compressing the key frames; A fifth module for optimizing the non-key frame data based on the multi-dimensional optimization information obtained by the third module; A sixth module for compressing the optimized non-key frame data.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Encoding and decoding method, device and system for intra-frame prediction mode of chrominance component
CN110971897A
Video stream acquisition method based on deep learning
CN119299703A
Apparatus, articles of manufacture, and methods for improved adaptive loop filtering in video encoding
US20220109889A1
Hybrid neural network based end-to-end image and video coding method
US20230096567A1
End-to-end video compression method and system based on deep learning, and storage medium
WO2021164176A1