Lightweight Encoding Method and System Based on Neighbor Consistency

Through a lightweight encoding method based on near-neighbor consistency, using Gaussian processing and graph neural network model to extract the scale features and prototype centers of video images, the compatibility and information reduction problems of video image encoding in the prior art are solved, and efficient image compression and information retention are achieved.

CN120128711BActive Publication Date: 2025-07-22CHANGSHA CHAOCHUANG ELECTRONICS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510607685.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-22
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing lightweight video image encoding methods have limitations in compatibility and reduction of information, which affect subsequent image processing operations.

Method used

Based on the consistency of the nearest neighbors, multiple smooth images are generated through Gaussian processing, features of different scales are extracted, prototype centers are generated using clustering methods, graph neural network models are constructed, prototype centers and connections of video images are predicted, and Gaussian filters of the best scale are used for encoding.

Benefits of technology

Lossless image compression is realized, encoding efficiency is improved, computing complexity is reduced, rich image information is retained, and information loss is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128711B_ABST
    Figure CN120128711B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight encoding method and system based on neighbor consistency. The method includes: S1: extracting features from images in video image samples at different scales; S2: generating final prototype centers at different scales by using a clustering method based on the scale features; S3: constructing a graph for the video image samples by using the scale features and the final prototype centers; S4: constructing a graph neural network model and training the graph obtained in step S3; S5: using the trained graph neural network model to predict the prototype centers and connections of the video image to be encoded; S6: determining the optimal scale according to the connections of the video image to be encoded, generating a new smoothed image, and encoding the new smoothed image. Based on the similarity of adjacent image blocks in the high and low resolution spaces in the video image, the present invention extracts scale features; and selects images with obvious features by using the graph structure, which is beneficial to lossless compression of images and improves the encoding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video image coding, and particularly to a lightweight coding method and system based on neighborhood consistency. Background Art

[0002] Encoding video images means converting video data into a form suitable for storage and transmission. In this process, steps such as quantization and transformation are used to process video images, and the best compression method is adopted to minimize the file size while ensuring image quality. Video image coding is widely used, which helps to improve the information transmission efficiency and provides users with a high-quality viewing experience.

[0003] At the same time, in order to minimize the computing resources required for encoding and the content stored, lightweight coding of video images has gained wide popularity. This coding method is suitable for resource-constrained environments, greatly reducing the encoding computational complexity and improving the real-time processing ability. Currently, existing image lightweight coding methods include the JPEG XS international standard and static coding based on the HEVC standard. These methods have limitations in practical applications, have compatibility issues, and the reduced amount of information after encoding affects subsequent image processing operations. Summary of the Invention

[0004] Neighborhood consistency means that pixel points and image blocks adjacent in position are similar, and this characteristic is applicable to images with different resolutions. Utilizing neighborhood consistency helps to remove redundant information in images, compress images, and is also beneficial for distinguishing different types of images. In view of this, based on the similarity shown by adjacent image blocks in the high and low resolution spaces, the present invention extracts features at different scales and selects the best scale, and provides a lightweight coding method and system based on neighborhood consistency.

[0005] The technical solution proposed by the present invention is a lightweight coding method based on neighborhood consistency, including the following steps:

[0006] S1: Perform Gaussian processing on the images in the video image sample to obtain multiple smoothed images, extract features from each smoothed image at different scales to obtain the scale features of each smoothed image;

[0007] S2: Based on the scale features of each smoothed image, use the clustering method to generate the final prototype centers at different scales;

[0008] S3: For the video image sample, use the scale features of each smoothed image and the final prototype centers at different scales to construct a graph;

[0009] S4: Construct a graph neural network model, train the graph obtained in step S3 to obtain a trained graph neural network model;

[0010] S5: Use the trained graph neural network model to predict the prototype centers and connection lines of the video image to be encoded;

[0011] S6: Determine the optimal scale according to the connection lines of the video image to be encoded, apply a Gaussian filter with a standard deviation corresponding to the optimal scale to process the video image to be encoded, obtain a new smoothed image, and encode the new smoothed image.

[0012] Optionally, the S1 includes:

[0013] S11: Perform Gaussian processing on the images in the video image samples to obtain multiple smoothed images, and construct a multi-scale space based on the multiple smoothed images;

[0014] S12: In the multi-scale space, extract features from each smoothed image at different scales to obtain the scale features of each smoothed image.

[0015] Optionally, the S11 includes:

[0016] S111: Apply Gaussian filters with standard deviations arranged in ascending order to the images in the video image samples, with each standard deviation corresponding to a Gaussian filter, to generate multiple smoothed images:

[0017] ;

[0018] Among them, represents an image, represents a Gaussian filter with a standard deviation of , represents the smoothed image generated by the Gaussian filter ; the value range of the standard deviation is , and the interval of the value of the standard deviation is ;

[0019] S112: Construct a multi-scale space according to the multiple smoothed images, and use the standard deviation of each smoothed image as the scale of the corresponding smoothed image.

[0020] Optionally, the S2 includes:

[0021] S21: Initialize multiple prototype centers at different scales to obtain multiple initial prototype centers;

[0022] S22: Calculate the distances from the scale features of each smoothed image to each initial prototype center:

[0023] ;

[0024] Among them, Indicates the scale The initial prototype center at Indicates the smooth image numbered in the video image sample; Indicates at the scale The scale feature of the i-th smooth image extracted at; Indicates the modulo operation;

[0025] Divide into the category represented by the nearest prototype center ;

[0026] S23: In each category, calculate the average value of the scale features of all smooth images, and use the average value to update the initial prototype center to which it belongs to obtain the updated prototype center;

[0027] S24: Based on the updated prototype center, repeat steps S22 and S23 until the stop condition is met, and generate the final prototype center of smooth images at different scales. The stop condition is:

[0028] ;

[0029] Among them, Indicates the clustering threshold.

[0030] Optionally, the S3 includes:

[0031] Construct a graph, specifically including:

[0032] Vertex definition: Use the scale features of the extracted smooth images and the final prototype center as vertices;

[0033] Edge definition: Connect the scale feature of the smooth image with the final prototype center to which it belongs, and the edge is defined as the distance from the scale feature of the smooth image to the final prototype center to which it belongs.

[0034] Optionally, the S5 includes:

[0035] Perform multi-scale processing on the video image to be encoded, extract features at different scales, and obtain the multi-scale features of the video image to be encoded;

[0036] Input the multi-scale features of the video image to be encoded into the trained graph neural network model to predict the prototype center and edges of the video image to be encoded.

[0037] Optionally, the S6 includes:

[0038] Traverse the edges of the video image to be encoded and find the edge with the shortest distance;

[0039] The scale corresponding to the edge as the optimal scale;

[0040] For the encoded video image, apply a Gaussian filter with a standard deviation of to generate a new smoothed image;

[0041] Use the discrete cosine transform to encode the new smoothed image.

[0042] The present invention also provides a lightweight encoding system based on neighbor consistency, which is characterized by including:

[0043] Feature extraction module: Perform Gaussian processing on the images in the video image samples to obtain multiple smoothed images, extract features from each smoothed image at different scales to obtain the scale features of each smoothed image;

[0044] Prototype center module: At different scales, initialize the prototype center, calculate the distance from the scale features of each smoothed image to the initial prototype center, divide the categories of the scale features of each smoothed image by a clustering method, and update the prototype center to generate the final prototype center;

[0045] Graph construction module: Based on the scale features of each smoothed image and the final prototype centers at different scales, define vertices, define connections, and construct a graph;

[0046] Graph neural network module: Construct a graph neural network model and perform training to obtain a trained graph neural network model;

[0047] Prediction module: Use the trained graph neural network model to predict the prototype center and connections of the video image to be encoded;

[0048] Image encoding module: Determine the optimal scale according to the connections of the video image to be encoded, generate a new smoothed image, and encode the new smoothed image.

[0049] Beneficial effects:

[0050] Based on the fact that adjacent image blocks in the video image show similarity in the high and low resolution spaces, the present invention is beneficial for extracting features at different scales; using the graph structure to select the scales with obvious features to process the image is beneficial for realizing lossless compression of the effective information of the image and improving the encoding efficiency.

[0051] The present invention uses Gaussian filters with different standard deviations for images to generate smoothed images, which is beneficial for the preliminary denoising processing of images; the constructed multi-scale space makes full use of the similarity of the features of the same type of images at different scales, extracts detailed features from images with different resolutions, and facilitates the subsequent screening of significantly different features; the clustering method is used for the extracted features to solve the prototype center, which is beneficial for dividing the features of different types of pictures; at the same time, the iterative solution method can effectively improve the accuracy of the prototype center; the graph structure can well represent the relationship between features and the prototype center, and the graph neural network model is used to screen the images processed at the scales with obvious features; moreover, the training of the graph neural network model can determine which features are appropriate and eliminate incorrect features; finally, the traditional coding method is used to encode the smoothed image with the best scale, which helps to represent richer image information using a lightweight coding method and avoid information loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 FIG. is a schematic flowchart of a lightweight coding method based on neighbor consistency provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited in any way. Any transformation or replacement based on the teachings of the present invention falls within the protection scope of the present invention.

[0054] Embodiment 1:

[0055] The lightweight coding method based on neighbor consistency, as Figure 1 shown, includes the following steps:

[0056] S1: Perform Gaussian processing on the images in the video image samples to obtain multiple smoothed images, and extract features from each smoothed image at different scales to obtain the scale features of each smoothed image;

[0057] S11: Perform Gaussian processing on the images in the video image samples to obtain multiple smoothed images, and construct a multi-scale space based on the multiple smoothed images;

[0058] S111: Apply Gaussian filters with standard deviations arranged in ascending order to the images in the video image samples. Each standard deviation corresponds to a Gaussian filter to generate multiple smoothed images:

[0059] ;

[0060] wherein, represents an image, represents a Gaussian filter with a standard deviation of and denotes the smoothed image generated by a Gaussian filter ; the standard deviation ranges from , where denotes the minimum value of the standard deviation , denotes the maximum value of the standard deviation ; the interval of the standard deviation is ;

[0061] S112: Construct a multi-scale space based on multiple smoothed images, and use the standard deviation of each smoothed image as the scale of the corresponding smoothed image;

[0062] S12: In the multi-scale space, extract features from each smoothed image at different scales to obtain the scale features of each smoothed image.

[0063] It should be noted that in the embodiments of the present invention, the scale features usually include the position, size, and direction of corners or edges.

[0064] S2: Based on the scale features obtained in step S1, use a clustering method to determine the final prototype centers at different scales;

[0065] S21: Initialize multiple prototype centers at different scales to obtain multiple initial prototype centers;

[0066] S22: Calculate the distance from the extracted scale features to the initial prototype centers:

[0067] ;

[0068] where denotes the initial prototype center at scale , denotes the smoothed image numbered in the video image samples, denotes the scale feature of the i-th smoothed image extracted at scale ; denotes the modulo operation;

[0069] Divide the scale features of each smoothed image into the category represented by the initial prototype center with the closest distance;

[0070] S23: In each category, calculate the average value of the scale features of all smoothed images, and use the average value to update the initial prototype center to which it belongs;

[0071] S24: Based on the updated prototype centers, repeat steps S22 and S23 until the following condition is met:

[0072] ;

[0073] Among them, represents the clustering threshold;

[0074] Generate the final prototype centers of smoothed images at different scales.

[0075] S3: For the video image samples, construct a graph using the scale features extracted in step S1 and the final prototype centers obtained in step S2; specifically including:

[0076] Define the extracted scale features as vertices and the final prototype centers as vertices, connect the scale features of the smoothed image with the final prototype centers to which they belong, and define the connection line as the distance from the scale feature of the smoothed image to the final prototype center to which it belongs.

[0077] S4: Construct a graph neural network model, train the graph obtained in step S3, and obtain a trained graph neural network model;

[0078] S5: Use the trained graph neural network model to predict the prototype centers and connection lines of the video image to be encoded; specifically including:

[0079] Perform multi-scale processing on the video image to be encoded, extract features at different scales, and obtain the multi-scale features of the video image to be encoded;

[0080] Input the multi-scale features of the video image to be encoded into the trained graph neural network model to predict the prototype centers and connection lines of the video image to be encoded.

[0081] S6: Determine the optimal scale according to the connection lines of the video image to be encoded, apply a Gaussian filter with a standard deviation corresponding to the optimal scale to the video image to be encoded to obtain a new smoothed image, and encode the new smoothed image with the optimal scale;

[0082] Traverse the connection lines of the video image to be encoded to find the connection line with the shortest distance;

[0083] Take the scale corresponding to the connection line as the optimal scale;

[0084] For the video image to be encoded, apply a Gaussian filter with a standard deviation of to generate a new smoothed image;

[0085] Use discrete cosine transform to encode the new smoothed image.

[0086] Embodiment 2:

[0087] The present invention also provides a lightweight encoding system based on neighbor consistency, including the following modules:

[0088] Feature extraction module: Perform Gaussian processing on the images in the video image samples to obtain multiple smoothed images, extract features from each smoothed image at different scales, and obtain the scale features of each smoothed image;

[0089] Prototype center module: At different scales, initialize the prototype center, calculate the distance from the scale features of each smoothed image to the initial prototype center, divide the categories of the scale features of each smoothed image through a clustering method, and update the prototype center to generate the final prototype center;

[0090] Graph construction module: Based on the scale features of each smoothed image and the final prototype centers at different scales, define vertices, define connections, and construct a graph;

[0091] Graph neural network module: Construct a graph neural network model and perform training to obtain a trained graph neural network model;

[0092] Prediction module: Use the trained graph neural network model to predict the prototype center and connections of the video image to be encoded;

[0093] Image encoding module: According to the connections of the video image to be encoded, determine the optimal scale, generate a new smoothed image, and encode the new smoothed image.

[0094] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the term "including", "comprising" or any other variant thereof in this article is intended to cover a non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, device, article or method including that element.

[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0096] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A lightweight coding method based on neighbor consistency, characterized in that The method includes the following steps: S1: Perform Gaussian processing on the images in the video image samples to obtain multiple smoothed images, extract features from each smoothed image at different scales to obtain the scale features of each smoothed image; S11: Perform Gaussian processing on the images in the video image samples to obtain multiple smoothed images, and construct a multi-scale space based on the multiple smoothed images; S12: In the multi-scale space, extract features from each smoothed image at different scales to obtain the scale features of each smoothed image; S2: Based on the scale features of each smoothed image, use a clustering method to generate the final prototype centers at different scales; S3: For the video image samples, construct a graph using the scale features of each smoothed image and the final prototype centers at different scales; S4: Construct a graph neural network model, train the graph obtained in step S3 to obtain a trained graph neural network model; S5: Use the trained graph neural network model to predict the prototype centers and connections of the video image to be encoded; For the video image to be encoded, perform multi-scale processing, extract features at different scales to obtain the multi-scale features of the video image to be encoded; Input the multi-scale features of the video image to be encoded into the trained graph neural network model to predict the prototype centers and connections of the video image to be encoded; S6: According to the connections of the video image to be encoded, determine the optimal scale, apply a Gaussian filter with a standard deviation corresponding to the optimal scale to process the video image to be encoded to obtain a new smoothed image, and encode the new smoothed image.

2. The lightweight encoding method based on neighbor consistency according to claim 1, wherein, The S11 includes: S111: Apply Gaussian filters with standard deviations arranged in ascending order to the images in the video image samples, each standard deviation corresponding to a Gaussian filter, to generate multiple smoothed images: ; Among them, represents an image, represents a Gaussian filter with a standard deviation of ; represents the smoothed image generated by the Gaussian filter ; The value range of the standard deviation is , and the interval of the value of the standard deviation is ; S112: Construct a multi-scale space based on the multiple smoothed images, and use the standard deviation of each smoothed image as the scale of the corresponding smoothed image.

3. The lightweight coding method based on neighbor consistency according to claim 1, characterized in that The step S2 includes: S21: Initialize multiple prototype centers at different scales to obtain multiple initial prototype centers; S22: Calculate the distances from the scale features of each smoothed image to each initial prototype center: ; Among them, represents the initial prototype center at scale ; represents the smoothed image with the number in the video image sample; represents the scale feature of the i-th smoothed image extracted at scale ; represents the modulo operation. Classify the scale features of each smoothed image into the category represented by the initial prototype center with the closest distance; S23: In each category, calculate the average value of the scale features of all smoothed images, and use the average value to update the initial prototype center to which it belongs to obtain an updated prototype center; S24: Based on the updated prototype centers, repeat steps S22 and S23 until the stop condition is met to generate the final prototype centers of the smoothed images at different scales. The stop condition is: ; Among them, represents the clustering threshold.

4. The lightweight encoding method based on neighbor consistency according to claim 3, wherein The S3 includes: Construct a graph, specifically including: Vertex definition: Use the extracted scale features of the smoothed images and the final prototype centers as vertices; Connection definition: Connect the scale features of the smoothed images with the final prototype centers to which they belong, and the connection is defined as the distance from the scale feature of the smoothed image to the final prototype center to which it belongs.

5. The lightweight encoding method based on neighbor consistency according to claim 1, characterized in that The step S6 includes: Traverse the connections of the video image to be encoded to find the connection with the shortest distance; Take the scale corresponding to the connection line as the optimal scale; For the encoded video image, apply a Gaussian filter with a standard deviation of to generate a new smoothed image; Encode the new smoothed image using the discrete cosine transform.

6. A lightweight coding system based on neighbor consistency, characterized in that Includes: Feature extraction module: Perform Gaussian processing on the images in the video image samples to obtain multiple smoothed images, extract features from each smoothed image at different scales, and obtain the scale features of each smoothed image; Prototype center module: At different scales, initialize the prototype center, calculate the distance from the scale feature of each smoothed image to the initialized prototype center, divide the categories of the scale features of each smoothed image through a clustering method, and update the prototype center to generate the final prototype center; Graph construction module: Based on the scale features of each smoothed image and the final prototype centers at different scales, define vertices, define connections, and construct a graph; Graph neural network module: Construct a graph neural network model and train it to obtain a trained graph neural network model; Prediction module: Use the trained graph neural network model to predict the prototype center and connections of the video image to be encoded; Image encoding module: Determine the optimal scale according to the connections of the video image to be encoded, generate a new smoothed image, and encode the new smoothed image; To implement the lightweight encoding method based on neighbor consistency as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Method for determining scene type of remote sensing image

    CN108154107A

  • Lightweight monocular depth estimation method and device based on self-supervised deep learning

    CN119494866A