Image encoding, decoding and transmission method and system for neural network

By performing three-dimensional convolutional feature and residual feature extraction, downsampling and Gaussian filtering on the image sequence, the downsampling parameters and interpolation coefficients are optimized, and the problem of insignificant information loss and features in neural network image encoding and decoding is solved, and the accuracy of efficient image encoding and decoding and recognition tasks is improved.

CN119544979BActive Publication Date: 2025-05-09CHANGSHA CHAOCHUANG ELECTRONICS TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510088669.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-09
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The image encoding and decoding method based on neural networks has problems such as loss of important information, lack of obvious image features after encoding, and unsatisfactory image recovery after decoding.

Method used

By extracting three-dimensional convolutional features and residual features from the image sequence, downsampling and Gaussian filtering are performed to construct a smooth feature map. Based on the target classification results of the smooth feature map, the downsampling parameters and interpolation coefficients are optimized, the transmission vector is constructed for transmission, and the upsampling feature map is optimized at the receiving end.

Benefits of technology

Effectively preserve the recognizable features and effective information of the image, improve the encoding and decoding efficiency, reduce encoding losses, improve the accuracy of image recognition tasks, and achieve the best image decoding effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119544979B_ABST
    Figure CN119544979B_ABST
Patent Text Reader

Abstract

The present invention discloses an image encoding, decoding and transmission method and system for neural network, the method comprising: S1: extracting residual features from an image sequence and then downsampling; S2: using a Gaussian filter to process the time and space pooling downsampling matrix obtained in step S1 to obtain a smooth feature map; S3: optimizing downsampling parameters, obtaining an updated smooth feature map and transmitting it; S4: at the receiving end, interpolating the updated smooth feature map in the time direction, and then interpolating in the height and width directions, superimposing the height and width interpolation feature map with the updated smooth feature map to obtain an upsampling feature map; S5: comparing the distribution difference between the updated smooth feature map and the upsampling feature map, optimizing the interpolation coefficient, and obtaining the optimal upsampling feature map. The present invention can retain the recognizable features and effective information of the image, improve the encoding and decoding efficiency, and achieve the best image encoding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image coding and decoding, and in particular to an image coding and decoding and transmission method and system for neural network. Background Art

[0002] Image coding and decoding refers to the process of converting an image into a digital signal and restoring the digital signal to the original image. Image coding and decoding technology is widely used in the field of image transmission and storage, which is conducive to improving the image transmission rate, saving transmission bandwidth, and reducing storage memory usage in wireless network environments. Among them, image coding technology can achieve efficient image compression, reduce the amount of calculation, and improve processing efficiency, and image decoding technology can achieve image information recovery and ensure the integrity of image information.

[0003] Neural networks are computational models that simulate the workings of the human brain's nervous system, and use a large number of interconnected neurons to achieve nonlinear mapping from input data to output data. Neural networks are widely used in many fields. For example, neural networks can recognize images based on image features. Neural networks have strong feature learning capabilities and good anti-interference performance and robustness. Using neural networks in the field of image processing is conducive to giving full play to the capabilities of neural networks, providing better image quality and higher processing efficiency. However, the current image encoding and decoding methods based on neural networks face the problems of important information loss, unclear image features after encoding, and unsatisfactory image recovery after decoding. Summary of the invention

[0004] In view of this, the present invention provides an image encoding, decoding and transmission method and system for a neural network, which optimizes the downsampling parameters and upsampling interpolation coefficients based on the classification accuracy of the neural network, which is beneficial to retaining the recognizable features and effective information of the image, improving the encoding and decoding efficiency, and achieving optimal image encoding and decoding.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] An image encoding, decoding and transmission method for a neural network comprises the following steps:

[0007] S1: extract three-dimensional convolution features from the image sequence, extract residual features based on the three-dimensional convolution features of the image sequence, and then downsample the residual features to obtain a temporal and spatial pooling downsampling matrix;

[0008] S2: Use Gaussian filter to process the temporal and spatial pooling downsampling matrix to obtain a smooth feature map;

[0009] S3: Based on the target classification result of the smooth feature map, optimize the downsampling parameters, obtain the updated smooth feature map, and construct the transmission vector for transmission;

[0010] S4: At the receiving end, the received transmission vector is restored to an updated smooth feature map, and the updated smooth feature map is interpolated in the time direction to obtain a time interpolation feature map; the time interpolation feature map is interpolated in the height and width directions to obtain a height and width interpolation feature map, and the height and width interpolation feature map is superimposed with the updated smooth feature map to obtain an upsampled feature map;

[0011] S5: Construct a metric network, compare the distribution differences between the updated smooth feature map and the upsampled feature map, optimize the interpolation coefficient, and obtain the optimal upsampled feature map.

[0012] Optionally, in step S1, extracting three-dimensional convolution features from the image sequence, extracting residual features based on the three-dimensional convolution features of the image sequence, and then downsampling the residual features to obtain a temporal and spatial pooling downsampling matrix, including:

[0013] S11: Extract 3D convolution features from image sequences:

[0014] ;

[0015] in, Indicates the extraction of three-dimensional convolution features, express An image sequence of dimensions, represents the number of time frames of the image sequence, Indicates the height of each frame image. Indicates the width of each frame image. Representing 3D convolutional features of image sequences;

[0016] S12: 3D convolutional features based on image sequences , extract residual features ;

[0017] S13: Residual features Downsampling:

[0018] ;

[0019] in, Indicates that the maximum pooling is used for downsampling. represents the temporal pooling step, represents the height pooling step, represents the width pooling step, express The temporal pooling downsampling matrix, express The temporal and height pooling downsampling matrices, express Temporal and spatial pooling downsampling matrices;

[0020] Temporal pooling step , height pooling step length , width pooling step The parameters are collectively referred to as downsampling parameters.

[0021] Optionally, in step S12, based on the three-dimensional convolution feature of the image sequence , extract residual features ,include:

[0022] S121: Through the first layer of convolution, the first intermediate feature is obtained:

[0023] ;

[0024] in, represents the convolution kernel of the first layer of convolution, represents the first intermediate feature;

[0025] Will Through the second layer of convolution, the second intermediate feature is obtained:

[0026] ;

[0027] in, represents the convolution kernel of the second layer of convolution, represents the second intermediate feature;

[0028] S122: Concatenate with the second intermediate feature and perform a nonlinear transformation:

[0029] ;

[0030] in, express and Add element by element, Indicates that the activation function is used for nonlinear transformation. represents the residual feature;

[0031] if and The dimensions do not match, so The convolution adjusts the dimension so that and Dimensions match.

[0032] Optionally, in step S2, a Gaussian filter is used to process the temporal and spatial pooling downsampling matrix to obtain a smooth feature map, including:

[0033] Perform Gaussian smoothing on the temporal and spatial pooling downsampling matrices and extract 3D convolutional features:

[0034] ;

[0035] in, express The smooth feature map obtained by extracting the three-dimensional convolution features after processing with the Gaussian filter, Represents a Gaussian smoothing filter:

[0036] ;

[0037] in, Respectively represent the time, height, and width coordinates in the time and space pooling downsampling matrix, represents a Gaussian filter that filters in time, height, and width directions, represents the coefficients of the Gaussian smoothing filter, Represents pi.

[0038] Optionally, in step S3, based on the target classification result of the smooth feature map, optimizing the downsampling parameters to obtain an updated smooth feature map, and constructing a transmission vector for transmission, including:

[0039] S31: Based on the target classification results of the smoothed feature map, construct the objective function:

[0040] ;

[0041] in, represents the objective function constructed based on the target classification result of the smoothed feature map, represents the cross entropy loss function, Indicates types of true classification labels, Indicates The network predicts the probability of the type, represents the regularization hyperparameter, which is used to balance the classification loss and the regularization term. represents the regularization term, which is used to constrain the downsampling parameters and prevent over-adjustment;

[0042] The cross entropy loss function is expressed as:

[0043] ;

[0044] in, Indicates the number of categories;

[0045] S32: Objective function Solve and optimize downsampling parameters;

[0046] S33: Based on the optimized downsampling parameters, call step S13 to update the temporal and spatial pooling downsampling matrices, and then call step S2 to obtain an updated smoothed feature map;

[0047] S34: Flatten the updated smooth feature map in row order to obtain a one-dimensional vector ; For one-dimensional vector Add redundant symbols, construct the transmission vector, and transmit:

[0048] ;

[0049] in, represents the constructed transmission vector, Represents a redundant symbol, Represents the application of Reed-Solomon encoding to a one-dimensional vector.

[0050] Optionally, in step S32, the objective function Solve and optimize downsampling parameters, including:

[0051] S321: Calculate the gradient of the objective function with respect to the downsampling parameters:

[0052] ;

[0053] in, It means to find the gradient of the objective function with respect to the time pooling step. It means to find the gradient of the objective function with respect to the height pooling step size, It means to find the gradient of the objective function with respect to the width pooling step size, It means to find partial derivative;

[0054] Collectively referred to as the gradient of the objective function with respect to the downsampling parameters;

[0055] S322: Update the downsampling parameters using the gradient descent method:

[0056] ;

[0057] in, represents the learning rate;

[0058] S323: Repeat calling step S321 and step S322 until the objective function converges to the optimal value.

[0059] Optionally, in step S4, at the receiving end, the received transmission vector is restored to an updated smooth feature map, the updated smooth feature map is interpolated in the time direction to obtain a time interpolation feature map; the time interpolation feature map is interpolated in the height and width directions to obtain a height and width interpolation feature map, and the height and width interpolation feature map is superimposed with the updated smooth feature map to obtain an up-sampled feature map, including:

[0060] For the received transmission vector Apply Reed-Solomon decoding;

[0061] If decoding is not possible, it means that there is an error in the transmission process and the received transmission vector is discarded. ;

[0062] If it can be decoded, the recovered one-dimensional vector is obtained after decoding , the restored one-dimensional vector Convert to a two-dimensional matrix to obtain an updated smooth feature map;

[0063] The updated smoothed feature map is interpolated in the time direction to obtain the time interpolation feature map:

[0064] ;

[0065] in, Represents the coordinate value of the time interpolation feature map in the time direction, Indicates that the coordinate value in the updated smooth feature map is The pixel amplitude, Indicates that the coordinate value in the updated smooth feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude of

[0066] Interpolate the time interpolation feature map in the height and width directions to obtain the height and width interpolation feature map:

[0067] ;

[0068] in, Indicates that the coordinate values ​​in the height and width interpolation feature map are The pixel amplitude, Represents the coordinate values ​​of the height and width interpolation feature map in the height direction, Represents the coordinate values ​​of the height and width interpolation feature map in the width direction, represents the interpolation coefficient, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude of

[0069] The height and width interpolated feature maps are superimposed with the updated smoothed feature map to obtain the upsampled feature map.

[0070] Optionally, in step S5, constructing a metric network, comparing the distribution difference between the updated smoothed feature map and the upsampled feature map, optimizing the interpolation coefficient, and obtaining the optimal upsampled feature map, includes:

[0071] S51: Construct a metric network, compare the distribution differences between the updated smoothed feature map and the upsampled feature map, and calculate the distribution consistency evaluation:

[0072] ;

[0073] in, Represents the metric loss, used to update the smooth feature map And upsampled feature maps Distribution consistency evaluation, Represents the updated smooth feature map in the metric network The cluster distribution of Represents the upsampled feature map in the metric network The cluster distribution of Represents the Euclidean distance between the updated smoothed feature map and the upsampled feature map in the metric network;

[0074] S52: Fine-tune interpolation coefficients:

[0075] ;

[0076] in, represents the learning rate;

[0077] Call step S4 to update the upsampled feature map ;

[0078] S53: Iterate steps S51 and S52 until the metric loss convergence;

[0079] At this time, the interpolation coefficients obtained are That is the optimized interpolation coefficient, and the updated up-sampled feature map obtained is the optimal up-sampled feature map.

[0080] The present invention also provides an image encoding, decoding and transmission system for a neural network, comprising:

[0081] Downsampling module: extracts the three-dimensional convolution features of the image sequence, extracts the residual features, downsamples the residual features, and obtains the temporal and spatial pooling downsampling matrix;

[0082] Gaussian filter module: Gaussian smoothing is used to process the temporal and spatial pooling downsampling matrix to obtain a smooth feature map;

[0083] Smooth feature map update module: constructs the objective function, solves the objective function, optimizes the downsampling parameters, updates the temporal and spatial pooling downsampling matrices, and updates the smooth feature map;

[0084] Upsampling module: interpolate the updated smooth feature map in the time direction to obtain a time interpolation feature map; interpolate the time interpolation feature map in the height and width directions to obtain a height and width interpolation feature map, and superimpose the height and width interpolation feature map with the updated smooth feature map to obtain an upsampled feature map;

[0085] Upsampling feature map optimization module: construct a metric network, calculate the distribution consistency evaluation of the updated smooth feature map and the upsampling feature map, fine-tune the interpolation coefficient, and update the upsampling feature map; iteratively process to obtain the optimized interpolation coefficient and the optimal upsampling feature map.

[0086] Beneficial effects:

[0087] The present invention performs downsampling and Gaussian filtering based on an image sequence to construct a smooth feature map, which is beneficial to eliminating irregular noise interference in the image; the downsampling parameters are optimized based on the target classification result of the smooth feature map, and the obtained updated smooth feature map fully retains the recognizable features and effective information of the image, avoids the loss of important information in the image sequence, is beneficial to improving coding efficiency, reducing coding losses, and helps to improve the accuracy of image-based recognition tasks; based on the distribution difference between the updated smooth feature map and the upsampled feature map, the upsampled interpolation coefficient is optimized, which is beneficial to not destroying important image information in the decoding stage, ensuring the consistency of features between the decoded image and the encoded image, and achieving optimal image decoding.

[0088] In the downsampling processing stage, the present invention adopts residual features to reduce the probability of gradient vanishing and gradient explosion in the neural network, and can improve the convergence performance and robustness of the neural network model training; Gaussian smoothing is performed on the time and space pooling downsampling matrix, which is conducive to reducing the influence of random interference in the downsampling matrix; the objective function constructed based on the target classification result fully considers the classification loss and the downsampling parameter constraints, and is convenient for using the gradient descent method to quickly iterate and solve the optimal downsampling parameters; in the decoding recovery stage, the pixel amplitude changes in the time direction, height and width direction are fully considered, and the upsampling feature map is obtained by interpolation; by measuring the metric loss, the optimized interpolation coefficient and the optimal upsampling feature map are obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 A flowchart of an image encoding, decoding and transmission method for a neural network provided by one embodiment of the present invention.

[0090] Figure 2 A three-dimensional convolution feature for an image sequence provided by an embodiment of the present invention , extract residual features Flowchart diagram.

[0091] Figure 3 A schematic diagram of a flow chart for comparing the distribution differences of an updated smoothed feature map and an upsampled feature map and optimizing interpolation coefficients provided in an embodiment of the present invention.

[0092] Figure 4 This is a smoothed feature map obtained by processing in step S2 of an embodiment of the present invention.

[0093] Figure 5 It is an updated smoothed feature map obtained by processing in step S3 of an embodiment of the present invention.

[0094] Figure 6 This is an upsampled feature map obtained by processing in step S4 of an embodiment of the present invention.

[0095] Figure 7 This is the optimal up-sampling feature map obtained by processing in step S5 of an embodiment of the present invention. DETAILED DESCRIPTION

[0096] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.

[0097] Embodiment 1:

[0098] A method for encoding, decoding and transmitting an image for a neural network, such as Figure 1 As shown, the following steps are included:

[0099] S1: Extract three-dimensional convolution features from the image sequence, extract residual features based on the three-dimensional convolution features of the image sequence, and then downsample the residual features to obtain the temporal and spatial pooling downsampling matrix, including:

[0100] S11: Extract 3D convolution features from image sequences:

[0101] ;

[0102] in, Indicates the extraction of three-dimensional convolution features, express An image sequence of dimensions, represents the number of time frames of the image sequence, Indicates the height of each frame image. Indicates the width of each frame image. Representing 3D convolutional features of image sequences;

[0103] S12: 3D convolutional features based on image sequences , extract residual features ,like Figure 2 As shown:

[0104] S121: Through the first layer of convolution, the first intermediate feature is obtained:

[0105] ;

[0106] in, represents the convolution kernel of the first layer of convolution, represents the first intermediate feature;

[0107] Will Through the second layer of convolution, the second intermediate feature is obtained:

[0108] ;

[0109] in, represents the convolution kernel of the second layer of convolution, represents the second intermediate feature;

[0110] S122: Concatenate with the second intermediate feature and perform a nonlinear transformation:

[0111] ;

[0112] in, express and Add element by element, Indicates that the activation function is used for nonlinear transformation. represents the residual feature;

[0113] if and The dimensions do not match, so The convolution adjusts the dimension so that and Dimension matching;

[0114] S13: Residual features Downsampling:

[0115] ;

[0116] in, Indicates that the maximum pooling is used for downsampling. represents the temporal pooling step, represents the height pooling step, represents the width pooling step, express The temporal pooling downsampling matrix, express The temporal and height pooling downsampling matrices, express Temporal and spatial pooling downsampling matrices;

[0117] Temporal pooling step , height pooling step length , width pooling step The parameters are collectively referred to as downsampling parameters.

[0118] S2: Use Gaussian filter to process the temporal and spatial pooling downsampling matrix to obtain a smooth feature map, including:

[0119] Perform Gaussian smoothing on the temporal and spatial pooling downsampling matrices and extract 3D convolutional features:

[0120] ;

[0121] in, express The smooth feature map obtained by extracting the three-dimensional convolution features after processing with the Gaussian filter, Represents a Gaussian smoothing filter:

[0122] ;

[0123] in, Respectively represent the time, height, and width coordinates in the time and space pooling downsampling matrix, represents a Gaussian filter that filters in time, height, and width directions, represents the coefficients of the Gaussian smoothing filter, Represents pi.

[0124] In the embodiment of the present invention, Figure 4 The figure shows the smoothed feature map obtained by step S2. The outline of the target is visible in the smoothed feature map. Due to the downsampling process in step S1, the resolution of the smoothed feature map is low and the medium-detail features are not visible.

[0125] S3: Based on the target classification result of the smooth feature map, optimize the downsampling parameters, obtain the updated smooth feature map, and construct the transmission vector for transmission, including;

[0126] S31: Based on the target classification results of the smoothed feature map, construct the objective function:

[0127] ;

[0128] in, represents the objective function constructed based on the target classification result of the smoothed feature map, represents the cross entropy loss function, Indicates types of true classification labels, Indicates The network prediction probability of each type is represents the regularization hyperparameter, which is used to balance the classification loss and the regularization term. represents the regularization term, which is used to constrain the downsampling parameters and prevent over-adjustment;

[0129] The cross entropy loss function is expressed as:

[0130] ;

[0131] in, Indicates the number of categories;

[0132] S32: Objective function Solve and optimize the downsampling parameters:

[0133] S321: Calculate the gradient of the objective function with respect to the downsampling parameters:

[0134] ;

[0135] in, It means to find the gradient of the objective function with respect to the time pooling step. It means to find the gradient of the objective function with respect to the height pooling step size, It means to find the gradient of the objective function with respect to the width pooling step size, It means to find partial derivative;

[0136] Collectively referred to as the gradient of the objective function with respect to the downsampling parameters;

[0137] S322: Update the downsampling parameters using the gradient descent method:

[0138] ;

[0139] in, represents the learning rate;

[0140] S323: Repeat step S321 and step S322 until the objective function converges to an optimal value;

[0141] S33: Based on the optimized downsampling parameters, call step S13 to update the temporal and spatial pooling downsampling matrices, and then call step S2 to obtain an updated smoothed feature map;

[0142] S34: Flatten the updated smooth feature map in row order to obtain a one-dimensional vector ; For one-dimensional vector Add redundant symbols, construct the transmission vector, and transmit:

[0143] ;

[0144] in, represents the constructed transmission vector, Represents a redundant symbol, Represents the application of Reed-Solomon encoding to a one-dimensional vector.

[0145] In the embodiment of the present invention, Figure 5 The figure shows the updated smooth feature map obtained by step S3. Compared with the smooth feature map obtained by step S2, the updated smooth feature map has improved clarity, indicating that optimizing the downsampling parameters is more conducive to identifying the target category based on the updated smooth feature map.

[0146] S4: At the receiving end, the received transmission vector is restored to an updated smooth feature map, the updated smooth feature map is interpolated in the time direction to obtain a time interpolation feature map; the time interpolation feature map is interpolated in the height and width directions to obtain a height and width interpolation feature map, and the height and width interpolation feature map is superimposed with the updated smooth feature map to obtain an up-sampled feature map, including:

[0147] For the received transmission vector Apply Reed-Solomon decoding;

[0148] If decoding is not possible, it means that there is an error in the transmission process and the received transmission vector is discarded. ;

[0149] If it can be decoded, the recovered one-dimensional vector is obtained after decoding , the restored one-dimensional vector Convert to a two-dimensional matrix to obtain an updated smooth feature map;

[0150] The updated smoothed feature map is interpolated in the time direction to obtain the time interpolation feature map:

[0151] ;

[0152] in, Represents the coordinate value of the time interpolation feature map in the time direction, Indicates that the coordinate value in the updated smooth feature map is The pixel amplitude, Indicates that the coordinate value in the updated smooth feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude of

[0153] Interpolate the time interpolation feature map in the height and width directions to obtain the height and width interpolation feature map:

[0154] ;

[0155] in, Indicates that the coordinate values ​​in the height and width interpolation feature map are The pixel amplitude, Represents the coordinate values ​​of the height and width interpolation feature map in the height direction, Represents the coordinate values ​​of the height and width interpolation feature map in the width direction, represents the interpolation coefficient, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is = The pixel amplitude of

[0156] The height and width interpolated feature maps are superimposed with the updated smoothed feature map to obtain the upsampled feature map.

[0157] In the embodiment of the present invention, Figure 6 The figure shows the up-sampled feature map obtained by step S4. After up-sampling, the image contrast is improved and the granularity is obvious.

[0158] S5: Construct a metric network, compare the distribution differences between the updated smooth feature map and the upsampled feature map, optimize the interpolation coefficient, and obtain the optimal upsampled feature map, such as Figure 3 As shown, including:

[0159] S51: Construct a metric network, compare the distribution differences between the updated smoothed feature map and the upsampled feature map, and calculate the distribution consistency evaluation:

[0160] ;

[0161] in, Represents the metric loss, used to update the smooth feature map And upsampled feature maps Distribution consistency evaluation, Represents the updated smooth feature map in the metric network The cluster distribution of Represents the upsampled feature map in the metric network The cluster distribution of Represents the Euclidean distance between the updated smoothed feature map and the upsampled feature map in the metric network;

[0162] S52: Fine-tune interpolation coefficients:

[0163] ;

[0164] in, represents the learning rate;

[0165] Call step S4 to update the upsampled feature map ;

[0166] S53: Iterate steps S51 and S52 until the metric loss convergence;

[0167] At this time, the interpolation coefficients obtained are That is the optimized interpolation coefficient, and the updated up-sampled feature map obtained is the optimal up-sampled feature map.

[0168] In the embodiment of the present invention, Figure 7 The figure shows the optimal up-sampling feature map obtained by step S5. Compared with the up-sampling feature map obtained by step S4, the optimal up-sampling feature map has less obvious granularity, indicating that the optimized interpolation coefficient makes the amplitude change of the up-sampling feature map smoother.

[0169] In the embodiment of the present invention, after extracting residual features from the image sequence in step S1, temporal and spatial pooling downsampling is performed to obtain a downsampling matrix, and the obtained temporal and spatial pooling downsampling matrix removes redundant information in the image sequence and only retains important feature information; after Gaussian filtering is performed on the temporal and spatial pooling downsampling matrix in step S2, irregular noise information in the image sequence is removed, noise interference is avoided, and the target feature is more prominent; step S3 optimizes the downsampling parameters guided by the target classification result, and the obtained updated smooth feature map further highlights the target feature, removes irrelevant information in the image sequence, and reduces the required bandwidth during information transmission; step S4 interpolates the updated smooth feature map in the time, height and width directions to restore detail information in a direction to obtain an upsampling feature map; step S5 uses the distribution consistency of the updated smooth feature map and the upsampling feature map as an evaluation criterion to optimize the interpolation coefficient to obtain an upsampling feature map consistent with the downsampling feature map, that is, in the upsampling refinement process, the characteristic of the image is not changed; step S5 obtains the optimal upsampling feature map through iterative steps for user use.

[0170] Embodiment 2: The present invention also provides an image encoding, decoding and transmission system for a neural network, comprising the following five modules:

[0171] Downsampling module: extracts the three-dimensional convolution features of the image sequence, extracts the residual features, downsamples the residual features, and obtains the temporal and spatial pooling downsampling matrix;

[0172] Gaussian filter module: Gaussian smoothing is used to process the temporal and spatial pooling downsampling matrix to obtain a smooth feature map;

[0173] Smooth feature map update module: constructs the objective function, solves the objective function, optimizes the downsampling parameters, updates the temporal and spatial pooling downsampling matrices, and updates the smooth feature map;

[0174] Upsampling module: interpolate the updated smooth feature map in the time direction to obtain a time interpolation feature map; interpolate the time interpolation feature map in the height and width directions to obtain a height and width interpolation feature map, and superimpose the height and width interpolation feature map with the updated smooth feature map to obtain an upsampled feature map;

[0175] Upsampling feature map optimization module: construct a metric network, calculate the distribution consistency evaluation of the updated smooth feature map and the upsampling feature map, fine-tune the interpolation coefficient, and update the upsampling feature map; iteratively process to obtain the optimized interpolation coefficient and the optimal upsampling feature map.

[0176] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variants thereof in this article are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0177] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0178] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. An image encoding, decoding and transmission method for a neural network, characterized in that: The method comprises: S1: extracting three-dimensional convolution features from the image sequence, extracting residual features based on the three-dimensional convolution features of the image sequence, and then downsampling the residual features to obtain a temporal and spatial pooling downsampling matrix; including: S11: Extract 3D convolution features from image sequences: ; in, Indicates the extraction of three-dimensional convolution features, express An image sequence of dimensions, represents the number of time frames of the image sequence, Indicates the height of each frame image. Indicates the width of each frame image. Representing 3D convolutional features of image sequences; S12: 3D convolutional features based on image sequences , extract residual features ; S13: Residual features Downsampling: ; in, Indicates that the maximum pooling is used for downsampling. represents the temporal pooling step, represents the height pooling step, represents the width pooling step, express The temporal pooling downsampling matrix, express The temporal and height pooling downsampling matrices, express Temporal and spatial pooling downsampling matrices; Temporal pooling step , height pooling step , width pooling step The parameters are collectively referred to as downsampling parameters; S2: Use Gaussian filter to process the temporal and spatial pooling downsampling matrix to obtain a smooth feature map; S3: Based on the target classification result of the smooth feature map, optimize the downsampling parameters, obtain the updated smooth feature map, and construct the transmission vector for transmission; including: S31: Based on the target classification results of the smoothed feature map, construct the objective function: ; in, represents the objective function constructed based on the target classification result of the smoothed feature map, represents the cross entropy loss function, Indicates types of true classification labels, Indicates The network prediction probability of each type is represents the regularization hyperparameter, which is used to balance the classification loss and the regularization term. represents the regularization term, which is used to constrain the downsampling parameters and prevent over-adjustment; The cross entropy loss function is expressed as: ; in, Indicates the number of categories; S32: For the objective function Solve and optimize downsampling parameters; S33: Based on the optimized downsampling parameters, call step S13 to update the temporal and spatial pooling downsampling matrices, and then call step S2 to obtain an updated smoothed feature map; S34: Flatten the updated smooth feature map in row order to obtain a one-dimensional vector ; For one-dimensional vector Add redundant symbols, construct the transmission vector, and transmit: ; in, represents the constructed transmission vector, Represents a redundant symbol, Reed-Solomon encoding is applied to a one-dimensional vector. S4: At the receiving end, the received transmission vector is restored to an updated smooth feature map, and the updated smooth feature map is interpolated in the time direction to obtain a time interpolation feature map; the time interpolation feature map is interpolated in the height and width directions to obtain a height and width interpolation feature map, and the height and width interpolation feature map is superimposed with the updated smooth feature map to obtain an upsampled feature map; S5: Construct a metric network, compare the distribution differences between the updated smooth feature map and the upsampled feature map, optimize the interpolation coefficient, and obtain the optimal upsampled feature map.

2. The image encoding, decoding and transmission method for a neural network according to claim 1, characterized in that: The step S12 comprises: S121: Through the first layer of convolution, the first intermediate feature is obtained: ; in, represents the convolution kernel of the first layer of convolution, represents the first intermediate feature; Will Through the second layer of convolution, the second intermediate feature is obtained: ; in, represents the convolution kernel of the second layer of convolution, represents the second intermediate feature; S122: Concatenate with the second intermediate feature and perform a nonlinear transformation: ; in, express and Add element by element, Indicates that the activation function is used for nonlinear transformation. Represents the residual feature.

3. The image encoding, decoding and transmission method for a neural network according to claim 1, characterized in that: The step S2 comprises: Perform Gaussian smoothing on the temporal and spatial pooling downsampling matrices and extract 3D convolutional features: ; in, express The smooth feature map obtained by extracting the three-dimensional convolution features after processing with the Gaussian filter, Represents a Gaussian smoothing filter: ; in, Respectively represent the time, height, and width coordinates in the time and space pooling downsampling matrix, represents a Gaussian filter that filters in time, height, and width directions, represents the coefficients of the Gaussian smoothing filter, Represents pi.

4. The image encoding, decoding and transmission method for a neural network according to claim 1, characterized in that: The step S32 comprises: S321: Calculate the gradient of the objective function with respect to the downsampling parameters: ; in, It means to find the gradient of the objective function with respect to the time pooling step. It means to find the gradient of the objective function with respect to the height pooling step size, It means to find the gradient of the objective function with respect to the width pooling step size, It means to find partial derivative; Collectively referred to as the gradient of the objective function with respect to the downsampling parameters; S322: Update the downsampling parameters using the gradient descent method: ; in, represents the learning rate; S323: Repeat calling step S321 and step S322 until the objective function converges to the optimal value.

5. The image encoding, decoding and transmission method for a neural network according to claim 1, characterized in that: The step S4 comprises: For the received transmission vector Apply Reed-Solomon decoding; If decoding is not possible, it means that there is an error in the transmission process and the received transmission vector is discarded. ; If it can be decoded, the recovered one-dimensional vector is obtained after decoding , the restored one-dimensional vector Convert to a two-dimensional matrix to obtain an updated smooth feature map; The updated smoothed feature map is interpolated in the time direction to obtain the time interpolation feature map: ; in, Represents the coordinate value of the time interpolation feature map in the time direction, Indicates that the coordinate value in the updated smooth feature map is The pixel amplitude, Indicates that the coordinate value in the updated smooth feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude of Interpolate the time interpolation feature map in the height and width directions to obtain the height and width interpolation feature map: ; in, Indicates that the coordinate values ​​in the height and width interpolation feature map are The pixel amplitude, Represents the coordinate values ​​of the height and width interpolation feature map in the height direction, Represents the coordinate values ​​of the height and width interpolation feature map in the width direction, represents the interpolation coefficient, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude, Indicates that the coordinate value in the time interpolation feature map is The pixel amplitude of The height and width interpolated feature maps are superimposed with the updated smoothed feature map to obtain the upsampled feature map.

6. The image encoding, decoding and transmission method for a neural network according to claim 5, characterized in that: The step S5 comprises: S51: Construct a metric network, compare the distribution differences between the updated smoothed feature map and the upsampled feature map, and calculate the distribution consistency evaluation: ; in, Represents the metric loss, used to update the smooth feature map And upsampled feature maps Distribution consistency evaluation, Represents the updated smooth feature map in the metric network The cluster distribution of Represents the upsampled feature map in the metric network The cluster distribution of Represents the updated smooth feature map in the metric network And upsampled feature maps The Euclidean distance of S52: Fine-tune interpolation coefficients: ; in, represents the learning rate; Call step S4 to update the upsampled feature map ; S53: Iterate steps S51 and S52 until the metric loss Convergence; at this point, the interpolation coefficients obtained That is the optimized interpolation coefficient, and the updated up-sampled feature map obtained is the optimal up-sampled feature map.

7. An image encoding, decoding and transmission system for a neural network, characterized in that: include: Downsampling module: extracts the three-dimensional convolution features of the image sequence, extracts the residual features, downsamples the residual features, and obtains the temporal and spatial pooling downsampling matrix; Gaussian filter module: Gaussian smoothing is used to process the temporal and spatial pooling downsampling matrix to obtain a smooth feature map; Smooth feature map update module: constructs the objective function, solves the objective function, optimizes the downsampling parameters, updates the temporal and spatial pooling downsampling matrices, and updates the smooth feature map; Upsampling module: interpolate the updated smooth feature map in the time direction to obtain a time interpolation feature map; interpolate the time interpolation feature map in the height and width directions to obtain a height and width interpolation feature map, and superimpose the height and width interpolation feature map with the updated smooth feature map to obtain an upsampled feature map; Upsampling feature map optimization module: construct a metric network, calculate the distribution consistency evaluation of the updated smooth feature map and the upsampling feature map, fine-tune the interpolation coefficient, and update the upsampling feature map; iterate the process to obtain the optimized interpolation coefficient and the optimal upsampling feature map; To implement an image encoding, decoding and transmission method for a neural network as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Video compression method and electronic equipment

    CN115941966A

  • Method for solving pseudo-edge and checkerboard problems based on MBLLEN improved enhanced network

    CN117291794A

  • Laser radar data processing method and system

    CN119087392A