3D video super-division method based on implicit neural representation, medium and equipment
Through the 3D video super-segment method based on implicit neural representation, the spatio-temporal information of the video sequence is obtained and feature expansion is performed, and the problem of inefficient video super-segment in the prior art is solved, and the efficient improvement of video in space and time dimensions is achieved.
Patent Information
- Application Number
- CN202510338470.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
The existing two-stage video super-scoring method is relatively low in terms of super-scoring efficiency and cannot effectively improve the spatial and temporal resolution of the video.
The 3D video super-segment method based on implicit neural representation is adopted to obtain the spatiotemporal information of the video sequence, perform feature expansion, and use the 3DINR decoding function to map and predict the RGB values under the target spatiotemporal coordinates to generate high frame rate and high resolution video sequences.
In a single-stage step, the spatial and temporal resolution of the video is simultaneously improved, the spatial and temporal correlation is effectively utilized, high-quality video is generated, and the video is infinitely amplified in the spatial and temporal dimensions.
Smart Images

Figure CN120219172A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly relates to a 3D video super-resolution method, medium and device based on implicit neural representation. Background Art
[0002] With the rapid development of digital media technology, the demand for high-definition and high-frame-rate video content is increasing day by day, especially in the scenarios of video super-resolution (VSR) and video frame interpolation (VFI). However, limited by shooting devices, storage and transmission conditions, existing video data often cannot directly meet these requirements. Therefore, how to improve the spatial resolution and temporal resolution of videos through technical means has become an important research direction in the current video processing field.
[0003] Most traditional super-resolution (SR) methods for 3D videos mainly focus on improving the spatial resolution of videos, ignoring the temporal dimension information in the video sequence, which may result in the generated super-resolution videos still appearing blurry or unnatural in visual effects. To avoid this problem, a two-stage super-resolution method is usually adopted, that is, the temporal resolution is improved by performing temporal interpolation in the first stage, and then the spatial resolution is improved by spatial super-resolution. However, the two-stage super-resolution method is relatively inefficient in terms of super-resolution efficiency.
[0004] It should be noted that the information disclosed in this background art section is only intended to increase the understanding of the overall background of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those of ordinary skill in the art. Summary of the Invention
[0005] To solve the technical problem that the existing two-stage super-resolution method is relatively inefficient in terms of super-resolution efficiency, the present invention provides a 3D video super-resolution method based on implicit neural representation. The 3D video super-resolution method based on implicit neural representation includes the following steps: Obtain a video sequence, regard the video sequence as a three-dimensional continuous feature map, and extract the spatio-temporal information of the video sequence to obtain a 3D feature vector; Perform feature expansion on the 3D feature vector to obtain an expanded 3D feature vector; Input the target spatio-temporal coordinates, and combine the expanded 3D feature vector to predict the RGB values at the target spatio-temporal coordinates to obtain a high-frame-rate and high-resolution video sequence.
[0006] Further, a 3D residual network encoder is used for extracting the spatio-temporal information.
[0007] Further, the 3D feature vector at least includes RGB values, the number of frames, and the resolution size.
[0008] Further, the feature expansion of the 3D feature vector includes the following steps: Establish a three-dimensional coordinate based on the 3D feature vector, and divide the coordinate values according to the 3D feature vector to ensure that each pixel is assigned a spatio-temporal coordinate; Concatenate the 3D feature vectors of each pixel with the 3D feature vectors of the pixels adjacent to it in the up, down, left, right, front, and back directions of the spatio-temporal coordinate to obtain the expanded 3D feature vector.
[0009] Further, the expanded 3D feature vector is expressed as:
[0010] where, is the vector concatenation operation, is the spatio-temporal coordinate, is the feature vector after expansion at coordinates. is the feature vector before expansion at coordinates, and refers to the distance of each coordinate cell. If is outside the boundary, we fill it with vector.
[0011] Further, use the 3DINR decoding function to decode and predict the RGB value at the target spatio-temporal coordinate.
[0012] Further, the prediction of the RGB value at the target spatio-temporal coordinate includes the following steps: Given the target spatio-temporal coordinate, input the expanded 3D feature vector into the 3DINR decoding function; The 3DINR decoding function maps the RGB value at the target spatio-temporal coordinate according to the target spatio-temporal coordinate and the expanded 3D feature vector.
[0013] Further, the 3DINR decoding function is expressed as:
[0014] where, is the function parameter, is the target spatio-temporal coordinate, is the coordinate predicted RGB value, is the one closest to the spatio-temporal coordinate in the subspaces of the front upper left, front upper right, front lower left, front lower right, rear upper left, rear upper right, rear lower left, and rear lower right feature vector, is the coordinates, is the coordinate and the volume between, where is the coordinate diagonal (i.e., 000 111, 100 011). is the sum of all volumes.
[0015] Furthermore, the present invention also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the 3D video super-resolution method based on implicit neural representation described in any one of the above.
[0016] Furthermore, the present invention also provides a device including at least one processor and a memory communicatively connected to the processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the processor to execute the 3D video super-resolution method based on implicit neural representation described in any one of the above.
[0017] Based on the above, for the 3D video super-resolution method, medium and device provided by the present invention, compared with the prior art, the technical solution regards a low-frame-rate and low-resolution video sequence as a three-dimensional continuous feature map, simultaneously extracts the spatial and temporal information of the video sequence in a single-stage step, uses the spatial and temporal information to map and predict the RGB values at the target spatio-temporal coordinates, generates a high-frame-rate and high-resolution video sequence, and makes more effective use of spatio-temporal correlation; and since the coordinates are continuous, it is possible to generate video sequences with arbitrary frame rates and arbitrary resolutions, thereby achieving infinite magnification of the video in the spatial and temporal dimensions. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the following description, the positional relationships of the components in the drawings, unless otherwise specified, are all based on the directions shown by the components in the drawings.
[0019] Figure 1 is a schematic flowchart of a 3D video super-resolution method based on implicit neural representation provided by an embodiment of the present invention; Figure 2 is a schematic framework diagram of a 3D video super-resolution method based on implicit neural representation provided by an embodiment of the present invention. Detailed Implementation Manner
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more. In addition, the term "comprising" and any of its variations mean "at least including".
[0022] Please refer to Figure 1 and Figure 2 , Figure 1 which is a schematic flowchart of a 3D video super-resolution method based on implicit neural representation provided by an embodiment of the present invention; Figure 2 which is a schematic framework diagram of a 3D video super-resolution method based on implicit neural representation provided by an embodiment of the present invention.
[0023] To solve the technical problem that the existing two-stage super-resolution method is relatively inefficient in terms of super-resolution efficiency, or to achieve at least one of the above advantages or other advantages, an embodiment of the present invention provides a 3D video super-resolution method based on implicit neural representation. As shown in the figure, the 3D video super-resolution method based on implicit neural representation includes the following steps: Step S10: Obtain a video sequence, regard the video sequence as a three-dimensional continuous feature map, and extract the spatio-temporal information of the video sequence to obtain a 3D feature vector. Preferably, a 3D residual network encoder is used to extract the spatio-temporal information of the video sequence.
[0024] In specific implementation, obtain or provide a video sequence with low frame rate and low resolution, and extract the frame data of the video. Arrange the extracted frame data in chronological order to form a continuous 3D matrix. Use a 3D convolutional network to obtain spatio-temporal information from the continuous 3D matrix. Determine the structure and dimension of the initial feature map according to the output of the 3D convolutional network. Perform normalization processing on the initial feature map to adjust the distribution range of the feature values. Extract deeper spatio-temporal features layer by layer through multi-layer 3D convolutional operations. Determine whether the dimension of the feature map meets the preset conditions: If it meets the conditions, perform feature fusion on the feature maps that meet the conditions to merge feature information at different levels. Generate the final 3D feature vector according to the fused feature map.
[0025] If it does not meet the conditions, return to the multi-layer 3D convolutional operation to extract deeper spatio-temporal features layer by layer again.
[0026] Specifically, regard the video sequence with low frame rate and low resolution as a three-dimensional coordinate, and since the coordinates are continuous, the video sequence can be roughly regarded as a three-dimensional continuous feature map. Use a 3D residual network encoder to extract spatio-temporal information from the video sequence, which can simultaneously extract the time resolution (frame rate) and spatial resolution (size) of the video sequence, so as to improve the time resolution and spatial resolution simultaneously in a single stage, make more effective use of spatio-temporal correlation, and generate high-quality videos. Preferably, the 3D feature vector includes at least RGB values, the number of frames, and the resolution size.
[0027] Exemplarily, assume that a video sequence with low frame rate and low resolution is provided , input it into the 3D residual network encoder to extract features from the video sequence, and obtain a 3D feature vector , where C and are the number of channels (RGB), T is the number of video frames, and H and W are the height and width of the video.
[0028] Step S20: Perform feature expansion on the 3D feature vector to obtain an expanded 3D feature vector.
[0029] Specifically, performing feature expansion on the 3D feature vector includes the following steps: S201: Establish a three-dimensional coordinate according to the 3D feature vector, and divide the coordinate values according to the 3D feature vector to ensure that each pixel is assigned a spatio-temporal coordinate; S202: Concatenate the 3D feature vectors of each pixel with the pixels adjacent to it in the up, down, left, right, front, and back directions of the spatio-temporal coordinate to obtain an expanded 3D feature vector.
[0030] In specific implementation, the starting point and ending point of the time dimension are determined according to the extracted time resolution to generate the time dimension coordinate axis. The height and width of the space dimension are determined according to the extracted spatial resolution to generate the space dimension coordinate axis. The time dimension coordinate axis and the space dimension coordinate axis are integrated to construct a three-dimensional spatio-temporal coordinate system. If the time range and space range of the target video change, the axis ranges of the three-dimensional spatio-temporal coordinate system are updated.
[0031] The coordinate information of each pixel in the feature map is determined according to the 3D feature vector, ensuring that each pixel is assigned a corresponding spatio-temporal coordinate. To enrich the information contained in each pixel, for each pixel, taking its coordinate as the center, the pixels adjacent to it in the up, down, left, right, front, and back directions are concatenated to obtain an extended 3D feature vector. If the number of pixels in the adjacent range is insufficient, pixels are supplemented by interpolation to ensure that the number of pixels reaches the preset length.
[0032] Exemplarily, according to A three-dimensional coordinate is established, and feature expansion is performed to obtain an extended 3D feature vector .
[0033] Of course, it can be understood that the core purpose of feature expansion is to expand the number of channels (RGB) in the 3D feature vector. For example, a certain pixel is 256-dimensional in RGB, and through feature expansion, it is expanded to 512 dimensions and mapped to 3 dimensions in subsequent steps.
[0034] Furthermore, the extended 3D feature vector is expressed as:
[0035] Wherein, is the vector concatenation operation, is the spatio-temporal coordinate, is the feature vector after expansion under the coordinate. is the feature vector before expansion under the coordinate, and refers to the distance of each coordinate cell. If is outside the boundary, we fill it with the vector.
[0036] Step S30: Input the target spatio-temporal coordinate, and combine the extended 3D feature vector to predict the RGB value under the target spatio-temporal coordinate to obtain a high-frame-rate and high-resolution video sequence.
[0037] In specific implementation, the 3DINR decoding function is used to decode and predict the RGB value under the target spatio-temporal coordinate. It specifically includes the following steps: S301: Given the target spatio-temporal coordinates, input the 3D feature vector after expansion into the 3D INR decoding function; S302: The 3D INR decoding function maps the RGB values at the target spatio-temporal coordinates based on the target spatio-temporal coordinates and the expanded 3D feature vector.
[0038] To generate the final high frame rate and high resolution video sequence , it is necessary to train the function parameters to achieve the mapping relationship from the 3D feature vector to the target spatio-temporal coordinates. Then the training process of the function parameters is as follows: Use the low resolution and low frame rate video sequence as the input, and extract the 3D feature vector of the video sequence through the 3D convolutional network; Expand the 3D feature vector and establish a 3D spatio-temporal coordinate system; Use the 3D INR decoding function to predict the RGB values at each spatio-temporal coordinate, compare with the high resolution video frame, calculate the loss and perform backpropagation optimization.
[0039] Further, calculating the loss and performing backpropagation optimization includes the following steps: Calculate the loss: Substitute the predicted output and gt into the loss function to calculate the loss value. This loss value reflects the performance of the current model on the given dataset, that is, the accuracy of the model prediction; Backpropagation optimization: Initialize the gradients of all parameters to zero to avoid gradient calculation errors. Starting from the output layer, calculate the gradients of the parameters of each layer layer by layer backward. In this process, the chain rule is mainly used to calculate the gradients. The chain rule allows a complex derivative to be decomposed into a series of simple derivatives multiplied. For each layer, calculate the gradient of the parameters of this layer, that is, the partial derivative of the loss function with respect to the parameters of this layer. Pass the gradient to the previous layer and continue to calculate the gradients of the parameters of the previous layer. This process continues until the input layer; Parameter update: After obtaining the gradients of the parameters of each layer, use an optimization algorithm (such as gradient descent, Adam, etc.) to update the parameters. The optimization algorithm will adjust the values of the parameters according to the magnitude and direction of the gradients to reduce the value of the loss function; Repeat the above process multiple times until the value of the loss function reaches a small value or no longer decreases significantly. In this process, the model will gradually learn the features of the data and improve the prediction accuracy.
[0040] Substitute the function parameters into the 3D INR decoding function, then the 3D INR decoding function is expressed as:
[0041] Where, is a function parameter, is the target spatio-temporal coordinate, is the coordinate predicted RGB value, is the one in the subspaces of the upper left front, upper right front, lower left front, lower right front, upper left rear, upper right rear, lower left rear, and lower right rear that is closest to the spatio-temporal coordinate nearest eigenvector, is coordinate of is the coordinate and the volume between, where is coordinate diagonal (i.e., 000 111, 100 011). is the sum of all volumes.
[0042] The 3DINR decoding function represents the mapping relationship between the expanded 3D feature vector and the RGB value at the target spatio-temporal coordinate, and can generate video sequences with arbitrary frame rates and resolutions. No matter how large the target spatio-temporal coordinate is, the RGB value at the target spatio-temporal coordinate can be predicted through the 3DINR decoding function, realizing the infinite magnification of the video in the spatial and temporal dimensions.
[0043] In some preferred embodiments, the present invention further provides a computer-readable storage medium storing computer instructions, and when the computer is executed by a processor, the 3D video super-resolution method based on implicit neural representation according to any one of the above is implemented.
[0044] In some preferred embodiments, the present invention further provides a device including at least one processor and a memory communicatively connected to the processor, where the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the processor executes the 3D video super-resolution method based on implicit neural representation according to any one of the above.
[0045] In summary, for the 3D video super-resolution method, medium, and device provided by the present invention, compared with the prior art, the present technical solution regards the low-frame-rate and low-resolution video sequence as a three-dimensional continuous feature map, simultaneously extracts the spatial and temporal information of the video sequence in a single-stage step, uses the spatial and temporal information to map and predict the RGB value at the target spatio-temporal coordinate, generates a high-frame-rate and high-resolution video sequence, and more effectively utilizes the spatio-temporal correlation; and since the coordinates are continuous, video sequences with arbitrary frame rates and resolutions can be generated, thus realizing the infinite magnification of the video in the spatial and temporal dimensions.
[0046] Although terms such as 3D feature vectors are used more frequently in this text, the possibility of using other terms is not excluded. The use of these terms is only for the convenience of describing and explaining the essence of the present invention; interpreting them as any additional restrictions is contrary to the spirit of the present invention.
[0047] In addition, those skilled in the art should understand that although there are many problems in the prior art, each embodiment or technical solution of the present invention can be improved in only one or several aspects, and it is not necessary to solve all the technical problems listed in the prior art or the background art at the same time. Those skilled in the art should understand that the content not mentioned in a claim should not be regarded as a limitation to that claim.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A 3D video super-resolution method based on implicit neural representation, characterized in that: include Acquire a video sequence, regard the video sequence as a three-dimensional continuous feature map and extract the spatiotemporal information of the video sequence to obtain a 3D feature vector; Performing feature expansion on the 3D feature vector to obtain an expanded 3D feature vector; The target space-time coordinates are input, and the RGB values under the target space-time coordinates are predicted in combination with the expanded 3D feature vector to obtain a video sequence with high frame rate and high resolution.
2. The 3D video super-resolution method based on implicit neural representation according to claim 1, characterized in that: A 3D residual network encoder is used to extract the spatiotemporal information.
3. The 3D video super-resolution method based on implicit neural representation according to claim 2, characterized in that: The 3D feature vector at least includes RGB value, frame number and resolution size.
4. The 3D video super-resolution method based on implicit neural representation according to claim 1, characterized in that: The feature expansion of the 3D feature vector comprises the following steps: Establishing three-dimensional coordinates according to the 3D feature vector, and dividing the coordinate values according to the 3D feature vector to ensure that each pixel is assigned a spatiotemporal coordinate; The 3D feature vector of each pixel is concatenated with the 3D feature vectors of the adjacent pixels in the space-time coordinates up, down, left, right, front and back to obtain an expanded 3D feature vector.
5. The 3D video super-resolution method based on implicit neural representation according to claim 4, characterized in that: The expanded 3D feature vector Expressed as: in, is a vector concatenation operation, is the space-time coordinate, It is expanded in The eigenvectors in the coordinates. is before expansion The eigenvectors in the coordinates, and Refers to the distance of each coordinate cell. Outside the boundary, we use Vector to fill.
6. The 3D video super-resolution method based on implicit neural representation according to claim 1, characterized in that: The 3DINR decoding function is used to decode and predict the RGB value of the target space-time coordinates.
7. The 3D video super-resolution method based on implicit neural representation according to claim 6, characterized in that: Predicting the RGB value of the target spatiotemporal coordinates comprises the following steps: Given the target space-time coordinates, combined with the expanded 3D feature vector, input the 3DINR decoding function; The 3DINR decoding function maps the RGB values under the target space-time coordinates according to the target space-time coordinates and the expanded 3D feature vector.
8. The 3D video super-resolution method based on implicit neural representation according to claim 7, characterized in that: The 3DINR decoding function is expressed as: in, is the function parameter, is the target space-time coordinate, is the coordinate Predicted RGB values, It is the space-time coordinate in the subspace of left front upper, right front upper, left front lower, right front lower, left back upper, right back upper, left back lower, right back lower Recent Eigenvector, yes The coordinates of is the coordinate and The volume between yes Coordinate diagonal (i.e. 000 111,100 011). is the sum of all volumes.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer is executed by a processor, the computer implements the 3D video super-resolution method based on implicit neural representation according to any one of claims 1 to 8.
10. A device, characterized in that: It includes at least one processor and a memory communicatively connected to the processor, wherein the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable the processor to perform the 3D video super-resolution method based on implicit neural representation as described in any one of claims 1-8.