Image processing method, apparatus, device, medium and product

CN122265056BActive Publication Date: 2026-09-04MOORE THREADS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610729628.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-09-04
Estimated Expiration
2046-05-26

AI Technical Summary

Technical Problem

[0004]然而,现有TAA方法在确定历史帧与当前帧的融合权重时,通常依赖运动矢量、邻域颜色方差等物理特征

Benefits of technology

通过基于光流信息对第i-1输出图像进行变形处理,实现历史图像与当前图像在空间上的对齐。在此基础上,根据第i-1变形图像与第i抖动图像的图像特征之间的特征相关度确定融合权重。由于图像特征能够从更丰富的维度表征图像内容,基于特征相关度确定的融合权重具有更强的代表性。当第i-1变形图像中存在伪影时,伪影区域的特征相关度较低,融合权重可以相应减小第i-1变形图像的融合比例,侧重于融合第i抖动图像,从而使得生成的第i输出图像能够有效去除伪影。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265056B_ABST
    Figure CN122265056B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method and device, equipment, medium and product, and relates to the technical field of images. The method comprises the following steps: performing deformation processing on an i-1 output image based on optical flow information to obtain an i-1 deformed image, the i-1 output image being an image displayed at an i-1 moment, i being a positive integer; determining a fusion weight based on the feature correlation between the image features of the i-1 deformed image and the image features of an i jitter image, the i jitter image being an image obtained by jitter sampling an original image at an i moment; and performing image fusion on the i-1 deformed image and the i jitter image based on the fusion weight to obtain an i output image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image technology, and in particular to an image processing method, apparatus, device, medium and product. Background Technology

[0002] Jagged edges refer to the unevenness of image edges, appearing as stepped or pixelated patterns, caused by insufficient sampling rate during image rendering or sampling. Anti-aliasing technology aims to eliminate or reduce these jagged edges, thereby improving image quality.

[0003] In related technologies, Temporal Anti-Aliasing (TAA) is commonly used. It continuously fuses accumulated frame information, combines historical frames with the current frame, and uses historical information to smooth the jagged areas, thereby effectively reducing the jagged effect.

[0004] However, existing TAA methods typically rely on physical features such as motion vectors and neighborhood color variance when determining the fusion weights between historical and current frames. When faced with image content whose complexity exceeds the preset rules, this can easily lead to weight judgment biases, resulting in artifacts in the fusion result. Summary of the Invention

[0005] This application provides an image processing method, apparatus, device, medium, and product. The technical solutions provided by this application include the following aspects.

[0006] On one hand, embodiments of this application provide an image processing method, the method comprising: Based on optical flow information, the (i-1)th output image is deformed to obtain the (i-1)th deformed image. The (i-1)th output image is the image rendered and displayed at the (i-1)th time, where i is a positive integer. Based on the feature correlation between the image features of the (i-1)th deformed image and the image features of the i-th jittered image, the fusion weight is determined. The i-th jittered image is obtained by jitter sampling of the original image at time i. Based on the fusion weights, the (i-1)th deformed image and the ith jittery image are fused to obtain the ith output image.

[0007] On the other hand, embodiments of this application provide an image processing apparatus, the apparatus comprising: The deformation module is configured to perform deformation processing on the (i-1)th output image based on optical flow information to obtain the (i-1)th deformed image, wherein the (i-1)th output image is the image rendered and displayed at the (i-1)th time, and i is a positive integer. The determination module is configured to determine the fusion weight based on the feature correlation between the image features of the (i-1)th deformed image and the image features of the i-th jittered image, wherein the i-th jittered image is obtained by jitter sampling of the original image at time i. The fusion module is configured to perform image fusion on the (i-1)th deformed image and the i-th jittery image based on the fusion weights to obtain the i-th output image.

[0008] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described image processing method.

[0009] On the other hand, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the above-described image processing method.

[0010] On the other hand, embodiments of this application provide a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor reads from the computer-readable storage medium and executes the computer program to implement the above-described image processing method.

[0011] The technical solution provided in this application can bring the following beneficial effects: By deforming the (i-1)th output image based on optical flow information, spatial alignment between the historical image and the current image is achieved. Furthermore, fusion weights are determined based on the feature correlation between the (i-1)th deformed image and the i-th jittered image. Since image features can represent image content from richer dimensions, fusion weights determined based on feature correlation are more representative. When artifacts exist in the (i-1)th deformed image, the feature correlation of the artifact region is low, and the fusion weight can correspondingly reduce the fusion ratio of the (i-1)th deformed image, focusing on fusing the i-th jittered image, thereby enabling the generated i-th output image to effectively remove artifacts. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of a jagged image provided in an exemplary embodiment of this application; Figure 2 This is a flowchart illustrating an embodiment of an image processing method provided in this application; Figure 3 This is a schematic diagram of a deformed image and the original image provided in an illustrative embodiment of this application; Figure 4 This is a schematic diagram of jitter sampling provided in an illustrative embodiment of this application; Figure 5 This is a schematic diagram of a deformed image and a jittered image provided in an illustrative embodiment of this application; Figure 6 This is a flowchart illustrating another embodiment of the image processing method provided in this application; Figure 7 This is a model structure diagram of an image processing model provided in an illustrative embodiment of this application; Figure 8 This is a flowchart illustrating the image processing model training process provided in one embodiment of this application; Figure 9 This is a model structure diagram of an image processing model provided in another illustrative embodiment of this application; Figure 10 This is a schematic diagram of a single-step training phase and a sequential training phase provided in an illustrative embodiment of this application; Figure 11 This is a schematic diagram of the structure of an image processing apparatus provided in an exemplary embodiment of this application; Figure 12 This is a structural block diagram of a computer device provided in an illustrative embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0014] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0015] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0016] It should be understood that although the terms “first,” “second,” etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word “if” as used herein may be interpreted as “when…” or “in response to determination.”

[0017] It should be noted that this application may display a prompt interface, pop-up window, or output voice prompt information before collecting user, processor, computer device, and other related data, and during the process of collecting user-related data. This prompt interface, pop-up window, or voice prompt information is used to inform the user that their related data is being collected. This ensures that the application only begins executing the steps related to collecting user-related data after receiving confirmation from the user regarding the prompt interface or pop-up window; otherwise (i.e., without receiving confirmation from the user), the steps related to collecting user-related data end, meaning no user-related data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the laws, regulations, and standards of relevant countries and regions.

[0018] First, the terms used in the embodiments of this application will be introduced.

[0019] Jagged edges: In the field of image rendering, jagged edges refer to the visual phenomenon where, during the sampling process, the sampling rate is insufficient to smoothly reproduce continuous image signals, causing what should be smooth straight or curved edges in an image to appear as stepped or pixelated blocks. The root cause of jagged edges is that real-world image signals are spatially continuous, while the sampling and display process of digital images is based on a discrete pixel grid. When a continuous signal is sampled through a discrete grid, the sampling points cannot completely record all the details of the continuous signal, especially at image edges where signal intensity changes drastically. Discrete sampling cannot accurately reproduce this continuous transition, thus visually forming stepped, distorted edges, i.e., jagged edges.

[0020] like Figure 1The diagram illustrates a jagged image provided in an illustrative embodiment of this application. The image shows a triangular image being sampled and rendered. Because the edge lines of the triangle extend continuously diagonally in space, while the sampling points are distributed in a discrete grid pattern, the sampling points cannot completely cover all the continuous positions traversed by the edge lines. Within the grid cells traversed by the edge lines, some sampling points are located inside the triangle, and some are located outside the triangle, resulting in the pixels traversed by the edge lines being discretized and marked as belonging to or not belonging to the triangle. Thus, the originally smooth diagonal edge, under the effect of discrete sampling, presents a stepped pixel block arrangement, forming a visually jagged effect.

[0021] Anti-aliasing technology refers to techniques used to eliminate or reduce jagged edges in images. Its core idea is to smooth image edges or optimize the sampling process to approximate the ideal effect of continuous signals. For example, TAA (Transient Aliasing) technology fuses pixel information through multi-frame accumulation, effectively eliminating jagged edges and flickering in dynamic scenes; Super Sampling Anti-Aliasing (SSAA) technology renders the image at a higher resolution during the rendering stage and then scales it down to the target resolution, effectively increasing the sampling rate and reducing jagged edges. All of these techniques can make image edge transitions more natural, reduce stair-step visual distortion, and thus improve image quality.

[0022] A jitter frame is an image frame generated during image rendering by applying small, regular spatial offsets to the camera viewpoint, projection matrix, or sampling grid. This offset is typically at the sub-pixel level, aiming to stagger the sampling positions between different frames, thereby capturing richer sub-pixel information over time. In subsequent image fusion or accumulation processing, fusing jitter frames with historical frames can effectively improve the equivalent sampling rate and eliminate or reduce aliasing. Jitter frames are also known as jitter-sampled images or jitter images.

[0023] Artifacts refer to non-realistic distorted patterns or noise that appear in images during image processing, compression, or reconstruction due to algorithmic defects, insufficient sampling, information loss, or computational errors. Artifacts manifest in various forms, including ghosting, trailing shadows, flickering, and abnormal textures. Ghosting is typically caused by improper multi-frame fusion resulting in historical image retention; trailing shadows are caused by trajectory blurring due to inaccurate motion compensation; and flickering originates from brightness or color abrupt changes due to insufficient inter-frame consistency. The presence of artifacts degrades image visual quality and interferes with information recognition; therefore, artifact suppression is a crucial process in image processing.

[0024] Ajazz refers to the unevenness of image edges, appearing as stepped or pixelated patterns, caused by insufficient sampling rate during image rendering or sampling. Anti-aliasing techniques aim to eliminate or mitigate this jaggedness, improving image quality. Related technologies typically employ the Transform-Aliasing (TAA) method, which continuously fuses accumulated frame information, combining historical frames with the current frame. Historical information is used to smooth jagged areas, effectively reducing the jagged effect. However, existing TAA methods often rely on physical features such as motion vectors and neighborhood color variance when determining the fusion weights between historical and current frames. When dealing with image content whose complexity exceeds preset rules, weight judgment biases can easily occur, leading to artifacts in the fusion result.

[0025] To address the aforementioned issues, this application proposes an image processing method that aligns the historical image with the current image spatially by deforming the (i-1)th output image based on optical flow information. Furthermore, fusion weights are determined according to the feature correlation between the (i-1)th deformed image and the i-th jittered image. Since image features can represent image content from richer dimensions, fusion weights determined based on feature correlation are more representative. When artifacts exist in the (i-1)th deformed image, the feature correlation of the artifact region is low, and the fusion weight can correspondingly reduce the fusion ratio of the (i-1)th deformed image, focusing on fusing the i-th jittered image, thereby effectively removing artifacts from the generated i-th output image.

[0026] The solution provided in this application can be used in computer devices with image processing needs, which can be terminals or servers. Terminals can be electronic devices such as mobile phones, tablets, in-vehicle terminals (vehicle infotainment systems), wearable devices, and PCs (Personal Computers). Servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud servers, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In the following embodiments, for ease of description, the image processing method is illustrated by an example of execution by a computer device.

[0027] like Figure 2 As shown, a flowchart of an image processing method provided by an exemplary embodiment of this application is illustrated. The method can be executed by the aforementioned computer device and includes at least one of the following steps.

[0028] Step 210: Based on optical flow information, deform the (i-1)th output image to obtain the (i-1)th deformed image. The (i-1)th output image is the image rendered and displayed at the (i-1)th time, where i is a positive integer.

[0029] In some embodiments, a computer device performs image processing on a series of original images sequentially. The images in the sequence are ordered chronologically, meaning the order of the images refers to their temporal sequence rather than their numerical arrangement within the image sequence, and the image content of adjacent sequences is related in chronological order. For ease of explanation, an example is described using time counting from 0 to the i-th time, where i is a positive integer.

[0030] For illustration purposes, the i-th output image corresponds to the image output by the computer device at time i, rather than the i-th image in the output order.

[0031] Here, the original image refers to the rendered image generated through rendering operations that has not yet undergone anti-aliasing processing. When processing the original image at time i, combining historical image data with the original image data at the current time can effectively improve the anti-aliasing effect. Therefore, the computer device needs to acquire the historical image data corresponding to time i.

[0032] In some embodiments, the computer device acquires the aforementioned historical image data based on the (i-1)th output image. The (i-1)th output image refers to the image rendered, anti-aliased, and finally displayed by the computer device at time (i-1). By using the image data of the (i-1)th output image as historical image data and performing image fusion processing with the image data of the original image at time (i), the image quality of the current frame can be improved.

[0033] However, if the (i-1)th output image is directly fused with the original image data at time i, significant content differences will arise between the two times due to potential scene motion or viewpoint changes. To compensate for these content differences, the computer device needs to perform motion compensation on the (i-1)th output image using optical flow information.

[0034] In some embodiments, the computer device performs warping processing on the (i-1)th output image based on optical flow information to obtain the (i-1)th warped image.

[0035] Indicatively, the (i-1)th deformed image is the image obtained by performing motion compensation on the (i-1)th output image based on the motion trajectory described by the optical flow information. It is used to characterize the estimated image corresponding to the current original image, which is inferred from the displayed image at the previous moment, considering the image motion, thereby achieving spatial alignment between the historical image and the image at the current moment.

[0036] Step 220: Determine the fusion weight based on the feature correlation between the image features of the (i-1)th deformed image and the image features of the ith jittered image. The ith jittered image is obtained by jittering the original image at time i.

[0037] After obtaining the (i-1)th deformed image used as historical image data, the computer device needs to further determine how to fuse them. If the original image at time i is directly fused with the (i-1)th deformed image, since the (i-1)th deformed image is an estimated image corresponding to the original image at time i, which is inferred from historical frames based on optical flow information, although there is a motion compensation relationship between the two in terms of content, they tend to be consistent in terms of sampling space. The obtained image content data is relatively similar, and the effect after fusion processing is limited, making it difficult to fully exert the anti-aliasing effect.

[0038] As an illustration, the image information acquired at the same or similar sampling positions in the original image at time i and the deformed image at time i-1 are highly repetitive, making it difficult to provide complementary sub-pixel information. Therefore, directly fusing the two cannot effectively expand the diversity of the sampling samples, resulting in limited smoothing ability of the fusion process for jagged edges, and making it difficult to fully realize the intended effect of anti-aliasing technology.

[0039] like Figure 3 The figure illustrates a schematic diagram of a deformed image and an original image provided in an illustrative embodiment of this application. As can be seen, the deformed image and the original image are quite similar in content, and their sampling point distributions overlap significantly. Due to the lack of pixel-level offset introduction, the amount of new information that can be accumulated during the fusion process is limited, making it difficult for the sampling points to collect effective complementary data, thus limiting the effectiveness of the anti-aliasing effect.

[0040] In some embodiments, the computer device uses the i-th jittered image as the current image data. The i-th jittered image refers to the original image rendered at time i, in which pixel offsets are introduced through jitter sampling, allowing the jittered image to collect different sampled data than the original image, thereby obtaining richer image information. When the i-th jittered image is fused with the (i-1)-th deformed image, this pixel offset helps to collect effective complementary data, achieving a higher anti-aliasing effect during image fusion.

[0041] Optionally, the i-th jittered image refers to the image generated by sampling after applying a small translational jitter to the camera viewpoint or projection matrix when the computer device renders the original image at time i. This application does not impose limitations on this.

[0042] like Figure 4 As shown, it illustrates a schematic diagram of jitter sampling provided in an illustrative embodiment of this application. As shown, image 401 is the original image at time i, and image 402 is the jittered image at time i. The jittered image at time i is an image generated by the computer device by applying translation jitter to the camera viewpoint when rendering the original image at time i. Compared with the original image at time i, the pattern content in the jittered image at time i is shifted to the right.

[0043] Combining the jittered image as the current image with the deformed image allows us to obtain more image information. For example... Figure 5 The diagram illustrates a distorted image and a jittered image provided in an illustrative embodiment of this application. The left image is the (i-1)th distorted image, and the right image is the i-th jittered image. Compared to the sampling point distribution in the (i-1)th distorted image, the i-th jittered image, due to the introduction of pixel-level jitter offset, exhibits a different spatial distribution of sampling points. This distribution allows the i-th jittered image to capture sampling location information that the (i-1)th distorted image fails to cover within the same spatial area. When the two are fused, the complementary sampling data provided by the i-th jittered image can effectively fill the gaps in the sampling grid of the (i-1)th distorted image, thereby achieving more accurate fusion of image edges in the fusion result and improving the anti-aliasing effect.

[0044] After determining the (i-1)th deformed image used as historical image data and the i-th jittery image used as current image data, the computer device needs to determine the weights that each should occupy during fusion. The appropriateness of the fusion weights directly affects the jagged edge smoothing effect and artifact suppression capability of the final image.

[0045] In some embodiments, in order to determine the fusion weights, the computer device extracts image features of the (i-1)th deformed image and image features of the i-th jittery image, respectively.

[0046] Image features refer to structured information that can characterize image content from multiple dimensions. Compared to single-dimensional image information, image features can express richer and more comprehensive image content attributes.

[0047] In some embodiments, after determining the image features of the (i-1)th deformed image and the image features of the i-th jittery image, the computer device calculates the feature correlation between the two sets of image features and determines the fusion weight based on the feature correlation.

[0048] Feature correlation is used to measure the similarity between the (i-1)th deformed image and the i-th jittered image at the feature level. Generally speaking, the higher the feature correlation, the closer the content of the (i-1)th deformed image is to the content of the i-th jittered image, meaning that the historical image information still has a high degree of credibility in the current frame; conversely, the lower the feature correlation, the greater the change in content between frames, and the lower the credibility of the historical image information.

[0049] In some embodiments, the computer device determines the fusion weight based on the calculated feature correlation. When the feature correlation is high, a larger fusion weight is assigned to the (i-1)th deformed image; when the feature correlation is low, the fusion weight of the (i-1)th deformed image is reduced accordingly to avoid introducing artifacts or ghosting from erroneous historical information.

[0050] In some embodiments, the fusion weight determined by the computer device is a single value. In other embodiments, the fusion weight determined by the computer device is a plurality of weights per pixel, and the number of such weights matches the resolution of the i-th jittery image, i.e., each pixel corresponds to an independent fusion weight. This application does not limit the specific form of the fusion weights.

[0051] Step 230: Based on the fusion weight, perform image fusion on the (i-1)th deformed image and the ith jittery image to obtain the ith output image.

[0052] After determining the fusion weights, the computer device performs image fusion processing on the (i-1)th deformed image and the i-th jittery image according to the fusion weights.

[0053] In some embodiments, when a computer device performs image fusion processing, it fuses the pixel values ​​of corresponding pixels in the (i-1)th deformed image and the ith jittery image.

[0054] To illustrate, for each pixel position in the image, the pixel value of the (i-1)th deformed image at that position and the pixel value of the i-th jittery image at that position are obtained respectively. The two are then combined according to the fusion weight to obtain the new pixel value of the i-th output image at the same position.

[0055] In some embodiments, the computer device performs image fusion by weighted fusion of the i-th jittery image and the (i-1)-th deformed image to obtain the i-th output image.

[0056] As an illustration, the i-th output image can be calculated using the following formula:

[0057] in, For the (i-1)th deformed image, For the i-th jittery image, To integrate weights, The value range is [0, 1]. When When the value is larger, historical information contributes more to the current frame's output image, which is beneficial for enhancing the smoothing effect of time accumulation; when... When the value is smaller, the current jittery image contributes more, which is beneficial for quickly responding to scene changes.

[0058] It is worth noting that the embodiments of this application do not impose specific restrictions on the method of image fusion based on fusion weights by computer devices.

[0059] By fusing the (i-1)th deformed image and the ith jittery image, the ith output image inherits effective information from the historical images to smooth jagged edges, and dynamically adjusts the participation level of the historical images based on feature correlation. This effectively suppresses artifacts caused by weight misjudgment while maintaining anti-aliasing performance. Finally, the computer outputs or stores the ith output image as the result of rendering and displaying the ith frame.

[0060] In summary, by deforming the (i-1)th output image based on optical flow information, spatial alignment between the historical image and the current image is achieved. Furthermore, the fusion weights are determined based on the feature correlation between the (i-1)th deformed image and the i-th jittered image. Since image features can represent image content from richer dimensions, the fusion weights determined based on feature correlation are more representative. When artifacts exist in the (i-1)th deformed image, the feature correlation of the artifact region is low, and the fusion weights can correspondingly reduce the fusion ratio of the (i-1)th deformed image, focusing on fusing the i-th jittered image, thereby enabling the generated i-th output image to effectively remove artifacts.

[0061] Optionally, the image processing method provided in this application embodiment can be executed by a chip, which can be a graphics processing unit (GPU) chip, a neural network processing unit (NPU) chip, or other chips that can be used for image processing. This application embodiment does not limit this.

[0062] In some embodiments, in order to improve the accuracy of the determined fusion weights, the computer device first determines the image features of the (i-1)th deformed image and the i-th jittery image, and then determines the fusion weights based on the feature correlation between the two image features. The correlation between high-level image features is used to more reliably measure the structural similarity between the two images, thereby improving the accuracy of the fusion weights and the image fusion quality.

[0063] like Figure 6As shown, a flowchart of an image processing method provided by another illustrative embodiment of this application is illustrated. The method can be executed by the computer device described above. Step 220 in the above-described steps can also be implemented as steps 221, 222 and 223.

[0064] Step 221: Extract features from the (i-1)th deformed image and the i-th jittered image to obtain the first image feature corresponding to the (i-1)th deformed image and the second image feature corresponding to the i-th jittered image.

[0065] When determining the fusion weight based on the feature correlation between the image features of the (i-1)th deformed image and the image features of the i-th jittery image, the computer device first needs to obtain the image features of the (i-1)th deformed image and the i-th jittery image.

[0066] Among them, the image features obtained by computer equipment, compared with the original pixel values ​​or single, simple local features, such as motion vectors, neighborhood color variance, etc., can represent the image content from a richer dimension and provide a more reliable basis for subsequent correlation calculation.

[0067] In some embodiments, the computer device extracts features from the (i-1)th deformed image and the i-th jittery image using a feature extraction network of an image processing model, respectively, to obtain first image features and second image features.

[0068] Specifically, this image processing model includes a feature extraction network. The computer device inputs the (i-1)th deformed image and the i-th jittery image into the feature extraction network, which performs layer-by-layer convolution or transformation processing on the input images to extract high-dimensional structured feature representations from the original pixel space. The outputs of the feature extraction network are denoted as the first image feature and the second image feature, respectively.

[0069] Step 222: Determine the feature correlation coefficient based on the first image features and the second image features.

[0070] After obtaining the first image features and the second image features, the computer device needs to quantify the degree of similarity between them. To do this, the computer device needs to further determine the feature correlation between the first image features and the second image features.

[0071] Among them, the feature correlation coefficient is used to characterize the consistency between the (i-1)th deformed image and the ith jittery image at the feature level: the higher the coefficient, the closer the two images are in the feature space, that is, the higher the credibility of historical information in the current frame; the lower the coefficient, the more significant the changes in the content between frames, and the lower the reference value of historical information.

[0072] In some embodiments, the image processing model further includes a correlation determination network, through which a computer device determines the feature correlation coefficients based on a first image feature and a second image feature.

[0073] In illustrative terms, a computer device inputs first image features and second image features into a correlation determination network of an image processing model. This correlation determination network calculates the similarity between two feature vectors or feature maps, such as a dot product, cosine similarity, or a learnable metric function, and outputs one or more feature correlation coefficients. This application does not limit the specific method by which the image processing model determines the feature correlation coefficients.

[0074] In some embodiments, when a computer device determines the feature correlation coefficient based on the first image feature and the second image feature through the correlation determination network of the image processing model, it needs to determine the independent feature attributes of the first image feature and the second image feature, as well as the correlation feature attributes between the two.

[0075] In some embodiments, the computer device generates a first autocorrelation feature of the first image feature, a second sub-correlation feature of the second image feature, and a cross-correlation feature of the first image feature and the second image feature based on the first image feature and the second image feature.

[0076] Among them, the first autocorrelation feature is used to characterize the internal structure information and feature distribution characteristics of the (i-1)th deformed image itself; the second autocorrelation feature is used to characterize the internal structure information and feature distribution characteristics of the i-th jittery image itself; and the cross-correlation feature is used to characterize the cross-image association information between the (i-1)th deformed image and the i-th jittery image, such as the feature correspondence between the two at corresponding positions or in the neighborhood.

[0077] In some embodiments, the first autocorrelation feature is the square of the first image feature, and the second autocorrelation feature is the square of the second image feature. In other embodiments, the first autocorrelation feature is the first image feature itself, and the second autocorrelation feature is the second image feature itself. In still other embodiments, the first and second autocorrelation features are generated by an image processing model through learning, and there is no direct numerical calculation relationship between them and the first and second image features.

[0078] It is worth noting that the embodiments of this application do not limit the specific generation method of the first autocorrelation feature and the second autocorrelation feature.

[0079] In some embodiments, the cross-correlation feature is the product of a first image feature and a second image feature. In other embodiments, the cross-correlation feature is the cosine similarity or other similarity measure between the first image feature and the second image feature.

[0080] It is worth noting that the specific calculation method for the cross-correlation features is not limited in the embodiments of this application.

[0081] In some embodiments, after obtaining the first autocorrelation feature, the second autocorrelation feature, and the cross-correlation feature, the computer device inputs the spliced ​​feature obtained by splicing the first autocorrelation feature, the second autocorrelation feature, and the cross-correlation feature into the correlation determination network of the image processing model to obtain the feature correlation coefficient output by the correlation determination network.

[0082] Among them, the splicing feature integrates the inherent attributes of a single image with the interrelationship between two images, which can provide a more comprehensive information basis for correlation judgment.

[0083] Optionally, the computer device uses the `cat` function to concatenate the first autocorrelation feature, the second autocorrelation feature, and the cross-correlation feature into a concatenated feature, and then inputs the concatenated feature into the correlation determination network of the image processing model. The correlation determination network performs feature transformation and mapping on the concatenated feature through convolutional layers, fully connected layers, or attention mechanisms, and finally outputs the feature correlation coefficient. This application embodiment does not impose limitations on this.

[0084] By first determining the first autocorrelation feature, the second autocorrelation feature, and the cross-correlation feature, and then determining the feature correlation coefficient, compared to directly calculating similarity based on the original image features, the correlation determination network can simultaneously utilize the structural information of the image itself and the interrelationships between images, thereby outputting a more accurate feature correlation coefficient and improving the accuracy of the determined feature correlation coefficient.

[0085] Step 223: Determine the fusion weights based on the first image features, the second image features, and the feature correlation coefficients.

[0086] After obtaining the first image features, the second image features, and the feature correlation coefficients, the computer device needs to integrate multiple pieces of information to generate the final fusion weights.

[0087] In some embodiments, the image processing model further includes a weight determination network, and the computer device determines the fusion weights based on the first image features, the second image features, and the feature correlation coefficient through the weight determination network of the image processing model.

[0088] Specifically, the computer device inputs the first image feature, the second image feature, and the feature correlation coefficient into the weight determination network. The network outputs fusion weights by combining the content information of the image features themselves and the similarity reflected by the feature correlation coefficient through a pre-trained mapping function.

[0089] Among them, the computer equipment uses an image processing model to determine the fusion weights. Compared with the method of determining the fusion weights by relying solely on feature correlation coefficients, the weight determination network can more finely adjust the contribution of historical information in different regions and under different conditions by simultaneously utilizing the original image features and correlation coefficients. This allows it to further suppress the generation of artifacts while ensuring anti-aliasing effects.

[0090] However, although the feature correlation coefficient reflects the degree of consistency between two images at the feature level, it is essentially a similarity-based metric and does not contain direct decision information regarding the fusion weights. For example, when textured regions and flat regions have the same feature correlation coefficient, using the same fusion weights is unlikely to achieve ideal fusion results. Therefore, a single similarity metric is insufficient to cover the aforementioned differentiated needs. Thus, it is necessary to introduce an additional learnable parameter to fine-tune the fusion weights based on the content information of the image features themselves.

[0091] In some embodiments, when determining the fusion weights, the computer device inputs the first image features and the second image features into the weight determination network of the image processing model to obtain the fusion coefficients output by the weight determination network.

[0092] The fusion coefficient is used to characterize the initial weight benchmark learned by the weight determination network based on the image features of two images, reflecting the fusion tendency of the network based on the image content information of the current frame and historical frames.

[0093] Optionally, the weight determination network can adjust the fusion coefficient by analyzing the texture complexity in image features: when the image region has rich texture, detail information is more important, and the network tends to reduce the fusion coefficient of historical frames to reduce motion blur caused by mismatched historical information; when the image region is relatively flat, the network tends to increase the fusion coefficient of historical frames to enhance the smoothing effect of time accumulation. Based on motion amplitude information, when a large-amplitude inter-frame motion is detected, the network reduces the fusion coefficient of historical images to avoid introducing artifacts from erroneous historical information. This embodiment of the application does not impose limitations on this.

[0094] In some embodiments, after determining the fusion coefficients, the computer device determines the fusion weights based on the fusion coefficients and the feature correlation coefficients.

[0095] Among them, the feature correlation coefficient provides an objective measure of the two frames at the feature level, while the fusion coefficient provides supplementary information learned by the weight determination network based on the image content. By combining the two, the final fusion weight not only utilizes the objective quantitative result of feature correlation, but also integrates the prior knowledge and scene adaptation ability learned by the network from image features. This allows for a more refined adjustment of the contribution of historical information in different regions and under different conditions, further suppressing the generation of artifacts while ensuring anti-aliasing effect.

[0096] By first determining the fusion coefficient based on the first image features and the second image features, and then determining the fusion weight based on the fusion coefficient and the feature correlation coefficient, compared to directly inputting the first image features, the second image features, and the feature correlation coefficient into the network to output the fusion weight, the fusion coefficient can make fuller use of the content information of the image features, thus improving the accuracy of the fusion coefficient. At the same time, by combining the feature correlation coefficient for correction, the information of the image itself and the relevant information between images can be better balanced in image fusion, thereby obtaining higher quality fusion weights, improving anti-aliasing effect and suppressing artifacts.

[0097] In some embodiments, the computer device performs a weighted summation or multiplication of the fusion coefficients and feature correlation coefficients, or combines them using other preset fusion functions, to obtain the fusion weights. This application does not limit this approach.

[0098] In summary, computer equipment uses image processing models to determine image features and the correlation coefficients between these features, and then determines the fusion weights based on these features and correlation coefficients. The image processing model can adaptively adjust the weight decision process according to the content of the input image, without requiring manually preset fixed rules or parameters. By automatically learning and extracting effective feature information from images using the model, the determination of fusion weights becomes more accurate and reliable, thus achieving high-quality fusion weights in various image scenarios.

[0099] In some embodiments, the image processing model can realize the complete processing process from input image to output image processing result. That is, the image processing model has complete anti-aliasing processing capability and can independently complete the anti-aliasing processing of input image.

[0100] like Figure 7 The diagram illustrates the model structure of an image processing model provided in an illustrative embodiment of this application, wherein the deformation module 701 is used to implement the deformation processing described in the above embodiment. Specifically, the deformation module 701 performs deformation processing on the (i-1)th output image based on optical flow information to obtain the (i-1)th deformed image.

[0101] The first feature extraction network 702 is used to perform feature extraction on the (i-1)th deformed image as described in the above embodiment, and output the first image feature. The second feature extraction network 703 is used to perform feature extraction on the i-th jittery image as described in the above embodiment, and output the second image feature.

[0102] The first feature extraction network 702 and the second feature extraction network 703 can be the same feature extraction network, such as a Siamese network with shared weights, or they can be two different feature extraction networks. This application embodiment does not limit this.

[0103] The correlation determination network 704 is used to implement the feature correlation coefficient 705 based on the first image feature and the second image feature as described in the above embodiment. Specifically, the correlation determination network 704 calculates the correlation between the first image feature and the second image feature and outputs the feature correlation coefficient 705.

[0104] The fusion coefficient determination network 706 is used to implement the fusion coefficient determination 707 based on the first image features and the second image features described in the above embodiment. Specifically, the fusion coefficient determination network 706 integrates the first image features and the second image features to output the fusion coefficient 707.

[0105] The weight determination network 708 is used to implement the process of determining the fusion weights based on the fusion coefficient 707 and the feature correlation coefficient 705 as described in the above embodiments. Specifically, the weight determination network 708 performs a multiplication operation on the fusion coefficient 707 and the feature correlation coefficient 705, and outputs the first fusion weight 709 and the second fusion weight 710.

[0106] The first fusion weight 709 is the fusion weight directly output by the weight determination network 708, used to indicate the fusion weight occupied by the (i-1)th deformed image in the image fusion process. The second fusion weight 710 refers to the fusion weight occupied by the i-th jittery image in the image fusion process. The second fusion weight 710 is determined by the difference between 1 and the first fusion weight 709, that is: second fusion weight 710 = 1 - first fusion weight 709.

[0107] After determining the first fusion weight 709 and the second fusion weight 710, the image processing model performs image fusion on the (i-1)th deformed image and the i-th jittered image according to the aforementioned fusion weights. Specifically, the product of the (i-1)th deformed image and the first fusion weight 709 is added to the product of the i-th jittered image and the second fusion weight 710 to obtain the i-th output image.

[0108] In summary, through the above model structure, the image processing model can adaptively determine the fusion weights based on the content of the input image, thereby improving the accuracy and reliability of image processing.

[0109] In some embodiments, the computer device needs to train an image processing model. The training process of the image processing model is described below.

[0110] like Figure 8 As shown, a flowchart of an image processing model training process provided in an illustrative embodiment of this application is illustrated. This process can be executed by the aforementioned computer device and includes at least one of the following steps.

[0111] Step 810: Input the sample jitter image sequence into the image processing model to obtain the predicted output image sequence output by the image processing model.

[0112] In some embodiments, for training an image processing model, the computer device first needs to construct a training dataset, which includes a sequence of sample jittered images.

[0113] Among them, the sample jitter image sequence refers to a multi-frame image sequence obtained by applying jitter sampling to the original image sequence.

[0114] In some embodiments, when performing image processing, the image processing model needs to generate the i-th output image based on the (i-1)-th output image and the i-th jitter image. For the first image to be processed (i.e., the case where i=1), there is no 0-th output image, and the computer device uses the 0-th jitter image as the initial value of the 0-th output image.

[0115] To illustrate, during the model training phase, the image processing model uses the 0th output image (i.e., the 0th jittered image) and the 1st jittered image as initial inputs to generate the 1st output image; subsequently, the 1st output image is used as the historical input for the next frame, and together with the 2nd jittered image, it is input into the model to generate the 2nd output image, and so on.

[0116] In some embodiments, the computer device inputs a sample jittered image sequence into an image processing model to be trained. The model processes the input jittered image sequence according to the image processing method described in the above embodiments and outputs a corresponding predicted output image sequence.

[0117] Step 820: Determine the first model loss based on the predicted output image sequence and the ground truth image sequence corresponding to the sample jitter image sequence.

[0118] In some embodiments, after obtaining the predicted output image sequence, the computer device compares the predicted output image sequence with the ground truth image sequence corresponding to the sample jittered image sequence. The ground truth image sequence corresponding to the sample jittered image sequence refers to the corresponding high-quality reference image used as the reference output result for the image processing model.

[0119] In some embodiments, the computer device obtains the ground truth image sequence corresponding to the sample jitter image sequence through high sampling rate rendering. In other embodiments, the computer device obtains the ground truth image sequence by accumulating multiple frames of jitter images and performing noise reduction processing. In still other embodiments, the computer device obtains the ground truth image sequence through manual annotation or from a public dataset. This application does not limit the specific method of obtaining the ground truth image sequence.

[0120] In some embodiments, when determining the first model loss, the computer device determines the first model loss by calculating the difference between the predicted output image and the ground truth image. The first model loss is used to quantify the degree of deviation between the predicted result of the current output of the image processing model and the real target.

[0121] Optionally, the computer device obtains the first model loss by employing loss functions such as mean squared error loss, perceptual loss, or adversarial loss. This application does not limit this approach.

[0122] Step 830: Train the image processing model based on the first model loss.

[0123] In some embodiments, the computer device trains the network parameters of the image processing model using a first model loss.

[0124] Indicatively, the computer device calculates the gradient of the loss function with respect to the model parameters using the backpropagation algorithm, and updates the model parameters using gradient descent to reduce the initial model loss. Through multiple iterations of training, the image processing model gradually learns the mapping relationship from jittery image sequences to high-quality output images, ultimately obtaining a fully trained image processing model.

[0125] It is worth noting that the above-described method of training the image processing model based on the first model loss is merely an illustrative example, and the embodiments of this application do not limit the specific method of training the image processing model based on the first model loss.

[0126] By inputting sample jittery image sequences into an image processing model and obtaining predicted output image sequences, then determining the first model loss based on the predicted output image sequences and the ground truth image sequences, and finally training the image processing model based on this loss, the model can learn the mapping relationship from jittery image sequences to high-quality output images. This training method enables the model to automatically learn processing strategies from data, improving the model's adaptability and generalization ability, thereby achieving good anti-aliasing effects in different image scenarios.

[0127] In some embodiments, the image processing model struggles to converge effectively because the sample jitter image sequence is directly input into the model for training. To reduce the learning difficulty and accelerate the convergence process, the image processing model introduces an additional residual prediction branch in its weight determination network. This weight determination network is also used to determine image residuals, which are used to compensate for the predicted output images in the predicted output image sequence, thereby correcting the deviation between the model's prediction and the true target.

[0128] By incorporating residuals as part of the learning objective, computer devices enable image processing models to learn to compensate for the differences between predicted and true results simultaneously, rather than fitting the complete output image from scratch. This eliminates the need to directly generate true results, thereby reducing training difficulty and improving convergence speed and prediction accuracy.

[0129] In some embodiments, before training the image processing model based on the first model loss, the computer device first inputs the sample jitter image sequence into the image processing model to obtain the predicted output image sequence and the image residual sequence output by the image processing model.

[0130] The image residual sequence is generated by the weight determination network of the image processing model and is used to compensate for the difference between the predicted output image and the real target. By introducing a residual prediction branch, the model does not need to learn the complete output image from scratch, but instead learns to compensate for the difference between the predicted result and the real result simultaneously, without directly generating the ground truth result, thus reducing the training difficulty.

[0131] In some embodiments, each pixel in the predicted output image corresponds to an independent image residual value. In other embodiments, an entire predicted output image corresponds to a single image residual value. This application does not limit the granularity of the image residual values.

[0132] In some embodiments, the computer device determines the second model loss based on the predicted output image sequence, the image residual sequence, and the ground truth image sequence corresponding to the sample jitter image sequence.

[0133] The second model loss takes into account both the difference between the predicted output image and the ground truth image, as well as the image residual.

[0134] In some embodiments, when determining the second model loss based on the predicted output image sequence, the image residual sequence, and the ground image sequence corresponding to the sample jitter image sequence, the computer device first compensates the predicted output image sequence based on the image residual sequence to obtain the compensated predicted output image sequence.

[0135] Optionally, the image residual sequence output by the weight determination network of the image processing model is used to correct the deviation in the predicted output image sequence. The computer device adds or weights the predicted output image sequence and the corresponding image residual sequence pixel by pixel, so that the residual information supplements and corrects the prediction result, thereby obtaining the compensated predicted output image sequence.

[0136] This application does not impose specific limitations on the method by which computer devices compensate for the predicted output image sequence based on the image residual sequence.

[0137] As an illustration, the i-th predicted output image after compensation from the computer device can be calculated using the following formula:

[0138] in, For the (i-1)th sample deformed image, For the i-th sample jitter image, To integrate weights, The image residual is used. According to the above calculation method, after the computer device performs weighted fusion of the deformed image of the (i-1)th sample and the jittery image of the ith sample based on the fusion weight, it further superimposes the image residual onto the fusion result to compensate and correct the predicted output image.

[0139] By compensating the predicted output image sequence, the model learns not only the output image itself during training, but also the difference between the output result and the real target. This eliminates the need to directly generate the real target result, which helps reduce the learning difficulty of the model.

[0140] In some embodiments, after obtaining the compensated predicted output image sequence, the computer device determines the image sequence loss based on the compensated predicted output image sequence and the ground truth image sequence.

[0141] Optionally, the computer device calculates the difference between the compensated predicted output image and the ground truth image, for example, by using loss functions such as mean squared error loss, perceptual loss, or absolute error loss to obtain an image sequence loss. This image sequence loss is used to quantify the degree of deviation between the compensated prediction result and the true target.

[0142] This application does not impose specific limitations on the method by which a computer device determines the image sequence loss based on the compensated predicted output image sequence and the ground truth image sequence.

[0143] In some embodiments, the computer device determines the residual sequence loss based on the image residual sequence.

[0144] Image residuals are used to correct biases in the predicted output image sequence, reflecting the difference between the predicted output image and the ground truth image from the image processing model. To enable the model to learn accurate residual estimation capabilities, this difference also needs to be included in the loss function for supervised learning.

[0145] Optionally, the computer device performs preset calculations on each image residual in the image residual sequence to determine the residual sequence loss. For example, the image residual can be multiplied by a preset weighting coefficient, or the square of the image residual can be calculated, and then the results at each position can be summed or averaged. This application embodiment does not limit the specific calculation method of the residual sequence loss.

[0146] In some embodiments, the computer device determines the second model loss based on image sequence loss and residual sequence loss.

[0147] The second model loss combines the two losses to simultaneously supervise the prediction quality of the output image and the prediction accuracy of the residual, enabling the model to take into account the learning objectives of both tasks during training, thereby achieving a better overall training effect.

[0148] Optionally, the computer device can combine the image sequence loss and the residual sequence loss using a weighted summation method, which involves multiplying the image sequence loss by a first weight coefficient and the residual sequence loss by a second weight coefficient, and then adding the two together to obtain the second model loss. Alternatively, other combination methods can be used, such as taking the maximum or average of the two values. This application does not limit this approach.

[0149] In some embodiments, using fixed weights on the residual sequence loss when determining the second model loss may cause the model to over-rely on the residual prediction branch, weakening the autonomous learning ability of the output image prediction branch and even affecting the model's convergence stability. Therefore, the computer device employs a dynamic weighting strategy for the residual sequence loss. Using dynamic weights adaptively adjusts the contribution ratio of the residual sequence loss to the total loss according to the training process, thereby guiding the model to focus on overall output quality in the early stages of training and gradually enhancing the accuracy of residual prediction in the later stages, thus improving the overall performance of the model.

[0150] In some embodiments, when training an image processing model, the computer device divides the training process into a fixed-weight training phase and a dynamic-weight training phase, depending on whether the weights of the residual sequence loss are fixed.

[0151] The computer equipment divides the training process into two stages. In the fixed-weight training stage, the image processing model can be initially trained under relatively stable loss constraints, avoiding unstable gradient changes introduced by dynamic weights in the early stages of training. In the dynamic-weight training stage, as the weight of the residual sequence loss gradually increases, the image processing model pays more attention to the residual prediction task, thereby reducing the model's dependence on fixed weight settings and enabling the model to learn the mapping relationship between the input and output images more precisely.

[0152] In some embodiments, during the fixed-weight training phase, the computer device determines the second model loss based on the image sequence loss, the residual sequence loss, and the fixed residual weights of the residual sequence loss.

[0153] To illustrate, when determining the second model loss, the computer device adds the image sequence loss and the residual sequence loss multiplied by a fixed residual weight. This fixed residual weight remains unchanged during training, providing the model with a stable supervision signal, enabling it to learn both output image prediction and residual prediction tasks in a balanced manner in the early stages of training.

[0154] In some embodiments, during the dynamic weight training phase, the computer device determines a second model loss based on the image sequence loss, the residual sequence loss, and the dynamic residual weights of the residual sequence loss, wherein the dynamic residual weights gradually increase during the dynamic weight training phase, and the dynamic weight training phase follows the fixed weight training phase.

[0155] Optionally, after the model completes its initial learning in the fixed-weight training phase, the computer device enters the dynamic-weight training phase, where the weights of the residual sequence loss gradually increase according to a preset scheduling strategy. By introducing residual supervision signals, the second model loss can more directly guide the model to learn the residual prediction task, thereby improving the training effect.

[0156] In some embodiments, the computer device gradually increases the dynamic residual weight of the residual sequence loss using a linear growth method. In other embodiments, the computer device gradually increases the dynamic residual weight of the residual sequence loss using an exponential growth method. This application does not limit the specific growth method of the dynamic residual weight.

[0157] The formula for calculating residual sequence loss is illustrated below:

[0158] in, This refers to the i-th predicted output image. This refers to the corresponding truth output image. For image residuals, For dynamic residual weights, The function is used to calculate the absolute value loss. The computer device increases the value linearly step by step. The value of is used to gradually increase the weight of residual loss in residual sequence loss.

[0159] By dividing the model training process into a fixed-weight training phase and a dynamic-weight training phase, and gradually increasing the dynamic residual weights of the residual sequence loss during the dynamic-weight training phase, the model's dependence on the residual branch can be gradually reduced after the model training reaches a certain level. This allows the model to focus more on its ability to generate high-quality output images without relying on the residuals. This enables the image processing model to generate accurate predictions independently during the inference phase, even without using the residual branch, thereby improving the model's generalization ability and inference efficiency.

[0160] In some embodiments, the dynamic residual weights of the residual sequence loss stop increasing after the model training reaches a preset condition. For example, when the model's prediction accuracy reaches a preset threshold, or when the number of training iterations reaches a preset number of steps, the dynamic residual weights will no longer increase and will maintain their current value. This application does not limit the specific conditions under which the dynamic residual weights stop increasing.

[0161] By compensating the predicted output image sequence based on the image residual sequence and using the compensated predicted output image sequence to determine the second model loss, compared to directly using the predicted output image to calculate the second model loss, calculating the loss based on the compensated predicted output image enables the model to simultaneously optimize the residual prediction task and the output image prediction task, thereby improving the overall training effect of the model and the final image processing quality.

[0162] In some embodiments, the computer device trains the image processing model based on a second model loss.

[0163] Optionally, the computer device calculates the gradient of the second model loss with respect to the model parameters using the backpropagation algorithm, and updates the model parameters using gradient descent to reduce the second model loss. This application embodiment does not limit the specific method by which the computer device trains the image processing model based on the second model loss.

[0164] like Figure 9 As shown, this diagram illustrates the model structure of an image processing model provided in another illustrative embodiment of this application, wherein the deformation module 901 is used to implement the deformation processing described in the above embodiment. Specifically, the deformation module 901 performs deformation processing on the (i-1)th output image based on optical flow information to obtain the (i-1)th deformed image.

[0165] The first feature extraction network 902 is used to perform feature extraction on the (i-1)th deformed image as described in the above embodiment, and output the first image feature. The second feature extraction network 903 is used to perform feature extraction on the i-th jittery image as described in the above embodiment, and output the second image feature.

[0166] The first feature extraction network 902 and the second feature extraction network 903 can be the same feature extraction network, such as a Siamese network with shared weights, or they can be two different feature extraction networks. This application embodiment does not limit this.

[0167] The correlation determination network 904 is used to implement the determination of the feature correlation coefficient 906 based on the first image feature and the second image feature as described in the above embodiment. Specifically, the correlation determination network 904 calculates the correlation between the first image feature and the second image feature and outputs the feature correlation coefficient 906.

[0168] The fusion coefficient determination network 905 is used to implement the determination of fusion coefficient 907 and image residual 908 based on the first image features and the second image features as described in the above embodiment. Specifically, the fusion coefficient determination network 905 integrates the first image features and the second image features to output fusion coefficient 907 and image residual 908.

[0169] Among them, the fusion coefficient 907 serves as the initial weight benchmark, reflecting the fusion tendency of the model based on the image feature content; the image residual 908 is used to compensate for the deviation between the predicted output image and the real target, reflecting the gap between the model output and the ideal output.

[0170] The weight determination network 909 is used to implement the process of determining the fusion weights based on the fusion coefficient 907 and the feature correlation coefficient 906 as described in the above embodiments. Specifically, the weight determination network 909 performs a multiplication operation on the fusion coefficient 907 and the feature correlation coefficient 906, and outputs the first fusion weight 910 and the second fusion weight 911.

[0171] Wherein, the first fusion weight 910 is the fusion weight directly output by the weight determination network 909, used to indicate the fusion weight occupied by the (i-1)th deformed image in the image fusion process; the second fusion weight 911 refers to the fusion weight occupied by the i-th jittery image in the image fusion process. The second fusion weight 911 is determined by the difference between 1 and the first fusion weight 910, that is: second fusion weight 911 = 1 - first fusion weight 910.

[0172] After determining the first fusion weight 910 and the second fusion weight 911, the image processing model performs image fusion on the (i-1)th deformed image and the i-th jittery image according to the aforementioned fusion weights, and further introduces image residual 908 to compensate for the fusion result. Specifically, the product of the (i-1)th deformed image and the first fusion weight 910 is added to the product of the i-th jittery image and the second fusion weight 911 to obtain a preliminary fusion result; then the image residual 908 is added to or weighted with this preliminary fusion result to obtain the final i-th output image.

[0173] In summary, through the above model structure, the image processing model can not only adaptively determine the fusion weights based on the content of the input image, but also compensate and correct the fusion result through image residuals, thereby further improving the accuracy and reliability of image processing.

[0174] By introducing an additional residual prediction branch into the image processing model, the model can use image residuals to correct the deviation between the predicted result and the real target, without having to directly fit the complete output image. This reduces the learning difficulty of the model, helps to obtain better convergence results, and thus improves the training efficiency and performance of the image processing model.

[0175] In some embodiments, since the image processing model is trained directly using sequence data, the model needs to learn both inter-frame temporal dependencies and single-frame image processing capabilities simultaneously, which is quite challenging and can easily lead to insufficient continuous processing capabilities. To reduce the learning difficulty of the model, a single-step training method is first used to pre-train the model before using the second model loss. In the single-step training phase, the output of the previous frame is not used as the input of the next frame, meaning that each training sample is independent and does not constitute a temporal dependency.

[0176] In some embodiments, during the single-step training phase, the computer device inputs the i-th jittered image and the (i-1)-th deformed image into the image processing model for the i-th jittered image in the sample jittered image sequence, and obtains the i-th predicted output image and the i-th image residual output by the image processing model. The (i-1)-th deformed image is obtained by deforming the (i-1)-th sample output image based on optical flow information.

[0177] In a schematic manner, the computer device inputs the jittered image of the i-th sample and the deformed image of the (i-1)-th sample into the image processing model. Following the image processing flow described in the above embodiment, the image processing model extracts the image features of the two images respectively, calculates the feature correlation coefficient, determines the fusion weight based on the fusion coefficient and the feature correlation coefficient, and finally outputs the i-th predicted output image and the i-th image residual.

[0178] Here, the i-th predicted output image is the model's prediction result of the current frame output image, and the i-th image residual is used to compensate for the deviation between the predicted output image and the ground truth image.

[0179] It is worth noting that during the single-step training phase, the output image of the i-1th sample does not come from the prediction output of the model at the previous time step. For example, it can come from the sample output image in the pre-built training dataset or the prediction output image generated in the previous training phase, thereby cutting off the temporal dependence between frames and allowing the model to focus on learning the mapping relationship of a single step.

[0180] like Figure 10 As shown, it illustrates a schematic diagram of a single-step training phase and a sequential training phase provided in an illustrative embodiment of this application, showing the data flow in the sequential training phase and the single-step training phase.

[0181] In the single-step training phase, for the training of the i-th frame, the computer inputs the i-th sample deformed image and the i-th sample jitter image into the image processing model to obtain the i-th predicted output image. It should be noted that the i-th sample deformed image is obtained by deforming the i-th sample output image based on optical flow information, but the i-th sample output image is not the predicted image output by the image processing model in the previous frame. In the single-step training phase, the training of each frame is independent; the prediction result output in the previous frame is not used as input for the training of the next frame. Specifically, the first predicted output image is only used to calculate the loss of the first frame and does not participate in the generation of the input for the second frame. This training method eliminates the temporal dependency between frames, reduces the learning difficulty of the model, and allows the model to focus on learning the mapping relationship of a single frame first.

[0182] In the sequence training phase, for the training of the i-th frame, the computer device inputs the i-th sample deformed image and the i-th sample jitter image into the image processing model to obtain the i-th predicted output image. Unlike the single-step training phase, in the sequence training phase, the prediction result of the previous frame is used as the historical input for the next frame. Specifically, the first predicted output image, after being deformed by adding optical flow information, is used as the historical input for the training of the second frame; the second predicted output image, after being deformed by adding optical flow information, is used as the historical input for the training of the third frame, and so on. Through this time-dependent transmission method, the model can learn the ability to process continuously between frames, thereby better utilizing historical information for anti-aliasing processing in actual inference.

[0183] In summary, the single-step training phase focuses on learning the mapping relationship of a single frame, while the sequence training phase further learns the temporal dependencies between frames based on this. Together, the two phases constitute a complete model training process.

[0184] In some embodiments, the computer device determines the third model loss based on the i-th predicted output image, the i-th image residual, and the i-th ground truth image corresponding to the i-th sample jitter image.

[0185] Here, the i-th predicted output image is the model's prediction result of the current frame's output image; the i-th image residual is used to compensate for the deviation between the predicted output image and the true target; and the i-th ground truth image serves as a high-quality reference image for the supervision signal. The computer device uses this to determine the third model loss used for training the model.

[0186] Here, the i-th predicted output image is the prediction result of the image processing model for the current frame output image; the i-th image residual is used to compensate for the deviation between the predicted output image and the real target; and the i-th ground truth image is used as a high-quality reference image for the supervision signal during model training. Based on the above three pieces of information, the computer device calculates the third model loss used to train the image processing model.

[0187] In some embodiments, when determining the third model loss, the computer device first compensates the i-th predicted output image based on the i-th image residual to obtain the compensated i-th predicted output image.

[0188] Optionally, the computer device adds or weights the i-th predicted output image and the i-th image residual pixel by pixel, so that the residual information corrects the prediction result, thereby obtaining a compensated predicted output image that is closer to the real target. This application embodiment does not specifically limit the method by which the computer device compensates the i-th predicted output image based on the i-th image residual.

[0189] The computer device compensates for the i-th predicted output image, enabling the image processing model to learn the difference between the predicted result and the real target during training, reducing the learning difficulty of the model directly fitting the complete output image.

[0190] In some embodiments, after compensating the i-th predicted output image, the computer device determines the image loss based on the compensated i-th predicted output image and the i-th ground truth image.

[0191] Optionally, the computer device calculates the difference between the compensated i-th predicted output image and the i-th ground truth image, for example, by using loss functions such as mean squared error loss, absolute error loss, or perceptual loss to obtain the image loss. This image loss is used to quantify the degree of deviation between the compensated prediction result and the true target, reflecting the output quality of the model after introducing residual compensation.

[0192] This application does not impose specific limitations on the method by which a computer device determines image loss based on the compensated i-th predicted output image and the i-th true image.

[0193] In some embodiments, after determining the image loss, the computer device determines the residual loss based on the residual of the i-th image.

[0194] The image residual is used to correct biases in the predicted output image, reflecting the difference between the predicted output image and the ground truth image. To enable the model to learn accurate residual estimation capabilities, this difference also needs to be included in the loss function for supervision.

[0195] Optionally, the computer device performs a preset calculation on the i-th image residual to determine the residual loss. For example, the image residual can be multiplied by a preset weighting coefficient, or the square of the image residual can be calculated, and then the results at each position can be summed or averaged. This application embodiment does not limit the specific calculation method of the residual loss.

[0196] In some embodiments, the computer device determines a third model loss based on image loss and residual loss.

[0197] The third model loss combines the two losses to simultaneously supervise the prediction quality of the output image and the prediction accuracy of the residual, enabling the model to take into account the learning objectives of both tasks during training, thereby achieving a better overall training effect.

[0198] Optionally, the computer device can combine the image loss and the residual loss using a weighted summation method, multiplying the image loss by a first weight coefficient and the residual loss by a second weight coefficient, and then adding the two together to obtain the third model loss; alternatively, other combination methods can be used, such as taking the maximum or average of the two. This application does not limit this approach.

[0199] By compensating the i-th predicted output image based on the i-th image residual and using the compensated i-th predicted output image to determine the third model loss, compared to directly using the predicted output image to calculate the third model loss, calculating the loss based on the compensated i-th predicted output image enables the model to simultaneously optimize the residual prediction task and the output image prediction task, thereby improving the overall training effect of the model and the final image processing quality.

[0200] In some embodiments, during the single-step training phase, the sample output image is selected from the pre-generated sample output image and the compensated predicted output image.

[0201] The pre-generated sample output image refers to image data that is pre-calculated and stored using a preset generation algorithm before training begins. For example, a high sampling rate rendering method can be used to preprocess each frame of the training dataset to generate high-quality, alias-free images as pre-generated sample output images. The pre-generated sample output image remains fixed during training and does not change with updates to the model parameters.

[0202] The compensated predicted output image refers to the image obtained after residual compensation of the predicted images output by the image processing model in the first few training iterations during multiple training rounds. Illustratively, after a certain number of training rounds, the model has acquired a certain predictive ability. At this point, the compensated predicted output image from the previous round can be used as the historical input for the current round to generate the deformed image of the (i-1)th sample.

[0203] In some embodiments, during actual training, the computer device dynamically selects between using pre-generated sample output images or compensated predicted output images based on the training rounds or the model's convergence level. Illustratively, in the early stages of training, when the model's predictive ability is weak, pre-generated sample output images are used as historical input to provide stable supervision signals. In the later stages of training, as the model's predictive ability improves, compensated predicted output images are gradually introduced, allowing the model to gradually adapt to a closed-loop processing method that uses its own output as historical input.

[0204] In other embodiments, the computer device trains the image processing model using pre-generated sample output images in the first round of training; in subsequent rounds of training, images are randomly selected from the pre-generated sample output images and the compensated predicted output images as input images for training the image processing model. This application does not limit the specific method of selecting sample output images.

[0205] By selecting sample output images from pre-generated sample output images and compensated predicted output images, the image processing model can gradually utilize its own generated results as historical information during the training process, thereby laying the foundation for the subsequent transition from the single-step training stage to the sequential training stage, improving the smoothness of the training process and the model's adaptability to temporal dependencies.

[0206] In some embodiments, the computer device trains the image processing model based on a third model loss.

[0207] Optionally, the computer device calculates the gradient of the third model loss with respect to the model parameters using the backpropagation algorithm, and updates the model parameters using gradient descent or its variants to reduce the third model loss. This application does not limit the specific method by which the computer device trains the image processing model based on the third model loss.

[0208] Through the single-step training phase, the model can learn the mapping ability from jittery images to high-quality output images without paying attention to the temporal information between multiple frames. This provides good parameter initialization for the subsequent sequence training phase, thereby reducing the overall training difficulty and accelerating convergence.

[0209] In some embodiments, when a computer device generates a pre-generated sample output image, it generates an i-th pre-generated sample output image based on the i-th sample jitter image, the i-th ground truth image corresponding to the i-th sample jitter image, and a noise image.

[0210] The noisy image contains random noise values, which are used as mixing weights to fuse the jittery image of the i-th sample with the ground truth image of the i-th sample. The pre-generated sample output image refers to the image data pre-calculated and stored by a preset algorithm before training, serving as a source of historical information in the early stages of training.

[0211] Optionally, the computer device fuses the jittered image of the i-th sample, the ground truth image of the i-th sample, and the noisy image, for example, by using weighted summation, linear interpolation, or a deep learning-based fusion method, to obtain the output image of the i-th pre-generated sample. This application embodiment does not limit the specific method by which the computer device generates the output image of the i-th pre-generated sample.

[0212] By generating pre-generated sample output images and using them as one of the sources of historical information during training, a stable supervision signal can be provided to the model in the early stages of training. At the same time, the diversity of training data can be increased by using noisy images, thereby improving the model's generalization ability.

[0213] In some embodiments, the computer device generates an initial noise image, the size of which is smaller than the size of the i-th sample jitter image.

[0214] The initial noise image refers to an initial noise matrix generated through random sampling, the size of which is smaller than the size of the i-th sample jitter image. Optionally, each pixel value in the initial noise image follows a standard normal distribution, a uniform distribution, or other preset random distribution. The width and height of the initial noise image can be set to one-twentieth and one-tenth of the width and height of the i-th sample jitter image, respectively. This application does not limit the size range of the initial noise image or the distribution type of its pixel values.

[0215] Since the size of the initial noisy image is smaller than the size of the target image, it needs to be adjusted to match the size of the sample jitter image through interpolation and scaling.

[0216] In some embodiments, the computer device interpolates the initial noise image based on the size of the i-th sample jitter image to obtain an interpolated noise image.

[0217] Interpolation refers to the process of calculating the pixel value at an unknown location based on the known pixel values ​​using a preset interpolation algorithm, which is used to enlarge the size of an initial noisy image.

[0218] Optionally, the interpolation method includes bilinear interpolation, nearest neighbor interpolation, or bicubic interpolation. Through interpolation, the initial noise image is enlarged to the same size as the jitter image of the i-th sample, thus matching the size of the noise image with the jitter image to be processed. This application does not limit the specific interpolation method used.

[0219] By interpolating the initial noisy image, the changes in adjacent pixel values ​​in the interpolated noisy image can be smoothed out, thus simulating the distribution pattern of pixels with similar characteristics in local regions of a real image. Compared to a completely random distribution of pixel values, the noisy image after interpolation has correlation in local regions, which can avoid the lack of local structural features in the mixed image caused by overly random pixel values. Therefore, when the computer device mixes the i-th sample jitter image with the i-th ground truth image based on the interpolated noisy image, the generated pre-generated sample output image has more reasonable local consistency, which is beneficial for the model to learn effective image features and improve training results.

[0220] In some embodiments, the interpolated noise image is randomly scaled to obtain a noise image.

[0221] Random scaling refers to adjusting the pixel values ​​in the interpolated noisy image within a random range to modify the numerical distribution of the noisy image. Random scaling further increases the diversity of the noisy image, resulting in a richer feature distribution in the generated pre-generated sample output image, thereby improving the generalization ability of the model training.

[0222] In some embodiments, the computer device performs random filtering on the pixel values ​​in the interpolated noisy image. Random filtering refers to filtering or removing pixel values ​​according to preset threshold conditions, such as filtering out pixel values ​​that are greater than a certain upper limit or less than a certain lower limit, in order to change the distribution characteristics of the noisy image.

[0223] In other embodiments, the computer device adds a mask to the pixel values ​​in the interpolated noisy image. Illustratively, all pixel values ​​within a certain region of the image are uniformly set to a fixed value of 0.5 to perform a masking operation on that region.

[0224] In other embodiments, the computer device performs non-linear enhancement processing on the pixel values ​​in the interpolated noisy image, such that larger pixel values ​​are further amplified and smaller pixel values ​​are further reduced.

[0225] Schematic, a random scaling factor matrix a is generated, in which each element independently follows a uniform distribution U(1, 1.1), and the offset factor matrix is ​​calculated. Interpolated noise image Each pixel value in the data corresponds to and The element values ​​at the same position are compared, and the pixel values ​​are transformed according to the following formula:

[0226] in, The function is used to restrict the calculation result to the interval (0, 1). This indicates element-wise multiplication. After the above transformation, the originally larger pixel values ​​are further increased, and the originally smaller pixel values ​​are further decreased, thus achieving a non-linear enhancement of the pixel value distribution.

[0227] It is worth noting that the specific method of randomly scaling the interpolated noise image is not limited in the embodiments of this application.

[0228] By interpolating the initial noisy image, the changes in adjacent pixel values ​​in the interpolated noisy image can be made smoother, thereby simulating the distribution pattern of pixels with similar features in local areas of the real image. Then, random scaling can further increase the diversity of the noisy image, so that the generated pre-generated sample output image has a richer feature distribution, thereby improving the generalization ability of the model training.

[0229] Optionally, the model training process provided in the above embodiments can be executed by a chip, which can be a GPU chip, NPU chip, or other chip with model inference and model training functions. This application embodiment does not limit this.

[0230] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0231] like Figure 11 As shown, it illustrates a schematic diagram of an image processing apparatus provided in an illustrative embodiment of this application. The apparatus includes a deformation module 1110, a determination module 1120, and a fusion module 1130.

[0232] The deformation module 1110 is configured to perform deformation processing on the (i-1)th output image based on optical flow information to obtain the (i-1)th deformed image. The (i-1)th output image is the image rendered and displayed at the (i-1)th time, where i is a positive integer. The determination module 1120 is configured to determine the fusion weight based on the feature correlation between the image features of the (i-1)th deformed image and the image features of the i-th jittered image, wherein the i-th jittered image is obtained by jitter sampling of the original image at time i. The fusion module 1130 is configured to perform image fusion on the (i-1)th deformed image and the ith jitter image based on the fusion weights to obtain the ith output image.

[0233] The determining module 1120 is also configured to extract features from the (i-1)th deformed image and the i-th jittered image to obtain the first image feature corresponding to the (i-1)th deformed image and the second image feature corresponding to the i-th jittered image; The feature correlation coefficient is determined based on the first image features and the second image features; The fusion weights are determined based on the first image features, the second image features, and the feature correlation coefficient.

[0234] The determining module 1120 is further configured to extract features from the (i-1)th deformed image and the i-th jittery image respectively through the feature extraction network of the image processing model, so as to obtain the first image feature corresponding to the (i-1)th deformed image and the second image feature corresponding to the i-th jittery image; Based on the first image features and the second image features, the feature correlation coefficient is determined by the correlation determination network of the image processing model; Based on the first image features, the second image features, and the feature correlation coefficient, the fusion weights are determined by the weight determination network of the image processing model.

[0235] The determining module 1120 is also configured to generate a first autocorrelation feature of the first image feature, a second autocorrelation feature of the second image feature, and a cross-correlation feature of the first image feature and the second image feature based on the first image feature and the second image feature; The concatenated features, obtained by concatenating the first autocorrelation feature, the second autocorrelation feature, and the cross-correlation feature, are input into the correlation determination network of the image processing model to obtain the feature correlation coefficient output by the correlation determination network.

[0236] The determination module 1120 is also configured to input the first image features and the second image features into the weight determination network of the image processing model to obtain the fusion coefficients output by the weight determination network; The fusion weights are determined based on the fusion coefficient and the feature correlation coefficient.

[0237] In some embodiments, the image processing apparatus further includes a training module.

[0238] The training module is configured to input sample jittered image sequences into the image processing model to obtain the predicted output image sequence output by the image processing model; The first model loss is determined based on the predicted output image sequence and the ground image sequence corresponding to the sample jitter image sequence. The image processing model is trained based on the first model loss.

[0239] In some embodiments, the weight determination network is also used to determine image residuals, which are used to compensate for predicted output images in the predicted output image sequence; The training module is also configured to input the sample jittered image sequence into the image processing model before training the image processing model based on the first model loss, so as to obtain the predicted output image sequence and the image residual sequence output by the image processing model. The second model loss is determined based on the predicted output image sequence, the image residual sequence, and the ground image sequence corresponding to the sample jitter image sequence. The image processing model is trained based on the second model loss.

[0240] The training module is also configured to compensate the predicted output image sequence based on the image residual sequence to obtain the compensated predicted output image sequence. Based on the compensated predicted output image sequence and the ground image sequence, determine the image sequence loss; Determine the residual sequence loss based on the image residual sequence; The second model loss is determined based on image sequence loss and residual sequence loss.

[0241] The training module is also configured to determine the second model loss based on the image sequence loss, residual sequence loss, and fixed residual weights of the residual sequence loss during the fixed weight training phase; and to determine the second model loss based on the image sequence loss, residual sequence loss, and dynamic residual weights of the residual sequence loss during the dynamic weight training phase, wherein the dynamic residual weights gradually increase during the dynamic weight training phase, and the dynamic weight training phase is located after the fixed weight training phase.

[0242] In some embodiments, before training the image processing model based on the second model loss... The training module is also configured to, during the single-step training phase, input the i-th sample jitter image and the i-1-th sample deformed image into the image processing model for the i-th sample jitter image in the sample jitter image sequence, and obtain the i-th predicted output image and the i-th image residual output by the image processing model. The i-1-th sample deformed image is obtained by deforming the i-1-th sample output image based on optical flow information. The loss of the third model is determined based on the i-th predicted output image, the i-th image residual, and the i-th ground image corresponding to the i-th sample jitter image. The image processing model is trained based on the third model loss.

[0243] The training module is also configured to compensate the i-th predicted output image based on the i-th image residual to obtain the compensated i-th predicted output image; Based on the compensated i-th predicted output image and the i-th ground image, determine the image loss; Determine the residual loss based on the residual of the i-th image; The third model loss is determined based on image loss and residual loss.

[0244] In some embodiments, during the single-step training phase, the sample output image is selected from the pre-generated sample output image and the compensated predicted output image.

[0245] The training module is also configured to generate the i-th pre-generated sample output image based on the i-th sample jitter image, the i-th ground truth image corresponding to the i-th sample jitter image, and the noise image.

[0246] The training module is also configured to generate an initial noise image, the size of which is smaller than the size of the i-th sample jitter image; Based on the size of the i-th sample jitter image, the initial noise image is interpolated to obtain the interpolated noise image; The interpolated noise image is randomly scaled to obtain the noise image.

[0247] In summary, by deforming the (i-1)th output image based on optical flow information, spatial alignment between the historical image and the current image is achieved. Furthermore, the fusion weights are determined based on the feature correlation between the (i-1)th deformed image and the i-th jittered image. Since image features can represent image content from richer dimensions, the fusion weights determined based on feature correlation are more representative. When artifacts exist in the (i-1)th deformed image, the feature correlation of the artifact region is low, and the fusion weights can correspondingly reduce the fusion ratio of the (i-1)th deformed image, focusing on fusing the i-th jittered image, thereby enabling the generated i-th output image to effectively remove artifacts.

[0248] Please refer to Figure 12 This diagram illustrates a structural block diagram of a computer device 1200 provided in an exemplary embodiment of this application. The computer device 1200 may be a smartphone, computer, tablet computer, or other terminal or server with storage capabilities. The computer device 1200 includes a processor 1201 and a memory 1202. The processor 1201 includes the image processing apparatus shown in the above embodiment.

[0249] Processor 1201 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 1201 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 1201 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1201 may integrate a GPU, which is responsible for rendering and drawing the content required to be displayed on the screen.

[0250] The memory 1202 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices.

[0251] In some embodiments, the computer device 1200 may also optionally include a peripheral device interface 1203 and at least one peripheral device.

[0252] Those skilled in the art will understand that Figure 12 The structure shown does not constitute a limitation on the computer device 1200 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0253] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0254] On the other hand, embodiments of this application provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, which is loaded and executed by the processor to implement the image processing method provided in the embodiments of this application as described above.

[0255] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the image processing method provided in the embodiments of this application as described above.

[0256] On the other hand, embodiments of this application provide a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, executes the computer instructions, and causes the processor of the computer device to load and execute the image processing methods provided in the above-described method embodiments.

[0257] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An image processing method, characterized in that, The method includes: Based on optical flow information, the (i-1)th output image is deformed to obtain the (i-1)th deformed image. The (i-1)th output image is the image rendered and displayed at the (i-1)th time, where i is a positive integer. Feature extraction is performed on the (i-1)th deformed image and the i-th jittered image to obtain the first image feature corresponding to the (i-1)th deformed image and the second image feature corresponding to the i-th jittered image. The i-th jittered image is obtained by jittering the original image at time i to introduce pixel offset. The image feature refers to the structured information that characterizes the image content from multiple dimensions. The feature correlation coefficient is determined based on the first image features and the second image features; The fusion weights are determined based on the first image features, the second image features, and the feature correlation coefficients. Based on the fusion weights, the (i-1)th deformed image and the ith jittery image are fused to obtain the ith output image.

2. The method according to claim 1, characterized in that, The step of extracting features from the (i-1)th deformed image and the i-th jittered image to obtain the first image feature corresponding to the (i-1)th deformed image and the second image feature corresponding to the i-th jittered image includes: The feature extraction network of the image processing model is used to extract features from the (i-1)th deformed image and the i-th jittered image respectively, so as to obtain the first image feature corresponding to the (i-1)th deformed image and the second image feature corresponding to the i-th jittered image; The step of determining the feature correlation coefficient based on the first image features and the second image features includes: Based on the first image features and the second image features, the correlation coefficient of the features is determined by the correlation determination network of the image processing model; The step of determining the fusion weight based on the first image features, the second image features, and the feature correlation coefficient includes: Based on the first image features, the second image features, and the feature correlation coefficient, the fusion weights are determined by the weight determination network of the image processing model.

3. The method according to claim 2, characterized in that, The step of determining the feature correlation coefficient based on the first image features and the second image features using the correlation determination network of the image processing model includes: Based on the first image features and the second image features, a first autocorrelation feature of the first image features, a second autocorrelation feature of the second image features, and a cross-correlation feature of the first image features and the second image features are generated. The concatenated feature, obtained by concatenating the first autocorrelation feature, the second autocorrelation feature, and the cross-correlation feature, is input into the correlation determination network of the image processing model to obtain the feature correlation coefficient output by the correlation determination network.

4. The method according to claim 2, characterized in that, The step of determining the fusion weights based on the first image features, the second image features, and the feature correlation coefficients through the weight determination network of the image processing model includes: The first image feature and the second image feature are input into the weight determination network of the image processing model to obtain the fusion coefficients output by the weight determination network; The fusion weight is determined based on the fusion coefficient and the feature correlation coefficient.

5. The method according to claim 2, characterized in that, The training process of the image processing model includes: Input the sample jittered image sequence into the image processing model to obtain the predicted output image sequence output by the image processing model; The first model loss is determined based on the predicted output image sequence and the ground truth image sequence corresponding to the sample jitter image sequence; The image processing model is trained based on the first model loss.

6. The method according to claim 5, characterized in that, The weight determination network is also used to determine image residuals, which are used to compensate for the predicted output images in the predicted output image sequence. The method further includes: The sample jitter image sequence is input into the image processing model to obtain the predicted output image sequence and the image residual sequence output by the image processing model; The second model loss is determined based on the predicted output image sequence, the image residual sequence, and the ground value image sequence corresponding to the sample jitter image sequence. The image processing model is trained based on the second model loss.

7. The method according to claim 6, characterized in that, The step of determining the second model loss based on the predicted output image sequence, the image residual sequence, and the ground value image sequence corresponding to the sample jitter image sequence includes: The predicted output image sequence is compensated based on the image residual sequence to obtain the compensated predicted output image sequence. Based on the compensated predicted output image sequence and the ground truth image sequence, determine the image sequence loss; Based on the image residual sequence, determine the residual sequence loss; The second model loss is determined based on the image sequence loss and the residual sequence loss.

8. The method according to claim 7, characterized in that, The step of determining the second model loss based on the image sequence loss and the residual sequence loss includes: During the fixed-weight training phase, the second model loss is determined based on the image sequence loss, the residual sequence loss, and the fixed residual weights of the residual sequence loss. During the dynamic weight training phase, the second model loss is determined based on the image sequence loss, the residual sequence loss, and the dynamic residual weights of the residual sequence loss, wherein the dynamic residual weights gradually increase during the dynamic weight training phase, and the dynamic weight training phase follows the fixed weight training phase.

9. The method according to claim 6, characterized in that, Before training the image processing model based on the second model loss, the method further includes: In the single-step training phase, for the i-th sample jitter image in the sample jitter image sequence, the i-th sample jitter image and the (i-1)-th sample deformed image are input into the image processing model to obtain the i-th predicted output image and the i-th image residual output by the image processing model. The (i-1)-th sample deformed image is obtained by deforming the (i-1)-th sample output image based on optical flow information. The third model loss is determined based on the i-th predicted output image, the i-th image residual, and the i-th ground value image corresponding to the i-th sample jitter image. The image processing model is trained based on the third model loss.

10. The method according to claim 9, characterized in that, The step of determining the third model loss based on the i-th predicted output image, the i-th image residual, and the i-th ground truth image corresponding to the i-th sample jitter image includes: The i-th predicted output image is compensated based on the i-th image residual to obtain the compensated i-th predicted output image. Based on the compensated i-th predicted output image and the i-th ground truth image, determine the image loss; Based on the residual of the i-th image, determine the residual loss; The third model loss is determined based on the image loss and the residual loss.

11. The method according to claim 10, characterized in that, In the single-step training phase, the sample output image is selected from the pre-generated sample output image and the compensated predicted output image.

12. The method according to claim 11, characterized in that, The method further includes: Based on the jittered image of the i-th sample, the ground truth image corresponding to the jittered image of the i-th sample, and the noise image, the output image of the i-th pre-generated sample is generated.

13. The method according to claim 12, characterized in that, The method further includes: Generate an initial noise image, the size of which is smaller than the size of the i-th sample jitter image; Based on the size of the i-th sample jitter image, the initial noise image is interpolated to obtain the interpolated noise image; The interpolated noise image is randomly scaled to obtain the noise image.

14. An image processing apparatus, characterized in that, The device includes: The deformation module is configured to perform deformation processing on the (i-1)th output image based on optical flow information to obtain the (i-1)th deformed image, wherein the (i-1)th output image is the image rendered and displayed at the (i-1)th time, and i is a positive integer. The determining module is configured to extract features from the (i-1)th deformed image and the i-th jittered image to obtain a first image feature corresponding to the (i-1)th deformed image and a second image feature corresponding to the i-th jittered image. The i-th jittered image is obtained by jittering the original image at time i to introduce pixel offset. Image features refer to structured information representing image content from multiple dimensions. A feature correlation coefficient is determined based on the first image feature and the second image feature. A fusion weight is determined based on the first image feature, the second image feature, and the feature correlation coefficient. The fusion module is configured to perform image fusion on the (i-1)th deformed image and the i-th jittery image based on the fusion weights to obtain the i-th output image.

15. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the image processing method as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the image processing method as described in any one of claims 1 to 13.

17. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, which a processor reads from and executes to implement the image processing method as described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Video frame processing method and device

    CN111524166A

  • Image processing method and device, electronic equipment, medium and program product

    CN120976015A