Image processing method and image processing device

By using the high-resolution first image frame to repair the low-resolution second image frame to generate high-resolution images, the problem of high hardware resource consumption in the prior art is solved, and the balance between high refresh rate and high image quality is achieved.

CN120198288APending Publication Date: 2025-06-24LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510228123.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When the existing super-resolution model processes multiple low-resolution image frames, the computing volume is huge, resulting in large hardware resource consumption and the inability to balance refresh rate and image quality.

Method used

By acquiring a high-resolution first image frame and a low-resolution second image frame, the second image frame is repaired using the first image frame to generate a high-resolution image corresponding to the second image frame, reducing the computational amount and hardware resource consumption.

Benefits of technology

It achieves the improvement of image quality while ensuring the refresh rate of game videos, reducing the amount of computing and hardware resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198288A_ABST
    Figure CN120198288A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method and an image processing device. The method comprises the following steps: acquiring a first image frame; according to a mapping relation of image contents between the first image frame and the second image frame, repairing the second image frame by using the first image frame to obtain a high-resolution image corresponding to the second image frame; wherein the resolution of the first image frame is greater than a preset resolution threshold, and the resolution of the second image frame is smaller than the preset resolution threshold; the first image frame and the second image frame meet a similarity condition; and outputting a high-resolution image frame corresponding to the second image frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to an image processing method and an image processing apparatus. Background Art

[0002] In order to improve the refresh rate and the picture quality of video content such as live broadcasts and games, usually a super-resolution model is used to process and dynamically synthesize high-resolution image frames from multiple low-resolution image frames. In this process, the super-resolution model needs to repeatedly search among multiple low-resolution image frames, predict texture details, and generate high-resolution image frames based on the predicted texture details. This method has a huge amount of computation, which will cause a large consumption of hardware resources, limit the refresh rate, and make it impossible to achieve a balance between the refresh rate and the image quality. Summary of the Invention

[0003] The technical solution of this application is implemented as follows:

[0004] An embodiment of this application provides an image processing method, which includes:

[0005] Obtain a first image frame;

[0006] According to the mapping relationship of the image content between the first image frame and the second image frame, use the first image frame to perform restoration processing on the second image frame to obtain a high-resolution image corresponding to the second image frame; wherein, the resolution of the first image frame is greater than a preset resolution threshold, the resolution of the second image frame is less than the preset resolution threshold; the first image frame and the second image frame meet the similarity condition;

[0007] Output a high-resolution image frame corresponding to the second image frame.

[0008] An embodiment of this application provides an image processing apparatus, including:

[0009] An obtaining module, configured to obtain a first image frame;

[0010] A processing module, configured to perform restoration processing on the second image frame by using the first image frame according to the mapping relationship of the image content between the first image frame and the second image frame, to obtain a high-resolution image corresponding to the second image frame; wherein, the resolution of the first image frame is greater than a preset resolution threshold, the resolution of the second image frame is less than the preset resolution threshold; the first image frame and the second image frame meet the similarity condition;

[0011] An output module, configured to output a high-resolution image frame corresponding to the second image frame. Description of the Drawings

[0012] Figure 1 It is a schematic flowchart of an image processing method provided by an embodiment of this application;

[0013] Figure 2 It is a schematic flowchart of an image processing method based on high - resolution image frames provided by an embodiment of the present application;

[0014] Figure 3 It is a schematic diagram of an algorithm framework of a super - resolution model provided by an embodiment of the present application;

[0015] Figure 4 It is a schematic diagram of the composition structure of an image processing device provided by an embodiment of the present application. Detailed implementation manners

[0016] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0017] In order to make the purpose, technical solutions, and advantages of the present application clearer, the present application will be further described in conjunction with the accompanying drawings. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0019] In the following description, "some embodiments\other embodiments" are involved, which describe subsets of all possible embodiments. However, it can be understood that "some embodiments\other embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0020] In the following description, the terms "first\second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0021] In the related art, for the scenario of real-time video output, in order to improve the refresh rate and picture quality of video image frames, there is a method of performing fine rendering on the main objects of interest (such as the foreground) in the picture and rough rendering on the background in the picture. For example, for a real-time output game video, each image frame needs to be rendered frame by frame, and the game engine needs to be modified to distinguish the main characters and secondary characters. Although the amount of calculation is reduced and high refresh rate is satisfied, the image quality cannot be guaranteed. Here, only the game scenario is taken as an example for illustration, and other types of video real-time output scenarios will not be exemplified one by one.

[0022] Based on the problems existing in the related art, the embodiments of the present application provide an image processing method, which can be applied to the processing of game video image frames or other types of video image frames. The present application takes game video image frames as an example for illustration. As Figure 1 shown, it is a schematic flowchart of an image processing method provided by an embodiment of the present application. The method includes the following steps:

[0023] S101. Obtain a first image frame.

[0024] It should be noted that the first image frame can be one or more image frames in a game video. In the case where the first image frame includes multiple ones, each first image frame may be continuous or discontinuous. The first image frame can be an image frame with a resolution greater than a preset resolution threshold, that is, the first image frame can be a high-resolution image frame.

[0025] In some embodiments, the first image frame can be obtained by real-time rendering of an electronic device running a game, or can be obtained by the electronic device running the game from the game engine after rendering by the game engine during the running of the game.

[0026] S102. According to the mapping relationship of the image content between the first image frame and the second image frame, use the first image frame to perform restoration processing on the second image frame to obtain a high-resolution image corresponding to the second image frame.

[0027] In some embodiments, the resolution of the second image frame is less than a preset resolution threshold, that is, the second image frame can be a low-resolution image frame. By using the high-resolution first image frame to perform restoration processing on the low-resolution second image frame, it is possible to supplement the detailed content in the second image frame with the detailed content in the first image frame, so as to obtain a high-resolution image corresponding to the second image frame.

[0028] In some embodiments, the first image frame and the second image frame meet the similarity condition, which may be that the similarity between the first image frame and the second image frame is greater than the similarity threshold. The similarity threshold may be 75%, 80%, etc. In practice, the first image frame and the second image frame may be image frames with a temporal relationship. For example, the first image frame and the second image frame may be adjacent image frames in the same game video, and the first image frame may be the previous image frame or the next image frame of the second image frame; the first image frame may also be the image frame separated from the second image frame by two, three,..., N image frames, where N is less than the total number of image frames in the game video.

[0029] In some embodiments, the mapping relationship of the image content between the first image frame and the second image frame may be the mapping relationship of corresponding image elements, image details, etc. in the first image frame and the second image frame. For example, the sky background included in the first image frame has a mapping relationship with the sky background included in the second image frame; the facial contour of the game character included in the first image frame has a mapping relationship with the picture contour of the game character included in the second image frame. The mapping relationship between the first image frame and the second image frame herein is only an exemplary illustration, and the present application is not limited thereto.

[0030] S103. Output the high-resolution image frame corresponding to the second image frame.

[0031] In some embodiments, after obtaining the high-resolution image frame corresponding to the second image frame, the second image frame may be output to the display interface of the electronic device running the game for the user to view. In implementation, the first image frame may also be output and displayed, and according to the temporal relationship between the first image frame and the second image frame, the high-resolution images corresponding to the first image frame and the second image frame may be output to the display interface of the electronic device in sequence.

[0032] It should be noted that for the same game video, there may be multiple second image frames that need to be repaired, and the first image frames having a mapping relationship with each second image frame may be different or partially the same. The repair process for each second image frame may be performed in sequence according to the temporal relationship of each second image frame, or two or more second image frames may be repaired simultaneously.

[0033] In an embodiment of the present application, a first image frame is obtained; according to the mapping relationship of the image content between the first image frame and the second image frame, the second image frame is repaired using the first image frame to obtain a high-resolution image corresponding to the second image frame; wherein, the resolution of the first image frame is greater than a preset resolution threshold, and the resolution of the second image frame is less than the preset resolution threshold; the first image frame and the second image frame satisfy a similarity condition; the high-resolution image frame corresponding to the second image frame is output. Thus, since there is a mapping relationship between the image content of the high-resolution first image frame and the low-resolution second image frame, and the first image frame and the second image frame satisfy the similarity condition, therefore, by repairing the second image frame using the first image frame, a high-resolution image frame corresponding to the second image frame can be directly generated through texture detail transfer without generating new image details, thereby reducing the computational amount and avoiding excessive consumption of hardware resources, and the refresh rate of the game video can be guaranteed while improving the image quality.

[0034] In some embodiments of the present application, during the process of repairing the second image frame using the first image frame, the image of the first region of the first image frame can be covered over the image of the second region of the second image frame.

[0035] Wherein, the image of the first region of the first image frame and the image of the second region of the second image frame have a content mapping relationship. The positions of the first region of the first image frame and the second region of the second image frame may be different. For example, the first region may be at the lower left of the first image frame, and the second region may be at the lower right of the second image frame.

[0036] In some embodiments, the image of the first region in the first image frame and the image of the second region in the second image frame may be the same or similar. For example, if the image of the first region is a game character, then the image of the second region may also be the same game character.

[0037] In some embodiments, covering the image of the first region of the first image frame over the image of the second region of the second image frame may be replacing the image of the second region of the second image frame with the image of the first region of the first image frame, or superimposing the image of the first region of the first image frame and the image of the second region of the second image frame.

[0038] It can be understood that by covering the image of the first region of the first image frame over the image of the second region of the second image frame that has a content mapping relationship with it, the supplement of the details of the second image frame based on the first image frame can be directly realized, the speed of repairing the second image frame is accelerated, and the refresh rate of the game video is guaranteed.

[0039] In some embodiments of the present application, obtaining the first image frame, that is, step S101 above can be implemented by the following steps S1011 to S1012, and each step will be described separately below.

[0040] S1011. Receive a plurality of first candidate image frames.

[0041] In some embodiments, the first candidate image frames can be multiple image frames in a game video provided by a game engine, and the first candidate image frames can be image frames with an image resolution less than a preset resolution; in other embodiments, the first candidate image frames can also be image frames with an image resolution greater than or equal to the preset resolution. In this case, the high-resolution first candidate image frames can be obtained by rendering by the game engine.

[0042] S1012. Determine the first image frame from the plurality of first candidate image frames.

[0043] In some embodiments, the first image frame determined from the plurality of first candidate image frames may include one or more. When the first candidate image frames are less than the preset resolution, a first target candidate image frame that meets the key frame condition can be determined from the plurality of first candidate image frames, and then the first target candidate image frame is rendered to obtain a high-resolution first image frame.

[0044] Here, the key frame condition can be preset according to the game content, scene type, etc. in the image frame. For example, the first target candidate image frame that meets the key frame condition can be a first candidate image frame that contains image content that can attract the user's special attention. The image content that can attract the user's special attention can include game characters, skill special effects, etc.; the first target candidate image frame that meets the key frame condition can also be an image frame determined from the first candidate image frames according to the sampling frequency corresponding to the scene type. Among them, the first target candidate image frame can be considered as a low-resolution image frame that has guiding or reference significance for the repair of the second image frame.

[0045] In other embodiments, when the image resolution of the first candidate image frames is greater than the preset resolution threshold, the first image frame that meets the key frame condition can also be directly selected from the multiple first candidate image frames.

[0046] It can be understood that by determining the first image frame from the plurality of first candidate image frames, it is convenient to subsequently repair the second image that has a content mapping relationship with the first image frame to obtain a high-resolution image corresponding to the second image frame.

[0047] In some embodiments of the present application, in the process of determining the first image frame from a number of first candidate image frames, the first image frame with attribute information characterizing a high-resolution image frame can be determined from the number of first candidate image frames based on the attribute information corresponding to the first candidate image frames.

[0048] Among them, the attribute information corresponding to the first candidate image frame may include a specific event in the first candidate image frame, and the specific event may indicate that the first candidate image frame can be used as a high-resolution image frame or a base image frame for generating a high-resolution image frame. The specific event may be an event that the user particularly pays attention to. For example, the specific event may include a close-up of a person, a scene transition, the start of a battle, the release of a special move, etc.

[0049] In some embodiments, the selection rule of the first image frame can be preset. For example, a low-resolution image frame containing a specific event can be determined as a candidate image frame to be rendered into a high-resolution image, or a high-resolution image frame containing a specific event can be directly used as the high-resolution image frame.

[0050] In some embodiments, the attribute information of the first candidate image frame can be analyzed to determine whether a specific event such as a close-up of a person or a scene transition is included in the first candidate image frame. If it is included, the first candidate image frame can be rendered to obtain the first image frame; in the case where the image resolution of the first candidate image frame is greater than the preset resolution, if the first candidate image frame includes a specific event, the first candidate image frame can be used as the first image frame.

[0051] It can be understood that determining the first image frame from a number of first candidate image frames according to the attribute information of the first candidate image frames realizes the mutual association between the selection of high-resolution image frames and the attributes of the image frames, making the finally determined first image frame change dynamically according to the content of the image frame, and ensuring the continuity of the content of the finally obtained high-resolution second image frame.

[0052] In some embodiments, in the process of determining the first image frame from a number of first candidate image frames, the first target candidate image frame can also be determined from a number of first candidate image frames with time sequence based on the target offset frame number, and the first target candidate image frame can be rendered to obtain the first image frame.

[0053] Among them, the target offset frame number can be a preset sampling frame number interval. For example, the target offset frame number can be 3 frames, 5 frames, 10 frames, etc. The target offset frame number can be determined according to the total number of frames of the game corresponding to the first candidate image, and the target offset frame number is less than the total number of frames of the game. The first candidate image frames with time sequence can be multiple first candidate image frames that are continuous in time.

[0054] In some embodiments, when the target offset number of frames is determined, a first target candidate image frame can be collected from a plurality of first candidate image frames with time series at intervals of the target offset frames, and the first target candidate image frame can be rendered to obtain a first image frame.

[0055] It can be understood that based on the target offset number of frames, the first target candidate image frame can be quickly determined from a plurality of first candidate image frames with time series, improving the determination efficiency of the first image frame.

[0056] In some embodiments, in the process of determining the first image frame from a plurality of first candidate image frames, the first target candidate image frame can also be determined from a plurality of first candidate image frames with time series based on the target acquisition frequency, and the first target candidate image frame can be rendered to obtain the first image frame.

[0057] Among them, the target sampling frequency can be determined in advance according to the scene type and application type of the game corresponding to the first candidate image frame. The scene type can include battle scenes and non-battle scenes; the application type can include real-time battle games, live games (or competitive commentary games), etc.

[0058] Here, the target sampling frequency corresponding to the battle scene can be greater than the target sampling frequency corresponding to the non-battle scene; the target sampling frequency corresponding to the real-time battle game can be greater than the target sampling frequency corresponding to the live game. In this way, the first target candidate image frame determined from a plurality of first candidate image frames with time series based on the target sampling frequency determined in advance according to the scene type and application type of the game can be an image frame that attracts special attention or is of interest to the user, so that after the first image frame finally rendered based on the first target candidate image frame is used to repair the second image frame, a high-resolution image frame that meets the user's expectations can be obtained.

[0059] In some embodiments, the scene type of a plurality of first candidate images with time series can be recognized, and according to the recognized result of the scene class and the pre-established correspondence between the game scene type and the sampling frequency, the target sampling frequency can be determined; or, the application type of the game corresponding to the first candidate image with time series can be obtained, and according to the application type and the pre-established correspondence between the game application type and the sampling frequency, the target sampling frequency can be determined.

[0060] In some embodiments, the target sampling frequency defines the sampling time interval of the first target candidate image frame. According to the target sampling frequency, the first target candidate image frame can be selected from the first candidate image frames with time series. Subsequently, by rendering the first target candidate image frame, a high-resolution first image frame can be obtained.

[0061] In some embodiments, in the process of determining the first image frame from a number of first candidate image frames, the first target candidate image frame that meets the image feature conditions can also be determined from the number of first candidate image frames based on the image feature information corresponding to the number of first candidate image frames, and the first target candidate image frame is rendered to obtain the first image frame.

[0062] Among them, the image feature information may include the image content of the first candidate image frame, the image complexity, the degree of content change of adjacent first candidate image frames, etc. The image content may be a pre-set game character, skill special effect, etc.; the image complexity can be determined according to the color, texture, edges, etc. of the image frame. For example, the complexity of the image frame can be measured by statistical features, information entropy, self-similarity, etc. determined according to the color, texture, edges, etc. of the image frame; the degree of content change of adjacent first candidate image frames can measure the difference degree of the image content in adjacent first candidate image frames.

[0063] In some embodiments, it is possible to determine whether the first candidate image meets the image feature conditions according to the image content, image complexity, degree of content change of adjacent first candidate image frames, etc., and when it is determined that the first candidate image meets the image feature conditions, the first candidate image is rendered to obtain the first image.

[0064] It can be understood that by determining the first target candidate image frame that meets the image feature conditions from a number of first candidate image frames based on the image feature information corresponding to the first candidate image frames, the first image frame obtained by rendering based on the first target candidate image frame can include image features that meet the image feature conditions, and the image content that the user is concerned about can be obtained by performing restoration processing on the second image frame according to the first image frame subsequently, improving the user experience.

[0065] In some embodiments, if in the first operating resource environment, the target sampling frequency is the first acquisition frequency; if in the second operating resource environment, the target sampling frequency is the second acquisition frequency.

[0066] Among them, the operating resource environment may include the current hardware resource state of the electronic device on which the game corresponding to the first candidate image runs, such as the remaining battery power of the electronic device, the load state of the processor, etc.

[0067] In some embodiments, the first operating resource environment is different from the second operating resource environment. For example, if the first operating resource environment indicates that the current battery of the electronic device is in a low power state and the processor is in a high load state, then the second operating resource environment indicates that the current battery of the electronic device is in a fully charged state and the processor is in a low load state; if the first operating resource environment indicates that the current battery of the electronic device is in a fully charged state and the processor is in a low load state, then the second operating resource environment indicates that the current battery of the electronic device is in a low power state and the processor is in a high load state.

[0068] Further, in the case where the first operating resource environment indicates that the current battery of the electronic device is in a low power state and the processor is in a high load state, and the second operating resource environment indicates that the current battery of the electronic device is in a fully charged state and the processor is in a low load state, the first acquisition frequency may be less than the second acquisition frequency; conversely, the first acquisition frequency may be greater than the second acquisition frequency.

[0069] It can be understood that when the electronic device is in the first operating resource environment, the first acquisition frequency is set as the target sampling frequency; and when the electronic device is in the second operating resource environment, the second acquisition frequency is set as the target sampling frequency, realizing the automatic adjustment of the target sampling frequency in different operating resource environments, which can make the finally determined target sampling frequency adapt to the current operating resource environment of the electronic device, so as to ensure the normal operation of the electronic device while increasing the rendering frequency of the first target candidate image frame.

[0070] In some embodiments of the present application, in the process of rendering the first target candidate image frame to obtain the first image frame, the target area in the display screen of the electronic device on which the game corresponding to the first candidate image frame runs may be determined; the image content corresponding to the target area in the first target candidate image frame is obtained; the image content of the target area is rendered to obtain the first image content; and the first image content and the second image content are fused to obtain the first image frame.

[0071] Among them, the target area may be the area that the user pays attention to or is interested in. For example, in the process of rendering the first target image, the area that the user is interested in or pays attention to may be detected, so as to determine the target area in the display screen of the electronic device. The second image content is the image content of the other areas in the first target candidate image frame except the target area.

[0072] It should be noted that the resolution of the first target candidate image frame is less than the preset resolution threshold, that is, the first target candidate image frame is a low-resolution image frame, and the second image content is the unrendered image content in the first target candidate image frame. Therefore, the second image content is still low-resolution.

[0073] In some embodiments, after determining the image content in the target area, only the image content in the target area may be rendered to obtain first image content with rich details. Then, the first image content and the second image content are subjected to a fusion process, such as stitching the first image content and the second image content, so as to obtain a first image frame.

[0074] It can be understood that by rendering the image content corresponding to the target area in the display screen of the electronic device in the first target candidate image, instead of rendering the image content in other areas except the target area in the first target candidate image frame, the image content concerned by the user can be retained, and at the same time, the rendering amount can be reduced and the efficiency of image processing can be improved.

[0075] In some embodiments of the present application, the first target candidate image frame meeting the image feature conditions may be a first candidate image frame whose contained image content has a target similarity with the target image content. Among them, the target image content may be the image content that the user particularly concerns or the content that can arouse the user's interest. For example, it may be a game character, a scene boundary, etc. The target similarity may be 95%, 98%, etc.

[0076] The electronic device may detect characters, skills, scene boundaries, etc. in the first candidate image frame in real time through artificial intelligence technology. For example, if it is detected that image content such as an explosion special effect, intense actions, and a subtitle prompt "Boss battle starts" appears in the picture of the first candidate image frame, the similarity between the image content and the target image content may be calculated, and when it is determined that the calculated similarity reaches the target similarity, the first candidate image frame may be determined as the first target candidate image frame.

[0077] It can be understood that by comparing the similarity between the content contained in the first candidate image frame and the target image content, the determined first target candidate image frame can contain the target image content concerned by the user, ensuring that the high-resolution image obtained after repairing the second image frame retains the key or concerned content.

[0078] In some embodiments, the first target candidate image frame meeting the image feature conditions may also be a first candidate image frame whose contained image content has an image complexity greater than an image complexity threshold. Among them, the complexity threshold may be preset to 60%, 70%, 75%, etc. The complexity of the first image frame may be determined according to the color, texture, edges, etc. of the first candidate image frame. For example, the complexity of the first image frame may be measured by statistical features, information entropy, self-similarity, etc. determined according to the color, texture, edges, etc. of the first candidate image frame.

[0079] The electronic device can analyze the scene texture richness, edge density, and color change in the first candidate image frame, calculate the complexity of the first candidate image frame based on the scene texture richness, edge density, and color change, and determine the first candidate image frame as the first target candidate image frame when it is determined that the complexity is greater than the image complexity threshold.

[0080] It can be understood that comparing the image complexity of the image content included in the first candidate image frame with the image complexity threshold can make the determined first target candidate image frame contain more details of the image content, so that the high-resolution image corresponding to the second image frame can be restored more accurately.

[0081] In some other embodiments, the first target candidate image frame that meets the image feature conditions can also be the first candidate image frame whose degree of image content transformation between the first candidate image frame adjacent in time series is greater than the transformation degree threshold. Among them, the degree of image content transformation can measure the difference degree of the image content in adjacent first candidate image frames, and the transformation degree threshold can be preset, such as 50%, 60%, etc. When determining...

[0082] The electronic device can perform optical flow estimation based on the current first candidate image frame and the first candidate image frame adjacent to it, determine the inter-frame difference degree according to the optical flow estimation result, determine the degree of content transformation between the current first candidate image frame and the adjacent first candidate image frame based on the inter-frame difference degree, and when it is determined that the degree of content transformation is greater than the transformation degree threshold, the current first candidate image frame can be determined as the first target candidate image frame.

[0083] It can be understood that comparing the degree of image content transformation between the first candidate image frame and other first candidate image frames adjacent in time series with the transformation degree threshold can make the determined first target candidate image frame capture the dynamically changing image content and ensure that the high-resolution image frame corresponding to the final second image frame has richer image content.

[0084] In some embodiments of the present application, according to the mapping relationship of the image content between the first image frame and the second image frame, the second image frame is repaired using the first image frame to obtain the high-resolution image corresponding to the second image frame. That is, the above step S102 can be implemented through the following step S1021, and the following will describe this step.

[0085] S1021: Input the first image frame and the second image frame into the target intelligent engine, so that the target intelligent engine repairs the second image frame using the first image frame according to the mapping relationship of the image content between the first image frame and the second image frame, and obtains the high-resolution image corresponding to the second image frame.

[0086] Among them, the target intelligent engine can be an artificial intelligence engine, and the artificial intelligence engine can include a trained super-resolution model. The first image frame and the second image frame are input into the super-resolution model, and the super-resolution model can analyze the image content of the first image frame and the second image frame, determine the mapping relationship of the image content between the two, and according to this mapping relationship, use the image content of the first image frame to perform restoration processing on the image content of the second image frame, and finally output the high-resolution image corresponding to the second image frame.

[0087] It can be understood that by inputting the first image frame and the second image frame into the target intelligent engine and using the target intelligent engine to perform restoration processing on the second image frame based on the first image frame, the restoration result of the second image frame can be obtained quickly.

[0088] In some embodiments of the present application, the target intelligent engine includes a first feature extraction module, a second feature extraction module, and an attention fusion module. Based on this, the implementation method for the target intelligent engine to perform restoration processing on the second image frame using the first image frame according to the mapping relationship of the image content between the first image frame and the second image frame to obtain the high-resolution image corresponding to the second image frame may include: using the first feature extraction module to extract features from the first image frame to obtain a first image feature matrix; and using the second feature extraction module to extract features from the second image frame to obtain a second image feature matrix; using the attention fusion module to determine, based on the additional information of the second image frame, the second image features in the second image feature matrix that have a mapping relationship with the first image features in the first image feature matrix, and performing superposition processing on the first image features and the corresponding second image features to obtain the high-resolution image frame corresponding to the second image frame.

[0089] Among them, the additional information includes the temporal information between the first image frame and the second image frame, and / or the image information of the second image frame. The temporal information can be the motion vector and optical flow information of the first image frame and the second image frame; the image information can be depth information, semantic labels, region of interest indicators, etc. The additional information of the second image frame can be output by the game engine or determined by the target intelligent engine itself.

[0090] In some embodiments, the attention fusion module has an attention mechanism, and this attention mechanism can focus on the influence of the additional information of the second image frame on the mapping relationship between the image features in the second image feature matrix and the image features in the first image feature matrix.

[0091] In some embodiments, the additional information may include motion feature information between the first image frame and the second image frame, as well as additional feature information of the second image frame. Therefore, based on the additional feature information, the attention fusion module can more accurately and quickly determine the second image features in the second image feature matrix that have a mapping relationship with the first image features in the first image feature matrix.

[0092] In some embodiments, the superimposing process of the first image features and the corresponding second image features may be to replace the second image features in the second image frame with the first image features, or to simply add the first image features and the second image features or perform a weighted sum, so as to obtain the updated image features corresponding to the original second image features in the second image frame.

[0093] It can be understood that based on the additional information of the second image frame, the attention fusion module can accurately and quickly determine the content mapping relationship between the first image frame and the second image frame, thereby accelerating the repair speed of the second image frame and avoiding content distortion of the high-resolution image frame of the finally obtained second image frame.

[0094] In the embodiments of the present application, a high-resolution image frame is obtained; according to the mapping relationship of the image content between the first image frame and the second image frame, the second image frame is repaired using the first image frame to obtain a high-resolution image corresponding to the second image frame; wherein, the resolution of the first image frame is greater than a preset resolution threshold, and the resolution of the second image frame is less than the preset resolution threshold; the first image frame and the second image frame satisfy the similarity condition; the high-resolution image frame corresponding to the second image frame is output. In this way, since there is a mapping relationship between the image content of the high-resolution first image frame and the low-resolution second image frame, and the first image frame and the second image frame satisfy the similarity condition, therefore, by using the first image frame to repair the second image frame, a high-resolution image frame corresponding to the second image frame can be directly generated through texture detail transfer without generating new image details, thereby reducing the calculation amount and avoiding excessive consumption of hardware resources, and the refresh rate of the game video can be guaranteed while improving the image quality.

[0095] Next, the implementation process of the application embodiments in an actual application scenario will be introduced.

[0096] As Figure 2 shown, it is a schematic flowchart of an image processing method based on a high-resolution image frame provided by the present application. The method includes:

[0097] S201. Obtain a high-resolution image frame (equivalent to the "first image frame" in other embodiments) and a low-resolution image frame (equivalent to the "second image frame") in the game video.

[0098] A low - resolution image frame can be a game image frame output by a game engine. The low - resolution image frame is an image frame that needs its resolution to be enhanced. The low - resolution image frame provides the basic structure and information of the image, but lacks high - resolution details and needs to enhance the details from it.

[0099] A high - resolution image frame can be a game image frame obtained after rendering by a game engine, or it can be obtained after the electronic device running the game gets a low - resolution version of the game image frame from the game engine and then renders it. The high - resolution image frame retains a high level of details and can guide the process of enhancing the resolution of the low - resolution image frame.

[0100] In some embodiments, by modifying the functions of the game engine, the game engine can identify and respond to specific events in the game to intelligently select when to render a high - resolution image frame for use as a reference benchmark for subsequent low - resolution frames.

[0101] Generally, a high - resolution image frame can be automatically selected based on content changes, such as scene transitions or significant visual changes. Specific events are identified by analyzing the differences between consecutive frames, such as pixel differences, changes in color histograms, or changes in other image features. It is also possible to directly obtain status information from the game engine, such as scene numbers, specific plot identifiers, etc., and identify specific events based on this status information.

[0102] Here, specific events can include scene transitions (the game switches from one environment to another, such as from indoors to outdoors, usually accompanied by a large amount of visual information changes. Rendering a high - resolution image frame can capture and emphasize the details and atmosphere of the new scene), plot climaxes (such as battles, explosions, or important conversations. These moments may require higher image quality to enhance the player's immersion), user interactions (specific operations of the player, such as using special abilities or triggering certain game mechanisms, may cause significant visual changes), character close - ups (when there are character interactions, plot progressions, or important conversations, a close - up shot of the character is taken and a high - resolution image frame is rendered to enhance the character's expressiveness and emotional transmission), the start of a battle (a battle scene usually contains a large amount of action and details. Rendering a high - resolution image frame can capture and highlight these exciting moments), unleashing a powerful move (when a character releases a special skill or a powerful move, rendering a high - resolution image frame can show the magnificence and shock of the skill effect), etc.

[0103] In some embodiments, a fixed offset can also be selected at a fixed refresh rate (equivalent to the "target offset frame number" in other embodiments). For example, high - resolution image frames are rendered at time intervals such as every 16 frames, 30 frames, etc., and other image frames are not rendered and serve as low - resolution image frames.

[0104] In some embodiments, a deep learning model, such as a convolutional neural network, can be pre-trained by manually calibrating appropriate key frames, and this pre-trained deep learning model can be used to analyze the current and previous frames to predict which frames are key frames (high-resolution image frames), thereby dynamically selecting key frames.

[0105] In some embodiments, different key frame selection rules may also be determined based on the game scene:

[0106] For high-frequency interactions or action scenes, where the player or opponent characters often move at high speed and the battle scenes are intensive, keyframe selection can be more frequent. For example, after entering the battle mode, the keyframe output interval can be shortened (for example, a keyframe is output every 1 second, but a keyframe is output every 0.5 seconds during the battle), or instant triggering can be performed based on events such as character health and skill release to ensure the clarity of key battle scenes.

[0107] For non-combat scenes (open world, exploration, story cutscenes, etc.), the characters in the screen move or interact relatively slowly, and the scene changes are not as compact as battles. The keyframe output frequency can be moderately reduced to reduce rendering and computing loads. During story dialogues or important shots, high-resolution keyframes can be inserted to highlight the story atmosphere or character expression details.

[0108] For live games or competitive commentary modes, unlike ordinary player games, this scene may include a global perspective, multi-perspective switching or continuous commentary screen. When the "scene switch" or "perspective switch" occurs, a key frame can be output immediately to make the picture of the new perspective quickly become clear; if the live broadcast screen is relatively static (the commentator does not switch lenses at high speed in the spectator mode), a relatively low frequency of key frame rendering can be maintained.

[0109] S202. Input the high-resolution image frame and the low-resolution image frame into a super-resolution model to use the high-resolution image frame to perform super-resolution processing on the low-resolution image frame (equivalent to "using the first image frame to perform repair processing on the second image frame" in other embodiments) to obtain a high-resolution image corresponding to the low-resolution image frame.

[0110] In some embodiments, temporal motion information and additional information of the low-resolution image frames may also be input into the super-resolution model to better determine image content in the high-resolution image that has a mapping relationship with image content in the low-resolution image.

[0111] Temporal motion information can include motion vectors, optical flow, or other forms of temporal data, which are often used for motion blur effects. The motion vector for each image pixel describes the distance that image pixel moves from one frame to the next. This can be output from the engine as a motion vector map.

[0112] The additional information may include a depth map, semantic labels, and region-of-interest indicators. The depth map provides information about the distances between various objects in the scene and the camera to the super-resolution model, which helps the super-resolution model understand the three-dimensional structure of the scene, thereby more precisely restoring image details, especially at object edges and overlapping regions; the semantic labels provide information about the category to which each pixel in the image belongs, such as people, vehicles, buildings, etc. This information enables the super-resolution model to more intelligently process different types of textures and edges because the restoration of textures and edges of different object types may require different processing methods; the region-of-interest indicators mark important or areas that require special attention in the image. These indicators can help the super-resolution model focus resources and attention on these areas, ensuring that these areas can receive priority and more refined processing during super-resolution processing.

[0113] In some embodiments, the super-resolution model may be a deep learning network model, such as Figure 3 As shown in the schematic diagram of the algorithm framework of a super-resolution model provided by an embodiment of the present application, the super-resolution model includes a first feature extraction module 301, a second feature extraction module 302, a motion information processing module 303, an attention fusion module 304, and an upsampling module 305. The first feature extraction module 301 is used to extract features from the input first high-resolution image frame 306; the second feature extraction module 302 is used to extract features from the input first low-resolution image frame 307; the motion information processing module 303 is used to perform a spatial transformation on the input motion vector / optical flow information / depth map 308. The motion vector / optical flow information / depth map 308 can be used as an input to the super-resolution model or not. When it is not used as an input to the super-resolution model, the motion information processing module 303 may not be enabled; the attention fusion module 304 is used to fuse the image features after feature extraction of the two image frames based on the input motion vector / optical flow information / depth map 308; the upsampling module 305 can perform interpolation processing on the fused image feature matrix to obtain the high-resolution image 309 corresponding to the first low-resolution image frame.

[0114] In some embodiments, the super-resolution model can be pre-trained. Multiple displays with different resolutions can be simulated on a mobile phone to obtain training data (including high-resolution frames and low-resolution frames) at the same time. The obtained high-resolution frames and low-resolution frames are input into the network to obtain the high-definition pictures corresponding to the low-resolution frames, and compared with the marked high-resolution frames. The loss function value is determined through the mean square error, the perceptual loss based on the feature maps of the pre-trained network, and the adversarial loss when improving the perceptual quality based on the generative adversarial network. After training, the super-resolution model can be applied to various types of games.

[0115] In the game scenario of this solution, since the key frames themselves are already of high resolution, which is equivalent to providing the "most accurate detail reference" for the super-resolution model, the super-resolution model only needs to "learn how to map the texture of the high-resolution key frames to the corresponding positions of the current low-resolution image frames" more, rather than repeatedly searching and inferring the missing high-frequency information among multiple frames. In this way, the dependence on multiple frames during training is greatly reduced, and it is also easier to adapt to new game scenarios or perform quasi-real-time processing in the engine.

[0116] The image processing method based on high-resolution image frames provided by this application outputs high-resolution image frames at specific events (or fixed cycles, etc.), and most other image frames are rendered at low resolution and then quickly enlarged by means of a super-resolution model, balancing the game frame rate and visual quality; and through the high-quality reference of key frames, even in the case of limited hardware resources, high-definition visual effects can be achieved in the game. Compared with rendering high-resolution images throughout the process, the processing requirements can be effectively reduced, and the fluency of the game can be maintained.

[0117] This application also provides an image processing device. Figure 4 As shown in the schematic composition structure diagram of an image processing device provided by an embodiment of this application, Figure 4 as shown, the image processing device 400 includes:

[0118] An acquisition module 401, configured to acquire a first image frame;

[0119] A processing module 402, configured to perform restoration processing on the second image frame by using the first image frame according to the mapping relationship of the image content between the first image frame and the second image frame, to obtain a high-resolution image corresponding to the second image frame; wherein, the resolution of the first image frame is greater than a preset resolution threshold, the resolution of the second image frame is less than the preset resolution threshold; the first image frame and the second image frame satisfy a similarity condition;

[0120] An output module 403, configured to output a high-resolution image frame corresponding to the second image frame.

[0121] In some embodiments, the processing module 402 is further configured to cover the image of the first region of the first image frame to the image of the second region of the second image frame;

[0122] wherein, the image of the first region of the first image frame has a content mapping relationship with the image of the second region of the second image frame.

[0123] In some embodiments, the acquisition module 401 includes:

[0124] A first receiving sub-module, configured to receive a plurality of first candidate image frames;

[0125] A first determination sub-module, configured to determine a first image frame from a plurality of the first candidate image frames.

[0126] In some embodiments, the first determination sub-module is further configured to:

[0127] Based on the attribute information corresponding to the first candidate image frames, determine a first image frame from a plurality of the first candidate image frames, where the attribute information characterizes a high-resolution image frame;

[0128] Based on a target offset number of frames, determine the first target candidate image frame from a plurality of the first candidate image frames with time sequence, and perform rendering processing on the first target candidate image frame to obtain the first image frame;

[0129] Based on a target acquisition frequency, determine the first target candidate image frame from a plurality of the first candidate image frames with time sequence, and perform rendering processing on the first target candidate image frame to obtain the first image frame;

[0130] Based on the image feature information corresponding to a plurality of the first candidate image frames, determine a first target candidate image frame that meets the image feature conditions from a plurality of the first candidate image frames, and perform rendering processing on the first target candidate image frame to obtain the first image frame.

[0131] In some embodiments, the first determination sub-module is further configured to: if in a first operating resource environment, determine the target sampling frequency as a first acquisition frequency; if in a second operating resource environment, determine the target sampling frequency as a second acquisition frequency.

[0132] In some embodiments, the first determination sub-module is further configured to: determine a target area in a display screen of an electronic device on which a game corresponding to the first candidate image frame runs; obtain image content corresponding to the target area in the first target candidate image frame; render the image content of the target area to obtain first image content; perform fusion processing on the first image content and second image content to obtain the first image frame; the second image content is image content of other areas in the first target candidate image frame except the target area.

[0133] In some embodiments, the image feature conditions select one of the following combinations: the image content included has a target similarity with target image content; the image complexity of the image content included is greater than an image complexity threshold; the degree of change in image content between the first candidate image frames adjacent in time sequence is greater than a change degree threshold.

[0134] In some embodiments, the processing module 402 is further configured to input the first image frame and the second image frame into a target intelligent engine, so that the target intelligent engine performs restoration processing on the second image frame by using the first image frame according to the mapping relationship of the image content between the first image frame and the second image frame, and obtains a high-resolution image corresponding to the second image frame.

[0135] In some embodiments, the target intelligent engine includes a first feature extraction module, a second feature extraction module, and an attention fusion module; the processing module 402 is further configured to:

[0136] Extract features from the first image frame by using the first feature extraction module to obtain a first image feature matrix; and extract features from the second image frame by using the second feature extraction module to obtain a second image feature matrix;

[0137] Use the attention fusion module to determine, based on the additional information of the second image frame, the second image features in the second image feature matrix that have a mapping relationship with the first image features in the first image feature matrix, and perform superposition processing on the first image features and the corresponding second image features to obtain a high-resolution image frame corresponding to the second image frame;

[0138] Wherein, the additional information includes the timing information between the first image frame and the second image frame, and / or the image information of the second image frame.

[0139] It should be noted that the description of the virtual device in the embodiments of the present application is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments, so details are not described herein. For the technical details not disclosed in this embodiment, please refer to the description of the method embodiments of the present application for understanding.

[0140] The present application provides an electronic device, which includes a central processing unit, a neural network processor, and a display screen. The central processing unit acquires a first image frame and a second image frame, and sends the first image frame and the second image frame to the neural network processor; the neural network processor performs restoration processing on the second image frame by using the first image frame according to the mapping relationship of the image content between the first image frame and the second image frame, and obtains a high-resolution image corresponding to the second image frame; wherein, the resolution of the first image frame is greater than a preset resolution threshold, and the resolution of the second image frame is less than the preset resolution threshold; the first image frame and the second image frame meet the similarity condition; the central processing unit also outputs the high-resolution image frame corresponding to the second image frame to the display screen for display.

[0141] It should be noted that the description of the electronic device in the embodiments of the present application is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. Therefore, it will not be elaborated here. For the technical details not disclosed in this embodiment, please refer to the description of the method embodiments of the present application for understanding.

[0142] It should be noted that in the embodiments of the present application, if the above image processing method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the relevant solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or an electronic device, etc.) to execute all or part of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0143] It should be noted that in this article, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitation, the element defined by the statement "including at least one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element.

[0144] In the several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the discussed components can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0145] In addition, in each embodiment of the present application, each functional unit can be entirely integrated into one processing unit, or each unit can be separately regarded as one unit, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0146] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROMs, magnetic disks, or optical discs.

[0147] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a product to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROMs, magnetic disks, or optical discs.

[0148] The above is only the implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image processing method, comprising: Acquire a first image frame; According to the mapping relationship between the image contents of the first image frame and the second image frame, the second image frame is repaired by using the first image frame to obtain a high-resolution image corresponding to the second image frame; wherein the resolution of the first image frame is greater than a preset resolution threshold, and the resolution of the second image frame is less than the preset resolution threshold; and the first image frame and the second image frame meet a similarity condition; Output a high-resolution image frame corresponding to the second image frame.

2. The method according to claim 1, wherein the step of performing restoration processing on the second image frame by using the first image frame comprises: Overlaying the image of the first image frame located in the first area with the image of the second image frame located in the second area; The image of the first image frame located in the first area and the image of the second image frame located in the second area have a content mapping relationship.

3. The method according to claim 1, wherein acquiring the first image frame comprises: receiving a plurality of first candidate image frames; A first image frame is determined from a plurality of first candidate image frames.

4. The method according to claim 3, wherein determining the first image frame from the plurality of first candidate image frames comprises one of the following: Based on the attribute information corresponding to the first candidate image frame, determining a first image frame whose attribute information represents a high-resolution image frame from among the plurality of first candidate image frames; Based on the target offset frame number, determining a first target candidate image frame from a plurality of first candidate image frames having a time sequence, and rendering the first target candidate image frame to obtain the first image frame; Based on the target acquisition frequency, determining a first target candidate image frame from a plurality of first candidate image frames having a time sequence, and rendering the first target candidate image frame to obtain the first image frame; Based on the image feature information corresponding to the first candidate image frames, a first target candidate image frame that meets the image feature condition is determined from the first candidate image frames, and the first target candidate image frame is rendered to obtain the first image frame.

5. The method according to claim 4, further comprising: If in the first operating resource environment, the target sampling frequency is the first acquisition frequency; If it is in the second operating resource environment, the target sampling frequency is the second acquisition frequency.

6. The method according to claim 4, wherein rendering the first target candidate image frame to obtain the first image frame comprises: Determine a target area on a display screen of an electronic device on which the game is running corresponding to the first candidate image frame; Acquire image content corresponding to the target area in the first target candidate image frame; Rendering the image content of the target area to obtain first image content; The first image content and the second image content are fused to obtain the first image frame; the second image content is the image content of other areas of the first target candidate image frame excluding the target area.

7. The method according to claim 4, wherein the image feature condition is selected from one of the following combinations: The included image content has a target similarity with the target image content; The image complexity of the included image content is greater than an image complexity threshold; The degree of image content change between the first candidate image frames adjacent in time sequence is greater than a change degree threshold.

8. The method according to any one of claims 1 to 7, wherein according to the mapping relationship between the image contents of the first image frame and the second image frame, the second image frame is restored using the first image frame to obtain a high-resolution image corresponding to the second image frame, comprising: The first image frame and the second image frame are input into a target intelligent engine, so that the target intelligent engine uses the first image frame to repair the second image frame according to a mapping relationship between the image contents of the first image frame and the second image frame, so as to obtain a high-resolution image corresponding to the second image frame.

9. The method according to claim 8, wherein the target intelligent engine comprises a first feature extraction module, a second feature extraction module and an attention fusion module; The target intelligent engine uses the first image frame to perform restoration processing on the second image frame according to a mapping relationship between image contents of the first image frame and the second image frame to obtain a high-resolution image corresponding to the second image frame, including: Using the first feature extraction module to extract features from the first image frame to obtain a first image feature matrix; and using the second feature extraction module to extract features from the second image frame to obtain a second image feature matrix; Determine, by the attention fusion module based on the additional information of the second image frame, a second image feature in the second image feature matrix that has a mapping relationship with the first image feature in the first image feature matrix, and superimpose the first image feature and the corresponding second image feature to obtain a high-resolution image frame corresponding to the second image frame; The additional information includes timing information between the first image frame and the second image frame, and / or image information of the second image frame.

10. An image processing device, comprising: An acquisition module, used for acquiring a first image frame; a processing module, configured to perform a restoration process on the second image frame using the first image frame according to a mapping relationship between image contents of the first image frame and the second image frame, so as to obtain a high-resolution image corresponding to the second image frame; wherein the resolution of the first image frame is greater than a preset resolution threshold, and the resolution of the second image frame is less than the preset resolution threshold; and the first image frame and the second image frame meet a similarity condition; An output module is used to output a high-resolution image frame corresponding to the second image frame.