Image processing method, electronic device, computer storage medium, and program product

CN122845860APending Publication Date: 2026-09-29BEIJING YOUKU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610999520.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]但因用户操作相较于视频播放的滞后性和不确定性,导致截取的视频帧经常存在着诸如模糊、拖影、或者非用户理想视频帧等问题,难以满足用户实际需求

Benefits of technology

[0009]根据本申请实施例提供的方案,当用户在视频播放过程中截取某一视频帧后,并不直接以该视频帧为最终画面,而是获取与该视频帧在时间轴上相邻近的多个参考帧,并将截取的视频帧本身与多个参考帧共同纳入候选范围,从中选取质量更为理想的候选视频帧。在此基础上,对该候选视频帧进行图像处理(例如图像增强处理),获得画质经过优化的目标视频帧,并将其展示给用户。由此,一方面,通过引入与截取的视频帧在时域上邻近的多个参考帧,有效扩展了候选帧的来源范围,能够弥补因用户截图操作的滞后性或不稳定性而导致的截取帧模糊、拖影或偏离用户预期内容等问题;另一方面,通过对筛选出的候选视频帧进行图像处理,进一步提升了目标视频帧的画面质量,使最终呈现的截图效果更为清晰、细腻,切实满足用户对截图画质的实际需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845860A_ABST
    Figure CN122845860A_ABST
Patent Text Reader

Abstract

An image processing method, an electronic device, a computer storage medium and a program product are provided, and the image processing method comprises: obtaining a target video frame based on a video frame intercepted in a video playing process, the target video frame being obtained by performing image processing on a candidate video frame selected from a plurality of reference frames corresponding to the video frame and the video frame, the plurality of reference frames being video frames adjacent in time domain to the video frame; and displaying the target video frame. Through the embodiment of the present application, the quality of the screenshot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to an image processing method, electronic device, computer storage medium, and computer program product. Background Technology

[0002] In the field of video playback, it has become common practice to extract a video frame during playback and then perform subsequent operations such as content review and user sharing based on that frame.

[0003] However, due to the lag and uncertainty of user operation compared to video playback, the captured video frames often have problems such as blurriness, ghosting, or non-ideal video frames, making it difficult to meet the actual needs of users. Summary of the Invention

[0004] In view of this, embodiments of this application provide an image processing solution to at least partially solve the above-mentioned problems.

[0005] According to a first aspect of the embodiments of this application, an image processing method is provided, comprising: obtaining a target video frame based on a video frame captured during video playback, wherein the target video frame is obtained by image processing of a plurality of reference frames corresponding to the video frame and candidate video frames selected from the video frame, wherein the plurality of reference frames are video frames that are temporally adjacent to the video frame; and displaying the target video frame.

[0006] According to a second aspect of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, which causes the processor to perform an operation corresponding to the method described in the first aspect.

[0007] According to a third aspect of the embodiments of this application, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0008] According to a fourth aspect of the embodiments of this application, a computer program product is provided, including computer instructions that instruct a computing device to perform an operation corresponding to the method described in the first aspect.

[0009] According to the solution provided in this application, when a user captures a video frame during video playback, the captured frame is not directly used as the final image. Instead, multiple reference frames adjacent to the captured frame on the timeline are obtained, and the captured video frame and the multiple reference frames are included in the candidate range. From these, a candidate video frame with more ideal quality is selected. Based on this, image processing (e.g., image enhancement processing) is performed on the candidate video frame to obtain a target video frame with optimized image quality, which is then displayed to the user. Thus, on the one hand, by introducing multiple reference frames that are adjacent to the captured video frame in the time domain, the source range of candidate frames is effectively expanded, which can compensate for problems such as blurry frames, ghosting, or deviation from the user's expected content caused by the lag or instability of the user's screenshot operation; on the other hand, by performing image processing on the selected candidate video frames, the image quality of the target video frame is further improved, making the final screenshot effect clearer and more delicate, effectively meeting the user's actual needs for screenshot image quality. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0011] Figure 1 A schematic diagram of an exemplary system to which the embodiments of this application are applicable; Figure 2A This is a flowchart illustrating the steps of an image processing method according to an embodiment of this application. Figure 2B for Figure 2A A schematic diagram of the first method of video frame extraction in the illustrated embodiment; Figure 2C for Figure 2A A schematic diagram of the second method of video frame extraction in the illustrated embodiment; Figure 2D for Figure 2A A schematic diagram of the third type of video frame extraction process in the illustrated embodiment; Figure 2E for Figure 2A A schematic diagram of an image processing entry mark in the illustrated embodiment; Figure 3 This is a flowchart of another image processing method according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.

[0013] The specific implementation of the embodiments of this application will be further described below with reference to the accompanying drawings.

[0014] Figure 1 An exemplary system applicable to embodiments of this application is shown. For example... Figure 1 As shown, the system 100 may include a cloud server 102, a communication network 104, and / or one or more user devices 106. Figure 1 The example in the text shows multiple user devices.

[0015] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 can perform any suitable function. For example, in some embodiments, the cloud server 102 can be used to store video data and send it to the user device 106 when a video acquisition request is received. As an optional example, in some embodiments, the cloud server 102 can acquire the target video frame and respond to the processing request, such as a sharing or forwarding request, sent by the user device 106 based on the target video frame, and return the response result to the user device 106.

[0016] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user equipment 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud server 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user equipment 106 and the cloud server 102, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communication link, or any suitable combination of such links.

[0017] User equipment 106 may include any one or more user devices suitable for interacting with a user and capable of playing video. In some embodiments, user equipment 106 may obtain a target video frame based on video frames captured during video playback. The target video frame is obtained by image processing of a plurality of reference frames corresponding to the captured video frames and candidate video frames selected from the captured video frames. The plurality of reference frames are video frames that are temporally adjacent to the captured video frames. The target video frame is then displayed. In some embodiments, user equipment 106 may include any suitable type of device. For example, in some embodiments, user equipment 106 may include mobile devices, tablet computers, laptop computers, desktop computers, wearable computers, game consoles, media players, vehicle entertainment systems, and / or any other suitable type of user equipment.

[0018] On a user's device, the user can play videos through applications with video playback capabilities (including but not limited to video playback apps, browsers, mini-programs, etc.). The interactive interface used to play video content in these applications can all be considered as the video playback interface.

[0019] In addition to rendering the video image on the user's device screen, the video playback interface also provides access to functions such as playback control and progress adjustment. During video viewing, users often need to capture the currently playing frame. This capture can be initiated by the user, for example, by clicking the screenshot icon on the video playback interface; by using a shortcut key; or by using the system-level screenshot function provided by the user's operating system. However, because the initiation time and the playback time of the video frame the user actually wants to capture are often out of sync, the user's screenshot operation has a lag. Furthermore, in some situations, the user's actions are unpredictable, making it easy to capture video frames that do not meet the user's actual needs, and easily resulting in phenomena such as blurring and ghosting.

[0020] Therefore, this application provides an image processing scheme, which will be described below through embodiments.

[0021] Reference Figure 2A This document illustrates a flowchart of the steps of an image processing method according to an embodiment of this application. The image processing method of this embodiment includes the following steps: Step S202: Obtain the target video frame based on the video frames captured during video playback.

[0022] The target video frame is obtained by image processing of multiple reference frames corresponding to the captured video frame and candidate video frames selected from the captured video frame. The multiple reference frames are video frames that are temporally adjacent to the video frame.

[0023] As mentioned earlier, because there is a lag between the time when a user initiates a screenshot operation and the actual time when the video is played, and because the user operation itself is sometimes uncertain, for example, a user expects to capture the precise moment when a person's expression reaches its peak, but the actual time of triggering the screenshot may be slightly earlier or later, resulting in a deviation between the captured video frame and the image that the user actually wants to save.

[0024] Therefore, in the image processing scheme of this application embodiment, multiple reference frames are obtained based on the video frame captured by the user. These multiple reference frames are temporally adjacent to the captured video frame, such as video frames adjacent on the video playback timeline. Temporal proximity ensures that the reference frames and the captured video frame are close in the video playback sequence, resulting in a high temporal correlation in their image content. In practical applications, a time window can be extended forward and backward from the captured video frame as a reference, and multiple reference frames can be obtained from the video stream within this time window. Alternatively, a certain number of video frames can be extracted forward and backward from the captured video frame as reference frames. These reference frames are still temporally adjacent to the captured video frame. The specific settings of the time window or the number of extractions can be flexibly set by those skilled in the art according to actual needs, and this application embodiment does not limit this. However, it is not limited to the above methods; methods that obtain reference frames only forward or only backward from the captured video frame are also applicable to the scheme of this application embodiment.

[0025] In one example, after capturing a video frame, multiple adjacent video frames within a preset time range before and after the captured frame can be obtained as reference frames. A scene consistency check is performed on the captured video frame and each reference frame to determine whether the multiple reference frames belong to the same scene content as the captured video frame. This scene consistency check can exclude video frames that are inconsistent with the user's capture intention. For example, when the reference frame maintains the same character, background, or shooting scene as the captured video frame, the reference frame can be retained; when the reference frame has switched to other shots, black screens, caption pages, transition shots, or other obviously different scene content, the reference frame is excluded.

[0026] Optionally, image content features of the captured video frame and each reference frame can be extracted separately to generate image fingerprints representing the content of the scene. Then, the difference or similarity between the image fingerprints of each reference frame and the image fingerprints of the captured video frame is calculated. If the difference is less than a preset threshold, or the similarity is greater than a preset threshold, the reference frame and the captured video frame are determined to belong to the same scene, and subsequent operations are performed; otherwise, the reference frame is filtered out. The captured video frame itself, along with multiple reference frames, are included in the candidate range, forming a set of video frames anchored to the captured video frame and covering the time domain before and / or after it. Candidate video frames with more ideal image quality or content quality are then selected from this set. This expands the source of candidate video frames beyond the single moment when the user triggers the screenshot, extending to multiple video frames within a time domain interval before and after that moment, thereby effectively improving the quality of the selected candidate video frames.

[0027] While it's possible to directly use captured video frames as a base and perform subsequent operations to acquire multiple reference frames without displaying them, ensuring the final target video frame still has high quality and good effect, one alternative approach is to first display the captured video frames during video playback on the video playback interface to allow users to more intuitively view the captured video frames. Then, based on this, multiple reference frames corresponding to the captured video frames are acquired, and candidate video frames are selected from the captured video frames and multiple reference frames. Displaying the captured video frames allows users to immediately and intuitively confirm the content of the currently captured scene, and provides an interactive trigger for subsequent candidate video frame selection and image processing operations.

[0028] In one alternative approach, displaying video frames captured during video playback on the video playback interface can be achieved by: displaying the captured video frames through a second interface on top of the first interface where the video is playing; wherein the second interface also displays an option to trigger screenshot processing for the captured video frames. By playing the video and displaying the video frames through the first and second interfaces respectively, playback and display are separated, avoiding interference with the user's normal video viewing. Furthermore, while displaying the captured video frames, the second interface also displays an option to trigger subsequent optimization processing (i.e., the screenshot processing mentioned above), providing the user with a flexible way to process the captured video frames and a more intuitive understanding of the entry point for subsequent optimization processing.

[0029] Based on this, in one alternative approach, in addition to the first interface where the video is playing, video frames captured during video playback are displayed through a second interface. This can be achieved using one of the following three display formats: The first display format: On top of the first interface where the video is playing, a second interface with a smaller display area than the first interface displays video frames captured during video playback.

[0030] The first interface is the original full-screen or large-screen interface that hosts video playback. The video content plays continuously within this interface for the user to watch without interrupting the viewing experience. The second interface can be an auxiliary display layer displayed on top of the first interface in an appropriate form. Its area is smaller than the first interface and does not obscure the main video content. This display method has a low degree of interference with the user's current viewing behavior. While confirming the captured video frames, the user can still perceive the continuous playback status of the video and watch it, thereby reducing the impact of function triggering on the user experience.

[0031] An example of this display method is as follows: Figure 2B As shown, Figure 2B The first interface shows the display when the video is playing normally, namely the first interface 2B-1; Figure 2B The second interface shows that after the video frame capture operation is performed, the captured video frame is presented on top of the video playback interface (first interface) 2B-1 in a smaller interface (second interface) 2B-2.

[0032] The second display format: Based on the first interface where the video is playing, the second interface displays video frames captured during the video playback process. The second interface has at least a first display area and a second display area. The first display area is used to continue playing the video, and the second display area is used to display the captured video frames.

[0033] In this method, the user is redirected from the current video playback interface (the first interface) to a second interface that simultaneously plays the video and displays captured video frames. The second interface handles both video playback continuation and previewing the captured video frames, dividing the screen space into two independent areas. The video continues playing in the first display area, while the captured video frames are displayed in the second display area. This method ensures that the user can view the complete video while simultaneously displaying captured video frames on the same screen.

[0034] An example of this display method is as follows: Figure 2C As shown, Figure 2C The first interface shows the display when the video is playing normally, namely the first interface 2C-1; Figure 2C The second interface shows that after the video frame capture operation is performed, the user jumps from the first interface 2C-1 to the second interface 2C-2, so that the video is played and the captured video frame is displayed in the second interface 2C-2 at the same time.

[0035] The third display format: Based on the first interface where the video is playing, a second interface is used to cover the first interface, and video frames captured during the video playback are displayed in the display area of ​​the second interface.

[0036] In this approach, the second interface replaces the first interface in a fully covered manner, guiding the user's visual focus to the viewing and processing of the captured video frames. This approach is suitable for interactive scenarios where a higher level of focus is required for the processing experience of captured video frames.

[0037] An example of this display method is as follows: Figure 2D As shown, Figure 2D The first interface shows the presentation when the video is playing normally, namely the first interface 2D-1; Figure 2D The second interface shows the transition from the first interface 2D-1 to the second interface 2D-2 after the video frame capture operation. The captured video frame is displayed in a larger area in the second interface 2D-2.

[0038] The above display methods provide a flexible way to display captured video frames, and the settings can be appropriately selected according to the actual needs of the scenario.

[0039] Furthermore, the second interface also displays an option to trigger screenshot processing for the captured video frames. This option allows users to conveniently initiate subsequent intelligent image processing flows for the captured video frames while viewing them, without needing to jump to an additional operation interface, thus reducing the complexity of the operation path. Those skilled in the art can flexibly set this screenshot processing option according to specific product forms and interaction design requirements. In the example of this application embodiment, the screenshot processing option can be illustrated as "intelligent image selection," which follows the settings of the captured video frames, such as... Figures 2B-2D The settings in the second interface used to display captured video frames.

[0040] Based on the above-mentioned screenshot processing options settings, in one optional mode, after displaying the captured video frame and screenshot processing options, a trigger operation for the screenshot processing options can be received. In response to the trigger operation, screenshot processing operations including obtaining multiple reference frames corresponding to the video frame, selecting candidate video frames from the video frame and reference frames, performing image processing on the candidate video frames, obtaining the target video frame and displaying the target video frame can be triggered.

[0041] The triggering operation can be initiated by the user. For example, the user can perform actions such as clicking, touching, or long-pressing on the screenshot processing options. By allowing the user to actively trigger the screenshot processing flow, the user has the flexibility to control the timing of the processing. This also allows the image processing capabilities provided by this solution to be integrated into the user's screenshot behavior in an optional and gradual manner, rather than interrupting the user's existing operating habits.

[0042] In one alternative approach, upon receiving a trigger operation, a page transition can also be triggered in response to that operation. Optionally, corresponding to the three aforementioned display formats, the following three transition logics can be employed: The first approach, in response to the triggered operation, redirects the user from a second screen (smaller than the first screen) to a third screen, displaying the captured video frames and the progress of the screenshot processing. An example is shown below. Figure 2B The third interface is shown in 2B-3.

[0043] The second approach, in response to the triggering operation, redirects from a second interface with a first display area and a second display area to a third interface, where the captured video frames and the progress of the screenshot processing operation are displayed. An example is... Figure 2C The third interface is shown in 2C-3.

[0044] The third approach involves updating the display of a second interface that covers the first interface in response to a trigger operation. The updated second interface displays the captured video frames and the progress of the screenshot processing operation. An example is... Figure 2D The third interface is shown in 2D-3.

[0045] The third interface, or the updated second interface, can be used to display the screenshot processing process and results. In this interface, the progress of the screenshot processing operation is presented to the user visually, such as through a progress bar, progress prompts, or loading status indicators. This allows the user to perceive the progress of the process, providing clear feedback during the waiting period and improving the user's perception of the smoothness of the processing.

[0046] In an alternative approach, displaying video frames captured during video playback based on the video playback interface may include: capturing and displaying video frames during video playback based on image processing entry markers displayed on the video playback interface.

[0047] The image processing entry marker can be a visual interactive element set on the video playback interface. It can be displayed in any appropriate manner to remind the user that the current video has an entry point capable of triggering the image processing solution provided in this application. When the user triggers the entry marker, the system responsively captures the currently playing frame and displays the captured video frame to the user, for example, in any of the three display methods described above. Combining the screenshot trigger with the image processing entry point simplifies the user's operation path from taking a screenshot to triggering subsequent image processing, helping to improve the discoverability and ease of use of the function. In one example, an example of this image processing entry marker is shown in 2E. Figure 2E At the same time Figures 2B-2C The first interface shows the entry marker, with the prompt text indicating "Use Smart Screenshot". This is an optional setting; as described above, the solution of this application embodiment can be achieved without this setting.

[0048] Based on the captured video frames, multiple reference frames corresponding to the captured video frames can be obtained, and candidate video frames can be selected from the captured video frames and multiple reference frames.

[0049] In one alternative approach, the selection of candidate video frames can be achieved by: obtaining multiple reference frames corresponding to the captured video frame; performing a quality assessment on the captured video frame and the multiple reference frames; and selecting candidate video frames from the captured video frame and the multiple reference frames based on the quality assessment results. For ease of description, in the embodiments of this application, the term "video frame set" is used to refer to multiple video frames, including the captured video frame and the multiple reference frames.

[0050] Quality assessment allows for the identification of video frames with relatively high image quality from a set of video frames, which can then be used as input for subsequent image processing. While multiple reference frames in temporal proximity are generally related to the captured video frames in terms of content, their actual image quality varies. For example, some frames may exhibit strong motion blur, while others may correspond to optimal expressions or movements. By quantitatively assessing the quality of video frames within the set, high-quality video frames can be identified based on objective data, rather than relying on randomness or heuristic rules.

[0051] The specific dimensions of quality assessment can be flexibly set by those skilled in the art according to the actual application scenario. For example, multiple dimensions such as the clarity, blur level, quality of the focus object (including but not limited to the completeness of facial expressions, eye opening and closing status, and focus clarity), and content similarity with the captured video frame can be comprehensively scored to obtain a multi-dimensional comprehensive quality score, and then candidate video frames can be selected based on this score. In addition, the specific implementation of quality assessment can also be implemented by those skilled in the art by deploying corresponding lightweight quality assessment models or algorithms on user devices according to actual needs.

[0052] In one alternative approach, obtaining multiple reference frames corresponding to the captured video frame and performing a quality assessment on the captured video frame and the multiple reference frames may include: obtaining multiple reference frames corresponding to the captured video frame; filtering reference frames from the multiple reference frames whose image differences from the video frame are greater than a preset difference threshold based on the captured video frame; and performing a quality assessment on the captured video frame and the filtered reference frames.

[0053] Among the reference frames extracted within the time-domain window, there may be video frames whose content differs significantly from the captured frame, such as those that have undergone camera cuts or scene jumps. Although these video frames are temporally adjacent to the captured video frames, their content differs greatly. Including them in the candidate selection would lead to invalid processing and waste data processing resources. Therefore, by setting a preset difference threshold, such reference frames with excessively different content can be filtered out before entering the quality assessment stage. This ensures that the video frames in the multiple reference frames are consistent with the scene of the captured video frame before quality assessment, improving the efficiency of quality assessment and saving system resources. The preset difference threshold can be configured by those skilled in the art according to actual needs, and this application embodiment does not limit this. For example, filtering can be based on one or more of visual feature differences, pixel differences, semantic differences, etc., and a corresponding preset difference threshold can be set.

[0054] Further optionally, filtering reference frames from multiple reference frames based on the captured video frame where the image difference from the video frame is greater than a preset difference threshold can be implemented by: using a perceptual hash algorithm to filter reference frames from multiple reference frames where the image difference is greater than a preset difference threshold based on the captured video frame.

[0055] Perceptual hashing is a class of algorithms used to measure the similarity of image content. It maps the visual content of an image to a fixed-length hash string and quantifies the degree of content difference between two images by calculating the Hamming distance between their hash strings. Perceptual hashing measures similarity in the content space rather than the pixel space, making it robust to factors such as changes in lighting, slight scaling, and compression distortion. It also boasts low computational complexity and high speed, making it suitable for rapidly screening large batches of inter-frame similarities under the constraint of limited computing resources on user devices.

[0056] By performing perceptual hash filtering before quality assessment, the consistency between video frames in the video frame set and the scene can be verified at the content space level. This ensures that the video frames entering the quality assessment stage are highly correlated with the extracted video frames in terms of content, establishing a reliable basis for subsequent quality assessment. Consequently, the output results of the entire "screening-assessment-image processing (including but not limited to image enhancement)" chain are more stable and reliable.

[0057] The candidate video frame obtained through the above processing may be one of multiple reference video frames, or it may be the extracted video frame itself. This candidate video frame has high quality within the set of video frames, including the extracted video frame and multiple reference video frames.

[0058] The candidate video frames obtained through the aforementioned steps already possess relatively good content quality. However, the video stream itself is affected by factors such as encoding bitrate, transmission loss, and ambient lighting conditions, which may still result in image quality issues such as insufficient sharpness and loss of detail in dark areas. To further improve the visual quality of the final screenshots presented to the user, this step applies additional image processing to the candidate video frames, resulting in a quantitative improvement in image quality.

[0059] Image processing is the process of optimizing the visual attributes of candidate video frames, including but not limited to sharpness, contrast, brightness equalization, and detail restoration, using image processing algorithms to obtain an output image with better visual quality. Image processing can be performed on the entire image or on a local area, and different algorithms or combinations of algorithms can be used depending on the processing objective.

[0060] In one alternative approach, image processing may include targeted local enhancement of certain content regions in the candidate video frame, such as the focal point of the target object (e.g., facial features), while simultaneously applying gentler adjustments to the background regions. This improves image clarity and detail while avoiding image artifacts or unnaturalness caused by over-processing. This region-specific differentiated processing strategy strikes a balance between image quality enhancement and image realism, resulting in a more natural and refined visual presentation of the final target video frame. In practical applications, those skilled in the art can deploy corresponding algorithms or lightweight models on user devices to implement image processing.

[0061] In one example, the brightness distribution information of candidate video frames can be analyzed first to determine the proportion or distribution characteristics of dark areas, mid-brightness areas, and bright areas in the image. Then, based on this brightness distribution information, enhancement parameters suitable for the current candidate video frame can be automatically determined to enhance the contrast, depth, and detail of the image. For example, for candidate video frames that are generally grayish and lack clear depth, the brightness and darkness can be enhanced to make people, backgrounds, and image details clearer. For areas with high brightness, the enhancement magnitude can be limited to avoid overexposure, whitening, or loss of detail in bright areas. For color information, color shifts can be constrained to ensure the enhanced target video frame maintains a natural appearance. When candidate video frames have insufficient brightness, weak detail, significant noise, or poor overall visual quality, an image enhancement model can be invoked to enhance the candidate video frames, improving their brightness, sharpness, detail, dynamic range, or overall visual quality.

[0062] After the image enhancement model outputs an enhanced image, it can be fused with candidate video frames, rather than directly replacing the candidate video frames with the enhanced image. For example, fusion weights can be determined based on the brightness, texture, subject, or other image region features of the candidate video frames. For instance, dark areas can utilize more information from the enhanced image to improve shadow details; bright areas can retain more original information from the candidate video frames to reduce overexposure, whitening, or unnatural appearances. This improves image quality while reducing the risk of image distortion caused by over-enhancement.

[0063] For candidate video frames containing people, the system can also identify facial areas, skin-tone related areas, or the main body of the person, and perform local sharpening or detail enhancement on these areas to make the main body of the person clearer and more natural. At the same time, the enhancement intensity of the background area can be limited to avoid over-sharpening the background and causing noise, edge jaggedness, or visual distortion.

[0064] After image enhancement is completed, the target video frame is obtained and displayed to the user.

[0065] Furthermore, in some cases, there may be occlusion or side edges in the video playback window, but the user's true intention in cropping the video frame is usually to crop the content of the playing video. Therefore, image processing may optionally include occlusion and / or side edge detection on candidate video frames; if, based on the detection results, it is determined that occlusion and / or side edges exist in the candidate video frames, then occlusion and / or side edge removal processing is performed.

[0066] The target video frames obtained through the above processing are of high quality and can effectively meet user needs.

[0067] Step S204: Display the target video frame.

[0068] After obtaining the target video frame, it can be displayed in the relevant display interface. Furthermore, the target video frame can be saved, shared, forwarded, or subjected to subsequent image processing.

[0069] In this embodiment, when a user captures a video frame during playback, that frame is not used directly as the final image. Instead, multiple reference frames that are temporally adjacent to the captured frame are obtained. The captured frame and these reference frames are then included in a candidate pool, from which a candidate frame of higher quality is selected. Based on this, image processing is performed on the candidate frame to obtain a target video frame with optimized image quality, which is then displayed to the user. Thus, on the one hand, by introducing multiple reference frames that are temporally adjacent to the captured frame, the source range of candidate frames is effectively expanded, compensating for problems such as blurry frames, ghosting, or deviations from the user's expected content caused by the lag or instability of the user's screenshot operation. On the other hand, by performing image processing on the selected candidate video frames, the image quality of the target video frame is further improved, resulting in a clearer and more detailed screenshot, effectively meeting the user's actual needs for screenshot image quality.

[0070] Reference Figure 3 This document illustrates a flowchart of another image processing method according to an embodiment of the present application. The image processing method of this embodiment includes: Step S302: Based on the video playback interface, display the video frames captured during the video playback process.

[0071] Step S304: Obtain multiple reference frames corresponding to the captured video frame, and select candidate video frames from the captured video frame and the multiple reference frames.

[0072] Among them, multiple reference frames are video frames that are temporally adjacent to the captured video frames.

[0073] Step S306: Perform image processing on the candidate video frames to obtain the target video frame, and display the target video frame.

[0074] The specific implementation of steps S302-S306 can be referred to the description of the relevant parts in the foregoing embodiments, and will not be repeated here.

[0075] Step S308: In the target video frame display interface, display the trigger options for comparing the captured video frame and the target video frame.

[0076] The trigger option provides users with an interactive entry point to compare the original captured video frames with the target video frames, allowing them to intuitively perceive the difference in image quality before and after image processing, thus forming an intuitive understanding and judgment of the image processing effect provided by this solution. Providing the comparison trigger option simultaneously in the target video frame display interface avoids users switching between multiple interfaces to complete the comparison operation, improving the accessibility of the comparison function.

[0077] In one example, the trigger option is set as follows: Figure 2B The fourth interface in the text, 2B-4. Figure 2C The fourth interface in the middle, 2C-4. Figure 2D The fourth interface in 2D-4 is shown in the lower right corner of the circular dashed box.

[0078] Step S310: In response to the operation of the trigger option, display a comparison image of the captured video frame and the target video frame.

[0079] In one alternative approach, this step can be implemented in the following two ways: The first method: Receive instantaneous operation on the trigger option, display a comparison image of the captured video frame and the target video frame; and end the display of the comparison image after receiving a cancel operation on the comparison image.

[0080] A momentary action can be a brief, triggered action by a user, such as a single click or tap, to display a comparison image. After the momentary action is triggered, the comparison image is displayed and remains displayed until the user actively cancels the display. For example, clicking the close button on the interface or window containing the comparison image, or clicking outside the comparison image area. This method is suitable for scenarios where users want to examine the differences between the original captured video frame and the target video frame in detail over a longer period.

[0081] The second method is to receive continuous operations on the trigger option, display a comparison image of the captured video frame and the target video frame during the duration of the continuous operation, and end the display of the comparison image when the continuous operation is detected to have ended.

[0082] Continuous operation refers to a user's action of maintaining a triggered state by continuously pressing or long-pressing. The display of the comparison image is synchronized with the state of the continuous operation; that is, the comparison image is displayed during the continuous press, and disappears immediately after the press is released, and the interface returns to the state displaying the target video frame. This approach creates a direct and instantaneous mapping between the user's viewing of the comparison image and their hand operation, resulting in highly intuitive operation. Furthermore, after completing the comparison viewing, the user can naturally return to the display state of the target video frame without any additional operation, simplifying the interaction chain.

[0083] This embodiment, based on the effects achieved in the aforementioned embodiments, organically combines the display and comparison functions of the target video frame. Users can not only intuitively perceive the visual quality of the target video frame, but also quantify the difference in image quality before and after processing through comparison operations. This allows the image processing effect provided by this solution to be better presented at the user's perception level, effectively improving the user's screenshot satisfaction.

[0084] Reference Figure 4 The diagram shows a structural schematic of an electronic device according to Embodiment 5 of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device.

[0085] like Figure 4 As shown, the electronic device may include: a processor 402, a communications interface 404, a memory 406, and a communications bus 408.

[0086] in: The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408.

[0087] Communication interface 404 is used to communicate with other electronic devices or servers.

[0088] The processor 402 is used to execute program 410, specifically the relevant steps in the above method embodiments.

[0089] Specifically, program 410 may include program code that includes computer operation instructions.

[0090] Processor 402 may be a CPU, a GPU (Graphics Processing Unit), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The electronic device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0091] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0092] Program 410 may include multiple computer instructions. Specifically, program 410 may use multiple computer instructions to cause processor 402 to perform the operation corresponding to any of the methods described in the foregoing multiple method embodiments.

[0093] The specific implementation of each step in procedure 410 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0094] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in any of the foregoing method embodiments. The computer storage medium includes, but is not limited to, compact disc read-only memory (CD-ROM), random access memory (RAM), floppy disk, hard disk, or magneto-optical disk.

[0095] This application also provides a computer program product, including computer instructions that instruct a computing device to perform an operation corresponding to any of the methods in the above-described multiple method embodiments.

[0096] Furthermore, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used for training the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0097] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0098] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., Random Access Memory (RAM), Read-Only Memory (ROM), Flash Memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0099] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0100] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.

Claims

1. An image processing method, comprising: Based on video frames captured during video playback, a target video frame is obtained. The target video frame is obtained by image processing of multiple reference frames corresponding to the video frame and candidate video frames selected from the video frame. The multiple reference frames are video frames that are temporally adjacent to the video frame. Display the target video frame.

2. The method according to claim 1, wherein, The method further includes: Based on the video playback interface, video frames captured during video playback are displayed.

3. The method according to claim 2, wherein, The method of displaying video frames captured during video playback based on the video playback interface includes: Based on the first interface where the video is playing, a second interface displays video frames captured during the video playback process; The second interface also displays an option to trigger screenshot processing for the captured video frames.

4. The method according to claim 3, wherein, The method of displaying video frames captured during video playback through a second interface, based on the first interface where video playback is located, includes: Above the first interface where the video is playing, the video frames captured during the video playback are displayed on a second interface with a smaller display area than the first interface. or, Based on the first interface where the video is playing, a second interface is used to display video frames captured during the video playback process. The second interface has at least a first display area and a second display area. The first display area is used to continue playing the video, and the second display area is used to display the captured video frames. or, Based on the first interface where the video is playing, a second interface covers the first interface, and the video frames captured during the video playback are displayed in the display area of ​​the second interface.

5. The method according to claim 3 or 4, wherein, The method further includes: The system receives a trigger operation for the screenshot processing option and, in response to the trigger operation, triggers a screenshot processing operation based on video frames captured during video playback to obtain the target video frame.

6. The method according to claim 5, wherein, The method further includes: In response to the triggering operation, the system jumps from a second interface with a display area smaller than the first interface to a third interface, where the captured video frame and the progress of the screenshot processing operation are displayed. or, In response to the triggering operation, the system jumps from a second interface having a first display area and a second display area to a third interface, where the captured video frame and the progress of the screenshot processing operation are displayed. or, In response to the triggering operation, the display of the second interface covering the first interface is updated, and the captured video frame and the progress of the screenshot processing operation are displayed in the updated second interface.

7. The method according to any one of claims 1-4, wherein, The display of the target video frame includes: The display interface of the target video frame shows trigger options for comparing the captured video frame and the target video frame.

8. The method according to claim 7, wherein, The method further includes: Upon receiving an instantaneous operation on the trigger option, a comparison image of the captured video frame and the target video frame is displayed; and upon receiving a cancel display operation on the comparison image, the display of the comparison image is terminated. or, The system receives continuous operations on the trigger option, displays a comparison image of the captured video frame and the target video frame during the duration of the continuous operation, and ends the display of the comparison image when the continuous operation is detected to have ended.

9. The method according to any one of claims 2-4, wherein, The method of displaying video frames captured during video playback based on the video playback interface includes: Based on the image processing entry markers displayed on the video playback interface, video frames during video playback are captured and displayed.

10. The method according to any one of claims 1-4, wherein, The process of obtaining the target video frame based on video frames captured during video playback includes: Obtain multiple reference frames corresponding to the video frame, and perform quality assessment on the video frame and the multiple reference frames; Based on the quality assessment results, candidate video frames are selected from the video frames and the plurality of reference frames. The candidate video frames are processed to obtain the target video frame.

11. The method according to claim 10, wherein, The step of obtaining multiple reference frames corresponding to the video frame and performing quality assessment on the video frame and the multiple reference frames includes: Obtain multiple reference frames corresponding to the video frame, and filter reference frames from the multiple reference frames whose image differences from the video frame are greater than a preset difference threshold based on the video frame; A quality assessment is performed on the video frame and the filtered reference frame.

12. An electronic device, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation corresponding to the method as described in any one of claims 1-11.

13. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-11.

14. A computer program product comprising computer instructions that instruct a computing device to perform an operation corresponding to any one of the methods described in claims 1-11.