Video processing method and apparatus and terminal device

The terminal device recognizes the mask and normal image of the video frame, edits and processes the video frames to generate a light-blocking effect, solving the quality and efficiency problems under manual light-blocking method, and achieving high-quality and flexible video light-blocking effect.

WO2025161695A1PCT designated stage Publication Date: 2025-08-07BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/137805
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2024-12-09
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

In the prior art, artificial lighting methods have many limitations and high light adjustment complexity during video shooting, resulting in a low quality of lighting effects.

Method used

The terminal device recognizes the mask image and normal image of the video frame, saves and stores resources, and edits and processes the video frames based on these images and the specified light-touch effect to generate a target video frame, and displays the light-touch effect.

Benefits of technology

Improve the quality and efficiency of the lighting effect and reduce the dependence on manual lighting. Users can preview and adjust the lighting effect in real time without waiting for all video frame processing to end.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137805_07082025_PF_FP_ABST
    Figure CN2024137805_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a video processing method and apparatus and a terminal device. The method comprises: for first video material, identifying a first video frame of the first video material to obtain a first mask image and a first normal image, wherein the first mask image is used for indicating an image area of a target object in the first video frame, and the first normal image is used for indicating depth information of the first video frame; storing a first storage resource obtained on the basis of the first mask image and a second storage resource obtained on the basis of the first normal image; in response to a preview operation for the first video frame, acquiring the first mask image and the first normal image on the basis of the first storage resource and the second storage resource; on the basis of the first mask image, the first normal image, and a specified lighting effect, editing the first video frame to obtain a target video frame, wherein the target video frame is used for indicating that the lighting effect is presented on the first video frame; and displaying the target video frame.
Need to check novelty before this filing date? Find Prior Art

Description

Video processing method, device and terminal equipment

[0001] This application claims priority to Chinese Patent Application No. 202410139394.2 filed on January 31, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] The embodiments of the present disclosure relate to a video processing method, apparatus, and terminal device. Background Art

[0003] When shooting a video, you can get a video with a good lighting effect by lighting the subject at different angles.

[0004] Currently, when shooting videos, users can use artificial lighting to fill in the light for the subject. For example, users can use multiple colors of lights at multiple angles to fill in the light for the subject, thereby producing a video with a well-lit effect. However, artificial lighting has many limitations, resulting in lower quality lighting effects. Summary of the Invention

[0005] In a first aspect, the present disclosure provides a video processing method, the video processing method comprising:

[0006] For a first video material, identifying a first video frame of the first video material, and obtaining a first mask image and a first normal image, wherein the first mask image is used to indicate an image region of a target object in the first video frame, and the first normal image is used to indicate depth information of the first video frame;

[0007] saving a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image;

[0008] In response to a preview operation on the first video frame, acquiring the first mask image and the first normal image according to the first storage resource and the second storage resource;

[0009] Editing the first video frame according to the first mask image, the first normal image, and the specified lighting effect to obtain a target video frame, where the target video frame is used to indicate that the lighting effect is presented on the first video frame;

[0010] The target video frame is displayed.

[0011] In a second aspect, the present disclosure provides a video processing device, which includes an identification module, a storage module, an acquisition module, a processing module, and a display module, wherein:

[0012] The recognition module is configured to, for a first video material, identify a first video frame of the first video material to obtain a first mask image and a first normal image, wherein the first mask image is used to indicate an image region of a target object in the first video frame, and the first normal image is used to indicate depth information of the first video frame;

[0013] The saving module is used to save a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image;

[0014] The acquisition module is configured to, in response to a preview operation on the first video frame, acquire the first mask image and the first normal image according to the first storage resource and the second storage resource;

[0015] The processing module is configured to edit the first video frame according to the first mask image, the first normal image, and a specified lighting effect to obtain a target video frame, where the target video frame is used to indicate that the lighting effect is presented on the first video frame;

[0016] The display module is used to display the target video frame.

[0017] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;

[0018] The memory stores computer-executable instructions;

[0019] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video processing method as described in the first aspect and various possible aspects of the first aspect.

[0020] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video processing method as described in the first aspect and various possible aspects of the first aspect are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] FIG1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;

[0023] FIG2 is a flow chart of a video processing method provided by an embodiment of the present disclosure;

[0024] FIG3 is a schematic diagram of a preview operation provided by an embodiment of the present disclosure;

[0025] FIG4 is a schematic diagram of displaying a target video frame provided by an embodiment of the present disclosure;

[0026] FIG5 is a schematic diagram of a method for processing storage resources provided by an embodiment of the present disclosure;

[0027] FIG6 is a schematic diagram of identifying a first video material according to an embodiment of the present disclosure;

[0028] FIG7 is a schematic diagram of a method for determining a target video according to an embodiment of the present disclosure;

[0029] FIG8 is a schematic structural diagram of a video processing device provided by an embodiment of the present disclosure; and

[0030] FIG9 is a schematic structural diagram of a terminal device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0032] To facilitate understanding, the concepts involved in the embodiments of the present disclosure are explained below.

[0033] Terminal device: is a device with wireless transceiver function. The terminal device can be deployed on land, including indoors or outdoors, handheld, wearable or vehicle-mounted. The terminal device can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a vehicle-mounted terminal device, a wireless terminal in self-driving, a wireless terminal device in remote medical, a wireless terminal device in smart grid, a wireless terminal device in transportation safety, a wireless terminal device in smart city, a wireless terminal device in smart home, a wearable terminal device, etc. The terminal device involved in the embodiments of the present disclosure can also be called a terminal, user equipment (UE), an access terminal device, a vehicle-mounted terminal, an industrial control terminal, a UE unit, a UE station, a mobile station, a mobile station, a remote station, a remote terminal device, a mobile device, a UE terminal device, a wireless communication device, a UE agent or a UE device, etc. Terminal devices can also be fixed or mobile.

[0034] In the related art, when shooting a video, the user can illuminate the subject at different angles, and thus obtain a video with the lighting effect. Currently, the user can use artificial lighting to fill in the light for the subject. For example, the user can set lights of various colors in the environment where the object is located, and light the object by adjusting the angles of the lights of various colors. In this way, the user can obtain a video with the lighting effect through the shooting device. However, artificial lighting has many limitations (such as the environment where the object is located cannot be set with many lights, the lighting cost is high, etc.), and the complexity of light adjustment is high (such as the adjustment of light angle, the adjustment of light intensity and the matching of light color, etc.), which results in the low quality of the video with the lighting effect shot by the terminal device.

[0035] In order to solve the technical problems in the related art, the embodiment of the present disclosure provides a video processing method. For a first video material, the terminal device can determine the target mask corresponding to the first video frame, edit the first video frame according to the target mask to obtain a first mask image, and perform normal estimation processing on the first video frame to obtain a result of the normal estimation, and determine the result of the normal estimation as a first normal image. The terminal device can save a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image. In response to a preview operation on the first video frame, the first mask image and the first normal image are obtained according to the first storage resource and the second storage resource. The first video frame is edited according to the first mask image, the first normal image and the specified lighting effect to obtain a target video frame, and the target video frame is displayed. In the above method, since the terminal device can edit the specified lighting effect in the first video frame according to the first mask image and the first normal image of the first video frame, no manual lighting is required, and since the first storage resource can store the first mask image and the second storage resource can store the first normal image, when the user previews the lighting result, the terminal device can generate a lighting preview effect according to the first storage resource and the second storage resource, without waiting for all video frames of the first video material to be processed and without manual lighting processing, which can improve the lighting effect, and since the terminal device can perform lighting processing on the video frame in combination with the mask image and the normal image, the quality of the lighting of the video frame can be improved.

[0036] The application scenario of the embodiment of the present disclosure is described below with reference to FIG1 .

[0037] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. Please refer to Figure 1, which includes: a video frame in a first video material. The terminal device (not shown in Figure 1) can identify the video frame to obtain a mask image and a normal image of the video frame. The terminal device can input the video frame, the normal image of the video frame, the mask image of the video frame, and the lighting parameters corresponding to the specified lighting effect into a pre-trained video lighting model, and the video lighting model can output a target video frame corresponding to the video frame, wherein the target video frame has the specified lighting effect. In this way, the terminal device can perform lighting processing on each video frame in the first video material, and then obtain a lighting video corresponding to the first video material, thereby improving the efficiency of lighting, and the terminal device can arbitrarily adjust the lighting effect of the lighting video according to the lighting parameters, thereby improving the flexibility of video lighting.

[0038] It should be noted that FIG1 is only an example of the embodiment of the present disclosure and does not limit the embodiment of the present disclosure.

[0039] The following detailed description of the technical solution of the present disclosure and how the technical solution of the present disclosure solves the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following embodiments of the present disclosure are described in conjunction with the accompanying drawings.

[0040] FIG2 is a flow chart of a video processing method provided by an embodiment of the present disclosure. Referring to FIG2 , the method may include:

[0041] S201 : For a first video material, identify a first video frame of the first video material, and obtain a first mask image and a first normal image.

[0042] The execution subject of the embodiment of the present disclosure may be a terminal device, or a video processing device provided in the terminal device. The video processing device may be implemented by software, or by a combination of software and hardware, which is not limited in the embodiment of the present disclosure.

[0043] The first video material may be a video material to be lighted. For example, the terminal device may receive the first video material sent by another device, the terminal device may shoot the first video material in real time, or the terminal device may obtain the shot first video material from a storage system. This is not limited in the present embodiment.

[0044] The first video frame may be a video frame in the first video material. For example, the first video frame may be a video frame that is recognized by the terminal device, wherein the terminal device may sequentially recognize and process the video frames in the first video material based on the playback order of the video frames in the first video material. For example, the first video material may include video frame A and video frame B. If the terminal device recognizes and processes video frame A, video frame A may be the first video frame. If the terminal device recognizes and processes video frame B, video frame B may be the first video frame.

[0045] The first mask image is used to indicate the image area of ​​the target object in the first video frame. For example, the first mask image may be an image obtained by multiplying the first video frame by the mask corresponding to the first video frame. The terminal device may obtain the first mask image according to the following feasible implementation: determining a target mask corresponding to the first video frame, and editing the first video frame according to the target mask to obtain the first mask image.

[0046] The target mask may be a mask of the first video frame, and the target mask has the same size as that of the first video frame.

[0047] Optionally, the terminal device can determine the image area of ​​the target object in the first video frame in the first video frame, and determine the target mask based on the image area. For example, the image area corresponding to the target in the first video frame is area 1, and the area corresponding to area 1 in the target mask is area 2 (the target mask has the same size as the first video frame), the value of area 2 in the target mask is 1, and the values ​​of other areas are 0. For example, the target object can be a portrait, an animal image, a plant image, etc., which is not limited in the embodiment of the present disclosure. The image area can be an area surrounded by the outline of the target object, or it can be a rectangular area surrounding the target object, which is not limited in the embodiment of the present disclosure.

[0048] It should be noted that the terminal device can determine the image area of ​​the target object in the first video frame according to any feasible implementation method, and the embodiment of the present disclosure is not limited to this.

[0049] Since the pixel values ​​in the target mask are either 0 or 1, when the terminal device edits the first video frame based on the target mask, if the pixel value in the target mask is 0, the pixel color at that location in the first mask image is black, and if the pixel value in the target mask is 1, the pixel color at that location in the first mask image is white. For example, if the first video frame includes a portrait, the area in the first mask image corresponding to the first video frame containing the portrait can be white, while the rest of the background area can be black.

[0050] The first normal image is used to indicate depth information of the first video frame. For example, the first normal image may include normal vector information of each pixel in the first video frame, wherein the normal vector information may be represented by a three-dimensional vector (normal vector), and the three components of the normal vector information may be mapped to three color channels.

[0051] The terminal device may obtain the first normal image according to the following feasible implementation method: performing normal estimation processing on the first video frame to obtain a normal estimation result corresponding to the first video frame, and determining the normal estimation result corresponding to the first video frame as the first normal image. For example, the terminal device may determine normal vector information for each pixel in the first video frame, and then generate the first normal image based on the normal vector information of each pixel point, thereby improving the accuracy of the first normal image.

[0052] It should be noted that the terminal device can perform normal estimation processing on the first video frame image according to any feasible implementation method, and the embodiment of the present disclosure is not limited to this.

[0053] S202: Save a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image.

[0054] Among them, when the terminal device identifies the first video frame in the first video material, the terminal device can identify the video frames in the first video material frame by frame from beginning to end, obtain the mask image and normal image of each video frame, and the terminal device can save the mask image and normal image corresponding to each video frame.

[0055] The first storage resource may include the first mask image corresponding to the processed video frame and / or the first video encoded based on the first mask image. For example, the terminal device may identify the video frames in the first video material. If the terminal device completes processing of all video frames in the first video material, the first storage resource may include the first mask image corresponding to all video frames and the first video encoded with the first mask image (the video encoded with the mask image corresponding to all video frames in the first video material). If the terminal device completes processing of some video frames in the first video material, the first storage resource may include the first mask image corresponding to the processed some video frames.

[0056] The second storage resource may include the first normal image corresponding to the processed video frame and / or the first video encoded based on the first normal image. For example, the terminal device may identify the video frames in the first video material. If the terminal device completes processing of all video frames in the first video material, the second storage resource may include the first normal images corresponding to all video frames and the second video encoded with the first normal image (the video encoded with the normal images corresponding to all video frames in the first video material). If the terminal device completes processing of some video frames in the first video material, the second storage resource may include the first normal images corresponding to the processed some video frames.

[0057] Optionally, when the terminal device obtains the first mask image and the first normal image of the first video frame, the first mask image and the first normal image can be stored in the cache. For example, the terminal device can save the mask image as file 1 and the normal image as file 2 in the cache. After the terminal device obtains all the mask images, it can encode to obtain the first video. After the terminal device obtains all the normal images, it can encode to obtain the second video. In this way, when the user previews the lighting effect at any stage of the lighting processing, the terminal device can obtain the mask image and normal image corresponding to the previewed video frame, and then obtain a preview video frame with the lighting effect, without having to wait for the processing of all video frames to be completed, thereby improving the lighting efficiency and improving the user experience.

[0058] It should be noted that the terminal device can encode the first video after obtaining all the mask images, and encode the second video after obtaining all the normal images. The terminal device can also create an encoding thread, that is, the terminal device identifies the video frames and encodes the first video and the second video in parallel. After the terminal device completes the identification of all the video frames, the terminal device can obtain the first video and the second video. In this way, the generation efficiency of the first video and the second video can be improved.

[0059] It should be noted that in actual application, after the terminal device identifies the video frame in the first video material, it can obtain the normal image and the mask image. The terminal device can place the normal image at the position marked as 1 and the mask image at the position marked as 0. In this way, the terminal device can divert the normal image and the mask image according to the position mark, and then store the normal image and the mask image separately.

[0060] S203 : In response to a preview operation on the first video frame, obtain a first mask image and the first normal image according to the first storage resource and the second storage resource.

[0061] Optionally, the user's preview operation on the first video frame may specifically be: displaying multiple video frames of the first video material, and in response to a touch operation on the multiple video frames, determining the video frame corresponding to the touch operation as the first video frame.

[0062] It should be noted that the preview operation may be a user's touch operation on the preview control, or a user's touch operation on the video frame, which is not limited in the embodiments of the present disclosure.

[0063] In the process of the terminal device identifying the video frames in the first video material, the display page of the terminal device may include multiple video frames arranged in the playback order. If the user clicks on any video frame, the terminal device can display the illuminated image of the video frame in the preview area of ​​the page.

[0064] The preview operation is described below with reference to FIG3 .

[0065] Figure 3 is a schematic diagram of a preview operation provided by an embodiment of the present disclosure. Please refer to Figure 3, which includes: a display page of a terminal device (not shown in Figure 3). Among them, the display page may include a preview area and an image display area. The image display area can display multiple video frames such as video frame 1, video frame 2, video frame 3, video frame 4 and video frame 5 in the first video material, and the preview area is used to display the lighting results of the video frames. When the user clicks on video frame 4, the terminal device can determine video frame 4 as the first video frame, and the terminal device can determine the user's preview operation as: preview the lighting effect of video frame 4. In this way, the user can flexibly preview the lighting effect of the video frame to improve the user experience.

[0066] Among them, the terminal device can obtain the first mask image corresponding to the first video frame in the first storage resource and obtain the first normal image corresponding to the first video frame in the second storage resource according to the first video frame corresponding to the preview operation.

[0067] Among them, the terminal device obtains the first mask image and the first normal image based on the first storage resource and the second storage resource. Specifically, it can be: if all video frames in the first video material are processed, the first video in the first storage resource is decoded to obtain the first mask image, and the second video in the second storage resource is decoded to obtain the first normal image. If there are unprocessed video frames in the first video material, the first mask image is obtained in the first storage resource, and the first normal image is obtained in the second storage resource.

[0068] For example, if there are unprocessed video frames in the first video material, the first storage resource may include mask images corresponding to multiple processed video frames, and the second storage resource may include normal images corresponding to multiple processed video frames. The terminal device can obtain the mask image and normal image of the first video frame in the first storage resource and the second storage resource according to the identifier of the first video frame.

[0069] For example, if all video frames in the first video material are processed, the first storage resource may include a first video generated based on multiple mask images, and the second storage resource may include a second video generated based on multiple normal images. Therefore, the terminal device can decode the first video to obtain multiple mask images, and determine the first mask image corresponding to the first video frame among the multiple mask images. The terminal device can decode the second video to obtain multiple normal images, and determine the first normal image corresponding to the first video frame among the multiple normal images.

[0070] S204 : Edit the first video frame according to the first mask image, the first normal image, and the specified lighting effect to obtain a target video frame.

[0071] The target video frame is used to indicate a lighting effect to be presented on the first video frame. For example, if the lighting effect corresponding to the lighting parameter is a gradient light effect, the target video frame may include the gradient light effect; if the lighting effect corresponding to the lighting parameter is a breathing light effect, the target video frame may include the breathing light effect.

[0072] Among them, the terminal device can obtain the target video frame according to the following feasible implementation method: determine the lighting parameters corresponding to the specified lighting effect, edit the lighting area in the first video frame according to the lighting parameters, the first mask image and the first normal image, and obtain the target video frame.

[0073] The lighting parameters may indicate information such as the lighting direction, lighting intensity, lighting color, and lighting area, which is not limited in the embodiments of the present disclosure.

[0074] It should be noted that the terminal device can determine the specified lighting effect based on the user's operation (for example, the display page of the terminal device may include multiple controls, each control is associated with a group of lighting effects, and the terminal device can determine the specified lighting effect based on the control clicked by the user). The terminal device can also determine the specified lighting effect according to any feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0075] Optionally, the terminal device can obtain the target video frame according to the following feasible implementation method: input the first video frame, the first mask image, the first normal image and the lighting parameters into the pre-trained lighting model, and the lighting model can output the target video frame corresponding to the first video frame.

[0076] Among them, the lighting model can be a pre-trained model, and the training samples of the lighting model may include sample images, sample mask images corresponding to the sample images, sample normal images corresponding to the sample images, sample lighting parameters, and sample lighting images after the sample images are illuminated by the sample lighting parameters. The terminal device can train the lighting model based on the training samples. It should be noted that the training process of the lighting model by the terminal device will not be repeated here in the embodiments of the present disclosure.

[0077] Optionally, the terminal device can also input a special effects prop package into the lighting model, so that the target video frame can include not only the lighting effect, but also the special effects corresponding to the special effects prop package, thereby improving the display effect of the target video frame.

[0078] S205: Display the target video frame.

[0079] Optionally, after the terminal device obtains the target video frame, it can display the target video frame on the display page. For example, because the user previews the first video frame, after the terminal device obtains the target video frame corresponding to the first video frame, it can display the target video frame in the preview area of ​​the display page of the terminal device to improve the user experience.

[0080] Next, the process of displaying the target video frame on the terminal device will be described with reference to FIG. 4 .

[0081] FIG4 is a schematic diagram of a display target video frame provided by an embodiment of the present disclosure. Referring to FIG4 , it includes: a display page of a terminal device (not shown in FIG4 ). The display page may include a preview area and an image display area. The image display area may display multiple video frames, such as video frame 1, video frame 2, video frame 3, video frame 4, and video frame 5, in the first video material. The preview area is used to display the lighting results of the video frames.

[0082] Referring to Figure 4 , when a user clicks on video frame 3, the terminal device can determine the preview operation as previewing the lighting effect of video frame 3. The terminal device can generate a lighting image for video frame 3 and display it in the preview area. When the user clicks on video frame 2, the terminal device can generate a lighting image for video frame 2, cancel the display of the lighting image for video frame 3 in the preview area, and display the lighting image for video frame 2 in the preview area. In this way, the terminal device can display the lighting results of the video frames in real time, improving the user experience.

[0083] An embodiment of the present disclosure provides a video processing method. For a first video material, a terminal device can determine a target mask corresponding to a first video frame, edit the first video frame according to the target mask to obtain a first mask image, and perform normal estimation processing on the first video frame to obtain a result of the normal estimation, and determine the result of the normal estimation as a first normal image. The terminal device can save a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image. In response to a preview operation on the first video frame, the terminal device obtains the first mask image and the first normal image according to the first storage resource and the second storage resource. The terminal device edits the first video frame according to the first mask image, the first normal image and a specified lighting effect to obtain a target video frame, and displays the target video frame. In the above method, since the first storage resource can store the first mask image and the second storage resource can store the first normal image, when the user previews the lighting result, the terminal device can generate a lighting preview effect based on the first storage resource and the second storage resource. There is no need to wait for all video frames of the first video material to be processed, and there is no need to perform manual lighting processing. This can improve the lighting effect. Moreover, since the terminal device can perform lighting processing on the video frame in combination with the mask image and the normal image, the quality of the lighting of the video frame can be improved.

[0084] Based on the embodiment shown in FIG2 , the above-mentioned video processing method also includes a method for processing the first storage resource and the second storage resource. Below, in conjunction with FIG5 , the method for processing the first storage resource and the second storage resource by the terminal device is described in detail.

[0085] FIG5 is a schematic diagram of a method for processing storage resources provided by an embodiment of the present disclosure. Referring to FIG5 , the method flow may include:

[0086] S501: Generate a first video and a second video.

[0087] The first video may be a video associated with multiple mask images. Optionally, the terminal device may generate the first video based on multiple mask images corresponding to multiple video frames of the first video material. For example, in the process of recognizing video frames, the terminal device may save the mask image (first storage resource) corresponding to each video frame in the cache. Therefore, the terminal device may create a mask image encoding thread, and then perform video encoding on the multiple mask images. When the terminal device finishes recognizing all video frames, the terminal device may obtain all mask images, and then encode the first video.

[0088] The second video may be a video associated with multiple normal images. Optionally, the terminal device may generate a second video based on multiple normal images corresponding to multiple video frames of the first video material. For example, in the process of identifying video frames, the terminal device may save the normal image corresponding to each video frame in the cache (the second storage resource). Therefore, the terminal device may create an encoding thread for the normal image, and then perform video encoding on the multiple normal images. When the terminal device finishes identifying all video frames, the terminal device may obtain all normal images, and then encode the second video.

[0089] It should be noted that the video information of the first video is the same as the video information of the first video material, and the video information of the second video is the same as the video information of the first video material. For example, the frame rate of the first video is the same as the frame rate of the first video material, and the frame rate of the second video is also the same as the frame rate of the first video material.

[0090] It should be noted that, during the process of the terminal device identifying the video frames of the first video material, the terminal device can encode the mask image and the normal image in parallel. When all video frames are identified, the terminal device can obtain the first video and the second video. The terminal device can also encode multiple mask images to obtain the first video and encode multiple normal images to obtain the second video when all video frames are identified. The embodiments of the present disclosure do not limit this.

[0091] 6 , the process of the terminal device identifying the first video material will be described below.

[0092] Figure 6 is a schematic diagram of a method for identifying a first video material provided by an embodiment of the present disclosure. Please refer to Figure 6, which includes: a first video material and a terminal device. When the terminal device obtains the first video material, it can decode the first video material frame by frame from the beginning to the end to obtain the video frame in the first video material. The terminal device can perform image processing on the video frame to obtain the normal image and mask image corresponding to the video frame, and store the normal image and mask image in the cache. During the image processing process, the terminal device can create a video encoding thread, and then encode the mask image into the first video and encode the normal image into the second video. In this way, before the encoding of the first video and the second video is completed, the cache of the terminal device can include the normal image and the mask image. Therefore, the terminal device can generate the lighting image of the video frame in real time, reduce the time consumed by the user during preview, and improve the user experience.

[0093] S502: If all video frames in the first video material have been processed, delete the first mask image in the first storage resource and the first normal image in the second storage resource.

[0094] Among them, after the terminal device obtains the first video and the second video, it means that the terminal device has completed the recognition of all video frames in the first video material. Therefore, the terminal device can delete the mask image in the first storage resource and delete the normal image in the second storage resource. Since the normal image and the mask image occupy a large cache space, this can free up cache space, reduce cache occupancy, and improve cache utilization.

[0095] The disclosed embodiments provide a method for processing first and second storage resources. When a terminal device completes recognizing all video frames of a first video material, the terminal device can determine the first video based on multiple mask images corresponding to multiple video frames of the first video material, determine the second video based on multiple normal images corresponding to multiple video frames of the first video material, and delete the multiple mask images and the multiple normal images from a cache. In this way, the terminal device can free up cache space, reduce cache occupancy, and improve cache utilization.

[0096] Based on any of the above embodiments, after the terminal device completes the recognition of all video frames of the first video material, the above video processing method also includes a method for obtaining the target video in response to the preview operation of the first video material. Below, in conjunction with Figure 7, the method for determining the target video is described in detail.

[0097] FIG7 is a schematic diagram of a method for determining a target video provided by an embodiment of the present disclosure. Referring to FIG7 , the method includes:

[0098] S701: In response to a preview operation on a first video material, obtain a first video and a second video corresponding to the first video material.

[0099] Among them, after the terminal device generates the first video and the second video, in order to reduce the cache occupancy rate, the terminal device can delete the mask image of the first storage resource and the normal image of the second storage resource in the cache. Therefore, if the user previews the lighting effect of the first video material, the terminal device can obtain the first video and the second video in the cache, and then obtain the mask image and normal image corresponding to each video frame in the first video material.

[0100] It should be noted that the terminal device can obtain the first video and the second video corresponding to the first video material according to any feasible implementation method, and the embodiment of the present disclosure is not limited to this.

[0101] S702: Edit the first video material according to the specified lighting effect, the first video, and the second video to obtain a target video.

[0102] The target video may include a specified lighting effect. For example, if the specified lighting effect is a background breathing light effect, the background in the target video will have a breathing light effect; if the specified lighting effect is a gradient color effect, the target video will include a gradient color lighting effect.

[0103] Optionally, the terminal device can obtain the target video according to the following feasible implementation method: decode the first video material to obtain multiple second video frames of the first video material, decode the first video to obtain multiple third video frames of the first video, decode the second video to obtain multiple fourth video frames of the second video, and obtain the target video according to the specified lighting effect, multiple second video frames, multiple third video frames and multiple fourth video frames.

[0104] Among them, since the first video and the second video can be special effect resources in the process of lighting the video frames in the first video material, the terminal device can support the operation of decoding the video in the special effect resource. For example, in the actual application process, since the special effect resource can include a special effect prop package, the terminal device only needs to download the special effect prop package and does not need to decode the special effect prop package. However, since the first video and the second video in the embodiment of the present disclosure belong to lighting special effects, the terminal device can create a decoding thread for the first video and the second video. When the first video and the second video are included in the special effect resource, the terminal device can decode the first video to obtain multiple third video frames, and decode the second video to obtain multiple fourth video frames.

[0105] Optionally, for any second video frame in the first video material, the terminal device obtains the target video based on the specified lighting effect, multiple second video frames, multiple third video frames and multiple fourth video frames. Specifically, it can be: determine the third video frame and the fourth video frame with the same timestamp as the second video frame, and process the second video frame and the third video frame and the fourth video frame with the same timestamp as the second video frame according to the lighting parameters corresponding to the specified lighting effect to obtain video frames of the target video.

[0106] Among them, the terminal device can determine the timestamp of the second video frame in the first video material, determine the video frame with the same timestamp among multiple third video frames, and determine the video frame with the same timestamp among multiple fourth video frames, and then input the second video frame, the third video frame with the same timestamp, the fourth video frame with the same timestamp and the lighting parameters corresponding to the specified lighting effect into the lighting model. The lighting model can output the lighting result of the second video frame. The terminal device can repeat the above steps until all second video frames in the first video material are processed. The target video can be obtained. The terminal device can synthesize the three videos into a target video with a lighting effect to improve the quality of the lighting. In addition, the user interaction side will only see one video with a lighting effect, which improves the user experience.

[0107] The disclosed embodiments provide a method for determining a second video. In response to a preview operation on a first video material, the first video and the second video corresponding to the first video material are obtained. The first video material is edited according to a specified lighting effect, the first video, and the second video to obtain a target video. In this way, the first video material can be accompanied by two auxiliary videos (the first video and the second video) for joint decoding, so that the terminal device can decode the first video and the second video, thereby improving the efficiency of lighting. In addition, the terminal device can merge the three videos into a video with a lighting effect, thereby improving the quality of the lighting effect.

[0108] FIG8 is a schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure. Referring to FIG8 , the video processing device 800 includes an identification module 801, a storage module 802, an acquisition module 803, a processing module 804, and a display module 805, wherein:

[0109] The recognition module 801 is configured to, for a first video material, identify a first video frame of the first video material, and obtain a first mask image and a first normal image, wherein the first mask image is used to indicate an image region of a target object in the first video frame, and the first normal image is used to indicate depth information of the first video frame;

[0110] The saving module 802 is configured to save a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image;

[0111] The acquisition module 803 is configured to acquire the first mask image and the first normal image according to the first storage resource and the second storage resource in response to a preview operation on the first video frame;

[0112] The processing module 804 is configured to edit the first video frame according to the first mask image, the first normal image, and the specified lighting effect to obtain a target video frame, where the target video frame is used to indicate that the lighting effect is presented on the first video frame;

[0113] The display module 805 is configured to display the target video frame.

[0114] According to one or more embodiments of the present disclosure, the identification module 801 is specifically configured to:

[0115] determining a target mask corresponding to the first video frame;

[0116] The first video frame is edited according to the target mask to obtain the first mask image.

[0117] According to one or more embodiments of the present disclosure, the identification module 801 is specifically configured to:

[0118] performing normal estimation processing on the first video frame to obtain a normal estimation result corresponding to the first video frame;

[0119] A result of normal estimation corresponding to the first video frame is determined as the first normal image.

[0120] According to one or more embodiments of the present disclosure, the first storage resource includes a first mask image corresponding to a processed video frame and / or a first video encoded based on the first mask image;

[0121] The second storage resource includes a first normal image corresponding to the processed video frame and / or a second video encoded based on the first normal image.

[0122] According to one or more embodiments of the present disclosure, the processing module 804 is specifically configured to:

[0123] If all video frames in the first video material are processed, the first mask image in the first storage resource and the first normal image in the second storage resource are deleted.

[0124] According to one or more embodiments of the present disclosure, the acquisition module 803 is specifically configured to:

[0125] If all video frames in the first video material have been processed, decoding the first video in the first storage resource to obtain a first mask image, and decoding the second video in the second storage resource to obtain a first normal image;

[0126] If there are unprocessed video frames in the first video material, the first mask image is obtained from the first storage resource, and the first normal image is obtained from the second storage resource.

[0127] According to one or more embodiments of the present disclosure, the processing module 804 is specifically configured to:

[0128] Determining lighting parameters corresponding to the specified lighting effect;

[0129] The lighting area in the first video frame is edited according to the lighting parameters, the first mask image, and the first normal image to obtain the target video frame.

[0130] The video processing device provided in the embodiment of the present disclosure can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0131] FIG9 is a schematic diagram of the structure of a terminal device provided by an embodiment of the present disclosure. Please refer to FIG9 , which shows a schematic diagram of the structure of a terminal device 900 suitable for implementing an embodiment of the present disclosure. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The terminal device shown in FIG9 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0132] As shown in FIG9 , terminal device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 902 or programs loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of terminal device 900 are also stored in RAM 903. Processing device 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.

[0133] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the terminal device 900 to communicate with other devices wirelessly or by wire to exchange data. Although FIG9 shows a terminal device 900 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may alternatively be implemented or have.

[0134] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0135] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0136] The computer-readable medium may be included in the terminal device, or may exist independently without being incorporated into the terminal device.

[0137] The computer-readable medium carries one or more programs. When the one or more programs are executed by the terminal device, the terminal device executes the method shown in the above embodiment.

[0138] An embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, various methods that may be involved in the above embodiments are implemented.

[0139] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements various possible methods involved in the above embodiments when executed by a processor.

[0140] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0142] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0143] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0145] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0146] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0147] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0148] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thereby, the user can independently choose whether to provide personal information to the software or hardware such as the terminal device, application, server or storage medium that performs the operation of the technical solution of the present disclosure based on the prompt message. As an optional but non-limiting implementation method, in response to receiving an active request from the user, the method of sending the prompt message to the user can be, for example, a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the terminal device.

[0149] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0150] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws and regulations. Data may include information, parameters and messages, such as flow switching indication information.

[0151] In a first aspect, an embodiment of the present disclosure provides a video processing method, the video processing method comprising:

[0152] For a first video material, identifying a first video frame of the first video material, and obtaining a first mask image and a first normal image, wherein the first mask image is used to indicate an image region of a target object in the first video frame, and the first normal image is used to indicate depth information of the first video frame;

[0153] saving a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image;

[0154] In response to a preview operation on the first video frame, acquiring the first mask image and the first normal image according to the first storage resource and the second storage resource;

[0155] Editing the first video frame according to the first mask image, the first normal image, and the specified lighting effect to obtain a target video frame, where the target video frame is used to indicate that the lighting effect is presented on the first video frame;

[0156] The target video frame is displayed.

[0157] According to one or more embodiments of the present disclosure, identifying a first video frame of the first video material to obtain a first mask image includes:

[0158] determining a target mask corresponding to the first video frame;

[0159] The first video frame is edited according to the target mask to obtain the first mask image.

[0160] According to one or more embodiments of the present disclosure, identifying a first video frame of the first video material and obtaining a first normal image includes:

[0161] performing normal estimation processing on the first video frame to obtain a normal estimation result corresponding to the first video frame;

[0162] A result of normal estimation corresponding to the first video frame is determined as the first normal image.

[0163] According to one or more embodiments of the present disclosure, the first storage resource includes a first mask image corresponding to a processed video frame and / or a first video encoded based on the first mask image;

[0164] The second storage resource includes a first normal image corresponding to the processed video frame and / or a second video encoded based on the first normal image.

[0165] According to one or more embodiments of the present disclosure, after saving the first storage resource obtained based on the first mask image and the second storage resource obtained based on the first normal image, the method further includes:

[0166] If all video frames in the first video material are processed, the first mask image in the first storage resource and the first normal image in the second storage resource are deleted.

[0167] According to one or more embodiments of the present disclosure, acquiring the first mask image and the first normal image according to the first storage resource and the second storage resource includes:

[0168] If all video frames in the first video material have been processed, decoding the first video in the first storage resource to obtain a first mask image, and decoding the second video in the second storage resource to obtain a first normal image;

[0169] If there are unprocessed video frames in the first video material, the first mask image is obtained from the first storage resource, and the first normal image is obtained from the second storage resource.

[0170] According to one or more embodiments of the present disclosure, the editing process of the first video frame according to the first mask image, the first normal image, and the specified lighting effect to obtain the target video frame includes:

[0171] Determining lighting parameters corresponding to the specified lighting effect;

[0172] The lighting area in the first video frame is edited according to the lighting parameters, the first mask image, and the first normal image to obtain the target video frame.

[0173] In a second aspect, an embodiment of the present disclosure provides a video processing device, which includes an identification module, a storage module, an acquisition module, a processing module, and a display module, wherein:

[0174] The recognition module is configured to, for a first video material, identify a first video frame of the first video material to obtain a first mask image and a first normal image, wherein the first mask image is used to indicate an image region of a target object in the first video frame, and the first normal image is used to indicate depth information of the first video frame;

[0175] The saving module is used to save a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image;

[0176] The acquisition module is configured to, in response to a preview operation on the first video frame, acquire the first mask image and the first normal image according to the first storage resource and the second storage resource;

[0177] The processing module is configured to edit the first video frame according to the first mask image, the first normal image, and a specified lighting effect to obtain a target video frame, where the target video frame is used to indicate that the lighting effect is presented on the first video frame;

[0178] The display module is used to display the target video frame.

[0179] According to one or more embodiments of the present disclosure, the identification module is specifically configured to:

[0180] determining a target mask corresponding to the first video frame;

[0181] The first video frame is edited according to the target mask to obtain the first mask image.

[0182] According to one or more embodiments of the present disclosure, the identification module is specifically configured to:

[0183] performing normal estimation processing on the first video frame to obtain a normal estimation result corresponding to the first video frame;

[0184] A result of normal estimation corresponding to the first video frame is determined as the first normal image.

[0185] According to one or more embodiments of the present disclosure, the first storage resource includes a first mask image corresponding to a processed video frame and / or a first video encoded based on the first mask image;

[0186] The second storage resource includes a first normal image corresponding to the processed video frame and / or a second video encoded based on the first normal image.

[0187] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0188] If all video frames in the first video material are processed, the first mask image in the first storage resource and the first normal image in the second storage resource are deleted.

[0189] According to one or more embodiments of the present disclosure, the acquisition module is specifically configured to:

[0190] If all video frames in the first video material have been processed, decoding the first video in the first storage resource to obtain a first mask image, and decoding the second video in the second storage resource to obtain a first normal image;

[0191] If there are unprocessed video frames in the first video material, the first mask image is obtained from the first storage resource, and the first normal image is obtained from the second storage resource.

[0192] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0193] Determining lighting parameters corresponding to the specified lighting effect;

[0194] The lighting area in the first video frame is edited according to the lighting parameters, the first mask image, and the first normal image to obtain the target video frame.

[0195] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;

[0196] The memory stores computer-executable instructions;

[0197] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video processing method as described in the first aspect and various possible aspects of the first aspect.

[0198] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video processing method as described in the first aspect and various possible aspects of the first aspect are implemented.

[0199] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0200] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0201] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video processing method, comprising: For a first video material, identifying a first video frame of the first video material, and obtaining a first mask image and a first normal image, wherein the first mask image is used to indicate an image region of a target object in the first video frame, and the first normal image is used to indicate depth information of the first video frame; saving a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image; In response to a preview operation on the first video frame, acquiring the first mask image and the first normal image according to the first storage resource and the second storage resource; Editing the first video frame according to the first mask image, the first normal image, and the specified lighting effect to obtain a target video frame, where the target video frame is used to indicate that the lighting effect is presented on the first video frame; and The target video frame is displayed.

2. The method according to claim 1, wherein The identifying the first video frame of the first video material to obtain a first mask image includes: determining a target mask corresponding to the first video frame; The first video frame is edited according to the target mask to obtain the first mask image.

3. The method according to claim 1, wherein The identifying the first video frame of the first video material to obtain a first normal image includes: performing normal estimation processing on the first video frame to obtain a normal estimation result corresponding to the first video frame; A result of normal estimation corresponding to the first video frame is determined as the first normal image.

4. The method according to any one of claims 1 to 3, wherein: The first storage resource includes a first mask image corresponding to the processed video frame and / or a first video encoded based on the first mask image; The second storage resource includes a first normal image corresponding to the processed video frame and / or a second video encoded based on the first normal image.

5. The method according to claim 4, wherein After saving the first storage resource obtained based on the first mask image and the second storage resource obtained based on the first normal image, the method further includes: If all video frames in the first video material are processed, the first mask image in the first storage resource and the first normal image in the second storage resource are deleted.

6. The method according to claim 4, wherein: The acquiring the first mask image and the first normal image according to the first storage resource and the second storage resource includes: If all video frames in the first video material have been processed, decoding the first video in the first storage resource to obtain the first mask image, and decoding the second video in the second storage resource to obtain the first normal image; If there are unprocessed video frames in the first video material, the first mask image is obtained from the first storage resource, and the first normal image is obtained from the second storage resource.

7. The method according to any one of claims 1 to 6, wherein: The editing process of the first video frame according to the first mask image, the first normal image and the specified lighting effect to obtain a target video frame includes: Determining lighting parameters corresponding to the specified lighting effect; The lighting area in the first video frame is edited according to the lighting parameters, the first mask image, and the first normal image to obtain the target video frame.

8. A video processing device, comprising a recognition module, a storage module, an acquisition module, a processing module, and a display module, wherein: The recognition module is configured to recognize a first video frame of a first video material, obtain a first mask image and a first normal image, wherein the first mask image is used to indicate an image area of a target object in the first video frame, and the first normal image is used to indicate depth information of the first video frame; The saving module is configured to save a first storage resource obtained based on the first mask image and a second storage resource obtained based on the first normal image; The acquisition module is configured to acquire the first mask image and the first normal image according to the first storage resource and the second storage resource in response to a preview operation on the first video frame; The processing module is configured to edit the first video frame according to the first mask image, the first normal image and the specified lighting effect to obtain a target video frame, where the target video frame is used to indicate that the lighting effect is presented on the first video frame; The display module is configured to display the target video frame.

9. A terminal device comprising: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the video processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the video processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Video processing method and device and equipment

    CN112165632A

  • Image-based lighting effect processing method and device, equipment and storage medium

    CN112270759A

  • Image rendering method and device, electronic equipment and storage medium

    CN115830040A

  • Secondary light source creating device and reflection characteristics measuring device

    JP2001165772A

  • Graphics rendering method and apparatus

    WO2022228383A1