Video processing method and apparatus, computer device, and storage medium

By allowing users to select the lighting and shadow rendering area and lighting parameters, and combining a multi-task neural network model to process the lighting of video frames, this technology solves the problem of insufficient lighting and shadow aesthetics in video scenes in existing technologies, and achieves personalized video lighting effects and a rich user experience.

WO2025031315A9PCT designated stage expired Publication Date: 2026-04-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing technologies struggle to personalize lighting for video frames through 3D stereoscopic perception, resulting in insufficient lighting and shadow aesthetics in video scenes and a poor user experience.

Method used

A video processing method is provided that allows users to select the lighting and shadow rendering area and lighting parameters by displaying instruction information, and combines a multi-task neural network model to generate an initial normal map and segmentation map, performs lighting processing, and generates the target video.

Benefits of technology

It enables personalized lighting processing for video scenes, improves the aesthetic effects of light and shadow and user experience, supports diverse lighting scenes and dynamic lighting effects, and meets the needs of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109768_02042026_PF_FP_ABST
    Figure CN2024109768_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a video processing method and apparatus, a computer device, and a storage medium. The method comprises: in response to receiving a lighting processing request for a video to be processed, displaying instruction information, wherein the instruction information comprises first instruction information used for instructing a user to select a light and shadow rendering area and second instruction information used for instructing the user to select lighting parameter information; for a target video frame in the video to be processed, respectively determining a target light and shadow rendering area selected by the user on the basis of the first instruction information, and target lighting parameter information selected on the basis of the second instruction information, wherein the target light and shadow rendering area and the target lighting parameter information determined for the target video frame are used for performing lighting processing on a target video clip associated with the target video frame; and displaying a target video obtained after performing, on the basis of the target light and shadow rendering area and the target lighting parameter information, lighting processing on the video to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Video processing method and device, computer device, and storage medium

[0001] This application claims priority to Chinese Patent Application No. 202310996106.0, filed on August 8, 2023, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD

[0002] The present disclosure relates to a video processing method, device, computer device, and storage medium. BACKGROUND

[0003] With the rapid development of Internet technology, users have a demand for light processing when displaying videos using smart devices. Smart lighting aims to re-light the video frames in the video by three-dimensional perception of the video, increase the light and shadow aesthetics of the video scene, and make the light and dark relationship of the video more prominent, thereby increasing the atmosphere. Therefore, it is particularly important to propose a method for light processing of a video to meet the needs of users.

[0004] SUMMARY

[0005] The embodiments of the present disclosure at least provide a video processing method, device, computer device, and storage medium.

[0006] In a first aspect, the embodiments of the present disclosure provide a video processing method, comprising:

[0007] In response to receiving a light processing request for a to-be-processed video, display indication information, the indication information comprising first indication information for indicating a user to select a light and shadow rendering region, and second indication information for indicating the user to select light parameter information;

[0008] For a target video frame in the to-be-processed video, determine a target light and shadow rendering region selected by the user based on the first indication information, and target light parameter information selected based on the second indication information; wherein the target light and shadow rendering region and the target light parameter information determined for the target video frame are used for light processing of a target video segment associated with the target video frame;

[0009] Display a target video processed based on the target light and shadow rendering region and the target light parameter information.

[0010] In an optional implementation, displaying a target video processed based on the target light and shadow rendering region and the target light parameter information, comprises:

[0011] In response to the target light-and-shadow rendering region selected for the target video frame including a target region, when a target video clip associated with the target video frame in the target video is displayed, a special effect of lighting the target region content matching the target region in a video frame of the target video clip according to the target lighting parameter information is displayed;

[0012] The target region includes at least one of a background region, a foreground region, a target part region in the foreground region, and a global region. When the target region includes the background region, the target region content matching the background region includes background content of the background region and partial foreground content adjacent to the background content.

[0013] In an optional implementation, displaying the target video obtained by performing lighting processing on the to-be-processed video based on the target light-and-shadow rendering region and the target lighting parameter information includes:

[0014] In response to the target lighting parameter information including light source movement track information, when the target video clip associated with the target video frame in the target video is displayed, a dynamic light effect matching the light source movement track information is displayed on the target light-and-shadow rendering region.

[0015] In an optional implementation, performing lighting processing on the to-be-processed video based on the target light-and-shadow rendering region and the target lighting parameter information to obtain a target video includes:

[0016] Detecting a to-be-processed video frame in a target video clip included in the to-be-processed video, and generating an initial normal map and an initial segmentation map corresponding to the to-be-processed video frame;

[0017] Preprocessing the initial segmentation map to generate a target segmentation map corresponding to the to-be-processed video frame, and preprocessing the initial normal map to generate a target normal map corresponding to the to-be-processed video frame;

[0018] Performing lighting processing on the target light-and-shadow rendering region of the to-be-processed video frame according to the target lighting parameter information of the to-be-processed video frame based on the target normal map and the target segmentation map corresponding to the to-be-processed video frame, to generate a processed video frame;

[0019] Generating the target video based on the processed video frame.

[0020] In an optional implementation, the preprocessing the initial segmentation map to generate the target segmentation map corresponding to the to-be-processed video frame includes:

[0021] The target segmentation graph of the adjacent to-be-processed video frame before the to-be-processed video frame is fused with the initial segmentation graph of the to-be-processed video frame to generate the target segmentation graph corresponding to the to-be-processed video frame.

[0022] In an optional implementation, the preprocessing of the initial normal graph to generate the target normal graph corresponding to the to-be-processed video frame comprises:

[0023] The target normal graph of the adjacent to-be-processed video frame before the to-be-processed video frame is fused with the initial normal graph of the to-be-processed video frame to generate the target normal graph corresponding to the to-be-processed video frame.

[0024] In an optional implementation, the method further comprises:

[0025] In response to the user starting the following function, target object recognition is performed on the to-be-processed video frame to determine initial detection box information corresponding to an object region in the to-be-processed video frame.

[0026] The initial detection box information of the to-be-processed video frame is smoothed based on the target detection box information of the adjacent to-be-processed video frame before the to-be-processed video frame to generate target detection box information corresponding to the to-be-processed video frame.

[0027] The target light and shadow rendering region of the to-be-processed video frame is lighted based on the target normal graph and the target segmentation graph corresponding to the to-be-processed video frame and the target lighting parameter information of the to-be-processed video frame to generate a processed video frame.

[0028] The target light and shadow rendering region of the to-be-processed video frame is lighted based on the target detection box information, the target normal graph and the target segmentation graph corresponding to the to-be-processed video frame and the target lighting parameter information of the to-be-processed video frame to generate a processed video frame.

[0029] In an optional implementation, after the processed video frame is generated, the method further comprises:

[0030] In response to receiving an ambient light parameter, the ambient light of the processed video frame is adjusted based on the ambient light parameter to generate an adjusted video frame.

[0031] In a case where it is determined that the adjusted video frame has a target event, the ambient light parameter is adjusted to obtain a new ambient light parameter, and the step of adjusting the ambient light of the processed video frame based on the ambient light parameter to generate an adjusted video frame is returned, wherein the target event comprises an overexposure event or an overdarkness event.

[0032] determine the adjusted video frame as a target video frame in a case where it is determined that the adjusted video frame does not exist a target event;

[0033] The generating the target video based on the processed video frame comprises:

[0034] generating the target video based on the target video frame.

[0035] In a second aspect, the present disclosure provides a video processing apparatus, comprising:

[0036] The first display module is configured to, in response to receiving a light processing request for a to-be-processed video, display indication information, the indication information comprising first indication information for instructing a user to select a light and shadow rendering region, and second indication information for instructing the user to select light parameter information.

[0037] The determining module is configured to, for a target video frame in the to-be-processed video, determine a target light and shadow rendering region selected by the user based on the first indication information, and target light parameter information selected by the user based on the second indication information; wherein the target light and shadow rendering region and the target light parameter information determined for the target video frame are used for light processing of a target video clip associated with the target video frame.

[0038] The second display module is configured to display a target video obtained by performing light processing on the to-be-processed video based on the target light and shadow rendering region and the target light parameter information.

[0039] In a third aspect, the present disclosure provides a computer device, comprising a processor, a memory and a bus, the memory storing machine readable instructions executable by the processor, the processor and the memory being in communication through the bus when the computer device is running, and the machine readable instructions being executed by the processor to perform the steps of the first aspect or any possible implementation manner of the first aspect.

[0040] In a fourth aspect, the present disclosure provides a computer readable storage medium, the computer readable storage medium storing a computer program, the computer program being executed by a processor to perform the steps of the first aspect or any possible implementation manner of the first aspect.

[0041] In order to make the above objectives, features and advantages of the present disclosure more apparent, more comprehensible, the following will specifically describe preferred embodiments in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. The drawings incorporated into the specification and form a part of the specification, which show the embodiments consistent with the present disclosure, and are used to explain the technical solutions of the present disclosure together with the specification. It should be understood that the following drawings only show some of the embodiments of the present disclosure, and therefore should not be considered as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor.

[0043] FIG. 1 shows a flowchart of a video processing method provided by an embodiment of the present disclosure;

[0044] FIG. 2 shows an interface schematic diagram of a display interface in the video processing method provided by an embodiment of the present disclosure;

[0045] FIG. 3a shows an interface schematic diagram of a display interface indicating a light source position in the video processing method provided by an embodiment of the present disclosure;

[0046] FIG. 3b shows an interface schematic diagram of a display interface indicating a color bar in the video processing method provided by an embodiment of the present disclosure;

[0047] FIG. 3c shows an interface schematic diagram of a display interface indicating an intensity identification bar in the video processing method provided by an embodiment of the present disclosure;

[0048] FIG. 4 shows a flowchart of another video processing method provided by an embodiment of the present disclosure;

[0049] FIG. 5 shows a schematic diagram of a video processing device provided by an embodiment of the present disclosure;

[0050] FIG. 6 shows a structural schematic diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0051] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the following will combine the drawings in the embodiments of the present disclosure to make a clear and complete description of the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, but not all the embodiments. The components of the embodiments of the present disclosure described and shown in the drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the claimed present disclosure, but only represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present disclosure.

[0052] With the rapid development of Internet technology, when users display videos by using smart devices, there is a demand for light processing of the displayed videos. Among them, smart lighting aims to re-light the video frames in the video by three-dimensional perception of the video, increase the light and shadow aesthetics of the video scene, make the highlight relationship of the video picture more prominent, and increase the atmosphere.

[0053] Based on this, the present disclosure provides a video processing method, in response to receiving a light processing request for a to-be-processed video, display indication information, the indication information includes first indication information for indicating that a user selects a light rendering area, and second indication information for indicating that a user selects light parameter information;Through the indication information, the user determines the required target light rendering area and the target light parameter information, and provides personalized data support for subsequent light processing. Further, after determining the target light rendering area selected by the user based on the first indication information and the target light parameter information selected by the user based on the second indication information for the target video frame in the to-be-processed video, the to-be-processed video can be light processed based on the target light rendering area and the target light parameter information, to obtain a target video, and the target video is displayed, realizing video light processing, and different users can realize different light processing, and the flexibility of light processing is high.

[0054] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0055] The term "and / or" in this paper only describes a relationship, which means that there are three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, the term "at least one" in this paper means any one of the multiple or any combination of at least two of the multiple, for example, including at least one of A, B and C, which means including any one or more elements selected from the set consisting of A, B and C.

[0056] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, and scenario of use of personal information involved in the present disclosure should be informed to the user and authorized by the user in accordance with relevant laws and regulations.

[0057] For example, in response to receiving a user's active request, send prompt information to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware that performs the operation of the technical solution of the present disclosure, such as electronic devices, application programs, servers or storage media.

[0058] As an optional but non-limiting implementation, in response to receiving the active request of the user, the manner of sending the prompt information to the user may be, for example, a pop-up window manner, in which the prompt information may be presented in the form of text. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0059] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0060] For the convenience of understanding the present embodiment, first, a video processing method disclosed by the present embodiment is described in detail. The execution subject of the video processing method provided by the present embodiment is generally a computer device with certain computing power, which may, for example, include a terminal device, which can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, etc. In some possible implementation, the video processing method can be realized by a processor calling computer readable instructions stored in a memory.

[0061] The video processing method provided by the present embodiment is described below with the terminal device as an example.

[0062] Referring to FIG. 1, a flowchart of the video processing method provided by the present embodiment is shown, and the method includes S101-S103, wherein:

[0063] S101, in response to receiving a lighting processing request for a to-be-processed video, display indication information, the indication information including first indication information for indicating the user to select a light and shadow rendering area, and second indication information for indicating the user to select lighting parameter information.

[0064] The to-be-processed video can be a video collected by the camera of the terminal device in real time, or a video uploaded to the terminal device. After obtaining the to-be-processed video, a lighting processing request for the to-be-processed video is triggered, and the triggering of the lighting processing request can be in various ways, which are not limited by the present disclosure. For example, a "lighting" function button can be displayed on the display interface, and the "lighting" function button is clicked to trigger the lighting request. Alternatively, the to-be-processed video can also be long-pressed to trigger the lighting processing request, and the like.

[0065] The indication information includes first indication information and second indication information. The first indication information is used to indicate that the user selects a light and shadow rendering area, such as a light and shadow rendering area including but not limited to: a global area, a background area, a foreground area, a target part area in the foreground area, a self-defined local area, etc. Exemplarily, a plurality of light and shadow rendering areas can be displayed in the form of a list; or a sample card matched with each light and shadow rendering area can also be constructed, each sample card displaying a light and shadow rendering area and a matched lighting sample picture, and the sample cards are displayed in the form of cards for the user to select.

[0066] The second indication information is used to indicate that the user selects lighting parameter information, wherein the lighting parameter information can be set as needed, such as lighting parameter information including but not limited to: the number of light sources, the position of the light source, the color of the light source, the intensity of the light source, the range of the light source, the intensity of the dark part, the intensity of the dark light of the whole picture, the moving track of the light source, the moving speed of the light source, the rotating speed of the light source, the color changing speed of the light source, etc.

[0067] Exemplarily, the second indication information can include a number input box, so that the user can input the required number of light sources. Alternatively, the second indication information can include "please set the light source on the display interface", so that the user can trigger the generation of the required light source on the interface according to the second indication information, and generate one light source each time the trigger is generated. The trigger position is determined as the light source position, realizing the construction of at least one light source. A color bar (i.e. the second indication information matched with the color of the light source) can also be displayed, so that the user can perform a trigger operation on the color bar to determine the light source color of each light source. The display form of the second indication information has various forms, which are not limited in the present disclosure.

[0068] S102, for the target video frame in the to-be-processed video, respectively determine the target light and shadow rendering area selected by the user based on the first indication information, and the target lighting parameter information selected based on the second indication information; wherein the target light and shadow rendering area and the target lighting parameter information determined for the target video frame are used to perform lighting processing on a target video segment associated with the target video frame.

[0069] The target video frame can be a key frame in the to-be-processed video, and the key frame can be determined by decoding the to-be-processed video. In implementation, after obtaining the to-be-processed video, the to-be-processed video and the target video frame included in the to-be-processed video can be displayed, as shown in FIG. 2. By displaying the target video frame included in the to-be-processed video frame, the user can determine the matching target light and shadow rendering region and the target lighting parameter information for the target video frame. The target light and shadow rendering region and / or the target lighting parameter information corresponding to a plurality of target video frames can be the same or different, that is, the present disclosure supports determining the target light and shadow rendering region and the target lighting parameter information for each of the plurality of target video frames, and also supports the plurality of target video frames sharing the same target light and shadow rendering region and target lighting parameter information.

[0070] The target video frame can be associated with a target video segment. After determining the target light and shadow rendering region and the target lighting parameter information corresponding to the target video frame, the target light and shadow rendering region and the target lighting parameter information corresponding to the target video frame can be determined as the target light and shadow rendering region and the target lighting parameter information matched by the associated target video segment, so as to perform lighting processing on the video frames in the target video segment associated with the target video frame by using the target light and shadow rendering region and the target lighting parameter information determined for the target video frame.

[0071] In implementation, the user can select a target video frame, and determine the target light and shadow rendering region according to the first indication information and determine the target lighting parameter information according to the second indication information for the selected target video frame. The target lighting parameter information is the specific content of the lighting parameter, for example, the target lighting parameter information can include: light source color-yellow, light source position-coordinate (x, y), etc.

[0072] S103, display a target video obtained by performing lighting processing on the to-be-processed video based on the target light and shadow rendering region and the target lighting parameter information.

[0073] After determining the target light and shadow rendering region and the target lighting parameter information, the to-be-processed video is processed by lighting based on the target light and shadow rendering region and the target lighting parameter information, and the target video is obtained and displayed.

[0074] In an optional implementation, displaying a target video obtained by performing lighting processing on the to-be-processed video based on the target light and shadow rendering region and the target lighting parameter information includes: in response to the target light and shadow rendering region selected for the target video frame including a target region, when displaying a target video segment associated with the target video frame in the target video, displaying a special effect obtained by performing lighting on target region content matching the target region in the video frame of the target video segment according to the target lighting parameter information.

[0075] The target region includes at least one of the following: a background region, a foreground region, a target part region in the foreground region, and a global region; when the target region includes the background region, the target region content matched by the background region includes background content of the background region and partial foreground content adjacent to the background content.

[0076] When the target region includes the background region, when displaying the target video segment associated with the target video frame in the target video, a special effect is displayed in which the background content of the video frame of the target video segment and the partial foreground content adjacent to the background content are lighted according to the target lighting parameter information. In this disclosure, when the background region is lighted, in order to increase the realism of the lighting effect, the background content of the video frame and the partial foreground content adjacent to the background content can be lighted to achieve the visual effect that the background and the partial foreground are illuminated.

[0077] When the target region includes the foreground region or the target part region in the foreground region, when displaying the target video segment associated with the target video frame in the target video, a special effect is displayed in which the foreground content in the foreground region or the target part region is lighted according to the target lighting parameter information. The target part in the foreground region can be determined according to the first indication information, such as when the foreground region is a human body region, the target part can be a face part, a head part, etc.

[0078] When the target region includes the global region, when displaying the target video segment associated with the target video frame in the target video, a special effect is displayed in which the global content of the video frame of the target video segment is lighted according to the target lighting parameter information, that is, the entire video frame is lighted.

[0079] Here, by setting multiple target regions, flexible lighting processing of the video to be processed is realized, the diversity of video lighting is improved, and user experience is increased.

[0080] In an optional implementation, displaying the target video obtained by performing lighting processing on the video to be processed based on the target light and shadow rendering region and the target lighting parameter information includes: in response to the target lighting parameter information including light source movement track information, when displaying the target video segment associated with the target video frame in the target video, displaying a dynamic lighting effect matched with the light source movement track information on the target light and shadow rendering region.

[0081] In implementation, the light source moving track can be set, such as a preset track can be set as the light source moving track, such as a surrounding track around the head region, a surrounding track around the face region, etc. The user can also be prompted to draw the light source moving track on the display interface through the second indication information. When the set light source is multiple, the multiple light sources can correspond to different light source moving tracks, or can match the same light source moving track.

[0082] When the target video segment associated with the target video frame in the target video is displayed, the dynamic light effect matched with the light source moving track information can be displayed on the target light and shadow rendering region, such as the dynamic light effect can be the light effect of highlighting different regions of the face when the light source moving track information is the surrounding track around the face region.

[0083] Here, the light source moving track information can be set to achieve the effect of adding a dynamic light effect on the to-be-processed video, and improve the video lighting effect.

[0084] The following exemplary describes various lighting scenarios implemented by the method proposed in the present disclosure.

[0085] Scenario one: basic lighting-all-around lighting. In implementation, the target light and shadow rendering region can be determined as a global region according to the first indication information, and the target lighting parameter information when the global region is lighted can be determined according to the second indication information. In this scenario, one or more light sources can be added for each target video frame in the to-be-processed video frame, and one or more light source parameters can be determined for each light source, such as light source position, light source color, light source intensity, light source range, dark part intensity, and overall dark light degree. When the to-be-processed video is lighted in this scenario, the background and the foreground can be illuminated at the same time, and there can be a shadow effect on the background.

[0086] Exemplarily, a preset graph such as a circular graph can be set to represent the light source position, so as to realize the adjustment of the light source position through the moving operation of the preset graph on the display interface. Referring to FIG. 3a, the light source position is shown in FIG. 3a. For example, a color bar including gradient color information can also be set, the user can select the required light source color on the color bar, and the light source color can be adjusted through the sliding operation on the color gradient area. Referring to FIG. 3b, the color bar is shown in FIG. 3b. For another example, an intensity identifier bar can be set, the light source intensity can be adjusted by adjusting the position of the adjustment identifier on the intensity identifier bar. Referring to the intensity identifier bar shown in FIG. 3c. The adjustment mode of the light source range, the dark part intensity, and the overall dark light degree can refer to the adjustment mode of the light source intensity, which is not described here.

[0087] Scene two: base lighting - foreground lighting. When implemented, the target light and shadow rendering area can be determined as the foreground area or a target part on the foreground area according to the first instruction information, and the target lighting parameter information when the foreground area is lit can be determined according to the second instruction information. In this scene, one or more light sources can be added for each target video frame in the video frame to be processed, and one or more of the following light source parameters can be determined for each light source: light source position, light source color, light source intensity, light source range, dark part intensity, and overall dark light degree. When lighting the video to be processed in this scene, the lighting effect of lighting the front and side of the foreground can be achieved. When lighting, the light source position can be fixed to achieve a fixed position lighting effect, or the light source position can be set to move according to the foreground object to achieve a moving lighting effect.

[0088] Scene three: base lighting - background lighting. When implemented, the target light and shadow rendering area can be determined as the background area according to the first instruction information, and the target lighting parameter information when the background area is lit can be determined according to the second instruction information. In this scene, one or more light sources can be added for each target video frame in the video frame to be processed, and one or more of the following light source parameters can be determined for each light source: light source position, light source color, light source intensity, light source range, dark part intensity, and overall dark light degree. When lighting the video to be processed in this scene, the background can be illuminated while part of the foreground area is illuminated, improving the lighting effect, and the shape of the background light source can be specified, such as a sunset lamp or a straight light tube.

[0089] Scene four - dynamic light effect. When implemented, the target light and shadow rendering area can be determined according to the first instruction information, and the target light and shadow rendering area in this scene can be any of the foreground area, the background area, and the global area. The target lighting parameter information is determined according to the second instruction information, and the start point, end point, and track point of the light source movement trajectory information can be adjusted. The dynamic light effect can include a variety of effects, and the light source movement trajectory information and light source parameters can be determined according to the effect to be achieved by the dynamic light effect. The dynamic light effect can be set as needed, such as a rotating light effect, a rainbow light effect, a glowing butterfly light effect, a flashing light effect, a music light effect, etc.

[0090] For example, when the dynamic light effect includes a rotating light effect, the light source movement trajectory information can indicate that the light source rotates around the top of the foreground object or around the face area rendering. The light source parameters include: starting light source position, light source color, light source intensity, light source range, dynamic light source rotation speed, dynamic light effect color change speed, etc.

[0091] For example, when the dynamic light effect includes a rainbow light effect, the light source randomly and dynamically moves to illuminate the human face under the dynamic light effect. In this case, the light source moves along a random trajectory, and the light source parameters include light source color richness, light source intensity, light source range, and dynamic light source random movement speed.

[0092] For example, when the dynamic light effect is a glowing butterfly light effect, the shape of the light source is in the shape of a butterfly, and the light source parameters such as light source color and light source intensity are set. For example, when the dynamic light effect is a flashing light effect, the light source intensity can change at a predetermined frequency to achieve a flashing effect. For example, under a music light effect, the music beat can be identified, and the movement speed and intensity of the light source can be adjusted according to the music beat.

[0093] The following describes the lighting processing procedure in detail.

[0094] In an optional embodiment, based on the target light and shadow rendering area and the target lighting parameter information, the lighting processing is performed on the to-be-processed video to obtain a target video, including:

[0095] In step a1, the to-be-processed video frame in the target video segment included in the to-be-processed video is detected to generate an initial normal map and an initial segmentation map corresponding to the to-be-processed video frame.

[0096] In step a2, the initial segmentation map is preprocessed to generate a target segmentation map corresponding to the to-be-processed video frame, and the initial normal map is preprocessed to generate a target normal map corresponding to the to-be-processed video frame.

[0097] In step a3, based on the target normal map and the target segmentation map corresponding to the to-be-processed video, the target light and shadow rendering area of the to-be-processed video frame is processed according to the target lighting parameter information of the to-be-processed video frame to generate a processed video frame.

[0098] In step a4, the target video is generated based on the processed video frame.

[0099] In step a1, the to-be-processed video frame in the target video segment included in the to-be-processed video can be detected by using a multi-task neural network model to generate an initial normal map and an initial segmentation map corresponding to the to-be-processed video frame. The network structure of the multi-task neural network model can be determined as needed. The to-be-processed video frame can be each video frame included in the target video segment, or at least part of the video frames in the target video segment.

[0100] The size of the initial normal map and the initial segmentation map can be the same as the size of the to-be-processed video frame. The pixel information of each pixel point in the initial normal map represents the normal information of a pixel point at the same pixel position in the to-be-processed video frame. The pixel information of each pixel point in the initial segmentation map represents the object category to which a pixel point at the same pixel position in the to-be-processed video frame belongs. The object category can include a foreground category and a background category, or can include specific object categories and a background category, such as a human body, an animal, a plant, and the like.

[0101] In step a2, in order to improve the accuracy of the segmentation map, the initial segmentation map can be preprocessed, such as blurring the initial segmentation map and the like, to obtain a target segmentation map. In order to improve the accuracy of the normal map, the initial normal map can be preprocessed, such as guided filtering processing, denoising processing and the like, to obtain a target normal map. When the target normal map and the target segmentation map obtained by processing are used for lighting processing, the target video obtained is smoother and has better display effect. In addition, the initial segmentation map and the initial normal map can also be preprocessed using the target segmentation map and the target normal map of the historical to-be-processed video frame before the to-be-processed video frame.

[0102] In an optional implementation, the preprocessing of the initial segmentation map to generate the target segmentation map corresponding to the to-be-processed video frame includes fusing the target segmentation map of the adjacent to-be-processed video frame before the to-be-processed video frame with the initial segmentation map of the to-be-processed video frame to generate the target segmentation map corresponding to the to-be-processed video frame.

[0103] In order to alleviate the segmentation mutation problem between adjacent video frames and improve the video smoothing effect, the target segmentation map of the previous to-be-processed video frame can be fused with the initial segmentation map (or the processed segmentation map obtained after blurring and the like) of the current to-be-processed video frame to generate the target segmentation map corresponding to the to-be-processed video frame. The fusion processing can be fusion according to a first preset ratio, such as summing the pixel information at the same pixel position in the target segmentation map of the adjacent to-be-processed video frame before the to-be-processed video frame and the initial segmentation map of the to-be-processed video frame according to a ratio to obtain the fusion pixel information at the pixel position, and obtaining the target segmentation map corresponding to the to-be-processed video frame based on the fusion pixel information at each pixel position.

[0104] In the present disclosure, the initial segmentation map of the to-be-processed video frame is fused with the target segmentation map of the previous to-be-processed video frame to realize the fusion of the segmentation information in time sequence between adjacent to-be-processed video frames in the to-be-processed video, so as to improve the accuracy of the target segmentation map, so that the target video with better display effect can be obtained based on the target segmentation map.

[0105] In an optional implementation, the preprocessing of the initial normal map to generate the target normal map corresponding to the to-be-processed video frame comprises: fusing the target normal map of a neighboring to-be-processed video frame before the to-be-processed video frame and the initial normal map of the to-be-processed video frame to generate the target normal map corresponding to the to-be-processed video frame.

[0106] In implementation, the target normal map of the neighboring to-be-processed video frame before the to-be-processed video frame can be fused with the initial normal map of the to-be-processed video frame, or the initial normal map of the to-be-processed video frame can be first processed, such as guided filter processing, denoising processing, etc., and the processed normal map can be fused with the target normal map of the previous to-be-processed video frame.

[0107] The fusion processing can be fusion according to a second preset ratio, for example, pixel information at the same pixel position in the target normal map of the neighboring to-be-processed video frame before the to-be-processed video frame and the initial normal map (or the processed normal map) of the to-be-processed video frame is summed according to a ratio to obtain fusion normal information at the pixel position, and the target normal map corresponding to the to-be-processed video frame is obtained based on the fusion normal information at each pixel position.

[0108] In the present disclosure, the initial normal map of the to-be-processed video frame is fused with the target normal map of the previous to-be-processed video frame to realize fusion of normal information in time sequence between neighboring to-be-processed video frames in the to-be-processed video, so as to improve the accuracy of the target normal map, so that a target video with better display effect can be obtained based on the target normal map.

[0109] In step a3, after obtaining the target normal map and the target segmentation map corresponding to the to-be-processed video frame, the pixel values of the pixel points in the target segmentation map can be first adjusted based on the target light and shadow rendering area, for example, the pixel values of the target pixel points in the target segmentation map that match the target light and shadow rendering area are adjusted to a first value such as 1, and the pixel values of the other pixel points except the target pixel points are adjusted to a second value such as 0 to obtain an adjusted segmentation map. The adjusted segmentation map is fused with the target normal map to obtain a fused normal map, and then the to-be-processed video frame can be lighted using a light model such as a Lambert light model in combination with the target lighting parameter information of the to-be-processed video frame and the fused normal map to obtain a processed video frame.

[0110] For example, the normal vector of each pixel point in the to-be-processed video frame can be determined according to the fused normal map, and the reverse vector of the incident light can be determined according to the target lighting parameter information, and the light intensity of the reflection light corresponding to each pixel point in the to-be-processed video frame can be determined according to the reverse vector, the normal vector, and the determined light intensity of the incident light. Based on the light intensity of the reflection light of each pixel point, the pixel information of each pixel point in the to-be-processed video frame is adjusted to obtain a processed video frame.

[0111] In step a4, the plurality of processed video frames are encoded and arranged to obtain a target video. Alternatively, when there are other video frames that have not been processed by lighting, the plurality of processed video frames and the other video frames can be encoded and arranged according to the time sequence information to obtain a target video.

[0112] Here, by using the obtained target depth map and target segmentation map, and combining the target lighting parameter information, the lighting processing of the to-be-processed video frame is realized, so that a target video meeting the user's demand can be obtained subsequently.

[0113] In an optional implementation, a follow-up function can be preset, and when the user starts the follow-up function, the target object followed by the light source can be processed. In specific implementation, the method further includes:

[0114] In response to the user starting the follow-up function, target object recognition is performed on the to-be-processed video frame to determine initial detection box information corresponding to the object region in the to-be-processed video frame; and the initial detection box information of the to-be-processed video frame is smoothed based on the target detection box information of the adjacent to-be-processed video frame before the to-be-processed video frame to generate target detection box information corresponding to the to-be-processed video frame.

[0115] The user can start the follow-up function on the display interface, such as clicking a follow-up function button to start the follow-up function, and when the follow-up function is started, the target object of the follow-up function can also be set, so as to perform target object recognition on the to-be-processed video frame. The target object can be determined according to the video content, such as a face, a pet face, etc.

[0116] In response to the user starting the follow-up function, target object recognition is performed on the to-be-processed video frame to obtain initial detection box information of the object region (i.e., the region where the target object is located) in the to-be-processed video frame. The initial detection box information can include the pixel coordinate information of the four vertices of the detection box.

[0117] In order to alleviate the jitter problem caused by the position mutation of the target object in the adjacent to-be-processed video frame, the initial detection box information of the to-be-processed video frame can be smoothed based on the target detection box information of the adjacent to-be-processed video frame before the to-be-processed video frame. For example, the pixel coordinate information of the top-left corner indicated by the target detection box information of the adjacent to-be-processed video frame and the pixel coordinate information of the top-left corner indicated by the initial detection box information of the to-be-processed video frame are averaged to obtain the processed pixel coordinate information of the top-left corner. Similarly, the processed pixel coordinate information of the four corners can be obtained, and the processed pixel coordinate information of the four corners constitutes the target detection box information corresponding to the to-be-processed video frame.

[0118] Further, the target light and shadow rendering region of the to-be-processed video frame can be lighted based on the target detection box information corresponding to the to-be-processed video frame, the target normal map and the target segmentation map, according to the target lighting parameter information of the to-be-processed video frame, to generate a processed video frame.

[0119] The target light and shadow rendering region matches the target object included in the target detection box, such as when the target detection box includes a face region, the target light and shadow rendering region is the face region. In implementation, the pixel values of the pixel points of the target segmentation map can be adjusted based on the target detection box information, such as adjusting the pixel values of the target pixel points in the target detection box of the target segmentation map that match the target light and shadow rendering region to a first value such as 1, and adjusting the pixel values of other pixel points except the target pixel points to a second value such as 0, to obtain an adjusted segmentation map.

[0120] The adjusted segmentation map is then fused with the target normal map to obtain a fused normal map, and the to-be-processed video frame can be lighted using a lighting model such as a Lambert lighting model, in combination with the target lighting parameter information of the to-be-processed video frame and the fused normal map, to obtain a processed video frame.

[0121] Here, by setting a following function such as a face following function, the light processing can be performed on the object of the following function, and the initial detection box information of the to-be-processed video frame is smoothed using the target detection box information of the adjacent to-be-processed video frame before the to-be-processed video frame to generate the target detection box information corresponding to the to-be-processed video frame, so that the target object can be accurately lighted based on the target detection box information, the jitter problem of the target object is alleviated, and the video lighting effect is improved.

[0122] After obtaining the processed video frame, ambient light can also be added to the processed video frame to improve the video lighting effect.

[0123] In an alternative embodiment, after generating the processed video frame, the method further comprises: in response to receiving the ambient light parameter, performing ambient light adjustment on the processed video frame based on the ambient light parameter to generate an adjusted video frame; in a case where the adjusted video frame is determined to have a target event, adjusting the ambient light parameter to obtain a new ambient light parameter, and returning to the step of performing ambient light adjustment on the processed video frame based on the ambient light parameter to generate an adjusted video frame, wherein the target event comprises an overexposure event or an overdark event; in a case where the adjusted video frame is determined not to have a target event, determining the adjusted video frame as a target video frame.

[0124] The generating the target video based on the processed video frame comprises: generating the target video based on the target video frame.

[0125] In implementation, the user can input the ambient light parameter, and in response to the received ambient light parameter, perform ambient light adjustment on the processed video frame based on the ambient light parameter, i.e., shield the processed video frame with the selected ambient light to generate an adjusted video frame.

[0126] It is determined whether the adjusted video frame has a target event, i.e., whether there is an overexposure event and / or an overdark event, and if not, the adjusted video frame is determined as a target video frame; if so, the ambient light parameter needs to be re-determined, and ambient light adjustment is performed on the processed video frame with the new ambient light parameter until an adjusted video frame without a target event is obtained. The ambient light parameter can be re-determined according to the target event, for example, when the target event is an overexposure event, the re-determined ambient light parameter has a lower brightness than the previous ambient light parameter; conversely, if the target event is an overdark event, the re-determined ambient light parameter has a higher brightness than the previous ambient light parameter.

[0127] In implementation, it can be determined whether the adjusted video frame has a target event according to the following process: determining respective event threshold values corresponding to each pixel point in the adjusted video frame, the event threshold values comprising a first event threshold value matched with the overexposure event and a second event threshold value matched with the overdark event; determining a first number of overexposure pixel points in the adjusted video frame whose pixel values are greater than the first event threshold value; and determining whether the adjusted video frame has an overexposure event according to the first number of overexposure pixel points in the adjusted video frame; and determining a second number of overdark pixel points in the adjusted video frame whose pixel values are less than the second event threshold value; and determining whether the adjusted video frame has an overexposure event according to the second number of overdark pixel points in the adjusted video frame.

[0128] In implementation, for each pixel in the adjusted video frame, the pixel information at the pixel position in the to-be-processed video frame can be determined according to the pixel position of the pixel, and the first event threshold and the second event threshold corresponding to the pixel can be determined according to the pixel information and the set overexposure ratio and underexposure ratio. For example, the pixel information can be multiplied by the overexposure ratio to obtain a first product, and the first product can be added to the pixel information to obtain a sum value, which is taken as the first event threshold. In addition, the pixel information can be multiplied by the underexposure ratio to obtain a second product, and the pixel information can be subtracted by the second product to obtain a difference value, which is taken as the second event threshold. Then, the event thresholds corresponding to the respective pixels in the adjusted video frame can be obtained, and the event thresholds include the first event threshold and the second event threshold.

[0129] The first pixels with pixel values greater than the first event threshold in the adjusted video frame can be counted, the first pixels are determined as overexposed pixels, and the first number of the overexposed pixels is determined. If the first number is greater than a first preset number value, it is determined that there is an overexposure event, otherwise, it is determined that there is no overexposure event. Alternatively, the overexposure ratio can be determined according to the first number and the total number of pixels in the adjusted video frame. If the overexposure ratio is greater than a first preset ratio, it is determined that there is an overexposure event, otherwise, it is determined that there is no overexposure event.

[0130] In addition, the second pixels with pixel values less than the second event threshold in the adjusted video frame can be counted, the second pixels are determined as underexposed pixels, and the second number of the underexposed pixels is determined. If the second number is greater than a second preset number value, it is determined that there is an underexposure event, otherwise, it is determined that there is no underexposure event. Alternatively, the underexposure ratio can be determined according to the second number and the total number of pixels in the adjusted video frame. If the underexposure ratio is greater than a second preset ratio, it is determined that there is an underexposure event, otherwise, it is determined that there is no underexposure event.

[0131] Finally, the plurality of target video frames can be arranged according to the time sequence information to generate a target video.

[0132] The present disclosure adds ambient light in the processed video frame, so that the generated target video frame meets the scene requirements, and the underexposure detection or overexposure detection can be performed after the ambient light is added, so as to alleviate the problem of poor display effect caused by overexposure or underexposure of the adjusted video frame, and ensure the video processing effect.

[0133] Referring to the flowchart shown in FIG. 4, the video processing method is exemplarily described in combination with FIG. 4. The to-be-processed video can be a video including a user. The method can include:

[0134] First, the target video frame is displayed. Specifically, after receiving the to-be-processed video, the to-be-processed video frame in the to-be-processed video can be determined and displayed.

[0135] Second, normal mode, segmentation mode. Specifically, the to-be-processed video frame is input into the normal model and the segmentation model for detection to obtain an initial normal map and an initial segmentation map of the to-be-processed video frame.

[0136] Third, face detection. Here, if the user has enabled the face following function, face detection is performed, such as face detection using a face detection model to obtain initial detection box information corresponding to a face region. If the user has not enabled the face following function, face detection is not required.

[0137] Fourth, receiving lighting mode. The lighting mode is used to indicate a target light and shadow rendering region.

[0138] Fifth, structure information processing. Specifically, the initial normal map output by the normal model can be preprocessed, such as guided filter processing and denoising processing, to obtain a processed normal map. The initial segmentation map output by the segmentation model can be preprocessed, such as blurring processing, to obtain a processed segmentation map.

[0139] Sixth, time sequence information fusion and parameter smoothing. Specifically, the process of time sequence information fusion includes: fusing the target segmentation map of the adjacent to-be-processed video frame before the to-be-processed video frame and the processed segmentation map of the to-be-processed video frame to generate a target segmentation map corresponding to the to-be-processed video frame. The process also includes fusing the target normal map of the adjacent to-be-processed video frame before the to-be-processed video frame and the processed normal map of the to-be-processed video frame to generate a target normal map corresponding to the to-be-processed video frame.

[0140] The process of parameter smoothing includes: smoothing the initial detection box information of the to-be-processed video frame based on the target detection box information of the adjacent to-be-processed video frame before the to-be-processed video frame to generate target detection box information corresponding to the to-be-processed video frame.

[0141] The process of parameter smoothing is performed when the face following function is enabled; otherwise, if the face following function is not enabled, the process of parameter smoothing is not performed.

[0142] Seventh, light and shadow rendering. Specifically, target lighting parameter information is received, and light and shadow processing is performed according to the target lighting parameter information, the target segmentation map, and the target normal map to obtain a processed video frame.

[0143] If the face following function is enabled, light and shadow processing is performed according to the target lighting parameter information, the target detection box information, the target segmentation map, and the target normal map to obtain a processed video frame.

[0144] Eighth, ambient light adjustment. Specifically, an ambient light parameter is received, and ambient light adjustment is performed on the processed video frame based on the ambient light parameter to generate an adjusted video frame.

[0145] Ninth, overexposure event detection, overdark event detection. Specifically, it is judged whether the adjusted video frame has overexposure event and overdark event, if not, the tenth step is executed, if yes, the ambient light parameter is re-determined, and returns to the eighth step.

[0146] Tenth, output target video frame. Wherein the target video includes target video frame.

[0147] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process, and the specific execution order of each step should be determined by its function and possible internal logic.

[0148] Based on the same inventive concept, the video processing method in the embodiments of the present disclosure also provides a video processing device corresponding to the video processing method. Since the principle of solving problems of the device in the embodiments of the present disclosure is similar to the above-mentioned video processing method, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be described here.

[0149] Referring to FIG. 5, it is a schematic diagram of the architecture of a video processing device provided by the embodiments of the present disclosure, the device comprises: a first display module 501, a determination module 502, a second display module 503; wherein,

[0150] The first display module 501 is configured to, in response to receiving a lighting processing request for a to-be-processed video, display indication information, the indication information comprising first indication information for indicating that a user selects a light and shadow rendering region, and second indication information for indicating that the user selects lighting parameter information;

[0151] The determination module 502 is configured to determine, for a target video frame in the to-be-processed video, a target light and shadow rendering region selected by the user based on the first indication information, and target lighting parameter information selected based on the second indication information; wherein the target light and shadow rendering region and the target lighting parameter information determined for the target video frame are used for lighting processing of a target video segment associated with the target video frame;

[0152] The second display module 503 is configured to display a target video obtained by lighting processing of the to-be-processed video based on the target light and shadow rendering region and the target lighting parameter information.

[0153] In an optional implementation, the second display module 503, when displaying the target video obtained by lighting processing of the to-be-processed video based on the target light and shadow rendering region and the target lighting parameter information, is configured to:

[0154] in response to the target light-and-shadow rendering region selected for the target video frame comprising a target region, when a target video clip associated with the target video frame in the target video is displayed, a special effect of lighting the target region content matching the target region in a video frame of the target video clip according to the target lighting parameter information is displayed;

[0155] wherein the target region comprises at least one of the following: a background region, a foreground region, a target part region in the foreground region, and a global region; in a case where the target region comprises the background region, the target region content matching the background region comprises background content of the background region and partial foreground content adjacent to the background content.

[0156] In an optional implementation, the second display module 503, when displaying the target video obtained by lighting processing on the to-be-processed video based on the target light-and-shadow rendering region and the target lighting parameter information, is configured to:

[0157] in response to the target lighting parameter information comprising light source movement trajectory information, when the target video clip associated with the target video frame in the target video is displayed, a dynamic light effect matching the light source movement trajectory information is displayed on the target light-and-shadow rendering region.

[0158] In an optional implementation, the second display module 503, when displaying the target video obtained by lighting processing on the to-be-processed video based on the target light-and-shadow rendering region and the target lighting parameter information, is configured to:

[0159] detecting each to-be-processed video frame in a target video clip contained in the to-be-processed video, to generate an initial normal map and an initial segmentation map corresponding to the to-be-processed video frame;

[0160] preprocessing the initial segmentation map to generate a target segmentation map corresponding to the to-be-processed video frame, and preprocessing the initial normal map to generate a target normal map corresponding to the to-be-processed video frame;

[0161] lighting processing, according to the target lighting parameter information of the to-be-processed video frame, on the target light-and-shadow rendering region of the to-be-processed video frame based on the target normal map and the target segmentation map corresponding to the to-be-processed video frame, to generate a processed video frame;

[0162] generating the target video based on the processed video frame.

[0163] In an optional implementation, the second display module 503, when preprocessing the initial segmentation map to generate the target segmentation map corresponding to the to-be-processed video frame, is configured to:

[0164] The target segmentation graph of the adjacent to-be-processed video frame before the to-be-processed video frame is fused with the initial segmentation graph of the to-be-processed video frame to generate the target segmentation graph corresponding to the to-be-processed video frame.

[0165] In an optional implementation, the second display module 503 is configured to:

[0166] The target normal graph of the adjacent to-be-processed video frame before the to-be-processed video frame is fused with the initial normal graph of the to-be-processed video frame to generate the target normal graph corresponding to the to-be-processed video frame.

[0167] In an optional implementation, the apparatus further includes a detection module 504 configured to:

[0168] In response to the user starting the following function, the target object recognition is performed on the to-be-processed video frame to determine initial detection box information corresponding to an object region in the to-be-processed video frame.

[0169] The initial detection box information of the to-be-processed video frame is smoothed based on the target detection box information of the adjacent to-be-processed video frame before the to-be-processed video frame to generate target detection box information corresponding to the to-be-processed video frame.

[0170] The second display module 503 is configured to:

[0171] The target normal graph and the target segmentation graph corresponding to the to-be-processed video frame are used to perform light processing on the target light and shadow rendering region of the to-be-processed video frame according to the target light parameter information of the to-be-processed video frame to generate a processed video frame.

[0172] In an optional implementation, after the processed video frame is generated, the apparatus further includes an adjustment module 505 configured to:

[0173] In response to receiving an ambient light parameter, the ambient light adjustment is performed on the processed video frame based on the ambient light parameter to generate an adjusted video frame.

[0174] In a case where it is determined that the adjusted video frame has a target event, the ambient light parameter is adjusted to obtain a new ambient light parameter, and the step of adjusting the ambient light of the processed video frame based on the ambient light parameter to generate an adjusted video frame is returned to, wherein the target event includes an overexposure event or an overdark event.

[0175] In a case where it is determined that the adjusted video frame has no target event, the adjusted video frame is determined as a target video frame.

[0176] The generating the target video based on the processed video frame includes:

[0177] The generating the target video based on the processed video frame includes:

[0178] The description of the processing flow of each module in the device and the interaction flow between the modules can refer to the related description in the above method embodiments, and will not be described in detail here.

[0179] Based on the same technical concept, the present disclosure also provides a computer device. Referring to FIG. 6, it is a structural schematic diagram of a computer device 600 provided by the present disclosure. The computer device in the present disclosure can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like, or various forms of servers, such as independent servers or server clusters. The computer device shown in FIG. 6 is only an example, and should not bring any limitation to the functions and use range of the present disclosure.

[0180] As shown in FIG. 6, the computer device 600 can include a processing device (such as a central processor, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory device (ROM) 602 or programs loaded from a storage device 608 to a random access memory device (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0181] Generally, the following devices can be connected to the I / O interface 605: input devices 606, including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 607, including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 608, including, for example, a magnetic tape, a hard disk, and the like; and communication devices 609. The communication devices 609 can allow the computer device 600 to communicate with other devices wirelessly or through wires to exchange data. Although FIG. 6 shows the computer device 600 with various devices, it should be understood that all of the illustrated devices are not required to implement or have the computer device 600. More or fewer devices can alternatively be implemented or have the computer device 600.

[0182] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing a video processing method. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 609, or installed from the storage devices 608, or installed from the ROM 602. When the computer program is executed by the processing devices 601, the above-mentioned functions defined in the method of the embodiments of the present disclosure are performed.

[0183] Embodiments of the present disclosure also provide a computer readable storage medium, which stores a computer program, and the computer program is run by a processor to perform the steps of the video processing method described in the above method embodiments. The storage medium can be a volatile or non-volatile computer readable storage medium.

[0184] Embodiments of the present disclosure also provide a computer program product, which carries a program code, and the program code includes instructions for performing the steps of the video processing method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be described here.

[0185] The above computer program product can be specifically implemented by hardware, software or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium, and in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.

[0186] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the foregoing method embodiment, and will not be repeated here. In several embodiments provided in the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and another division can be made in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, and can be electrical, mechanical or other forms.

[0187] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0188] In addition, each functional unit in each embodiment of the present disclosure can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0189] If the functions are realized in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present disclosure essentially or the part of the prior art or the part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present disclosure. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various program code storage media.

[0190] Finally, it should be noted that the above-described embodiments are merely specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, and are not intended to limit the present disclosure. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can make modifications or easy changes to the technical solutions described in the foregoing embodiments, or easily think of changes or equivalent replacements for some of the technical features; and these modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A video processing method, comprising: in response to receiving a light processing request for a to-be-processed video, displaying indication information, the indication information including first indication information for indicating a user to select a light rendering region and second indication information for indicating the user to select light parameter information; for a target video frame in the to-be-processed video, determining a target light rendering region selected by the user based on the first indication information and target light parameter information selected based on the second indication information; wherein the target light rendering region and the target light parameter information determined for the target video frame are used for light processing on a target video segment associated with the target video frame; displaying a target video obtained by performing light processing on the to-be-processed video based on the target light rendering region and the target light parameter information.

2. The method of claim 1, wherein, The displaying of the target video obtained by performing light processing on the to-be-processed video based on the target light rendering region and the target light parameter information comprises: in response to the target light rendering region selected for the target video frame including a target region, when displaying a target video segment associated with the target video frame in the target video, displaying a special effect obtained by performing light processing on target region content in a video frame of the target video segment that matches the target region according to the target light parameter information; wherein the target region includes at least one of the following: a background region, a foreground region, a target part region in the foreground region, and a global region; in the case where the target region includes the background region, the target region content matching the background region includes background content of the background region and partial foreground content adjacent to the background content.

3. The method of claim 1, wherein, The displaying of the target video obtained by performing light processing on the to-be-processed video based on the target light rendering region and the target light parameter information comprises: in response to the target light parameter information including light source movement trajectory information, when displaying a target video segment associated with the target video frame in the target video, displaying a dynamic light effect matching the light source movement trajectory information on the target light rendering region.

4. The method according to any of claims 1 to 3, wherein, The performing of light processing on the to-be-processed video based on the target light rendering region and the target light parameter information to obtain a target video comprises: for a to-be-processed video frame in a target video segment included in the to-be-processed video, detecting the to-be-processed video frame to generate an initial normal map and an initial segmentation map corresponding to the to-be-processed video frame; preprocessing the initial segmentation map to generate a target segmentation map corresponding to the to-be-processed video frame, and preprocessing the initial normal map to generate a target normal map corresponding to the to-be-processed video frame; performing light processing on the target light rendering region of the to-be-processed video frame according to the target light parameter information of the to-be-processed video frame based on the target normal map and the target segmentation map corresponding to the to-be-processed video frame to generate a processed video frame; and generating the target video based on the processed video frame.

5. The method of claim 4, wherein, The preprocessing of the initial segmentation map comprises: The target segmentation map corresponding to the to-be-processed video frame is generated by fusing the target segmentation map of the adjacent to-be-processed video frame before the to-be-processed video frame and the initial segmentation map of the to-be-processed video frame.

6. The method of claim 4 or 5, wherein, The preprocessing of the initial normal map comprises: The target normal map corresponding to the to-be-processed video frame is generated by fusing the target normal map of the adjacent to-be-processed video frame before the to-be-processed video frame and the initial normal map of the to-be-processed video frame.

7. The method of any one of claims 4-6, further comprising: In response to the user starting the following function, performing target object recognition on the to-be-processed video frame to determine initial bounding box information corresponding to an object region in the to-be-processed video frame; Based on the target bounding box information of the adjacent to-be-processed video frame before the to-be-processed video frame, performing smoothing processing on the initial bounding box information of the to-be-processed video frame to generate target bounding box information corresponding to the to-be-processed video frame; The target light and shadow rendering area of the to-be-processed video frame is rendered according to the target normal map and the target segmentation map corresponding to the to-be-processed video frame and the target lighting parameter information of the to-be-processed video frame, and a processed video frame is generated. The target light and shadow rendering area of the to-be-processed video frame is rendered according to the target bounding box information, the target normal map and the target segmentation map corresponding to the to-be-processed video frame and the target lighting parameter information of the to-be-processed video frame, and a processed video frame is generated. After generating the processed video frame, further comprising:

8. The method according to any one of claims 4-7, wherein, In response to receiving an ambient light parameter, adjusting the processed video frame based on the ambient light parameter to generate an adjusted video frame; In a case where it is determined that the adjusted video frame has a target event, adjusting the ambient light parameter to obtain a new ambient light parameter, and returning to the step of adjusting the processed video frame based on the ambient light parameter to generate an adjusted video frame, wherein the target event includes an overexposure event or an overdark event; In a case where it is determined that the adjusted video frame does not have a target event, determining the adjusted video frame as a target video frame; The target video is generated based on the processed video frame, comprising: The target video is generated based on the target video frame.

9. A video processing apparatus, comprising: A first display module configured to display indication information in response to receiving a lighting processing request for a to-be-processed video, the indication information comprising first indication information for indicating that a user selects a light and shadow rendering area, and second indication information for indicating that a user selects lighting parameter information; ​ A determining module is configured to determine, for a target video frame in the to-be-processed video, a target light-and-shadow rendering region selected by a user based on the first indication information and target lighting parameter information selected based on the second indication information; wherein the target light-and-shadow rendering region and the target lighting parameter information determined for the target video frame are used for lighting processing of a target video clip associated with the target video frame. A second display module is configured to display a target video obtained by lighting processing of the to-be-processed video based on the target light-and-shadow rendering region and the target lighting parameter information.

10. A computer device comprising: A processor, a memory and a bus, the memory stores machine readable instructions executable by the processor, when the computer device is running, the processor and the memory communicate through the bus, and the machine readable instructions are executed by the processor to perform the steps of the video processing method according to any one of claims 1-8.

11. A computer-readable storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to perform the steps of the video processing method according to any one of claims 1-8.