Video processing method and device, electronic equipment and storage medium

By performing image segmentation and object motion trajectory analysis on video frames, the playback problems caused by damage or missing video frames are solved, and the video quality and user experience are improved.

CN120017911APending Publication Date: 2025-05-16KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510193998.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

During video transmission, decoding or playback, due to poor network environment or equipment failure, the video frame may be damaged or missing, resulting in poor video playback effect and poor user experience.

Method used

By segmenting images of multiple reference video frames of video to be processed, object mask information and object materials are generated, and target object locations are determined based on object motion trajectory information, and target video frames for replacing video frames to be replaced.

Benefits of technology

Improve the image quality and authenticity of video frames, reduce video mosaics and lags, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017911A_ABST
    Figure CN120017911A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method, and relates to the technical field of artificial intelligence, in particular to the technical field of chips, large models and computer vision. According to the specific implementation scheme, image segmentation is carried out on a plurality of reference video frames of a video to be processed to obtain respective frame image segmentation results of the plurality of reference video frames, and the frame image segmentation results comprise at least one piece of object mask information; generating at least one object material for at least one object according to the respective frame image segmentation results of the plurality of reference video frames; determining at least one piece of target object position information according to the at least one piece of object motion track information, the target object position information being used for indicating the position of the object in a to-be-replaced video frame of the to-be-processed video; and generating a target video frame for replacing the video frame to be replaced according to the at least one target object position information and the at least one object material. The invention further provides a video processing device, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of chips, large models and computer vision technology, and can be applied to the scene of abnormal video frame repair. More specifically, the present disclosure provides a video processing method, device, electronic device and storage medium. Background Art

[0002] With the development of computer vision technology, various devices can be used to collect and transmit videos. During the process of video transmission, storage, and video decoding and playback, one or more video frames in the video may be damaged, resulting in poor video playback effect. Summary of the invention

[0003] The present disclosure provides a video processing method, apparatus, device and storage medium.

[0004] According to one aspect of the present disclosure, a video processing method is provided, the method comprising: performing image segmentation on multiple reference video frames of a video to be processed to obtain frame image segmentation results of each of the multiple reference video frames, wherein the frame image segmentation results include at least one object mask information corresponding to at least one object; generating at least one object material for at least one object based on the frame image segmentation results of each of the multiple reference video frames; determining at least one target object position information based on at least one object motion trajectory information corresponding to at least one object, wherein the target object position information is used to indicate the position of the object in a video frame to be replaced of the video to be processed; generating a target video frame for replacing the video frame to be replaced based on at least one target object position information and at least one object material.

[0005] According to another aspect of the present disclosure, a video processing device is provided, which includes: an image segmentation module, used to perform image segmentation on multiple reference video frames of a video to be processed, and obtain frame image segmentation results of each of the multiple reference video frames, wherein the frame image segmentation results include at least one object mask information corresponding to at least one object; a first generation module, used to generate at least one object material for at least one object based on the frame image segmentation results of each of the multiple reference video frames; a determination module, used to determine at least one target object position information based on at least one object motion trajectory information corresponding to at least one object, wherein the target object position information is used to indicate the position of the object in the video frame to be replaced of the video to be processed; a second generation module, used to generate a target video frame for replacing the video frame to be replaced based on at least one target object position information and at least one object material.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided according to the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method provided according to the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method provided according to the present disclosure is implemented.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0011] Figure 1 is a schematic diagram of an exemplary system architecture to which a video processing method and apparatus can be applied according to an embodiment of the present disclosure;

[0012] Figure 2 is a flowchart of a video processing method according to an embodiment of the present disclosure;

[0013] Figure 3A and Figure 3B are schematic diagrams of reference video frames according to an embodiment of the present disclosure;

[0014] Figure 3C is a schematic diagram of a scene image according to an embodiment of the present disclosure;

[0015] Figure 3D is a schematic diagram of a target video frame according to an embodiment of the present disclosure;

[0016] Figure 4 is a block diagram of a video processing device according to an embodiment of the present disclosure; and

[0017] Figure 5 is a block diagram of an electronic device to which a video processing method can be applied according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0019] During video transmission, decoding or playback, one or more video frames may be missing due to various factors such as poor network environment or device failure, which may cause mosaic, video freeze, or even video screen distortion in the video played by the terminal device, ultimately leading to a poor user experience. You can use simple interpolation technology to try to restore missing or damaged video frames in the video. However, the video frames generated based on interpolation technology are quite different from the actual video frames.

[0020] Interpolation technology is to restore the pixels of the original scene in the video and to interpolate and enhance the pixels of the picture. It is performed on the basis of the original low-resolution image and it is difficult to generate high-quality video frames.

[0021] Therefore, in order to improve user experience, the present disclosure provides a video processing method, which will be described below.

[0022] Figure 1 is a schematic diagram of an exemplary system architecture to which a video processing method and apparatus can be applied according to an embodiment of the present disclosure. It should be noted that: Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0023] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0024] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptops, desktop computers, etc.

[0025] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0026] It should be noted that the video processing method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the video processing device provided in the embodiment of the present disclosure can generally be arranged in the server 105. The video processing method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the video processing device provided in the embodiment of the present disclosure can also be arranged in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. The video processing method provided in the embodiment of the present disclosure can also be executed by the terminal devices 101, 102, 103. Accordingly, the video processing device provided in the embodiment of the present disclosure can generally be arranged in the terminal devices 101, 102, 103.

[0027] It can be understood that the above describes the system architecture of the present disclosure, and the following will describe the method of the present disclosure.

[0028] Figure 2 is a flowchart of a video processing method according to an embodiment of the present disclosure.

[0029] like Figure 2 As shown, the method 200 may include operations S210 to S240.

[0030] In operation S210, image segmentation is performed on a plurality of reference video frames of a video to be processed to obtain frame image segmentation results of the plurality of reference video frames.

[0031] In the embodiment of the present disclosure, the video to be processed may include multiple reference video frames and video frames to be replaced. The video frame to be replaced may be a damaged video frame in the video to be processed. For example, the video frame to be replaced may be a video frame with a distorted screen. The video frame to be replaced may be an image. The reference video frame may be an image.

[0032] In the embodiment of the present disclosure, the frame image segmentation result of the reference video frame may be a result of semantic segmentation or a result of instance segmentation.

[0033] In the embodiment of the present disclosure, the frame image segmentation result may include at least one object mask information corresponding to at least one object. Each object mask information may correspond to an object and may indicate the area where the object is located.

[0034] In operation S220, at least one object material for at least one object is generated according to the frame image segmentation results of each of the plurality of reference video frames.

[0035] In the disclosed embodiment, an image can be generated based on multiple object mask information corresponding to the same object to obtain an object material for the object. Using multiple object mask information to generate an image can make the image contour more realistic, thereby making the processed video more realistic.

[0036] In operation S230, at least one target object position information is determined according to at least one object motion track information corresponding to at least one object.

[0037] In the embodiment of the present disclosure, the object motion trajectory information may indicate the positions of the object in a plurality of video frames of the video to be processed.

[0038] In the disclosed embodiment, the target object position information may indicate the position of the object in the video frame to be replaced. For example, the object motion trajectory information includes a plurality of trajectory points. Each trajectory point corresponds to a moment. According to the moment corresponding to the video frame to be replaced, a trajectory point may be determined from the object motion trajectory information. The coordinates of the trajectory point may be used as the target object position information.

[0039] In operation S240, a target video frame for replacing the to-be-replaced video frame is generated according to at least one target object position information and at least one object material.

[0040] In the embodiment of the present disclosure, in the target video frame generated according to the target object position information and the object material, the position of the object material in the video frame to be replaced is consistent with the position indicated by the target object position information.

[0041] Through the disclosed embodiments, the image generation is performed based on the multiple object mask information of the object, which can improve the quality of the object material and the image quality of the target video frame. The target object position information is determined based on the object motion trajectory information, and the position of the object material can be accurately determined, making the generated video frame more realistic. After the target video frame replaces the video frame to be replaced, the replaced video can be made more realistic and smooth, effectively improving the user experience.

[0042] It is understood that operations S210 to S220 and operation S230 may be performed sequentially. However, the embodiments of the present disclosure are not limited thereto, and the two groups of operations may also be performed in other orders, such as first performing operation S230 and then performing operations S210 to S220, or operations S210 to S220 and operation S230 may be performed in parallel.

[0043] It can be understood that the method of the present disclosure is described above, and a plurality of video frames of a video to be processed will be described below.

[0044] In some embodiments, multiple video frames of the video to be processed may be detected to obtain a detection result. The detection result may indicate whether there is a video frame to be replaced in the video to be processed. For example, a trained classification model may be used to detect multiple video frames to obtain multiple detection values. If the detection value is a first preset value, it may be determined that the video frame is a normal video frame. The first preset value may be, for example, 0. If the detection value is a second preset value, it may be determined that the video frame is an abnormal video frame to be replaced. The second preset value may be, for example, 1. The classification model may be a convolutional neural network (CNN).

[0045] In some embodiments, the plurality of reference video frames may include at least one preceding video frame of the video frame to be replaced. For example, 20 video frames preceding the video frame to be replaced may be used as 20 reference video frames.

[0046] In some embodiments, the plurality of reference video frames may include at least one subsequent video frame of the video frame to be replaced. For example, 20 video frames after the video frame to be replaced may be used as 20 reference video frames.

[0047] Through the embodiments of the present disclosure, the multiple reference video frames may be multiple video frames within a reference time period. The reference time period may be a time period before the video frame to be replaced, or a time period after the video frame to be replaced. The duration of the reference time period may be less than or equal to a preset duration threshold. The preset duration threshold may be, for example, 1 second. As a result, the change amplitude of the object within the reference time period is small, which can reduce the interference of the object change on the image generation and improve the image quality of the target video frame.

[0048] It can be understood that the above describes the method of determining the video frame to be replaced and the reference video frame in the present disclosure, and some methods of obtaining the object motion trajectory information will be described below.

[0049] In some embodiments, the above method may further include: processing a plurality of reference video frames using an optical flow algorithm to obtain at least one object motion trajectory information.

[0050] For example, the optical flow algorithm assumes that the brightness of multiple pixels of the same object remains unchanged in a short period of time. The object motion trajectory can be determined based on the brightness value of the pixel. The short period of time can be, for example, 1 to 2 seconds. The multiple reference video frames processed by the optical flow algorithm may include 20 preceding video frames and 20 following video frames of the video frame to be replaced. In these 40 reference video frames, the brightness values ​​of multiple pixels of the same object may remain unchanged.

[0051] For another example, semantic segmentation can be performed on multiple reference video frames respectively to obtain multiple reference segmentation results. The reference segmentation result may include at least one reference mask information. The reference mask information may indicate the area where the object is located in the reference video frame. Based on the reference mask information, one or more pixels of the object may be used as one or more feature points. The feature point may be a pixel at the center of the object. The brightness of the feature point in different reference video frames is the same. Based on the feature point, multiple positions may be determined in multiple reference video frames. Based on the multiple positions of the feature point in multiple reference video frames, the motion trajectory of the feature point may be determined as the object motion trajectory information.

[0052] Through the embodiments of the present disclosure, the motion trajectory information of the object can be efficiently determined using the optical flow algorithm, thereby improving the video processing efficiency.

[0053] It can be understood that the above description of the present disclosure is based on the example that the multiple reference video frames include at least one previous video frame and at least one subsequent video frame. However, the present disclosure is not limited thereto, and the multiple reference video frames may include only multiple previous video frames or only multiple subsequent video frames.

[0054] It can be understood that some methods for determining the object motion trajectory information are described above, and some methods for determining the frame image segmentation result will be described below.

[0055] In some embodiments, the frame image segmentation result may include at least one local image segmentation result. The local image segmentation result may be obtained by performing image segmentation on a local image of a reference video frame, which will be described below.

[0056] In some embodiments, in some implementations of the above operation S210, performing image segmentation on multiple reference video frames of the video to be processed to obtain image segmentation results of each of the multiple reference video frames may include: determining multiple reference position information respectively used for the multiple reference video frames according to at least one object motion trajectory information. The reference position information includes at least one reference object position information, and the reference object position information may indicate the position of the object in the reference video frame. For example, the reference object position information may be a coordinate in the reference video frame. The coordinate may be the coordinate of the center point or vertex of the object.

[0057] In some embodiments, in some implementations of the above operation S210, performing image segmentation on multiple reference video frames of the video to be processed to obtain image segmentation results of each of the multiple reference video frames may include: determining at least one partial image of each of the multiple reference video frames according to multiple reference position information for the multiple reference video frames. For example, the position indicated by the reference object position information may be used as the center point to determine a bounding box of a preset size. Based on the area surrounded by the bounding box, a partial image may be determined.

[0058] In some embodiments, in some implementations of the above operation S210, performing image segmentation on multiple reference video frames of the video to be processed to obtain image segmentation results of each of the multiple reference video frames may include: performing image segmentation on at least one local image of each of the multiple reference video frames to obtain at least one local image segmentation result of each of the multiple reference video frames. For example, semantic segmentation is performed on at least one local image of the reference video frame to obtain at least one local image segmentation result of the reference video frame. The local image segmentation result may include one or more object mask information. At least one local image segmentation result may be used as a frame image segmentation result.

[0059] Through the embodiments of the present disclosure, multiple local images are determined according to the motion trajectory, and the local images are segmented, which can improve the accuracy of image segmentation while reducing the computing resources required for segmentation, and can more accurately determine the object mask information. It can be understood that the frame image segmentation result can be different from the above-mentioned reference segmentation result. The frame image segmentation result includes at least one local image segmentation result obtained by segmenting the local image, which has higher accuracy. The reference segmentation result is the segmentation result of the entire image, which has global information, but is slightly less accurate.

[0060] It can be understood that the above describes the frame image segmentation results of the present disclosure, and some methods of generating object materials of the present disclosure will be described below.

[0061] In some embodiments, in some implementations of the above operation S220, generating at least one object material for at least one object according to the frame image segmentation results of each of the multiple reference video frames includes: generating at least one object material for at least one object using a large model according to the image segmentation results of each of the multiple reference video frames. For example, the large model can be a text-based large model or a picture-based large model. The picture-based large model can be a large model based on a stable diffusion model. As described above, the multiple reference video frames can include the same object. Thus, the frame image segmentation results of each of the multiple reference video frames can include multiple object mask information of the object. Based on the multiple object mask information, multiple initial materials corresponding to the object can be determined from the multiple reference video frames. The multiple initial materials are provided to the large model, and an object material for the object can be generated.

[0062] In some embodiments, the number of reference video frames used to generate the object material may be less than the number of reference video frames used to generate the object motion trajectory. For example, the object material may be generated using a large model based on the frame image segmentation results of the two preceding video frames of the video frame to be replaced and the frame image segmentation results of the two following video frames of the video frame to be replaced.

[0063] Through the embodiments of the present disclosure, more refined object materials can be generated using a large model, which can effectively improve the quality of the generated target video frames, thereby improving the user experience.

[0064] It can be understood that before, during or after obtaining the frame image segmentation result and the object material, the above operation S230 can be performed to determine at least one target object position information according to at least one object motion trajectory information. Next, operation S240 can be performed to generate a target video frame for replacing the video frame to be replaced according to at least one target object position information and at least one object material, which will be described below.

[0065] In some embodiments, in some implementations of the above operation S240, generating a target video frame for replacing a video frame to be replaced according to at least one target object position information and at least one object material may include: determining a scene image according to a plurality of reference video frames. Generating a target video frame according to the scene image, at least one object position information and at least one object material.

[0066] In some embodiments, determining the scene image according to the multiple reference video frames includes: determining a plurality of image data to be fused according to the multiple reference video frames. The image data to be fused may be determined according to video frame difference information between two adjacent reference video frames.

[0067] For example, by subtracting two adjacent reference video frames, video frame difference information can be obtained. The subtraction can be pixel-by-pixel subtraction. Next, the video frame difference information and the reference video are fused to obtain the image data to be fused. The fusion method can be pixel-by-pixel addition. In an example, the image data to be fused can be determined by the following formula:

[0068] J_i=P_i-P_i+1+ P_i (Formula 1)

[0069] J_i may be the i-th image data to be fused. P_i-P_i+1 may be the video frame difference information. P_i may be the i-th reference video frame. P_i+1 may be the i+1-th reference video frame. i may be an integer greater than or equal to 1 and less than N. N may be an integer greater than 1. N is the number of reference video frames.

[0070] In some embodiments, determining the scene image according to the multiple reference video frames may further include: averaging the multiple image data to be fused pixel by pixel to obtain the scene image. For example, the scene image may be obtained by the following formula:

[0071] (Formula 2)

[0072] B may be a scene image.

[0073] Next, the object material may be added to the position indicated by the target position information in the scene image to obtain a target video frame.

[0074] It can be understood that the above text describes the method of determining the image data to be fused in combination with formula 1. However, the present disclosure is not limited to this. In other embodiments, based on the above video frame difference information, the video frame consistency information can be determined in the i-th reference video frame. For example, after subtracting the i-th reference video frame and the i+1-th reference video frame, a consistent area with a pixel value of 0 can be obtained. The outline and position of the area can be represented by the video frame consistency information. According to N reference video frames, N-1 video frame consistency information can be determined. According to the N-1 video frame consistency information, N-1 consistent areas can be extracted from the 1st reference video frame to the N-1th reference video frame. The brightness value of the pixel in the image area outside the consistent area is set to a preset value, and an initial image can be obtained. According to the pixel-by-pixel averaging of the N-1 initial images, a scene image can be obtained. The scene image can also be called a background image. The preset value can be 0 or 255. The scene image can be a red, green, and blue (RGB) image.

[0075] Next, the method of the present disclosure will be further described in conjunction with multiple video frames.

[0076] Figure 3A and Figure 3B Each of them is a schematic diagram of a reference video frame according to an embodiment of the present disclosure.

[0077] like Figure 3A As shown, the reference video frame P31 may be a previous video frame of the video frame to be replaced. The object o30 in the reference video frame P31 is located at the bottom of the reference video frame P31.

[0078] like Figure 3B As shown in FIG. 1 , the reference video frame P33 may be a subsequent video frame of the video frame to be replaced. The object o30 in the reference video frame P33 is located in the middle of the reference video frame P33. FIG. 3A to FIG. 3B As shown, the object o30 is moving toward the upper right.

[0079] Next, the method 200 may be executed. In the process of executing the method 200, the following may be obtained: Figure 3C Scene image B30 is shown.

[0080] Figure 3C is a schematic diagram of a scene image according to an embodiment of the present disclosure.

[0081] like Figure 3C As shown, the scene image B30 includes consistent image regions in the reference video frame P31 and the reference video frame P33.

[0082] Figure 3D is a schematic diagram of a target video frame according to an embodiment of the present disclosure.

[0083] like Figure 3D As shown, the target video frame P32 can be used to replace the video frame to be replaced between the reference video frame P31 and the reference video frame P33.

[0084] It can be understood that the method of the present disclosure is described above, and the device of the present disclosure will be described below.

[0085] Figure 4 is a block diagram of a video processing device according to an embodiment of the present disclosure.

[0086] like Figure 4 As shown, the apparatus 400 may include an image segmentation module 410 , a first generation module 420 , a determination module 430 and a second generation module 440 .

[0087] The image segmentation module 410 is used to perform image segmentation on multiple reference video frames of the video to be processed to obtain frame image segmentation results of each of the multiple reference video frames. The frame image segmentation results include at least one object mask information corresponding to at least one object.

[0088] The first generating module 420 is configured to generate at least one object material for at least one object according to the frame image segmentation results of each of the plurality of reference video frames.

[0089] The determination module 430 is used to determine at least one target object position information according to at least one object motion track information corresponding to at least one object. The target object position information is used to indicate the position of the object in the video frame to be replaced of the video to be processed.

[0090] The second generating module 440 is configured to generate a target video frame for replacing a video frame to be replaced according to at least one target object position information and at least one object material.

[0091] In some embodiments, the frame image segmentation result includes at least one local image segmentation result. The image segmentation module includes: a first determination submodule, used to determine multiple reference position information respectively used for multiple reference video frames based on at least one object motion trajectory information. The reference position information includes at least one reference object position information, and the reference object position information is used to indicate the position of the object in the reference video frame. A second determination submodule is used to determine at least one local image of each of the multiple reference video frames based on the multiple reference position information for the multiple reference video frames. The image segmentation submodule is used to perform image segmentation on at least one local image of each of the multiple reference video frames to obtain at least one local image segmentation result of each of the multiple reference video frames.

[0092] In some embodiments, the device further includes: a processing module, configured to process a plurality of reference video frames using an optical flow algorithm to obtain at least one object motion trajectory information.

[0093] In some embodiments, the first generation module includes: a first generation submodule, which is used to generate at least one object material for at least one object using a large model according to the image segmentation results of each of the multiple reference video frames.

[0094] In some embodiments, the second generation module includes: a third determination submodule for determining a scene image according to a plurality of reference video frames. A second generation submodule for generating a target video frame according to the scene image, at least one object position information and at least one object material.

[0095] In some embodiments, the third determination submodule includes: a determination unit, configured to determine a plurality of image data to be fused according to a plurality of reference video frames. The image data to be fused is determined according to video frame difference information between two adjacent reference video frames. A pixel-by-pixel averaging unit, configured to average the plurality of image data to be fused pixel by pixel to obtain a scene image.

[0096] In some embodiments, the image data to be fused is obtained by fusing the reference video frame and the video frame difference information.

[0097] In some embodiments, the plurality of reference video frames include at least one of the following: at least one preceding video frame of the video frame to be replaced; and at least one succeeding video frame of the video frame to be replaced.

[0098] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0099] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0100] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0101] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 to a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0102] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0103] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as video processing methods. For example, in some embodiments, the video processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the video processing method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the video processing method in any other appropriate manner (eg, by means of firmware).

[0104] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0105] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0106] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROM) or flash memories, optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0107] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0108] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.

[0109] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.

[0110] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0111] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A video processing method, comprising: Performing image segmentation on a plurality of reference video frames of the video to be processed to obtain frame image segmentation results of each of the plurality of reference video frames, wherein the frame image segmentation results include at least one object mask information corresponding to at least one object; generating at least one object material for at least one object according to the frame image segmentation results of each of the plurality of reference video frames; Determine at least one target object position information according to at least one object motion trajectory information corresponding to at least one of the objects, wherein the target object position information is used to indicate the position of the object in the video frame to be replaced in the video to be processed; A target video frame for replacing the video frame to be replaced is generated according to at least one target object position information and at least one object material.

2. The method according to claim 1, wherein: The frame image segmentation result includes at least one local image segmentation result. The performing image segmentation on the multiple reference video frames of the video to be processed to obtain the image segmentation results of each of the multiple reference video frames comprises: Determine, according to at least one of the object motion trajectory information, a plurality of reference position information respectively used for the plurality of reference video frames, wherein the reference position information includes at least one reference object position information, and the reference object position information is used to indicate the position of the object in the reference video frame; Determining at least one partial image of each of the plurality of reference video frames according to a plurality of reference position information for the plurality of reference video frames; Image segmentation is performed on at least one of the local images of each of the multiple reference video frames to obtain at least one local image segmentation result of each of the multiple reference video frames.

3. The method according to claim 1 or 2, further comprising: The plurality of reference video frames are processed using an optical flow algorithm to obtain at least one piece of object motion trajectory information.

4. The method according to claim 1, wherein: Generating at least one object material for at least one object according to the frame image segmentation results of each of the plurality of reference video frames comprises: At least one object material for at least one of the objects is generated using a large model according to the image segmentation results of each of the plurality of reference video frames.

5. The method according to claim 1, wherein: The generating, according to at least one of the object position information and at least one of the object materials, a target video frame for replacing the video frame to be replaced comprises: Determining a scene image according to the plurality of reference video frames; The target video frame is generated according to the scene image, at least one piece of the object position information and at least one of the object materials.

6. The method according to claim 5, wherein: Determining the scene image according to the plurality of reference video frames comprises: Determine a plurality of image data to be fused according to the plurality of reference video frames, wherein the image data to be fused is determined according to video frame difference information between two adjacent reference video frames; The plurality of image data to be fused are averaged pixel by pixel to obtain the scene image.

7. The method according to claim 6, wherein: The image data to be fused is obtained by fusing the reference video frame and the video frame difference information.

8. The method according to claim 1, wherein: The plurality of reference video frames include at least one of the following: at least one preceding video frame of the video frame to be replaced; At least one subsequent video frame of the video frame to be replaced.

9. A video processing device, comprising: An image segmentation module, configured to perform image segmentation on a plurality of reference video frames of a video to be processed, and obtain frame image segmentation results of each of the plurality of reference video frames, wherein the frame image segmentation results include at least one object mask information corresponding to at least one object; A first generating module, configured to generate at least one object material for at least one object according to the frame image segmentation results of each of the plurality of reference video frames; A determination module, configured to determine at least one target object position information according to at least one object motion trajectory information corresponding to at least one of the objects, wherein the target object position information is used to indicate the position of the object in the to-be-replaced video frame of the to-be-processed video; The second generating module is used to generate a target video frame for replacing the video frame to be replaced according to at least one target object position information and at least one object material.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.