Video processing method, system, electronic device and storage medium

By using an adjustment model trained on image content samples, and adjusting the video using size adjustment parameters, the problem of incomplete content during video size adjustment is solved, thus achieving both the integrity of the video content and the improvement of the overall effect.

CN118175389BActive Publication Date: 2026-01-16ALIBABA (BEIJING) SOFTWARE SERVICES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410168087.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-05
Publication Date
2026-01-16
Estimated Expiration
2044-02-05

AI Technical Summary

Technical Problem

Existing technologies often result in incomplete video content and poor video quality when adjusting video size.

Method used

An adjustment model trained on image content samples is used to adjust the original video using size adjustment parameters to generate the target video, ensuring the integrity of the video content.

Benefits of technology

It achieves the goal of maintaining the integrity of video content while adjusting video size, thus improving the effect of video adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118175389B_ABST
    Figure CN118175389B_ABST
Patent Text Reader

Abstract

The application discloses a video processing method, system, electronic equipment and storage medium. The method can comprise: monitoring an original video to be adjusted; determining a size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; and adjusting the original video by using at least the size adjustment parameter in the adjustment model to obtain a target video, wherein the target video comprises original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video in a direction corresponding to the size adjustment parameter and by a distance. The application solves the technical problem of poor video adjustment effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of video processing, in particular, to a video processing method and system, an electronic device and a storage medium. BACKGROUND

[0002] At present, in the current digital media era, diversified video content is widely spread on various platforms. Due to the rich display form and high degree of spread of videos, videos have gradually replaced pictures to become the medium for users to record life and for businesses to market products. However, different media that display videos have different display ways for videos, and different display ways correspond to different video sizes. Therefore, how to adapt the same video to different video sizes for display is very important.

[0003] In the related art, if the size of a video needs to be adjusted, a commonly used way is to crop the video. However, this way will seriously damage the content in the video. For example, it will cause the characters or objects in the video to be incomplete, resulting in serious loss of information in the video. Therefore, there is still the technical problem of poor video adjustment effect.

[0004] At present, there is no effective solution to the above problems. SUMMARY

[0005] The embodiments of the present application provide a video processing method and system, an electronic device and a storage medium to at least solve the technical problem of poor video adjustment effect.

[0006] According to an aspect of the embodiments of the present application, a video processing method is provided. The method can include: monitoring an original video to be adjusted; determining a size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; and adjusting the original video by at least using the size adjustment parameter in the adjustment model to obtain a target video, wherein the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent a region offset from the video area boundary of the original video in a direction and distance corresponding to the size adjustment parameter.

[0007] According to another aspect of embodiments of the present application, another method for processing a video is provided. The method can include obtaining an original video to be adjusted from an e-commerce platform; determining a size adjustment parameter of the original video based on a media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust a video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; adjusting the original video by at least using the size adjustment parameter in the adjustment model to obtain a target video, wherein target video content of the target video includes original video content from the original video and matching video content that matches the original video content, the matching video content being displayed in a target video region, the target video region being used to represent a region that is offset from a video region boundary of the original video in a direction and by a distance corresponding to the size adjustment parameter; and pushing the target video to the media platform for playing.

[0008] According to another aspect of embodiments of the present application, another method for processing a video is provided. The method can include displaying an original video to be adjusted and a size adjustment parameter of the original video on an operation interface in response to an input instruction acting on the operation interface, wherein the size adjustment parameter is used to adjust a video size of the original video; and displaying a target video on the operation interface in response to a video adjustment instruction acting on the operation interface, wherein the target video is obtained by adjusting the original video by at least using the size adjustment parameter in an adjustment model, the adjustment model being trained based on image content samples, and target video content of the target video includes original video content from the original video and matching video content that matches the original video content, the matching video content being displayed in a target video region, the target video region being used to represent a region that is offset from a video region boundary of the original video in a direction and by a distance corresponding to the size adjustment parameter.

[0009] According to another aspect of embodiments of the present application, another method for processing a video is provided. The method can include displaying an original video to be adjusted and a size adjustment parameter of the original video on an operation interface in response to an input instruction acting on the operation interface, wherein the size adjustment parameter is used to adjust a video size of the original video; and displaying a target video on the operation interface in response to a video adjustment instruction acting on the operation interface, wherein the target video is obtained by adjusting the original video by at least using the size adjustment parameter in an adjustment model, the adjustment model being trained based on image content samples, and target video content of the target video includes original video content from the original video and matching video content that matches the original video content, the matching video content being displayed in a target video region, the target video region being used to represent a region that is offset from a video region boundary of the original video in a direction and by a distance corresponding to the size adjustment parameter.

[0010] According to another aspect of embodiments of the present application, a system for processing a video is provided. The system can include: an e-commerce platform configured to produce an original video to be adjusted; a video adjustment end configured to determine a size adjustment parameter of the original video based on a media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust a video size of the original video; input the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; and adjust the original video using at least the size adjustment parameter in the adjustment model to obtain a target video, wherein a target video content of the target video includes original video content from the original video and matching video content that matches the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region that is offset from a video region boundary of the original video in a direction and by a distance corresponding to the size adjustment parameter; and the media platform is configured to play the target video.

[0011] According to another aspect of embodiments of the present application, an electronic device is also provided. The electronic device can include a memory and a processor, the memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the video processing method of embodiments of the present application is implemented.

[0012] According to another aspect of embodiments of the present application, a processor is also provided. The processor is configured to run a program, and when the program is running, the video processing method of embodiments of the present application is executed.

[0013] According to another aspect of embodiments of the present application, a computer readable storage medium is also provided. The computer readable storage medium includes a stored program, and when the program is running, the device in which the storage medium is located is controlled to execute the video processing method of embodiments of the present application.

[0014] According to another aspect of embodiments of the present application, a computer program product is also provided. The computer program product includes a computer program, and when the computer program is executed by a processor, the video processing method of embodiments of the present application is implemented.

[0015] In the embodiment of the present application, if the size of a certain original video needs to be adjusted, it can be monitored whether the original video to be adjusted is received, and if the original video to be transmitted is monitored, the size adjustment parameter that needs to adjust the video size of the original video can be determined. The adjustment model can be trained in advance according to the image content sample, that is, the ability of the adjustment model to adjust the size of the original video can be trained according to the shape and appearance of the image content. The size adjustment parameter and the original video are input into the adjustment model, and the adjustment model is used to adjust the video size of the original video according to the size adjustment parameter to obtain a target video. In the target video, not only the original video content from the original video, but also the matching video content matched by the adjustment model according to the original video content can be included, and the matching video content can be displayed in the target video area, that is, the matching video content can be displayed in the area obtained by offsetting from the video area boundary of the original video along the direction and distance corresponding to the size adjustment parameter. Since the embodiment of the present application can accurately adjust the video size of the original video to the video size of the target video required by using the adjustment model according to the size adjustment parameter and the original video, the situation of missing video content caused by simply cropping the video is avoided, thereby achieving the purpose of ensuring the integrity of the video content, and further realizing the technical effect of improving the video adjustment effect, solving the technical problem of poor video adjustment effect.

[0016] It is easy to note that the above general description and the following detailed description are only for illustrating and explaining the present application, and do not constitute a limitation on the present application. BRIEF DESCRIPTION OF DRAWINGS

[0017] The drawings described herein are used to provide further understanding of the present application, constitute a part of the present application, and the illustrative embodiments of the present application and the description thereof are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 is a schematic diagram of an application scene of a video processing method according to an embodiment of the present application;

[0019] Figure 2 is a flowchart of a video processing method according to an embodiment of the present application;

[0020] Figure 3 is a flowchart of another video processing method according to an embodiment of the present application;

[0021] Figure 4 is a flowchart of another video processing method according to an embodiment of the present application;

[0022] Figure 5is a flowchart of another video processing method according to an embodiment of the present application;

[0023] Figure 6 is a flowchart of a video processing system according to an embodiment of the present application;

[0024] Figure 7 is a flowchart of a video size transformation method based on diffusion model picture completion according to an embodiment of the present application;

[0025] Figure 8 is a schematic diagram of a video size transformation system based on diffusion model picture completion according to an embodiment of the present application;

[0026] Figure 9 is a schematic diagram of a multi-stage cascading processing of an original video according to an embodiment of the present application;

[0027] FIG. 10(a) is a schematic diagram of an example of a target video obtained by adjusting an original video according to an embodiment of the present application;

[0028] FIG. 10(b) is a schematic diagram of another example of a target video obtained by adjusting an original video according to an embodiment of the present application;

[0029] Figure 11 is a schematic diagram of a video processing apparatus according to an embodiment of the present application;

[0030] Figure 12 is a schematic diagram of another video processing apparatus according to an embodiment of the present application;

[0031] Figure 13 is a schematic diagram of another video processing apparatus according to an embodiment of the present application;

[0032] Figure 14 is a schematic diagram of another video processing apparatus according to an embodiment of the present application;

[0033] Figure 15 is a structural block diagram of a computer terminal according to an embodiment of the present application;

[0034] Figure 16 is a block diagram of an electronic device for implementing a video processing method according to an embodiment of the present application;

[0035] Figure 17 is a hardware structural block diagram of a computer terminal (or mobile device) for implementing a video processing method according to an embodiment of the present application;

[0036] Figure 18 is a structural block diagram of a computing environment for implementing a video processing method according to an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0038] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0039] First, some of the nouns or terms that appear in the description of the embodiments of the present application are applicable to the following explanations:

[0040] Video size transformation refers to adjusting the resolution, frame rate and other parameters of a video to adapt to different display devices or transmission requirements;

[0041] Diffusion model refers to a mathematical model used to simulate and predict the diffusion trend of data in space or events;

[0042] Multi-stage cascade refers to combining multiple processing steps or modules together to form a multi-level processing flow;

[0043] Contrastive language-image pretraining (CLIP) model refers to a pre-training model released by OpenAI in 2022, which is used for natural language understanding and computer vision tasks;

[0044] Visual encoder is an algorithm or model used to convert image or video data into a digital representation;

[0045] Canny edge detection refers to an algorithm for detecting image edges, which can be used for image segmentation, object detection and other tasks;

[0046] Picture consistency refers to the consistency of parameters such as the quality, color, and brightness of a video or an image, and no obvious difference or incoordination occurs.

[0047] Embodiment 1

[0048] According to the embodiments of the present application, a video processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0049] The video processing method provided by the embodiments of the present application can be applied to the application scenarios as shown in Figure 1 but not limited thereto. In the application scenarios as shown in Figure 1 The server 10 can be a cloud. The server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client devices 20 can include, but are not limited to, smartphones, tablet computers, notebook computers, palm computers, personal computers, smart home devices, vehicle-mounted devices, etc. The client devices collectively constitute a client opposite the server. An interactive interface for obtaining a file upload request can be deployed on a graphical user interface on the client device. The interactive interface can be a generative dialogue interface. The client device 20 can interact with the user through the graphical user interface to implement the video processing method provided by the embodiments of the present application.

[0050] In the embodiments of the present application, the system composed of the client device and the server can execute the following steps: if the client has a demand for adjusting a certain original video, the original video required to be adjusted can be input on the interactive interface on the client device and can be sent to the server through the network, for example, the original video to be adjusted can be made in an e-commerce platform on the client device and can be sent to a video adjustment end in the server. After receiving the original video to be adjusted, the server can execute the following steps: step S102, monitoring the original video to be adjusted; step S104, determining the size adjustment parameter of the original video; step S106, inputting the size adjustment parameter and the original video into an adjustment model; step S108, adjusting the original video by at least using the size adjustment parameter in the adjustment model to obtain a target video. In the above process, the target video adjusted in the server can be sent to the client device through the network. The corresponding target video can be displayed on the interactive interface of the client device, for example, the target video can be played on a media platform on the client device.

[0051] The embodiment of the present application can accurately adjust the video size of the original video to the video size of the target video by using the adjustment model according to the size adjustment parameter and the original video, avoids the loss of video content caused by simply cutting the video, achieves the purpose of ensuring the integrity of the video content, and further realizes the technical effect of improving the video adjustment effect, and solves the technical problem of poor video adjustment effect.

[0052] The embodiment of the present application provides a video processing method as shown in the following Figure 2 application scenarios, the embodiment of the present application provides a video processing method as shown in the following Figure 2 application scenarios, the embodiment of the present application provides a video processing method as shown in the following Figure 2 application scenarios, the embodiment of the present application provides a video processing method as shown in the following

[0053] Step S202, monitoring the original video to be adjusted.

[0054] In the technical solution provided in the above step S202 of the present application, the original video can be a video whose size needs to be adjusted, can be an input video input by a client into a server for the server to adjust the size of the video, and can also be referred to as an original video or a video segment. If the scenario of adjusting the size of the video is an e-commerce scenario, the original video can be an e-commerce video. The video size usually refers to the width and height of the video, measured in pixels (px), and the video size determines the display size and ratio of the video on the screen. For example, common video sizes include a high-definition video size of 1920x1080 pixels and a standard-definition video size of 720x480 pixels, and the video size can be adjusted according to different display devices and resolutions.

[0055] In this embodiment, the original video to be adjusted can be monitored.

[0056] Optionally, if a user needs to adjust the video size of a certain original video, the user can perform corresponding operations in the corresponding client to trigger the operation of transmitting the original video to the server, and can transmit the original video to the server through the network, wherein the user can also be referred to as a user.

[0057] For example, if the original video is to be displayed in a certain media application program in the client, the requirements of the video size of the video to be displayed in the application program can be determined. If the requirements of the video size are inconsistent with the video size of the original video, the video size of the original video needs to be adjusted so that the video size meets the requirements of the application program. The user can transmit the original video to an interactive interface that can interact with the server, and the interactive interface can transmit the received original video to the server to adjust the size of the original video.

[0058] Step S204, determine the size adjustment parameter of the original video.

[0059] In the technical solution provided in the above step S204 of the present application, the size adjustment parameter can be used to adjust the video size of the original video, which can also be referred to as a video size adjustment value or a final target video size value to which the original video needs to be adjusted. If the video size of the original video needs to be extended, the size adjustment parameter at this time can also be referred to as an extension parameter.

[0060] In this embodiment, after monitoring that there is an original video to be adjusted, the size adjustment parameter of the original video can be determined.

[0061] Optionally, after determining that the video size of the original video needs to be adjusted, the size adjustment parameter to which the original video needs to be adjusted to the required video size can be determined.

[0062] Optionally, after monitoring the original video to be adjusted, the original video size (video original size) of the original video can be determined, and the target video size (specified output size) to which the original video needs to be adjusted can also be determined. Based on the original video size and the target video size, the size adjustment parameter is determined, wherein the video original size can also be referred to as the original size. The target video size can also be referred to as the target size.

[0063] Optionally, if the user inputs multiple specified output sizes to which the original video needs to be adjusted, that is, inputs multiple target video sizes, the target video sizes under the condition that one of the length or width of the original video size of the original video is maintained unchanged can be calculated, and the union of the regions represented by the multiple target video sizes is obtained to obtain the final extended size adjustment parameter.

[0064] Step S206, input the size adjustment parameter and the original video into the adjustment model.

[0065] In the technical solution provided in the above step S206 of the present application, the adjustment model can be trained based on image content samples, and can be a mathematical model for simulating and predicting the diffusion trend of objects in space or time in the video, such as a diffusion model. The image content sample can be the shape and appearance of the image of a common object in daily life, which is only used as an example and is not limited to the image content sample for training the adjustment model.

[0066] In this embodiment, after determining the size adjustment parameter of the original video, the size adjustment parameter and the original video can be input into the adjustment model, and the original size of the original video can be adjusted according to the size adjustment parameter through the adjustment model.

[0067] Optionally, the adjustment model can be trained in advance according to a large number of videos, that is, the adjustment model can learn the image content such as the shape and appearance of common objects in each frame of the video by learning in a large number of videos.

[0068] Optionally, after the adjustment model is trained by a large number of videos, the adjustment model can realize the completion of the video content in the video and the size transformation by using the knowledge learned in the training process when the video in a given small area.

[0069] Because the original video content in the video may be incomplete or even lost by simply cropping the original size of the original video, there is still a technical problem that the size of the video cannot be effectively adjusted. However, in the embodiment of the present application, the adjustment model can be input with the size adjustment parameter and the original video, and the adjustment model can be used to supplement the video content in the original video by using the learned shape and appearance of various common objects, and the original size of the original video can be adjusted according to the size adjustment parameter. The purpose of adjusting the size of the original video can be achieved, and the purpose of maintaining the integrity of the original video content in the video can be achieved, thereby realizing the technical effect that the size of the video can be effectively adjusted.

[0070] Step S208, at least using the size adjustment parameter to adjust the original video in the adjustment model to obtain a target video.

[0071] In the technical solution provided by the above step S208 of the present application, the target video can include original video content from the original video, and matching video content matched with the original video content. The target video can be the final extended video obtained by adjusting the size of the original video according to the needs of the client. The original video content can be the shape and appearance of the objects and persons contained in the original video. The matching video content can be displayed in the target video area, which can be represented by extended pixels obtained by extending some objects or persons in the original video. The target video area can be used to represent the area offset from the video area boundary of the original video along the direction and distance corresponding to the size adjustment parameter, which can also be called the extended area. The target video can also be called the output video or the extended video.

[0072] In this embodiment, after the size adjustment parameter and the original video are input into the adjustment model, the size of the original video can be adjusted in the adjustment model by using the size adjustment parameter to obtain a target video that meets the needs of the client.

[0073] Optionally, the adjustment model has a good visual content generation capability, that is, the adjustment model has the capability of matching the original video content in the video with the corresponding matching video content according to the learned image content. In the adjustment model, the original video can be adjusted so that the picture quality, picture texture, color, and target consistency of the original video can be improved.

[0074] In the embodiments of the present application, the diffusion model can be used to understand the content of the video, and the knowledge learned by the diffusion model can be used to expand the content of the original video in an imaginative way, which can greatly protect the integrity of the original content of the original video, thereby achieving the technical effect of effectively adjusting the size of the video.

[0075] Through the above steps S202 to S208 of the present application, if the size of a certain original video needs to be adjusted, it can be monitored whether the original video to be adjusted is received. If the original video to be transmitted is monitored, it can be determined that the size adjustment parameter of the video size of the original video needs to be adjusted. The size adjustment parameter and the original video are input into the adjustment model, and the adjustment model is used to adjust the video size of the original video according to the size adjustment parameter to obtain a target video. In the target video, not only the original video content from the original video, but also the matching video content matched by the adjustment model according to the original video content can be included, and the matching video content can be displayed in the target video area, that is, the matching video content can be displayed in the area obtained by offsetting from the video area boundary of the original video in the direction and distance corresponding to the size adjustment parameter. Since the embodiments of the present application can accurately adjust the video size of the original video to the video size of the target video required by using the adjustment model according to the size adjustment parameter and the original video, the loss of video content caused by simply cropping the video is avoided, thereby achieving the purpose of ensuring the integrity of the video content, and further achieving the technical effect of improving the effect of video adjustment, solving the technical problem of poor video adjustment effect.

[0076] The above method of the embodiment will be further introduced below.

[0077] As an optional implementation, in step S206, inputting the size adjustment parameter and the original video into the adjustment model comprises: generating a video generation parameter based at least on the size adjustment parameter and the original video, wherein the video generation parameter is used to make the adjustment model fill the matching video content in the target video area; and inputting the video generation parameter into the adjustment model.

[0078] In this embodiment, in the process of inputting the size adjustment parameter and the original video into the adjustment model, the video generation parameter can be generated based at least on the size adjustment parameter and the original video, and the video generation parameter can be input into the adjustment model, wherein the video generation parameter can be used to enable the adjustment model to fill the matching video content matching the original video in the target video region, which can include noise information, mask information, given video information, edge detection map, etc. of the original video. It should be noted that the above video generation parameter is only for illustration and is not limited herein.

[0079] Optionally, after inputting the size adjustment parameter and the original video into the adjustment model, the video generation parameter can be obtained by performing information encoding, random sampling noise, etc. on each frame of the original video and the region contained in the video size of the original video, i.e., on the given video frame and the given video region and the given video region. Thus, the above information is input into the trained adjustment model.

[0080] As an optional implementation, the video generation parameter is generated based at least on the size adjustment parameter and the original video, including: generating mask information based on the size adjustment parameter and the original video size of the original video, wherein the video generation parameter includes the mask information, and the mask information is used to at least enable the adjustment model to determine the target video region.

[0081] In this embodiment, in the process of generating the video generation parameter based at least on the size adjustment parameter and the original video, the corresponding mask information can be generated based on the original video size corresponding to the size adjustment parameter and the original video, wherein the video generation parameter can include the mask information, which can be a mask mask.

[0082] Optionally, the original video size of the original video and the video frame of the original video are information encoded. The trained adjustment model is used to adjust the video size of the original video, and four parts of information can be input into the adjustment model. One part can be noise input, i.e., the noise in the original video size of the original video and the video frame of the original video after information encoding can be extracted by random sampling. Another part can be mask information. Another part can be given video information, i.e., the original video content in each video frame in the original video. Another part can be the edge information of the original video.

[0083] For example, Gaussian noise in the original video size of the original video and the video frame of the original video after information encoding can be obtained by Gaussian sampling. It should be noted that the above method and process for determining the noise in the original video are only for illustration and are not limited herein.

[0084] Optionally, in the process of generating the mask information, it can be determined from the original video which areas need to be filled, which areas of the original video are given information, that is, which areas of the original video are target video areas and which areas are original video areas. Based on the above, the mask information of the original video is generated.

[0085] Optionally, the given video information input into the adjustment model can be information contained in each video frame in the input original video, which contains target features that can assist in video extension. The target features can be shape, appearance, and size of objects or tasks in the original video, and the like. It should be noted that the above target features are only illustrative and are not limited in this regard. As long as the features can assist in adjusting the video size of the original video, they are within the protection scope of the embodiments of the present application.

[0086] As an optional implementation, generating the mask information based on the size adjustment parameter and the original video size of the original video includes: determining the area of the target video area using the size adjustment parameter and the original video size; and generating the mask information based on the area of the target video area.

[0087] In this embodiment, in the process of generating the mask information based on the size adjustment parameter and the original video size of the original video, the area of the target video area can be determined using the size adjustment parameter and the original video size, and the corresponding mask information can be generated based on the area of the target video area.

[0088] Optionally, before generating the mask information, the extension parameter calculation can be performed, that is, the size adjustment parameter can be determined first.

[0089] Optionally, in the process of performing the extension parameter calculation, the video original size of the original video input by the user and the specified multiple output sizes can be first calculated to determine the minimum extension area, that is, to determine the target video area. The area of the target video area can be determined, that is, the area of the extension area for extending the top, bottom, left, and right of the original video, and the corresponding mask mask can be constructed and input into the adjustment model.

[0090] For example, if the video original size of the original video is 512x512 pixels, and the two target video sizes are 720x1280 pixels and 1280x720 pixels, the video original size can be first extended to 512x910 pixels and 910x512 pixels, and then the union can be calculated to obtain the final size of the original video that needs to be extended, that is, the size adjustment parameter of the original video can be 910x910.

[0091] It should be noted that the above process and method of determining the size adjustment parameter of the original video are only for illustration, and are not specifically limited here.

[0092] As an optional implementation, the method further includes: in a case where the size adjustment parameter is to enlarge the video size of the original video, extracting first video information from the original video, wherein the video generation parameter includes the first video information, and the first video information is used to cause the adjustment model to generate matching video content for extending the original video in the target video region.

[0093] In this embodiment, in a case where the size adjustment parameter is to enlarge the video size of the original video, the first video information can be extracted from the original video, and the first video information can be included in the video generation parameter. The first video information can be used to cause the adjustment model to generate matching video content for extending the original video in the target video region, and can be given video information, that is, various information contained in each video frame of the original video.

[0094] Optionally, the first video information can also be input to the adjustment model, that is, information in the video frame of the original video can be input to the adjustment model, and the information contains target features that can assist the adjustment model in extension.

[0095] In the embodiments of the present application, in order to ensure the accuracy of the adjustment model in filling the matching video content in the target video region during the adjustment of the video size of the original video, four parts of content can be input to the adjustment model, one part is noise information obtained by randomly sampling noise from the original video, one part is mask information indicating which region needs to be filled and which region is given information, one part is given video information, and one part is edge information. Through multiple aspects of content, the accuracy of the adjustment model in analyzing the original video content in the original video can be improved, and the accuracy of determining the matching video content matched with the original video content can also be improved, thereby achieving the technical effect of improving the accuracy of adjusting the video.

[0096] As an optional implementation, the first video information is extracted from the original video, including: extracting video features from the original video frame of the original video, wherein the first video information includes the video features, and the video features are used to represent objects in the original video content and / or scenes to which the original video content belongs.

[0097] In this embodiment, in the process of extracting the first video information from the original video, video features can be extracted from the original video frames of the original video, and the video features can be included in the first video information. The video features can be used to represent the objects in the original video content and / or the scenes described in the original video content, and the video features can be global features. The original video frames can be each video frame in the original video.

[0098] Optionally, if it is necessary to extract the video features in the original video, the original video can be subjected to a video feature extraction process.

[0099] Optionally, before the target video region corresponding to the original video is extended, the original video content contained in the original video can be analyzed and understood to extract a plurality of video information from the original video that can assist the model in adjusting the extension of the original video.

[0100] Optionally, the video frames of the original video and the video regions of the original video are input into an encoder to obtain the first video information, wherein the encoder can be a visual encoder.

[0101] As an optional implementation, the video features are extracted from the original video frames of the original video, including: encoding the original video content of the original video frames and the original video regions formed by the video region boundaries to obtain the video features.

[0102] In this embodiment, in the process of extracting the video features from the original video frames of the original video, the original video content of the original video frames and the original video regions formed by the video region boundaries can be encoded to obtain the video features.

[0103] Optionally, the original video content and the original video regions in the original video can be encoded by a visual encoder to obtain the video features.

[0104] Optionally, a pre-trained model is used for natural language understanding and computer vision tasks, for example, a visual encoder of CLIP can be used to process the video content in the original video to extract video features that can include information about the original characters, objects and scenes in the original video. The extraction of the video features is particularly important for subsequent size extension, which can greatly improve the consistency of the extension target.

[0105] It should be noted that the above method and process of extracting global features from the original video frames in the original video are only for illustration, and are not limited in this regard. Any process and method that can analyze and understand the information about the characters, objects and backgrounds contained in the original video frames are within the scope of protection of the embodiments of the present application.

[0106] As an optional implementation, the method further comprises: performing color detection on two adjacent original video frames in the original video to obtain color change information between the two adjacent original video frames; if the color change information is less than a color change threshold, determining that the two adjacent original video frames belong to the same shot segment; and extracting the video features from the original video frames of the original video, including: extracting the video features from the shot segment.

[0107] In this embodiment, color detection can be performed on two adjacent original video frames in the original video to obtain color change information between the two adjacent original video frames, and the size between the color change information and a color change threshold can be determined. If the color change information is less than the color change threshold, it is determined that the two adjacent original video frames belong to the same shot segment. Video features can be extracted from the shot segment, wherein the color change threshold can be a color change threshold, which can be pre-set according to an empirical value. It should be noted that the setting of the color change threshold and the threshold size is only for illustration, and is not specifically limited here.

[0108] Since a video often contains multiple shots in a real scene, in order to ensure the continuity between the shots and the consistency of the target in the video, the embodiments of the present application can perform shot splitting on the original video. In the process of performing shot splitting on the original video, whether adjacent video frames are the same shot segment can be determined by performing color detection on each video frame in the original video, and a preset color change threshold is used for shot differentiation to obtain multiple video shots for subsequent processing, thereby achieving the technical effect of improving the continuity between the shots and the consistency of the target in the video.

[0109] Optionally, in the process of video shot splitting, color detection can be performed on adjacent original video frames in the original video to obtain corresponding color change information. If the color change information is less than a color change threshold, it can be indicated that the adjacent frames are the same shot segment. If the color change information is greater than or equal to the color change threshold, it can be indicated that the adjacent frames are different shot segments. After video shot splitting, video feature extraction can be performed.

[0110] As an optional implementation, the first video information is extracted from the original video, including: performing edge detection on the original video to obtain edge information, wherein the first video information includes the edge information.

[0111] In this embodiment, in the process of extracting the first video information from the original video, edge detection can be performed on the original video to obtain edge information, wherein the first video information can include the edge information, which can be an edge detection map and can include texture information such as lines and contours of the target in the original video.

[0112] In the embodiments of the present application, in order to ensure the continuity between the extended area of the original video and the area of the original video, edge detection can be performed on the original video, that is, edge information containing lines, contours and the like of the target is extracted from the original video to provide certain assistance for subsequent size extension, thereby achieving the technical effect of improving the accuracy of video adjustment.

[0113] Optionally, edge detection is performed on the original video by an algorithm for detecting image edges, such as a Canny edge detection algorithm, to extract edge information that can assist in size extension from the original video.

[0114] It should be noted that the process and method of extracting edge information from the original video are only for illustration, and are not specifically limited herein, as long as the process and method of determining the texture information that can assist in size extension from the original video are within the protection scope of the embodiments of the present application.

[0115] As an optional implementation, the method further includes: in a case where the size adjustment parameter is to reduce the video size of the original video, determining second video information based on the size adjustment parameter and the original video size of the original video, wherein the video generation parameter includes the second video information, and the second video information is used to cause the adjustment model to generate matching video content for occluding the original video in the target video area.

[0116] In this embodiment, in a case where the size adjustment parameter is to reduce the video size of the original video, the second video information can be determined based on the size adjustment parameter and the original video size, and the second video information can be included in the video generation parameter. The second video information can be used to cause the adjustment model to generate matching video content that can occlude the original video in the target video area, that is, the part of the mask mask filled outside the original video.

[0117] Optionally, if it is determined that the video size of the original video needs to be reduced, the original video size and the size adjustment parameter for adjusting the original video can be used to determine the area of the original video that needs to be occluded when the original video is reduced, that is, the target video area. The matching video content that needs to cover the target video area and can match the original video content of the original video, that is, the video content that needs to be covered when the original video is reduced, can also be determined.

[0118] Due to different requirements of different media on the size of the video, the user often only makes the video once when making the video, and needs to be promoted to multiple media, which requires the video size conversion system to modify the size. If the video size conversion is to cut the video content by using the cropping tool, this form will seriously damage the integrity of the objects, characters and other targets in the video, and the viewer experience is poor. Therefore, how to perform size conversion without damaging the original video content is a problem worth exploring. In the embodiments of the present application, a diffusion model can be used to understand the original video content in the original video, and when the size of the original video is enlarged, the knowledge learned by the diffusion model can be used to imaginatively expand the content. When the size of the original video is reduced, the knowledge learned by the expansion model can be used to imaginatively cover the original video. This greatly protects the integrity of the original video content, and at the same time, the proposed multi-level cascade processing strategy is used to realize the size conversion of videos of any length. It provides a convenient video processing tool for users to promote videos, and improves the viewing experience of videos. Thus, the technical effect of effectively converting the video size of the video is realized.

[0119] As an optional implementation, the method further includes: determining a noise sampling strategy corresponding to the adjustment model, wherein the noise sampling strategy is used to represent a sampling rule of noise samples, and the noise samples are used to train the adjustment model with the image content samples; and sampling the original video by using the noise sampling strategy to obtain noise information, wherein the video generation parameter includes the noise information.

[0120] In this embodiment, the noise sampling strategy corresponding to the adjustment model can be determined, and the noise sampling strategy can be used to sample the noise in the original video to obtain the noise information in the original video. The noise sampling strategy can be used to represent a sampling rule of noise samples. The noise information can be included in the video generation parameter.

[0121] Optionally, if the adjustment model needs to more accurately identify and size-convert the original video, noise needs to be input to the adjustment model, that is, some noise information in the original video needs to be input to the adjustment model, so that the adjustment model can learn the original video content in the original video to determine the matching video content.

[0122] Optionally, the noise in the original video can be sampled by using the noise sampling strategy, for example, a random noise sampling strategy can be used to sample random noise in the original video to obtain the noise information in the original video, and the noise information can be input to the adjustment model.

[0123] For example, the Gaussian sampling method can be used to obtain the (random) Gaussian noise from the original video, and the (random) Gaussian noise can be input into the diffusion model. It should be noted that the above noise sampling strategy and the noise information obtained according to the noise sampling strategy are only for illustration, and are not specifically limited here, as long as the method and process of noise sampling on the original video and input into the adjustment model for size transformation of the original video are within the protection scope of the embodiments of the present application.

[0124] As an optional implementation, in step S208, the original video is adjusted in the adjustment model using the size adjustment parameter to obtain the target video, including: in the adjustment model, converting the video generation parameter into the to-be-decoded information matching the video content; decoding the to-be-decoded information to obtain the pixel of the matching video content; and generating the target video based on the pixel of the matching video content and the pixel of the original video content.

[0125] In this embodiment, in the process of adjusting the original video in the adjustment model using the size adjustment parameter, the video generation parameter can be converted into the to-be-decoded information matching the video content in the adjustment model, the to-be-decoded information can be decoded to obtain the pixel corresponding to the matching video content adapted to the original video, and the target video corresponding to the original video after video size transformation can be generated based on the pixel of the matching video content and the pixel of the original video content, wherein the to-be-decoded information can be the to-be-decoded information after extension in the adjustment model. The pixel of the matching video content can be the extended pixel representation of the video used for extending the original video.

[0126] Optionally, after the noise information, the mask information, the edge information, and the first video information are determined, the above four kinds of information can be input into the trained adjustment model, and the noise reduction process of the adjustment model can be performed.

[0127] For example, in the noise reduction process of the adjustment model, the above four kinds of information can be subjected to fifty-step noise reduction, so that the input random Gaussian noise can be converted into the extended to-be-decoded information. It should be noted that the number of times of noise reduction performed in the above noise reduction process in the adjustment model is only for illustration, and is not specifically limited here.

[0128] Optionally, the to-be-decoded information further includes the content of the input original video obtained by the feature encoder. The above to-be-decoded information can be decoded to obtain the output video.

[0129] Optionally, in the decoding process, based on the to-be-decoded information other than the content obtained by the feature encoder from the original video, the video extension pixel representation of the current input can be generated. Based on the video extension pixel representation at this time and the content obtained by the feature encoder after decoding, the output video can be generated.

[0130] As an optional implementation, the method further includes: performing super-resolution processing on the original video, wherein the video size of the original video after super-resolution processing is higher than the video size of the original video before super-resolution processing; and determining the video size of the original video after super-resolution processing as the video size of the original video.

[0131] In this embodiment, the original video can be subjected to super-resolution processing, and the video size of the original video after super-resolution processing can be determined as the video size of the original video, wherein the video size of the original video after super-resolution processing is higher than the video size of the original video before super-resolution processing. The super-resolution processing can be used to increase the resolution of the video frame. The video size of the original video after super-resolution processing can be the video super-resolution of the original video.

[0132] In the embodiments of the present application, due to the limitation of machine resources during training of the adjustment model, for example, the trained adjustment model is trained using a video with a size of 256x256, therefore, the video after size adjustment is also 256x256. In order to ensure that the resolution of the video after size adjustment can meet the size of the video after adjustment, the obtained video can be subjected to video super-resolution processing.

[0133] Optionally, using the trained super-resolution deep learning model, the original video can be subjected to several times of super-resolution, for example, four times of super-resolution, if the original video is 256x256, four times of super-resolution can generate a video with a resolution of 1024x1024, and the final target video can be obtained by cropping and scaling according to the size of the user.

[0134] It should be noted that the number of times and the process of super-resolution of the original video in the process of video super-resolution processing are only for illustration, and are not limited specifically herein.

[0135] As an optional implementation, the original video includes original video frames, the step S204 of determining the size adjustment parameter of the original video includes: determining the size adjustment parameter of the original video frame; the step S206 of inputting the size adjustment parameter and the original video into the adjustment model includes: inputting the size adjustment parameter and the original video frame into the adjustment model; the step S208 of adjusting the original video in the adjustment model at least by using the size adjustment parameter to obtain a target video includes: adjusting the original video frame in the adjustment model at least by using the size adjustment parameter to obtain a target video frame, wherein the target video includes the target video frame; the method further includes: in response to the original video frame having a next original video frame, determining the next original video frame of the original video frame as the original video frame, and returning to execute from the following steps; determining the size adjustment parameter of the original video frame until the original video frame does not have a next original video frame.

[0136] In this embodiment, the size adjustment parameter of the original video frame can be determined, and the size adjustment parameter and the original video frame can be input into the adjustment model. In the adjustment model, the original video frame can be adjusted at least by using the size adjustment parameter to obtain the target video frame in the target video. When the original video frame has a next original video frame, the next original video frame of the original video frame can be determined as the original video frame, and the above steps can be returned to continue to adjust the original video frame and determine the corresponding size adjustment parameter until the original video frame does not have a next original video frame, and the determination of the size adjustment parameter ends, wherein the original video can include the original video frame, and the target video can include the target video frame.

[0137] In the embodiments of the present application, since a limited number of videos, such as 16 frames of videos, can be processed at a time when the adjustment model is set, in order to enable the adjustment model to process videos of any length, a multi-stage cascading inference scheme can be proposed, and the process of determining the size adjustment parameter, inputting the size adjustment parameter and the original video frame into the adjustment model, and adjusting the original video to obtain the corresponding target video frame can be executed multiple times in a cascading manner. In order to ensure the consistency of the target in a longer video, the video frame generated in the previous stage can be used as input in each stage processing process, that is, the size adjustment parameter of the current original video frame can be determined by using the previous original video frame, and the current original video frame can be used to determine the size adjustment parameter of the next original video frame. Finally, the complete output of the longer video can be obtained, thereby achieving the purpose that the adjustment model can process videos of any length and realizing the technical effect of improving the flexibility of the size transformation of the adjustment model.

[0138] As an optional implementation, in response to the original video frame existing the next original video frame, the next original video frame of the original video frame is determined as the original video frame, comprising: in response to the original video frame existing the next original video frame, the next original video frame of the original video frame and the target video frame are determined as the original video frame.

[0139] In this embodiment, in the process of determining the next original video frame of the original video frame as the original video frame when the original video frame exists the next original video frame, the next original video frame of the original video frame and the target video frame can be determined as the original video frame.

[0140] In the embodiments of the present application, in order to ensure the consistency of the target in a video, the generated video frame of the previous stage will be used as input when performing each stage of processing, that is, the target video frame determined by the current original video frame can be input into the processing process of the next original video frame of the current original video frame, and the next target video frame corresponding to the next original frame is determined by using the target video frame corresponding to the next original video frame and the current original video frame, thereby realizing the technical effect of ensuring the consistency of the target in the video.

[0141] As an optional implementation, the method further comprises: determining an image content sample from the video sample; training the deep learning model using attribute information of different objects in the image content sample to obtain a diffusion model; and determining the diffusion model as the adjustment model.

[0142] In this embodiment, the image content sample can be determined from the video sample, and the deep learning model can be trained using the attribute information of different objects in the image content sample to obtain a diffusion model, and the diffusion model can be determined as the adjustment model, wherein the attribute information of the object can be the shape and appearance of the object, and the like, which is only used as an example and is not limited.

[0143] Optionally, in order to transform the video size of the original video, a large number of videos can be pre-set, and the deep learning model can be trained using the image content sample in each video frame of the large number of videos. The deep learning model can learn various attributes such as the shape and appearance of different objects appearing in the above-mentioned large number of videos, thereby obtaining a diffusion model. The diffusion model can be used to complete the less region from the learned knowledge when a less region is given, thereby realizing video size transformation.

[0144] For example, if it is an e-commerce scene, the diffusion model can be learned from a large number of e-commerce videos, and the model can learn the shape and appearance of common items from the e-commerce videos, thereby realizing the completion of the video content from the learned knowledge when a less region is given, thereby realizing size adjustment.

[0145] The embodiment of the present application further provides a video processing method, Figure 3 is a flow chart of a video processing method according to the embodiment of the present application, as shown in the figure, the method can comprise the following steps: Figure 3

[0146] Step S302, obtaining an original video to be adjusted from an e-commerce platform.

[0147] In the technical solution provided in the above step S302 of the present application, the original video to be adjusted can be obtained from the e-commerce platform.

[0148] Optionally, if the use scenario of video size transformation is an e-commerce scenario, the e-commerce platform can be included in the client, the original video to be adjusted can be made in the e-commerce platform, the video size of the original video can be adjusted to obtain a target video, and the target video can be sent to a media platform.

[0149] Step S304, determining a size adjustment parameter of the original video based on a media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust the video size of the original video.

[0150] In the technical solution provided in the above step S304 of the present application, after obtaining the original video to be adjusted from the e-commerce platform, the size adjustment parameter of the original video can be determined, wherein the size adjustment parameter can be used to adjust the video size of the original video.

[0151] Optionally, after determining that the video size of the original video needs to be adjusted, the size adjustment parameter required to adjust the video size of the original video to the required video size can be determined.

[0152] Optionally, after monitoring the original video to be adjusted, the original video size of the original video can be determined, and the target video size to which the original video needs to be adjusted can also be determined, and the size adjustment parameter can be determined based on the original video size and the target video size.

[0153] Optionally, if the user inputs a plurality of specified output sizes to which the original video needs to be adjusted, that is, inputs a plurality of target video sizes, the target video size under the condition that one of the length or width of the original video is kept unchanged can be calculated, and the union of the regions represented by the plurality of target video sizes is taken to obtain the final extended size adjustment parameter.

[0154] Step S306, inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples.

[0155] ​In the technical solution provided in step S306 of the present application, after determining the size adjustment parameter of the original video based on the media platform to which the original video is to be pushed, the size adjustment parameter and the original video can be input into the adjustment model, wherein the adjustment model is trained based on image content samples.

[0156] Optionally, after determining the size adjustment parameter of the original video, the size adjustment parameter and the original video can be input into the adjustment model, and the original size of the original video can be adjusted according to the size adjustment parameter through the adjustment model.

[0157] Optionally, the adjustment model can be trained in advance according to a large number of videos, that is, the adjustment model can learn the shapes and appearances of common objects and other image contents in each frame of the videos through learning in the large number of videos.

[0158] Optionally, after the adjustment model is trained through a large number of videos, the adjustment model can realize the size transformation by complementing the video content in the video with the knowledge learned in the training process when the video of a given small area is input.

[0159] In step S308, the original video is adjusted in the adjustment model by using at least the size adjustment parameter to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region offset from the video region boundary of the original video in a direction and a distance corresponding to the size adjustment parameter.

[0160] In the technical solution provided in step S308 of the present application, after the size adjustment parameter and the original video are input into the adjustment model, the video size of the original video can be adjusted in the adjustment model by using the size adjustment parameter to obtain a target video meeting the needs of the client, wherein the target video can include original video content from the original video and matching video content matched with the original video content, and the target video can be the final extended video obtained by adjusting the size of the original video according to the needs of the client. The original video content can be the shapes and appearances of objects and persons and other information contained in the original video. The matching video content can be displayed in a target video region and can be represented by extended pixels obtained by extending some objects or persons in the process of extending the original video. The target video region can be used to represent a region offset from the video region boundary of the original video in a direction and a distance corresponding to the size adjustment parameter, and can also be referred to as an extended region.

[0161] Optionally, the good visual content generation capability in the adjustment model, i.e., the capability of the adjustment model to match the original video content in the video with the corresponding matching video content according to the learned image content, is utilized. In the adjustment model, the original video can be adjusted so that the picture quality, picture texture, color, and target consistency of the original video can be improved.

[0162] Step S310, pushing the target video to the media platform for playing.

[0163] In the technical solution provided in the above step S310 of the present application, after the original video is adjusted in the adjustment model to obtain the target video, the target video can be pushed to the media platform for playing.

[0164] Optionally, after the original video is adjusted to obtain the target video that meets the size standard of the media platform, the target video can be sent to the media platform for playing.

[0165] Through the above steps S302 to S310 of the present application, the original video to be adjusted is obtained from the e-commerce platform; the size adjustment parameter of the original video is determined based on the media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust the video size of the original video; the size adjustment parameter and the original video are input into the adjustment model, wherein the adjustment model is trained based on image content samples; in the adjustment model, the original video is adjusted at least by using the size adjustment parameter to obtain the target video, wherein the target video content of the target video includes the original video content from the original video and the matching video content matched with the original video content, the matching video content is displayed in the target video region, and the target video region is used to represent the region offset from the video region boundary of the original video in the direction and distance corresponding to the size adjustment parameter; the target video is pushed to the media platform for playing, thereby achieving the technical effect of improving the video adjustment effect and solving the technical problem of poor video adjustment effect.

[0166] The present application also provides a video processing method, Figure 4 is a flowchart of a video processing method according to an embodiment of the present application, as shown in Figure 4 The method can include the following steps:

[0167] Step S402, in response to the input instruction acting on the operation interface, displaying the original video to be adjusted and the size adjustment parameter of the original video on the operation interface, wherein the size adjustment parameter is used to adjust the video size of the original video.

[0168] In the technical solution provided in step S402 of the present application, if it is necessary to adjust the video size of the original video, a corresponding input instruction can be executed on the operation interface, the original video is sent to the server for size adjustment, and the original video can be adjusted on the operation interface, and the size adjustment parameters of the original video can be displayed, wherein the size adjustment parameters can be used to adjust the video size of the original video.

[0169] Optionally, if it is necessary to adjust the original video, an input instruction can be executed on the operation interface, the original video is input to the operation interface, and the size to which the original video needs to be adjusted can be input.

[0170] Optionally, if the user inputs a plurality of specified output sizes to which the original video needs to be adjusted, that is, a plurality of target video sizes are input, the target video sizes under the condition that one of the length or width of the original video is maintained can be calculated, and the union of the regions represented by the plurality of target video sizes is obtained to obtain the final size adjustment parameter of the extension.

[0171] For example, the name and location of the original video to be adjusted in size can be input in the input box in the operation interface, a corresponding instruction is generated to call the original video from the location to display on the operation interface, and the original video can be adjusted in size. It should be noted that the above process of triggering the input instruction and operation is only an example and is not limited herein.

[0172] Step S404, in response to the video adjustment instruction acting on the operation interface, displaying the target video on the operation interface, wherein the target video is obtained by adjusting the original video using at least the size adjustment parameter in the adjustment model, the adjustment model is trained based on the image content sample, the target video content of the target video includes the original video content from the original video, and the matching video content matched with the original video content, the matching video content is displayed in the target video region, and the target video region is used to represent the region offset from the video region boundary of the original video along the direction and distance corresponding to the size adjustment parameter.

[0173] In the technical solution provided in the step S404 above, after the original video and the corresponding size adjustment parameter are displayed on the operation interface, a video adjustment instruction can be triggered on the operation interface, the video size of the original video can be adjusted, and the target video after the adjustment is displayed on the operation interface, wherein the target video can be obtained by adjusting the original video by at least using the size adjustment parameter in the adjustment model. The adjustment model can be trained based on image content samples. The target video content of the target video can include original video content from the original video and matching video content matched with the original video content. The matching video content is displayed in the target video area. The target video area can be used to represent an area offset from the video area boundary of the original video in a direction and distance corresponding to the size adjustment parameter.

[0174] Optionally, after the size adjustment parameter of the original video is determined, the size adjustment parameter and the original video can be input into the adjustment model, and the original size of the original video can be adjusted according to the size adjustment parameter by using the adjustment model.

[0175] Optionally, the adjustment model can be trained in advance according to a large number of videos, that is, the adjustment model can learn the shapes and appearances of common objects and other image contents in each frame of image in the videos by learning in the large number of videos.

[0176] Optionally, after the adjustment model is trained by using a large number of videos, the adjustment model can realize size transformation by using the knowledge learned in the training process to complete the video content in the video when a given video has a small area.

[0177] Optionally, by using the good visual content generation capability of the adjustment model, that is, by using the ability of the adjustment model to match corresponding matching video content to the original video content in the video according to the learned image content, the original video can be adjusted in the adjustment model, so that the picture quality, picture texture, color, and target consistency of the original video can be improved.

[0178] By the steps S402 to S404 provided in the application, the original video to be adjusted and the size adjustment parameter of the original video are displayed on the operation interface in response to the input instruction acting on the operation interface, wherein the size adjustment parameter is used to adjust the video size of the original video; the target video is displayed on the operation interface in response to the video adjustment instruction acting on the operation interface, wherein the target video is obtained by adjusting the original video by at least using the size adjustment parameter in the adjustment model, the adjustment model is obtained based on the image content sample, the target video content of the target video includes the original video content from the original video and the matching video content matched with the original video content, the matching video content is displayed in the target video area, and the target video area is used to represent the area offset from the video area boundary of the original video in the direction and distance corresponding to the size adjustment parameter. Thus, the technical effect of improving the video adjustment effect is achieved, and the technical problem of poor video adjustment effect is solved.

[0179] The application further provides a video processing method, Figure 5 The application provides a video processing method, Figure 5 The application provides a video processing method,

[0180] Step S502: displaying the original video to be adjusted on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device.

[0181] In the technical solution provided in the step S502, the original video to be adjusted can be displayed on the presentation screen of the virtual reality (VR) device or the augmented reality (AR) device.

[0182] Step S504: determining a size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video.

[0183] In the technical solution provided in the step S504, the size adjustment parameter of the original video can be determined, and the size adjustment parameter can be used to adjust the video size of the original video.

[0184] Optionally, after the original video to be adjusted is monitored, the original video size of the original video can be determined, and the target video size to which the original video needs to be adjusted can also be determined, and the size adjustment parameter is determined based on the original video size and the target video size.

[0185] Step S506: inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is obtained based on an image content sample.

[0186] In the technical solution provided in step S506 of the present application, after the size adjustment parameter of the original video is determined, the size adjustment parameter and the original video can be input into the adjustment model, wherein the adjustment model can be obtained based on image content samples.

[0187] Optionally, after the size adjustment parameter of the original video is determined, the size adjustment parameter and the original video can be input into the adjustment model, and the original size of the original video can be adjusted according to the size adjustment parameter through the adjustment model.

[0188] Optionally, the adjustment model can be trained in advance according to a large number of videos, that is, the adjustment model can learn the shapes and appearances of common objects and other image contents in each frame of the video through learning in a large number of videos. After the adjustment model is trained through a large number of videos, the adjustment model can complete the video content and realize size transformation based on the knowledge learned in the training process when a given video has a small area.

[0189] Step S508, in the adjustment model, at least the size adjustment parameter is used to adjust the original video to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region offset from the video region boundary of the original video in a direction and a distance corresponding to the size adjustment parameter.

[0190] In the technical solution provided in step S508 of the present application, after the size adjustment parameter and the original video are input into the adjustment model, the video size of the original video can be adjusted in the adjustment model by using the size adjustment parameter to obtain a target video meeting the needs of the client.

[0191] In the embodiment of the present application, the good visual content generation capability of the adjustment model is used, that is, the ability of the adjustment model to match the original video content in the video with corresponding matching video content according to the learned image content. In the adjustment model, the original video can be adjusted so that the picture quality, picture texture, color and target consistency of the original video can be improved. The diffusion model can be used to understand the content of the video, and the knowledge learned by the diffusion model can be used to expand the content of the original video with imagination, which can greatly protect the integrity of the original content of the original video, thereby realizing the technical effect that the size of the video can be effectively adjusted.

[0192] Step S510, driving the VR device or the AR device to display the target video.

[0193] In the technical solution provided in the foregoing step S510 of the present application, the VR device or the AR device can be driven to display the target video.

[0194] Optionally, after adjusting the video size of the original video to obtain the corresponding target video, the target video can be input to the VR device or the AR device, and the VR device or the AR device can be driven to display the target video.

[0195] By the foregoing steps S502 to S510 of the present application, the original video to be adjusted is displayed on the presentation screen of the virtual reality (VR) device or the augmented reality (AR) device; a size adjustment parameter of the original video is determined, wherein the size adjustment parameter is used to adjust the video size of the original video; the size adjustment parameter and the original video are input to an adjustment model, wherein the adjustment model is trained based on image content samples; in the adjustment model, the original video is adjusted at least by using the size adjustment parameter to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region offset from the video region boundary of the original video in a direction and a distance corresponding to the size adjustment parameter; and the VR device or the AR device is driven to display the target video, thereby achieving the technical effect of improving the video adjustment effect and solving the technical problem of poor video adjustment effect.

[0196] Embodiment 2

[0197] According to the embodiments of the present application, an embodiment of a video processing system is also provided, Figure 6 is a schematic diagram of a video processing system according to an embodiment of the present application, as Figure 6 shown, the video processing system 600 can include an e-commerce platform 601, a video adjustment end 602, and a media platform 603.

[0198] The e-commerce platform 601 is configured to produce an original video to be adjusted.

[0199] In this embodiment, the original video to be adjusted can be produced in the e-commerce platform 601.

[0200] Optionally, the user can produce an original video for displaying a certain product in the e-commerce platform 601 according to his own needs, such as the product to be displayed, the background of the displayed product, and the like, and can send the original video to the video adjustment end 602 to adjust the video size of the original video.

[0201] Optionally, the user can also input the required size to be adjusted to the video adjustment end 602 according to his own needs.

[0202] The video adjustment end 602 is configured to determine a size adjustment parameter of the original video based on a media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust a video size of the original video; input the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; and adjust the original video by using at least the size adjustment parameter in the adjustment model to obtain a target video, wherein a target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region offset from a video region boundary of the original video in a direction and a distance corresponding to the size adjustment parameter.

[0203] In this embodiment, in the video adjustment end 602, the size adjustment parameter of the original video can be determined, the size adjustment parameter and the original video can be input into the adjustment model, and the original size of the original video can be adjusted according to the size adjustment parameter by using the adjustment model. In the adjustment model, the video size of the original video can be adjusted by using the size adjustment parameter to obtain a target video meeting the needs of a client. The adjusted target video can be input into the media platform 603.

[0204] Optionally, after the original video to be adjusted is monitored, the original video size of the original video can be determined, and the target video size to which the original video needs to be adjusted can also be determined, and the size adjustment parameter can be determined based on the original video size and the target video size.

[0205] Optionally, the adjustment model can be trained in advance according to a large number of videos, that is, the adjustment model can learn the shapes and appearances of common objects and other image contents in each frame of the videos, so that the adjustment model can complete the video content in the video and realize size transformation based on the knowledge learned in the training process when a given video has a small region.

[0206] Optionally, the good visual content generation capability of the adjustment model, that is, the ability of the adjustment model to match the original video content in the video with corresponding matching video content according to the learned image content, can be used to adjust the original video in the adjustment model, so that the picture quality, picture texture, color, and target consistency of the original video can be improved.

[0207] The media platform 603 is configured to play the target video.

[0208] In this embodiment, the target video can be played in the media platform 603.

[0209] Optionally, after adjusting the video size of the original video according to the size standard of the video to be pushed by the media platform 603 to obtain a target video conforming to the size standard of the media platform, the target video can be sent to the media platform 603 for playing the target video in the media platform 603.

[0210] In this embodiment, a video processing system is provided. An original video to be adjusted is made through an e-commerce platform 601; a size adjustment parameter of the original video is determined based on a media platform to which the original video is to be pushed through a video adjustment end 602, wherein the size adjustment parameter is used to adjust the video size of the original video; the size adjustment parameter and the original video are input into an adjustment model, wherein the adjustment model is trained based on image content samples; in the adjustment model, the original video is adjusted at least by using the size adjustment parameter to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region offset from a video region boundary of the original video in a direction and a distance corresponding to the size adjustment parameter; and the target video is played through the media platform 603, thereby achieving the technical effect of improving the effect of video adjustment and solving the technical problem of poor effect of video adjustment.

[0211] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application, such as data for verification, are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0212] Embodiment 3

[0213] At present, in the current digital media era, video content is widely spread on various platforms. Because of its rich display form and high information dissemination degree, video has gradually replaced pictures to become a medium for users to record life and businesses to market. However, different forms of display in various media have different requirements for the size of the video. How to adapt the same original video to different size display areas is a basic but challenging problem, so the video size conversion technology becomes particularly important.

[0214] In the related art, the video size transformation technology is more to crop the original video content, but it will seriously damage the elements of the original video content, and will cause the incomplete of the characters, objects and the like in the video, resulting in information loss. Therefore, there is still a technical problem of poor video adjustment effect.

[0215] Further, the application provides a video size transformation method based on diffusion model picture completion. If the size of an original video needs to be adjusted, it can be monitored whether the original video to be adjusted is received. If the original video to be transmitted is monitored, the size adjustment parameter for adjusting the video size of the original video can be determined. The adjustment model can be trained in advance according to the image content sample, that is, the ability of the adjustment model to adjust the size of the original video can be trained according to the shape and appearance of some common image content. The size adjustment parameter and the original video are input into the adjustment model, and the adjustment model is used to adjust the video size of the original video according to the size adjustment parameter to obtain a target video. In the target video, not only the original video content from the original video, but also the matching video content matched by the adjustment model according to the original video content can be included, and the matching video content can be displayed in the target video area, that is, the matching video content can be displayed in the area obtained by offsetting from the video area boundary of the original video in the direction and distance corresponding to the size adjustment parameter. Since the embodiment of the application can accurately adjust the video size of the original video to the video size of the target video required by using the adjustment model according to the size adjustment parameter and the original video, the loss of video content caused by simply cropping the video is avoided, thereby achieving the purpose of ensuring the integrity of the video content, and further realizing the technical effect of improving the video adjustment effect, solving the technical problem of poor video adjustment effect.

[0216] The above method of the embodiment will be further introduced below.

[0217] In the embodiment of the application, a novel video size transformation system can be provided, which accurately converts video content from one size to another size while maintaining the integrity of the original content in the video by using diffusion model picture completion technology. The diffusion model can be learned in a large number of e-commerce videos, and the model can learn the shape and appearance of common objects, so as to realize the completion of video content from learned knowledge when a small area is given, and realize size transformation. In addition, the extension capability of the diffusion model in long video generation is further explored, and a multiple cascade inference strategy is proposed to realize the size extension of any long video.

[0218] Optionally, promoting commodities through short videos is a rapidly developing promotion form in recent years, and major Internet mainstream media all provide short video commodity promotion services. However, different media have different requirements for the size of the video. For merchants, when making a short video, it is often only made once, and needs to be promoted to multiple media, which requires a video size conversion system to modify the size. Typically, video size conversion uses a cropping tool to cut the video content, which can severely damage the integrity of objects, characters and other targets in the video, and the viewer experience is poor. Therefore, how to perform size conversion without damaging the original video content is a problem worth exploring. The present scheme proposes to use a diffusion model to understand the content of the video, and use the knowledge learned by the model to expand the content imaginatively, which greatly protects the integrity of the original content of the video. At the same time, using the proposed multi-level cascade processing strategy, the size conversion of videos of any length is realized, providing a convenient video processing tool for merchants to promote short videos, and improving the viewing experience of consumers.

[0219] Figure 7 is a flowchart of a video size conversion method based on a diffusion model picture completion according to an embodiment of the application, as shown in Figure 7 , the method can include the following steps:

[0220] Step S702, obtaining an original video to be size-converted.

[0221] In this embodiment, if a user needs to adjust the video size of a certain original video, the corresponding operation can be performed in the corresponding client to trigger the operation of transmitting the original video to the server, and the original video can be transmitted to the server through the network.

[0222] Step S704, based on the original video, performing expansion parameter calculation to obtain a size adjustment parameter.

[0223] In this embodiment, the expansion parameter calculation can be performed based on the original video to obtain the size adjustment parameter.

[0224] Optionally, if the user inputs multiple specified output sizes to which the original video needs to be adjusted, that is, inputs multiple target video sizes, the target video sizes under the condition that one of the length or width of the original video is kept unchanged can be calculated, and the union of the regions represented by the multiple target video sizes is obtained to obtain the final expanded size adjustment parameter.

[0225] Optionally, for the video original size input by the user and the specified multiple output sizes, first calculate the minimum expansion area to determine the area of the expansion area above, below, left and right of the original video, and construct a mask mask input to the subsequent expansion step.

[0226] For example, if the original video has a video original size of 512x512 pixels, and two target video sizes are 720x1280 pixels and 1280x720 pixels, respectively, the video original size can be first extended to 512x910 pixels and 910x512 pixels, and then the union can be taken to calculate the final size of the original video to be extended, that is, the size adjustment parameter of the original video can be 910x910.

[0227] Step S706, video shot splitting is performed on the original video.

[0228] In this embodiment, in a real scene, a video often contains multiple shots, in order to maintain the continuity between shots and the consistency of the target, the present scheme will first split the input video into shots. By color detection on the video frames, it is determined whether adjacent frames are the same shot segment, and a preset color change threshold is used for shot differentiation, so as to obtain multiple video shots for subsequent processing.

[0229] Optionally, in the process of video shot splitting, color detection can be performed on adjacent original video frames in the original video to obtain corresponding color change information. If the color change information is less than the color change threshold, it can be indicated that the adjacent frames are the same shot segment. If the color change information is greater than or equal to the color change threshold, it can be indicated that the adjacent frames are different shot segments. After video shot splitting, video feature extraction can be performed.

[0230] Step S708, video feature extraction is performed on the original video after video shot splitting to obtain the original video content.

[0231] In this embodiment, before video region extension, content analysis and understanding of the input video are needed to extract various video information that can assist in extension. A pre-trained CLIP visual encoder is used to extract global features from the video content, which contains information such as original characters, objects and scenes in the video. The extraction of these information is particularly important for subsequent size extension, which can greatly improve the consistency of the extension target.

[0232] Optionally, in order to ensure the continuity of the extended region and the existing region, Canny edge detection information is also extracted from the input video, which contains texture information such as lines and contours of the target, providing certain assistance for subsequent size extension.

[0233] Step S710, extension region generation is performed on the original video.

[0234] In this embodiment, extension region generation on the input video is a very important step.

[0235] Optionally,Figure 8 This is a schematic diagram of a video size change system based on diffusion model image completion according to an embodiment of this application, such as... Figure 8 As shown, the system can include an encoder 801, a diffusion model 802, and a feature encoder 803. Based on the superior generative capabilities of the diffusion model, a complete video resizing system can be built, enabling the transformation of input videos of any size and duration into videos of a specified size for various applications. The main process includes: video stretching parameter calculation, video shot breakdown, video feature extraction, stretching region generation, and video super-resolution. The entire video resizing process is fully automated; users only need to input the video and specify the output size to obtain the corresponding output video. The entire resizing process does not involve cropping of people, objects, or other targets.

[0236] Optionally, such as Figure 8 As shown, the encoder 801 first encodes information from the given video region and the given video frame, while randomly sampling noise. The trained diffusion model then inputs this information. The diffusion model 802 can receive sampled noise (noise information), mask information, given video information, and edge detection maps.

[0237] Optionally, random Gaussian noise is obtained through Gaussian sampling, a step consistent with the training process of the diffusion model, to obtain noise information. This determines which regions need to be filled and which regions are mask information for the given information. It also determines the information representing the original input video frame, which includes target features that can assist in the extension process; that is, it determines the given video information. Finally, it determines the Canny edge detection map representing the texture information of the input video, which is also an important feature for assisting in the extension. All of the above information is then input into the trained diffusion model 802. Alternatively, the original video input feature encoder 803 can be processed, and the processed content can be input into the diffusion model 802 to perform the diffusion model's noise reduction process. After 50 steps of noise reduction, the input Gaussian noise is converted into extended information to be decoded. The final step, the decoding process, generates the extended pixel representation of the current input video.

[0238] Step S712: Perform super-resolution processing on the video after generating the extended region to obtain the target video.

[0239] In this embodiment, due to machine resource limitations during diffusion model training, a 256x256 resolution video was used for training. Therefore, the stretched video is also 256x256 in size. In this step, the trained super-resolution deep learning model is used to perform a 4x super-resolution on the input video, generating a 1024x1024 resolution video. This video is then cropped and scaled according to the user's desired size to obtain the final stretched video.

[0240] Step S714, output the target video.

[0241] In this embodiment, the target video can be input into the corresponding media platform for display.

[0242] Optionally, Figure 9 is a schematic diagram of a multi-stage cascade processing of an original video according to an embodiment of the present application, as Figure 9 shown, if a video has a total of 451 frames (0 frame to 450 frame), which is coarse-grained, since the diffusion model is set to process only 16 frames of video at a time, in order to process an arbitrarily long video, a multi-stage cascade inference scheme can be proposed, which performs multiple times through the cascade manner for the above-mentioned five steps. In order to ensure the consistency of the target in the long video, the generated video frame of the previous stage is used as the input in each processing process. Based on the above steps, the video can be gradually processed in fine-grained manner, and finally the complete long video output is obtained.

[0243] Optionally, as Figure 9 shown, in order to ensure the consistency of the target in a video, the generated video frame of the previous stage will be used as the input in each processing process. That is, the target video frame determined by the current original video frame can be input into the processing process of the next original video frame of the current original video frame, and the next target video frame corresponding to the next original frame is determined by using the next original video frame and the target video frame corresponding to the current original video frame, thereby realizing the technical effect of ensuring the consistency of the target in the video.

[0244] For example, FIG. 10(a) is a schematic diagram of an example of a target video obtained by adjusting an original video according to an embodiment of the present application, as shown in FIG. 10(a), the original video A is an e-commerce video showing a shoe, and the merchant holds the shoe in front of the lens to show it. However, the original video size of the original video A cannot completely supplement the shoe and the hand movement of the merchant, therefore, the original video A can be resized, and the appearance of the shoe in the obtained target video A is complete, and the complete hand movement of the merchant can be shown.

[0245] For another example, FIG. 10(b) is a schematic diagram of another example of a target video obtained by adjusting an original video according to an embodiment of the present application, as shown in FIG. 10(b), the original video B is an e-commerce video in which a merchant is showing a cabinet. However, the original video size of the original video B cannot completely supplement the appearance of the cabinet, therefore, the original video B can be resized, and the other side cabinet door in the obtained target video B is completely described.

[0246] In the embodiment of the present application, if it is necessary to adjust the size of a certain original video, it can be monitored whether the original video to be adjusted is received. If the original video to be transmitted is monitored, it can be determined that the size adjustment parameter for adjusting the video size of the original video is needed. The adjustment model can be trained in advance according to the image content sample, that is, the ability of the adjustment model to adjust the size of the original video can be trained according to the shape and appearance of some common image content. The size adjustment parameter and the original video are input into the adjustment model, and the adjustment model is used to adjust the video size of the original video according to the size adjustment parameter to obtain a target video. In the target video, not only the original video content from the original video, but also the matching video content matched by the adjustment model according to the original video content can be included, and the matching video content can be displayed in the target video area, that is, the matching video content can be displayed in the area obtained by offsetting from the video area boundary of the original video in the direction and distance corresponding to the size adjustment parameter. Since the embodiment of the present application can accurately adjust the video size of the original video to the video size of the target video required by using the adjustment model according to the size adjustment parameter and the original video, the situation of missing video content caused by simply cropping the video is avoided, so that the purpose of ensuring the integrity of the video content is achieved, and the technical effect of improving the video adjustment effect is realized, and the technical problem of poor video adjustment effect is solved.

[0247] Embodiment 4

[0248] According to the embodiments of the present application, a video processing method for implementing the above-mentioned video processing method is also provided. Figure 2 A video processing device for implementing the video processing method shown in the above-mentioned video processing method is also provided.

[0249] Figure 11 is a schematic diagram of a video processing device according to the embodiments of the present application, as shown in the above-mentioned video processing device 1100 can include a monitoring unit 1102, a first determination unit 1104, a first input unit 1106 and a first adjustment unit 1108. Figure 11 The monitoring unit 1102 is configured to monitor the original video to be adjusted.

[0250] The first determination unit 1104 is configured to determine the size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video.

[0251] The first input unit 1106 is configured to input the size adjustment parameter and the original video into the adjustment model, wherein the adjustment model is trained based on the image content sample.

[0252]

[0253] ​The first adjustment unit 1108 is used to adjust the original video in the adjustment model using at least the size adjustment parameters to obtain the target video. The target video includes the original video content from the original video and the matching video content that matches the original video content. The matching video content is displayed in the target video area. The target video area is used to represent the area offset from the video area boundary of the original video along the direction and distance corresponding to the size adjustment parameters.

[0254] The monitoring unit 1102, the first determining unit 1104, the first input unit 1106, and the first adjusting unit 1108 mentioned above correspond to steps S202 to S208 in Embodiment 1. The four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 1704) and processed by one or more processors (e.g., processors 1702a, 1702b, ..., 1702n). The above units can also be part of a device and run in the computer terminal 170 provided in Embodiment 8.

[0255] According to embodiments of this application, a method for implementing the above is also provided. Figure 3 The video processing method shown is a video processing device.

[0256] Figure 12 This is a schematic diagram of a video processing apparatus according to an embodiment of this application, such as... Figure 12 As shown, the video processing device 1200 may include: an acquisition unit 1202, a second determination unit 1204, a second input unit 1206, a second adjustment unit 1208, and a first playback unit 1210.

[0257] The acquisition unit 1202 is used to acquire the original video to be adjusted from the e-commerce platform.

[0258] The second determining unit 1204 is used to determine the size adjustment parameters of the original video based on the media platform to which the original video is to be pushed, wherein the size adjustment parameters are used to adjust the video size of the original video.

[0259] The second input unit 1206 is used to input the size adjustment parameters and the original video into the adjustment model, wherein the adjustment model is trained based on image content samples.

[0260] The second adjusting unit 1208 is configured to adjust the original video in the adjusting model by using at least the size adjusting parameter to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matched video content matched with the original video content, the matched video content is displayed in a target video area, and the target video area is used to represent a region offset from a video area boundary of the original video in a direction and at a distance corresponding to the size adjusting parameter.

[0261] The first playing unit 1210 is configured to push the target video to the media platform for playing.

[0262] It should be noted that the above-mentioned obtaining unit 1202, the second determining unit 1204, the second input unit 1206, the second adjusting unit 1208 and the first playing unit 1210 correspond to steps S302 to S310 in Embodiment 1, and the five units have the same instances and application scenarios as the corresponding steps, but are not limited to the above-mentioned disclosed content in Embodiment 1. It should be noted that the above-mentioned units can be hardware components or software components stored in a memory (for example, the memory 1704) and processed by one or more processors (for example, the processors 1702a, 1702b, …, 1702n), and the above-mentioned units can also be a part of the device and can run in the computer terminal 170 provided in Embodiment 8.

[0263] According to the embodiments of the present application, a video processing device for implementing the above-mentioned video processing method is also provided. Figure 4

[0264] Figure 13 is a schematic diagram of a video processing device according to an embodiment of the present application. As shown in the figure, the video processing device 1300 can include a first display unit 1302 and a second display unit 1304. Figure 13

[0265] The first display unit 1302 is configured to display, in response to an input instruction acting on an operation interface, an original video to be adjusted and a size adjusting parameter of the original video on the operation interface, wherein the size adjusting parameter is used to adjust the video size of the original video.

[0266] ​​The second display unit 1304 is configured to display a target video on the operation interface in response to a video adjustment instruction acting on the operation interface, wherein the target video is obtained by adjusting the original video by using at least the size adjustment parameter in the adjustment model, the adjustment model is trained based on the image content sample, the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent a region offset from a video area boundary of the original video in a direction corresponding to the size adjustment parameter and by a distance.

[0267] It should be noted that the first display unit 1302 and the second display unit 1304 correspond to steps S402 to S404 in Embodiment 1, and the two units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the units can be hardware components or software components stored in a memory (for example, the memory 1704) and processed by one or more processors (for example, the processors 1702a, 1702b, …, 1702n), or the units can be a part of the computer terminal 170 provided in Embodiment 8 and run in the computer terminal 170.

[0268] According to the embodiments of the present application, a video processing method and a video processing device are also provided. Figure 5 The video processing device is used for implementing the video processing method as shown in the embodiments of the present application.

[0269] Figure 14 FIG. 1 is a schematic diagram of a video processing device according to the embodiments of the present application, as shown in the figure, the video processing device 1400 can include a first display unit 1402, a third determination unit 1404, a third input unit 1406, a third adjustment unit 1408, and a second display unit 1410. Figure 14

[0270] The first display unit 1402 is configured to display an original video to be adjusted on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device.

[0271] The third determination unit 1404 is configured to determine a size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video.

[0272] The third input unit 1406 is configured to input the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on an image content sample.

[0273] ​The third adjusting unit 1408 is configured to adjust the original video by using at least the size adjustment parameter in the adjustment model to obtain a target video, wherein the target video content of the target video comprises original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region offset from a video region boundary of the original video in a direction and at a distance corresponding to the size adjustment parameter.

[0274] The second display unit 1410 is configured to drive the VR device or the AR device to display the target video.

[0275] It should be noted that the first display unit 1402, the third determining unit 1404, the third input unit 1406, the third adjusting unit 1408 and the second display unit 1410 correspond to steps S502 to S510 in Embodiment 1, and the five units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware components or software components stored in a memory (for example, the memory 1704) and processed by one or more processors (for example, the processors 1702a, 1702b, …, 1702n), or the above units can be a part of the device and can run in the computer terminal 170 provided in Embodiment 8.

[0276] In the processing apparatus of the video, if the size of the original video needs to be adjusted, whether the original video to be adjusted is received can be monitored, and if the original video to be transmitted is monitored, the size adjustment parameter for adjusting the video size of the original video can be determined. The adjustment model can be trained in advance according to the image content sample, that is, the ability of the adjustment model to adjust the size of the original video can be trained according to the shape and appearance of some common image content. The size adjustment parameter and the original video are input into the adjustment model, and the adjustment model is used to adjust the video size of the original video according to the size adjustment parameter to obtain a target video. In the target video, not only the original video content from the original video, but also the matching video content matched by the adjustment model according to the original video content can be included, and the matching video content can be displayed in the target video area, that is, the matching video content can be displayed in the area obtained by offsetting from the video area boundary of the original video in the direction and distance corresponding to the size adjustment parameter. Since the embodiment of the present application can accurately adjust the video size of the original video to the video size of the target video required by using the adjustment model according to the size adjustment parameter and the original video, the video content loss caused by simply cropping the video is avoided, so that the purpose of ensuring the integrity of the video content is achieved, and the technical effect of improving the video adjustment effect is realized, and the technical problem of poor video adjustment effect is solved.

[0277] Embodiment 5

[0278] The embodiment of the present application can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Alternatively, in the embodiment, the above-mentioned computer terminal can be replaced by a mobile terminal or other terminal device.

[0279] Alternatively, in the embodiment, the above-mentioned computer terminal can be located in at least one network device of a plurality of network devices of a computer network.

[0280] In the embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the video processing method: monitoring the original video to be adjusted; determining the size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on an image content sample; and adjusting the original video by at least using the size adjustment parameter in the adjustment model to obtain a target video, wherein the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video in a direction and distance corresponding to the size adjustment parameter.

[0281] Optionally, Figure 15 is a structural block diagram of a computer terminal according to an embodiment of the present application. As shown in the figure, the computer terminal A can include one or more (only one is shown in the figure) processors 1502, a memory 1504, and a transmission device 1506. Figure 15

[0282] The memory can be configured to store software programs and modules, such as program instructions / modules corresponding to the video processing method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the video processing method described above. The memory can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal A through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0283] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: generating video generation parameters based on at least the size adjustment parameter and the original video, wherein the video generation parameters are used to make the adjustment model fill in matching video content in the target video region; and inputting the video generation parameters into the adjustment model.

[0284] Optionally, the processor can further execute program codes of the following steps: generating mask information based on the size adjustment parameter and the original video size of the original video, wherein the video generation parameters include the mask information, and the mask information is used to at least determine the target video region by the adjustment model.

[0285] Optionally, the processor can further execute program codes of the following steps: determining the area of the target video region by using the size adjustment parameter and the original video size; and generating the mask information based on the area of the target video region.

[0286] Optionally, the processor can further execute program codes of the following steps: extracting first video information from the original video when the size adjustment parameter is to enlarge the video size of the original video, wherein the video generation parameters include the first video information, and the first video information is used to make the adjustment model generate matching video content for extending the original video in the target video region.

[0287] ​Optionally, the processor can further execute program codes of the following steps: extracting video features from the original video frames of the original video, wherein the first video information comprises the video features, and the video features are used to represent objects in the original video content and / or scenes to which the original video content belongs.

[0288] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: encoding the original video content of the original video frame and the original video region formed by the video region boundary to obtain video features.

[0289] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: performing color detection on two adjacent original video frames in the original video to obtain color change information between the two adjacent original video frames; if the color change information is less than a color change threshold, determining that the two adjacent original video frames belong to the same shot segment; and extracting video features from the original video frames of the original video, including: extracting video features from the shot segment.

[0290] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: performing edge detection on the original video to obtain edge information, wherein the first video information comprises the edge information.

[0291] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: in the case that the size adjustment parameter is to reduce the video size of the original video, determining the second video information based on the size adjustment parameter and the original video size of the original video, wherein the video generation parameter comprises the second video information, and the second video information is used to enable the adjustment model to generate matching video content for obscuring the original video in the target video region.

[0292] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: determining a noise sampling strategy corresponding to the adjustment model, wherein the noise sampling strategy is used to represent a sampling rule of noise samples, and the noise samples are used to train the adjustment model with image content samples; and sampling the original video by using the noise sampling strategy to obtain noise information, wherein the video generation parameter comprises the noise information.

[0293] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: converting the video generation parameter into to-be-decoded information of the matching video content in the adjustment model; decoding the to-be-decoded information to obtain pixels of the matching video content; and generating the target video by combining the pixels of the matching video content with pixels of the original video content.

[0294] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: performing super-resolution processing on the original video, wherein the video size of the super-resolution original video is higher than the video size of the original video before super-resolution; and determining the video size of the super-resolution original video as the video size of the original video.

[0295] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: the original video includes original video frames, determining the size adjustment parameter of the original video frame; inputting the size adjustment parameter and the original video frame into the adjustment model; in the adjustment model, at least using the size adjustment parameter to adjust the original video frame to obtain a target video frame, wherein the target video includes the target video frame; in response to the original video frame existing next original video frame, determining the next original video frame of the original video frame as the original video frame, returning to execute from the following steps; determining the size adjustment parameter of the original video frame until the original video frame does not exist next original video frame.

[0296] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: in response to the original video frame existing next original video frame, determining the next original video frame of the original video frame and the target video frame as the original video frame.

[0297] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: determining an image content sample from a video sample; training a deep learning model using attribute information of different objects in the image content sample to obtain a diffusion model; and determining the diffusion model as the adjustment model.

[0298] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining an original video to be adjusted from an e-commerce platform; determining a size adjustment parameter of the original video based on a media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on an image content sample; in the adjustment model, at least using the size adjustment parameter to adjust the original video to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matching the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video in a direction and distance corresponding to the size adjustment parameter; and pushing the target video to the media platform for playing.

[0299] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: in response to an input instruction acting on the operation interface, displaying the original video to be adjusted and the size adjustment parameter of the original video on the operation interface, wherein the size adjustment parameter is used to adjust the video size of the original video; in response to a video adjustment instruction acting on the operation interface, displaying the target video on the operation interface, wherein the target video is obtained by adjusting the original video by at least using the size adjustment parameter in the adjustment model, the adjustment model is obtained based on image content samples, the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in the target video area, and the target video area is used to represent an area offset from the video area boundary of the original video in the direction and distance corresponding to the size adjustment parameter.

[0300] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: displaying the original video to be adjusted on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; determining a size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is obtained based on image content samples; adjusting the original video by at least using the size adjustment parameter in the adjustment model to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video in the direction and distance corresponding to the size adjustment parameter; and driving the VR device or the AR device to display the target video.

[0301] The embodiment of the application provides a video processing method. In the embodiment of the application, if the size of an original video needs to be adjusted, whether the original video to be adjusted is received is monitored, and if the original video to be transmitted is monitored, a size adjustment parameter that needs to be used to adjust the size of the original video is determined. An adjustment model can be trained in advance according to image content samples, that is, the ability of the adjustment model to adjust the size of the original video can be trained according to the shape and appearance of some common image content. The size adjustment parameter and the original video are input into the adjustment model, the adjustment model is used to adjust the size of the original video according to the size adjustment parameter, and a target video is obtained. In the target video, not only original video content from the original video, but also matching video content matched by the adjustment model according to the original video content can be included, and the matching video content can be displayed in a target video area, that is, the matching video content can be displayed in a region obtained by offsetting from the video area boundary of the original video along a direction and a distance corresponding to the size adjustment parameter. Since the embodiment of the application can accurately adjust the size of the original video to the size of the target video required by using the adjustment model according to the size adjustment parameter and the original video, the loss of video content caused by simply cutting the video is avoided, the purpose of ensuring the integrity of the video content is achieved, and the technical effect of improving the video adjustment effect is achieved, and the technical problem of poor video adjustment effect is solved.

[0302] Those skilled in the art can understand that Figure 15 The structure shown is only schematic, and the computer terminal A can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a mobile Internet device (Mobile Internet Devices, referred to as MID), a PAD, and the like. Figure 15 This does not limit the structure of the computer terminal A. For example, the computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 15 For example, the computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 15 For example, the computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.

[0303] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (Read-Only Memory, referred to as ROM), a random access memory (Random Access Memory, referred to as RAM), a magnetic disk or an optical disk, etc.

[0304] Embodiment 6

[0305] The embodiments of the present application further provide a computer readable storage medium. Optionally, in the embodiments, the computer readable storage medium can be used to store program codes executed by the video processing method provided in the first embodiment.

[0306] Optionally, in the embodiments, the computer readable storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0307] Optionally, in the embodiments, the computer readable storage medium is configured to store program codes for monitoring the original video to be adjusted; determining a size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; and adjusting the original video in the adjustment model by using at least the size adjustment parameter to obtain a target video, wherein the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region offset from the video region boundary of the original video along a direction and a distance corresponding to the size adjustment parameter.

[0308] Optionally, the computer readable storage medium can further store program codes for generating a video generation parameter based on at least the size adjustment parameter and the original video, wherein the video generation parameter is used to enable the adjustment model to fill the matching video content in the target video region; and inputting the video generation parameter into the adjustment model.

[0309] Optionally, the computer readable storage medium can further store program codes for generating mask information based on the size adjustment parameter and the original video size of the original video, wherein the video generation parameter includes the mask information, and the mask information is used to enable the adjustment model to determine the target video region.

[0310] Optionally, the computer readable storage medium can further store program codes for determining the area of the target video region by using the size adjustment parameter and the original video size; and generating the mask information based on the area of the target video region.

[0311] Optionally, the computer readable storage medium can further execute program codes of the following steps: extracting the first video information from the original video, in the case that the size adjustment parameter is to enlarge the video size of the original video, wherein the video generation parameter comprises the first video information, and the first video information is used to make the adjustment model generate the matching video content for extending the original video in the target video region.

[0312] Optionally, the computer readable storage medium can further execute program codes of the following steps: extracting the video features from the original video frames of the original video, wherein the first video information comprises the video features, and the video features are used to represent the objects in the original video content and / or the scenes to which the original video content belongs.

[0313] Optionally, the computer readable storage medium can further execute program codes of the following steps: encoding the original video content of the original video frame and the original video region formed by the video region boundary to obtain the video features.

[0314] Optionally, the computer readable storage medium can further execute program codes of the following steps: performing color detection on two adjacent original video frames in the original video to obtain color change information between the two adjacent original video frames; determining that the two adjacent original video frames belong to the same shot segment if the color change information is less than a color change threshold; and extracting the video features from the original video frames of the original video, comprising: extracting the video features from the shot segment.

[0315] Optionally, the computer readable storage medium can further execute program codes of the following steps: performing edge detection on the original video to obtain edge information, wherein the first video information comprises the edge information.

[0316] Optionally, the computer readable storage medium can further execute program codes of the following steps: determining the second video information based on the size adjustment parameter and the original video size of the original video, in the case that the size adjustment parameter is to reduce the video size of the original video, wherein the video generation parameter comprises the second video information, and the second video information is used to make the adjustment model generate the matching video content for occluding the original video in the target video region.

[0317] Optionally, the computer readable storage medium can further execute program codes of the following steps: determining a noise sampling strategy corresponding to the adjustment model, wherein the noise sampling strategy is used to represent a sampling rule of noise samples, and the noise samples are used to train the adjustment model with the image content samples; and sampling the original video by using the noise sampling strategy to obtain noise information, wherein the video generation parameter comprises the noise information.

[0318] Optionally, the computer readable storage medium can further execute program codes of the following steps: in the adjustment model, converting the video generation parameter into to-be-decoded information matching the video content; decoding the to-be-decoded information to obtain pixels of the video content matching the video content; and generating the target video by combining the pixels of the video content matching the video content with the pixels of the original video content.

[0319] Optionally, the computer readable storage medium can further execute program codes of the following steps: performing super-resolution processing on the original video, wherein the video size of the original video after the super-resolution processing is higher than the video size of the original video before the super-resolution processing; and determining the video size of the original video after the super-resolution processing as the video size of the original video.

[0320] Optionally, the computer readable storage medium can further execute program codes of the following steps: the original video includes an original video frame, determining a size adjustment parameter of the original video frame; inputting the size adjustment parameter and the original video frame into the adjustment model; in the adjustment model, adjusting the original video frame by at least using the size adjustment parameter to obtain a target video frame, wherein the target video includes the target video frame; in response to the original video frame having a next original video frame, determining the next original video frame of the original video frame as the original video frame, and returning to execute from the following steps: determining the size adjustment parameter of the original video frame until the original video frame does not have a next original video frame.

[0321] Optionally, the computer readable storage medium can further execute program codes of the following steps: in response to the original video frame having a next original video frame, determining the next original video frame of the original video frame and the target video frame as the original video frame.

[0322] Optionally, the computer readable storage medium can further execute program codes of the following steps: determining an image content sample from the video sample; training the deep learning model by using attribute information of different objects in the image content sample to obtain a diffusion model; and determining the diffusion model as the adjustment model.

[0323] Optionally, the computer readable storage medium can further execute program codes of the following steps: obtaining the original video to be adjusted from the e-commerce platform; determining the size adjustment parameter of the original video based on the media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into the adjustment model, wherein the adjustment model is trained based on image content samples; adjusting the original video by at least using the size adjustment parameter in the adjustment model to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video in a direction and distance corresponding to the size adjustment parameter; and pushing the target video to the media platform for playing.

[0324] Optionally, the computer readable storage medium can further execute program codes of the following steps: displaying the original video to be adjusted and the size adjustment parameter of the original video on the operation interface in response to an input instruction acting on the operation interface, wherein the size adjustment parameter is used to adjust the video size of the original video; and displaying the target video on the operation interface in response to a video adjustment instruction acting on the operation interface, wherein the target video is obtained by adjusting the original video by at least using the size adjustment parameter in the adjustment model, the adjustment model is trained based on image content samples, the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video in a direction and distance corresponding to the size adjustment parameter.

[0325] Optionally, the computer readable storage medium can further execute program codes of the following steps: displaying the original video to be adjusted on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; determining the size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into the adjustment model, wherein the adjustment model is trained based on image content samples; adjusting the original video by at least using the size adjustment parameter in the adjustment model to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video in a direction and distance corresponding to the size adjustment parameter; and driving the VR device or the AR device to display the target video.

[0326] Embodiment 7

[0327] Embodiments of this application may provide an electronic device that may include a memory and a processor.

[0328] Figure 16 This is a block diagram of an electronic device for a video processing method according to an embodiment of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0329] like Figure 16 As shown, device 1600 includes a computing unit 1601, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1602 or a computer program loaded into random access memory (RAM) 1603 from storage unit 1608. RAM 1603 may also store various programs and data required for the operation of device 1600. The computing unit 1601, ROM 1602, and RAM 1603 are interconnected via bus 1604. Input / output (I / O) interface 1605 is also connected to bus 1604.

[0330] Multiple components in device 1600 are connected to I / O interface 1605, including: input unit 1606, such as keyboard, mouse, etc.; output unit 1604, such as various types of monitors, speakers, etc.; storage unit 1608, such as disk, optical disk, etc.; and communication unit 1609, such as network card, modem, wireless transceiver, etc. Communication unit 1609 allows device 1600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0331] The computing unit 1601 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 1601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1601 performs various methods and processes described above, such as the processing method of a video. For example, in some embodiments, the processing method of a video can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1600 via the ROM 1602 and / or the communication unit 1609. When the computer program is loaded into the RAM 1603 and executed by the computing unit 1601, one or more steps of the processing method of a video described above can be performed. Alternatively, in other embodiments, the computing unit 1601 can be configured to perform the processing method of a video by any other appropriate means (e.g., by means of firmware).

[0332] Embodiment 8

[0333] The embodiments of the present application further provide a computer program product. Optionally, in the embodiments, the computer program product can include a computer program, and the computer program, when executed by a processor, implements the processing method of a video of the embodiments of the present application.

[0334] According to the embodiments of the present application, a processing method of a video is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0335] The method embodiments provided by the embodiments of the present application can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 17 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a processing method of a video according to the embodiments of the present application, like Figure 17As shown, the computer terminal 170 (or mobile device) can include one or more (shown as 1702a, 1702b, …, 1702n) processors 1702 (the processor 1702 can include, but not limited to, a microcontroller unit (MCU) or a field programmable gate array (FPGA) processing device, etc.), a memory 1704 for storing data, and a transmission device 1706 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports in the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 17 The structure shown is only a schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 170 can also include more or less components than those shown in the figure, or have a different configuration than that shown in the figure. Figure 17 For example, the computer terminal 170 can also include more or less components than those shown in the figure, or have a different configuration than that shown in the figure. Figure 17 For example, the computer terminal 170 can also include more or less components than those shown in the figure, or have a different configuration than that shown in the figure.

[0336] Figure 17 The hardware structure diagram shown not only can be used as an exemplary block diagram of the above-mentioned computer terminal 170 (or mobile device), but also can be used as an exemplary block diagram of the above-mentioned server, in an alternative embodiment, Figure 18 The above-mentioned computer terminal 170 (or mobile device) is shown as an embodiment of a computing node in a computing environment 1801. Figure 17

[0337] Figure 18 The structure block diagram of the computing environment of a video processing method according to an embodiment of the present application is shown as Figure 18 As shown, the computing environment 1801 includes a plurality of (shown as 1810-1, 1810-2, …) computing nodes (such as servers) running on a distributed network. The computing nodes all contain local processing and memory resources, and the end user 1802 can remotely run an application program or store data in the computing environment 1801. The application program can be provided as a plurality of services 1820-1, 1820-2, 1820-3 and 1820-4 in the computing environment 1601, representing services “F”, “G”, “I” and “H” respectively.

[0338] ​End users 1802 can provide and access services through a web browser or other software application on a client, in some embodiments, provisioning and / or requests of end users 1802 can be provided to ingress gateway 1830. Ingress gateway 1830 can include a corresponding proxy to handle provisioning and / or requests for services (one or more services provided in computing environment 1801).

[0339] Services are provided or deployed according to various virtualization technologies supported by computing environment 1801. In some embodiments, services can be provided according to virtualization based on virtual machines (VMs), virtualization based on containers, and / or the like. Virtualization based on virtual machines can be to emulate a real computer by initializing a virtual machine to execute programs and applications without directly accessing any actual hardware resources. While a virtual machine virtualization machine, according to virtualization based on containers, a container can be launched to virtualize an entire operating system so that multiple workloads can run on a single operating system instance.

[0340] In one embodiment of virtualization based on containers, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, as shown in Figure 18 Service 1820-2 can be equipped with one or more Pods 1840-1, 1840-2,..., 1840-N (collectively, Pods) as shown. A Pod can include a proxy 1845 and one or more containers 1842-1, 1842-2,..., 1842-M (collectively, containers). The one or more containers in a Pod handle requests related to one or more corresponding functions of the service, and proxy 1845 generally controls network functions related to the service, such as routing, load balancing, and the like. Other services can also be equipped with Pods similar to the Pods.

[0341] In operation, executing a user request from end user 1802 can require invoking one or more services in computing environment 1801, and executing one or more functions of a service can require invoking one or more functions of another service. As shown in Figure 18 Service “F” 1820-1 receives a user request from end user 1802 from ingress gateway 1830, service “F” 1820-1 can invoke service “G” 1820-2, and service “G” 1820-2 can request service “I” 1820-3 to execute one or more functions.

[0342] The computing environment described above can be a cloud computing environment, with the allocation of resources managed by a cloud service provider, allowing the development of functionality without the need to consider implementing, tuning or scaling servers. The computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be split into a set of functions that can automatically scale independently, rather than scaling a single hardware device to handle potential loads.

[0343] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an Application Specific Integrated Circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0344] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, fully on a machine as a standalone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0345] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a special-purpose standard product (ASSP), a system on a chip that includes one or both of a general- purpose processing portion and a special- purpose circuit portion, a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations of each to implement the various embodiments. These various embodiments can each be implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0346] Program code to implement methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, implements the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0347] In the context of the present application, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0348] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a Cathode Ray Tube (CRT) or a Liquid Crystal Display (LCD) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0349] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0350] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0351] It should be noted that the above-mentioned sequence numbers of the embodiments of the present application are only for description, not representing the advantages and disadvantages of the embodiments.

[0352] In the above-described embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0353] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.

[0354] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e. they can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0355] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0356] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes or the whole or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory, random access memory, mobile hard disk, magnetic disk or optical disk, and various program code storage media.

[0357] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A method of processing a video, characterized by, The method comprises: monitoring an original video to be adjusted; determining a size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust a video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; in the adjustment model, adjusting the original video at least by using the size adjustment parameter to obtain a target video, wherein the target video comprises original video content from the original video and matching video content matching the original video content, the matching video content being displayed in a target video region, the target video region being used to represent a region offset from a video region boundary of the original video in a direction and distance corresponding to the size adjustment parameter; wherein inputting the size adjustment parameter and the original video into the adjustment model comprises: generating video generation parameters based at least on the size adjustment parameter and the original video, wherein the video generation parameters comprise noise information, mask information, video features, and edge information, the noise information being obtained by randomly sampling noise on the encoded original video, the mask information being determined from the original video based on the size adjustment parameter to obtain the target video region, the video features being used to represent objects in the original video content and / or scenes to which the original video content belongs, and the edge information being used to represent texture information of the objects; and inputting the video generation parameters into the adjustment model; in the adjustment model, adjusting the original video at least by using the size adjustment parameter to obtain a target video, comprising: performing a noise reduction operation on the video generation parameters in the adjustment model to obtain to-be-decoded information; and decoding the to-be-decoded information to obtain the target video.

2. The method of claim 1, wherein, The video generation parameters are used to cause the adjustment model to fill the matching video content in the target video region.

3. The method of claim 2, wherein, Generating video generation parameters based at least on the size adjustment parameter and the original video comprises: generating the mask information based on the size adjustment parameter and an original video size of the original video, wherein the mask information is used to cause the adjustment model to at least determine the target video region.

4. The method of claim 3, wherein, Generating mask information based on the size adjustment parameter and the original video size of the original video comprises: determining an area of the target video region by using the size adjustment parameter and the original video size; and generating the mask information based on the area of the target video region.

5. The method of claim 2, wherein, The method further comprises: in a case where the size adjustment parameter is to enlarge the video size of the original video, extracting first video information from the original video, wherein the video generation parameters comprise the first video information, and the first video information is used to cause the adjustment model to generate the matching video content for extending the original video in the target video region.

6. The method of claim 5, wherein, Extracting first video information from the original video comprises: extracting the video feature from an original video frame of the original video, wherein the first video information comprises the video feature, and the video feature is used to represent an object in the original video content and / or a scene to which the original video content belongs.

7. The method of claim 6, wherein, extracting the video feature from an original video frame of the original video, comprising: encoding the original video content of the original video frame and an original video region formed by the video region boundary to obtain the video feature.

8. The method of claim 6, wherein, The method further comprises: performing color detection on two adjacent original video frames in the original video to obtain color change information between the two adjacent original video frames; if the color change information is less than a color change threshold, determining that the two adjacent original video frames belong to the same shot segment; extracting the video feature from an original video frame of the original video, comprising: extracting the video feature from the shot segment.

9. The method of claim 5, wherein, extracting first video information from the original video, comprising: performing edge detection on the original video to obtain the edge information, wherein the first video information comprises the edge information.

10. The method of claim 2, wherein, The method further comprises: in a case where the size adjustment parameter is to reduce the video size of the original video, determining second video information based on the size adjustment parameter and the original video size of the original video, wherein the video generation parameter comprises the second video information, and the second video information is used to enable the adjustment model to generate the matching video content used to occlude the original video in the target video region.

11. The method of claim 2, wherein, The method further comprises: determining a noise sampling strategy corresponding to the adjustment model, wherein the noise sampling strategy is used to represent a sampling rule of a noise sample, and the noise sample is used to train the adjustment model with the image content sample; sampling the original video using the noise sampling strategy to obtain noise information, wherein the video generation parameter comprises the noise information.

12. The method of claim 2, wherein, in the adjustment model, adjusting the original video using the size adjustment parameter to obtain a target video, comprising: in the adjustment model, converting the video generation parameter into the to-be-decoded information of the matching video content; generating the target video based on the to-be-decoded information, comprising: decoding the to-be-decoded information to obtain pixels of the matching video content; generating the target video by combining the pixels of the matching video content with pixels of the original video content.

13. The method of claim 1, wherein, The method further comprises: performing super-resolution processing on the original video, wherein the video size of the original video after super-resolution processing is higher than the video size of the original video before super-resolution processing; determining the video size of the original video after super-resolution processing as the video size of the original video.

14. The method of claim 1, wherein, The original video comprises an original video frame, determining the size adjustment parameter of the original video comprises determining the size adjustment parameter of the original video frame; inputting the size adjustment parameter and the original video into an adjustment model comprises inputting the size adjustment parameter and the original video frame into the adjustment model; In the adjustment model, the original video is adjusted at least by using the size adjustment parameter to obtain a target video, comprising: in the adjustment model, the original video frame is adjusted at least by using the size adjustment parameter to obtain a target video frame, wherein the target video comprises the target video frame; The method further comprises: in response to the original video frame existing next original video frame, determining the next original video frame of the original video frame as the original video frame, and returning to execute from the following steps; Determine the size adjustment parameter of the original video frame until the original video frame does not exist the next original video frame.

15. The method of claim 14, wherein, In response to the original video frame existing next original video frame, determining the next original video frame of the original video frame as the original video frame, comprising: In response to the original video frame existing next original video frame, determining the next original video frame of the original video frame and the target video frame as the original video frame.

16. The method according to any one of claims 1 to 15, characterized in that, The method further comprises: Determine the image content sample from the video sample; Train a deep learning model using the attribute information of different objects in the image content sample to obtain a diffusion model; Determine the diffusion model as the adjustment model.

17. A method of processing a video, the method comprising: Comprise: Obtain the original video to be adjusted from the e-commerce platform; Determine the size adjustment parameter of the original video based on the media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust the video size of the original video; Input the size adjustment parameter and the original video into the adjustment model, wherein the adjustment model is trained based on the image content sample; In the adjustment model, the original video is adjusted at least by using the size adjustment parameter to obtain a target video, wherein the target video content of the target video comprises original video content from the original video and matching video content matched with the original video content, the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video along the direction and distance corresponding to the size adjustment parameter; Push the target video to the media platform for playing; Wherein, inputting the size adjustment parameter and the original video into the adjustment model comprises: generating video generation parameters based at least on the size adjustment parameter and the original video, wherein the video generation parameters comprise noise information, mask information, video features and edge information, the noise information is obtained by randomly sampling noise on the encoded original video, the mask information is obtained by determining the target video area from the original video based on the size adjustment parameter, the video features are used to represent the objects in the original video content and / or the scenes to which the original video content belongs, and the edge information is used to represent the texture information of the objects; input the video generation parameters into the adjustment model; In the adjustment model, the original video is adjusted by at least using the size adjustment parameter to obtain a target video, comprising: in the adjustment model, performing a noise reduction operation on the video generation parameter to obtain to-be-decoded information; decoding the to-be-decoded information to obtain the target video.

18. A method of processing a video, the method comprising: Comprise: In response to an input instruction acting on an operation interface, display an original video to be adjusted and a size adjustment parameter of the original video on the operation interface, wherein the size adjustment parameter is used to adjust the video size of the original video; In response to a video adjustment instruction acting on the operation interface, display a target video on the operation interface, wherein the target video is obtained by adjusting the original video by at least using the size adjustment parameter in an adjustment model, the adjustment model is trained based on image content samples, the target video content of the target video includes original video content from the original video and matching video content matching the original video content, the matching video content is displayed in a target video area, the target video area is used to represent an area offset from the video area boundary of the original video along the direction and distance corresponding to the size adjustment parameter, the target video is obtained by decoding to-be-decoded information, the to-be-decoded information is obtained by performing a noise reduction operation on the video generation parameter in the adjustment model, the video generation parameter includes noise information, mask information, video features and edge information, the noise information is obtained by randomly sampling noise on the encoded original video, the mask information is obtained from the original video based on the size adjustment parameter to determine the target video area, the video features are used to represent objects in the original video content and / or scenes to which the original video content belongs, and the edge information is used to represent texture information of the objects; the video generation parameter is generated based on at least the size adjustment parameter and the original video.

19. A method of processing a video, the method comprising: Comprise: Display an original video to be adjusted on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; Determine a size adjustment parameter of the original video, wherein the size adjustment parameter is used to adjust the video size of the original video; Input the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; In the adjustment model, the original video is adjusted by at least using the size adjustment parameter to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matching the original video content, and the matching video content is displayed in a target video area, and the target video area is used to represent an area offset from the video area boundary of the original video along the direction and distance corresponding to the size adjustment parameter; Drive the VR device or the AR device to display the target video; The size adjustment parameter and the original video are input into the adjustment model, including: generating video generation parameters based on at least the size adjustment parameter and the original video, wherein the video generation parameters include noise information, mask information, video features, and edge information, the noise information is obtained by randomly sampling noise on the encoded original video, the mask information is obtained by determining the target video region from the original video based on the size adjustment parameter, the video features are used to represent objects in the original video content and / or scenes to which the original video content belongs, and the edge information is used to represent texture information of the objects; and the video generation parameters are input into the adjustment model. In the adjustment model, the original video is adjusted based on at least the size adjustment parameter to obtain a target video, including: performing a noise reduction operation on the video generation parameters in the adjustment model to obtain to-be-decoded information; and decoding the to-be-decoded information to obtain the target video.

20. A system for processing video, characterized by Including: An e-commerce platform for making an original video to be adjusted; A video adjustment end for determining a size adjustment parameter of the original video based on a media platform to which the original video is to be pushed, wherein the size adjustment parameter is used to adjust the video size of the original video; inputting the size adjustment parameter and the original video into an adjustment model, wherein the adjustment model is trained based on image content samples; and adjusting the original video based on at least the size adjustment parameter in the adjustment model to obtain a target video, wherein the target video content of the target video includes original video content from the original video and matching video content matching the original video content, the matching video content is displayed in a target video region, and the target video region is used to represent a region offset from the video region boundary of the original video in a direction and distance corresponding to the size adjustment parameter. A media platform for playing the target video. The video adjustment end is configured to perform the following steps to input the size adjustment parameter and the original video into the adjustment model: generate video generation parameters based on at least the size adjustment parameter and the original video, wherein the video generation parameters include noise information, mask information, video features, and edge information, the noise information is obtained by randomly sampling noise on the encoded original video, the mask information is obtained by determining the target video region from the original video based on the size adjustment parameter, the video features are used to represent objects in the original video content and / or scenes to which the original video content belongs, and the edge information is used to represent texture information of the objects; and input the video generation parameters into the adjustment model. The video adjustment end is further configured to perform the following steps to generate the target video: performing a noise reduction operation on the video generation parameters in the adjustment model to obtain to-be-decoded information; and decoding the to-be-decoded information to obtain the target video.

21. An electronic device, comprising: Including: a memory storing an executable program; a processor configured to execute the program, wherein the program, when executed, performs the method of any one of claims 1 to 19.

22. A computer-readable storage medium, characterized in that, a computer readable storage medium storing a program, wherein the program, when executed, controls a device in which the storage medium is located to perform the method of any one of claims 1 to 19.

23. A computer program product, characterised in that, computer instructions that, when executed by a processor, perform the method of any one of claims 1 to 19.

Citation Information

Patent Citations

  • Image generation method and device and storage medium

    CN114299175A

  • Text image generation method and diffusion generation model training method

    CN116797868A

  • Video generation method and device, electronic equipment and storage medium

    CN117241100A