Video processing method, automatic question answering method, computing device, computer readable storage medium and computer program product

By obtaining the motion feature information of the initial video frame and its surrounding frames, the target video frame is generated, and the flickering problem caused by inconsistency between video frames is solved, and a smooth video processing effect is achieved.

CN120455730APending Publication Date: 2025-08-08HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410172118.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-06
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

When applying image processing technology to video processing, the prior art faces the flickering problem caused by inconsistency between video frames. Especially in video style migration tasks, existing methods are difficult to effectively maintain video consistency.

Method used

By obtaining the motion feature information of the initial video frame and its surrounding video frames, the target video frame is generated using the preset motion feature generation strategy to ensure the consistency between the video frames and avoid flickering.

Benefits of technology

It realizes rendering of incoherent reference videos into smooth target videos, solving the problem of inconsistency between video frames and improving the effect of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455730A_ABST
    Figure CN120455730A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video processing method, an automatic question and answer method, computing equipment, a computer readable storage medium and a computer program product. The video processing method comprises the following steps: acquiring a first initial video frame in an initial video and at least one second initial video frame corresponding to the first initial video frame; determining a reference video corresponding to the initial video, and determining a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame; generating motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy; and processing the first reference video frame according to the motion feature information to obtain a target video frame corresponding to the first reference video frame, and generating a target video according to at least one target video frame. And the first video frames are processed according to the motion feature information of each first video frame in the reference video, so that the reference video is rendered into a smooth target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a video processing method. Background Art

[0002] The field of image processing has experienced rapid development in recent years. In particular, the emergence of diffusion models trained on massive datasets has marked a major revolution in image synthesis technology. These diffusion models not only surpass generative adversarial networks in quality but also achieve remarkable results in a variety of fields, including image style transfer, super-resolution, and image editing.

[0003] However, applying these advanced image processing techniques to video processing presents the challenge of maintaining video consistency. This is particularly true in tasks like video style transfer, where each frame is processed independently, often leading to discontinuities between frames and the appearance of flickering in the resulting video. Therefore, a video processing method that effectively addresses this video consistency issue is urgently needed. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a video processing method, an automatic question-answering method, and a video processing method applied to a cloud server. One or more embodiments of this specification also relate to a video processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a video processing method is provided, including:

[0006] Acquire a first initial video frame in an initial video, and at least one second initial video frame corresponding to the first initial video frame;

[0007] Determining a reference video corresponding to the initial video, and determining a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video;

[0008] generating motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy;

[0009] The first reference video frame is processed according to each motion feature information to obtain a target video frame corresponding to the first reference video frame, and a target video is generated according to at least one target video frame.

[0010] According to a second aspect of the embodiments of this specification, a video processing method applied to a cloud server is provided, including:

[0011] A video processing instruction sent by a receiving device, wherein the video processing instruction carries an initial video and a reference video corresponding to the initial video;

[0012] Acquire a first initial video frame in the initial video, and at least one second initial video frame corresponding to the first initial video frame;

[0013] Obtaining a reference video corresponding to the initial video, and determining a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video;

[0014] generating motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy;

[0015] The first reference video frame is processed according to each motion feature information to obtain a target video frame corresponding to the first reference video frame, and a target video is generated according to at least one target video frame, and the target video is returned to the terminal side device.

[0016] According to a third aspect of the embodiments of this specification, there is provided an automatic question-answering method, comprising:

[0017] receiving a video modification instruction, wherein the video modification instruction carries an initial video and an initial question text;

[0018] Parsing the video modification instruction to determine an initial question text, and generating an initial video modification prompt text according to the initial question text;

[0019] Inputting the initial video and the initial video modification prompt text into a video modification model to obtain a reference video generated by the video modification model;

[0020] Acquire motion feature information between a target first video frame and each second video frame in the reference video according to the initial video, wherein the target first video frame is any video frame in the reference video, and the second video frame is at least one video frame corresponding to the target first video frame in the reference video;

[0021] Each target first video frame in the reference video is adjusted according to motion feature information corresponding to each target first video frame in the reference video to generate a target video, and video modification answer information is generated according to the target video and the initial question text.

[0022] According to a fourth aspect of the embodiments of this specification, a computing device is provided, including:

[0023] memory and processor;

[0024] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned video processing method, an automatic question-answering method, or a video processing method applied to a cloud server are implemented.

[0025] According to the fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned video processing method, an automatic question-answering method, or a video processing method applied to a cloud server.

[0026] According to the sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned video processing method, an automatic question-answering method, or a video processing method applied to a cloud server.

[0027] An embodiment of the present specification implements obtaining a first initial video frame in an initial video, and at least one second initial video frame corresponding to the first initial video frame; determining a reference video corresponding to the initial video, and determining a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video; generating motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy; processing the first reference video frame according to each motion feature information to obtain a target video frame corresponding to the first reference video frame, and generating a target video based on at least one target video frame.

[0028] By applying the scheme of the embodiments of this specification, the motion feature information between each frame in the reference video determined by the motion feature information between video frames in the initial video can effectively reflect the motion feature information between each frame in the reference video. Subsequently, the motion feature information between each frame in the reference video determined by the above-mentioned calculation of the feature information in the initial video can determine the motion relationship that should exist between each frame in the reference video through the smooth initial video, and then the non-smooth reference video can be processed through the motion feature information between each frame in the smooth video, thereby avoiding the video flickering problem caused by the discontinuity between frames, and thus realizing the rendering of the flickering reference video into a smooth target video. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flow chart of a video processing method provided by one embodiment of this specification;

[0030] Figure 2a It is a schematic diagram of a mixing table of a video frame and a left video frame process in a video processing method provided by one embodiment of the present specification;

[0031] Figure 2b It is a schematic diagram of a mixing table of a video frame and a right video frame process in a video processing method provided by one embodiment of the present specification;

[0032] Figure 3a Schematic diagram of a mixed table of a video frame and a left video frame result in a video processing method provided by an embodiment of the present specification;

[0033] Figure 3b Schematic diagram of a mixed table of a video frame and a right video frame result in a video processing method provided by one embodiment of this specification;

[0034] Figure 4 This is a flow chart of a method for processing video on a cloud server provided by one embodiment of this specification;

[0035] Figure 5 This is an architecture diagram of a video processing system provided by one embodiment of this specification;

[0036] Figure 6 This is a flow chart of an automatic question-answering method provided by one embodiment of this specification;

[0037] Figure 7 This is a process flow chart of a video optimization method provided by one embodiment of this specification;

[0038] Figure 8 This is a schematic structural diagram of a video processing device provided by one embodiment of this specification;

[0039] Figure 9 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0040] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0041] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0042] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0043] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0044] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. It is pre-trained on a large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.

[0045] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0046] First, the terms involved in one or more embodiments of this specification are explained.

[0047] Match backpropagation exploits the spatial continuity between image patches. Once the algorithm finds a matching region for one patch, it assumes that adjacent patches likely have a good match at a similar location in the target image. Therefore, this matching information is "backpropagated" to the adjacent patches as an initial guess for their matching regions. This significantly speeds up the matching process by reducing the scope and number of searches.

[0048] Random Search: Once a preliminary matching region is determined for an image block, the algorithm performs a random search around that matching region to find a better match. This random search is performed within a shrinking window, which allows for a more efficient and progressively refined match. This strategy assumes that matching regions are typically located near the current matching region.

[0049] Diffusion models, trained on massive datasets, have revolutionized image synthesis. They have proven to comprehensively outperform generative adversarial networks, even reaching a level of creative prowess comparable to that of human artists. However, extending these image processing techniques to video processing presents the challenge of maintaining smooth video. In particular, in video style transfer, since each frame in a video is processed independently, directly applying image processing methods often results in discontinuities, causing noticeable flickering in the generated video.

[0050] To address the aforementioned video flicker problem, many methods have been proposed to enhance the consistency of generated videos and avoid the video artifact problem. For example, full-video rendering methods involve processing each frame through a deep learning model. To enhance frame consistency, mechanisms specifically designed for video processing are employed. Key frame sequence rendering methods use a deep learning model to process a sequence of key video frames while simultaneously generating the remaining frames using interpolation. Single-frame rendering methods process only a single frame through a deep learning model, then render the complete video based on motion information extracted from the original video.

[0051] However, these methods still have drawbacks. In full-video rendering, existing methods struggle to ensure video continuity, and noticeable flickering can still occur in some cases. In keyframe sequence rendering, the content of adjacent keyframes remains inconsistent, resulting in abrupt transitions in the rendered video. In single-frame rendering, due to the limited information available in a single frame, frame tearing often occurs in high-speed motion videos.

[0052] In this specification, a video processing method, an automatic question-answering method, and a video processing method applied to a cloud server are provided. This specification also involves a video processing device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0053] Generally, the application scenario of this video processing method is that the user uploads the original video to the style conversion service. The provider of the style conversion service processes the original video uploaded by the user according to the deep learning model to generate a reference video that is similar to the original video but has a different style. Since the frames in the reference video directly output by the deep learning model may not be coherent, the reference video directly output by the deep learning model may have a flickering problem. Therefore, the video processing method provided in this specification is applied to process the video directly output by the deep learning model to ensure the coherence between the frames in the video, so as to avoid the flickering problem in the video after style conversion.

[0054] See also Figure 1 , Figure 1 A flowchart of a video processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0055] Step 102: Acquire a first initial video frame in an initial video, and at least one second initial video frame corresponding to the first initial video frame.

[0056] In practical applications, the initial video is a video before style conversion, the first initial video frame is any frame in the initial video, and the second initial video frame is a video frame near the first initial video frame.

[0057] Specifically, the initial video can be understood as a video uploaded by the user. The user converts the uploaded initial video into a target video of another style according to the style conversion service, for example, converting an original real landscape video into an anime-style landscape video.

[0058] The target video is a video after the initial video has undergone style conversion, which can be understood as a video whose target style specified by the user is the same as the initial video in content. In the process of converting the initial video uploaded by the user into a target video, it must first be processed by a deep learning model (preferably a large model such as a diffusion model) to obtain a reference video directly output by the deep learning model. The reference video is a video generated by the deep learning model based on the initial video, and the reference video can be understood as a video after the initial video has undergone style conversion. Since there may be no coherence between the video frames in the above-mentioned reference video, the above-mentioned reference video is then processed by the video processing method provided in this specification to obtain a target video with coherence between the video frames, and the target video is presented to the user to avoid presenting an incoherent video directly to the user.

[0059] It should be noted that the first initial video frame is any video frame in the initial video, and confirming the second initial video frame corresponding to the first initial video frame can be understood as confirming a preset number of frames near the first initial video frame. Specifically, the specific method for confirming the second initial video frame corresponding to the first initial video frame is preferably a sliding window with a preset step size, where the video frame at the center of the sliding window is the first initial video frame, and the remaining frames in the sliding window are the second initial video frames. It should be noted that after the sliding window slides to the target video frame, if the number of video frames surrounding the target video frame does not meet the preset step size, only the video frames that meet the step size are selected.

[0060] In one embodiment provided by the present application, the initial video uploaded by the user is video A, and video A has three video frames, namely initial video frame 0, initial video frame 1, and initial video frame 2. When the preset step size is 1, the sliding window will obtain a video frame before and after the target video frame. The sliding window begins to slide, and the initial position of the sliding window is initial video frame 0. Then, it is confirmed that the current first initial video frame is initial video frame 0, and the second initial video frame is a video frame adjacent to initial video frame 0, that is, initial video frame 1; then, the sliding window slides to the position of initial video frame 1, and the current first initial video frame is confirmed to be initial video frame 1, and the second initial video frame is a video frame adjacent to initial video frame 1, that is, initial video frame 0 and initial video frame 2; then, the sliding window slides to the position of initial video frame 2, and the current first initial video frame is confirmed to be initial video frame 2, and the second initial video frame is a video frame adjacent to initial video frame 2, that is, initial video frame 1. Then, it is confirmed that the current sliding window has slid to the last video frame of the video, and the sliding of the sliding window ends.

[0061] By acquiring the first initial video frame in the initial video and the second initial video frame surrounding the first initial video frame, it is possible to subsequently acquire the association relationship between the various video frames in the coherent initial video, so that the incoherent reference video can be processed subsequently through the association relationship between the various frames in the acquired initial video frames to generate a coherent target video.

[0062] Step 104: Determine a reference video corresponding to the initial video, and determine a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video.

[0063] In practical applications, the first reference video frame is a video frame at the same position as the first initial video frame in the reference video, and the second reference video frame is a video frame at the same position as the second initial video frame in the reference video frame.

[0064] Specifically, in the scheme for optimizing the reference video output by the deep learning model when the reference video has the same number of frames as the initial video, since the reference video has the same number of frames as the initial video, reconfirming the first reference video frame relative to the first initial video frame can be used to determine the video frame at the same position as the first initial video frame as the first reference video frame. Similarly, confirming the second reference video frame relative to the second initial video frame can be used to determine the video frame at the same position as the second initial video frame as the second reference video frame.

[0065] In one embodiment provided in the present application, the initial video A uploaded by the user has three video frames, namely initial video frame 0, initial video frame 1, and initial video frame 2. Subsequently, the deep learning model generates an animated reference video B based on the above initial video A. Reference video B also has three video frames, namely reference video frame 0, reference video frame 1, and reference video frame 2. Continuing with the above example, it is determined that there is a first initial video frame: initial video frame 1 and a second initial video frame: initial video frame 0 and initial video frame 2. Since initial video frame 1 is the second frame in the initial video, the first reference video frame corresponding to the first initial video frame is the second frame in the reference video, that is, reference video frame 1. Similarly, since initial video frame 0 and initial video frame 2 are the first and third frames in the initial video, respectively, the second reference video frame corresponding to the second initial video frame is the first and third frame in the reference video, that is, reference video frame 0 and reference video frame 2.

[0066] By confirming the first reference video frame and the second reference video frame corresponding to the above-mentioned first initial video frame and the second initial video frame in the reference video, it is possible to subsequently process the incoherent reference video through the association relationship between the frames in the acquired initial video frames to generate a coherent target video.

[0067] Step 106: Generate motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy.

[0068] In practical applications, the motion feature generation strategy is an information generation strategy for determining the connection information between the first initial video frame and the surrounding second initial video frames. The motion feature information is the connection information between the first initial video frame and the surrounding second initial video frames.

[0069] Specifically, the motion feature information between the first initial video frame and each second initial video frame can be understood as the corresponding position of each pixel point in the second initial video frame in the first initial video frame when each second initial video frame changes to the first initial video frame, which also reflects the change characteristics of each second initial video frame to the first initial video frame. Furthermore, it can be further understood as representing the motion characteristics of the entity in each second initial video frame changing to the corresponding entity in the first initial video frame.

[0070] By calculating the motion feature information between each frame in the reference video and determining it through the above-mentioned calculation of the feature information in the initial video, the motion relationship that should exist between each frame in the reference video can be determined through the smooth initial video, and then the non-smooth reference video can be processed through the motion feature information between each frame in the smooth initial video, thereby avoiding the video flickering problem caused by the discontinuity between frames, and thus realizing the rendering of the flickering reference video into a smooth target video.

[0071] Considering that the motion feature information obtained between video frames is the change relationship between each pixel between two video frames, the video frame includes at least one pixel point;

[0072] Furthermore, based on a preset motion feature generation strategy, generating motion feature information between the first initial video frame and each second initial video frame includes:

[0073] Determining a target second initial video frame, and determining at least one intermediate pixel point from the target second initial video frame, wherein the target second initial video frame is any one of the initial second video frames;

[0074] Acquire initial pixel motion feature information, and adjust the initial pixel motion feature information according to each intermediate pixel point and the first initial video frame to acquire target pixel motion feature information corresponding to each intermediate pixel point;

[0075] Motion feature information between the first initial video frame and the target second initial video frame is determined according to the motion feature information of each target pixel.

[0076] In actual applications, the intermediate pixel points are the pixel points in the target second initial video frame, the initial pixel motion feature information is the randomly initialized pixel motion feature information, and the target pixel motion feature information is the pixel motion feature information after adjustment. Specifically, the target pixel motion feature information is the intermediate pixel points in the target second initial video frame, and the positions of the corresponding pixel points in the first initial video frame.

[0077] By determining the corresponding position of each pixel point in the target second initial video frame in the first initial video frame, the motion feature information between the target second initial video frame and the first initial video frame can be determined, and then the non-smooth reference video can be processed according to the motion feature information between each frame in the smooth initial video, avoiding the video flickering problem caused by the incoherence between frames, and then rendering the flickering reference video into a smooth target video.

[0078] Considering the correspondence between pixels and the change relationship between pixels, it is necessary to compare two pixels and calculate the difference to determine the pixel that best matches another pixel. Therefore, the motion feature generation strategy includes a motion feature adjustment function.

[0079] Furthermore, adjusting the initial pixel motion feature information according to each intermediate pixel point and the first initial video frame includes:

[0080] Processing the intermediate pixel points according to the initial pixel motion feature information to obtain reference pixel points;

[0081] Obtaining a pixel difference value corresponding to the reference pixel point according to the motion feature adjustment function, the intermediate pixel point and the reference pixel point;

[0082] Adjusting the initial pixel motion feature information according to the pixel point difference value to obtain reference pixel motion feature information;

[0083] The reference pixel motion feature information is used as the initial pixel motion feature information, and the steps of processing the intermediate pixel points according to the initial pixel motion feature information to obtain the reference pixel points are continued until the adjustment stop condition is met.

[0084] In actual applications, the motion feature adjustment function is a function for adjusting the initial pixel motion feature information, the reference pixel point is the pixel point obtained after processing the intermediate pixel point based on the initial pixel motion feature information, the pixel point difference value is the difference between the pixel point after processing according to the initial pixel motion feature and the actual corresponding pixel point in the first initial video frame, and the reference pixel motion feature information is the pixel motion feature information after the initial pixel motion feature information is adjusted once.

[0085] Specifically, the reference pixel point can be understood as the pixel point coordinates of an information obtained after processing the coordinates of the intermediate pixel point according to the initial pixel motion feature information. According to the processed pixel point coordinates, the pixel point corresponding to the coordinates is determined in the first initial video frame, which is the reference pixel point.

[0086] It should be noted that the pixel points here include coordinate data and color data. When calculating the pixel difference subsequently, it is preferred to calculate the difference in pixel color. The specific method of adjusting the initial pixel motion feature information can be any method of obtaining pixel points, such as matching feedback, random search, etc. This manual does not impose any restrictions on this.

[0087] The adjustment stop condition may be, for example, the pixel difference reaching the target value, or the number of iterations of the algorithm reaching the target number, and this specification does not impose any restrictions thereon. Considering that the reference pixel motion feature information obtained after adjustment is stopped upon reaching the adjustment stop condition is not necessarily the reference pixel motion feature information with the smallest pixel difference, it is preferred that, after adjustment is stopped, the pixel motion feature information corresponding to the reference pixel with the smallest pixel difference is set as the target pixel motion feature information.

[0088] Considering that when processing a video, the efficiency of video processing can be improved on the basis of reducing computing power consumption, or the result of video processing can be further improved without considering computing power consumption, the motion feature adjustment function includes a motion feature region adjustment function or a motion feature overall adjustment function;

[0089] Furthermore, obtaining a pixel difference value corresponding to the reference pixel point according to the motion feature adjustment function, the intermediate pixel point, and the reference pixel point includes:

[0090] When the motion feature adjustment function is a motion feature area adjustment function, obtaining a pixel difference value according to the reference pixel and the intermediate pixel;

[0091] When the motion feature adjustment function is an overall motion feature adjustment function, the error pixel points corresponding to the intermediate pixel points in each reference second initial video frame are obtained, and the pixel difference value is obtained based on each error pixel point, the intermediate pixel point and the reference pixel point, wherein the reference second initial video frame is the second initial video frame other than the target second initial video frame.

[0092] In actual applications, the initial pixel point is the pixel point obtained by the coordinates of the reference pixel point in the first initial video frame, the motion feature area adjustment function is a calculation function that determines the pixel difference value based on the difference between the reference pixel point after the change and the intermediate pixel point before the change, the error pixel point is the pixel point corresponding to the intermediate pixel point in other second initial video frames, and the motion feature overall adjustment function is a calculation function that determines the pixel difference value based on the difference between the reference pixel point after the change and the intermediate pixel point before the change and the error pixel point in other second initial video frames.

[0093] Specifically, the error pixel point can be understood as the average value of the color of the pixels with the same coordinates as the middle pixel point in the other second initial video frames except the target second initial video frame, that is, the error pixel point.

[0094] In one embodiment provided in the present application, the motion feature adjustment function for calculating pixel difference values is a motion feature area adjustment function, and the specific calculation method is shown in Formula 1:

[0095]

[0096] Where x, y are the coordinates of the middle pixel point, S is the first initial video frame, T is the second initial video frame, and F is the initial pixel motion feature information. That is, the pixel difference with coordinates x, y, that is, the pixel difference of the middle pixel after processing according to the initial pixel motion feature information. S[F(x, y)] is the pixel point in the first initial video frame whose coordinates are processed by the initial pixel motion feature information, which is the reference pixel point. T[x, y] is the pixel point in the target second initial video frame whose coordinates are the coordinates of the middle pixel point, which is the middle pixel point. The calculation method is to calculate the color difference between the two pixels and square it.

[0097] In another embodiment provided by the present application, the motion feature adjustment function for calculating the pixel difference value is a motion feature overall adjustment function, and the specific calculation method is shown in Formula 2:

[0098]

[0099] Where x, y are the coordinates of the middle pixel point, S is the first initial video frame, T is the second initial video frame, and F is the initial pixel motion feature information. is the average value of the rest of the second initial video frames, and α is a preset weight parameter. The larger the parameter, the greater the weight of the target initial video frame, which means that more attention is paid to the similarity between the reference pixel and the intermediate pixel during the adjustment process. That is, the pixel difference value of the pixel with coordinates x, y, that is, the pixel difference value of the middle pixel after processing according to the initial pixel motion feature information. S[F(x, y)] is the pixel point in the first initial video frame whose coordinates are processed by the initial pixel motion feature information, which is the reference pixel point. T[x, y] is the pixel point in the target second initial video frame whose coordinates are the coordinates of the middle pixel point, which is the middle pixel point. It is the average value of the colors of the pixels with the same coordinates as the above coordinates in the other second initial video frames except the target second initial video frame according to the coordinates of the middle pixel point, wherein the calculation method is to respectively calculate the color differences between the middle pixel point and the reference pixel point and the error pixel point and square them, and perform weighted summation on the calculation results.

[0100] Considering that if the color of the pixel after the change is similar to that of the pixel before the change, the probability that the pixel after the change is the pixel after the change in the second initial video frame is greater, the pixel change mode (initial pixel motion characteristics) is adjusted by the intermediate pixel before the change and the reference pixel after the transformation, and the target pixel motion characteristics of each pixel can be automatically obtained.

[0101] Step 108: Process the first reference video frame according to each motion feature information to obtain a target video frame corresponding to the first reference video frame, and generate a target video according to at least one target video frame.

[0102] In practical applications, the target video frame is a video frame in the target video where the pixel gaps between adjacent target video frames are smooth. Specifically, the target video frame can be understood as a first reference video frame in the reference video fused with the features of a second reference video frame nearby to reduce the gap between the first reference video frame and the second reference video frame nearby. This avoids video flicker caused by discontinuities between frames in the video, and further enables rendering of the flickering reference video into a smooth target video.

[0103] Furthermore, processing the first reference video frame according to each motion feature information to obtain a target video frame corresponding to the first reference video frame includes:

[0104] Determining a target second reference video frame, wherein the target second reference video frame is any one of the second reference video frames;

[0105] Acquiring mapping motion feature information, wherein the mapping motion feature information is motion feature information between the first reference video frame and the target second reference video frame;

[0106] Processing the target second reference video frame according to the mapped motion feature information to generate a to-be-mixed video frame corresponding to the target second reference video frame;

[0107] According to a preset video frame mixing rule, the first reference video frame, and each to-be-mixed video frame, a target video frame corresponding to the first reference video frame is obtained.

[0108] In actual applications, the mapped motion feature information is the motion feature information between the target second initial video frame and the first initial video frame corresponding to the target second reference video frame. The video frame to be mixed is the video frame after the second reference video frame changes according to the mapped motion feature information. The video frame mixing rule is the rule for mixing the video frames to be mixed to generate the target video frame.

[0109] Specifically, since the second reference video frame in the reference video corresponds to the second initial video frame with the same video frame position in the initial video, and the first reference video frame in the reference video corresponds to the first initial video frame with the same video frame position in the initial video, the motion feature information between the first reference video frame and the target second reference video frame is the motion information between the first initial video frame and the target second reference video frame.

[0110] The video frame mixing rule can be understood as a mixing rule for mixing the video frames in a certain order. The specific mixing method between the video frames can be any method of fusing two video frames, for example,

[0111] Since the order in which video frames are mixed and the number of times they are mixed have a certain influence on the mixing result, the preset video frame mixing rules can effectively mix the video frames to be mixed after the transformation of each second reference video frame with the first reference video frame, so that the mixed target frame is similar to the first reference video frame as a whole, and the differences between the obtained target video frames are smooth.

[0112] Furthermore, processing the target second reference video frame according to the mapped motion feature information to generate a to-be-mixed video frame corresponding to the target second reference video frame includes:

[0113] Determining at least one pixel to be processed and a pixel region to be processed corresponding to each pixel to be processed from the target second reference video frame, wherein the pixel region includes at least one pixel;

[0114] Each pixel area to be processed is processed according to the mapped motion feature information, a target pixel point corresponding to each pixel point to be processed is determined, and a video frame to be mixed corresponding to the target second reference video frame is determined according to each target pixel point.

[0115] In practical applications, the pixel point to be processed is the pixel point in the target second reference video frame, the pixel area to be processed is the area composed of all pixel points within a preset distance around the pixel point to be processed, and the target pixel point is the pixel point obtained by processing each pixel point in the pixel area to be processed according to the mapped motion feature information.

[0116] Specifically, determining the pixel area to be processed corresponding to the pixel to be processed can be understood as obtaining all pixels within the preset range of the above-mentioned pixel to be processed according to the preset distance. For example, the coordinates of the pixel to be processed a are (8, 9) and the preset range is 1, then the obtained pixel area to be processed is the pixel area with the lower left corner being (7, 8) and the upper right corner being (9, 10).

[0117] It should be noted that, taking into account the situation that the area where the pixels to be processed are located is the edge of the video frame or any area to be processed where the preset distance cannot be obtained, black pixels are used to expand the video frame outward by the preset distance pixels. Since the value of black pixels stored in the computer is 0, using black pixels for expansion will not have any impact on the pixels in the area obtained by calculation.

[0118] After all the pixel points in the pixel area to be processed are processed by the corresponding mapping motion feature information, the method for obtaining the target pixel point can be to process all the pixel points in the pixel area to be processed by the corresponding mapping motion feature information to obtain the color corresponding to each pixel, and then take the average value of the colors corresponding to each pixel as the color of the target pixel point, and use the coordinates of the pixel point to be processed after the mapping motion feature information is processed as the coordinates of the target pixel to obtain the target pixel.

[0119] By generating a frame to be replaced by changing all pixels in an area corresponding to each pixel in each second reference video frame, it is possible to avoid the problem of excessive gaps between pixels in the converted video frame due to one-to-one conversion based on a single pixel in the video frame. In addition, compared to directly changing the entire video frame as a whole, dividing the video frame into multiple areas and changing the multiple areas can reduce the problem of excessive video memory caused by the change of the video frame, thereby improving the efficiency of video processing.

[0120] Considering that in actual applications, the video frame mixing strategy can be stored in a special table so that the video frames can be mixed faster with less computing power consumption when processing the video, the video frame mixing strategy includes a mixing table query strategy;

[0121] Furthermore, obtaining a target video frame corresponding to the first reference video frame according to a preset video frame mixing rule, the first reference video frame, and each video frame to be mixed includes:

[0122] Constructing a mixing table according to each first reference video frame and each to-be-mixed video frame corresponding to each first reference video frame;

[0123] Based on the mixing table query strategy, querying at least one intermediate mixing result corresponding to the first reference video frame from the mixing table;

[0124] The intermediate mixing results are mixed to obtain a target video frame corresponding to the first reference video frame.

[0125] In actual applications, the mixing table stores the intermediate mixing results of each first reference video frame under each preset sliding window size. The intermediate mixing result is the mixing result after preliminary mixing of each frame to be mixed. The mixing table query strategy is a query strategy for querying the intermediate mixing results of each first reference video frame based on the mixing table.

[0126] Specifically, the mixing table can be understood as a tree-like structure that stores intermediate mixing results after preliminary mixing of various results to be mixed under various sliding window sizes. In actual applications, the larger the sliding window is, the better the mixing effect is. Therefore, when processing, it is necessary to obtain video processing results when the sliding window is in multiple sizes for comparison to determine the video processing results with better effects. Therefore, compared with mixing the video frames to be mixed each time the sliding window is adjusted, the intermediate mixing results of the video frames to be mixed at each sliding size are stored in the storage table, so that the corresponding intermediate mixing results can be mobilized for mixing when generating the target video, so as to reduce the number of mixing times of the video frames to be mixed, thereby improving the efficiency of generating the target video.

[0127] Preferably, the hybrid table includes a process hybrid table and a result hybrid table; when the hybrid table is a process hybrid table, the corresponding hybrid table query strategy is a process hybrid table query strategy; when the hybrid table is a result hybrid table, the corresponding hybrid table query strategy is a result hybrid table query strategy.

[0128] The process mixing table can be understood as a tree structure that stores the intermediate mixing results of each first reference video frame under each preset sliding window size. The process mixing table query strategy can be understood as a strategy for determining the intermediate mixing results required for each first reference video frame according to the sliding window size selected by the user, such as Figure 2a and Figure 2b As shown, Figure 2a : is a schematic diagram of a mixing table of a video frame and a left video frame process in a video processing method provided by an embodiment of this specification. Figure 2b It is a schematic diagram of a mixing table of a video frame and a right video frame in a video processing method provided by an embodiment of this specification.

[0129] in, Figure 2a and Figure 2bThe mixing level in is the preset sliding window size. When the mixing level is 0, the corresponding sliding window size is 2 to the power of 0, that is, 1. When the mixing level is 1, the sliding window size is 1 to the power of 1, that is, 2. Similarly, the sliding window size corresponding to the mixing level 2 is 4, and the sliding window size corresponding to the mixing level 3 is 8. The target video frame is the serial number of the target video frame obtained after the first reference video frame is processed, that is, the serial number of the first reference video frame. In the embodiment shown in the figure, the video corresponds to 8 video frames. Taking the mixing process of the mixing level 2 and the target video frame 3 as an example, first obtain the intermediate mixing result of the target video frame 3 and the left video frame, and query Figure 2a As can be seen from the table, it is necessary to obtain the leaf node corresponding to the target frame 3 and the ancestor node of the leaf node in the mixing level 1 and the mixing level 2, that is, the intermediate mixing results S3, S2->S3 and (S0->S3)+(S1->S3); then the intermediate mixing results S3, S3<-S4 and (S3<-S5)+(S3<-S6) of the target video frame 3 and the right video frame are obtained in the same way; then the repeated S3 is removed and the above 5 intermediate mixing results are mixed to obtain the target video frame 3, that is, only 4 mixing sets are needed to obtain the target video frame 3, and to mix the video frames to be mixed obtained by the sliding window, it is necessary to mix the first reference video frame S3 and the 6 video frames to be mixed, S2->S3, S0->S3, S1->S3, S3<-S4, S3<-S5, S3<-S6, a total of 7 frames, and 6 mixings are required to obtain the target video frame 3.

[0130] The result mixing table can be understood as a tree structure that stores further intermediate mixing results of each first reference video frame under each preset sliding window size. The result mixing table query strategy can be understood as a strategy for determining the intermediate mixing results required for each first reference video frame according to the sliding window size selected by the user, such as Figure 3a and Figure 3b As shown, Figure 3a Schematic diagram of a mixed table of a video frame and a left video frame result in a video processing method provided by an embodiment of the present specification; Figure 3b It is a schematic diagram of a mixing table of a video frame and a right video frame result in a video processing method provided in one embodiment of this specification.

[0131] in, Figure 3a and Figure 3b The blending level and target video frame in the above Figure 2a and Figure 2bThe technical features are the same as those in the example, so they will not be repeated here. In the embodiment shown in the figure, the video corresponds to 8 video frames. Taking the mixing process of mixing level 2 and target video frame 3 as an example, first obtain the intermediate mixing result of target video frame 3 and the left video frame, and query Figure 2a As can be seen from the table, it is necessary to obtain the ancestor node of the leaf node corresponding to the target frame 3 in the mixing level 2, that is, the intermediate mixing result (S0->S3)+(S2->S3)+(S1->S3)+S3; then, in the same way, obtain the intermediate mixing result S3+(S3<-S5)+(S3<-S4)+(S3<-S6) of the target video frame 3 and the right video frame; then, mix the above two intermediate mixing results to obtain the video frame mixing result corresponding to the target video frame 3.

[0132] By generating the target video frame through the intermediate mixing result pairs stored in the result mixing table and the process mixing table, the number of mixing times required to generate the target video when the user adjusts the size of the sliding window can be effectively reduced, thereby improving the efficiency of video processing.

[0133] Considering that in addition to performing similarity mixing on each frame to generate a new video, key video frames of the video to be processed can also be extracted, and transition video frames can be inserted between the key video frames to achieve a smooth video. Therefore, before obtaining the first initial video frame in the initial video and at least one second initial video frame corresponding to the first initial video frame, the method further includes:

[0134] Acquire at least one initial key video frame in the initial video, and determine a reference key video frame corresponding to the initial key video frame from a reference video corresponding to the initial video;

[0135] Based on a preset key video frame feature generation strategy, key video frame feature information between each initial key video frame is generated;

[0136] Obtaining transition video frames corresponding to each reference key video frame according to each reference key video frame and key video frame feature information corresponding to each reference key video frame;

[0137] A target video is generated according to each reference key video frame and a transition video frame corresponding to each reference key video frame.

[0138] In practical applications, the initial key video frame is any key video frame in the initial video, the reference key video frame is the key video frame in the reference video corresponding to the initial key video frame, the key video frame feature generation strategy is an information generation strategy for generating motion information of two key video frames and an intermediate transition video frame, the key video frame feature information is information representing the motion features of the two key video frames and the intermediate transition video frame, and the transition video frame is a non-key video frame between the two key video frames.

[0139] Specifically, the initial key video frame can be understood as a video frame in which an entity in the initial video moves beyond a threshold value. The method for confirming the initial key video frame can be any method for confirming a key frame, and this specification does not impose any restrictions on this. The reference video for obtaining the reference key video frame is a video generated by a deep learning model based on the initial video. The reference video can only include the reference video key frames after processing the initial key video frames in the initial video, or it can be a reference video frame corresponding to all the initial video frames of the initial video. This application does not impose any restrictions on this.

[0140] The transition video frame can be understood as a non-key video frame between reference key video frames in the reference video. The above-mentioned transition video frame is generated by processing the two initial key video frames according to the key video frame feature information corresponding to the two initial key video frames to generate a transition video frame between the two initial key video frames. It should be noted that the key video feature information corresponding to the above-mentioned two key video frames may not only be one, but may include multiple key video feature information. In the case of including multiple key video feature information, at most the same number of transition video frames as the key video feature information can be generated between the above-mentioned two key video frames.

[0141] Considering that the key video frame feature information obtained between the key video frames is the possible change relationship of each pixel between two key video frames, the video frame includes at least one pixel point;

[0142] Based on the preset key video frame feature generation strategy, key video frame feature information between the initial key video frames is generated, including:

[0143] Determining a first initial key video frame and a second initial key video frame, and determining a reference transition video frame based on the first initial key video frame and the second initial key video frame, wherein the first initial key video frame is any one of the initial key video frames, and the second initial key video frame is a key video frame adjacent to the first initial key video frame;

[0144] Determining an initial transition pixel point in the reference transition video frame;

[0145] Acquiring initial pixel transition characteristic information, determining a first reference pixel and a second reference pixel based on the initial pixel transition characteristic information and the initial transition pixel point, and adjusting the initial pixel transition characteristic information according to the first reference pixel point, the second reference pixel point, and the initial transition pixel point to generate target pixel transition characteristic information corresponding to each of the first initial pixel point and the second reference pixel point;

[0146] Key video frame feature information between each initial key video frame is generated according to each target pixel transition feature information.

[0147] In actual applications, the first initial key video frame is any key frame of the initial video, the second initial key video frame is a key video frame adjacent to the first initial key video frame, the reference transition video frame is a non-key frame between the two initial key video frames, the initial transition pixel point is any pixel point in the above-mentioned reference transition video frame, the first reference pixel point is the pixel point corresponding to the initial transition pixel point in the first initial key video frame after processing according to the initial pixel transition feature information, the second reference pixel point is the pixel point corresponding to the initial transition pixel point in the second initial key video frame after processing according to the initial pixel transition feature information, the initial pixel transition feature information is randomly initialized pixel transition feature information, and the target pixel transition feature information is the adjusted initial pixel transition feature information.

[0148] Specifically, there can be multiple reference transition video frames. In the case of multiple reference transition video frames, multiple target pixel transition feature information will be generated between the first initial key video frame and the second initial key video frame, so that non-key frames up to the number of reference transition video frames can be inserted into the above two key video frames subsequently.

[0149] The pixel transition feature information is similar to the motion feature information mentioned above, and both represent feature information of changes between two video frames. The difference is that the pixel transition feature information here is feature information of changes between two initial key video frames and the reference transition video frame in between. That is, the output pixel transition feature information includes multiple groups, each of which contains feature information of changes with the first initial key video frame (first pixel transition feature information) and feature information of changes with the second initial key video frame (second pixel transition feature information). Accordingly, when generating the target video, the reference key video frames corresponding to the two adjacent initial key video frames can be processed separately, and the two obtained results can be mixed to obtain the intermediate non-key video frame (that is, the transition video frame).

[0150] Also, considering the correspondence between pixels and the change relationship between pixels, it is necessary to compare two pixels and calculate the difference to determine the pixel that best matches another pixel. Therefore, the key video frame feature generation strategy includes a transition video frame adjustment function.

[0151] Determining a first reference pixel and a second reference pixel based on the initial pixel transition feature information and the initial transition pixel point, and adjusting the initial pixel transition feature information according to the first reference pixel point, the second reference pixel point, and the initial transition pixel point to generate target pixel transition feature information corresponding to each of the first initial pixel point and the second reference pixel point, including:

[0152] Processing the initial transition pixel point according to the initial pixel transition feature information to obtain a first reference pixel point and a second reference pixel point;

[0153] Obtaining a transition pixel difference value corresponding to the reference pixel point according to the transition video frame adjustment function, the initial transition pixel point, the first reference pixel point, and the second reference pixel point;

[0154] Adjusting the initial pixel transition feature information according to the transition pixel point difference value to obtain reference pixel transition feature information;

[0155] The reference pixel transition characteristic information is used as the initial pixel transition characteristic information, and the steps of processing the initial transition pixel point according to the initial pixel transition characteristic information to obtain the first reference pixel point and the second reference pixel point are continued until the adjustment stop condition is reached.

[0156] In practical applications, the transition video frame adjustment function includes a first key video frame adjustment function and a second key video frame adjustment function, and the transition pixel point difference value includes a first transition pixel point difference value and a second transition pixel point difference value.

[0157] The first key video frame adjustment function is a function for calculating the first transition pixel difference between the first reference pixel and the initial transition pixel. The first transition pixel difference is a pixel difference used to adjust the change characteristic information between the first initial key video frames. The second key video frame adjustment function is a function for calculating the second transition pixel difference between the second reference pixel and the initial transition pixel. The second transition pixel difference is a pixel difference used to adjust the change characteristic information between the second initial key video frames.

[0158] It should be noted that, including coordinate data and color data, when subsequently calculating the pixel difference, it is preferred to calculate the difference in pixel color. The specific method of adjusting the initial pixel transition feature information can be any method of obtaining pixel points, such as matching feedback, random search, etc. In addition, the specific method of adjusting the initial pixel transition feature information can also include obtaining pixel points in adjacent key video frames, and this manual does not impose any restrictions on this.

[0159] In one embodiment provided in the present application, the first key video frame adjustment function and the second key video frame adjustment function

[0160] The combined key video frame adjustment function is shown in Formula 3:

[0161]

[0162] Where x, y are the coordinates of the initial transition pixel point, S l is the first initial key video frame, S r is the second initial key video frame, T is the reference transition video frame, F l is the change feature information between the first initial key video frame, F r is the change feature information between the initial key video frame and the second initial key video frame, α is a preset weight parameter. The larger the parameter, the greater the weight of the reference transition video frame, which means that more emphasis is placed on the similarity between the initial transition pixel point and the first reference pixel point and the second reference pixel point during the adjustment process. is the difference value of the first transition pixel with coordinates x, y, is the difference value of the second transition pixel with coordinates x, y, S l [F l (x, y)] is the pixel point in the first initial key video frame whose coordinates are processed by the initial pixel transition feature information and the coordinates of the initial transition pixel point are the first reference pixel point. r [F r (x, y)] is the pixel point whose coordinates in the second initial key video frame are processed by the initial pixel transition feature information to obtain the coordinates of the initial transition pixel point, which is the second reference pixel point. T[x, y] is the pixel point whose coordinates in the reference transition video frame are the coordinates of the intermediate pixel point, which is the initial transition pixel point.

[0163] The first transition pixel difference is calculated by calculating the difference between the initial transition pixel and the first reference pixel, squaring the difference, and then performing a weighted summation of the difference with the square of the difference between the first reference pixel and the second reference pixel. The second transition pixel difference is calculated by calculating the difference between the initial transition pixel and the second reference pixel, squaring the difference, and then performing a weighted summation of the difference with the square of the difference between the first reference pixel and the second reference pixel.

[0164] It should be noted that the adjustment stop condition here can be that each of the two pixel point differences reaches the target value, or that the number of iterations of the algorithm reaches the target number, etc., and this specification does not impose any restrictions on this. Considering that after the adjustment is stopped when the adjustment stop condition is reached, the reference pixel transition feature information obtained is not necessarily the reference pixel transition feature information with the smallest pixel point difference value, it is preferred that after the adjustment is stopped, the pixel transition feature information with the smallest sum of the two pixel point differences is set as the target pixel transition feature information.

[0165] By setting a function to calculate the reference frame and two key frames, the final target pixel transition feature information is obtained, which can improve the accuracy of the obtained target pixel transition information, so as to insert smooth transition video frames between the key video frames of the reference video, so that the flickering reference video can be rendered into a smoother target video.

[0166] Considering that the rendered target video needs to be sent to the user to present the rendering result to the user, after generating the target video, the method further includes:

[0167] The target video is sent to a target user.

[0168] In actual applications, the target user is the user who receives the target video. Specifically, the target user is usually the user who sent the initial video. The user sends the initial video and video modification information used to edit the initial video to the cloud-side device that provides the video processing method. The cloud-side device then generates a reference video similar to the first video and in accordance with the video modification information based on the initial video and video modification information sent by the user. The cloud-side device then performs the above steps on the video frames in the reference video to generate the target video of the process. The target video can then be sent to the target user so that the target user can view the modification results of the initial video. The target user can also be a different user from the user who sent the initial video, and this specification does not impose any restrictions on this.

[0169] Considering that the user may make further adjustments to the video after receiving the target video, after sending the target video to the target user, the method further includes:

[0170] receiving video adjustment information sent by the target user for the target video;

[0171] The target video is adjusted according to the video adjustment information to obtain the adjusted target video, and the adjusted target video is returned to the target user.

[0172] In actual applications, the video adjustment information is the information sent by the user to adjust the target video. The target video can be modified to generate a pending video that has not yet been processed in a process, and then the above steps are used to process the pending video to generate the adjusted target video.

[0173] The adjusted target video can be a video of a different style from the target video or a video that replaces an entity in the target video, or the target video can be left unchanged and a video of the same style or type can be generated again, etc. This application does not impose any restrictions on this.

[0174] When the video adjustment information sent by the user is to regenerate a video of the same style or type, the target video can be processed according to the same prompt words used by the user to process the initial video to generate a reference video, generating a reference video of the same style or type as the target video, and then the reference video can be processed again to generate a smooth adjusted target video, and this adjusted target video has the same style or type as the previous target video. Alternatively, the prompt words used to process the initial video to generate a reference video can be obtained, and then the initial video can be processed again according to the prompt words to generate a reference video of the same style or type as the target video, and then the reference video can be processed again to generate a smooth adjusted target video, and this adjusted target video has the same style or type as the previous target video, etc. This manual does not impose any restrictions on this.

[0175] In one embodiment provided in this specification, the target video is a video of a white rabbit eating grass, and the video adjustment information sent by the user is "change the white rabbit to a black rabbit". The above target video is then adjusted according to the adjustment information to generate an adjusted target video of a black rabbit eating grass, and the adjusted target video is returned to the user.

[0176] In another embodiment provided in this specification, the target video is a real video of a white rabbit eating grass, and the video adjustment information sent by the user is "change this video to anime style". The above-mentioned target video is then adjusted according to the adjustment information to generate an adjusted target video of an anime-style white rabbit eating grass video, and the adjusted target video is returned to the user.

[0177] In another embodiment provided in this specification, the target video is a real video of a white rabbit eating grass, and the video adjustment information sent by the user is "I think this video is not very good, give me another similar video", and then the above-mentioned target video is adjusted according to the adjustment information to generate another adjusted target video of a white rabbit eating grass video, and the adjusted target video is returned to the user.

[0178] By receiving video adjustment information fed back by the user based on the target video, adjusting the target video, and then returning the adjusted target video to the user, the target video can be further adjusted according to the user's needs, and the adjusted target video is returned to the user, so that the user can adjust the target video again, and further, the target video can be adjusted according to the adjustment information sent by the user until the user stops adjusting, so that the target video received by the user can be more in line with the user's needs, thereby improving the user experience.

[0179] Considering that adding certain prompt interactions when returning the video to the user can improve the user experience, after generating the target video, the method further includes:

[0180] Based on the target video, generating video adjustment prompt information;

[0181] Sending the target video and the video adjustment prompt information to a target user;

[0182] receiving video adjustment information sent by a target user based on the video adjustment prompt information;

[0183] The target video is adjusted according to the video adjustment information to obtain the adjusted target video, and the adjusted target video is returned to the target user.

[0184] In practical applications, video adjustment prompt information is information that prompts the user to adjust the target video. Specifically, the video adjustment prompt information can be understood as information that prompts the user to perform common adjustments to the target video. For example, after returning the target video to the target user, at least one interactive button is generated in the interactive interface, such as Button 1, "Generate another video of the same type," or Button 2, "Generate a video of a different style." The user clicks the corresponding button to generate video adjustment information. In addition, the user can also edit the video adjustment information to adjust the target video, etc. This manual does not impose any restrictions on this.

[0185] By generating video adjustment prompt information, the user can adjust the target video according to the above video adjustment prompt information, and some suggestions for adjusting the video are given to the user during interaction with the user, thereby improving the user experience.

[0186] By applying the solution of the embodiments of this specification, the motion feature information determined by calculating between video frames in the initial video through a motion feature generation strategy that does not require a deep learning model or a machine learning model can reduce the computing power consumption when processing the video while ensuring the smoothness of the generated video. Furthermore, by using the motion feature information between each frame in the reference video determined by the above calculation of the feature information in the initial video, the motion relationship that should exist between each frame in the reference video can be determined through the smooth initial video, and then the non-smooth reference video can be processed through the motion feature information between each frame in the smooth video, thereby avoiding the video flickering problem caused by the discontinuity between frames, and thus achieving the rendering of the flickering reference video into a smooth target video.

[0187] In addition, a special tree-like mixing table is used to store the mixing rules required for each frame, so that the video can be processed at a more efficient computing speed during video processing, thereby reducing the consumption of time and computing power, thereby improving the efficiency of video processing.

[0188] In addition, the key video frames in the smooth initial video are processed through the key video frame feature generation strategy to obtain the features of the transition video frames between the key video frames in the smooth video, and then the key video frames of the non-smooth reference video can be extracted, the key video frames of the reference video are obtained and processed, and smooth transition video frames are inserted between the key video frames of the reference video, so that the flickering reference video can be rendered into a smoother target video.

[0189] See also Figure 4 , Figure 4 A flowchart of a video processing method applied to a cloud server provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0190] Step 402: Receive a video processing instruction sent by a terminal-side device, wherein the video processing instruction carries an initial video and a reference video corresponding to the initial video.

[0191] Step 404: Acquire a first initial video frame in the initial video and at least one second initial video frame corresponding to the first initial video frame.

[0192] Step 406: Obtain a reference video corresponding to the initial video, and determine a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video.

[0193] Step 408: Generate motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy.

[0194] Step 410: Process the first reference video frame according to each motion feature information to obtain a target video frame corresponding to the first reference video frame, generate a target video based on at least one target video frame, and return the target video to the terminal device.

[0195] The above is a schematic diagram of the video processing method applied to the cloud server according to this embodiment. It should be noted that the technical solution of the video processing method applied to the cloud server and the technical solution of the video processing method described above are based on the same concept. For details not described in detail in the technical solution of the video processing method applied to the cloud server, please refer to the description of the technical solution of the video processing method described above.

[0196] By applying the solution of the embodiments of this specification, the motion feature information determined by calculating between video frames in the initial video through a motion feature generation strategy that does not require a deep learning model or a machine learning model can reduce the computing power consumption when processing the video while ensuring the smoothness of the generated video. Furthermore, by using the motion feature information between each frame in the reference video determined by the above calculation of the feature information in the initial video, the motion relationship that should exist between each frame in the reference video can be determined through the smooth initial video, and then the non-smooth reference video can be processed through the motion feature information between each frame in the smooth video, thereby avoiding the video flickering problem caused by the discontinuity between frames, and thus achieving the rendering of the flickering reference video into a smooth target video.

[0197] In addition, a special tree-like mixing table is used to store the mixing rules required for each frame, so that the video can be processed at a more efficient computing speed during video processing, thereby reducing the consumption of time and computing power, thereby improving the efficiency of video processing.

[0198] See also Figure 5 , Figure 5 1 shows an architecture diagram of a video processing system provided by an embodiment of the present specification. The video processing system may include a client 100 and a server 200.

[0199] The client 100 is configured to send a video processing instruction carrying an initial video and a reference video corresponding to the initial video to the server 200;

[0200] The server 200 is used to obtain the first initial video frame in the initial video, and at least one second initial video frame corresponding to the first initial video frame. Obtain a reference video corresponding to the initial video, and determine the first reference video frame corresponding to the first initial video frame and the second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video. Based on a preset motion feature generation strategy, generate motion feature information between the first initial video frame and each second initial video frame. Process the first reference video frame according to each motion feature information, obtain a target video frame corresponding to the first reference video frame, and generate a target video based on at least one target video frame; send the target video to the client 100;

[0201] The client 100 is further configured to receive the target video sent by the server 200 .

[0202] By applying the scheme of the embodiments of this specification, the motion feature information between each frame in the reference video determined by the motion feature information between video frames in the initial video can effectively reflect the motion feature information between each frame in the reference video. Subsequently, the motion feature information between each frame in the reference video determined by the above-mentioned calculation of the feature information in the initial video can determine the motion relationship that should exist between each frame in the reference video through the smooth initial video, and then the non-smooth reference video can be processed through the motion feature information between each frame in the smooth video, thereby avoiding the video flickering problem caused by the discontinuity between frames, and thus realizing the rendering of the flickering reference video into a smooth target video.

[0203] The video processing system may include multiple clients 100 and a server 200. The clients 100 may be referred to as end-side devices, and the server 200 may be referred to as cloud-side devices. Multiple clients 100 may establish communication connections through the server 200. In a video processing scenario, the server 200 provides video processing services between the multiple clients 100. The multiple clients 100 may act as either senders or receivers, communicating through the server 200.

[0204] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the video processing scenario, the user can publish a data stream to the server 200 through the client 100, and the server 200 can generate a target video based on the data stream and push the target video to other clients with which communication has been established.

[0205] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.

[0206] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application. The client 100 can be based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0207] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that support background training for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server that integrates a blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0208] It is worth noting that the video processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server to execute the video processing methods provided in the embodiments of this specification. In other embodiments, the video processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0209] See also Figure 6 , Figure 6 A flowchart of an automatic question-answering method provided by an embodiment of this specification is shown, which specifically includes the following steps:

[0210] Step 602: Receive a video modification instruction, wherein the video modification instruction carries an initial video and an initial question text.

[0211] Step 604: parse the video modification instruction to determine the initial question text, and generate an initial video modification prompt text based on the initial question text.

[0212] Step 606: Input the initial video and the initial video modification prompt text into the video modification model to obtain a reference video generated by the video modification model.

[0213] Step 608: Obtain motion feature information between the target first video frame and each second video frame in the reference video based on the initial video, wherein the target first video frame is any video frame in the reference video, and the second video frame is at least one video frame corresponding to the target first video frame in the reference video.

[0214] Step 610: Adjust each target first video frame in the reference video according to the motion feature information corresponding to each target first video frame in the reference video to generate a target video, and generate video modification answer information according to the target video and the initial question text.

[0215] In actual applications, the video modification instruction is an instruction to modify the initial video, the initial question text is the text entered by the user to modify the initial video, the initial video modification prompt text is the text input into the model to modify the initial video, and the video modification model is a model that modifies the initial video according to the above-mentioned initial video modification prompt text.

[0216] Specifically, the user sends a video modification instruction containing an initial video and an initial question text for editing the initial video to a cloud-side device that provides a video processing method. The cloud-side device then generates an initial video modification prompt text based on the initial question text sent by the user, and inputs the initial video modification prompt text and the initial video into the video modification model together to obtain a reference video output by the video modification model. Since the reference video directly output by the video has the problem of video discontinuity, the reference video is processed in the manner described in the above-mentioned video processing method for each frame in the reference video to obtain a smoother target video. Video modification answer information is then generated based on the target video, and the video modification answer information is returned to the user. The user can further modify the target video again until the user stops modifying the video.

[0217] The video modification model can be understood as a deep learning model used for video modification (preferably a large model such as a diffusion model is used). The video modification model can be used to modify the style or content of the original video based on the text sent by the user. This manual does not impose any restrictions on this.

[0218] The above is a schematic scheme of the automatic question-answering method of this embodiment. It should be noted that the technical scheme of this automatic question-answering method and the technical scheme of the above-mentioned video processing method are based on the same concept. For details not described in detail in the technical scheme of the automatic question-answering method, please refer to the description of the technical scheme of the above-mentioned video processing method.

[0219] By applying the solution of the embodiment of this specification, a video modification instruction sent by a user is received, the initial video that the user needs to modify is modified, and a reference video modified according to the target video is generated. Then, the motion feature information determined by calculating between the video frames in the initial video through a motion feature generation strategy that does not require a deep learning model or a machine learning model can be reduced while reducing the computing power consumption when processing the video while ensuring the smoothness of the generated video. Furthermore, the motion feature information between the frames in the reference video determined by the above calculation of the feature information in the initial video can determine the motion relationship that should exist between the frames in the reference video through the smooth initial video, and then the non-smooth reference video can be processed through the motion feature information between the frames in the smooth video, thereby avoiding the video flickering problem caused by the discontinuity between frames, and thus rendering the flickering reference video into a smooth target video.

[0220] The following combined Figure 7 , taking the application of the video processing method provided in this specification in video optimization as an example, the video processing method is further explained. Figure 7 A flowchart of a video optimization method according to an embodiment of the present disclosure is shown, which specifically includes the following steps.

[0221] Step 702: Obtain the original video and the reference video after style adjustment.

[0222] Step 704: Set a sliding window with a step size of 2, query the tree-structured mixing table, and determine the mixing method of each target video frame in the reference video and other video frames in the sliding window.

[0223] Step 706: Obtain motion feature information of the other video frames in the sliding window and the target video frame according to the above-identified mixing method, and process the other video frames in the sliding window according to the obtained motion feature information.

[0224] It should be noted that the specific method of obtaining the motion feature information of other video frames and the target video frame in the sliding window is to determine the target video frame in the reference video and the corresponding video frames of other video frames in the original video, and then obtain the motion feature information of other video frames and the target video frame based on the corresponding video frames in the original video.

[0225] Step 708: Mix the other processed video frames and the target video frame according to the above mixing method to obtain a processed video frame.

[0226] Step 710: Process all target video frames in the reference video to obtain multiple processed video frames, and splice the processed video frames into the target video.

[0227] By applying the solution of the embodiments of this specification, the motion feature information determined by calculating between video frames in the initial video through a motion feature generation strategy that does not require a deep learning model or a machine learning model can reduce the computing power consumption when processing the video while ensuring the smoothness of the generated video. Furthermore, by using the motion feature information between each frame in the reference video determined by the above calculation of the feature information in the initial video, the motion relationship that should exist between each frame in the reference video can be determined through the smooth initial video, and then the non-smooth reference video can be processed through the motion feature information between each frame in the smooth video, thereby avoiding the video flickering problem caused by the discontinuity between frames, and thus achieving the rendering of the flickering reference video into a smooth target video.

[0228] In addition, a special tree-like mixing table is used to store the mixing rules required for each frame, so that the video can be processed at a more efficient computing speed during video processing, thereby reducing the consumption of time and computing power, thereby improving the efficiency of video processing.

[0229] Corresponding to the above method embodiment, this specification also provides a video processing device embodiment, Figure 8 FIG. 1 shows a schematic diagram of the structure of a video processing device provided by an embodiment of this specification. Figure 8 As shown, the device includes:

[0230] The initial video frame acquisition module 802 is configured to acquire a first initial video frame in an initial video and at least one second initial video frame corresponding to the first initial video frame;

[0231] a reference video frame confirmation module 804 configured to determine a reference video corresponding to the initial video, and to determine a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video;

[0232] The motion feature information generating module 806 is configured to generate motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generating strategy;

[0233] The target video generation module 808 is configured to process the first reference video frame according to each motion feature information, obtain a target video frame corresponding to the first reference video frame, and generate a target video according to at least one target video frame.

[0234] Optionally, the video frame includes at least one pixel;

[0235] The motion feature information generating module 806 is further configured to:

[0236] Determining a target second initial video frame, and determining at least one intermediate pixel point from the target second initial video frame, wherein the target second initial video frame is any one of the initial second video frames;

[0237] Acquire initial pixel motion feature information, and adjust the initial pixel motion feature information according to each intermediate pixel point and the first initial video frame to acquire target pixel motion feature information corresponding to each intermediate pixel point;

[0238] Motion feature information between the first initial video frame and the target second initial video frame is determined according to the motion feature information of each target pixel.

[0239] Optionally, the motion feature generation strategy includes a motion feature adjustment function;

[0240] The motion feature information generating module 806 is further configured to:

[0241] Processing the intermediate pixel points according to the initial pixel motion feature information to obtain reference pixel points;

[0242] Obtaining a pixel difference value corresponding to the reference pixel point according to the motion feature adjustment function, the intermediate pixel point and the reference pixel point;

[0243] Adjusting the initial pixel motion feature information according to the pixel point difference value to obtain reference pixel motion feature information;

[0244] The reference pixel motion feature information is used as the initial pixel motion feature information, and the steps of processing the intermediate pixel points according to the initial pixel motion feature information to obtain the reference pixel points are continued until the adjustment stop condition is met.

[0245] Optionally, the motion feature adjustment function includes a motion feature region adjustment function or a motion feature overall adjustment function;

[0246] The motion feature information generating module 806 is further configured to:

[0247] Determining an initial pixel point corresponding to the reference pixel point in the first initial video frame;

[0248] When the motion feature adjustment function is a motion feature area adjustment function, obtaining a pixel difference value according to the reference pixel and the intermediate pixel;

[0249] When the motion feature adjustment function is an overall motion feature adjustment function, the error pixel points corresponding to the intermediate pixel points in each reference second initial video frame are obtained, and the pixel difference value is obtained based on each error pixel point, the intermediate pixel point and the reference pixel point, wherein the reference second initial video frame is the second initial video frame other than the target second initial video frame.

[0250] Optionally, the target video generation module 808 is further configured to:

[0251] Determining a target second reference video frame, wherein the target second reference video frame is any one of the second reference video frames;

[0252] Acquiring mapping motion feature information, wherein the mapping motion feature information is motion feature information between the first reference video frame and the target second reference video frame;

[0253] Processing the target second reference video frame according to the mapped motion feature information to generate a to-be-mixed video frame corresponding to the target second reference video frame;

[0254] According to a preset video frame mixing rule, the first reference video frame, and each to-be-mixed video frame, a target video frame corresponding to the first reference video frame is obtained.

[0255] Optionally, the target video generation module 808 is further configured to:

[0256] Determining at least one pixel to be processed and a pixel region to be processed corresponding to each pixel to be processed from the target second reference video frame, wherein the pixel region includes at least one pixel;

[0257] Each pixel area to be processed is processed according to the mapped motion feature information, a target pixel point corresponding to each pixel point to be processed is determined, and a video frame to be mixed corresponding to the target second reference video frame is determined according to each target pixel point.

[0258] Optionally, the video frame mixing strategy includes a mixing table query strategy;

[0259] The target video generation module 808 is further configured to obtain a target video frame corresponding to the first reference video frame according to a preset video frame mixing rule, the first reference video frame, and each video frame to be mixed, including:

[0260] Constructing a mixing table according to each first reference video frame and each to-be-mixed video frame corresponding to each first reference video frame;

[0261] Based on the mixing table query strategy, querying at least one intermediate mixing result corresponding to the first reference video frame from the mixing table;

[0262] The intermediate mixing results are mixed to obtain a target video frame corresponding to the first reference video frame.

[0263] Optionally, the video processing device further includes a frame insertion module configured to:

[0264] Acquire at least one initial key video frame in the initial video, and determine a reference key video frame corresponding to the initial key video frame from a reference video corresponding to the initial video;

[0265] Based on a preset key video frame feature generation strategy, key video frame feature information between each initial key video frame is generated;

[0266] Obtaining transition video frames corresponding to each reference key video frame according to each reference key video frame and key video frame feature information corresponding to each reference key video frame;

[0267] A target video is generated according to each reference key video frame and a transition video frame corresponding to each reference key video frame.

[0268] Optionally, the video frame includes at least one pixel;

[0269] The interpolation module is further configured to:

[0270] Determining a first initial key video frame and a second initial key video frame, and determining a reference transition video frame based on the first initial key video frame and the second initial key video frame, wherein the first initial key video frame is any one of the initial key video frames, and the second initial key video frame is a key video frame adjacent to the first initial key video frame;

[0271] Determining an initial transition pixel point in the reference transition video frame;

[0272] Acquiring initial pixel transition characteristic information, determining a first reference pixel and a second reference pixel based on the initial pixel transition characteristic information and the initial transition pixel point, and adjusting the initial pixel transition characteristic information according to the first reference pixel point, the second reference pixel point, and the initial transition pixel point to generate target pixel transition characteristic information corresponding to each of the first initial pixel point and the second reference pixel point;

[0273] Key video frame feature information between each initial key video frame is generated according to each target pixel transition feature information.

[0274] Optionally, the key video frame feature generation strategy includes a transition video frame adjustment function;

[0275] The interpolation module is further configured to:

[0276] Processing the initial transition pixel point according to the initial pixel transition feature information to obtain a first reference pixel point and a second reference pixel point;

[0277] Obtaining a transition pixel difference value corresponding to the reference pixel point according to the transition video frame adjustment function, the initial transition pixel point, the first reference pixel point, and the second reference pixel point;

[0278] Adjusting the initial pixel transition feature information according to the transition pixel point difference value to obtain reference pixel transition feature information;

[0279] The reference pixel transition characteristic information is used as the initial pixel transition characteristic information, and the steps of processing the initial transition pixel point according to the initial pixel transition characteristic information to obtain the first reference pixel point and the second reference pixel point are continued until the adjustment stop condition is reached.

[0280] Optionally, the video processing device further includes a target video sending module configured to:

[0281] The target video is sent to a target user.

[0282] Optionally, the video processing device further includes a target video adjustment module configured to:

[0283] receiving video adjustment information sent by the target user for the target video;

[0284] The target video is adjusted according to the video adjustment information to obtain the adjusted target video, and the adjusted target video is returned to the target user.

[0285] Optionally, the video processing device further includes a user prompt adjustment module configured to:

[0286] Based on the target video, generating video adjustment prompt information;

[0287] Sending the target video and the video adjustment prompt information to a target user;

[0288] receiving video adjustment information sent by a target user based on the video adjustment prompt information;

[0289] The target video is adjusted according to the video adjustment information to obtain the adjusted target video, and the adjusted target video is returned to the target user.

[0290] The above is a schematic scheme of a video processing device of this embodiment. It should be noted that the technical scheme of the video processing device and the technical scheme of the above-mentioned video processing method are based on the same concept. For details not described in detail in the technical scheme of the video processing device, please refer to the description of the technical scheme of the above-mentioned video processing method.

[0291] By applying the solution of the embodiments of this specification, the motion feature information determined by calculating between video frames in the initial video through a motion feature generation strategy that does not require a deep learning model or a machine learning model can reduce the computing power consumption when processing the video while ensuring the smoothness of the generated video. Furthermore, by using the motion feature information between each frame in the reference video determined by the above calculation of the feature information in the initial video, the motion relationship that should exist between each frame in the reference video can be determined through the smooth initial video, and then the non-smooth reference video can be processed through the motion feature information between each frame in the smooth video, thereby avoiding the video flickering problem caused by the discontinuity between frames, and thus achieving the rendering of the flickering reference video into a smooth target video.

[0292] In addition, a special tree-like mixing table is used to store the mixing rules required for each frame, so that the video can be processed at a more efficient computing speed during video processing, thereby reducing the consumption of time and computing power, thereby improving the efficiency of video processing.

[0293] In addition, the key video frames in the smooth initial video are processed through the key video frame feature generation strategy to obtain the features of the transition video frames between the key video frames in the smooth video, and then the key video frames of the non-smooth reference video can be extracted, the key video frames of the reference video are obtained and processed, and smooth transition video frames are inserted between the key video frames of the reference video, so that the flickering reference video can be rendered into a smoother target video.

[0294] Figure 9The block diagram of a computing device 900 according to one embodiment of the present disclosure is shown. Components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.

[0295] The computing device 900 also includes an access device 940 that enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 902.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0296] In one embodiment of the present specification, the above components of the computing device 900 and Figure 9 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 9 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0297] The computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 may also be a mobile or stationary server.

[0298] Among them, the processor 920 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned video processing method, an automatic question-answering method, or a video processing method applied to a cloud server.

[0299] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device is based on the same concept as the aforementioned video processing method, an automatic question-answering method, or the technical solution of the video processing method applied to a cloud server. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned video processing method, an automatic question-answering method, or the technical solution of the video processing method applied to a cloud server.

[0300] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned video processing method, an automatic question-answering method, or a video processing method applied to a cloud server.

[0301] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the aforementioned video processing method, an automatic question-answering method, or the technical solution of the video processing method applied to a cloud server. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the video processing method, an automatic question-answering method, or the technical solution of the video processing method applied to a cloud server.

[0302] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned video processing method, an automatic question-answering method, or a video processing method applied to a cloud server.

[0303] The above is an illustrative solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the technical solution of the aforementioned video processing method, an automatic question-answering method, or a video processing method applied to a cloud server. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the aforementioned video processing method, an automatic question-answering method, or a video processing method applied to a cloud server.

[0304] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0305] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0306] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0307] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0308] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A video processing method, comprising: Acquire a first initial video frame in an initial video, and at least one second initial video frame corresponding to the first initial video frame; Determining a reference video corresponding to the initial video, and determining a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video; generating motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy; The first reference video frame is processed according to each motion feature information to obtain a target video frame corresponding to the first reference video frame, and a target video is generated according to at least one target video frame.

2. The method according to claim 1, wherein the video frame comprises at least one pixel; Generating motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy includes: Determining a target second initial video frame, and determining at least one intermediate pixel point from the target second initial video frame, wherein the target second initial video frame is any one of the initial second video frames; Acquire initial pixel motion feature information, and adjust the initial pixel motion feature information according to each intermediate pixel point and the first initial video frame to acquire target pixel motion feature information corresponding to each intermediate pixel point; Motion feature information between the first initial video frame and the target second initial video frame is determined according to the motion feature information of each target pixel.

3. The method of claim 2, wherein the motion feature generation strategy comprises a motion feature adjustment function; Adjusting the initial pixel motion feature information according to each intermediate pixel point and the first initial video frame includes: Processing the intermediate pixel points according to the initial pixel motion feature information to obtain reference pixel points; Obtaining a pixel difference value corresponding to the reference pixel point according to the motion feature adjustment function, the intermediate pixel point and the reference pixel point; Adjusting the initial pixel motion feature information according to the pixel point difference value to obtain reference pixel motion feature information; The reference pixel motion feature information is used as the initial pixel motion feature information, and the steps of processing the intermediate pixel points according to the initial pixel motion feature information to obtain the reference pixel points are continued until the adjustment stop condition is met.

4. The method according to claim 3, wherein the motion feature adjustment function comprises a motion feature region adjustment function or a motion feature overall adjustment function; Obtaining a pixel difference value corresponding to the reference pixel point according to the motion feature adjustment function, the intermediate pixel point, and the reference pixel point, including: When the motion feature adjustment function is a motion feature area adjustment function, obtaining a pixel difference value according to the reference pixel and the intermediate pixel; When the motion feature adjustment function is an overall motion feature adjustment function, the error pixel points corresponding to the intermediate pixel points in each reference second initial video frame are obtained, and the pixel difference value is obtained based on each error pixel point, the intermediate pixel point and the reference pixel point, wherein the reference second initial video frame is the second initial video frame other than the target second initial video frame.

5. The method according to claim 1, wherein the step of processing the first reference video frame according to each motion feature information to obtain a target video frame corresponding to the first reference video frame comprises: Determining a target second reference video frame, wherein the target second reference video frame is any one of the second reference video frames; Acquiring mapping motion feature information, wherein the mapping motion feature information is motion feature information between the first reference video frame and the target second reference video frame; Processing the target second reference video frame according to the mapped motion feature information to generate a to-be-mixed video frame corresponding to the target second reference video frame; According to a preset video frame mixing rule, the first reference video frame, and each to-be-mixed video frame, a target video frame corresponding to the first reference video frame is obtained.

6. The method according to claim 5, wherein processing the target second reference video frame according to the mapped motion feature information to generate a to-be-mixed video frame corresponding to the target second reference video frame comprises: Determining at least one pixel to be processed and a pixel region to be processed corresponding to each pixel to be processed from the target second reference video frame, wherein the pixel region includes at least one pixel; Each pixel area to be processed is processed according to the mapped motion feature information, a target pixel point corresponding to each pixel point to be processed is determined, and a video frame to be mixed corresponding to the target second reference video frame is determined according to each target pixel point.

7. The method of claim 5, wherein the video frame blending strategy comprises a blending table lookup strategy; Obtaining a target video frame corresponding to the first reference video frame according to a preset video frame mixing rule, the first reference video frame, and each to-be-mixed video frame, including: Constructing a mixing table according to each first reference video frame and each to-be-mixed video frame corresponding to each first reference video frame; Based on the mixing table query strategy, querying at least one intermediate mixing result corresponding to the first reference video frame from the mixing table; The intermediate mixing results are mixed to obtain a target video frame corresponding to the first reference video frame.

8. The method according to claim 1, further comprising, before obtaining a first initial video frame in an initial video and at least one second initial video frame corresponding to the first initial video frame: Acquire at least one initial key video frame in the initial video, and determine a reference key video frame corresponding to the initial key video frame from a reference video corresponding to the initial video; Based on a preset key video frame feature generation strategy, key video frame feature information between each initial key video frame is generated; Obtaining transition video frames corresponding to each reference key video frame according to each reference key video frame and key video frame feature information corresponding to each reference key video frame; A target video is generated according to each reference key video frame and a transition video frame corresponding to each reference key video frame.

9. The method according to claim 8, wherein the video frame comprises at least one pixel; Based on the preset key video frame feature generation strategy, key video frame feature information between the initial key video frames is generated, including: Determining a first initial key video frame and a second initial key video frame, and determining a reference transition video frame based on the first initial key video frame and the second initial key video frame, wherein the first initial key video frame is any one of the initial key video frames, and the second initial key video frame is a key video frame adjacent to the first initial key video frame; Determining an initial transition pixel point in the reference transition video frame; Acquiring initial pixel transition characteristic information, determining a first reference pixel and a second reference pixel based on the initial pixel transition characteristic information and the initial transition pixel point, and adjusting the initial pixel transition characteristic information according to the first reference pixel point, the second reference pixel point, and the initial transition pixel point to generate target pixel transition characteristic information corresponding to each of the first initial pixel point and the second reference pixel point; Key video frame feature information between each initial key video frame is generated according to each target pixel transition feature information.

10. The method of claim 9, wherein the key video frame feature generation strategy comprises a transition video frame adjustment function; Determining a first reference pixel and a second reference pixel based on the initial pixel transition feature information and the initial transition pixel point, and adjusting the initial pixel transition feature information according to the first reference pixel point, the second reference pixel point, and the initial transition pixel point to generate target pixel transition feature information corresponding to each of the first initial pixel point and the second reference pixel point, including: Processing the initial transition pixel point according to the initial pixel transition feature information to obtain a first reference pixel point and a second reference pixel point; Obtaining a transition pixel difference value corresponding to the reference pixel point according to the transition video frame adjustment function, the initial transition pixel point, the first reference pixel point, and the second reference pixel point; Adjusting the initial pixel transition feature information according to the transition pixel point difference value to obtain reference pixel transition feature information; The reference pixel transition characteristic information is used as the initial pixel transition characteristic information, and the steps of processing the initial transition pixel point according to the initial pixel transition characteristic information to obtain the first reference pixel point and the second reference pixel point are continued until the adjustment stop condition is reached.

11. The method according to any one of claims 1 to 10, further comprising, after generating the target video: The target video is sent to a target user.

12. The method according to claim 11, after sending the target video to the target user, the method further comprises: receiving video adjustment information sent by the target user for the target video; The target video is adjusted according to the video adjustment information to obtain the adjusted target video, and the adjusted target video is returned to the target user.

13. The method according to any one of claims 1 to 10, further comprising, after generating the target video: Based on the target video, generating video adjustment prompt information; Sending the target video and the video adjustment prompt information to a target user; receiving video adjustment information sent by a target user based on the video adjustment prompt information; The target video is adjusted according to the video adjustment information to obtain the adjusted target video, and the adjusted target video is returned to the target user.

14. An automatic question-answering method, comprising: receiving a video modification instruction, wherein the video modification instruction carries an initial video and an initial question text; Parsing the video modification instruction to determine an initial question text, and generating an initial video modification prompt text according to the initial question text; Inputting the initial video and the initial video modification prompt text into a video modification model to obtain a reference video generated by the video modification model; Acquire motion feature information between a target first video frame and each second video frame in the reference video according to the initial video, wherein the target first video frame is any video frame in the reference video, and the second video frame is at least one video frame corresponding to the target first video frame in the reference video; Each target first video frame in the reference video is adjusted according to motion feature information corresponding to each target first video frame in the reference video to generate a target video, and video modification answer information is generated according to the target video and the initial question text.

15. A video processing method, applied to a cloud server, comprising: A video processing instruction sent by a receiving device, wherein the video processing instruction carries an initial video and a reference video corresponding to the initial video; Acquire a first initial video frame in the initial video, and at least one second initial video frame corresponding to the first initial video frame; Obtaining a reference video corresponding to the initial video, and determining a first reference video frame corresponding to the first initial video frame and a second reference video frame corresponding to each second initial video frame, wherein the number of video frames of the reference video is the same as the number of video frames of the initial video; generating motion feature information between the first initial video frame and each second initial video frame based on a preset motion feature generation strategy; The first reference video frame is processed according to each motion feature information to obtain a target video frame corresponding to the first reference video frame, and a target video is generated according to at least one target video frame, and the target video is returned to the terminal side device.

16. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 13, 14 or 15 are implemented.

17. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13, 14 or 15.

18. A computer program product comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13, 14 or 15.