A video processing method and related device

By 3D video clips, setting up scene models, and adding instructions to process according to special effects, the problem of complexity and error-prone items in video special effects processing is solved, and more efficient and accurate special effects processing is achieved.

CN115633134BActive Publication Date: 2025-06-20SHENZHEN SHANJIAN INTELLIGENT SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211176674.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-26
Publication Date
2025-06-20
Estimated Expiration
2042-09-26

AI Technical Summary

Technical Problem

The prior art is difficult and prone to errors in video special effects processing, especially in the process of increasing items into video, especially due to complexity caused by changes in view angles and directions between video frames.

Method used

By obtaining video clips in the video file, performing two-dimensional scene three-dimensionalization, and establishing a scene model. When special effects addition instructions are detected, the video file is processed according to the scene model and instructions to add and process item elements.

Benefits of technology

It reduces the difficulty of video special effects processing, improves the simplicity and speed of processing, ensures stable addition and correct projection of item elements, and enhances the accuracy of special effects processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115633134B_ABST
    Figure CN115633134B_ABST
Patent Text Reader

Abstract

The present invention discloses a video processing method and related devices. The method includes: obtaining a video file, where the video file includes a plurality of video segments; for each of the video segments, taking the video segment as a processing segment, and performing three-dimensionalization of a two-dimensional scene based on the image frames in the processing segment to obtain a scene model corresponding to the processing segment; when a special effect addition instruction corresponding to a preset item element is detected, processing the video file according to the scene model and the special effect addition instruction to obtain a special effect video. The present invention provides a convenient and fast special effect processing method for videos, improving the efficiency of video special effect processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimedia technology, and particularly relates to a video processing method and related devices. Background Art

[0002] With the development of Internet entertainment, more and more people are involved in Internet audio production. With the development and popularization of technology, the thresholds for audio production and video production are getting lower and lower. In video production, an essential link is the special effect processing of videos. For example, adjusting the light, adding items, or removing items.

[0003] For the type of adjusting light, as long as the video frames to be adjusted are determined, they can be processed in a unified manner. However, for the latter two types, due to the changes in perspective and direction between each frame of the video, when adding items to the video, it is necessary to continuously adjust parameters such as the perspective and position of the items in each image frame, which is rather cumbersome and error-prone. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to reduce the difficulty of video special effects. In view of the deficiencies of the prior art, a video processing method and related devices are provided.

[0005] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0006] A video processing method, the method comprising:

[0007] Obtain a video file, wherein the video file includes a plurality of video segments;

[0008] For each of the video segments, use the video segment as a processing segment, and perform three-dimensionalization of a two-dimensional scene based on the image frames in the processing segment to obtain a scene model corresponding to the processing segment;

[0009] When a special effect addition instruction for a preset item element is detected, process the video file according to the scene model and the special effect addition instruction to obtain a special effect video.

[0010] The video processing method, wherein for each of the video segments, using the video segment as a processing segment, and performing three-dimensionalization of a two-dimensional scene based on the image frames in the processing segment to obtain a scene model corresponding to the processing segment includes:

[0011] For each of the video segments, use the video segment as a processing segment, input the processing segment into a trained three-dimensional scene model, and control the three-dimensional scene model to perform three-dimensionalization on the image frames in the processing segment to obtain a scene model corresponding to the video segment.

[0012] The video processing method, wherein, before taking each of the video segments as a processing segment and performing two-dimensional scene three-dimensionalization according to the image frames in the processing segment to obtain a scene model corresponding to the processing segment, further includes:

[0013] Obtain a training video for a training model;

[0014] Perform data conversion on the training frames in the training video to obtain a five-dimensional function, wherein the five-dimensional function includes position coordinates and observation vectors;

[0015] Input the five-dimensional vector into a preset three-dimensional prediction model to obtain the voxel color and voxel density corresponding to the five-dimensional feature;

[0016] Render all the voxel colors and the voxel densities to obtain a prediction model;

[0017] Based on a preset loss function, calculate the loss value between the prediction model and the training model;

[0018] Based on the loss value, train the three-dimensional prediction model until the three-dimensional prediction model converges.

[0019] The video processing method, wherein, when a special effect addition instruction for a preset item element is detected, processing the video file according to the scene model and the special effect addition instruction to obtain a special effect video includes:

[0020] Perform object recognition on the scene model to obtain a plurality of tracking objects;

[0021] When a special effect addition instruction for a preset item element is detected, determine a tracking element and an insertion coordinate in the tracking objects according to the indication coordinates in the special effect addition instruction;

[0022] Determine the processing parameters corresponding to the item element according to the scene model, the tracking element, and the indication coordinates;

[0023] Process the video file according to the insertion coordinate, the processing parameters, and the item element to obtain a special effect video.

[0024] The video processing method, wherein the processing parameters include perspective parameters and projection parameters; the determining the processing parameters corresponding to the item element according to the scene model, the tracking element, and the indication coordinates includes:

[0025] Determine the perspective parameters corresponding to the item element according to the tracking element and the perspective information corresponding to the scene model; and,

[0026] Determine the projection parameters corresponding to the item element according to the light information of the tracking element in the scene model.

[0027] The video processing method, wherein the determining the projection parameters corresponding to the item element according to the light information of the tracking element in the scene model includes:

[0028] Determine the light surface corresponding to the tracking element according to the light source distribution information in the scene model, wherein the light surface includes a light-receiving surface, a side-light surface, and a backlight surface;

[0029] Calculate the light propagation function and the luminance transfer function according to the luminance values of the light-receiving surface, the side-light surface, and the backlight surface in the tracking element, and the luminance value of the light source in the scene model;

[0030] Calculate the projection parameters corresponding to the item element according to the indication coordinates, the light propagation function, and the luminance transfer function.

[0031] The video processing method, wherein the special effect video includes a plurality of special effect images; the processing the video file according to the insertion coordinates, the processing parameters, and the item element to obtain the special effect video includes:

[0032] When the scene model includes a mirror object, generate a mirror element corresponding to the item element, mirror coordinates corresponding to the mirror element, and mirror parameters according to the insertion coordinates and the world coordinates of the mirror object;

[0033] Process the video file according to the mirror information and the item information to obtain a special effect video, wherein the mirror information includes the mirror element, the mirror coordinates, and the mirror parameters, and the item information includes the item element, the insertion coordinates, and the processing parameters.

[0034] The video processing method, wherein the special effect video includes a special effect processing video and a special effect supplementary video; the processing the video file according to the insertion coordinates, the processing parameters, and the item element to obtain the special effect video includes:

[0035] Compare the to-be-processed model corresponding to the to-be-processed segment with the scene model to determine the comparison model corresponding to the scene model;

[0036] Generate a special effect supplementary instruction corresponding to the comparison model according to the special effect addition instruction;

[0037] Process the to-be-processed segment corresponding to the comparison model according to the special effect supplementary instruction to obtain a special effect supplementary video; and,

[0038] Process the processing segment according to the insertion coordinates, the processing parameters, and the item elements to obtain a special effect processed video.

[0039] A computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in any of the video processing methods described above.

[0040] A terminal device includes: a processor, a memory, and a communication bus; a computer-readable program executable by the processor is stored on the memory;

[0041] The communication bus realizes the connection and communication between the processor and the memory;

[0042] When the processor executes the computer-readable program, it implements the steps in any of the video processing methods described above.

[0043] Beneficial effects: According to the image frames, the present invention creates a scene model for different video segments in a video file. When the user needs to add an item to the video file, the item element is added to the scene model. Since the scene model is derived from the image frames, the image frames can be regarded as the result of projecting the scene model from a certain perspective. Therefore, based on the scene model with the added item, the two-dimensional image frames can be modified. Thus, the item can enter from the three-dimensional scene model into the two-dimensional frame image, improving the simplicity and speed of special effect processing. Description of the Drawings

[0044] Figure 1 It is a flowchart of the video processing method provided by the present invention.

[0045] Figure 2 It is a schematic diagram of splitting a video file into video segments in the video processing method provided by the present invention.

[0046] Figure 3 It is a schematic diagram of the display interface in the video processing method provided by the present invention.

[0047] Figure 4 It is a schematic diagram of the light surface in the video processing method provided by the present invention.

[0048] Figure 5 It is a schematic diagram of the structure principle of the terminal device provided by the present invention. Detailed Embodiments

[0049] The present invention provides a video processing method. To make the objectives, technical solutions and effects of the present invention clearer and more explicit, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0050] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0051] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0052] As Figure 1 shown, the present embodiment provides a video processing method. For the convenience of description, a common server is used as the execution subject for description. Here, the server can be replaced by devices with data processing functions such as tablets and computers. The video processing method includes the following steps:

[0053] S10. Obtain a video file.

[0054] Specifically, first, obtain the video file to be processed. The video file can be composed of one or more video segments. The video file can be sourced from local, cloud, or client transmission. The division criteria for video segments and video segments lie in whether the source shots are the same. In the video file, different video segments may be marked. At this time, as long as the video file is directly split according to the marks, several video segments can be obtained. For video files without marks, the video file can be split according to shot boundaries to obtain several video segments. The determination of shot boundaries can be achieved by means of video background, area changes of people or objects in image frames, etc. For example, if the previous image frame contains a person and the next image frame does not contain a person, then the boundary between the previous image frame and the next image frame is determined as the shot boundary.

[0055] S20. For each of the said video segments, take the video segment as the processing segment, and perform three-dimensionalization of the two-dimensional scene according to the image frames in the processing segment to obtain a scene model corresponding to the processing segment.

[0056] Specifically, since the scenes corresponding to different video segments are different, the subsequent scene models obtained by three-dimensionalization are also different. Here, for each video segment, take the video segment as the processing segment, and perform three-dimensionalization of the two-dimensional scene according to the image frames in this video segment to obtain a scene model corresponding to the video segment.

[0057] The present invention can adopt the method of establishing a three-dimensional scene model based on depth maps. Although this method is simple, the acquisition of depth maps requires on-site acquisition, which has high requirements for equipment and environment and is difficult. However, for multiple images that change within a continuous time, more information will be provided between images and images. Therefore, for video segments, the three-dimensionalization can be directly performed according to multiple image frames in the video segment to obtain a scene model.

[0058] This embodiment provides a pre-trained three-dimensional scene model. Take the video segment as the processing segment and input it into the three-dimensional scene model, and control the three-dimensional scene model to perform three-dimensionalization on the image frames in the input video segment to obtain a scene model corresponding to the video segment. The training process of the three-dimensional scene model includes:

[0059] A10. Obtain a training video for the training model.

[0060] Specifically, pre-obtain a training video for a preset training model. The training video can be obtained by shooting real objects, can also be made into a two-dimensional video according to the training model, or can be a training video formed by a combination of a large number of images with known camera parameters.

[0061] A20. Perform data conversion on the training frames in the training video to obtain a five-dimensional function, where the five-dimensional function includes training coordinates and observation vectors.

[0062] Specifically, for each image frame in the training video as a training frame, perform data conversion on it to convert it into a five-dimensional function. The five-dimensional function includes the training coordinates corresponding to the image frame in space and also includes observation vectors, where the observation vectors include observation angles and starting coordinates of the observation.

[0063] A30. Input the five-dimensional vector into a preset three-dimensional prediction model to obtain the voxel color and voxel density corresponding to the five-dimensional feature.

[0064] Specifically, input the five-dimensional feature vector into the three-dimensional prediction model. The three-dimensional prediction model can convert it into the attribute values of voxels in the three-dimensional model, such as voxel color and voxel density.

[0065] It can be expressed by the formula F Θ : (x, d) → (c, σ)

[0066] where x = {x, y, z} represents three-dimensional coordinates; d = {θ, φ} represents the two-dimensional observation vector; c = {r, g, b}, representing the color of the voxel related to the viewing angle; σ represents the voxel density. The three-dimensional prediction model can adopt an MLP network.

[0067] A40. Render all the voxel colors and the voxel densities to obtain a prediction model.

[0068] Specifically, during the three-dimensional model modeling process, knowing the voxel color and voxel density can realize the rendering and output of the three-dimensional model to obtain a prediction model.

[0069] A50. Calculate the loss value between the prediction model and the training model based on a preset loss function.

[0070] Specifically, after obtaining the prediction model, to evaluate the accuracy of the three-dimensional prediction model, calculate the loss value between the prediction model and the training model based on a preset loss function, which is the value indicating that the three-dimensional prediction model is inaccurate.

[0071] A60. Train the three-dimensional prediction model based on the loss value until the three-dimensional prediction model converges.

[0072] Specifically, the loss value is then backpropagated into the three-dimensional prediction model to adjust the parameters within the three-dimensional prediction model. The processes of training, loss calculation, and adjustment are repeated until the three-dimensional prediction model meets the preset convergence conditions, achieving model convergence. The preset convergence conditions may include that the accuracy of the three-dimensional prediction model reaches a threshold, or the number of training times reaches a target number, etc.

[0073] S30. When a special effect addition instruction for a preset item element is detected, the video file is processed according to the scene model and the special effect addition instruction to obtain a special effect video.

[0074] Specifically, a number of item elements are preset in advance. The item elements can be pre-designed by the designer or manually added by the user. The item elements may include parameters such as the shape, size, and color of various items such as vases and balls.

[0075] When the user needs to add an item element to the video file, the item element to be added and the indicated position of the item element to be added can be selected through an external device, thereby generating a special effect addition instruction for sending to the server. In one way of generating a special effect addition instruction in this embodiment, as Figure 3 shown, on the display interface connected to the server, an image frame in the video file is displayed on the left, and the item elements that can be added to the video file are listed on the right. The user can directly drag the item element on the right to the displayed image frame and then release the mouse. The coordinates when the mouse is released are the indicated coordinates, that is, the coordinates where the user expects to add the item element.

[0076] Since the corresponding scene model is created for different video segments, when the item element and the indicated position are obtained, the item element is moved onto the scene model according to the indicated coordinates. In this processing segment, each image frame can be regarded as being projected from the scene model from a specific perspective. Therefore, after the item element is moved onto the scene model, based on the angle of the scene model projected by each image frame and the scene model after adding the item element, the two-dimensional information corresponding to the item element is added to the image frame, thereby obtaining the special effect image corresponding to the image frame. After all the image frames are converted into special effect images with added item elements, a special effect video with added item elements for the entire video segment is obtained.

[0077] Further, if the indicated coordinates are coordinates determined for the world coordinates of the scene model, then based on the indicated coordinates, the unique position corresponding to the item element can be determined to add the item element. However, since the displayed image is two-dimensional, it is not easy for the user to determine the position they expect to determine, and the specified coordinates are mostly two-dimensional coordinates. And during the video progress, the viewing angle is often adjusted. Therefore, it is not stable to determine the position in the scene model only by two-dimensional coordinates. Therefore, in this embodiment, the items in the scene model are used as targets to fix the positions of the item elements. A process of processing a video segment in this embodiment is as follows:

[0078] B10. Perform object recognition on the scene model to obtain a number of tracking objects;

[0079] Specifically, since the scene model is modeled by voxels and not modeled according to the existing items, after obtaining the scene model, first perform item recognition on the scene model to identify the items in the scene model, and make these items that are already in the scene model become tracking objects. Common final objects such as walls, tables, people, and chairs.

[0080] B20. When detecting a special effect addition instruction corresponding to a preset item element, determine the tracking element and the insertion coordinates in the tracking object according to the indicated coordinates in the special effect addition instruction.

[0081] Specifically, when detecting a special effect addition instruction corresponding to an item element, according to the coordinates corresponding to the indicated coordinates on the displayed image frame, the two-dimensional coordinates where the user hopes to add the item element on the image frame can be determined. According to the two-dimensional coordinates, determine the tracking element in the tracking object. As Figure 3 shown, in this embodiment, if the item corresponding to the indicated coordinates input by the user in the image frame is the sky, then use the sky as the tracking element.

[0082] Based on the fixation of the tracking element and the indication coordinates, the item element can relatively stably determine its fixed position in the scene model. When there is a change in the viewing angle in the video clip, for example, adjusted from the 11 o'clock direction to the 10 o'clock direction, the projection of the item element in each image frame is relatively fixed. For example, if the tracking element selected in the previous text is the sky and the item element is the sun, then the position where the sun is projected in the viewing angle of image frame 1 is in the area of the sky, and the position where it is projected in the viewing angle of image frame 2 is also in the area of the sky. This method can better improve the stability of the insertion position of the item element. Another example is that the item element is placed on the surface of a certain plane, such as the surface of a wall. Taking the surface of the wall as a fixed surface, based on the function and indication coordinates of this fixed surface in the world coordinate system, the item element can be fixed at a unique coordinate. In this embodiment, this coordinate is used as the insertion coordinate. In addition, the user can also send a correction instruction to adjust the insertion coordinate corresponding to the item element to fix the position of the item element.

[0083] B30. Determine the processing parameters corresponding to the item element according to the scene model, the tracking element, and the indication coordinates.

[0084] Specifically, according to the scene model, the tracking element, and the indication coordinates, the item element can determine a relatively stable position coordinate, and use the display parameters corresponding to this insertion coordinate, such as light, projection, etc., as the processing parameters corresponding to the item element. In this embodiment, taking the processing parameters including perspective parameters and projection parameters as an example, the perspective parameter is the perspective relationship between the item element and the scene model, and the projection parameter is the light and shade distribution displayed by the item element under the light of the scene model. Therefore, the perspective parameters corresponding to the item element can be determined according to the perspective information corresponding to the tracking element and the scene model. At the same time, the projection parameters corresponding to the item element are determined according to the light information of the tracking element in the scene model.

[0085] When determining the projection parameters, since the item element depends on the tracking element, the light surface corresponding to the tracking element can be determined first according to the light source distribution information in the scene model. The light surface refers to the division according to the different degrees of light reception, generally including the light-receiving surface, the side-light surface, and the back-light surface. Taking Figure 4 as an example, the light source in the scene model is distributed in the upper right corner, so the light-receiving surface (the surface marked 3 in the figure), the side-light surface (the surface marked 2 in the figure), and the back-light surface (the surface marked 1 in the figure) in the tracking element can be determined.

[0086] According to the brightness values of the light-receiving surface, side-light surface, and backlight surface in the tracking element and the brightness value of the light source in the scene model, calculate the light propagation function and the brightness transfer function. The light propagation function is a function representing the light propagation path from the light source to the tracking element; the brightness transfer function is a function of the change in the brightness values of the light-receiving surface, side-light surface, and backlight surface during the process of light propagating to the tracking element. Finally, according to the insertion coordinates corresponding to the item element, insert the item element into the scene model, and calculate the brightness presented on different surfaces of the item element in the scene model according to the light source transfer function and the brightness transfer function, that is, the projection parameters.

[0087] Therefore, according to the scene model, the tracking element, and the indication coordinates, the item element can determine a relatively stable position coordinate, and use the display parameters corresponding to this position coordinate, such as light, projection, etc., as the processing parameters corresponding to the item element.

[0088] B40. Process the video file according to the insertion coordinates, the processing parameters, and the item element to obtain a special effect video.

[0089] Specifically, after obtaining the processing parameters corresponding to the item element, the item element can be inserted into the scene model more realistically and appropriately. According to the perspective information, the scene model inserted with the item element and the mirror element can be projected again to obtain a projection image, and this projection image replaces the image frame. However, in this way, phenomena such as item movement appear in the video clip, and the new projection image cannot well retain the information of the original image frame. Therefore, in this embodiment, according to the perspective corresponding to different image frames of the video clip, the item element is projected and inserted into this image frame. After processing each image frame, a special effect video is obtained.

[0090] Furthermore, if the scene model includes mirror objects, such as mirrors and lakes, when adding an item element to the scene model, in the actual scene, a corresponding mirror image will appear on the mirror. Therefore, in this embodiment, when the scene model contains mirror objects, the mirrors in the scene model need to be processed.

[0091] First, when it is detected that there are mirror objects in the scene model, according to the insertion coordinates corresponding to the item element and the world coordinates corresponding to the mirror object, generate a mirror element corresponding to the item element, as well as the mirror coordinates and mirror parameters corresponding to the mirror element. The mirror parameters refer to parameters similar to the processing parameters corresponding to the item element, and may include projection parameters, perspective parameters, etc. Take the item element, insertion coordinates, and processing parameters as item information, and take the mirror element, mirror coordinates, and mirror parameters as mirror information. Process the video file according to the mirror information and the item information to obtain a special effect video. The processing method for the image frames in the video file has been described above, so it will not be elaborated one by one.

[0092] Furthermore, there may be a situation where scenes are shared among video segments. For example, video segment 1 is for scene A, video segment 2 is for scene B, and video segment 3 is still for scene A. If the user only inserts an item element for video segment 1, video segment 3 should also be processed with special effects. Therefore, after processing the video segment, it further includes:

[0093] C10. Compare the model to be processed corresponding to the segment to be processed with the scene model to determine the comparison model corresponding to the scene model.

[0094] Specifically, the segment to be processed is the video segment other than the video segment selected by the user for processing among all video segments. The model to be processed is the model obtained by three-dimensionalizing the segment to be processed.

[0095] Both the model to be processed and the scene model are three-dimensional models, so they can be compared to determine the model to be processed that is more similar to the scene model as its corresponding comparison model. For the comparison of three-dimensional models, coordinate system normalization can be used first to convert the three-dimensional models into models in the same coordinate system, and then methods such as appearance comparison and geometric similarity can be used to compare the model similarity between the scene model and the model to be processed in the same coordinate system. Then select the model to be processed whose model similarity meets the preset threshold as the comparison model corresponding to the scene model.

[0096] If geometric similarity is used to compare the similarity between the model to be processed and the scene model, since multiple dimensions such as brightness, color, and topological structure are used in geometric similarity to evaluate the model similarity between the two models, when calculating the model similarity, different weights can be set for different dimensions. For each dimension, calculate the product of the single-dimension similarity under this dimension and the corresponding weight, and then sum all the single-dimension similarities with added weights to obtain the model similarity. For example, in video shooting, as time goes by, the brightness will change, so the weight value corresponding to the brightness dimension is relatively low, for example, less than 30%, while the color is relatively stable, so the weight value corresponding to this dimension is relatively high, for example, greater than 50%.

[0097] C20. Generate a special effect supplement instruction corresponding to the comparison model according to the special effect addition instruction.

[0098] Specifically, the special effect addition instruction is for the scene model. Since the three-dimensional models corresponding to different video segments are different, it is necessary to convert the special effect addition instruction into an instruction for the comparison model. First, establish a conversion function between the coordinate systems of the scene model and the comparison model, and then based on the conversion function, convert the special effect addition instruction to obtain the special effect supplement instruction.

[0099] C30. Process the segment to be processed corresponding to the special effect supplement instruction to obtain a special effect supplement video; and, process the processed segment according to the insertion coordinates, the processing parameters, and the item element to obtain a special effect processed video.

[0100] Specifically, finally, based on the special effect supplement instruction, process the segment to be processed corresponding to the comparison model to obtain a special effect supplement video with item elements added in the segment to be processed. At the same time, process the processed segment according to the insertion coordinates, the processing parameters, and the item element to obtain a special effect processed video. Since this process is the same as the process for processing video segments described above, it will not be elaborated here. Replace the position of the special effect processed video with the processed segment in the video file, and replace the position of the special effect supplement video with the segment to be processed in the video file to obtain the processed special effect video.

[0101] The present invention creates a scene model for a video file. When a user needs to add an item element, simulate it on the scene model, and then change the item element from three - dimensional to two - dimensional according to the angle of image frame projection to obtain its two - dimensional information and add it to the image frame. In this process, the user only needs to determine the position to be added and the item element to be added, which reduces the difficulty of the entire video special effect processing and improves the processing speed. In addition, during the implementation of the present invention, tracking elements are also combined to improve the stability of the position of the item element, the influence of mirror objects on video special effects, and the processing method when different video segment scenes overlap, improving the processing accuracy and further reducing the processing threshold.

[0102] Based on the above video processing method, the present invention also provides a terminal device, as Figure 3 shown, which includes at least one processor 20; a display screen 21; and a memory 22. It may also include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can communicate with each other through the bus 24. The display screen 21 is set to display a preset user - guiding interface in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical commands in the memory 22 to execute the method in the above - mentioned embodiments.

[0103] In addition, when the logical commands in the above - mentioned memory 22 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer - readable storage medium.

[0104] The memory 22 is configured as a computer-readable storage medium and can be set to store software programs and computer-executable programs, such as program commands or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, commands or modules stored in the memory 22, that is, implements the methods in the above embodiments.

[0105] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include high-speed random access memory and may also include non-volatile memory. For example, various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes may also be a transient computer-readable storage medium.

[0106] In addition, the specific processes of loading and executing multiple commands in the above computer-readable storage medium and the terminal device have been described in detail in the above methods and will not be repeated here.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video processing method, characterized in that, The method includes: Obtaining a video file, where the video file includes a plurality of video segments; For each of the video segments, using the video segment as a processing segment, and performing three-dimensionalization of a two-dimensional scene based on the image frames in the processing segment to obtain a scene model corresponding to the processing segment; When a special effect addition instruction corresponding to a preset item element is detected, processing the video file according to the scene model and the special effect addition instruction to obtain a special effect video; The step of when a special effect addition instruction corresponding to a preset item element is detected, processing the video file according to the scene model and the special effect addition instruction to obtain a special effect video includes: Performing object recognition on the scene model to obtain a plurality of tracking objects; When a special effect addition instruction corresponding to a preset item element is detected, determining a tracking element and an insertion coordinate in the tracking objects according to the indication coordinates in the special effect addition instruction; Determining processing parameters corresponding to the item element according to the scene model, the tracking element, and the indication coordinates; Processing the video file according to the insertion coordinate, the processing parameters, and the item element to obtain a special effect video; The special effect video includes a special effect processing video and a special effect supplementary video; the step of processing the video file according to the insertion coordinate, the processing parameters, and the item element to obtain a special effect video includes: Comparing a to-be-processed model corresponding to a to-be-processed segment with the scene model to determine a comparison model corresponding to the scene model; Generating a special effect supplementary instruction corresponding to the comparison model according to the special effect addition instruction; Processing the to-be-processed segment corresponding to the comparison model according to the special effect supplementary instruction to obtain a special effect supplementary video; and Processing the processing segment according to the insertion coordinate, the processing parameters, and the item element to obtain a special effect processing video.

2. The video processing method according to claim 1, characterized in that, The step of for each of the video segments, using the video segment as a processing segment, and performing three-dimensionalization of a two-dimensional scene based on the image frames in the processing segment to obtain a scene model corresponding to the processing segment includes: For each of the video segments, using the video segment as a processing segment, inputting the processing segment into a trained three-dimensional scene model, and controlling the three-dimensional scene model to perform three-dimensionalization on the image frames in the processing segment to obtain a scene model corresponding to the video segment.

3. The video processing method according to claim 2, characterized in that, Before the step of for each of the video segments, using the video segment as a processing segment, and performing three-dimensionalization of a two-dimensional scene based on the image frames in the processing segment to obtain a scene model corresponding to the processing segment, it further includes: Obtaining a training video for a training model; Performing data conversion on training frames in the training video to obtain a five-dimensional function, where the five-dimensional function includes position coordinates and observation vectors; Inputting the five-dimensional function into a preset three-dimensional prediction model to obtain voxel colors and voxel densities corresponding to the five-dimensional features; Rendering all the voxel colors and the voxel densities to obtain a prediction model; Calculate a loss value between the prediction model and the training model based on a preset loss function; Train the three-dimensional prediction model based on the loss value until the three-dimensional prediction model converges.

4. The video processing method according to claim 1, characterized in that, The processing parameters include perspective parameters and projection parameters; the determining of the processing parameters corresponding to the item element according to the scene model, the tracking element, and the indication coordinates includes: Determine the perspective parameters corresponding to the item element according to the tracking element and the perspective information corresponding to the scene model; and, Determine the projection parameters corresponding to the item element according to the light information of the tracking element in the scene model.

5. The video processing method according to claim 4, characterized in that, The determining of the projection parameters corresponding to the item element according to the light information of the tracking element in the scene model includes: Determine a light surface corresponding to the tracking element according to the light source distribution information in the scene model, where the light surface includes a light-receiving surface, a side-light surface, and a backlight surface; Calculate a light propagation function and a brightness transfer function according to the brightness values of the light-receiving surface, the side-light surface, and the backlight surface in the tracking element, and the brightness value of the light source in the scene model; Calculate the projection parameters corresponding to the item element according to the indication coordinates, the light propagation function, and the brightness transfer function.

6. The video processing method according to claim 1, characterized in that, The special effect video includes a plurality of special effect images; the processing of the video file according to the insertion coordinates, the processing parameters, and the item element to obtain a special effect video includes: When the scene model includes a mirror object, generate a mirror element corresponding to the item element, mirror coordinates corresponding to the mirror element, and mirror parameters according to the insertion coordinates and the world coordinates of the mirror object; Process the video file according to the mirror information and the item information to obtain a special effect video, where the mirror information includes the mirror element, the mirror coordinates, and the mirror parameters, and the item information includes the item element, the insertion coordinates, and the processing parameters.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the video processing method according to any one of claims 1 to 6.

8. A terminal device, characterized in that, Including: A processor, a memory, and a communication bus; A computer-readable program executable by the processor is stored on the memory; The communication bus realizes connection communication between the processor and the memory; When the processor executes the computer-readable program, it implements the steps in the video processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video picture processing method, electronic terminal and device

    CN109089058A

  • Video processing method and device and electronic equipment

    CN110650368A

  • Voxel model and image generation method and device, and storage medium

    CN114119838A