Video processing method, video processing device and electronic device

By analyzing the source video to generate special-effect videos with three-dimensional dynamic objects consistent with the scene, the problem of insufficient attractiveness of existing video advertising forms is solved, and more attractive and interactive advertising implantation is achieved, improving the user experience.

CN113298926BActive Publication Date: 2025-07-04ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010081332.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-02-06
Publication Date
2025-07-04
Estimated Expiration
2040-02-06

AI Technical Summary

Technical Problem

Existing video advertising forms, such as patch advertising and hard advertising, are unattractive and affect the user experience, while soft-implanted advertising is not easily attracting user attention.

Method used

By analyzing the scene of the source video, a special effect video containing three-dimensional dynamic objects is generated to keep it consistent with the scene of the source video, including spatial structure, light and shadow conditions, occlusion interaction relationships and camera motion information, to implant three-dimensional dynamic objects that match the scene.

Benefits of technology

It improves user attraction, and the three-dimensional dynamic objects interact effectively with the source video scene, does not affect the user's viewing experience, expands the implanted space, and enhances the exposure of advertisements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113298926B_ABST
    Figure CN113298926B_ABST
Patent Text Reader

Abstract

A video processing method, a video processing device, and an electronic device are disclosed. The video processing method includes: obtaining a source video; parsing the scene of the source video to obtain video information related to the scene of the source video where a three-dimensional dynamic object is to be implanted; and generating a special effect video including the three-dimensional dynamic object based on the video information. In this way, the attractiveness to users can be enhanced through the three-dimensional dynamic object, and at the same time, the viewing experience of users will not be affected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and more specifically, to a video processing method, a video processing device, and an electronic device. Background Art

[0002] Currently, there are two common forms of video advertisements. The first is the patch advertisements around the played video, for example, the patch advertisements on the left and right sides of the video. This form of video advertisement is monotonous and lacks attraction to users. The other is the hard advertisements inserted at the front, middle, and rear of the video. This form of video advertisement will pause the original video playback, causing great interruption to the user's viewing and poor experience.

[0003] In contrast, soft-implanted advertisements replace some objects displayed in the video with advertisement objects, so that the advertisement objects are not displayed separately from the video content, which can give users a better experience. However, although soft-implanted advertisements reduce the interruption to users' video viewing, they are not easily noticed by users themselves, thus reducing the attractiveness to users.

[0004] Therefore, it is desirable to provide an improved video processing solution. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a video processing method, a video processing device, and an electronic device, which can generate a special effect video containing three-dimensional dynamic objects consistent with the scene of the source video by parsing the scene of the source video, thereby enhancing the attractiveness to users through the three-dimensional dynamic objects without affecting the user's viewing experience.

[0006] According to one aspect of this application, a video processing method is provided, including: obtaining a source video; parsing the scene of the source video to obtain video information related to the scene of the source video where a three-dimensional dynamic object is to be implanted; and generating a special effect video containing the three-dimensional dynamic object based on the video information.

[0007] In the above video processing method, parsing the scene of the source video to obtain video information related to the scene of the source video where a three-dimensional dynamic object is to be implanted includes: performing three-dimensional reconstruction on the scene of the source video to obtain spatial structure information of the scene of the source video where a three-dimensional dynamic object is to be implanted; and / or performing light estimation on the scene of the source video to obtain light and shadow condition information of the scene of the source video where a three-dimensional dynamic object is to be implanted.

[0008] In the above video processing method, generating a special effect video from the three-dimensional dynamic object based on the video information includes: generating the special effect video based on the spatial structure information and / or the light and shadow condition information, so that the three-dimensional dynamic object conforms to the spatial structure and / or the light and shadow condition of the scene.

[0009] In the above video processing method, parsing the scene of the source video to obtain video information related to the scene of the source video where the three-dimensional dynamic object is to be implanted includes: performing image segmentation on the scene of the source video to obtain mask information for representing the occlusion interaction relationship between the three-dimensional dynamic object and other subjects in the scene.

[0010] In the above video processing method, generating a special effect video from the three-dimensional dynamic object based on the video information includes: generating the special effect video based on the mask information, so that the three-dimensional dynamic object has transparency based on the occlusion interaction relationship.

[0011] In the above video processing method, parsing the scene of the source video to obtain video information related to the scene of the source video where the three-dimensional dynamic object is to be implanted includes: performing camera calibration on the scene of the source video to obtain the motion information of the camera for shooting the scene.

[0012] In the above video processing method, generating a special effect video from the three-dimensional dynamic object based on the video information includes: generating the special effect video based on the motion information of the camera, so that the motion of the three-dimensional dynamic object corresponds to the motion information of the camera.

[0013] In the above video processing method, after generating a special effect video including the three-dimensional dynamic object based on the video information, it further includes: synchronously playing the special effect video and the source video.

[0014] According to another aspect of the present application, there is provided a video processing apparatus, including: an acquisition unit for acquiring a source video; an analysis unit for parsing the scene of the source video acquired by the acquisition unit to obtain video information related to the scene of the source video where the three-dimensional dynamic object is to be implanted; and a generation unit for generating a special effect video including the three-dimensional dynamic object based on the video information obtained by the analysis unit.

[0015] In the above video processing apparatus, the analysis unit is configured to: perform three-dimensional reconstruction on the scene of the source video to obtain the spatial structure information of the scene of the source video where the three-dimensional dynamic object is to be implanted; and / or perform light estimation on the scene of the source video to obtain the light and shadow condition information of the scene of the source video where the three-dimensional dynamic object is to be implanted.

[0016] In the above video processing apparatus, the generating unit is configured to: generate the special effect video based on the spatial structure information and / or the light and shadow condition information, so that the three-dimensional dynamic object conforms to the spatial structure and / or the light and shadow condition of the scene.

[0017] In the above video processing method, the parsing unit is configured to: perform image segmentation on the scene of the source video to obtain mask information for representing the occlusion interaction relationship between the three-dimensional dynamic object and other subjects in the scene.

[0018] In the above video processing method, the generating unit is configured to: generate the special effect video based on the mask information, so that the three-dimensional dynamic object has transparency based on the occlusion interaction relationship.

[0019] In the above video processing method, the parsing unit is configured to: perform camera calibration on the scene of the source video to obtain the motion information of the camera for shooting the scene.

[0020] In the above video processing method, the generating unit is configured to: generate the special effect video based on the motion information of the camera, so that the motion of the three-dimensional dynamic object corresponds to the motion information of the camera.

[0021] In the above video processing apparatus, it further includes: a playing unit, configured to synchronously play the special effect video and the source video after generating the special effect video including the three-dimensional dynamic object based on the video information.

[0022] According to another aspect of the present application, there is provided an electronic device, including: a processor; and a memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the video processing method as described above.

[0023] According to yet another aspect of the present application, there is provided a computer-readable medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the video processing method as described above.

[0024] The video processing method, video processing apparatus, and electronic device provided by the present application can generate a special effect video including a three-dimensional dynamic object consistent with the scene of the source video by parsing the scene of the source video, thereby enhancing the attractiveness to users through the three-dimensional dynamic object without affecting the viewing experience of users. Description of the Drawings

[0025] The above and other objects, features, and advantages of the present application will become more apparent by describing the embodiments of the present application in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0026] Figure 1 The flowchart of the video processing method according to an embodiment of the present application is illustrated.

[0027] Figure 2 The flowchart of an application example of the video processing method according to an embodiment of the present application is illustrated.

[0028] Figure 3 The block diagram of the video processing apparatus according to an embodiment of the present application is illustrated.

[0029] Figure 4 The block diagram of the electronic device according to an embodiment of the present application is illustrated. Detailed implementation manners

[0030] Hereinafter, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0031] Application Overview

[0032] Currently, the most common soft-implanted advertisement is to place corresponding advertisement objects at the shooting site during video shooting. However, the subsequent flexibility of this method is not good. Therefore, the advertisement form of secondary implantation has been increasingly applied.

[0033] The advertisement form of secondary implantation is usually of two types: the first is to detect the implantable advertisement area in the video and synchronously superimpose the newly implanted advertisement in the form of an object at the same position, size, and shape in this area; the second is to detect the area in the video where the advertisement booth can be expanded (such as a blank rooftop, etc.), construct a new advertisement booth (such as a billboard, etc.), and superimpose the new advertisement booth together with the corresponding advertisement on the video.

[0034] However, the advertisement implantation areas of the above two schemes are usually non-main areas of the video and are not likely to attract the attention of users watching the video.

[0035] In addition, since the advertisement main bodies implanted in the above schemes, whether they are objects or images, are in a static form, it is not easy to attract the attention of users watching the video. In addition, the static advertisement main bodies cannot effectively interact with other main bodies in the video, further reducing the advertisement effect of the implanted advertisement.

[0036] In view of the above technical problems, the basic concept of the present application is to implant a three-dimensional dynamic object in a video, and by analyzing the scene of the source video, make the implanted three-dimensional dynamic object consistent with the scene of the source video, so as to enhance the attraction to users through the effective interaction between the three-dimensional dynamic object and the scene of the source video.

[0037] Specifically, the video processing method, video processing device and electronic device provided by the present application first obtain a source video, then analyze the scene of the source video to obtain video information related to the scene of the source video where the three-dimensional dynamic object is to be implanted, and finally generate a special effect video containing the three-dimensional dynamic object based on the video information.

[0038] Therefore, the video processing method, video processing device and electronic device provided by the present application can implant a dynamic three-dimensional object in the source video and make the implanted three-dimensional dynamic object conform to the picture of the original scene of the source video. In this way, the three-dimensional dynamic object itself is more likely to attract the attention of users watching the source video compared to static objects.

[0039] In addition, compared with static objects, the three-dimensional dynamic object can also effectively interact with the scene of the source video. For example, it can effectively interact with other subjects existing in the source video, thereby further enhancing the attraction to users.

[0040] Moreover, by analyzing the scene of the source video to obtain relevant video information, the implanted three-dimensional dynamic object can be made to conform to the picture of the original scene of the source video, thus not affecting the user experience of watching the source video.

[0041] Furthermore, by making the implanted three-dimensional dynamic object conform to the picture of the original scene of the source video, the scene for implanting the three-dimensional dynamic object in the source video can be expanded, not limited to static positions suitable for implanting static objects, etc.

[0042] It should be noted that in the video processing method, video processing device and electronic device provided by the present application, the three-dimensional dynamic object can be any implanted object implanted in the source video for advertising purposes, not limited to soft implants, and can also be other implanted objects in specific scenes of the video, such as implanted pictures, implanted characters, etc. By placing brand advertisements on this three-dimensional dynamic object, the attention of movie-watching users to the advertisement can be aroused, and it is more novel and interesting compared to static advertisement forms.

[0043] In addition, in the video processing method, video processing device, and electronic device provided in this application, the three-dimensional dynamic object is not limited to an object implanted in the source video for advertising purposes, but can also be other objects in the source video that the user is expected to notice. For example, it can be a specific logo or the like in the source video that the user is expected to notice. Alternatively, the three-dimensional dynamic object can also be an object related to the content of the scene in the source video, such as a person or an item.

[0044] After introducing the basic principle of this application, various non-limiting embodiments of this application will be specifically introduced below with reference to the accompanying drawings.

[0045] Exemplary Method

[0046] Figure 1 The flowchart of the video processing method according to an embodiment of this application is illustrated.

[0047] As Figure 1 shown, the video processing method according to an embodiment of this application includes the following steps.

[0048] S110, Obtain a source video. The source video can be any video that the user watches, such as a movie, a TV drama, a video clip, etc.

[0049] S120, Parse the scene of the source video to obtain video information related to the scene of the source video where the three-dimensional dynamic object is to be implanted. The video information refers to the video information that enables the three-dimensional dynamic object implanted in the source video to conform to the picture of the scene in the source video where the three-dimensional dynamic object is implanted.

[0050] According to the differences in the scene, the video information may include different types of information, such as the spatial structure information of the scene, the light and shadow condition information, and the movement information of the camera that shoots the scene, etc.

[0051] S130, Generate a special effect video including the three-dimensional dynamic object based on the video information. That is, after obtaining the video information in step S120, perform processing on the display effect of the three-dimensional dynamic object based on the video information, so that when generating the special effect video including the three-dimensional dynamic object, the three-dimensional dynamic object presented in the special effect video can be consistent with the scene of the source video, so as not to affect the user's viewing of the source video.

[0052] Next, various different types of video information and the processing of the three-dimensional dynamic object based on the video information will be further described in detail.

[0053] Since the scenes in the source video may have various three-dimensional spatial structures, such as streets, walls, etc., and the three-dimensional dynamic objects implanted in the scenes need to conform to the spatial structures of the scenes where they are located. For example, a person needs to walk along a street and avoid walls, etc. Therefore, by performing three-dimensional reconstruction on the scenes of the source video, the spatial structure information of the scenes can be obtained. Alternatively, the spatial structure information of the scenes can also be obtained by manually constructing a three-dimensional rough model of the scenes.

[0054] In this way, based on the obtained spatial structure information of the scenes, the three-dimensional dynamic objects need to be processed to make them conform to the spatial structures of the scenes.

[0055] In addition, the scenes in the source video may also have various light and shadow effects. For example, in the presence of a light source, various objects in the source video have light and shadow effects. For example, the side facing the light source has a lighting effect, while the side facing away from the light source has a shadow effect. Therefore, when implanting three-dimensional dynamic objects into the scenes, they need to be consistent with the light and shadow conditions of the scenes.

[0056] Specifically, the environmental high-dynamic range image of the scene can be obtained through light estimation. Compared with ordinary images, it has more environmental image details, so that the accurate light and shadow conditions of the scene can be reflected. Then, by processing the three-dimensional dynamic objects, they can be made to conform to the light and shadow conditions.

[0057] That is, in the video processing method according to the embodiments of the present application, the video information obtained by parsing the scenes of the source video related to the scenes of the source video where three-dimensional dynamic objects are to be implanted includes: obtaining the spatial structure information of the scenes of the source video where three-dimensional dynamic objects are to be implanted by performing three-dimensional reconstruction on the scenes of the source video; and / or obtaining the light and shadow condition information of the scenes of the source video where three-dimensional dynamic objects are to be implanted by performing light estimation on the scenes of the source video.

[0058] Moreover, in the above video processing method, generating a special effect video from the three-dimensional dynamic objects based on the video information includes: generating the special effect video based on the spatial structure information and / or light and shadow condition information, so that the three-dimensional dynamic objects conform to the spatial structures and / or light and shadow conditions of the scenes.

[0059] In the embodiments of the present application, the video information further includes the occlusion interaction relationship between the three-dimensional dynamic object and other entities in the scene. Since there are already many occlusion relationships in the video frame itself. For example, for a table placed in front of a wall, the table will occlude the wall. In this case, if a three-dimensional dynamic object is newly implanted, new occlusion relationships will be generated for both the table and the wall. For example, if the newly implanted three-dimensional dynamic object is behind the table, the table will occlude the three-dimensional dynamic object, and at the same time, the three-dimensional dynamic object will also occlude the wall.

[0060] Regarding the occlusion interaction relationship between the three-dimensional dynamic object and other entities in the scene, in the embodiments of the present application, mask information is used to represent this occlusion interaction relationship. For example, for the occlusion interaction relationship between the three-dimensional dynamic object and the table and the wall as described above, the table can be represented by a blue mask, and the wall can be represented by a red mask. Then, when the three-dimensional dynamic object moves into the range of the area of the blue mask, it means that the three-dimensional dynamic object is about to be occluded by the table, and the pixels of the three-dimensional dynamic object located in the blue area can be processed as transparent, so as to represent that the three-dimensional dynamic object is occluded by the table.

[0061] That is, in the embodiments of the present application, since the special effect video containing the three-dimensional dynamic object is played by being superimposed with the source video, in order to represent that the three-dimensional dynamic object is occluded by other entities in the source video, the occluded part of the three-dimensional dynamic object can be processed as transparent. In this way, when the special effect video containing the three-dimensional dynamic object is superimposed with the source video for playing, the audience will feel that the three-dimensional dynamic object is occluded by other entities.

[0062] Therefore, through the mask information used to represent the occlusion interaction relationship between the three-dimensional dynamic object and other entities in the scene, the transparency of the three-dimensional dynamic object can be processed, so as to present the effect of the occlusion interaction relationship between the three-dimensional dynamic object and other entities in the scene to the audience through the transparency of the three-dimensional dynamic object based on the occlusion interaction relationship.

[0063] Therefore, in the video processing method according to the embodiments of the present application, the video information obtained by parsing the scene of the source video to be related to the scene of the source video where the three-dimensional dynamic object is to be implanted includes: obtaining mask information for representing the occlusion interaction relationship between the three-dimensional dynamic object and other entities in the scene by performing image segmentation on the scene of the source video.

[0064] Moreover, in the above video processing method, generating a special effect video from the three-dimensional dynamic object based on the video information includes: generating the special effect video based on the occlusion interaction relationship between the three-dimensional dynamic object and other entities in the scene, so that the three-dimensional dynamic object has transparency based on the occlusion interaction relationship.

[0065] In the source video, it is very likely that there is a situation where the perspective of the scene itself moves, that is, the camera shooting the scene is in a moving state. In this case, when implanting a three-dimensional dynamic object, if the motion state of the camera is not considered, the three-dimensional dynamic object will be distorted. For example, if in a certain scene, the camera shooting the scene is in a state of moving to the right, then the scene itself will have a relative leftward movement. Correspondingly, when implanting the three-dimensional dynamic object into the scene, the three-dimensional dynamic object should also have a corresponding leftward movement state.

[0066] In addition, the relative position of the camera shooting the scene will also affect the shooting angle of the scene, thereby affecting the overall orientation of the scene. Therefore, it is also necessary to determine the relative position of the camera to determine the orientation of the three-dimensional dynamic object implanted in the scene.

[0067] Therefore, in order to obtain the motion information of the camera, that is, the position and motion trajectory of the camera, the scene can be calibrated for the camera to reverse the position and motion trajectory of the camera from the scene itself, and then the three-dimensional dynamic object can be processed based on the position and motion trajectory of the camera, so that the motion of the three-dimensional dynamic object conforms to the motion of the camera.

[0068] Therefore, in the video processing method according to the embodiments of the present application, parsing the scene of the source video to obtain video information related to the scene of the source video for implanting a three-dimensional dynamic object includes: calibrating the camera for the scene of the source video to obtain the motion information of the camera for shooting the scene.

[0069] Moreover, in the above video processing method, generating a special effect video from the three-dimensional dynamic object based on the video information includes: generating the special effect video based on the motion information of the camera, so that the motion of the three-dimensional dynamic object corresponds to the motion information of the camera.

[0070] In addition, in the embodiments of the present application, according to the characteristics of the scene of the source video, in addition to the above-mentioned spatial structure information, light and shadow condition information, occlusion interaction relationship information, and camera movement information, other video information may also need to be obtained by parsing the scene. For example, when there is a reflective object (such as a mirror) in the scene, the three-dimensional moving object will generate a mirror image for the reflective object. Therefore, it is necessary to perform planar texture detection processing on the scene to obtain the mirror texture of the three-dimensional dynamic object, so that the three-dimensional dynamic object in the special effect video has the mirror planar texture characteristics conforming to the reflective object.

[0071] In this way, through the video processing method according to the embodiments of the present application as described above, a special effect video containing the three-dimensional dynamic object is obtained. Then, by synchronously playing the special effect video with the source video, when the user watches the source video, they can see the special effect video containing the three-dimensional dynamic object.

[0072] That is to say, in the video processing method according to the embodiments of the present application, after generating a special effect video containing the three-dimensional dynamic object based on the video information, it further includes: synchronously playing the special effect video with the source video.

[0073] In summary, the video processing method according to the embodiments of the present application implants a three-dimensional dynamic object into the source video, which is more likely to attract the user's attention than a static three-dimensional object, and can effectively interact with other subjects in the source video, improving the user's attention.

[0074] In addition, by parsing the spatial structure, light and shadow conditions, etc. of the scene of the source video, the implanted three-dimensional dynamic object fits the scene of the source video and will not cause a sense of incongruity, thus not affecting the user's viewing experience.

[0075] Moreover, since the implanted three-dimensional dynamic object can be dynamically integrated into the scene of the source video, the implanted space in the source video is expanded. For example, it is not limited to the blank static area where static objects or images can be implanted, but can include dynamic areas that interact with other subjects.

[0076] Therefore, when the video processing method according to the embodiments of the present application is applied to implant advertisements, not only the advertisement occupancy is expanded, but also a more attractive dynamic three-dimensional advertisement can be realized, enhancing the advertisement exposure.

[0077] Application Example

[0078] Figure 2 The flowchart of an application example of the video processing method according to the embodiments of the present application is illustrated.

[0079] As Figure 2As shown, in this application example, the video processing method according to an embodiment of the present application is used to embed advertisements in a video.

[0080] First, in step S210, the video clip to be implanted is obtained. Then, in step S220, the video scene is parsed, which includes: S221, using a segmentation algorithm to generate a mask video of the occlusion relationship; S222, using a 3D reconstruction algorithm or manually constructing a 3D rough model of the video scene; and S223, using a lighting estimation algorithm to perceive light and shadow / texture to generate an environmental HDR map.

[0081] Next, in step S230, the camera position and motion trajectory are inversely deduced using the camera calibration algorithm. Then, in step S240, the 3D model is superimposed, including: S241, performing spatial structure fitting processing according to the 3D structure data in step S222; and S242, performing light and shadow / environment fitting processing according to the light and shadow / texture data in step S223.

[0082] In step S250, the model animation is played and the original video is played at the same time for synchronous recording. Then, in step S260, special effect materials are generated, including S261, processing the occlusion relationship according to the mask video in step S221 to generate the implanted special effect material file. Finally, in step S270, advertisements are placed.

[0083] Thus, in this application example, due to the use of animated 3D models, it is interesting and easier to attract users' attention. Moreover, by using algorithms to analyze information such as the spatial structure of the video and light and shadow conditions, the implanted model fits the original video scene without any violation. In addition, by using the segmentation algorithm to output the occlusion relationship mask, the implanted model can have an occlusion interaction relationship with the subject in the video. Therefore, this application example not only expands the advertising booth, but also introduces more attractive dynamic 3D implants, enhancing the exposure of the advertisement.

[0084] Exemplary Device

[0085] Figure 3 The figure shows a block diagram of a video processing device according to an embodiment of the present application.

[0086] like Figure 3 As shown, the video processing device 300 according to the embodiment of the present application includes: an acquisition unit 310, used to acquire a source video; a parsing unit 320, used to parse the scene of the source video acquired by the acquisition unit 310 to obtain video information related to the scene of the source video to be implanted with a three-dimensional dynamic object; and a generation unit 330, used to generate a special effects video containing the three-dimensional dynamic object based on the video information obtained by the parsing unit 320.

[0087] In one example, in the above video processing device 300, the parsing unit 320 is configured to: obtain the spatial structure information of the scene of the source video for implanting the three-dimensional dynamic object by performing three-dimensional reconstruction on the scene of the source video; and / or, obtain the light and shadow condition information of the scene of the source video for implanting the three-dimensional dynamic object by performing light estimation on the scene of the source video.

[0088] In one example, in the above video processing device 300, the generating unit 330 is configured to: generate the special effect video based on the spatial structure information and / or the light and shadow condition information, so that the three-dimensional dynamic object conforms to the spatial structure and / or the light and shadow condition of the scene.

[0089] In one example, in the above video processing device 300, the parsing unit 320 is configured to: obtain the mask information for representing the occlusion interaction relationship between the three-dimensional dynamic object and other subjects in the scene by performing image segmentation on the scene of the source video.

[0090] In one example, in the above video processing device 300, the generating unit 330 is configured to: generate the special effect video based on the mask information, so that the three-dimensional dynamic object has transparency based on the occlusion interaction relationship.

[0091] In one example, in the above video processing device 300, the parsing unit 320 is configured to: obtain the motion information of the camera for shooting the scene by performing camera calibration on the scene of the source video.

[0092] In one example, in the above video processing device 300, the generating unit 330 is configured to: generate the special effect video based on the motion information of the camera, so that the motion of the three-dimensional dynamic object corresponds to the motion information of the camera.

[0093] In one example, in the above video processing device 300, it further includes: a playback unit, configured to synchronously play the special effect video and the source video after generating the special effect video including the three-dimensional dynamic object based on the video information.

[0094] Here, those skilled in the art can understand that the specific functions and operations of each unit and module in the above video processing device 300 have been introduced in detail in the description of the Figure 1 above video processing method, and therefore, the repeated description thereof will be omitted.

[0095] As described above, the video processing apparatus 300 according to an embodiment of the present application can be implemented in various terminal devices, such as a production system for advertisement implantation. In one example, the video processing apparatus 300 according to an embodiment of the present application can be integrated into a terminal device as a software module and / or a hardware module. For example, the video processing apparatus 300 can be a software module in the operating system of the terminal device, or can be an application developed for the terminal device; of course, the video processing apparatus 300 can also be one of many hardware modules of the terminal device.

[0096] Alternatively, in another example, the video processing apparatus 300 and the terminal device can also be separate devices, and the video processing apparatus 300 can be connected to the terminal device through a wired and / or wireless network, and transmit interaction information in accordance with a predefined data format.

[0097] Exemplary Electronic Device

[0098] Next, Figure 4 an electronic device according to an embodiment of the present application will be described.

[0099] Figure 4 A block diagram of an electronic device according to an embodiment of the present application is illustrated.

[0100] As Figure 4 shown, the electronic device 10 includes one or more processors 11 and a memory 12.

[0101] The processor 11 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 10 to perform desired functions.

[0102] The memory 12 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 11 can run the program instructions to implement the video processing methods of various embodiments of the present application described above and / or other desired functions. Various contents such as video information, models of three-dimensional dynamic objects, and special effect videos can also be stored in the computer-readable storage media.

[0103] In one example, the electronic device 10 may further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0104] The input device 13 may include, for example, a keyboard, a mouse, and the like.

[0105] The output device 14 may output various information to the outside, including the generated special effect video, etc. The output device 14 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and the like.

[0106] Of course, for simplicity, Figure 4 only some of the components related to the present application in the electronic device 10 are shown, and components such as a bus, an input / output interface, and the like are omitted. In addition, according to specific application scenarios, the electronic device 10 may further include any other appropriate components.

[0107] Exemplary Computer Program Product and Computer Readable Storage Medium

[0108] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the video processing method according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.

[0109] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on a first user computing device, partially on the first user device, executed as an independent software package, partially on the first user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0110] Furthermore, an embodiment of the present application may also be a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the video processing method according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.

[0111] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0112] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. Additionally, the above-disclosed specific details are only for illustrative and easy-to-understand purposes and are not limitations. The above details do not limit the present application to necessarily implement using the above specific details.

[0113] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with each other.

[0114] It should also be noted that in the devices, equipment, and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present application.

[0115] The above description of the disclosed aspects enables any person skilled in the art to make or use the present application. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0116] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit embodiments of the present application to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A video processing method, characterized in that, including: obtaining a source video; obtaining video information related to the scene of the source video for implanting a three-dimensional dynamic object by parsing the scene of the source video; and generating a special effect video including the three-dimensional dynamic object based on the video information; The obtaining video information related to the scene of the source video for implanting a three-dimensional dynamic object by parsing the scene of the source video includes: obtaining spatial structure information of the scene of the source video for implanting a three-dimensional dynamic object by performing three-dimensional reconstruction on the scene of the source video; Generating a special effect video from the three-dimensional dynamic object based on the video information includes: generating the special effect video based on the spatial structure information so that the three-dimensional dynamic object conforms to the spatial structure of the scene; The obtaining video information related to the scene of the source video for implanting a three-dimensional dynamic object by parsing the scene of the source video further includes: obtaining motion information of the camera for shooting the scene by performing camera calibration on the scene of the source video; Generating a special effect video from the three-dimensional dynamic object based on the video information further includes: generating the special effect video based on the motion information of the camera so that the motion of the three-dimensional dynamic object corresponds to the motion information of the camera.

2. The video processing method according to claim 1, wherein The obtaining video information related to the scene of the source video for implanting a three-dimensional dynamic object by parsing the scene of the source video further includes: obtaining light and shadow condition information of the scene of the source video for implanting a three-dimensional dynamic object by performing light estimation on the scene of the source video.

3. The video processing method according to claim 2, wherein Generating a special effect video from the three-dimensional dynamic object based on the video information includes: generating the special effect video based on the light and shadow condition information so that the three-dimensional dynamic object conforms to the light and shadow conditions of the scene.

4. The video processing method according to claim 1, wherein The obtaining video information related to the scene of the source video for implanting a three-dimensional dynamic object by parsing the scene of the source video includes: obtaining mask information for representing the occlusion interaction relationship between the three-dimensional dynamic object and other objects in the scene by performing image segmentation on the scene of the source video.

5. The video processing method according to claim 4, wherein Generating a special effect video from the three-dimensional dynamic object based on the video information includes: generating the special effect video based on the mask information so that the three-dimensional dynamic object has transparency based on the occlusion interaction relationship.

6. The video processing method according to claim 1, wherein After generating a special effect video including the three-dimensional dynamic object based on the video information, it further includes: synchronously playing the special effect video and the source video.

7. A video processing device, characterized in that, including: an obtaining unit for obtaining a source video; an analysis unit for obtaining video information related to the scene of the source video for implanting a three-dimensional dynamic object by parsing the scene of the source video obtained by the obtaining unit; and a generating unit for generating a special effect video including the three-dimensional dynamic object based on the video information obtained by the analysis unit; The analysis unit is specifically configured to obtain spatial structure information of the scene of the source video for implanting a three-dimensional dynamic object by performing three-dimensional reconstruction on the scene of the source video; A generating unit, specifically configured to generate the special effect video based on the spatial structure information, so that the three-dimensional dynamic object conforms to the spatial structure of the scene; An analysis unit, specifically further configured to obtain the motion information of the camera for shooting the scene by performing camera calibration on the scene of the source video; A generating unit, specifically further configured to generate the special effect video based on the motion information of the camera, so that the motion of the three-dimensional dynamic object corresponds to the motion information of the camera.

8. An electronic device, comprising: A processor; And A memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the video processing method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method for implanting advertisement in video

    CN103024480A

  • Method for intelligently implanting video contents based on Faster R-CNN model

    CN107493488A

  • Method and device for implanting push information into video and electronic equipment

    CN109842811A