Method and device for synthesizing special effects in virtual shooting, electronic device and storage medium

By using the motion capture system and inertial measurement unit to calculate the special effects posture in real time, the problem of low efficiency in the integration of virtual scenes and actors' performances in virtual shooting is solved, and efficient special effects addition and cost reduction are achieved.

CN119865566BActive Publication Date: 2025-09-23YOUKU CULTURE TECH (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510038488.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-09-23
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

In existing virtual shooting technologies, the integration of virtual scenes and actors' performances is inefficient, requiring directors and other personnel to communicate and adjust multiple times in the later stages, resulting in low efficiency.

Method used

During the virtual shooting process, the position and posture data of the actors and cameras are obtained through the motion capture system and inertial measurement unit, the posture of the special effects is calculated in real time, and the target special effects are added to the real scene captured by the camera to generate a composite picture.

Benefits of technology

It enables instant viewing of the combination of special effects and actors' performances on the virtual shooting scene, improves shooting efficiency, reduces the requirements for rendering performance, and reduces the cost of special effects production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119865566B_ABST
    Figure CN119865566B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and device for synthesizing special effects in virtual shooting, an electronic device, and a storage medium. The method comprises: in a virtual shooting process, obtaining a real-scene picture currently captured by a camera, tracking data currently captured by a motion capture system, and posture change data of a target object currently captured by an inertial measurement unit; determining target posture data corresponding to a target special effect based on the position data and posture change data of the target object; generating a special effects picture aligned with the real-scene picture based on the posture data of the camera, the target posture data corresponding to the target special effect, and the special effects data of the target special effect; synthesizing the special effects picture with the real-scene picture to obtain a synthesized picture, in which the target object in the synthesized picture has the target special effect. Thus, target special effects can be efficiently and automatically added to the real-scene picture during the virtual shooting process, so that the combined effect of the added special effects and the actor's performance can be instantly viewed at the scene of the virtual shooting, thereby improving shooting efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of virtual photography, and in particular to a method and device for synthesizing special effects in virtual photography, an electronic device, and a storage medium. Background Art

[0002] Virtual filming involves rendering a virtual scene in real time on a screen like a light-emitting diode (LED), with actors and other entities acting as the real scene. This allows cameras to capture both the real and virtual scenes simultaneously, creating the illusion of placing the actors in a virtual environment. This technology, combining real-time virtual image rendering with actor performance, brings a whole new experience to film and television production.

[0003] Special effects such as "spells" are usually added to TV dramas and movies with themes such as fairy tales, magic, and science fiction. In the relevant technology, the virtual scene and the real scene of the actors' performance are integrated during the virtual shooting process. After the virtual shooting is completed, the "spells" and other special effects are post-produced based on the opinions of the director and other personnel and the content of the script. This method usually requires the director and other personnel to communicate and adjust multiple times during the post-production process, which is less efficient. Summary of the Invention

[0004] In view of this, the present disclosure proposes a method and device for synthesizing special effects in virtual shooting, an electronic device and a storage medium, which can efficiently and automatically add targeted special effects to the real-scene pictures taken by the camera during the virtual shooting process, so that the combined effect of the added special effects and the actor's performance can be viewed instantly at the virtual shooting scene, thereby improving shooting efficiency.

[0005] According to one aspect of the present disclosure, a method for synthesizing special effects in virtual shooting is provided, comprising: in a virtual shooting process, obtaining a real-scene picture currently captured by a camera, tracking data currently captured by a motion capture system, and posture change data of a target object to which special effects are to be added currently captured by an inertial measurement unit, wherein the real-scene picture comprises a virtual scene displayed on a screen and a real scene in front of the screen, and the tracking data comprises the posture data of the camera and the position data of the target object; determining target posture data corresponding to a target special effect to be added based on the position data of the target object and the posture change data; generating a special effects picture aligned with the real-scene picture based on the posture data of the camera, the target posture data corresponding to the target special effect, and the special effects data of the target special effect; synthesizing the special effects picture with the real-scene picture to obtain a synthesized picture, in which the target object has the target special effect.

[0006] In one possible implementation, the posture change data includes the acceleration and angular velocity of the target object from the last acquisition moment to the current acquisition moment, and the position data of the target object includes the position data of the tracking point set on the target object in the global coordinate system; wherein, determining the target posture data corresponding to the target special effect to be added based on the position data of the target object and the posture change data includes: determining the initial posture data corresponding to the target special effect at the current acquisition moment based on the target posture data corresponding to the target special effect at the last acquisition moment and the posture change data; wherein, the target posture data corresponding to the target special effect at the last acquisition moment is determined based on the position data and posture change data of the target object collected at the last acquisition moment; based on the position data of the tracking point set on the target object in the global coordinate system and the position data of the tracking point set on the target object in the local coordinate system of the target object, the initial posture data corresponding to the target special effect at the current acquisition moment is optimized to obtain the target posture data corresponding to the target special effect at the current acquisition moment.

[0007] In one possible implementation, the special effects picture aligned with the real scene picture is generated based on the posture data of the camera, the target posture data corresponding to the target special effect, and the special effects data of the target special effect, including: determining the relative posture relationship between the camera and the target special effect based on the posture data of the camera and the target posture data of the target special effect; based on the relative posture relationship, mapping the special effects data of the target special effect to the imaging plane corresponding to the real scene picture to obtain a special effects picture aligned with the real scene picture, and the position of the target special effect in the special effects picture matches the position of the target object in the real scene picture.

[0008] In one possible implementation, the real scene includes an entity in front of the screen, and the tracking data also includes posture data of the entity; wherein, synthesizing the special effects picture with the real scene picture to obtain a synthesized picture includes: determining the occlusion relationship between the entity and the target special effect under the camera perspective based on the posture data of the camera, the posture data of the entity, and the target posture data of the target special effect, the occlusion relationship being used to indicate the distance of the entity from the camera relative to the target special effect; and synthesizing the special effects picture with the real scene picture based on the occlusion relationship to obtain a synthesized picture.

[0009] In one possible implementation, the special effects picture and the real scene picture are synthesized according to the occlusion relationship to obtain a synthesized picture, including: when the occlusion relationship indicates that the entity is closer to the camera relative to the target special effect, according to the entity area occupied by the entity in the real scene picture, the area in the special effects picture that overlaps with the entity area is eliminated to obtain the target special effects picture; and the target special effects picture is synthesized with the real scene picture to obtain a synthesized picture.

[0010] In one possible implementation, synthesizing the special effects picture with the real scene picture according to the occlusion relationship to obtain a synthesized picture includes: when the occlusion relationship indicates that the entity is farther away from the camera relative to the target special effect, synthesizing the special effects picture with the real scene picture to obtain a synthesized picture.

[0011] In a possible implementation manner, the method further includes: displaying the synthesized picture; and / or, in response to a recording operation on the synthesized picture, recording the synthesized picture.

[0012] According to another aspect of the present disclosure, a device for synthesizing special effects in virtual shooting is provided, including: an acquisition module for acquiring, during the virtual shooting process, a real-scene picture currently captured by a camera, tracking data currently captured by a motion capture system, and posture change data of a target object to which special effects are to be added currently captured by an inertial measurement unit, wherein the real-scene picture includes a virtual scene displayed on a screen and a real scene in front of the screen, and the tracking data includes the posture data of the camera and the position data of the target object; a determination module for determining target posture data corresponding to a target special effect to be added based on the position data of the target object and the posture change data; a generation module for generating a special effects picture aligned with the real-scene picture based on the posture data of the camera, the target posture data corresponding to the target special effect, and the special effects data of the target special effect; and a synthesis module for synthesizing the special effects picture with the real-scene picture to obtain a synthesized picture, in which the target object has the target special effect.

[0013] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0014] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.

[0015] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0016] According to various aspects of the present disclosure, by using the position data of the target object collected by the motion capture system and the posture change data collected by the IMU, more accurate target posture data of the target special effect can be determined, thereby generating a more accurate special effect picture aligned with the real scene picture, and then synthesizing the special effect picture with the real scene picture, it is possible to efficiently and automatically add the target special effect to the real scene picture collected by the camera during the virtual shooting process to generate a synthetic picture, so as to show the pictures before and after adding the target special effect to the director, producer, special effects production personnel, etc. in real time, so that these personnel can instantly view the combined effect of the added special effects and the actor's performance at the virtual shooting site, thereby improving shooting efficiency. At the same time, the target special effect is added to the real scene picture taken by the camera without being rendered to the screen displaying the virtual scene, which greatly reduces the requirements for rendering performance, effectively reduces the cost of special effects production, and is more economical.

[0017] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0019] Figure 1 A block diagram of a system for synthesizing special effects in virtual shooting according to an embodiment of the present disclosure is shown.

[0020] Figure 2 A schematic diagram illustrating an application scenario of a system for synthesizing special effects in virtual shooting according to an embodiment of the present disclosure is shown.

[0021] Figure 3 A flowchart of a method for synthesizing special effects in virtual shooting according to an embodiment of the present disclosure is shown.

[0022] Figure 4 A schematic diagram illustrating occluded special effects synthesis using a special effects rendering engine according to an embodiment of the present disclosure is shown.

[0023] Figure 5 A schematic diagram illustrating a human body region recognition process according to an embodiment of the present disclosure is shown.

[0024] Figure 6A schematic diagram of generating a composite image using a special effects rendering engine according to an embodiment of the present disclosure.

[0025] Figure 7 A block diagram of a device for synthesizing special effects in virtual shooting according to an embodiment of the present disclosure is shown.

[0026] Figure 8 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0027] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0028] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0029] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C. In the description of this disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0030] It should be understood that the terms "first," "second," and the like in the claims, specification, and drawings of the present disclosure are used to distinguish between different objects, rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0031] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0032] In order to address the problem of low efficiency in special effects generation caused by the production of special effects such as "spells" based on the opinions of the director and other personnel and the content of the script in the later stage of virtual shooting, the efficiency of special effects generation can be improved by synthesizing special effects in real time during the virtual shooting process. In related technologies, Unreal Engine is used to install the RenderStream plug-in, and Disguise (a virtual shooting production platform) is used to start two Unreal Engines to render special effects and virtual scenes respectively, so that special effects and virtual scenes can be displayed on the screen in real time during virtual shooting to achieve synthetic recording; however, this method requires the Disguise central control and multiple rendering machines to work together, which is cumbersome to operate; and the special effects and virtual scenes are rendered separately using two Unreal Engines to avoid frame drops, and no relevant tracking data is saved. If the special effects need to be replaced or modified later, the production cycle will be as long and inefficient as the original solution.

[0033] In the disclosed embodiment, a scheme for synthesizing special effects in virtual filming is proposed (described in detail below). During the virtual filming process, the position, posture, and other data of the camera, actor, and special effects display during the filming process can be obtained by using a motion capture system (such as OptiTrack). The relative position between the special effects image and the real scene image captured by the camera is aligned through spatial position calculation, and the special effects occlusion position is processed through entity recognition. Finally, the image synthesis is performed, thereby achieving real-time addition of target special effects to the real scene image captured by the camera (the image includes the virtual scene displayed on the screen and the real scene in front of the screen). The synthesized image before and after the addition of the target special effects can be displayed to the director, producer, special effects production personnel, etc., so that these personnel can immediately view the combination of special effects and actor performances at the virtual filming site. The special effects can be adjusted to meet needs during the virtual filming process, thereby improving efficiency. At the same time, the generated special effects are added to the real scene image captured by the camera without being rendered to the screen displaying the virtual scene, which greatly reduces the requirements for rendering performance, effectively reduces the cost of special effects production, and is more economical.

[0034] Figure 1 FIG. 1 is a block diagram of a virtual shooting special effects synthesis system according to an embodiment of the present disclosure; FIG. Figure 1 As shown, the system includes:

[0035] A screen (such as an LED screen) is used to display the virtual scene. Depending on the actual shooting requirements, the LED screen can be configured as a flat screen, a curved screen, a tri-fold screen, or a special-shaped screen with a multi-faceted three-dimensional structure.

[0036] The camera is used to perform real-time shooting during the virtual shooting process, wherein, during the virtual shooting process, the screen displays the virtual scene in real time, and a real scene can be arranged in front of the screen. The real scene may include entities such as actors, props, and scenery. The actors can perform in front of the screen, and the camera can simultaneously capture the virtual scene displayed on the screen and the real scene in front of the screen; illustratively, during the virtual shooting process, the camera can shoot in real time at a preset frame rate, and can transmit the captured images to the special effects rendering engine in real time through an SDI interface, etc.

[0037] A motion capture system (i.e., an action capture system) is used to capture tracking data (such as position data and posture data) of entities in front of a camera or screen (such as actors) and target objects to which special effects are to be added in real time. A motion capture system is a system that captures information about changes in the motion state of an object. For example, a motion capture system may be OptiTrack, Mosys, etc. As an example, a motion capture system may include multiple motion capture trackers (such as reflective markers, sensors, etc.). These motion capture trackers may be installed on entities and target objects in front of a camera or screen to track the real-time posture data of the camera, entity, and target object, and transmit the tracked posture data to the special effects rendering engine in real time.

[0038] An inertial measurement unit (IMU) is used to be set on a target object and collect the target object's posture change data in real time (such as the acceleration and angular velocity of the three axes when the target object moves in three-dimensional space). IMU is a device used to measure the motion state of an object, usually composed of multiple sensors. It is mainly used to measure the acceleration and angular velocity of the three axes of an object in three-dimensional space, so as to be able to calculate the position, speed and posture of the object's movement. Among them, the target object's posture change data collected by the IMU can be used to determine the target posture data of the target special effect to be added to the target object, so that the determined target posture data of the target special effect can be sent to the special effects rendering engine.

[0039] The special effects rendering engine is used to add the desired special effects to the picture taken by the camera in real time, and render to generate the picture with the added special effects; illustratively, the special effects rendering engine can add target special effects to the real scene based on the real scene taken by the camera in real time, the tracking data collected by the motion capture system, and the target posture data of the target special effects, and generate a synthetic picture with the added special effects.

[0040] For example, the system may further include a display device for displaying the images captured by the camera, and for displaying the composite images generated by the special effects rendering engine after adding special effects. The number of display devices may be one or more, and may include, for example, a monitoring screen viewed by directors, producers, etc., and a display panel corresponding to the special effects production tools used by special effects production personnel. The images after adding special effects can be displayed on multiple display devices, allowing different staff members to preview and monitor the composite effects in real time.

[0041] Exemplarily, the above system may further include: a virtual scene rendering engine, configured to render a virtual scene in real time and output the rendered virtual scene to a screen for display.

[0042] Among them, since the images with added special effects rendered by the special effects rendering engine are displayed on a display device for viewing by directors, producers, special effects production personnel, etc., without being displayed on a screen used to display virtual scenes, and the size of the display device is usually much smaller than that of the screen, the requirements for rendering performance are relatively low. Therefore, an engine with lower rendering performance can be used as a special effects rendering engine, or a single engine can be used as both a special effects rendering engine and a virtual scene rendering engine to meet the needs of real-time special effects generation, thereby reducing the cost of special effects generation. As an example, the special effects rendering engine and the virtual scene rendering engine can be implemented based on the same Unreal Engine, without the need to configure two Unreal Engines, thereby effectively reducing the cost of special effects production and being more economical.

[0043] It should be noted that the above system is merely an example, and the components in the above system may be integrated, or the above system may be configured with more or fewer components, without limitation. For example, the components in the above system may be time-synchronized, for example, by generating a synchronization signal using a synchronization signal generator, thereby achieving time synchronization between components such as the motion capture system, camera, IMU, screen, virtual scene rendering engine, display device, and special effects rendering engine.

[0044] Figure 2 A schematic diagram of an application scenario of a virtual shooting special effects synthesis system according to an embodiment of the present disclosure is shown as follows: Figure 2 As shown, the above Figure 1 In the system shown in , during the virtual shooting process, the actor is in front of the LED screen, and the camera frames the actor and the LED screen, thereby capturing a real scene including the virtual scene displayed on the screen and the real scene in front of the screen (including the actor). The captured real scene can be transmitted to the special effects rendering engine;

[0045] The motion capture system can track the pose data of the camera, the actor, and the target objects on the actor (such as the actor's hands, feet and other human body parts, or the actor's entire body) in real time, and can transmit the tracking data (camera pose data (i.e., camera pose data) and character pose data (i.e., actor pose data)) to the special effects rendering engine. At the same time, the tracking data (camera pose data) can be transmitted to the virtual scene rendering engine;

[0046] The IMU is set on the target object and collects the target object's posture change data in real time. In this way, the posture change data collected by the IMU can be used to determine the initial posture data of the target special effect, and the position data of the target object collected by the motion capture system can be used to optimize the initial posture data of the target special effect to obtain more accurate target posture data of the target special effect. The target posture data of the target special effect (i.e., special effect posture data) can then be transmitted to the special effect rendering engine;

[0047] After obtaining the real-life scene captured by the camera, the special effects rendering engine can align the relative positions of the special effects image and the real-life scene captured by the camera through spatial position calculation based on the tracking data captured by the motion capture system and the target pose data of the target special effect. That is, the relative position of the special effect and the real scene is calculated so that the positions of the special effect and the real scene can be matched, so as to add the target special effect corresponding to the target object to the real-life scene in real time and generate a composite picture with the added special effect. In some examples, the human body area recognition can also be used to process the partial area where the person occludes the special effect. That is, the occlusion relationship can be calculated using the relative position of the person and the special effect, thereby realizing special effects synthesis with occlusion effect. In other examples, the special effects rendering engine can perform color conversion on the real-life scene captured by the camera, the special effects image, the composite picture after adding the special effect, etc. to eliminate color difference. In other examples, the special effects rendering engine can perform special effects video synthesis recording in response to the recording operation, that is, record the composite picture after adding the special effect. In other examples, the composite picture after adding the special effect can also be displayed on a display device for these personnel to preview and monitor.

[0048] The virtual scene rendering engine can render the image of the virtual scene in real time based on tracking data (camera pose data) and display the rendered image on the LED screen. Specifically, the virtual scene rendering engine can provide the virtual camera in the virtual scene rendering engine with the pose data of the real camera in three-dimensional space. The virtual scene rendering engine relies on real-time rendering and other technologies to synchronize the movement and composition of the virtual camera with the real camera, so that the rendered virtual scene image conforms to the perspective captured by the real camera.

[0049] In practical applications, the method for synthesizing special effects in virtual shooting of the embodiment of the present disclosure can be deployed on various terminal devices through software or hardware modification. The terminal device can be deployed with a rendering engine. The terminal device involved in the embodiment of the present disclosure can refer to a device with a wireless connection function and / or a wired connection function. The wireless connection function means that it can be connected to other devices through wireless connection methods such as wifi and Bluetooth. The terminal device involved in the embodiment of the present disclosure can also communicate with other devices through a wired connection function. The terminal device involved in the embodiment of the present disclosure can be a touch screen, a non-touch screen, or a screenless terminal. The touch screen can be controlled by clicking, sliding, etc. on the display screen with a finger or a stylus. The non-touch screen device can be connected to an input device such as a mouse, keyboard, touch panel, etc., and the terminal device can be controlled by the input device. For example, a device without a screen can be a Bluetooth speaker without a screen. For example, the terminal device of the present application can include but is not limited to user equipment (UE), mobile devices, mobile terminals, handheld devices, tablet computers, laptops, PDAs, computing devices, etc.

[0050] The method for synthesizing special effects in virtual shooting of the embodiment of the present disclosure can also be deployed on a server, on which a rendering engine can be deployed. The server can be located in the cloud or locally, and can be a physical device or a virtual device, such as a virtual machine, a container, etc., with a wireless communication function, wherein the wireless communication function can be set in the chip (system) or other parts or components of the server. It can refer to a device with a wireless connection function, and the function of wireless connection means that it can be connected to other servers or terminal devices through wireless connection methods such as Wi-Fi and Bluetooth. The server involved in the embodiment of the present disclosure can also have the function of communicating via a wired connection. For example, the server of the embodiment of the present disclosure can be located in the cloud, communicate with the terminal device, receive the real scene picture currently captured by the camera, the tracking data currently captured by the motion capture system, and the position change data of the target object currently captured by the inertial measurement unit, and use the method for synthesizing special effects in virtual shooting deployed on the server based on the above-mentioned real scene picture, tracking data and position change data to generate a synthetic picture with special effects added, and return it to the terminal device to display the generated synthetic picture to the user in the terminal device.

[0051] Figure 3 FIG. 1 is a flow chart showing a method for synthesizing special effects in virtual shooting according to an embodiment of the present disclosure. Figure 3 As shown, the method includes: steps S11 to S14.

[0052] In step S11, during the virtual shooting process, the real scene currently captured by the camera, the tracking data currently captured by the motion capture system, and the posture change data of the target object to be added with special effects currently captured by the inertial measurement unit are obtained. The real scene includes the virtual scene displayed on the screen and the real scene in front of the screen. The tracking data includes the posture data of the camera and the position data of the target object.

[0053] In practical applications, during the virtual shooting process, the camera, motion capture system, and inertial measurement unit can synchronously collect data at a certain frame rate and transmit the data to an electronic device used to execute the screen special effects synthesis method of the embodiment of the present disclosure. For example, the electronic device can be deployed with the above-mentioned special effects rendering engine to implement subsequent special effects synthesis processing to generate a synthetic screen with special effects. Of course, the electronic device can also be deployed with a virtual scene rendering engine to render the image of the virtual scene and display the image of the virtual scene on the screen.

[0054] The target object is an object associated with a special effect, such as the sender or receiver of a special effect, or the target to which the special effect is applied. A target object can represent entities such as actors, props, or scenery in a virtual filmed real-world scene, or it can be a component of these entities. For example, a target object can be an actor's eyes, hands, or feet. As an example, when filming a TV series or film about fantasy, magic, or science fiction, the special effect could be a "spell." For example, if the target object is an actor, the special effect could be "ice" appearing around the actor's body. Another example is if the target object is the actor's eyes, the special effect could be a "laser" emitted from the eyes. Or if the target object is the actor's hands, the special effect could be a "fireball" with attacking abilities appearing in the hands. Another example is if the target object is stone, the special effect could be a "dazzling light" emitted from the stone.

[0055] One or more motion capture trackers can be pre-installed on the camera and the target object, wherein the location and number of the motion capture trackers can be configured according to the conditions; for example, if the target object is an actor, the motion capture tracker can be installed on the back of the actor's body; for another example, if the target object is the actor's hands, feet, limbs or other body parts, the motion capture tracker can be installed on the corresponding body parts; for another example, if the target object is a leaf in a real scene, the motion capture tracker can be installed on the leaf. The motion capture system can capture the pose data of each motion capture tracker in a global coordinate system (such as a world coordinate system) in real time. The pose data can be used as the pose data of the entity (i.e., the camera or target object) on which the corresponding motion capture tracker is installed. The position data of the target object used below can be the position data contained in the pose data. Different motion capture trackers can have different numbers, so that the pose data or position data collected by the motion capture system can be distinguished based on the numbers of the motion capture trackers installed on different entities.

[0056] Considering that the accuracy of the target object's pose data (equivalent to the pose data of the target special effect to be added to the target object) will affect the accuracy of subsequently adding the target special effect to the real scene, or in other words, if the target object's pose data is not accurate enough, there will be a large error between the display position of the target special effect in the subsequently generated special effect scene and the target object in the real scene, thereby affecting the display effect of the synthesized special effect. Using a motion capture system alone to collect the pose data of the target object will have certain errors (sometimes even missing data), and using data collected by an inertial measurement unit alone to calculate the pose will produce cumulative errors over time. Therefore, the embodiment of the present disclosure not only uses the motion capture system to track the position data of the target object, but also uses the inertial measurement unit to collect the pose change data of the target object in real time. In this way, based on the position data of the target object tracked by the motion capture system and the pose change data of the target object collected by the inertial measurement unit in real time, more accurate pose data of the target object can be determined. Among them, an inertial measurement unit can be set on the target object to be added with special effects in advance according to actual needs, and the pose change data collected by the inertial measurement unit is used as the pose change data of the entire target object. It is understandable that the target special effect to be added to the target object will change with the change of the position and posture of the target object. Therefore, using more accurate position and posture data of the target object can determine more accurate position and posture data of the target special effect.

[0057] In step S12, target posture data corresponding to the target special effect to be added is determined based on the position data and posture change data of the target object.

[0058] Among them, after obtaining the position data of the target object currently collected by the motion capture system, the posture change data corresponding to the current frame can be found (equivalent to the posture change data between the current frame and the previous frame), and then the posture change data is used to perform forward recursion on the target posture data and speed of the target object in the previous frame to obtain the initial posture data and speed prediction result of the target object in the current frame. Then, according to the position data (i.e., the three-dimensional coordinates in the world coordinate system) of the tracking point (i.e., the motion capture tracker) set on the target object of the current frame, the initial posture data of the target object of the current frame is optimized to obtain the optimized target posture data of the target object in the current frame, and then based on the target posture data of the target object in the current frame, the target posture data of the target special effect corresponding to the current frame can be determined. Among them, the position data of the target object currently collected by the motion capture system includes the position data of the tracking point set on the target object in the global coordinate system (such as the world coordinate system), which is equivalent to collecting the three-dimensional coordinates of a point on the surface of the target object in the world coordinate system.

[0059] Among them, the posture change data currently collected by the inertial measurement unit may include the acceleration and angular velocity of the target object from the last collection moment (i.e., the last frame) to the current collection moment (i.e., the current frame). It should be understood that by integrating the acceleration and angular velocity, the posture change (i.e., position change and attitude change) of the target object from the last frame to the current frame can be calculated, and then the target posture data of the target object in the previous frame can be combined to obtain the target posture data of the target object in the current frame. The target posture data of the target object determined in this way can represent the overall posture of the target object, and then the target posture data representing the overall posture of the target special effect can be determined.

[0060] It can be understood that the target special effect can be in the same posture as the target object, that is, the relative posture relationship between the target special effect and the target object can indicate that there is no difference in the posture of the two, or in other words, the posture of the target object is also the posture of the target special effect. For example, the target object is the actor's hand, and the target special effect is the "fireball" in the actor's hand. At the same time, the posture of the "fireball" and the "hand" are consistent. In actual applications, the target special effect may also be at a certain distance from the target object, that is, there is a difference in the posture data of the two. Then the relative posture relationship between the target special effect and the target object indicates the difference in posture between the two. After obtaining the target posture data of the target object, the target posture data of the target special effect can be inferred based on the relative posture relationship between the target special effect and the target object. The size of the difference indicated by the relative posture relationship between the target special effect and the target object can be set according to needs. For example, the target object is the actor's foot, and the target special effect is the "glowing footprints left by the previous step" when the actor walks. At the same moment, there is a difference in posture between the "glowing footprints left by the previous step" and the "foot". The second relative posture relationship between the "glowing footprints left by the previous step" and the "foot" can be set to a position difference of 0.5 meters along the direction of the footsteps.

[0061] Since the target object is an entity in the real scene, its posture data is easier to obtain and calculate. Therefore, after obtaining the target posture data of the target object, the target posture data of the target special effect can be further determined by combining the relative posture relationship between the target special effect and the target object. For example, the relative posture relationship between the target special effect and the target object is relatively fixed. For example, when a "fireball" is displayed on the actor's hand, the relative posture relationship between the target special effect "fireball" and the target object "hand" remains unchanged, thereby showing the effect of a "fireball" always on the hand; for another example, when an actor walks normally, the speed is usually uniform, and the relative posture relationship between the "luminous footprint left by the previous step" and the "foot" remains unchanged; in this way, after the relative posture relationship between the target special effect and the target object is preset, as the posture of the target object changes, the latest target posture data of the target special effect can be determined based on the latest target posture data of the target object and the relative posture relationship between the two.

[0062] In a possible implementation, determining the target pose data corresponding to the target special effect to be added based on the position data and pose change data of the target object may include:

[0063] Determining initial pose data corresponding to the target special effect at the current acquisition moment based on target pose data and pose change data corresponding to the target special effect at the previous acquisition moment; wherein the target pose data corresponding to the target special effect at the previous acquisition moment is determined based on the position data and pose change data of the target object acquired at the previous acquisition moment;

[0064] Based on the position data of the tracking points set on the target object in the global coordinate system and the position data of the tracking points set on the target object in the local coordinate system of the target object, the initial pose data corresponding to the target special effect at the current acquisition moment is optimized to obtain the target pose data corresponding to the target special effect at the current acquisition moment.

[0065] The target pose data corresponding to the target special effect at the previous capture moment is determined based on the target object's position data and pose change data at the previous capture moment. This determination process is the same as the target pose data corresponding to the target special effect at the current capture moment. It should be understood that the target object can have an initialized pose at the beginning of filming, that is, the target special effect can have initialized target pose data.

[0066] As mentioned above, the posture of the target special effect and the target object can be the same, which means that the target posture data corresponding to the target special effect at the last acquisition moment is also the target posture data corresponding to the target object at the last acquisition moment. Therefore, the integral deduction can be directly performed based on the posture change data to obtain the posture change of the target object from the last acquisition moment to the current acquisition moment, and then the posture change is accumulated with the target posture data corresponding to the target special effect at the last acquisition moment to obtain the initial posture data corresponding to the target special effect at the current acquisition moment; or, there can also be a relatively fixed posture difference between the target special effect and the target object. In this case, the target special effect is initialized. The standard pose data can be determined based on the initialization pose data of the target object combined with the relative pose relationship between the target special effect and the target object (that is, the relatively fixed pose difference between the two). It should be understood that although there is a pose difference between the target object and the target special effect, the pose change of the target object from the last acquisition moment to the current acquisition moment is also the pose change of the target special effect from the last acquisition moment to the current acquisition moment. Therefore, the initial pose data corresponding to the target special effect at the current acquisition moment determined by the target pose data corresponding to the target special effect at the last acquisition moment and the pose change data is the actual pose that already includes the above-mentioned pose difference.

[0067] It can be understood that the position of the tracking point on the target object is fixed. If a local coordinate system is established with a certain point on the target object (such as the tracking point, the center point, or the IMU location, etc.) as the origin, then the position data of the tracking point in the local coordinate system is known. Then, if the pose of the target object is represented by the rotation matrix R and the translation matrix t, and the position of a tracking point of the target object (a motion capture tracker) in the global coordinate system is P W , the position of the target object in the local coordinate system is P R , where P W But the position data of the tracking points set on the target object collected by the motion capture system, P R For example, the P of a tracking point on the target object can be obtained by modeling the target object and each tracking point set on the target object and establishing a local coordinate system based on the modeling result of the target object. R , P W and P R Should satisfy P W =R*P R +t, and when the pose estimation of the target object is inaccurate, P W ≠R*P R +t, so we can adjust the pose of the target object so that P W =R*P R +t, thus, P can be satisfied W =R*P R+t is the target, and the initial pose data corresponding to the target special effect is optimized to obtain the data that satisfies P W =R*P R +t target pose data (that is, the target pose data for the target special effect), wherein those skilled in the art can use the optimization algorithm known in the art to optimize the initial pose data corresponding to the target special effect to obtain the data that satisfies P W =R*P R +t target pose data, which is not limited in the embodiment of the present disclosure. In this way, more accurate target pose data of the target special effect can be obtained, which is conducive to improving the position accuracy of the target special effect in the subsequent special effect screen. It should be understood that the initial pose data calculated using the pose change data collected by the IMU is usually not much different from the actual pose of the target object. Compared with directly using P W and P R To find the satisfying P W =R*P R For R and t of +t, the target pose data of the target object obtained by optimizing the initial pose data calculated by using the pose change data collected by IMU is more computationally efficient and requires less calculation, and the target pose data that is more consistent with the actual pose of the target object can be found more quickly.

[0068] In actual applications, if the position data of the target object is not collected at the current moment (i.e., the position data of the target object is lost), that is, the motion capture system does not track the position data of the target object at the current moment, then the initial position data of the target special effect at the current moment determined based on the target posture data corresponding to the target special effect at the previous collection moment and the posture change data can be directly used as the target posture data of the target special effect at the current collection moment, thereby avoiding position jumps caused by missing tracking data, affecting the subsequent calculation of the relative positions of special effects, cameras, and entities, so as to improve the stability of the motion capture system tracking data and prevent tracking data jumps. However, it should be noted that since the use of data collected by the IMU to calculate the posture will produce cumulative errors, this method can be applied for a short period of time. If the position data of the target object is not collected for more than a certain period of time, for example, the position of the motion capture tracker set on the target object can be adjusted to avoid losing the position data of the target object for a long time.

[0069] In step S13, a special effects picture aligned with the real scene picture is generated according to the posture data of the camera, the target posture data corresponding to the target special effect, and the special effect data of the target special effect.

[0070] This step S13 can be performed by, for example, the above Figure 1The special effects rendering engine in the system shown is executed. For example, the special effects rendering engine can be synchronized with the frame rate and timestamp of the camera, and the rendering engine frame rate and timestamp can be locked to ensure that the special effects picture rendered by the special effects rendering engine has the same frame rate and timestamp as the real-scene picture captured by the camera, so as to facilitate the subsequent synthesis of the special effects picture with the corresponding real-scene picture. It is understandable that during a virtual shooting process, the frame rate of the camera usually remains unchanged. Therefore, after a synchronization, the frame rate of the special effects rendering engine can remain unchanged, and special effects can be directly added to the real-scene picture captured in real time.

[0071] Among them, the camera's posture data may include the camera's position data and posture data, the camera's position data may indicate the position of the camera's lens optical center in the world coordinate system, and the camera's posture data may indicate the camera's lens' shooting direction in the world coordinate system. The target posture data of the target special effect may include the position data and posture data of the target special effect corresponding to the real scene (i.e., in the world coordinate system). Thus, given the camera's posture data and the target posture data corresponding to the target special effect, the relative posture relationship between the camera and the target special effect can be known, that is, the target special effect under the camera's shooting angle of view can be known. Based on the relative posture relationship between the camera and the target special effect, the special effects data corresponding to the target special effect can be mapped to the camera's imaging plane (i.e., the imaging plane of the real-scene picture currently captured by the camera), and a special effects picture aligned with the real-scene picture can be obtained. Thus, the above-mentioned special effects picture aligned with the real-scene picture is generated based on the camera's posture data, the target posture data corresponding to the target special effect, and the special effects data of the target special effect, including:

[0072] Determine the relative posture relationship between the camera and the target special effect based on the posture data collected by the camera and the target posture data of the target special effect;

[0073] Based on the relative posture relationship, the special effect data of the target special effect is mapped to the imaging plane corresponding to the real scene picture to obtain a special effect picture aligned with the real scene picture, and the position of the target special effect in the special effect picture matches the position of the target object in the real scene picture.

[0074] Among them, the special effects data of the target special effects can be used to indicate the style, form, color, size and other specific content of the target special effects. Mapping the special effects data of the target special effects to the imaging plane corresponding to the real scene based on the relative posture relationship can be understood as using a camera to shoot and image the target special effects with a certain posture displayed on the target object. Therefore, the position of the target special effect in the special effects picture obtained by mapping the special effects data of the target special effects to the imaging plane corresponding to the real scene using the relative posture relationship between the camera and the target special effect matches the position of the target special effect in the real scene, such as the positions of the two overlap or maintain a certain distance. Among them, mapping the special effects data of the target special effect to the imaging plane based on the relative posture relationship, that is, mapping the target posture of the target special effect and the style, form, color and other contents indicated by the special effects data to the imaging plane, so that the mapped target special effect presents a certain posture and the style, form and color indicated by the special effects data in the imaging plane. In actual applications, special effects producers can prepare special effects data for target special effects in advance according to actual needs, and can pre-set the association between the target special effects and the motion capture tracker set on the target object. In this way, based on the association, they can know the target special effects currently to be added to any target object.

[0075] Optionally, since a three-dimensional virtual scene is constructed in the rendering engine during the virtual shooting process, and models such as a screen model and a virtual camera model are constructed; wherein the screen model is synchronized with the content displayed on the real screen, and the virtual camera model is synchronized with the field of view of the real camera; thus, the special effects rendering engine can construct a target special effects model based on the special effects data of the target special effects, and place the target special effects model in the above-mentioned three-dimensional virtual scene, so that the relative posture relationship between the target special effects model and the virtual camera model is synchronized with the relative posture relationship between the target special effects and the real camera, that is, after the target special effects model is placed in the three-dimensional virtual scene, the real When the relative posture of the camera and the target special effect changes, the relative posture between the target special effect model and the virtual camera model will also change synchronously; in this way, based on the relative posture relationship between the real camera and the target special effect, the relative posture relationship between the target special effect model and the virtual camera model in the three-dimensional virtual scene can be determined, and then combined with the posture of the virtual camera model in the three-dimensional virtual scene (that is, the shooting angle of the virtual camera model), the target special effect model can be mapped to the imaging plane of the virtual camera model, that is, based on the target special effect model under the shooting angle of the virtual camera model, a special effect picture aligned with the real scene picture is generated.

[0076] Among them, the above-mentioned virtual camera model and target special effect model can also be constructed in the special effect rendering engine, and the camera pose data and the target pose data of the target special effect can also be uniformly converted into the rendering engine coordinate system of the special effect rendering engine to realize the generation of a special effect picture aligned with the real scene picture based on the relative pose relationship between the camera and the target special effect, and this embodiment of the present disclosure is not limited to this. It should be understood that since the real scene picture is a picture shot by a real camera, the perspective of the real camera when shooting the real scene picture is the same as the perspective of the virtual camera model. According to the above method, the relative pose relationship between the real camera and the target special effect when the real camera shoots the real scene picture can be determined, that is, the pose of the target special effect model placed in the three-dimensional virtual scene can be determined. After the target special effect model is placed according to the pose, the position of the target special effect model in the perspective of the virtual camera model is the position where the target special effect needs to be added to the real scene picture. Thus, a special effect picture aligned with the real scene picture can be generated, that is, the position of the target special effect in the generated special effect picture can be matched with the position of the target object in the real scene picture, such as the positions of the two overlapping or maintaining a certain distance.

[0077] In step S14 , the special effect picture and the real scene picture are synthesized to obtain a synthesized picture, in which the target object has the target special effect.

[0078] It should be understood that the resolution and size of the special effects image generated by step S13 can be the same as the real scene image. As an example, a special effects rendering engine can be used to configure the special effects image on a layer of the real scene image to achieve the synthesis of the special effects image and the real scene image to obtain a synthesized image. The embodiments of this disclosure do not limit the image synthesis method.

[0079] Consider that different devices may use different color spaces, resulting in color differences in the images output by different devices. For example, the color space used by the special effects production tool to produce the target special effects may differ from the color space used by the camera. As a result, there may be color differences between the special effects image generated by the special effects rendering engine and the real-scene image used by the camera. At the same time, the special effects rendering engine will synthesize the special effects image and the real-scene image to render a composite image, and the color space used by the special effects rendering engine for rendering may also differ from the color space of the special effects production tool and the camera. Therefore, in order to eliminate the color difference between the target special effect and other content in the generated composite image, the color difference between the target special effect and other content in the image can be eliminated before generating the composite image. Therefore, in one possible implementation, before synthesizing the special effects picture with the real scene picture, the method may further include: obtaining a first mapping relationship between the color space used by the camera and the color space used by the special effects rendering engine, and a second mapping relationship between the color space used by the special effects production tool to produce the target special effect and the color space used by the special effects rendering engine; then, based on the first mapping relationship, the real scene picture can be color converted; based on the second mapping relationship, the special effects picture can be color converted, that is, the real scene picture and the special effects picture are converted to the color space of the special effects rendering engine, and then the special effects picture converted to the color space of the special effects rendering engine is synthesized with the real scene picture.

[0080] Among them, the color space can be RGB color space, CMYK color space, Lab color space, YCbcr color space, etc., and the mapping relationship between different color spaces can be represented by a conversion matrix. For example, the color calibration mapping relationship LUT (Look-Up Table) of the camera and the color calibration mapping relationship lookup table of the special effects production tool can be obtained in advance from the manufacturer of the device or pre-calibrated. The lookup table contains the mapping relationship between the color space used by the corresponding device and other color spaces. In this way, color conversion is performed on both the real scene picture and the special effects picture, so that the color space used by the camera corresponding to the real scene picture and the color space used by the special effects production tool corresponding to the special effects picture are uniformly converted into the color space used by the special effects rendering engine to eliminate the color difference between the real scene picture and the special effects picture, thereby obtaining a composite picture without color difference.

[0081] In practical applications, after obtaining the composite picture, the composite picture can be displayed, for example, by Figure 1The system is executed on a display device. For example, the real-life scene and the synthesized scene can be displayed simultaneously on the display device, allowing directors, producers, special effects creators, and others to better view and compare the effects before and after adding special effects. As an example, the display device can be a display panel corresponding to a special effects creation tool, and the synthesized scene can be displayed on the display panel so that special effects creators can preview the generated special effects. As another example, the display device can be a director's monitor screen, and the synthesized scene can be displayed on the director's monitor screen so that the director can monitor the generated special effects in real time.

[0082] In one possible implementation, the method may further include: recording the composite picture in response to a recording operation for the composite picture. For example, a special effects producer or the like may trigger the recording operation via a recording function button displayed on a display device; as an example, a function button for recording the picture may be displayed in a display panel corresponding to the special effects production tool, and the special effects producer or the like may click the button to trigger the special effects rendering engine to record the composite picture, i.e., the special effects rendering engine stores the composite picture; the recorded composite picture includes the real-scene picture content and the added special effects, so that during the virtual shooting process, the composite picture with the added special effects can be recorded in real time to obtain one or more segments of video data with the added special effects.

[0083] It should be noted that post-production staff can also use special effects production tools to import new special effects according to needs, and make post-production adjustments to the target special effects in the synthetic images generated during the virtual shooting process (that is, the special effects layer images on the real-scene images) by inserting and modifying them frame by frame.

[0084] According to the method of the embodiment of the present disclosure, by using the position data of the target object collected by the motion capture system and the posture change data collected by the IMU, more accurate target posture data of the target special effect can be determined, thereby generating a more accurate special effect picture aligned with the real scene picture, and then synthesizing the special effect picture with the real scene picture, it is possible to efficiently and automatically add the target special effect to the real scene picture collected by the camera to generate a synthetic picture during the virtual shooting process, so as to show the pictures before and after adding the target special effect to the director, producer, special effects production personnel, etc. in real time, so that these personnel can instantly view the combined effect of the added special effects and the actor's performance at the virtual shooting site, thereby improving shooting efficiency. At the same time, the target special effect is added to the real scene picture taken by the camera without being rendered to the screen displaying the virtual scene, which greatly reduces the requirements for rendering performance, effectively reduces the cost of special effects production, and is more economical.

[0085] Considering that the special effects images generated by the above-mentioned special effects synthesis method can often only be on the top layer of the overall image, if part of the special effects are required to be blocked by entities in front of the screen (such as characters, scenery, props, etc. in front of the screen), it cannot be achieved because the special effects image layer is overlaid as a whole on the real scene image captured by the camera. Therefore, the embodiment of the present disclosure also proposes, based on the above-mentioned special effects synthesis scheme in virtual shooting, to achieve the occlusion effect of entity-blockable special effects in the special effects rendering engine through entity position positioning, special effects position positioning, entity recognition, etc., which is conducive to making the target special effects in the synthesized image appear more natural.

[0086] As described above, the real scene includes an entity in front of the screen, and the tracking data collected by the tracking system may also include the entity's posture data. That is, a motion capture tracker may be fixedly set on the entity to track the entity's posture data. Based on this, in one possible implementation, in the above step S14, the special effects image and the real scene image are synthesized to obtain a synthesized image, which may include:

[0087] Step S141, determining an occlusion relationship between the entity and the target special effect from the camera's perspective based on the camera's pose data, the entity's pose data, and the target pose data of the target special effect, where the occlusion relationship indicates how far the entity is from the camera relative to the target special effect.

[0088] Step S142 : synthesize the special effects image and the real scene image according to the occlusion relationship to obtain a synthesized image.

[0089] In step S141, the known camera pose data means the position and viewing angle of the camera in the world coordinate system. Therefore, the known entity pose data means the position and posture of the known entity in the world coordinate system. The known target pose data of the target special effect means the position and posture of the known target special effect in the world coordinate system. Therefore, based on the phase pose relationship between the entity and the camera and the relative pose relationship between the target special effect and the camera, the distance of the entity from the camera relative to the target special effect under the camera's viewing angle can be obtained.

[0090] As described above, the virtual camera model and the target special effect model can be constructed in the special effect rendering engine. Thus, the special effect rendering engine is used to perform the step S141. Figure 4As shown, the camera pose (camera pose data), entity pose (i.e., entity pose data) and special effect pose (target pose data of target special effect) are converted to the special effect rendering engine coordinate system, so that the virtual camera model and target special effect model in the special effect rendering engine are synchronized with the camera pose and special effect pose, respectively, and an entity model synchronized with the entity pose (for example, a character model synchronized with the character pose) can also be constructed to determine the distance of the entity model from the virtual camera model to the target special effect model under the perspective of the virtual camera model, that is, to determine the distance of the entity from the target special effect to the camera under the perspective of the camera, so that the occlusion relationship can be obtained, and then the special effect synthesis processing of step S142 is executed.

[0091] It should be understood that if the entity is closer to the camera than the target special effect, the entity may block the special effect, while if the entity is farther from the camera than the target special effect, the entity is likely not to block the special effect. Therefore, in step S142, the special effect image and the real scene image are synthesized according to the occlusion relationship to obtain a synthesized image, which may include:

[0092] When the occlusion relationship indicates that the entity is closer to the camera than the target special effect, the area in the special effect picture that overlaps with the entity area is eliminated based on the entity area occupied by the entity in the real scene picture to obtain the target special effect picture; the target special effect picture is synthesized with the real scene picture to obtain a synthesized picture; or

[0093] When the occlusion relationship indicates that the entity is farther away from the camera than the target special effect, the special effect picture and the real scene picture are synthesized to obtain a synthesized picture.

[0094] For example, Figure 5 As shown, if the entity includes a person in front of the screen, after obtaining the camera image (i.e., the real-scene image captured by the camera), the real-scene image can be first detected to detect the specific position of the person in the real-scene image, and then fine human body segmentation can be performed based on the detection result (i.e., the specific position of the person in the real-scene image) to obtain the human body area occupied by the person in the real-scene image, wherein the human body area can be represented as a human body mask for subsequent processing of special effects images. The human body mask can specifically represent the outline of the human body area cut out of the real-scene image, and the human body mask can use a grayscale image. By adopting a two-step strategy of first detection and then segmentation, the recognition accuracy of the human body area is higher.

[0095] Among them, human body segmentation can be performed in a deep learning manner. For example, after the real scene picture is input into the human body segmentation model, the pixel points corresponding to the human body position in the current frame real scene picture and its outer bounding box will be generated. At the same time, the score and category of each bounding box will be output. After obtaining the result, the first wave of screening is performed according to the score of the bounding box to remove the low-quality mis-segmentation results. Then the segmentation results of non-human categories are screened out, and finally the segmentation results are optimized, such as filling in positions such as holes to obtain the segmented human body area. It should be understood that when the entity also includes other objects in front of the screen, the corresponding object detection model and object segmentation model can also be used to obtain the entity area where any entity in the real scene image is located.

[0096] In practical applications, the above step S142 can be performed in a special effects rendering engine, for example, Figure 6 As shown, the special effects rendering engine can determine the human body area occupied by the character in the camera's real scene image (i.e., the real scene image captured by the camera) when the occlusion relationship indicates that the character is closer to the camera than the target special effect (i.e., the character may block the special effect), and eliminate the area in the special effects image that overlaps with the human body area. That is, when the target special effect is far from the camera relative to the character, the human body is cut to obtain a character mask, and the area blocked by the human body in the special effects image is erased by the generated character mask. The target special effects image obtained after erasure is then synthesized with the camera's real scene image (such as placing the target special effects image on a layer of the real scene image) to obtain a synthetic image with an occlusion effect; and when the occlusion relationship indicates that the character is farther from the camera than the target special effect (representing that the target special effect is in front of the character, and the character is likely not to block the target special effect at this time), the special effects image can be directly synthesized with the real scene image to obtain a synthetic image. After obtaining the synthetic image, the synthetic image can be previewed, monitored, and recorded.

[0097] It should be understood that when the entity is closer to the camera than the target special effect, the entity area in the real-scene picture may not overlap with the target special effect displayed in the special effects picture. That is to say, although the area in the special effects picture that overlaps with the entity area can be eliminated, the eliminated area may not contain the specific content of the target special effect. In this case, the target special effects picture obtained is equivalent to the original special effects picture.

[0098] Among them, the specific implementation method of the picture synthesis in step S14 in the above-mentioned embodiment of the present disclosure can be specifically referred to to implement the processing of synthesizing the target special effect picture or the special effect picture with the real scene picture in the above-mentioned step S142. For example, the target special effect picture and the real scene picture can be converted to the color space of the special effect rendering engine, and then the target special effect picture converted to the color space of the special effect rendering engine is synthesized with the real scene picture, which will not be repeated here.

[0099] According to the disclosed embodiments, a certain occlusion effect can be created between the target special effects and the entity in the generated composite image, making the display effect of the target special effects added in the composite image more natural and realistic, and increasing the universality of special effects addition scenes in virtual photography while not requiring higher rendering performance from the device. Furthermore, by calculating the relative positional relationship between the entity, special effects, and camera, the relative occlusion relationship between the entity and the target special effects in the virtual and real space is determined, and a set of special effects image processing occlusion solutions is proposed through entity recognition and position recognition, which can prevent the occluded portion of the special effects image from being displayed without affecting the display of the entity.

[0100] Figure 7 A block diagram of a device for synthesizing special effects in virtual shooting according to an embodiment of the present disclosure is shown. Figure 7 As shown, the device includes:

[0101] Acquisition module 701 is used to acquire, during the virtual shooting process, the real scene currently captured by the camera, the tracking data currently captured by the motion capture system, and the pose change data of the target object to be added with special effects currently captured by the inertial measurement unit. The real scene includes the virtual scene displayed on the screen and the real scene in front of the screen. The tracking data includes the pose data of the camera and the position data of the target object.

[0102] A determination module 702 is configured to determine target pose data corresponding to a target special effect to be added based on the position data of the target object and the pose change data;

[0103] A generating module 703 is configured to generate a special effect picture aligned with the real scene picture based on the camera's posture data, the target posture data corresponding to the target special effect, and the special effect data of the target special effect;

[0104] The synthesis module 704 is configured to synthesize the special effect image and the real scene image to obtain a synthesized image, in which the target object has the target special effect.

[0105] In one possible implementation, the posture change data includes the acceleration and angular velocity of the target object from the last acquisition moment to the current acquisition moment, and the position data of the target object includes the position data of the tracking point set on the target object in the global coordinate system; wherein, determining the target posture data corresponding to the target special effect to be added based on the position data of the target object and the posture change data includes: determining the initial posture data corresponding to the target special effect at the current acquisition moment based on the target posture data corresponding to the target special effect at the last acquisition moment and the posture change data; wherein, the target posture data corresponding to the target special effect at the last acquisition moment is determined based on the position data and posture change data of the target object collected at the last acquisition moment; based on the position data of the tracking point set on the target object in the global coordinate system and the position data of the tracking point set on the target object in the local coordinate system of the target object, the initial posture data corresponding to the target special effect at the current acquisition moment is optimized to obtain the target posture data corresponding to the target special effect at the current acquisition moment.

[0106] In one possible implementation, the special effects picture aligned with the real scene picture is generated based on the posture data of the camera, the target posture data corresponding to the target special effect, and the special effects data of the target special effect, including: determining the relative posture relationship between the camera and the target special effect based on the posture data of the camera and the target posture data of the target special effect; based on the relative posture relationship, mapping the special effects data of the target special effect to the imaging plane corresponding to the real scene picture to obtain a special effects picture aligned with the real scene picture, and the position of the target special effect in the special effects picture matches the position of the target object in the real scene picture.

[0107] In one possible implementation, the real scene includes an entity in front of the screen, and the tracking data also includes posture data of the entity; wherein, synthesizing the special effects picture with the real scene picture to obtain a synthesized picture includes: determining the occlusion relationship between the entity and the target special effect under the camera perspective based on the posture data of the camera, the posture data of the entity, and the target posture data of the target special effect, the occlusion relationship being used to indicate the distance of the entity from the camera relative to the target special effect; and synthesizing the special effects picture with the real scene picture based on the occlusion relationship to obtain a synthesized picture.

[0108] In one possible implementation, the special effects picture and the real scene picture are synthesized according to the occlusion relationship to obtain a synthesized picture, including: when the occlusion relationship indicates that the entity is closer to the camera relative to the target special effect, according to the entity area occupied by the entity in the real scene picture, the area in the special effects picture that overlaps with the entity area is eliminated to obtain the target special effects picture; and the target special effects picture is synthesized with the real scene picture to obtain a synthesized picture.

[0109] In one possible implementation, synthesizing the special effects picture with the real scene picture according to the occlusion relationship to obtain a synthesized picture includes: when the occlusion relationship indicates that the entity is farther away from the camera relative to the target special effect, synthesizing the special effects picture with the real scene picture to obtain a synthesized picture.

[0110] In a possible implementation, the apparatus further includes: a display module configured to display the synthesized picture; and / or a recording module configured to record the synthesized picture in response to a recording operation on the synthesized picture.

[0111] According to the device of the embodiment of the present disclosure, by using the position data of the target object collected by the motion capture system and the posture change data collected by the IMU, more accurate target posture data of the target special effect can be determined, thereby generating a more accurate special effect picture aligned with the real scene picture, and then synthesizing the special effect picture with the real scene picture, it is possible to achieve efficient and automatic addition of target special effects to the real scene picture collected by the camera to generate a synthetic picture during the virtual shooting process, so as to show the pictures before and after adding the target special effects to the director, producer, special effects production personnel, etc. in real time, so that these personnel can instantly view the combined effect of the added special effects and the actor's performance at the virtual shooting site, thereby improving shooting efficiency. At the same time, the target special effects are added to the real scene picture taken by the camera without being rendered to the screen displaying the virtual scene, which greatly reduces the requirements for rendering performance, effectively reduces the cost of special effects production, and is more economical.

[0112] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0113] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0114] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0115] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0116] Figure 8 FIG1 shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Figure 8 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0117] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.

[0118] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0119] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0120] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0121] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0122] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0123] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0124] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0125] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0126] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0127] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for synthesizing special effects in virtual shooting, characterized in that: include: During the virtual shooting process, the real scene currently captured by the camera, the tracking data currently captured by the motion capture system, and the pose change data of the target object to be added with special effects currently captured by the inertial measurement unit are obtained. The real scene includes the virtual scene displayed on the screen and the real scene in front of the screen. The tracking data includes the pose data of the camera and the position data of the target object. Determining target pose data corresponding to a target special effect to be added according to the position data of the target object and the pose change data; Generate a special effects picture aligned with the real scene picture according to the pose data of the camera, the target pose data corresponding to the target special effect, and the special effects data of the target special effect; synthesizing the special effect picture with the real scene picture to obtain a synthesized picture, wherein the target object in the synthesized picture has the target special effect; The pose change data includes the acceleration and angular velocity of the target object from the last acquisition moment to the current acquisition moment, and the position data of the target object includes the position data of the tracking point set on the target object in the global coordinate system; The step of determining target pose data corresponding to a target special effect to be added based on the position data of the target object and the pose change data includes: Determining initial pose data corresponding to the target special effect at the current acquisition moment based on target pose data corresponding to the target special effect at the previous acquisition moment and the pose change data; wherein the target pose data corresponding to the target special effect at the previous acquisition moment is determined based on the position data and pose change data of the target object acquired at the previous acquisition moment; Based on the position data of the tracking point set on the target object in the global coordinate system and the position data of the tracking point set on the target object in the local coordinate system of the target object, the initial pose data corresponding to the target special effect at the current acquisition moment is optimized to obtain the target pose data corresponding to the target special effect at the current acquisition moment.

2. The method according to claim 1, characterized in that The generating, based on the camera's posture data, the target posture data corresponding to the target special effect, and the special effect data of the target special effect, a special effect picture aligned with the real scene picture, comprises: Determining a relative posture relationship between the camera and the target special effect based on the posture data of the camera and the target posture data of the target special effect; Based on the relative posture relationship, the special effect data of the target special effect is mapped to the imaging plane corresponding to the real scene picture to obtain a special effect picture aligned with the real scene picture, and the position of the target special effect in the special effect picture matches the position of the target object in the real scene picture.

3. The method according to claim 1 or 2, characterized in that The real scene includes an entity in front of the screen, and the tracking data also includes posture data of the entity; wherein the synthesizing the special effect picture and the real scene picture to obtain a synthesized picture includes: Determining, based on the camera's pose data, the entity's pose data, and the target pose data of the target special effect, an occlusion relationship between the entity and the target special effect from the camera's perspective, the occlusion relationship being used to indicate how far the entity is from the camera relative to the target special effect; According to the occlusion relationship, the special effect picture and the real scene picture are synthesized to obtain a synthesized picture.

4. The method according to claim 3, characterized in that The synthesizing the special effect picture and the real scene picture according to the occlusion relationship to obtain a synthesized picture includes: When the occlusion relationship indicates that the entity is closer to the camera than the target special effect, based on the entity area occupied by the entity in the real scene, the area in the special effect picture that overlaps with the entity area is eliminated to obtain the target special effect picture; The target special effect picture and the real scene picture are synthesized to obtain a synthesized picture.

5. The method according to claim 3, characterized in that The synthesizing the special effect picture and the real scene picture according to the occlusion relationship to obtain a synthesized picture includes: When the occlusion relationship indicates that the entity is farther from the camera than the target special effect, the special effect picture and the real scene picture are synthesized to obtain a synthesized picture.

6. The method according to claim 1, wherein The method further comprises: Displaying the composite picture; and / or, In response to a recording operation on the composite picture, the composite picture is recorded.

7. A device for synthesizing special effects in virtual shooting, characterized in that: include: An acquisition module is used to acquire, during the virtual shooting process, the real scene currently captured by the camera, the tracking data currently captured by the motion capture system, and the pose change data of the target object to be added with special effects currently captured by the inertial measurement unit. The real scene includes the virtual scene displayed on the screen and the real scene in front of the screen. The tracking data includes the pose data of the camera and the position data of the target object; A determination module, configured to determine target posture data corresponding to a target special effect to be added based on the position data of the target object and the posture change data; A generating module, configured to generate a special effects picture aligned with the real scene picture based on the pose data of the camera, the target pose data corresponding to the target special effect, and the special effects data of the target special effect; a synthesis module, configured to synthesize the special effect picture with the real scene picture to obtain a synthesized picture, wherein the target object in the synthesized picture has the target special effect; The pose change data includes the acceleration and angular velocity of the target object from the last acquisition moment to the current acquisition moment, and the position data of the target object includes the position data of the tracking point set on the target object in the global coordinate system; The step of determining target pose data corresponding to a target special effect to be added based on the position data of the target object and the pose change data includes: Determining initial pose data corresponding to the target special effect at the current acquisition moment based on target pose data corresponding to the target special effect at the previous acquisition moment and the pose change data; wherein the target pose data corresponding to the target special effect at the previous acquisition moment is determined based on the position data and pose change data of the target object acquired at the previous acquisition moment; Based on the position data of the tracking point set on the target object in the global coordinate system and the position data of the tracking point set on the target object in the local coordinate system of the target object, the initial pose data corresponding to the target special effect at the current acquisition moment is optimized to obtain the target pose data corresponding to the target special effect at the current acquisition moment.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 6 when executing the instructions stored in the memory.

9. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video processing method and device

    CN112544070A

  • Method for surveying position coordinate

    JP2009097985A

  • Information processing device and method, and computer-readable storage medium

    WO2024078384A1