Virtual reality based video generation method and apparatus, device and medium

By acquiring and fusing depth maps of the real foreground and virtual background, and adjusting the display parameters of video frames, the problem of real cameras being unable to acquire depth information of the virtual background is solved, achieving a natural combination of the virtual background and the real foreground, and improving the display effect.

CN116527863BActive Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-04-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, real cameras cannot acquire depth information of virtual backgrounds, resulting in virtual backgrounds lacking realism, a stiff combination of real foreground and virtual background, and poor display effects.

Method used

By acquiring depth maps of the real foreground and virtual background, fusing the depth maps to generate a fused depth map, adjusting the display parameters of the video frames, increasing the depth-of-field effect, and generating a target video with a depth-of-field effect.

Benefits of technology

It enhances the realism of virtual backgrounds, making the combination of virtual backgrounds and real-world foregrounds in videos more natural and resulting in better display effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116527863B_ABST
    Figure CN116527863B_ABST
Patent Text Reader

Abstract

The application discloses a virtual reality-based video generation method and device, equipment and medium, and relates to the field of virtual reality. The method comprises the following steps: obtaining a target video frame from a video frame sequence, wherein the video frame sequence is obtained by a real camera collecting a target scene, the target scene comprises a real foreground and a virtual background, and the virtual background is displayed on a physical screen in a real environment; obtaining a real foreground depth map and a virtual background depth map of the target video frame; fusing the real foreground depth map and the virtual background depth map to obtain a fused depth map; adjusting display parameters of the target video frame according to the fused depth map to generate a depth-of-field effect diagram of the target video frame; and generating a target video with a depth-of-field effect based on the depth-of-field effect diagram of the target video frame. The target video obtained by the application has a good depth-of-field effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual reality, and in particular to a method, apparatus, device and medium for video generation based on virtual reality. Background Technology

[0002] Virtual production refers to computer-aided production and visual filmmaking methods. Virtual production includes various approaches, such as visualization, performance capture, hybrid virtual production, and live LED wall in-camera production.

[0003] The technology involves placing an LED (Light Emitting Diode) wall behind actors, props, and other real-world foreground elements during filming, projecting a real-time virtual background onto the LED wall. A real-world camera simultaneously captures the actors, props, and the content displayed on the LED wall, inputting the captured data into a computer. The computer then outputs the footage from the real-world camera in real time.

[0004] However, the display quality of each frame captured by a real camera is poor, which in turn leads to poor video quality when shot by a real camera. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for generating video based on virtual reality. The method updates depth information for the virtual background, resulting in a more natural video with better display quality. The technical solution is as follows:

[0006] According to one aspect of this application, a video generation method based on virtual reality is provided, the method comprising:

[0007] The target video frame is obtained from a video frame sequence, which is obtained by capturing the target scene with a real camera. The target scene includes a real foreground and a virtual background, and the virtual background is displayed on a physical screen in the real environment.

[0008] The real foreground depth map and virtual background depth map of the target video frame are obtained. The real foreground depth map includes the depth information from the real foreground to the real camera, and the virtual background depth map includes the depth information from the virtual background to the real camera after being mapped onto the real environment.

[0009] The real foreground depth map and the virtual background depth map are fused to obtain a fused depth map, which includes the depth information of each reference point in the target scene from the real environment to the real camera;

[0010] Adjust the display parameters of the target video frame according to the fused depth map to generate a depth-of-field effect map of the target video frame;

[0011] Based on the depth-of-field effect map of the target video frame, a target video with a depth-of-field effect is generated.

[0012] According to another aspect of this application, a virtual reality-based video generation apparatus is provided, the apparatus comprising:

[0013] The acquisition module is used to acquire target video frames from a video frame sequence, wherein the video frame sequence is obtained by capturing a target scene using a real camera, and the target scene includes a real foreground and a virtual background, wherein the virtual background is displayed on a physical screen in a real environment;

[0014] The acquisition module is further configured to acquire the real foreground depth map and the virtual background depth map of the target video frame. The real foreground depth map includes the depth information from the real foreground to the real camera, and the virtual background depth map includes the depth information from the virtual background to the real camera after being mapped onto the real environment.

[0015] A fusion module is used to fuse the real foreground depth map and the virtual background depth map to obtain a fused depth map, wherein the fused depth map includes the depth information of each reference point in the target scene from the real environment to the real camera;

[0016] The update module is used to adjust the display parameters of the target video frame according to the fused depth map and generate a depth-of-field effect map of the target video frame;

[0017] The update module is also used to generate a target video with a depth effect based on the depth effect map of the target video frame.

[0018] According to another aspect of this application, a computer device is provided, comprising: a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, wherein the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the virtual reality-based video generation method as described above.

[0019] According to another aspect of this application, a computer storage medium is provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the program code being loaded and executed by a processor to implement the virtual reality-based video generation method as described above.

[0020] According to another aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the virtual reality-based video generation method described above.

[0021] The beneficial effects of the technical solutions provided in this application include at least the following:

[0022] A real-world camera captures the target scene, generating a sequence of video frames. The target video frames are then obtained from this sequence, and their depth information is updated to ensure greater accuracy. A target video with a depth-of-field effect is then generated based on these frames. Because the virtual background's depth information is more accurate, the video combining the virtual background and the real foreground appears more natural and has a better display effect. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A schematic diagram of a computer system provided in an exemplary embodiment of this application is shown;

[0025] Figure 2 This illustration shows a schematic diagram of a virtual reality-based depth-of-field effect map generation method provided in an exemplary embodiment of this application;

[0026] Figure 3 A flowchart illustrating a virtual reality-based video generation method provided in an exemplary embodiment of this application is shown.

[0027] Figure 4 A schematic diagram of the interface of a virtual reality-based video generation method provided in an exemplary embodiment of this application is shown;

[0028] Figure 5A schematic diagram of the interface of a virtual reality-based video generation method provided in an exemplary embodiment of this application is shown;

[0029] Figure 6 A schematic diagram of the interface of a virtual reality-based video generation method provided in an exemplary embodiment of this application is shown;

[0030] Figure 7 A schematic diagram of the interface of a virtual reality-based video generation method provided in an exemplary embodiment of this application is shown;

[0031] Figure 8 A flowchart illustrating a virtual reality-based video generation method provided in an exemplary embodiment of this application is shown.

[0032] Figure 9 This illustration shows a schematic diagram of the computation of depth information provided in an exemplary embodiment of this application;

[0033] Figure 10 This illustration shows a schematic diagram of a virtual reality-based depth-of-field effect map generation method provided in an exemplary embodiment of this application;

[0034] Figure 11 This illustration shows a schematic diagram of virtual reality-based video generation and production provided in an exemplary embodiment of this application;

[0035] Figure 12 A schematic diagram of the interface for generating a depth-of-field effect map based on virtual reality provided in an exemplary embodiment of this application is shown.

[0036] Figure 13 A schematic diagram of a computer device provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0038] Depth of field (DOF) refers to the range of relatively sharp images in front of and behind the focal point of a camera. In optics, especially in video recording or photography, it describes the range of distances in space where a clear image can be formed. Camera lenses can only focus light to a fixed distance; images further away from this point will gradually blur. However, within a certain specific distance, the degree of blur is imperceptible to the naked eye; this specific distance is called the depth of field.

[0039] Realistic foreground: Physical objects in a real-world environment, typically including actors and surrounding real-world set elements. Objects close to the camera are used as the foreground for the camera.

[0040] Virtual background: A pre-designed virtual environment. Virtual sets generally include scenes that are difficult to create in reality, as well as fantasy-like scenes. After being processed by a program engine, the results are output to an LED screen, serving as the background for cameras behind the real-world set.

[0041] YUV: A color encoding method used in video processing components. "Y" represents luminance (or luma), which is the grayscale value, while "U" and "V" represent chrominance (or chroma), which describe color and saturation and are used to specify the color of a pixel.

[0042] Intrinsic and extrinsic parameters: These include the camera's intrinsic and extrinsic parameters. Intrinsic parameters are those related to the camera's inherent characteristics, such as focal length and pixel size. Extrinsic parameters are the camera's parameters in the world coordinate system, such as position and rotation direction. Intrinsic and extrinsic parameters are used to map points in the world coordinate system to the pixels captured by the camera.

[0043] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0044] Virtual production technology is used in the production of video content such as movies and television shows. This technology involves setting up an LED screen in the video shooting location to display a virtual background, and placing a real foreground—referring to real-world objects or creatures—in front of the LED screen. A camera simultaneously captures both the real foreground and the virtual background to obtain the video.

[0045] However, in related technologies, real cameras can only acquire background information from the real camera to the LED curtain wall, but cannot acquire depth information of the virtual background, resulting in the virtual background lacking realism and the combination of real foreground and virtual background being awkward.

[0046] Figure 1 A schematic diagram of a computer system provided in an exemplary embodiment of this application is shown. The computer system 100 includes a computer device 110, a physical screen 120, a reality camera 130, and a depth camera 140.

[0047] A first video editing application is installed on computer device 110. This first application can be a small program within an app, a dedicated application, or a web client. A second application for generating a virtual world is also installed on computer device 110. Both the first and second applications can be small programs within an app, dedicated applications, or web clients. The first and second applications can be the same application or different applications. When the first and second applications are the same application, this application can simultaneously perform video editing and virtual world generation functions. When the first and second applications are different applications, data exchange is possible between them. For example, the second application provides the first application with the depth information of each virtual object in the virtual world relative to the virtual camera, and the first application updates the image of the video frame based on the aforementioned depth information.

[0048] The physical screen 120 is used to display a virtual background. The computer device 110 transmits data of the virtual world to the physical screen 120, and the physical screen 120 displays the virtual world as a virtual background.

[0049] The reality camera 130 is used to capture the real foreground 150 and the virtual background displayed on the physical screen 120. The reality camera 130 transmits the captured video to the computer device 110. The reality camera 130 can transmit the captured video to the computer device 110 in real time, or it can transmit the captured video to the computer device 110 at preset intervals.

[0050] Depth camera 140 is used to acquire a depth map of the real foreground 150. Depth camera 140 and real camera 130 are positioned differently. Depth camera 140 transmits the captured depth map of the real foreground 150 to computer device 110, and computer device 110 determines the depth information of the real foreground 150 based on the depth map.

[0051] Figure 2 This illustration shows a schematic diagram of a virtual reality-based depth-of-field rendering method provided in an exemplary embodiment of this application. The method can be... Figure 1 The computer system 100 shown is executing.

[0052] like Figure 2As shown, the real-world camera 210 captures the target scene to obtain a target video frame 220. The depth camera 250 also captures the target scene to obtain a depth map. Using a pixel-matching method, the pixels in the target video frame 220 and the depth map are matched, and the depth information from the depth map is provided to the target video frame 220 to obtain a real-world foreground depth map 260. On the other hand, the virtual camera 230 acquires the depth information of the rendered target corresponding to the virtual background to obtain a virtual background depth map 240. The depth information from the real-world foreground depth map 260 and the virtual background depth map 240 are fused to obtain a fused depth map 270. Based on the fused depth map 270, a depth-of-field effect is added to the target video frame 220 to obtain a depth-of-field effect map 280.

[0053] Figure 3 This illustration shows a flowchart of a virtual reality-based video generation method provided in an exemplary embodiment of this application. The method can be... Figure 1 The computer system 100 shown executes the method, which includes:

[0054] Step 302: Obtain the target video frame from the video frame sequence. The video frame sequence is obtained by capturing the target scene with a real camera. The target scene includes a real foreground and a virtual background. The virtual background is displayed on a physical screen in the real environment.

[0055] Alternatively, the camera transmits the captured video frame sequence to a computer device.

[0056] The target video frame is any image in the video frame sequence. For example, the video frame sequence includes 120 frames, and the 45th frame is randomly selected as the target video frame. Optionally, the video frame sequence is played at a preset frame rate to generate a video. For example, every 24 frames in the video frame sequence constitutes one second of video.

[0057] Optionally, the realistic foreground includes at least one of real objects, real creatures, and real people. For example, such as... Figure 4 As shown, the foreground 401 is an actor on the video shooting set.

[0058] A virtual background refers to virtual content displayed on a screen. Optionally, virtual content includes at least one of virtual environments, virtual characters, virtual objects, virtual props, and virtual images. This application does not specifically limit the content displayed as a virtual background. For example, such as... Figure 4 As shown, virtual background 402 is displayed on display screen 403. Virtual background 402 is a virtual image of the city.

[0059] Step 304: Obtain the real foreground depth map and virtual background depth map of the target video frame. The real foreground depth map includes the depth information from the real foreground to the real camera, and the virtual background depth map includes the depth information from the virtual background mapped to the real environment to the real camera.

[0060] Optionally, a real-world foreground depth map is determined using a depth map provided by a depth camera and a target video frame. For example, depth information of each pixel within the target video frame is obtained using the depth map provided by the depth camera. This depth information represents the distance from the real-world reference point corresponding to each pixel in the target video frame to the real-world camera. If the depth value of a first pixel is greater than a first depth threshold, the first pixel is determined to belong to the virtual background. If the depth value of a second pixel is greater than a second depth threshold, the second pixel is determined to belong to the real-world foreground. The first depth threshold is not less than the second depth threshold, and both the first and second depth thresholds can be set by a technician.

[0061] Optionally, a virtual background depth map is generated using a virtual camera, which is used to capture a rendered target corresponding to the virtual background in a virtual environment. For example, a computer device can obtain the distance from the rendered target to the virtual camera, convert this distance into a real-world distance, and obtain the virtual background depth map.

[0062] For example, such as Figure 5 As shown, Figure 5 The real-world foreground depth map is shown. The real-world foreground depth map includes depth information from the real-world foreground to the real-world camera, but does not include depth information from the virtual background mapped onto the real-world environment to the real-world camera.

[0063] For example, such as Figure 6 As shown, Figure 6 A virtual background depth map is shown. The virtual background depth map includes the depth information from the virtual background mapped onto the real environment to the real camera, but does not include the depth information from the real foreground to the real camera.

[0064] Step 306: Fuse the real foreground depth map and the virtual background depth map to obtain a fused depth map. The fused depth map includes the depth information of each reference point in the target scene from the real environment to the real camera.

[0065] Optionally, the first depth information of each pixel in the real foreground depth map is updated based on the second depth information of each pixel in the virtual background depth map to obtain a fused depth map.

[0066] The real-world foreground depth map includes a foreground region corresponding to the real-world foreground and a background region corresponding to the virtual background. This application embodiment can update the first depth information of pixels within the background region, or it can update the first depth information of pixels within the foreground region.

[0067] For example, based on the second depth information of the second pixel in the virtual background depth map that belongs to the background region, the first depth information of the first pixel in the real foreground depth map that belongs to the background region is updated to obtain a fused depth map.

[0068] For example, update the third depth information of the third pixel point belonging to the foreground region within the real-world foreground depth map, where the third pixel point is the pixel point corresponding to the target object.

[0069] For example, such as Figure 7 As shown, the fused depth map includes depth information from both the real foreground 701 and the virtual background 702. (Comparison) Figure 5 and Figure 7 Thus, compared to the real-world foreground depth map, the fused depth map also provides depth information for the virtual background 702.

[0070] Step 308: Adjust the display parameters of the target video frame according to the fused depth map to generate a depth-of-field effect map of the target video frame.

[0071] Optionally, the display parameters include at least one of sharpness, brightness, grayscale, contrast, and saturation. Optionally, the display parameters of the target video frame can be adjusted according to the actual needs of the technicians. For example, the brightness of the pixels corresponding to the virtual background in the fused depth map can be increased according to the actual needs of the technicians.

[0072] Optionally, a distance range is determined based on the preset aperture or preset focal length of the real camera. The distance range represents the distance from the reference point corresponding to the pixel with a sharpness greater than the sharpness threshold to the real camera. Based on the fused depth map and the distance range, the sharpness of each pixel in the target video frame is adjusted to generate a depth-of-field effect map of the target scene.

[0073] Optionally, based on the focus distance and fusion depth map of the real camera, the sharpness of the area corresponding to the virtual background within the target video frame is adjusted to generate the depth-of-field effect map of the target scene.

[0074] Optionally, the sharpness of each pixel within the target video frame can be adjusted according to preset conditions. These preset conditions are determined by technicians based on actual needs. For example, the sharpness of pixels within a preset area in the target video frame can be adjusted; this preset area can be set by the technicians themselves.

[0075] Step 310: Generate a target video with depth-of-field effect based on the depth-of-field effect map of the target video frame.

[0076] Optionally, the depth-of-field effect map of the target video frame is played at a preset frame rate to obtain a target video with a depth-of-field effect.

[0077] For example, the target video frame includes at least two video frames. The depth-of-field effect maps of the target video frames are arranged in chronological order to obtain a target video with a depth-of-field effect. Optionally, the depth-of-field effect maps of consecutive target video frames are arranged in chronological order to generate a target video with a depth-of-field effect.

[0078] In summary, in this embodiment, a real-world camera captures the target scene, generating a video frame sequence. Then, a target video frame is obtained from the video frame sequence, and its depth information is updated to make the depth information more accurate. A target video with a depth-of-field effect is then generated based on the target video frame. Because the depth information of the virtual background is more accurate, the video composed of the virtual background and the real foreground is more natural and has a better display effect.

[0079] Furthermore, the depth-of-field effect map is generated using parameters such as the focusing distance, focal length, and aperture of a real camera. Therefore, the virtual background in the depth-of-field effect map can simulate the shooting effect of a real camera, making the virtual background display more natural and closer to real objects.

[0080] In the following embodiments, the depth information of the background region in the real-world foreground depth map is updated to make the depth information of the background region more accurate. Taking the real-world scene depth map as an example, which includes two optional implementation methods, the real-world scene depth map can be obtained by setting up a depth camera or by setting up an auxiliary camera. Furthermore, the sharpness of the virtual background pixels can be updated using parameters of the real-world camera, such as focusing distance, aperture, and focal length.

[0081] Figure 8 This illustration shows a flowchart of a virtual reality-based video generation method provided in an exemplary embodiment of this application. The method can be derived from... Figure 1 The computer system 100 shown executes the method, which includes:

[0082] Step 801: Obtain the target video frame from the video frame sequence.

[0083] Alternatively, the camera transmits the captured video frame sequence to a computer device.

[0084] The target video frame is any image in the video frame sequence. For example, the video frame sequence includes 120 frames, and the 45th frame is randomly selected as the target video frame. Optionally, the video frame sequence is played at a preset frame rate to generate a target video with a depth-of-field effect. For example, every 24 frames in the video frame sequence constitute one second of video.

[0085] Optionally, the realistic foreground includes at least one of a real object, a real creature, or a real person.

[0086] A virtual background refers to virtual content displayed on a screen. Optionally, virtual content includes at least one of virtual environments, virtual characters, virtual objects, virtual props, and virtual images. This application does not specifically limit the content displayed as a virtual background.

[0087] Step 802: Obtain the real foreground depth map of the target video frame.

[0088] In one alternative implementation, depth information of the real foreground is acquired using a depth camera, and the method may include the following steps:

[0089] 1. Generate spatial offset information between the depth camera and the real camera based on the intrinsic and extrinsic parameters of the depth camera and the real camera.

[0090] Optionally, the intrinsic and extrinsic parameters of the depth camera include intrinsic parameters and extrinsic parameters. Intrinsic parameters are parameters related to the characteristics of the depth camera itself, such as focal length and pixel size. Extrinsic parameters are parameters of the depth camera in the world coordinate system, such as camera position and rotation direction.

[0091] Optionally, the intrinsic and extrinsic parameters of the real-world camera include intrinsic parameters and extrinsic parameters. Intrinsic parameters are those related to the characteristics of the real-world camera itself, such as focal length and pixel size. Extrinsic parameters are those of the real-world camera in the world coordinate system, such as the camera's position and rotation direction.

[0092] In one optional implementation, the spatial offset information refers to the mapping relationship between the camera coordinate system of the depth camera and the camera coordinate system of the real-world camera. For example, based on the intrinsic and extrinsic parameters of the depth camera, the depth mapping relationship of the depth camera is determined, which refers to the mapping relationship from the camera coordinate system of the depth camera to the real-world coordinate system; based on the intrinsic and extrinsic parameters of the real-world camera, the real-world mapping relationship of the real-world camera is determined, which refers to the mapping relationship from the camera coordinate system of the real-world camera to the real-world coordinate system; through the real-world coordinate system, the mapping relationship between the camera coordinate system of the depth camera and the camera coordinate system of the real-world camera is determined, thus obtaining the spatial offset information between the depth camera and the real-world camera.

[0093] Optionally, the depth camera and the reality camera may capture the target scene from different angles. Optionally, the depth camera and the reality camera may be positioned differently.

[0094] 2. Obtain the depth map captured by the depth camera.

[0095] Optionally, the depth camera transmits the acquired depth map to a computer device via a wired or wireless connection.

[0096] 3. Based on the spatial offset information, the depth information of the depth map is mapped onto the target video frame to obtain the real foreground depth map.

[0097] Since spatial offset information refers to the mapping relationship between the camera coordinate system of the depth camera and the camera coordinate system of the real camera, the correspondence between the depth map and the pixels on the target video frame can be obtained. Based on this correspondence, the depth information of each pixel on the depth map is mapped to each pixel on the target video frame to obtain the real foreground depth map.

[0098] In one alternative implementation, depth information of the real foreground is acquired via another reference camera, and the method may include the following steps:

[0099] 1. Based on the intrinsic and extrinsic parameters of the reference camera, obtain the first mapping relationship. The first mapping relationship is used to represent the mapping relationship between the camera coordinate system of the reference camera and the real coordinate system. The reference camera is used to capture the target scene from a second angle, which is different from the first angle.

[0100] Optionally, the intrinsic and extrinsic parameters of the reference camera include intrinsic parameters and extrinsic parameters. Intrinsic parameters are parameters related to the characteristics of the reference camera itself, such as focal length and pixel size. Extrinsic parameters are parameters of the reference camera in the world coordinate system, such as camera position and rotation direction.

[0101] Optionally, the first mapping relationship is also used to represent the positional correspondence between pixels on a reference image captured by a reference camera and real-world points. For example, the coordinates of pixel A on the reference image are (x1, x2), the first mapping relationship satisfies the function f, which is generated based on the intrinsic and extrinsic parameters of the reference camera, while the real-world point corresponding to pixel A in the real environment is (y1, y2, y3) = f(x1, x2).

[0102] 2. Based on the intrinsic and extrinsic parameters of the real camera, obtain the second mapping relationship, which is used to represent the mapping relationship between the camera coordinate system and the real coordinate system.

[0103] Optionally, the first mapping relationship is also used to represent the positional correspondence between pixels on the target video frame captured by the real camera and real-world points. For example, if the coordinates of pixel B on the reference image are (x3, x4), the first mapping relationship satisfies the following functional relationship. Functional Relationship It is generated based on the intrinsic and extrinsic parameters of a real camera, while the real-world point corresponding to pixel B in the real environment is...

[0104] 3. Reconstruct the reference image captured by the reference camera according to the first mapping relationship to obtain the reconstructed reference image.

[0105] Optionally, each pixel in the reference image is mapped to the real-world environment according to the first mapping relationship to obtain a reconstructed reference image. The reconstructed reference image includes the positions of reference points corresponding to each pixel in the reference image in the real-world environment.

[0106] 4. Reconstruct the target video frames captured by the real camera according to the second mapping relationship to obtain the reconstructed target scene image.

[0107] Optionally, each pixel in the target video frame is mapped to the real-world environment according to the second mapping relationship to obtain a reconstructed target scene image. The reconstructed target scene image includes the positions of reference points corresponding to each pixel in the target video frame in the real-world environment.

[0108] 5. Based on the parallax between the reconstructed reference image and the reconstructed target scene image, determine the depth information of each pixel within the target video frame to obtain the real foreground depth map.

[0109] Optionally, to facilitate the calculation of disparity, the reconstructed reference image and the reconstructed target scene image are mapped onto the same plane, and two pixels corresponding to the same real point on the reconstructed reference image and the reconstructed target scene image are determined; the depth information of each pixel in the target video frame is determined based on the disparity of the aforementioned two pixels, and a real foreground depth map is obtained.

[0110] For example, such as Figure 9 As shown, assuming the real camera and the reference camera are located on the same plane, the reconstructed reference image and the reconstructed target scene image are pre-mapped onto the plane containing the X-axis, such that the distance from reference point 903 to the X-axis is the same as the distance from reference point 903 to the X-axis. Reference point 903 forms pixel 901 on the target video frame through the center point 904 of the real camera, and reference point 903 forms pixel 902 on the reference image through the center point 905 of the reference camera. Figure 9In this context, f represents the focal length of the real camera and the reference camera, which are the same. z represents the distance from reference point 903 to the real camera, i.e., the depth information from reference point 903 to the real camera. x represents the distance from reference point 903 to the Z-axis. x1 represents the position of pixel 901 on the target video frame. xr represents the position of pixel 902 on the reference image. Then, through... Figure 9 The following equation can be obtained from similar triangles in the equation:

[0111]

[0112] Therefore, according to the above equation, we can obtain z = f*b / (x1-xr). Where (x1-xr) is the parallax.

[0113] It should be noted that the embodiments of this application do not specifically limit the method for obtaining the real foreground depth map of the target video frame. In addition to the two optional methods mentioned above, those skilled in the art can choose other methods to obtain the real foreground depth map of the target video frame according to actual needs, which will not be elaborated here.

[0114] Step 803: Obtain the virtual background depth map of the target video frame.

[0115] In one alternative implementation, a virtual background depth map is acquired using a virtual camera. This method may include the following steps:

[0116] 1. Obtain the rendering target corresponding to the virtual background in the virtual environment.

[0117] Optionally, the physical screen area captured by the real camera is determined based on the shooting angle and position of the real camera; the display content on the physical screen area is determined based on the physical screen area; and the rendering target in the virtual environment is obtained based on the aforementioned display content. For example, if the size of the physical screen is 30×4 (m×m), and the real camera is set at a position 30 meters away from the physical screen, with an angle of 90 degrees between the real camera and the physical screen, it is determined that the real camera has captured a portion of the physical screen, and the size of this portion of the physical screen is 20×3 (m×m).

[0118] 2. Generate a rendering target depth map of the rendering target in the virtual environment. The rendering target depth map includes virtual depth information, which is used to represent the distance from the rendering target to the virtual camera in the virtual environment.

[0119] For example, a computer device stores various data about a virtual environment, including the distances from a virtual camera to various points in the virtual environment. After determining the rendering target, the distance from the rendering target to the virtual camera can be directly determined to obtain a depth map of the rendering target.

[0120] 3. Convert the virtual depth information in the rendered target depth map into real depth information to obtain a virtual background depth map. The real depth information is used to represent the distance from the rendered target mapped to the real environment to the real camera.

[0121] Optionally, a first position of the virtual camera in the virtual environment is determined; a second position of the real camera in the real environment is determined; and the virtual depth information in the rendering target depth map is converted into real depth information based on the positional relationship between the first and second positions. For example, the virtual camera is located at position A in the virtual environment, and the real camera is located at position B in the virtual environment. In this case, the distance from point 1 in the virtual environment to the virtual camera is x. Let the distance from the rendering target mapped to the real environment to the real camera be y, then y = f(x), where f represents a functional relationship.

[0122] Optionally, the target depth map is an image obtained from the perspective of a virtual camera, while the virtual background depth map is an image obtained from the perspective of a real camera. In this case, coordinate transformation of the pixels in the target depth map is required. For example, the first position coordinates of the pixels in the target depth map on the physical screen are determined; based on the first position coordinates and the intrinsic and extrinsic parameters of the real camera, the first position coordinates are mapped to second position coordinates, which are position coordinates on the virtual background depth map; the depth information of the aforementioned pixels is filled into the pixels corresponding to the second position coordinates to obtain the virtual background depth map.

[0123] It should be noted that steps 802 and 803 are not in any particular order. Steps 802 and 803 can be executed simultaneously, or step 802 can be executed first and then step 803, or step 803 can be executed first and then step 802.

[0124] Step 804: In the virtual background depth map, determine the j-th second pixel point that corresponds to the i-th first pixel point belonging to the background region in the real foreground depth map.

[0125] Where i and j are positive integers, and the initial value of i can be any integer.

[0126] Optionally, based on the intrinsic and extrinsic parameters of the real camera, the screen coordinates of the i-th first pixel in the background region on the physical screen are determined; in the virtual environment, the coordinates of the virtual point corresponding to the i-th first pixel are determined based on the screen coordinates; based on the intrinsic and extrinsic parameters of the virtual camera, the coordinates of the virtual point are mapped onto the virtual background depth map to obtain the j-th second pixel.

[0127] Optionally, the position coordinates of the j-th second pixel are determined in the virtual background depth map based on the position coordinates of the i-th first pixel in the real foreground depth map. For example, if the position coordinates of the i-th first pixel in the real foreground depth map are (4, 6), then the position coordinates of the j-th second pixel in the virtual background depth map are also (4, 6).

[0128] Step 805: Use the second depth information of the j-th second pixel to replace the first depth information of the i-th first pixel in the real foreground depth map.

[0129] For example, the depth value in the second depth information of the j-th second pixel is used to replace the depth value in the first depth information of the i-th first pixel in the real foreground depth map. For instance, in the real foreground depth map, if the depth value of the i-th first pixel is 20 and the depth value of the j-th second pixel corresponding to the i-th first pixel is 80, then the depth value of the i-th first pixel is modified to 80.

[0130] In another optional implementation of this application, before replacing the first depth information of the i-th first pixel with the second depth information of the j-th second pixel, the second depth information of the j-th second pixel can be modified. For example, the depth value in the second depth information can be modified to a first target depth value, or the depth value in the second depth information can be increased by a second target depth value, or the depth value in the second depth information can be decreased by a third target depth value. The first target depth value, the second target depth value, and the third target depth value can all be set by a person skilled in the art. For example, assuming there are 3 second pixels with depth values ​​of 20, 43, and 36 respectively, the depth values ​​of these 3 second pixels can be uniformly set to 40.

[0131] For example, please refer to Figure 5 and Figure 7 , Figure 7 Compared to Figure 5 , Figure 7 The depth information of the virtual background has been replaced.

[0132] Step 806: Update i to i+1, repeat the above two steps until the first depth information of each first pixel in the background region of the real foreground depth map is traversed to obtain the fused depth map.

[0133] If the background region includes at least two first pixels, then iterate through each first pixel in the real foreground depth map that belongs to the background region until the first depth information of each first pixel in the background region is replaced with the second depth information of the second pixel.

[0134] Step 807: Adjust the display parameters of the target video frame according to the fused depth map to generate a depth-of-field effect map of the target video frame.

[0135] Optionally, the display parameters include at least one of sharpness, brightness, grayscale, contrast, and saturation.

[0136] Optionally, a distance range is determined based on the preset aperture or preset focal length of the real camera. This distance range represents the distance from the reference point corresponding to a pixel with a sharpness greater than a sharpness threshold to the real camera. Based on the fused depth map and the distance range, the sharpness of each pixel within the target video frame is adjusted to generate a depth-of-field effect map of the target video frame. For example, if the distance range is [0, 20], then the sharpness of pixels within the distance range is set to 100%, and the sharpness of pixels outside the distance range is set to 40%.

[0137] Optionally, based on the focus distance and fusion depth map of the real camera, the sharpness of the area corresponding to the virtual background within the target video frame is adjusted to generate a depth-of-field effect map of the target video frame.

[0138] Optionally, the sharpness of each pixel within the target video frame can be adjusted according to preset conditions. These preset conditions are determined by technicians based on actual needs. For example, the sharpness of pixels within a preset area in the target video frame can be adjusted; this preset area can be set by the technicians themselves.

[0139] Step 808: Generate a target video with depth-of-field effect based on the depth-of-field effect map of the target video frame.

[0140] Optionally, the depth-of-field effect map of the target video frame is played at a preset frame rate to obtain a target video with a depth-of-field effect.

[0141] For example, the target video frame includes at least two video frames. The depth-of-field effect maps of the target video frames are arranged in chronological order to obtain a target video with a depth-of-field effect. Optionally, the depth-of-field effect maps of consecutive target video frames are arranged in chronological order to generate a target video with a depth-of-field effect.

[0142] In summary, in this embodiment, a real-world camera captures the target scene, generating a video frame sequence. Then, a target video frame is obtained from the video frame sequence, and its depth information is updated to make the depth information more accurate. A target video with a depth-of-field effect is then generated based on the target video frame. Because the depth information of the virtual background is more accurate, the video composed of the virtual background and the real foreground is more natural and has a better display effect.

[0143] Furthermore, this embodiment provides multiple methods for acquiring the real-world foreground depth map, allowing technicians to adjust the acquisition method according to actual needs. It can acquire depth information of the real-world foreground not only through a depth camera but also through two real-world cameras, increasing the flexibility of the solution. Moreover, because the depth information of the background region in the fused depth map is obtained by updating a virtual background depth map, and the depth information of the virtual background depth map is generated by a virtual camera capturing the virtual environment, this method yields more accurate depth information, resulting in a depth-of-field effect that better meets actual requirements and provides superior performance.

[0144] In the following embodiments, considering that in some scenarios the real-world foreground is not easily movable, but it is desirable to modify the display effect of the real-world foreground in the target video, the depth information of the real-world foreground can also be adjusted to meet preset requirements.

[0145] Figure 10 This illustration shows a flowchart of a depth information update method based on virtual reality provided in an exemplary embodiment of this application. The method can be derived from... Figure 1 The computer system 100 shown executes the method, which includes:

[0146] Step 1001: In the foreground region of the real-world foreground depth map, determine the third pixel point belonging to the target object.

[0147] The third pixel is the pixel corresponding to the target object. The target object is an object in the real environment.

[0148] Optionally, a real-world foreground depth map is determined using a depth map provided by a depth camera and a target video frame. For example, depth information of each pixel within the target video frame is obtained using the depth map provided by the depth camera. This depth information represents the distance from the real-world reference point corresponding to each pixel in the target video frame to the real-world camera. If the depth value of a first pixel is greater than a first depth threshold, the first pixel is determined to belong to the virtual background. If the depth value of a second pixel is greater than a second depth threshold, the second pixel is determined to belong to the real-world foreground. The first depth threshold is not less than the second depth threshold, and both the first and second depth thresholds can be set by a technician.

[0149] Optionally, pixels within the depth threshold range in the foreground region can be designated as the third pixel. The depth threshold range can be set by the technician.

[0150] Optionally, a pixel within the foreground region that belongs to the target object region can be designated as the third pixel. The target object region can be set by the technician. Alternatively, the target object region can be defined within the foreground region using a selection box.

[0151] The third pixel is any pixel in the foreground area.

[0152] Step 1002: In response to the depth value update instruction, update the depth value of the third pixel.

[0153] Optionally, in response to a depth value setting command, the depth value of the third pixel is set to a first preset depth value. The first preset depth value can be set by a technician according to actual needs. For example, in some scenarios, the target object is difficult to move, or a larger depth value is desired, but space constraints prevent moving the target object to the desired location. In such cases, the depth value of the third pixel corresponding to the target object can be uniformly set to the first preset depth value, ensuring that the depth information of the target object in the depth-of-field rendering matches the actual requirements.

[0154] Optionally, in response to a depth value increase command, a second preset depth value is added to the depth value of the third pixel. This second preset depth value can be set by a technician according to actual needs. Optionally, in response to a depth value decrease command, a third preset depth value is decreased to the depth value of the third pixel. This third preset depth value can be set by a technician according to actual needs. For example, in some scenarios, it may be desirable for the target object to exchange positions with other objects, or for the target object to move in front of or behind other objects. In such cases, the depth value of the third pixel corresponding to the target object can be changed so that the generated depth-of-field effect map can reflect the positional relationship between the target object and other objects. For example, in the foreground, there is a reference object. The depth value of the pixel corresponding to the reference object is 10, indicating that the reference object is 10 meters away from the camera. The depth value of the pixel corresponding to the target object is 15, indicating that the target object is 15 meters away from the camera. However, the actual requirement is that the depth map of the foreground should show that the distance between the target object and the camera is less than the distance between the reference object and the camera. Therefore, we can choose to reduce the depth value of the third pixel corresponding to the target object by 8. Then, the depth value of the third pixel corresponding to the target object is 7, which can meet the above requirement.

[0155] In a specific example, there is a tree in the foreground, 20 meters away from the camera. Since moving the tree is inconvenient, but we want the target video to show that the tree is 40 meters away, technicians can input a depth value setting command to uniformly set the depth value of the third pixel corresponding to the tree to 40. In this way, both the resulting depth map and the target video display will show that the tree is 40 meters away from the camera. Moreover, achieving this display effect does not require moving the tree, making the operation simple and efficient.

[0156] In another specific example, there are tree A and tree B in the real foreground. Tree A is 20 meters away from the camera, and tree B is 25 meters away. The technician wants tree A to appear behind tree B in the target video, but directly moving tree A or tree B is not a practical solution. In this case, the technician can use a depth setting command to directly set the depth value of tree A to 30. In this case, the target video will display tree A 30 meters away from the camera, while tree B will remain 25 meters away, achieving the desired effect of tree A being behind tree B. Alternatively, the technician can use a depth increase command to increase the depth value of tree A by 15, resulting in a depth value of 35. In this case, the target video will also display tree A 35 meters away from the camera, while tree B will remain 25 meters away, again achieving the desired effect of tree A being behind tree B. Alternatively, technicians can reduce the depth value of tree B by 10 using a depth reduction command, resulting in a depth value of 15 for tree B. In this case, in the display effect of the target video, tree A is still 20 meters away from the real camera, and tree B is 15 meters away from the real camera, satisfying the display effect of tree A being behind tree B.

[0157] In summary, this embodiment can modify each third pixel in the foreground region so that the depth information of the pixels in the foreground region meets the requirements. This not only reduces the movement of objects in the real foreground but also makes the depth information of the real foreground more accurate.

[0158] Moreover, when it is inconvenient to move the target object, the depth information of each third pixel in the foreground area can be directly adjusted according to the actual needs of the technicians, so that the foreground area in the target video or depth-of-field effect map can present the display effect desired by the technicians.

[0159] Figure 11This illustration shows a schematic diagram of a virtual reality-based depth-of-field rendering method provided in an exemplary embodiment of this application. The method is implemented as a plugin in UE4 (Unreal Engine 4). Optionally, the method can also be implemented as a plugin in Unity3D (a real-time 3D interactive content creation and operation platform, belonging to the creation engine and development tools). This application embodiment does not specifically limit the application platform of the method. Figure 11 In the illustrated embodiment, the method for generating depth-of-field effect maps based on virtual reality is implemented through a virtual depth-of-field creation plugin 1101. The specific steps are as follows:

[0160] 1. Plugin implements two-thread synchronous processing.

[0161] 1.1 Realistic Foreground Depth Processing Thread: This thread processes data from both the real-world camera 1102 and the depth camera 1104. This includes converting the original YUV data to RGBA and using OpenCV (a cross-platform computer vision and machine learning library) to process the data, combining the depth information provided by the depth camera 1104 to obtain the depth information of the real-world foreground from the real-world camera 1102. Within this thread, a shared texture from DX11 is used to copy the virtual background depth map to the current thread, fusing them into a fused depth map 1107 containing depth information from both the real-world foreground and the virtual background. Finally, a Compute Shader (a computer technology that enables parallel processing on a GPU, i.e., Graphics Processing Unit) is used to obtain the depth-of-field effect 1108 based on the data captured by the real-world camera 1102 and the fused depth map 1107.

[0162] 1.2 Virtual Background Depth Processing Thread: Obtain depth information from the Render Target corresponding to the virtual camera 1103, generate a virtual background depth map 1106, and copy it to the shared texture in the real foreground depth processing thread.

[0163] 2. In the real-world camera data processing thread, the selection scheme for the real-world foreground depth map is first determined. This application's embodiments include the following two optional implementation methods:

[0164] 2.1 Selecting a depth camera: Calibrate the intrinsic and extrinsic parameters of depth camera 1104. Based on the intrinsic and extrinsic parameters of depth camera 1104 and reality camera 1102, map the depth map captured by depth camera 1104 onto reality camera 1102 to obtain the depth map of reality camera 1102 in the corresponding scene, i.e., obtain the real foreground depth map 1105.

[0165] 2.2. Selecting an auxiliary camera: Calibrate the intrinsic and extrinsic parameters of the auxiliary camera, and perform stereo correction on the image based on the auxiliary camera's intrinsic and extrinsic parameters to obtain a corrected mapping relationship. Perform stereo correction on the image based on the intrinsic and extrinsic parameters of the real camera 1102 to obtain a corrected mapping relationship. Then, in the data processing of each frame, use the aforementioned two mapping relationships to reconstruct the data provided by the two cameras respectively, and then generate a disparity map. Based on the disparity map, obtain the depth information of the target video frame to obtain the real foreground depth map 1105.

[0166] 3. In the virtual background depth processing thread, after obtaining the rendering target captured by the virtual camera 1103, the depth information is obtained from the rendering target, and this depth information is converted into the corresponding real linear distance to obtain the virtual background depth map 1106. Then, it is synchronously copied to another texture in the real foreground depth processing thread.

[0167] 4. In the real-world camera data processing thread, the data is fused into a fusion depth map 1107. This allows the image from the real-world camera to determine the sharpness of the virtual background pixels based on the focus distance; and to determine the maximum and minimum distances for displaying sharp pixels, as well as the degree of blur in blurred areas, based on the set aperture or focal length. By matching the fusion depth map 1107 with these parameters, the display method of the pixels can be determined, generating the final depth-of-field effect map 1108.

[0168] Please refer to Figure 12 This illustration shows a schematic diagram of a virtual reality-based video generation apparatus according to an embodiment of this application. The above functions can be implemented in hardware or by hardware executing corresponding software. The apparatus 1200 includes:

[0169] The acquisition module 1201 is used to acquire target video frames from a video frame sequence, wherein the video frame sequence is obtained by capturing a target scene using a real camera, and the target scene includes a real foreground and a virtual background, wherein the virtual background is displayed on a physical screen in a real environment;

[0170] The acquisition module 1201 is further configured to acquire the real foreground depth map and the virtual background depth map of the target video frame. The real foreground depth map includes the depth information from the real foreground to the real camera, and the virtual background depth map includes the depth information from the virtual background to the real camera after being mapped onto the real environment.

[0171] The fusion module 1202 is used to fuse the real foreground depth map and the virtual background depth map to obtain a fused depth map. The fused depth map includes the depth information of each reference point in the target scene from the real environment to the real camera.

[0172] The update module 1203 is used to adjust the display parameters of the target video frame according to the fused depth map and generate a depth effect map of the target video frame.

[0173] The update module 1203 is also used to generate a target video with a depth effect based on the depth effect map of the target video frame.

[0174] In an optional design of this application, the real foreground depth map includes a background region corresponding to the virtual background; the fusion module 1202 is further configured to update the first depth information of the first pixel in the real foreground depth map belonging to the background region according to the second depth information of the second pixel in the virtual background depth map belonging to the background region, so as to obtain the fused depth map.

[0175] In an optional design of this application, the acquisition module 1201 is further configured to: determine, in the virtual background depth map, a j-th second pixel corresponding to the i-th first pixel in the real foreground depth map that belongs to the background region, where i and j are positive integers; use the second depth information of the j-th second pixel to replace the first depth information of the i-th first pixel in the real foreground depth map; update i to i+1, and repeat the above two steps until the first depth information of each first pixel in the real foreground depth map belonging to the background region is traversed to obtain the fused depth map.

[0176] In an optional design of this application, the acquisition module 1201 is further configured to determine the screen coordinates of the i-th first pixel in the background region on the physical screen based on the intrinsic and extrinsic parameters of the real camera; in the virtual environment, determine the coordinates of the virtual point corresponding to the i-th first pixel based on the screen coordinates; and map the coordinates of the virtual point onto the virtual background depth map based on the intrinsic and extrinsic parameters of the virtual camera to obtain the j-th second pixel. The virtual camera is used to capture a rendering target corresponding to the virtual background in the virtual environment.

[0177] In an optional design of this application, the real foreground depth map includes a foreground region corresponding to the real foreground; the fusion module 1202 is further configured to update the third depth information of the third pixel point belonging to the foreground region within the real foreground depth map, wherein the third pixel point is a pixel point corresponding to the target object.

[0178] In an optional design of this application, the fusion module 1202 is further configured to determine a third pixel belonging to the target object in the foreground region of the real foreground depth map; and update the depth value of the third pixel in response to a depth value update instruction.

[0179] In an optional design of this application, the fusion module 1202 is further configured to set the depth value of the third pixel to a first preset depth value according to a depth value setting instruction; or, to increase the depth value of the third pixel by a second preset depth value according to a depth value increase instruction; or, to decrease the depth value of the third pixel by a third preset depth value according to a depth value decrease instruction.

[0180] In an optional design of this application, the acquisition module 1201 is further configured to generate spatial offset information between the depth camera and the real camera based on the intrinsic and extrinsic parameters of the depth camera and the intrinsic and extrinsic parameters of the real camera; acquire the depth map captured by the depth camera; and map the depth information of the depth map onto the target video frame based on the spatial offset information to obtain the real foreground depth map.

[0181] In an optional design of this application, the real-world camera is used to capture the target scene from a first angle; the acquisition module 1201 is further used to acquire a first mapping relationship based on the intrinsic and extrinsic parameters of a reference camera, the first mapping relationship representing the mapping relationship between the camera coordinate system of the reference camera and the real-world coordinate system, the reference camera being used to capture the target scene from a second angle, the second angle being different from the first angle; acquire a second mapping relationship based on the intrinsic and extrinsic parameters of the real-world camera, the second mapping relationship representing the mapping relationship between the camera coordinate system of the real-world camera and the real-world coordinate system; reconstruct the reference image captured by the reference camera based on the first mapping relationship to obtain a reconstructed reference image; reconstruct the target video frame captured by the real-world camera based on the second mapping relationship to obtain a reconstructed target scene image; determine the depth information of each pixel in the target video frame based on the parallax between the reconstructed reference image and the reconstructed target scene image to obtain the real-world foreground depth map.

[0182] In an optional design of this application, the acquisition module 1201 is further configured to acquire a rendering target in a virtual environment corresponding to the virtual background; generate a rendering target depth map of the rendering target in the virtual environment, the rendering target depth map including virtual depth information, the virtual depth information being used to represent the distance from the rendering target to the virtual camera in the virtual environment; convert the virtual depth information in the rendering target depth map into real depth information to obtain the virtual background depth map, the real depth information being used to represent the distance from the rendering target mapped to the real environment to the real camera.

[0183] In an optional design of this application, the update module 1203 is further configured to determine a distance range based on the preset aperture or preset focal length of the real camera, wherein the distance range represents the distance from the reference point corresponding to a pixel with a sharpness greater than a sharpness threshold to the real camera; and adjust the sharpness of each pixel in the target video frame based on the fused depth map and the distance range to generate the depth-of-field effect map of the target video frame.

[0184] In an optional design of this application, the update module 1203 is further configured to adjust the sharpness of the region corresponding to the virtual background within the target video frame based on the focus distance of the real camera and the fusion depth map, and generate the depth-of-field effect map of the target video frame.

[0185] In an optional design of this application, the acquisition module 1201 is further configured to acquire at least two depth-of-field effect images corresponding to the video frame sequence; the update module 1203 is further configured to arrange the at least two depth-of-field effect images in chronological order to obtain the depth-of-field video corresponding to the video frame sequence.

[0186] In summary, in this embodiment, a real-world camera captures the target scene, generating a video frame sequence. The target video frame is then obtained from the video frame sequence, and its depth information is updated to ensure greater accuracy. The addition of depth information from the virtual background results in a more natural and visually appealing image formed by the combination of the virtual background and the real foreground.

[0187] Figure 13 This is a schematic diagram illustrating the structure of a computer device according to an exemplary embodiment. The computer device 1300 includes a Central Processing Unit (CPU) 1301, a system memory 1304 including Random Access Memory (RAM) 1302 and Read-Only Memory (ROM) 1303, and a system bus 1305 connecting the system memory 1304 and the CPU 1301. The computer device 1300 also includes a basic input / output system (I / O system) 1306 that facilitates information transfer between various devices within the computer device, and a mass storage device 1307 for storing the operating system 1313, application programs 1314, and other program modules 1315.

[0188] The basic input / output system 1306 includes a display 1308 for displaying information and an input device 1309 for user input, such as a mouse or keyboard. Both the display 1308 and the input device 1309 are connected to the central processing unit 1301 via an input / output controller 1310 connected to the system bus 1305. The basic input / output system 1306 may also include the input / output controller 1310 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1310 also provides output to a display screen, printer, or other types of output devices.

[0189] Mass storage device 1307 is connected to central processing unit 1301 via a mass storage controller (not shown) connected to system bus 1305. Mass storage device 1307 and its associated computer device readable media provide non-volatile storage for computer device 1300. That is, mass storage device 1307 may include computer device readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.

[0190] Without loss of generality, computer device readable media can include computer device storage media and communication media. Computer device storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technique for storing information such as computer device readable instructions, data structures, program modules, or other data. Computer device storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, digital video disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer device storage media are not limited to the above-mentioned types. The system memory 1304 and mass storage device 1307 described above can be collectively referred to as memory.

[0191] According to various embodiments of this disclosure, computer device 1300 can also be connected to and operated on a remote computer device on a network, such as the Internet. That is, computer device 1300 can be connected to network 1312 via network interface unit 1311 connected to system bus 1305, or network interface unit 1311 can be used to connect to other types of networks or remote computer device systems (not shown).

[0192] The memory also includes one or more programs, which are stored in the memory. The central processing unit 1301 executes the one or more programs to implement all or part of the steps of the above-described virtual reality-based video generation method.

[0193] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the virtual reality-based video generation method provided in the above-described method embodiments.

[0194] This application also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the virtual reality-based video generation method provided in the above-described method embodiments.

[0195] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the virtual reality-based video generation method provided in the above embodiments.

[0196] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0197] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0198] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A video generation method based on virtual reality, characterized in that, The method includes: The target video frame is obtained from a video frame sequence, which is obtained by capturing the target scene with a real camera. The target scene includes a real foreground and a virtual background, and the virtual background is displayed on a physical screen in the real environment. The real foreground depth map and virtual background depth map of the target video frame are obtained. The real foreground depth map includes the depth information from the real foreground to the real camera and the foreground region corresponding to the real foreground. The virtual background depth map includes the depth information from the virtual background to the real camera after being mapped to the real environment. In the foreground region of the real foreground depth map, determine the third pixel point belonging to the target object; In response to the depth value update command, update the depth value of the third pixel; The real foreground depth map and the virtual background depth map are fused to obtain a fused depth map, which includes the depth information of each reference point in the target scene from the real environment to the real camera; Adjust the display parameters of the target video frame according to the fused depth map to generate a depth-of-field effect map of the target video frame; Based on the depth-of-field effect map of the target video frame, a target video with a depth-of-field effect is generated.

2. The method according to claim 1, characterized in that, The real-world foreground depth map includes a background region corresponding to the virtual background; The process of fusing the real foreground depth map and the virtual background depth map to obtain a fused depth map includes: Based on the second depth information of the second pixel point belonging to the background region in the virtual background depth map, the first depth information of the first pixel point belonging to the background region in the real foreground depth map is updated to obtain the fused depth map.

3. The method according to claim 2, characterized in that, The step of updating the first depth information of the first pixel in the real foreground depth map belonging to the background region based on the second depth information of the second pixel in the virtual background depth map belonging to the background region, to obtain the fused depth map, includes: In the virtual background depth map, the j-th second pixel point corresponding to the i-th first pixel point belonging to the background region in the real foreground depth map is determined, where i and j are positive integers; The first depth information of the i-th first pixel is replaced in the real foreground depth map using the second depth information of the j-th second pixel. Update i to i+1, and repeat the above two steps until the first depth information of each first pixel point belonging to the background region in the real foreground depth map is traversed to obtain the fused depth map.

4. The method according to claim 3, characterized in that, The step of determining the j-th second pixel point corresponding to the i-th first pixel point belonging to the background region in the real foreground depth map in the virtual background depth map includes: Based on the intrinsic and extrinsic parameters of the real camera, determine the screen coordinates of the i-th first pixel in the background region on the physical screen; In the virtual environment, the coordinates of the virtual point corresponding to the i-th first pixel are determined based on the screen coordinates; Based on the intrinsic and extrinsic parameters of the virtual camera, the coordinates of the virtual point are mapped onto the virtual background depth map to obtain the j-th second pixel point. The virtual camera is used to capture the rendering target corresponding to the virtual background in the virtual environment.

5. The method according to claim 1, characterized in that, The step of updating the depth value of the third pixel in response to the depth value update instruction includes: According to the depth value setting instruction, the depth value of the third pixel is set to the first preset depth value; Alternatively, according to the depth value increment instruction, a second preset depth value is added to the depth value of the third pixel; Alternatively, according to the depth value reduction instruction, the depth value of the third pixel is reduced by a third preset depth value.

6. The method according to any one of claims 1 to 5, characterized in that, The acquisition of the real-world foreground depth map of the target scene includes: Based on the intrinsic and extrinsic parameters of the depth camera and the intrinsic and extrinsic parameters of the real camera, spatial offset information between the depth camera and the real camera is generated; Obtain the depth map captured by the depth camera; Based on the spatial offset information, the depth information of the depth map is mapped onto the target video frame to obtain the real foreground depth map.

7. The method according to any one of claims 1 to 5, characterized in that, The real-world camera is used to capture the target scene from a first angle; The acquisition of the real-world foreground depth map of the target scene includes: Based on the intrinsic and extrinsic parameters of the reference camera, a first mapping relationship is obtained. The first mapping relationship is used to represent the mapping relationship between the camera coordinate system and the real coordinate system of the reference camera. The reference camera is used to capture the target scene from a second angle, which is different from the first angle. Based on the intrinsic and extrinsic parameters of the real camera, a second mapping relationship is obtained. The second mapping relationship is used to represent the mapping relationship between the camera coordinate system of the real camera and the real coordinate system. The reference image captured by the reference camera is reconstructed according to the first mapping relationship to obtain a reconstructed reference image; the target video frame captured by the real camera is reconstructed according to the second mapping relationship to obtain a reconstructed target scene image. Based on the parallax between the reconstructed reference image and the reconstructed target scene image, the depth information of each pixel within the target video frame is determined to obtain the real foreground depth map.

8. The method according to any one of claims 1 to 5, characterized in that, The step of obtaining the virtual background depth map of the target scene includes: Obtain the rendering target in the virtual environment corresponding to the virtual background; Generate a rendering target depth map of the rendering target in the virtual environment, the rendering target depth map including virtual depth information, the virtual depth information being used to represent the distance from the rendering target to the virtual camera in the virtual environment; The virtual depth information in the rendered target depth map is converted into real depth information to obtain the virtual background depth map. The real depth information is used to represent the distance from the rendered target mapped to the real environment to the real camera.

9. The method according to any one of claims 1 to 5, characterized in that, The step of adjusting the display parameters of the target video frame based on the fused depth map to generate a depth-of-field effect map of the target video frame includes: Based on the preset aperture or preset focal length of the real camera, a distance range is determined. The distance range is used to represent the distance from the reference point corresponding to the pixel with a sharpness greater than the sharpness threshold to the real camera. Based on the fused depth map and the distance range, the sharpness of each pixel in the target video frame is adjusted to generate the depth-of-field effect map of the target video frame.

10. The method according to any one of claims 1 to 5, characterized in that, The step of adjusting the display parameters of the target video frame based on the fused depth map to generate a depth-of-field effect map of the target video frame includes: Based on the focus distance of the real camera and the fusion depth map, the sharpness of the area corresponding to the virtual background within the target video frame is adjusted to generate the depth-of-field effect map of the target video frame.

11. The method according to any one of claims 1 to 5, characterized in that, The process of generating a target video with a depth-of-field effect based on the depth-of-field effect map of the target video frame includes: The depth-of-field effect diagram of the target video frame is played at a preset frame rate to obtain the target video with depth-of-field effect.

12. A video generation device based on virtual reality, characterized in that, The device includes: The acquisition module is used to acquire target video frames from a video frame sequence, wherein the video frame sequence is obtained by capturing a target scene using a real camera, and the target scene includes a real foreground and a virtual background, wherein the virtual background is displayed on a physical screen in a real environment; The acquisition module is further configured to acquire the real foreground depth map and the virtual background depth map of the target video frame. The real foreground depth map includes the depth information from the real foreground to the real camera and the foreground region corresponding to the real foreground. The virtual background depth map includes the depth information from the virtual background to the real camera after being mapped to the real environment. A fusion module is configured to determine a third pixel belonging to the target object in the foreground region of the real foreground depth map; and update the depth value of the third pixel in response to a depth value update instruction. The fusion module is used to fuse the real foreground depth map and the virtual background depth map to obtain a fused depth map. The fused depth map includes the depth information of each reference point in the target scene from the real environment to the real camera. The update module is used to adjust the display parameters of the target video frame according to the fused depth map and generate a depth-of-field effect map of the target video frame; The update module is also used to generate a target video with a depth effect based on the depth effect map of the target video frame.

13. The apparatus according to claim 12, characterized in that, The real foreground depth map includes a background region corresponding to the virtual background; the fusion module is used to update the first depth information of the first pixel in the real foreground depth map belonging to the background region based on the second depth information of the second pixel in the virtual background depth map belonging to the background region, so as to obtain the fused depth map.

14. The apparatus according to claim 13, characterized in that, The acquisition module is used to determine, in the virtual background depth map, the j-th second pixel point corresponding to the i-th first pixel point belonging to the background region in the real foreground depth map, where i and j are positive integers; The first depth information of the i-th first pixel is replaced in the real foreground depth map using the second depth information of the j-th second pixel. Update i to i+1, and repeat the above two steps until the first depth information of each first pixel point belonging to the background region in the real foreground depth map is traversed to obtain the fused depth map.

15. The apparatus according to claim 14, characterized in that, The acquisition module is used to determine the screen coordinates of the i-th first pixel in the background area on the physical screen based on the internal and external parameters of the real camera. In the virtual environment, the coordinates of the virtual point corresponding to the i-th first pixel are determined based on the screen coordinates; Based on the intrinsic and extrinsic parameters of the virtual camera, the coordinates of the virtual point are mapped onto the virtual background depth map to obtain the j-th second pixel point. The virtual camera is used to capture the rendering target corresponding to the virtual background in the virtual environment.

16. The apparatus according to claim 12, characterized in that, The fusion module is configured to set the depth value of the third pixel to a first preset depth value according to a depth value setting instruction; or, add a second preset depth value to the depth value of the third pixel according to a depth value increase instruction; or, decrease the depth value of the third pixel by a third preset depth value according to a depth value decrease instruction.

17. The apparatus according to any one of claims 12 to 16, characterized in that, The acquisition module is used to generate spatial offset information between the depth camera and the real camera based on the intrinsic and extrinsic parameters of the depth camera and the intrinsic and extrinsic parameters of the real camera; and to acquire the depth map collected by the depth camera. Based on the spatial offset information, the depth information of the depth map is mapped onto the target video frame to obtain the real foreground depth map.

18. The apparatus according to any one of claims 12 to 16, characterized in that, The real-world camera is used to capture the target scene from a first angle; The acquisition module is used to acquire a first mapping relationship based on the intrinsic and extrinsic parameters of a reference camera. The first mapping relationship is used to represent the mapping relationship between the camera coordinate system and the real coordinate system of the reference camera. The reference camera is used to capture the target scene from a second angle, which is different from the first angle. Based on the intrinsic and extrinsic parameters of the real camera, a second mapping relationship is obtained. The second mapping relationship is used to represent the mapping relationship between the camera coordinate system of the real camera and the real coordinate system. The reference image captured by the reference camera is reconstructed according to the first mapping relationship to obtain a reconstructed reference image; the target video frame captured by the real camera is reconstructed according to the second mapping relationship to obtain a reconstructed target scene image. Based on the parallax between the reconstructed reference image and the reconstructed target scene image, the depth information of each pixel within the target video frame is determined to obtain the real foreground depth map.

19. The apparatus according to any one of claims 12 to 16, characterized in that, The acquisition module is used to acquire the rendering target in the virtual environment corresponding to the virtual background; Generate a rendering target depth map of the rendering target in the virtual environment, the rendering target depth map including virtual depth information, the virtual depth information being used to represent the distance from the rendering target to the virtual camera in the virtual environment; The virtual depth information in the rendered target depth map is converted into real depth information to obtain the virtual background depth map. The real depth information is used to represent the distance from the rendered target mapped to the real environment to the real camera.

20. The apparatus according to any one of claims 12 to 16, characterized in that, The update module is used to determine a distance range based on the preset aperture or preset focal length of the real camera. The distance range represents the distance from the reference point corresponding to a pixel with a sharpness greater than a sharpness threshold to the real camera. Based on the fused depth map and the distance range, the sharpness of each pixel in the target video frame is adjusted to generate the depth-of-field effect map of the target video frame.

21. The apparatus according to any one of claims 12 to 16, characterized in that, The update module is used to adjust the sharpness of the area corresponding to the virtual background within the target video frame based on the focus distance of the real camera and the fusion depth map, and generate the depth-of-field effect map of the target video frame.

22. The apparatus according to any one of claims 12 to 16, characterized in that, The update module is used to play the depth-of-field effect map of the target video frame at a preset frame rate to obtain the target video with depth-of-field effect.

23. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the virtual reality-based video generation method as described in any one of claims 1 to 11.

24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the virtual reality-based video generation method as described in any one of claims 1 to 11.

25. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the video generation method based on virtual reality as described in any one of claims 1 to 11.