Virtual camera-based image generation method, apparatus, device, medium and program product

By deploying multiple physical cameras around the video conference screen, determining the spatial parameters of the virtual camera and the physical camera, acquiring scene images and rendering them, the problem of the inflexible adjustment of cameras in existing video conferencing systems is solved, enabling personalized shooting by the virtual camera and improving the user experience.

CN119520724BActive Publication Date: 2026-01-13CHINA TELECOM CLOUD TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411697977.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2026-01-13
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Existing video conferencing systems cannot flexibly adjust the camera's shooting angle according to user needs, making it difficult to meet personalized requirements and resulting in a poor user experience.

Method used

By deploying multiple physical cameras around the video session screen, determining the target spatial parameters between the virtual camera and each physical camera, acquiring multiple scene images, and using image rendering algorithms to simulate the shooting of the virtual camera, a rendered image is generated.

Benefits of technology

It enables flexible virtual camera settings, allowing users to place the camera anywhere on the video session screen according to their needs, thus meeting personalized requirements and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119520724B_ABST
    Figure CN119520724B_ABST
Patent Text Reader

Abstract

The application relates to a virtual camera-based image generation method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: in response to a triggering operation on a target position on a video session screen, taking the target position as a virtual arrangement position of a virtual camera on the video session screen, and arranging a plurality of physical cameras around the video session screen; determining target space parameters between the virtual camera and each physical camera in a target coordinate system based on a physical arrangement position of each physical camera and the virtual arrangement position; obtaining a plurality of scene images, each scene image being obtained by photographing a target object by a corresponding physical camera; performing image rendering based on the target space parameters corresponding to each physical camera, the plurality of scene images and virtual camera parameters of the virtual camera, obtaining a rendered image, and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image generation method, apparatus, computer device, computer-readable storage medium, and computer program product based on a virtual camera. Background Technology

[0002] With the development of internet technology, video conferencing has become an indispensable part of people's daily communication and remote collaboration. Currently, most video conferencing systems are equipped with cameras, and can be roughly divided into two categories based on the type of camera: one is a conferencing system with only a single camera, which transmits the image captured by that camera to the other party in real time; the other is a conferencing system with multiple cameras, which transmits the images captured by different cameras to the other party in real time as needed.

[0003] However, both of these types of conferencing systems transmit real images captured by physical cameras, which can only shoot from a fixed angle and cannot be flexibly adapted to different needs, making it difficult to meet users' personalized requirements. Summary of the Invention

[0004] Therefore, it is necessary to provide an image generation method, apparatus, computer device, computer-readable storage medium, and computer program product based on a virtual camera that can flexibly shoot based on a virtual camera, meet users' personalized needs, and improve user experience, in order to address the above-mentioned technical problems.

[0005] In a first aspect, this application provides an image generation method based on a virtual camera, including:

[0006] In response to a trigger operation on a target location on the video session screen, the target location is used as the virtual deployment location of the virtual camera on the video session screen, and multiple physical cameras are deployed around the video session screen.

[0007] Based on the physical deployment position of each physical camera and the virtual deployment position, the target spatial parameters between the virtual camera and each physical camera in the target coordinate system are determined.

[0008] Multiple scene images are acquired, each scene image being captured by a corresponding physical camera of the target object;

[0009] Based on the target space parameters corresponding to each physical camera, multiple scene images, and the virtual camera parameters of the virtual camera, image rendering is performed to obtain a rendered image. The rendered image is obtained by simulating the virtual camera taking a picture of the target object at the virtual camera position.

[0010] Secondly, this application also provides an image generation apparatus based on a virtual camera, comprising:

[0011] The location determination module is used to respond to a trigger operation on a target location on the video session screen and use the target location as the virtual deployment location of the virtual camera on the video session screen, wherein multiple physical cameras are deployed around the video session screen.

[0012] The parameter determination module is used to determine the target spatial parameters between the virtual camera and each physical camera in the target coordinate system based on the physical deployment position of each physical camera and the virtual deployment position.

[0013] The image acquisition module is used to acquire multiple scene images, each of which is obtained by capturing the target object with a corresponding physical camera;

[0014] The image rendering module is used to render images based on the target space parameters corresponding to each physical camera, multiple scene images, and the virtual camera parameters of the virtual camera to obtain a rendered image. The rendered image is obtained by simulating the virtual camera taking a picture of the target object at the virtual camera position.

[0015] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0016] In response to a trigger operation on a target location on the video session screen, the target location is used as the virtual deployment location of the virtual camera on the video session screen, and multiple physical cameras are deployed around the video session screen.

[0017] Based on the physical deployment position of each physical camera and the virtual deployment position, the target spatial parameters between the virtual camera and each physical camera in the target coordinate system are determined.

[0018] Multiple scene images are acquired, each scene image being captured by a corresponding physical camera of the target object;

[0019] Based on the target space parameters corresponding to each physical camera, multiple scene images, and the virtual camera parameters of the virtual camera, image rendering is performed to obtain a rendered image. The rendered image is obtained by simulating the virtual camera taking a picture of the target object at the virtual camera position.

[0020] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0021] In response to a trigger operation on a target location on the video session screen, the target location is used as the virtual deployment location of the virtual camera on the video session screen, and multiple physical cameras are deployed around the video session screen.

[0022] Based on the physical deployment position of each physical camera and the virtual deployment position, the target spatial parameters between the virtual camera and each physical camera in the target coordinate system are determined.

[0023] Multiple scene images are acquired, each scene image being captured by a corresponding physical camera of the target object;

[0024] Based on the target space parameters corresponding to each physical camera, multiple scene images, and the virtual camera parameters of the virtual camera, image rendering is performed to obtain a rendered image. The rendered image is obtained by simulating the virtual camera taking a picture of the target object at the virtual camera position.

[0025] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0026] In response to a trigger operation on a target location on the video session screen, the target location is used as the virtual deployment location of the virtual camera on the video session screen, and multiple physical cameras are deployed around the video session screen.

[0027] Based on the physical deployment position of each physical camera and the virtual deployment position, the target spatial parameters between the virtual camera and each physical camera in the target coordinate system are determined.

[0028] Multiple scene images are acquired, each scene image being captured by a corresponding physical camera of the target object;

[0029] Based on the target space parameters corresponding to each physical camera, multiple scene images, and the virtual camera parameters of the virtual camera, image rendering is performed to obtain a rendered image. The rendered image is obtained by simulating the virtual camera taking a picture of the target object at the virtual camera position.

[0030] The aforementioned image generation method, apparatus, computer device, computer-readable storage medium, and computer program product based on a virtual camera, responds to a trigger operation on a target location on a video session screen, using the target location as the virtual placement position of the virtual camera on the video session screen, with multiple physical cameras deployed around the perimeter of the video session screen. That is, as the user selects a location on the video session screen, that location is automatically set as the placement position of the virtual camera. Thus, the user can flexibly place the virtual camera at any location on the video session screen according to actual needs, improving the flexibility of the virtual camera setup process. Based on the physical and virtual placement positions of each physical camera, target spatial parameters between the virtual camera and each physical camera are determined in the target coordinate system to reflect the spatial correlation between the virtual and physical cameras; multiple scene images are acquired, each scene image being captured by the corresponding physical camera of the target object; based on the target spatial parameters corresponding to each physical camera, the multiple scene images, and the virtual camera parameters of the virtual camera, image rendering is performed to obtain a rendered image. In other words, after photographing the target object using a physical camera, the system can capture detailed angular information about the object from the angles of the physical cameras positioned at different locations. Then, by combining the spatial relationship between the physical and virtual cameras and the parameters of the virtual camera, it simulates a virtual camera photographing the target object from a virtual camera position. The entire process allows for flexible shooting based on the user's actual needs, meeting personalized requirements and improving the user experience. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is an application environment diagram of an image generation method based on a virtual camera in one embodiment;

[0033] Figure 2 This is an application environment diagram of the image generation method based on a virtual camera in another embodiment;

[0034] Figure 3 This is a flowchart illustrating an image generation method based on a virtual camera in one embodiment;

[0035] Figure 4 This is a schematic diagram of a video session scenario in one embodiment;

[0036] Figure 5This is a schematic diagram of the target space parameter determination steps in one embodiment;

[0037] Figure 6 This is a schematic diagram of a video session screen in a Cartesian coordinate system in one embodiment;

[0038] Figure 7 This is a schematic diagram of a video session screen in polar coordinates in one embodiment;

[0039] Figure 8 This is a schematic diagram of the virtual camera focusing in one embodiment;

[0040] Figure 9 This is a schematic diagram of an image preview pop-up window in one embodiment;

[0041] Figure 10 This is a schematic diagram of an image preview pop-up window in another embodiment;

[0042] Figure 11 This is a schematic diagram illustrating the generation of a rendered image in one embodiment;

[0043] Figure 12 This is a structural block diagram of an image generation device based on a virtual camera in one embodiment;

[0044] Figure 13 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0046] The image generation method based on a virtual camera provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with multiple camera modules 104 deployed around its outer edge. For example, at least two camera modules 104 are arranged around the perimeter of terminal 102. Figure 1 A camera module 104 is installed on each of the four edges of the terminal 102. Each camera module 104 contains at least one physical camera, such as... Figure 2 In another application scenario shown, Figure 2The numbers 1 to 4 represent the four camera modules, and the numbers 5 to 8 represent the physical cameras. A physical camera is a camera with an independent physical camera device. A virtual camera (VCAM) is a device or software that uses computer vision and graphics technology to simulate the working principle of a real camera. That is, it does not have a physical camera device, but it can simulate the shooting function of a physical camera to perform simulated shooting.

[0047] In some embodiments, in response to a trigger operation on a target location on a video session screen in the terminal 102, the target location is used as the virtual deployment location of a virtual camera on the video session screen. Multiple camera modules 104 are deployed around the video session screen, and each camera module 104 contains at least one physical camera. Based on the physical deployment location and virtual deployment location of each physical camera, target spatial parameters between the virtual camera and each physical camera in the target coordinate system are determined. The terminal 102 acquires the corresponding scene image sent by each camera module 104. Each scene image is obtained by the corresponding physical camera capturing the target object. Based on the target spatial parameters corresponding to each physical camera, the multiple scene images, and the virtual camera parameters of the virtual camera, the terminal 102 performs image rendering to obtain a rendered image. The rendered image is obtained by simulating the virtual camera capturing the target object at the virtual camera location.

[0048] Terminal 102 is a terminal with real-time video call functionality. For example, terminal 102 can be considered as the initiating terminal for virtual camera shooting. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, etc. IoT devices can be smart TVs, smart in-vehicle devices, projection devices, etc. Camera module 104 is used to send scene images captured by the physical camera in camera module 104 to terminal 102. Camera module 104 includes a physical camera, which can be an RGB (Red, Green, Blue; color camera) camera or an RGB-D (Red, Green, Blue-Depth, color depth) camera. An RGB camera, also known as a color camera, can simultaneously capture the three main color channels of an image—red, green, and blue—to generate a complete color image. An RGB-D camera is a camera that combines RGB color images and depth images using three-dimensional visual sensing technology. It acquires the three-dimensional information of a scene by collecting color images and depth information from the scene.

[0049] In one exemplary embodiment, such as Figure 3 As shown, an image generation method based on a virtual camera is provided, which can be applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps S302 to S308. Wherein:

[0050] In step S302, in response to the triggering operation of the target position on the video session screen, the target position is used as the virtual deployment position of the virtual camera on the video session screen, and multiple physical cameras are deployed around the video session screen.

[0051] The video session screen is the screen where a video session is in progress, including but not limited to conference sessions, teaching sessions, training sessions, live broadcasts, and social sessions. The video session can be based on naked-eye 3D (3D) or non-3D. For example, in a naked-eye 3D session, the video session screen is a screen equipped with a naked-eye 3D display. A naked-eye 3D display is a visual experience technology that allows viewers to directly view stereoscopic 3D images with their naked eyes without wearing any glasses or helmets. Triggering operations include but are not limited to mouse clicks, mouse swipes, and touchscreen operations; correspondingly, the target position can be determined based on the mouse or based on the user's touch on the video session screen. Multiple physical cameras are deployed around the video session screen. For example, based on the shape of the video session screen, at least one physical camera is deployed in different directions to obtain detailed information about the target object from different perspectives. For example, if the video session screen is rectangular, at least one physical camera is deployed on each border of the rectangle. Figure 2 The deployment.

[0052] Virtual deployment location is used to characterize the deployment location of virtual cameras on the video session screen.

[0053] Optionally, during a video session, if the conditions for starting the virtual camera are met, the terminal (which can be regarded as the initiating terminal that initiates the virtual camera shooting) responds to the triggering operation on the target position on the video session screen and uses the target position as the virtual deployment position of the virtual camera on the video session screen.

[0054] For example, during a video session, a virtual camera shooting control is displayed on the video session screen of the terminal, and in response to a trigger operation on the virtual camera shooting control, it is determined that the virtual camera start-up conditions are met.

[0055] For example, during a video session, the terminal receives a voice command from any member of the video session to start the virtual camera, and determines that the conditions for starting the virtual camera have been met.

[0056] For example, such as Figure 4The diagram illustrates a video conversation scenario in one embodiment. Terminal 1 of video call participant A and terminal 2 of video call participant B conduct a video conversation via the internet. The screens of terminals 1 and 2 simultaneously display the video conversation images of A and B. Taking video call participant A as the main focus, after clicking the virtual camera activation control on terminal 1, and then clicking a target location on the video conversation screen of terminal 1, terminal 1 determines that the target location is the virtual deployment location of the virtual camera, used for image simulation of video call participant A.

[0057] Step S304: Based on the physical and virtual deployment positions of each physical camera, determine the target spatial parameters between the virtual camera and each physical camera in the target coordinate system.

[0058] For each physical camera, the target spatial parameters of the physical camera and the virtual camera are used to characterize the spatial positional relationship between the physical camera and the virtual camera in the target coordinate system. The target control parameters include, but are not limited to, distance and angle. The target coordinate system can be at least one of a rectangular coordinate system and a polar coordinate system.

[0059] For example, the terminal acquires the physical deployment position of each physical camera on the video session screen and selects a target coordinate system from multiple preset coordinate systems. For each physical camera, the terminal determines the target spatial parameters corresponding to that physical camera based on its physical deployment position and virtual deployment position. These target spatial parameters refer to the target spatial parameters between the virtual camera and each physical camera in the target coordinate system. For example, for each physical camera, the terminal determines at least one of the distance and angle between the physical camera and the virtual camera in the target coordinate system based on its physical deployment position and virtual deployment position.

[0060] Step S306: Acquire multiple scene images, each scene image being captured by a corresponding physical camera of the target object.

[0061] The target object can be a target in the scene where the terminal is located. In some embodiments, the target object can be the party initiating the video call using the virtual camera. In some embodiments, the target object can also be the environment where the terminal is located. In other embodiments, the target object can also be a party not initiating the video call using the virtual camera; the specific method is not limited.

[0062] Step S308: Based on the target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera, image rendering is performed to obtain a rendered image. The rendered image is obtained by simulating the virtual camera taking pictures of the target object at the virtual camera position.

[0063] The virtual camera parameters can be default and do not need to be set manually, or they can be parameters entered manually as needed. The virtual camera can have parameters such as pitch angle and rotation angle.

[0064] For example, the terminal generates a rendered image based on the target space parameters corresponding to each physical camera, multiple scene images, and the virtual camera parameters of the virtual camera, using an image rendering algorithm. The rendered image is obtained by simulating a virtual camera taking a picture of the target object at the virtual camera's location.

[0065] Among these, the image rendering algorithm can be NeRF (Neural Radiance Fields, an implicit field representation method), which can learn a continuous volume representation of a scene through a deep neural network and use it for high-quality novel perspective synthesis. Alternatively, the image rendering algorithm can be Spacetime Gaussian Feature Splatting: this is a novel perspective synthesis method for dynamic scenes, designed to achieve high resolution, realistic rendering effects, real-time rendering, and compact storage.

[0066] In the aforementioned image generation method based on virtual cameras, in response to a trigger operation on a target location on the video session screen, the target location is used as the virtual placement position of the virtual camera on the video session screen, while multiple physical cameras are placed around the perimeter of the video session screen. That is, as the user selects a location on the video session screen, that location is automatically set as the placement position of the virtual camera. Therefore, the user can place the virtual camera at any location on the video session screen according to actual needs, flexibly setting up the virtual camera and improving the flexibility of the virtual camera setup process. Based on the physical and virtual placement positions of each physical camera, target spatial parameters between the virtual camera and each physical camera are determined in the target coordinate system to reflect the spatial correlation between the virtual and physical cameras. Multiple scene images are acquired, each scene image being captured by the corresponding physical camera of the target object. Based on the target spatial parameters corresponding to each physical camera, the multiple scene images, and the virtual camera parameters of the virtual camera, image rendering is performed to obtain a rendered image. In other words, after photographing the target object using a physical camera, the system can capture detailed angular information about the object from the angles of the physical cameras positioned at different locations. Then, by combining the spatial relationship between the physical and virtual cameras and the parameters of the virtual camera, it simulates a virtual camera photographing the target object from a virtual camera position. The entire process allows for flexible shooting based on the user's actual needs, meeting personalized requirements and improving the user experience.

[0067] In some embodiments, such as Figure 5The diagram illustrates the target spatial parameter determination steps in one embodiment. The target coordinate system includes a Cartesian coordinate system and a polar coordinate system. Based on the physical and virtual deployment positions of each physical camera, the target spatial parameters between the virtual camera and each physical camera in the target coordinate system are determined, including:

[0068] Step S502: Based on the virtual deployment location, determine the first virtual coordinates of the virtual camera in the rectangular coordinate system and the second virtual coordinates in the polar coordinate system.

[0069] For example, such as Figure 6 The diagram shown illustrates a video conference screen in a Cartesian coordinate system in one embodiment. The center of the video conference screen is taken as the origin of the Cartesian coordinate system. The x-axis and y-axis are determined based on the top / bottom and left / right orientations of the video conference screen to establish the Cartesian coordinate system. Specifically, the top / bottom center of the video conference screen is the x-axis, the right is the positive x-axis direction, and the left is the negative x-axis direction. The left / right center of the video conference screen is the y-axis, the top is the positive y-axis direction, and the bottom is the negative y-axis direction. The coordinates of the screen center are (0,0). The dimensions of the video conference screen are set as follows: length and width are 2Xc and 2Yc, respectively. If the virtual camera is at point A on the video conference screen, its first virtual coordinates in the Cartesian coordinate system are (Xa, Ya).

[0070] like Figure 7 The diagram shown illustrates a video session screen in polar coordinates in one embodiment. The center of the video session screen is designated as the pole of the polar coordinate system, and a ray Ox is drawn horizontally to the right from the pole as the polar axis to establish the polar coordinate system. Therefore, the second virtual coordinate of the virtual camera in the polar coordinate system is (ra, θa), where ra = ... θa is the angle between the coordinates and the polar coordinates.

[0071] Step S504: For each physical camera, based on the physical deployment location of the physical camera, determine the first physical coordinates of the physical camera in the rectangular coordinate system and the second physical coordinates in the polar coordinate system.

[0072] Reference Figure 2 The deployment of physical cameras in the middle, combined with Figure 6 We know that the first physical coordinates of physical cameras 5-8 in the Cartesian coordinate system are (-Xc, 0), (Xc, 0), (0, Yc), and (0, -Yc), respectively. (Refer to...) Figure 7 In the polar coordinate system, the second physical coordinates of physical cameras 5-8 in the polar coordinate system are (Xc, π), (Xc, 0), (Yc, π / 2), and (Yc, 3π / 2), respectively.

[0073] Step S506: For each physical camera, determine the first spatial parameters between the physical camera and the virtual camera based on the first physical coordinates and the first virtual coordinates, and determine the second spatial parameters between the physical camera and the virtual camera based on the second physical coordinates and the second virtual coordinates.

[0074] For example, for each physical camera, the angle between the physical camera and the virtual camera is calculated based on the first physical coordinates and the first virtual coordinates, serving as the corresponding first spatial parameter. The angle between the physical camera and the virtual camera is determined based on the second physical coordinates and the second virtual coordinates, serving as the corresponding second spatial parameter.

[0075] Continuing with the example above, the first physical coordinates of physical cameras 5-8 in the rectangular coordinate system are (-Xc, 0), (Xc, 0), (0, Yc), and (0, -Yc), respectively, and the first virtual coordinate of the virtual camera is (Xa, Ya). The distances from the virtual camera to the physical camera 5-8 are then denoted as the first spatial parameters, namely d1, d2, d3, and d4.

[0076] The second physical coordinates of the physical cameras 5-8 in the polar coordinate system are (Xc, π), (Xc, 0), (Yc, π / 2), and (Yc, 3π / 2). The second virtual coordinates of the virtual camera are (ra, a); then the angles between the virtual camera and the physical cameras 5-8 are all recorded as the second spatial parameters, denoted as α1, α2, α3, and α4 respectively.

[0077] Step S508: For each physical camera, the first spatial parameter and the second spatial parameter between the physical camera and the virtual camera are used as the target spatial parameter between the physical camera and the virtual camera.

[0078] In this embodiment, by determining the coordinates of the virtual camera and each physical camera in the Cartesian coordinate system and their respective coordinates in the polar coordinate system, the spatial positional relationship between the virtual camera and each physical camera can be accurately and comprehensively obtained. As a result, accurate simulation rendering can be performed based on this and the scene images captured by the physical cameras, ensuring the accuracy of the rendered images generated by the virtual camera.

[0079] In some embodiments, the virtual camera parameters of the virtual camera include at least the virtual camera's pitch angle, rotation angle, and focal point position parameters for focusing on the target object.

[0080] like Figure 8 The diagram shown illustrates the focusing of a virtual camera in one embodiment. Point O is the center of the video session screen, the focal point is the target object, and the distance between point O and the focal point is L.

[0081] For example, the pitch and rotation angles of the virtual camera can be set to 0 by default, for instance. The parameters of the virtual camera can be manually set as needed.

[0082] For example, virtual camera parameters also include field of view, resolution, and frame rate, which can be consistent with those of a physical camera.

[0083] In this embodiment, the virtual camera parameters include at least the virtual camera's pitch angle, rotation angle, and focal point position parameters for focusing on the target object. This accurately reflects the virtual camera's shooting parameters, ensuring the quality of the rendered image.

[0084] In some embodiments, image rendering is performed based on the target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera to obtain a rendered image, including: calling a trained image processing model, inputting the target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera into the image processing model to obtain a rendered image.

[0085] The trained image processing model is a neural network-based model, which can be constructed using the NeRF algorithm or the spatiotemporal Gaussian feature spraying algorithm.

[0086] For example, after the trained image processing model obtains the target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera, it determines the ray direction of each pixel in each scene image and the color and lighting information of each pixel in each scene image based on the scene images and the camera parameters of the physical cameras. The image processing model then fits the ray direction of each pixel in the rendered image based on the target space parameters and virtual camera parameters corresponding to each physical camera. Based on the pose of the physical camera corresponding to each scene image, the image processing model determines the intersection point (i.e., the intersection point of each pixel in the scene image with the scene in the scene image) corresponding to each pixel in each scene image. Based on the ray direction, color, and lighting information of each pixel in each scene image, the ray direction of each pixel in the rendered image, and the intersection point of each pixel in each scene image, the image processing model determines the color and lighting information of each intersection point corresponding to each pixel in the rendered image. Thus, the rendering color of each pixel in the rendered image is determined, and the rendered image is generated based on the rendering colors of each pixel.

[0087] In this embodiment, by calling the trained image processing model, image fitting (rendering) can be performed accurately and efficiently based on the target space parameters corresponding to each physical camera, multiple scene images, and the virtual camera parameters of the virtual camera, so as to obtain an accurate rendered image.

[0088] In some embodiments, the method further includes: displaying an image preview pop-up at a virtual deployment location on the video session screen, the image preview pop-up including a rendered image; the method further includes: in response to a triggering operation on another location on the video session screen, displaying a new image preview pop-up at another location, the new preview pop-up including a new rendered image, the new rendered image being obtained by simulating a virtual camera taking a picture at the other location.

[0089] For example, such as Figure 9 The image shown is a schematic diagram of an image preview pop-up window in one embodiment. Figure 9 The illustration shows that after a virtual camera is set at virtual deployment location A in the video session screen, an image preview pop-up is displayed at the virtual deployment location, showing the rendered image. Alternatively, a corresponding image preview pop-up can be displayed around the virtual deployment location, for example, above the virtual deployment location. Alternatively, the virtual deployment location can be displayed using a camera icon.

[0090] Triggering operations on other locations on the video call screen can be moving operations, dragging operations, etc. For example, in the rendered image preview state, the party initiating the video call using the virtual camera can drag the virtual camera to any location on the video call screen by left-clicking and holding the virtual camera position and the image preview screen. During the dragging process, the virtual camera position and the rendered image in the image preview pop-up are updated in real time, and the virtual camera position is always kept aligned with the image preview pop-up.

[0091] During the drag operation, it is necessary to ensure that the focus point position parameters of the virtual camera remain unchanged. For example, if the virtual camera focuses on a point 1.2 meters away from the center of the screen at the same height as point O, the focus point must remain unchanged during the drag operation. At the same time, it may be necessary to change the tilt and rotation angles of the virtual camera to ensure this.

[0092] For example, such as Figure 10 The image shown is a schematic diagram of an image preview pop-up window in another embodiment. The current virtual camera's virtual deployment position is... Figure 10 Position A in the image displays the front view of the target object. By dragging and dropping, the virtual placement position can be changed from position A to... Figure 10 After position B in the image, refer to the new rendered image corresponding to position B; the target object will be displayed slightly to the left. Alternatively, you can change the position from A to B by dragging. Figure 10 After position C, refer to the new rendered image corresponding to position C, which shows the target object to the right.

[0093] In this embodiment, after generating a rendered image corresponding to the virtual deployment location, the user can directly select other locations to generate new rendered images as new virtual deployment locations, thereby improving the user experience.

[0094] In some embodiments, the new rendered image generation step includes: updating the virtual camera parameters of the virtual camera to obtain updated virtual camera parameters; using other locations as new virtual deployment locations, and determining new spatial parameters between the virtual camera and each physical camera in the target coordinate system based on the new virtual deployment locations and the physical deployment locations of each physical camera; and performing image rendering based on the new spatial parameters corresponding to each physical camera, multiple scene images, and the updated virtual camera parameters to obtain a new rendered image.

[0095] For example, the terminal obtains new pitch angles and new rotation angles for the virtual camera. Based on the focal point position parameters used to focus on the target object, the new pitch angles, and the new rotation angles, updated virtual camera parameters are obtained. Based on the new virtual deployment positions and the physical deployment positions of each physical camera, new spatial parameters between the virtual camera and each physical camera in the target coordinate system are determined. Based on the new spatial parameters corresponding to each physical camera, multiple scene images, and the updated virtual camera parameters, image rendering is performed to obtain a new rendered image.

[0096] For example, after obtaining new spatial parameters, the terminal can also acquire new scene images again, and perform image rendering based on the new spatial parameters corresponding to each physical camera, the new scene images, and the updated virtual camera parameters to obtain new rendered images.

[0097] In this embodiment, after obtaining the new virtual deployment location, it is necessary to update the virtual camera parameters of the virtual camera in a timely manner so as to render a new rendering image that matches it.

[0098] In a specific embodiment, the implementation process refers to the following steps:

[0099] Step 1: In response to the trigger operation on the target location on the video session screen, the target location is used as the virtual deployment location of the virtual camera on the video session screen, and multiple physical cameras are deployed around the video session screen.

[0100] Step 2: Based on the virtual deployment location, determine the first virtual coordinates of the virtual camera in the Cartesian coordinate system and the second virtual coordinates in the polar coordinate system. For each physical camera, based on the physical deployment location of the physical camera, determine the first physical coordinates of the physical camera in the Cartesian coordinate system and the second physical coordinates in the polar coordinate system. For each physical camera, determine the first spatial parameter between the physical camera and the virtual camera based on the first physical coordinates and the first virtual coordinates, and determine the second spatial parameter between the physical camera and the virtual camera based on the second physical coordinates and the second virtual coordinates. For each physical camera, use the first spatial parameter and the second spatial parameter between the physical camera and the virtual camera as the target spatial parameter between the physical camera and the virtual camera.

[0101] Step 3: Acquire multiple scene images, each captured by a corresponding physical camera of the target object. Determine the virtual camera parameters, which include at least the virtual camera's pitch angle, rotation angle, and focus point position parameters for focusing on the target object.

[0102] Step 4: Invoke the trained image processing model, inputting the target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera into the image processing model to obtain the rendered image. An image preview pop-up window, including the rendered image, is displayed at the virtual deployment location on the video session screen.

[0103] Specifically, such as Figure 11 The diagram illustrates the generation of a rendered image in one embodiment. When the virtual deployment position is set to point A, four first spatial parameters d1, d2, d3, and d4 in Cartesian coordinates and second spatial parameters α1, α2, α3, and α4 in polar coordinates are obtained, following the method mentioned earlier. The scene image (four real-time video images), d1, d2, d3, d4, α1, α2, α3, α4, and virtual camera parameters (set virtual camera parameters, default values) are input into the image processing model (image processing algorithm model, such as NeRF / spatiotemporal Gaussian feature spraying), and the fitted rendered image (virtual camera image), i.e., the image captured by the virtual camera at position A, is output.

[0104] Step 5: In response to trigger operations at other locations on the video session screen, generate new rendered images. This involves updating the virtual camera parameters to obtain updated virtual camera parameters; using these other locations as new virtual deployment locations, and determining new spatial parameters between the virtual camera and each physical camera in the target coordinate system based on the new virtual deployment locations and the physical deployment locations of each physical camera; and rendering images based on the new spatial parameters corresponding to each physical camera, multiple scene images, and the updated virtual camera parameters to obtain new rendered images. Display new image preview pop-ups at these other locations, including the new rendered images.

[0105] In this embodiment, in response to a trigger operation on a target location on the video session screen, the target location is used as the virtual placement location of the virtual camera on the video session screen, and multiple physical cameras are placed around the perimeter of the video session screen. That is, as the user selects a location on the video session screen, that location is automatically set as the placement location of the virtual camera. Therefore, the user can place the virtual camera at any location on the video session screen according to actual needs, flexibly setting up the virtual camera and improving the flexibility of the virtual camera setup process. Based on the physical and virtual placement locations of each physical camera, target spatial parameters between the virtual camera and each physical camera are determined in the target coordinate system to reflect the spatial correlation between the virtual and physical cameras. Multiple scene images are acquired, each scene image being captured by the corresponding physical camera of the target object. Based on the target spatial parameters corresponding to each physical camera, the multiple scene images, and the virtual camera parameters of the virtual camera, image rendering is performed to obtain a rendered image. In other words, after photographing the target object using a physical camera, the system can capture detailed angular information about the object from the angles of the physical cameras positioned at different locations. Then, by combining the spatial relationship between the physical and virtual cameras and the parameters of the virtual camera, it simulates a virtual camera photographing the target object from a virtual camera position. The entire process allows for flexible shooting based on the user's actual needs, meeting personalized requirements and improving the user experience.

[0106] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0107] Based on the same inventive concept, this application also provides a virtual camera-based image generation apparatus for implementing the above-described virtual camera-based image generation method. The solution provided by this apparatus is similar to the implementation described in the above-described method; therefore, the specific limitations in one or more embodiments of the virtual camera-based image generation apparatus provided below can be found in the limitations of the virtual camera-based image generation method described above, and will not be repeated here.

[0108] In one exemplary embodiment, such as Figure 12 As shown, an image generation device 1200 based on a virtual camera is provided, including: a position determination module 1202, a parameter determination module 1204, an image acquisition module 1206, and an image rendering module 1208, wherein:

[0109] The location determination module 1202 is used to respond to the trigger operation of the target location on the video session screen, and use the target location as the virtual deployment position of the virtual camera on the video session screen, with multiple physical cameras deployed around the video session screen.

[0110] The parameter determination module 1204 is used to determine the target space parameters between the virtual camera and each physical camera in the target coordinate system based on the physical deployment position and virtual deployment position of each physical camera.

[0111] The image acquisition module 1206 is used to acquire multiple scene images, each scene image being obtained by a corresponding physical camera capturing the target object;

[0112] The image rendering module 1208 is used to render images based on the target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera to obtain a rendered image. The rendered image is obtained by simulating the virtual camera taking pictures of the target object at the virtual camera position.

[0113] In some embodiments, the target coordinate system includes a rectangular coordinate system and a polar coordinate system. The parameter determination module 1204 is used to determine, based on the virtual deployment position, a first virtual coordinate in the rectangular coordinate system and a second virtual coordinate in the polar coordinate system for the virtual camera; for each physical camera, based on the physical deployment position of the physical camera, a first physical coordinate in the rectangular coordinate system and a second physical coordinate in the polar coordinate system for the physical camera; for each physical camera, a first spatial parameter between the physical camera and the virtual camera is determined according to the first physical coordinate and the first virtual coordinate, and a second spatial parameter between the physical camera and the virtual camera is determined according to the second physical coordinate and the second virtual coordinate; for each physical camera, the first spatial parameter and the second spatial parameter between the physical camera and the virtual camera are used as the target spatial parameter between the physical camera and the virtual camera.

[0114] In some embodiments, the virtual camera parameters of the virtual camera include at least the virtual camera's pitch angle, rotation angle, and focal point position parameters for focusing on the target object.

[0115] In some embodiments, the image rendering module 1208 is used to call a trained image processing model, input the target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera into the image processing model to obtain a rendered image.

[0116] In some embodiments, the apparatus further includes a display module, configured to display an image preview pop-up at a virtual location on the video session screen, the image preview pop-up including a rendered image; and a display module configured to display a new image preview pop-up at another location in response to a trigger operation on another location on the video session screen, the new preview pop-up including a new rendered image, the new rendered image being obtained by simulating a virtual camera taking a picture at another location.

[0117] In some embodiments, the apparatus further includes an update module, which is used to update the virtual camera parameters of the virtual camera to obtain updated virtual camera parameters; a parameter determination module 1204, which is used to take other locations as new virtual deployment locations and determine new spatial parameters between the virtual camera and each physical camera in the target coordinate system based on the new virtual deployment locations and the physical deployment locations of each physical camera; and an image rendering module, which is used to perform image rendering based on the new spatial parameters corresponding to each physical camera, multiple scene images, and the updated virtual camera parameters to obtain a new rendered image.

[0118] The modules in the aforementioned virtual camera image generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0119] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 13 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for generating images from a virtual camera.

[0120] Those skilled in the art will understand that Figure 13 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0121] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0122] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0123] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0125] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0126] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0127] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for generating an image based on a virtual camera, characterized by, The method is applied to a terminal initiating virtual camera shooting, and the method comprises the following steps: In response to a triggering operation on a target position on a video session screen, the target position is taken as a virtual arrangement position of a virtual camera on the video session screen, and a plurality of physical cameras are arranged around the video session screen; Based on a physical arrangement position of each physical camera and the virtual arrangement position, target space parameters between the virtual camera and each physical camera in a target coordinate system are determined; A plurality of scene images are obtained, each scene image being obtained by a corresponding physical camera shooting a target object, and the target object being a video call party initiating virtual camera shooting; Based on the target space parameters corresponding to each physical camera, the plurality of scene images, and virtual camera parameters of the virtual camera, a rendering image is generated by an image rendering algorithm for new view synthesis, the rendering image being obtained by simulating the virtual camera shooting the target object at the virtual camera position, and the image rendering algorithm for new view synthesis being NeRF or spatiotemporal high-dimensional feature spraying; An image preview pop-up window is displayed at the virtual arrangement position on the video session screen, and the image preview pop-up window comprises the rendering image; In response to a triggering operation on another position on the video session screen, a new image preview pop-up window is displayed at the other position, and the new image preview pop-up window comprises a new rendering image, the new rendering image being obtained by simulating the virtual camera shooting at the other position.

2. The method of claim 1, wherein, The target coordinate system comprises a rectangular coordinate system and a polar coordinate system, and the target space parameters between the virtual camera and each physical camera in the target coordinate system are determined based on the physical arrangement position of each physical camera and the virtual arrangement position, comprising the following steps: Based on the virtual arrangement position, a first virtual coordinate of the virtual camera in the rectangular coordinate system and a second virtual coordinate of the virtual camera in the polar coordinate system are respectively determined; For each physical camera, a first physical coordinate of the physical camera in the rectangular coordinate system and a second physical coordinate of the physical camera in the polar coordinate system are respectively determined based on the physical arrangement position of the physical camera; For each physical camera, a first space parameter between the physical camera and the virtual camera is determined according to the first physical coordinate and the first virtual coordinate, and a second space parameter between the physical camera and the virtual camera is determined according to the second physical coordinate and the second virtual coordinate; For each physical camera, the first space parameter and the second space parameter between the physical camera and the virtual camera are taken as the target space parameters between the physical camera and the virtual camera.

3. The method of claim 1, wherein, The virtual camera parameters of the virtual camera at least comprise a pitch angle, a rotation angle of the virtual camera, and a focus point position parameter for focusing on the target object.

4. The method of claim 1, wherein, The rendering image is obtained by image rendering based on the target space parameters corresponding to each physical camera, the plurality of scene images, and the virtual camera parameters of the virtual camera, comprising the following steps: The trained image processing model is called, and target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera are input into the image processing model to obtain the rendered image.

5. The method of claim 1, wherein, The new rendered image generation step includes: The virtual camera parameters are updated to obtain updated virtual camera parameters; The other positions are taken as new virtual arrangement positions, and new space parameters between the virtual camera and each physical camera in a target coordinate system are determined according to the new virtual arrangement positions and the physical arrangement positions of each physical camera; Based on the new space parameters corresponding to each physical camera, the multiple scene images, and the updated virtual camera parameters, image rendering is performed to obtain a new rendered image.

6. An apparatus for generating an image based on a virtual camera, the apparatus comprising: The device includes: A position determination module configured to, in response to a triggering operation on a target position on a video session screen, take the target position as a virtual arrangement position of a virtual camera on the video session screen, and arrange multiple physical cameras around the video session screen; A parameter determination module configured to determine target space parameters between the virtual camera and each physical camera in a target coordinate system based on physical arrangement positions of each physical camera and the virtual arrangement position; An image acquisition module configured to acquire multiple scene images, each of which is obtained by photographing a target object by a corresponding physical camera, and the target object is a video call party that initiates photographing by the virtual camera; An image rendering module configured to generate a rendered image by an image rendering algorithm for new view synthesis based on target space parameters corresponding to each physical camera, multiple scene images, and virtual camera parameters of the virtual camera, the rendered image being obtained by simulating photographing of the target object by the virtual camera at the virtual camera position, and the image rendering algorithm for new view synthesis being NeRF or spatiotemporal high-strength feature spraying; An image preview pop-up window is displayed at the virtual arrangement position on the video session screen, and the image preview pop-up window includes the rendered image; In response to a triggering operation on another position on the video session screen, a new image preview pop-up window is displayed at the other position, and the new image preview pop-up window includes a new rendered image, the new rendered image being obtained by simulating photographing by the virtual camera at the other position.

7. The apparatus of claim 6, wherein, The target coordinate system comprises a rectangular coordinate system and a polar coordinate system, the parameter determination module is configured to determine a first virtual coordinate of the virtual camera in the rectangular coordinate system and a second virtual coordinate of the virtual camera in the polar coordinate system based on the virtual arrangement position; for each physical camera, determine a first physical coordinate of the physical camera in the rectangular coordinate system and a second physical coordinate of the physical camera in the polar coordinate system based on a physical arrangement position of the physical camera; for each physical camera, determine a first spatial parameter between the physical camera and the virtual camera according to the first physical coordinate and the first virtual coordinate, and determine a second spatial parameter between the physical camera and the virtual camera according to the second physical coordinate and the second virtual coordinate; and for each physical camera, take the first spatial parameter and the second spatial parameter between the physical camera and the virtual camera as a target spatial parameter between the physical camera and the virtual camera.

8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method, system and equipment for realizing stereo video communication

    CN101651841A

  • Controlled three-dimensional communication endpoint

    CN104782122A

  • Image processing apparatus, image capturing system, image processing method, and recording medium

    CN111226255A