Video image generation device and method

By analyzing real-time images and generating a three-dimensional portrait model, combined with suitable three-dimensional situational templates, the problem of fixed video image presentation in video conferencing is solved, and a more natural and immersive video conferencing experience is achieved.

CN120021246APending Publication Date: 2025-05-20INSTITUTE FOR INFORMATION INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410002796.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-01-02
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The prior art When multiple participating users conduct video conferences, the video image presentation method is fixed and cannot provide a good conference experience.

Method used

By analyzing multiple real-time images, splitting out the target image, generating a three-dimensional portrait model, and selecting a suitable three-dimensional situation template based on the number of users and the number of positions of the three-dimensional situation templates, synthesizing the three-dimensional portrait model into the spatial annotation position of the template, and generating video images of multiple users.

Benefits of technology

It solves the problem that video images are too rigid and unnatural, provides an immersive experience closer to real scenes, and improves the sense of participation and experience quality of video conferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120021246A_ABST
    Figure CN120021246A_ABST
Patent Text Reader

Abstract

The invention discloses a video image generation device and method. The apparatus analyzes a plurality of instant images corresponding to a plurality of users to segment a target image from each of the plurality of instant images. The device generates a three-dimensional portrait model corresponding to each of the plurality of users based on the target image of each of the plurality of instant images. The apparatus determines a first three-dimensional context template from a plurality of three-dimensional context templates based on a number of users of the plurality of users and a number of positions corresponding to each of the plurality of three-dimensional context templates. The device synthesizes the plurality of three-dimensional portrait models to a spatial labeling position of the first three-dimensional situation template to generate a video image corresponding to the plurality of users. The video image generation technology provided by the invention provides immersive experience which is closer to that of a real scene for users participating in a conference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video image generating device and method. Specifically, the present invention relates to a video image generating device and method for generating video images of multiple users. Background Art

[0002] In the prior art, when multiple participating users are having a video conference, the commonly used video image can usually only be presented in a fixed manner. For example, the video images of each of the multiple users are presented in a fixed two-dimensional arrangement or a traditional background in the video image. In such a case, the presented video image cannot provide a good conference experience for the users participating in the conference.

[0003] In view of this, how to provide a video image generating technology that can generate video images of multiple users is an urgent goal for the industry to strive for. Summary of the Invention

[0004] In view of the above, the present disclosure provides a video image generating device and method for solving the above problems.

[0005] An object of the present disclosure is to provide a video image generating device. The video image generating device includes a transceiver interface, a memory, and a processor. The processor is electrically connected to the transceiver interface and the memory. The memory is used to store a plurality of three-dimensional scenario templates, where each of the plurality of three-dimensional scenario templates corresponds to a position quantity and a spatial marked position. The processor analyzes a plurality of live images corresponding to a plurality of users to cut out a target image from each of the plurality of live images. The processor generates a three-dimensional human figure model corresponding to each of the plurality of users based on the target image of each of the plurality of live images. The processor determines a first three-dimensional scenario template from the plurality of three-dimensional scenario templates based on the user quantity of the plurality of users and the position quantity corresponding to each of the plurality of three-dimensional scenario templates. The processor synthesizes the plurality of three-dimensional human figure models into the spatial marked position of the first three-dimensional scenario template to generate a video image corresponding to the plurality of users.

[0006] Another object of the present disclosure is to provide a method for generating a video image, which is used for an electronic device. The electronic device stores a plurality of three-dimensional scene templates, and each of the plurality of three-dimensional scene templates corresponds to a number of positions and a spatially marked position. The method for generating a video image includes the following steps: analyzing a plurality of live images corresponding to a plurality of users to cut out a target image from each of the plurality of live images; generating a three-dimensional human model corresponding to each of the plurality of users based on the target image of each of the plurality of live images; determining a first three-dimensional scene template from the plurality of three-dimensional scene templates based on the number of users of the plurality of users and the number of positions corresponding to each of the plurality of three-dimensional scene templates; and synthesizing the plurality of three-dimensional human models into the spatially marked position of the first three-dimensional scene template to generate a video image corresponding to the plurality of users.

[0007] In an embodiment of the present invention, the first three-dimensional scene template further includes an environmental parameter, and the processor further performs the following operations: rendering the plurality of three-dimensional human models based on the environmental parameter to generate the video image corresponding to the plurality of users.

[0008] In an embodiment of the present invention, generating the video image corresponding to the plurality of users includes the following operations: inputting the environmental parameter of the first three-dimensional scene template, the plurality of three-dimensional human models, and the spatially marked position of the first three-dimensional scene template into a pre-trained diffusion model to generate the video image corresponding to the plurality of users; wherein the environmental parameter includes a lighting indication vector and a pose attribute vector.

[0009] In an embodiment of the present invention, the spatially marked position of the first three-dimensional scene template further corresponds to a human pose setting, and the processor further performs the following operations: synthesizing the plurality of three-dimensional human models corresponding to the human pose setting into the spatially marked position of the first three-dimensional scene template to generate the video image corresponding to the plurality of users.

[0010] In an embodiment of the present invention, the first three-dimensional scene template further includes a plurality of spatial perspectives, and the processor further performs the following operations: generating a perspective video image corresponding to each of the plurality of spatial perspectives based on the plurality of spatial perspectives; and transmitting the plurality of perspective video images to a playback device based on a perspective switching mechanism so that the playback device performs a playback operation.

[0011] In an embodiment of the present invention, the perspective switching mechanism is a first speaking position priority, a random playback, or a round-robin playback.

[0012] In an embodiment of the present invention, the multiple three-dimensional scenario templates are generated based on the following operations: inputting multiple two-dimensional images and a description text corresponding to each of the multiple two-dimensional images into a depth model to generate the multiple three-dimensional scenario templates, where the depth model is trained by multiple scene depth maps.

[0013] In an embodiment of the present invention, the processor further performs the following operations: based on the target image of each of the multiple real-time images, instantaneously render the multiple three-dimensional human models in the video image to update the video image.

[0014] In an embodiment of the present invention, the operation of segmenting out the target image further includes the following operations: generating an edge block information corresponding to each of the multiple target images based on the edge state of each of the multiple target images; and synthesizing the multiple edge block information and the multiple three-dimensional human models to the spatial annotation position of the first three-dimensional scenario template to generate the video image corresponding to the multiple users.

[0015] In an embodiment of the present invention, the processor further performs the following operations: receiving a perspective switching signal corresponding to a first user, where the perspective switching signal is used to indicate switching to a first spatial perspective; and generating a first perspective video image corresponding to the first spatial perspective based on the perspective switching signal.

[0016] In an embodiment of the present invention, the first three-dimensional scenario template further includes an environmental parameter, and the video image generation method further includes the following steps: rendering the multiple three-dimensional human models based on the environmental parameter to generate the video image corresponding to the multiple users.

[0017] In an embodiment of the present invention, generating the video image corresponding to the multiple users includes the following steps: inputting the environmental parameter of the first three-dimensional scenario template, the multiple three-dimensional human models, and the spatial annotation position of the first three-dimensional scenario template into a pre-trained diffusion model to generate the video image corresponding to the multiple users; where the environmental parameter includes a lighting indication vector and a pose attribute vector.

[0018] In an embodiment of the present invention, the spatial annotation position of the first three-dimensional scenario template further corresponds to a human pose setting, and the video image generation method further includes the following steps: synthesizing the multiple three-dimensional human models corresponding to the human pose setting to the spatial annotation position of the first three-dimensional scenario template to generate the video image corresponding to the multiple users.

[0019] In an embodiment of the present invention, the first three-dimensional scenario template further includes a plurality of spatial perspectives, and the video image generation method further includes the following steps: generating a perspective video image corresponding to each of the plurality of spatial perspectives based on the plurality of spatial perspectives; and transmitting the plurality of perspective video images to a playback device based on a perspective switching mechanism, so that the playback device performs a playback operation.

[0020] In an embodiment of the present invention, the perspective switching mechanism is a first speaking position priority, a random playback, or a round-robin playback.

[0021] In an embodiment of the present invention, the plurality of three-dimensional scenario templates are generated based on the following steps: inputting a plurality of two-dimensional images and a description text corresponding to each of the plurality of two-dimensional images into a depth model to generate the plurality of three-dimensional scenario templates, where the depth model is trained by a plurality of scene depth maps.

[0022] In an embodiment of the present invention, the video image generation method further includes the following steps: instantaneously rendering the plurality of three-dimensional human models in the video image based on the target images of the plurality of instant images, so as to update the video image.

[0023] In an embodiment of the present invention, segmenting out the target image further includes the following steps: generating edge block information corresponding to each of the plurality of target images based on the edge state of each of the plurality of target images; and synthesizing the plurality of edge block information and the plurality of three-dimensional human models to the spatial annotation positions of the first three-dimensional scenario template to generate the video image corresponding to the plurality of users.

[0024] In an embodiment of the present invention, the video image generation method further includes the following steps: receiving a perspective switching signal corresponding to a first user, where the perspective switching signal is used to indicate switching to a first spatial perspective; and generating a first perspective video image corresponding to the first spatial perspective based on the perspective switching signal.

[0025] The video image generation technology provided by this disclosure (including at least devices and methods) generates three-dimensional portrait models corresponding to each of the multiple users by segmenting target images from each of the multiple instant images. Next, the video image generation technology provided by this disclosure determines a suitable three-dimensional scenario template from the multiple three-dimensional scenario templates based on the number of users and the number of positions corresponding to each of the multiple three-dimensional scenario templates. Finally, the video image generation technology provided by this disclosure synthesizes the multiple three-dimensional portrait models into the spatial annotation positions of the three-dimensional scenario template to generate video images corresponding to the multiple users. The video image generation technology provided by this disclosure can correspondingly select a suitable three-dimensional scenario template and adaptively synthesize the three-dimensional portrait model into the three-dimensional scenario template, thus solving the drawbacks in the prior art that the video images are too rigid and the pictures are not natural, and providing users participating in the meeting with an immersive experience closer to the real scene.

[0026] The following elaborates on the detailed technology and implementation manners of this disclosure with reference to the accompanying drawings, so that those with ordinary knowledge in the technical field to which this disclosure belongs can understand the technical features of the claimed disclosure. Description of the Drawings

[0027] Figure 1 It is a schematic diagram of a video conferencing system depicting the first embodiment;

[0028] Figure 2 It is a schematic diagram of the architecture of a video image generation device depicting certain embodiments;

[0029] Figure 3A It is a schematic diagram of an instant image depicting certain embodiments;

[0030] Figure 3B It is a schematic diagram of an image after removing the background depicting certain embodiments;

[0031] Figure 4 It is a schematic diagram of a three-dimensional scenario template depicting certain embodiments;

[0032] Figure 5 It is a schematic diagram of a spatial perspective image depicting certain embodiments; and

[0033] Figure 6 It is a partial flowchart of a video image generation method depicting the second embodiment.

[0034]

Symbol Explanation

[0035] 1: Video conferencing system

[0036] 2: Video image generation device

[0037] U1, U2, U3: User devices

[0038] 21: Transceiver Interface

[0039] 23: Memory

[0040] 25: Processor

[0041] T1, T2, …, Tn: 3D Scenario Templates

[0042] 300: Live Image

[0043] 301: Image after Background Removal

[0044] TA1, TA2: Target Images

[0045] 400: 3D Scenario Template

[0046] P1, P2, P3, P4: Spatial Annotation Positions

[0047] 500: Spatial Perspective Image

[0048] 600: Method for Generating Video Image

[0049] S601, S603, S605, S607: Steps Detailed Implementation Manner

[0050] The following will explain a video image generating device and method provided by the present disclosure through implementation manners. However, these implementation manners are not used to limit that the present disclosure must be implemented in any environment, application, or manner as described in these implementation manners. Therefore, the description of the implementation manners is only for the purpose of explaining the present disclosure, rather than for limiting the scope of the present disclosure. It should be understood that in the following implementation manners and drawings, elements not directly related to the present disclosure have been omitted and not shown, and the dimensions of each element and the dimensional ratios between elements are only for illustration, rather than for limiting the scope of the present disclosure.

[0051] First, the applicable scenario of the present disclosure will be described. Its schematic diagram is depicted in Figure 1 . As Figure 1 shown, in the present disclosure, the video conferencing system 1 includes a video image generating device 2 and user devices U1, U2, U3. In this scenario, the user devices U1, U2, U3 can be connected to the video image generating device 2 through wired or wireless networks. It should be noted that the user devices U1, U2, U3 will continuously generate live images (e.g., at a frequency of 30 frames per second) and transmit the live images to the video image generating device 2, and the video image generating device 2 generates a video image and then transmits the video image back to the user devices U1, U2, U3.

[0052] It should be understood that Figure 1For illustration only, the present disclosure does not limit the number of user devices connected to the video image generation device 2. Those with ordinary knowledge in the art should be able to understand the implementation manners of other numbers of user devices based on the content of the present disclosure, which will not be elaborated herein.

[0053] A first implementation manner of the present disclosure is the video image generation device 2, and its architecture schematic diagram is depicted in Figure 2 . The video image generation device 2 includes a transceiver interface 21, a memory 23, and a processor 25. The processor 25 is electrically connected to the transceiver interface 21 and the memory 23. In this implementation manner, the memory 23 is used to store a plurality of three-dimensional scenario templates T1, T2,..., Tn.

[0054] In this implementation manner, each of the plurality of three-dimensional scenario templates T1, T2,..., Tn corresponds to a number of positions and a spatially marked position. It should be noted that each of the three-dimensional scenario templates T1, T2,..., Tn can correspond to a three-dimensional mesh model of a different scenario or a different environmental space, and each of the three-dimensional scenario templates T1, T2,..., Tn can correspond to a suitable number of participating users and a preset position in the space.

[0055] For easy understanding, please refer to Figure 4 a schematic diagram 400 of a three-dimensional scenario template illustrated in Figure 4 . As shown in

[0056] , it illustrates a three-dimensional scenario template of a living room space. In this example, the three-dimensional scenario template includes spatially marked positions P1, P2, P3, and P4 (i.e., the positions where the images of the participating users will be synthesized), and the number of positions is 4 (i.e., the maximum number of participating users is 4).

[0057] First, in this implementation manner, the processor 25 analyzes a plurality of instant images corresponding to a plurality of users (such as the users operating the user devices U1, U2, U3) to cut out target images from each of the plurality of instant images.

[0058] For ease of understanding, please refer to Figure 3A one of the instant image schematic diagrams 300 illustrated in Figure 3B and the image schematic diagram 301 after removing the background of the image illustrated in. In Figure 3A and Figure 3B the processor 25 receives the instant image 300 from the user device, and by performing a method of image matting operation, generates the image 301 after removing the background image and the segmented target images TA1 and TA2.

[0059] It should be noted that the processor 25 can perform a matting operation on the instant image by executing a high-resolution image matting algorithm (such as: RobustVideoMatting algorithm or Deep Image Matting algorithm, etc.) to generate a high-resolution target image.

[0060] Next, the processor 25 generates a three-dimensional human portrait model corresponding to each of the multiple users based on the target image of each of the multiple instant images.

[0061] In some embodiments, a three-dimensional human portrait model corresponding to a user (such as: virtual avatar, PV3D, tyleNerf, StyleSDF, EG3D, etc.) can be constructed through multiple instant images corresponding to a user via a trained diffusion model.

[0062] In some embodiments, the processor 25 can reduce the quantity of instant image data and the modeling time required for building the model based on the technology of RODIN Diffusion.

[0063] Next, the processor 25 determines a suitable three-dimensional context template (such as: the first three-dimensional context template referred to in some embodiments) from the three-dimensional context templates T1, T2,..., Tn based on the number of users among the multiple users and the number of positions corresponding to each of the multiple three-dimensional context templates. For ease of description, the first three-dimensional context template will be used as a substitute for the selected three-dimensional context template hereinafter.

[0064] For example, if the number of users participating in this meeting is 2 - 4, the processor 25 determines a three-dimensional context template that meets the number of positions from the three-dimensional context templates T1, T2,..., Tn (such as: Figure 4 the three-dimensional context template shown), and uses this three-dimensional context template as the background space of the current video image for subsequent synthesis.

[0065] In some embodiments, the processor 25 may also select a suitable three-dimensional scenario template from the three-dimensional scenario templates T1, T2, …, Tn based on other conditions (such as: the meeting theme, the ages of the participating users, the styles of the participating users, the distance relationships of various locations, the suitable spatial environment, etc.), which will not be elaborated here.

[0066] It should be noted that the three-dimensional scenario template can be generated by a pre-trained training model. Specifically, the processor 25 can input multiple two-dimensional images (such as: images corresponding to spaces) and descriptive texts corresponding to each of the multiple two-dimensional images (such as: descriptions corresponding to spaces) into a depth model to generate the multiple three-dimensional scenario templates, where the depth model is trained by multiple scene depth maps.

[0067] For example, the processor 25 can collect multiple two-dimensional images corresponding to multiple different spaces and with depth information. Then, the processor 25 can obtain a scene depth map by inputting the multiple two-dimensional images into a depth evaluation model. Then, the processor 25 combines the scene depth map with a pre-trained text-to-image model (such as: an AI model) to generate a three-dimensional grid model of a target space (such as: a study environment, a meeting room environment, or a classroom environment, etc.). Finally, the processor 25 labels the number of positions and the spatial annotation positions of the three-dimensional grid model to generate corresponding three-dimensional scenario templates.

[0068] Finally, the processor 25 synthesizes the multiple three-dimensional human models to the spatial annotation positions of the first three-dimensional scenario template to generate a video image corresponding to the multiple users.

[0069] For ease of understanding, please refer to Figure 4 . In this example, the processor 25 can synthesize user images (i.e., three-dimensional human models) corresponding to a quantity of 4 (or less than this quantity) to the spatial annotation positions P1, P2, P3, and P4 and generate a video image corresponding to the multiple users.

[0070] In some embodiments, in order to make the generated video image of better quality, the three-dimensional scenario template may further include environmental parameters (such as: lighting parameters, color tone parameters, shadows, spatial line positions, etc. of each area), and the processor 25 further renders the multiple three-dimensional human models based on the environmental parameters to generate the video image corresponding to the multiple users.

[0071] In some embodiments, in order to better integrate the target image into the three-dimensional scenario template, synthetic operations can be performed through a pre-trained diffusion model. Specifically, the foreground image (e.g., the target image cut from the live video) and the background image (e.g., the three-dimensional scenario template) can be respectively input into the diffusion model, and the multiple three-dimensional portrait models can be synthesized to the spatial annotation positions of the first three-dimensional scenario template by additionally inputting parameters (e.g., the illumination indication vector, the pose attribute vector).

[0072] For example, the processor 25 can use a controllable image synthesis (e.g., ControlCom-Image-Composition or Collage Diffusion) model for synthetic operations.

[0073] In some embodiments, the processor 25 generates the video image corresponding to the multiple users through the following operations: inputting the environmental parameters of the first three-dimensional scenario template (e.g., including an illumination indication vector and a pose attribute vector), the multiple three-dimensional portrait models, and the spatial annotation positions of the first three-dimensional scenario template into a pre-trained diffusion model to generate the video image corresponding to the multiple users.

[0074] In some embodiments, in order to make the video image more conform to the three-dimensional scenario template, each of the spatial annotation positions in the three-dimensional scenario template can be more corresponding to the portrait pose setting (i.e., the preset pose corresponding to the user). Specifically, the spatial annotation positions of the first three-dimensional scenario template are more corresponding to a portrait pose setting, and the processor 25 synthesizes the multiple three-dimensional portrait models corresponding to the portrait pose setting to the spatial annotation positions of the first three-dimensional scenario template to generate the video image corresponding to the multiple users.

[0075] For ease of understanding, please refer to Figure 4 . In this example, the spatial annotation positions P2 and P4 are corresponding to the portrait pose setting where the user is sitting (i.e., when the three-dimensional portrait model is synthesized to this position, it will be presented as a sitting portrait model). In addition, the spatial annotation positions P1 and P3 are corresponding to the portrait pose setting where the user is standing (i.e., when the three-dimensional portrait model is synthesized to this position, it will be presented as a standing portrait model).

[0076] It should be noted that in this disclosure, the portrait pose setting can further include other actions, such as: motion dynamics, specific postures, interaction relationships between positions, etc. This disclosure is not limited thereto.

[0077] In some embodiments, in order to improve the quality of the video film, each of the three-dimensional scenario templates T1, T2, …, Tn can further include multiple spatial perspectives, enabling the video image to be played through different perspectives (e.g., played in turn from different perspectives).

[0078] Specifically, the processor 25 generates a perspective video image corresponding to each of the multiple spatial perspectives. Then, based on a perspective switching mechanism, the processor 25 transmits the multiple perspective video images to a playback device for the playback device to perform a playback operation.

[0079] For example, the perspective switching mechanism can be a speaking position priority (i.e., switching to the perspective image corresponding to the speaker), a random playback, or a round-robin playback.

[0080] For ease of understanding, please refer to Figure 5 , Figure 5 FIG. 500 illustrates a schematic diagram of a spatial perspective image. In this example, the processor 25 can switch from other spatial perspectives to the spatial perspective such as in FIG. 500 and generate video images corresponding to the multiple users based on this spatial perspective.

[0081] In some embodiments, to improve the quality of the video film, the user can more freely adjust the perspective they want to view. Specifically, the processor 25 receives a perspective switching signal corresponding to a first user (e.g., a perspective switching signal transmitted by the user device), where the perspective switching signal is used to indicate switching to a first spatial perspective. Then, based on the perspective switching signal, the processor 25 generates a first perspective video image corresponding to the first spatial perspective.

[0082] In some embodiments, the processor 25 can further dynamically update the images of the users in the video image based on the real-time images of the users, so that the users participating in the meeting can interact more dynamically. Specifically, the processor 25 instantaneously renders the multiple three-dimensional portrait models in the video image based on the target images of each of the multiple real-time images to update the video image.

[0083] In some embodiments, the processor 25 segments the target images from the real-time images in a way that preserves edge information. Therefore, when synthesizing the target images into the three-dimensional scene template, the quality of the video image can be better by synthesizing the edge information into the three-dimensional scene template. Specifically, the processor 25 generates an edge block information corresponding to each of the multiple target images based on the edge state of each of the multiple target images. Then, the processor 25 synthesizes the multiple edge block information and the multiple three-dimensional portrait models into the spatial annotation positions of the first three-dimensional scene template to generate the video image corresponding to the multiple users.

[0084] As can be seen from the above description, the video image generating device 2 provided by the present disclosure generates three-dimensional portrait models corresponding to each of the multiple users by cutting out target images from each of the multiple real-time images. Next, the video image generating device 2 provided by the present disclosure determines a suitable three-dimensional scene template from the multiple three-dimensional scene templates based on the number of users and the number of positions corresponding to each of the multiple three-dimensional scene templates. Finally, the video image generating device 2 provided by the present disclosure synthesizes the multiple three-dimensional portrait models to the spatial annotation positions of the three-dimensional scene template to generate a video image corresponding to the multiple users. The video image generating device 2 provided by the present disclosure can correspondingly select a suitable three-dimensional scene template and adaptively synthesize the three-dimensional portrait model to the three-dimensional scene template, thus solving the drawbacks of the prior art that the video image is too rigid and the picture is not natural, and providing an immersive experience closer to the real scene for the users participating in the meeting.

[0085] The second embodiment of the present disclosure is a video image generating method, and its flowchart is depicted in Figure 6 . The video image generating method 600 is applicable to an electronic device, such as the video image generating device 2 described in the first embodiment. The electronic device stores multiple three-dimensional scene templates, and each of the multiple three-dimensional scene templates corresponds to a position number and a spatial annotation position, such as the three-dimensional scene templates T1, T2,..., Tn described in the first embodiment. The video image generating method 600 generates a video image corresponding to multiple users through steps S601 to S607.

[0086] In step S601, the electronic device analyzes multiple real-time images corresponding to multiple users to cut out a target image from each of the multiple real-time images.

[0087] Next, in step S603, the electronic device generates a three-dimensional portrait model corresponding to each of the multiple users based on the target image of each of the multiple real-time images.

[0088] Next, in step S605, the electronic device determines a first three-dimensional scene template from the multiple three-dimensional scene templates based on the number of users among the multiple users and the number of positions corresponding to each of the multiple three-dimensional scene templates.

[0089] Finally, in step S607, the electronic device synthesizes the multiple three-dimensional portrait models to the spatial annotation positions of the first three-dimensional scene template to generate a video image corresponding to the multiple users.

[0090] In some embodiments, where the first three-dimensional scenario template further includes an environmental parameter, the video image generation method 600 further includes the following steps: rendering the plurality of three-dimensional portrait models based on the environmental parameter to generate the video image corresponding to the plurality of users.

[0091] In some embodiments, where generating the video image corresponding to the plurality of users includes the following steps: inputting the environmental parameter of the first three-dimensional scenario template, the plurality of three-dimensional portrait models, and the spatial annotation positions of the first three-dimensional scenario template into a pre-trained diffusion model to generate the video image corresponding to the plurality of users; wherein, the environmental parameter includes a lighting indication vector and a pose attribute vector.

[0092] In some embodiments, where the spatial annotation positions of the first three-dimensional scenario template further correspond to a portrait pose setting, and the video image generation method further includes the following steps: synthesizing the plurality of three-dimensional portrait models corresponding to the portrait pose setting to the spatial annotation positions of the first three-dimensional scenario template to generate the video image corresponding to the plurality of users.

[0093] In some embodiments, where the first three-dimensional scenario template further includes a plurality of spatial perspectives, and the video image generation method 600 further includes the following steps: generating a perspective video image corresponding to each of the plurality of spatial perspectives based on the plurality of spatial perspectives; and transmitting the plurality of perspective video images to a playback device based on a perspective switching mechanism so that the playback device performs a playback operation.

[0094] In some embodiments, where the perspective switching mechanism is a speaking position priority, a random playback, or a round-robin playback.

[0095] In some embodiments, where the plurality of three-dimensional scenario templates are generated based on the following steps: inputting a plurality of two-dimensional images and a description text corresponding to each of the plurality of two-dimensional images into a depth model to generate the plurality of three-dimensional scenario templates, where the depth model is trained by a plurality of scene depth maps.

[0096] In some embodiments, the video image generation method 600 further includes the following steps: instantaneously rendering the plurality of three-dimensional portrait models in the video image based on the target image of each of the plurality of instant images to update the video image.

[0097] In some embodiments, the step of segmenting the target images further includes: generating edge block information corresponding to each of the multiple target images based on the edge state of each of the multiple target images; and synthesizing the multiple edge block information and the multiple three-dimensional portrait models to the spatial annotation positions of the first three-dimensional scene template to generate the video image corresponding to the multiple users.

[0098] In some embodiments, the video image generation method 600 further includes the following steps: receiving a perspective switching signal corresponding to a first user, where the perspective switching signal is used to indicate switching to a first spatial perspective; and generating a first perspective video image corresponding to the first spatial perspective based on the perspective switching signal.

[0099] In addition to the above steps, the second embodiment can also execute all the operations and steps of the video image generation device 2 described in the first embodiment, having the same functions and achieving the same technical effects. Those with ordinary knowledge in the technical field to which this disclosure pertains can directly understand how the second embodiment executes these operations and steps based on the above first embodiment, having the same functions and achieving the same technical effects, so details are not repeated.

[0100] In summary, the video image generation technology provided by this disclosure (at least including the device and method) segments the target images from each of the multiple live images to generate three-dimensional portrait models corresponding to each of the multiple users. Then, the video image generation technology provided by this disclosure determines a suitable three-dimensional scene template from the multiple three-dimensional scene templates based on the number of users and the number of positions corresponding to each of the multiple three-dimensional scene templates. Finally, the video image generation technology provided by this disclosure synthesizes the multiple three-dimensional portrait models to the spatial annotation positions of the three-dimensional scene template to generate the video image corresponding to the multiple users. The video image generation technology provided by this disclosure can correspondingly select a suitable three-dimensional scene template and adaptively synthesize the three-dimensional portrait models to the three-dimensional scene template, thus solving the drawbacks in the prior art that the video image is too rigid and the picture is not natural, and providing an immersive experience closer to the real scene for the users participating in the meeting.

[0101] The above embodiments are only used to illustrate some implementation aspects of this disclosure and to explain the technical features of this disclosure, rather than to limit the protection scope and range of this disclosure. Any changes or equivalent arrangements that can be easily completed by those with ordinary knowledge in the technical field to which this disclosure pertains fall within the scope claimed by this disclosure, and the scope of the right protection of this disclosure is subject to the claims.

Claims

1. A video image generating device, characterized in that: Include: A transceiver interface; A storage device for storing a plurality of three-dimensional situation templates, wherein each of the plurality of three-dimensional situation templates corresponds to a position number and a spatial annotation position; as well as A processor is electrically connected to the transceiver interface and the storage, and is used to perform the following operations: Analyzing a plurality of real-time images corresponding to a plurality of users to segment a target image from each of the plurality of real-time images; Based on the target image of each of the multiple real-time images, generate a three-dimensional portrait model corresponding to each of the multiple users; Determine a first three-dimensional situation template from the plurality of three-dimensional situation templates based on a number of users of the plurality of users and the number of positions corresponding to each of the plurality of three-dimensional situation templates; as well as The multiple three-dimensional portrait models are synthesized to the space annotation position of the first three-dimensional situation template to generate a video image corresponding to the multiple users.

2. The video image generating device according to claim 1, characterized in that: The first three-dimensional situation template further includes an environment parameter, and the processor further performs the following operations: The multiple three-dimensional portrait models are rendered based on the environmental parameters to generate the video images corresponding to the multiple users.

3. The video image generating device according to claim 1, characterized in that: Generating the video image corresponding to the plurality of users includes the following operations: Inputting an environmental parameter of the first three-dimensional situation template, the plurality of three-dimensional portrait models and the spatial annotation position of the first three-dimensional situation template into a pre-trained diffusion model to generate the video image corresponding to the plurality of users; The environmental parameter includes a lighting indication vector and a posture attribute vector.

4. The video image generating device according to claim 1, wherein: The spatial annotation position of the first three-dimensional situation template further corresponds to a portrait posture setting, and the processor further performs the following operations: The multiple three-dimensional portrait models corresponding to the portrait posture settings are synthesized to the space annotation position of the first three-dimensional situation template to generate the video images corresponding to the multiple users.

5. The video image generating device according to claim 1, characterized in that: The first three-dimensional scenario template further includes a plurality of spatial perspectives, and the processor further performs the following operations: Based on the multiple spatial perspectives, generating a perspective video image corresponding to each of the multiple spatial perspectives; and Based on a switching perspective mechanism, the plurality of perspective video images are transmitted to a playback device so that the playback device performs a playback operation.

6. The video image generating device according to claim 5, characterized in that: The switching perspective mechanism is a speaking position priority, a random play or a round-robin play.

7. The video image generating device according to claim 1, characterized in that: The multiple three-dimensional scenario templates are generated based on the following operations: A plurality of two-dimensional images and a description text corresponding to each of the plurality of two-dimensional images are input into a depth model to generate the plurality of three-dimensional scene templates, wherein the depth model is trained by a plurality of scene depth maps.

8. The video image generating device according to claim 1, wherein: The processor also performs the following operations: Based on the target image of each of the multiple real-time images, the multiple three-dimensional portrait models in the video image are rendered in real time to update the video image.

9. The video image generating device according to claim 1, characterized in that: Segmenting the target image also includes the following operations: Based on an edge state of each of the plurality of target images, generating edge block information corresponding to each of the plurality of target images; as well as The plurality of edge block information and the plurality of three-dimensional portrait models are synthesized into the space annotation position of the first three-dimensional situation template to generate the video image corresponding to the plurality of users.

10. The video image generating device according to claim 1, wherein: The processor also performs the following operations: receiving a perspective switching signal corresponding to a first user, wherein the perspective switching signal is used to indicate switching to a first spatial perspective; and Based on the perspective switching signal, a first perspective video image corresponding to the first spatial perspective is generated.

11. A method for generating a video image, characterized in that: For an electronic device, the electronic device stores a plurality of three-dimensional situation templates, each of the plurality of three-dimensional situation templates corresponds to a position number and a spatial annotation position, and the video image generation method comprises the following steps: Analyzing a plurality of real-time images corresponding to a plurality of users to segment a target image from each of the plurality of real-time images; Based on the target image of each of the multiple real-time images, generate a three-dimensional portrait model corresponding to each of the multiple users; Determine a first three-dimensional situation template from the plurality of three-dimensional situation templates based on a number of users of the plurality of users and the number of positions corresponding to each of the plurality of three-dimensional situation templates; as well as The multiple three-dimensional portrait models are synthesized to the space annotation position of the first three-dimensional situation template to generate a video image corresponding to the multiple users.

12. The video image generation method according to claim 11, characterized in that: The first three-dimensional situation template further includes an environmental parameter, and the video image generation method further includes the following steps: The multiple three-dimensional portrait models are rendered based on the environmental parameters to generate the video images corresponding to the multiple users.

13. The video image generation method according to claim 11, characterized in that: Generating the video image corresponding to the plurality of users comprises the following steps: Inputting an environmental parameter of the first three-dimensional situation template, the plurality of three-dimensional portrait models and the spatial annotation position of the first three-dimensional situation template into a pre-trained diffusion model to generate the video image corresponding to the plurality of users; The environmental parameter includes a lighting indication vector and a posture attribute vector.

14. The video image generation method according to claim 11, wherein: The space annotation position of the first three-dimensional situation template further corresponds to a portrait posture setting, and the video image generation method further includes the following steps: The multiple three-dimensional portrait models corresponding to the portrait posture settings are synthesized to the space annotation position of the first three-dimensional situation template to generate the video images corresponding to the multiple users.

15. The video image generation method according to claim 11, characterized in that: The first three-dimensional situation template also includes a plurality of spatial perspectives, and the video image generation method further includes the following steps: Based on the multiple spatial perspectives, generating a perspective video image corresponding to each of the multiple spatial perspectives; and Based on a switching perspective mechanism, the plurality of perspective video images are transmitted to a playback device so that the playback device performs a playback operation.

16. The video image generation method according to claim 15, characterized in that: The switching perspective mechanism is a speaking position priority, a random play or a round-robin play.

17. The video image generation method according to claim 11, characterized in that: The multiple three-dimensional situation templates are generated based on the following steps: A plurality of two-dimensional images and a description text corresponding to each of the plurality of two-dimensional images are input into a depth model to generate the plurality of three-dimensional scene templates, wherein the depth model is trained by a plurality of scene depth maps.

18. The video image generation method according to claim 11, characterized in that: The video image generation method further comprises the following steps: Based on the target image of each of the multiple real-time images, the multiple three-dimensional portrait models in the video image are rendered in real time to update the video image.

19. The video image generation method according to claim 11, characterized in that: Segmenting the target image also includes the following steps: Based on an edge state of each of the plurality of target images, generating edge block information corresponding to each of the plurality of target images; as well as The plurality of edge block information and the plurality of three-dimensional portrait models are synthesized into the space annotation position of the first three-dimensional situation template to generate the video image corresponding to the plurality of users.

20. The video image generation method according to claim 11, characterized in that: The video image generation method further comprises the following steps: receiving a perspective switching signal corresponding to a first user, wherein the perspective switching signal is used to indicate switching to a first spatial perspective; and Based on the perspective switching signal, a first perspective video image corresponding to the first spatial perspective is generated.