Method and apparatus for virtual human video fusion

By collecting and fusing depth information from virtual human videos and scene videos using binocular cameras, the problem of inaccurate virtual human position display was solved, achieving high-precision virtual human video fusion and improving both cost-effectiveness and viewing experience.

CN115811587BActive Publication Date: 2026-01-27AVIT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211438154.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2026-01-27
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

In existing technologies, the positional information of the video field is lost when virtual human videos are synthesized with background videos, making it impossible to accurately express the relative position of the virtual human in the video scene, and it is also not economical.

Method used

A binocular camera is used to capture scene video, obtain the first depth information of the scene objects, import the virtual human image and associate it with the second depth information, and fuse the two depth information to generate a virtual human fused video and transmit it to the display device.

Benefits of technology

It improves the accuracy of virtual human positioning in video scenes, enhances the viewing experience, and reduces bandwidth usage and hardware requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115811587B_ABST
    Figure CN115811587B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for virtual human video fusion, which comprises the following steps: collecting a scene video based on a binocular camera, processing the scene video to obtain first depth information of at least one object in the scene, importing multiple images of virtual humans into the scene video, associating each image with second depth information of the virtual human, fusing the images of the virtual human and the scene video based on the first depth information and the second depth information to obtain a virtual human fusion video and transmitting the virtual human fusion video to a display device for displaying the virtual human fusion video on an interactive interface. The application improves the accuracy of the position display of the virtual human in the video scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and more particularly to a method and apparatus for virtual human video fusion. Background Technology

[0002] Currently, with the rapid development of technology, virtual humans are accounting for an increasingly larger proportion in the video field.

[0003] In existing technologies, background video and virtual human video are imported into the broadcast console and composited. However, in this method, the position information of the video field is lost when the broadcast console composites the video, and the relative position information of the virtual human in the video scene cannot be accurately expressed. Summary of the Invention

[0004] In view of this, this application provides a method and apparatus for virtual human video fusion, which aims to improve the accuracy of virtual human position display in video scenes.

[0005] To achieve the above objectives, this application provides a method for virtual human video fusion, the method comprising:

[0006] Based on a binocular camera, scene video is acquired and processed to obtain the first depth information of at least one object in the scene;

[0007] Multiple images of virtual humans are imported into the scene video; each image is associated with the second depth information of the virtual human.

[0008] Based on the first depth information and the second depth information, the images of each virtual human and the scene video are fused to obtain a virtual human fused video, which is then transmitted to a display device for display on the interactive interface.

[0009] In one possible implementation of this application, processing the scene video to obtain first depth information of at least one object in the scene includes:

[0010] Based on the OpenCV algorithm library, edge extraction is performed on each frame of the scene image in the scene video to determine at least one object in the scene.

[0011] The first depth information of each object in each image is determined, and three-dimensional reconstruction is performed to obtain the current scene.

[0012] In one possible implementation of this application, the step of fusing the images of each virtual human and the scene video based on the first depth information and the second depth information to obtain a virtual human fused video and transmitting it to a display device includes:

[0013] Based on the first depth information and the second depth information, the image of the virtual human is fused with the corresponding scene image to obtain multiple fused images;

[0014] The merged images are combined to obtain a virtual human merged video, which is then transmitted to a display device.

[0015] In one possible implementation of this application, before importing multiple images of virtual humans into the scene video, the following steps are included:

[0016] Constructing a 3D model of a virtual human;

[0017] Obtain multiple images of the three-dimensional model.

[0018] In one possible implementation of this application, the processing of the scene video includes:

[0019] Extract multiple consecutive frames of scene images from the acquired scene video;

[0020] Determine whether the illuminance of each scene image is less than the preset illuminance;

[0021] If the illuminance of the scene image is less than the preset illuminance, it is marked as an image to be processed;

[0022] The image to be processed is subjected to image enhancement processing to achieve a preset illumination level.

[0023] In one possible implementation of this application, the image enhancement processing of the image to be processed to achieve a preset illumination includes:

[0024] The image to be processed is decomposed into an illuminated image and a reflected image;

[0025] The illumination image is subjected to illuminance enhancement processing to obtain the target illumination image;

[0026] The reflection image is denoised to obtain the target reflection image;

[0027] Based on the target illumination image and the target reflection image, the image to be processed is reconstructed to achieve the preset illumination.

[0028] In one possible implementation of this application, the step of fusing the images of each virtual human and the scene video based on the first depth information and the second depth information includes:

[0029] Construct a mask model for the edges of a virtual human;

[0030] Based on the mask model, the edges of the virtual human are denoised.

[0031] For example, to achieve the above objective, this application also provides an apparatus for virtual human video fusion, the apparatus comprising:

[0032] The acquisition module is used to acquire scene video based on a binocular camera and process the scene video to obtain the first depth information of at least one object in the scene;

[0033] The import module is used to import multiple images of virtual humans into the scene video; each image is associated with the second depth information of the virtual human.

[0034] The fusion module is used to fuse the images of each virtual human and the scene video based on the first depth information and the second depth information to obtain a virtual human fused video and transmit it to the display device so that the display device can display the virtual human fused video on the interactive interface.

[0035] Compared to existing technologies that import background video and virtual human video together into a broadcast console for video compositing, this method loses the positional information of the video field during video compositing, failing to accurately represent the relative position of the virtual human within the video scene. This application, based on a binocular camera, acquires scene video and processes it to obtain first depth information of at least one object in the scene; it imports multiple virtual human images into the scene video; each image is associated with the virtual human's second depth information; based on the first and second depth information, it fuses the images of each virtual human with the scene video to obtain a fused virtual human video, which is then transmitted to a display device for display on an interactive interface. This application inputs virtual human images into the scene video and fuses them according to the first and second depth information, accurately representing the relative position of the virtual human within the video scene. Therefore, this application improves the accuracy of displaying the virtual human's position within the video scene. Attached Figure Description

[0036] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating the first embodiment of the virtual human video fusion method of this application.

[0039] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0040] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0041] This application provides a method for virtual human video fusion, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the virtual human video fusion method of this application.

[0042] This application provides embodiments of a method for virtual human video fusion. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order. For ease of description, the following omits the various steps of the virtual human video fusion method, which includes:

[0043] Step S10: Based on the binocular camera, acquire scene video and process the scene video to obtain the first depth information of at least one object in the scene.

[0044] Step S20: Import multiple images of virtual humans into the scene video; each image is associated with the second depth information of the virtual human.

[0045] Step S30: Based on the first depth information and the second depth information, fuse the images of each virtual human and the scene video to obtain a virtual human fused video and transmit it to the display device so that the display device can display the virtual human fused video on the interactive interface.

[0046] In this embodiment, the specific application scenario is:

[0047] The application scenarios for virtual humans are gradually increasing. For example, virtual humans are often added to event commentary videos and indoor sports videos to enhance the realism and entertainment value of the activities. Existing methods for virtual human video fusion include exporting virtual human videos and scene videos to a control room for fusion. However, this method loses the positional information of the video field during video synthesis, failing to accurately represent the relative position of the virtual human within the video scene. Furthermore, the occluded portion of the scene video is determined by the size of the virtual human video frame; the scene video frame where the virtual human is placed is completely obscured, making it impossible to adjust the occlusion based on the positional relationship between the virtual human and objects. This results in low realism and a poor viewing experience. Additionally, transmitting two videos consumes significant bandwidth. Moreover, using a control room for virtual human video fusion is not economically viable.

[0048] The purpose of this application is to improve the accuracy of virtual human positioning in video scenes.

[0049] Specifically, in this application, the virtual human image is transmitted to a video containing depth information captured by a binocular camera. Fusion is then performed within the binocular camera, ensuring no loss of depth information and improving the accuracy of the virtual human's position display. Furthermore, occlusion of relevant parts can be performed based on the positional relationship between the virtual human and objects, resulting in high realism and enhanced viewing experience. The fused video is then transmitted to the display device; only one video needs to be transmitted, minimizing bandwidth usage. Moreover, video fusion can be performed solely through a binocular camera, eliminating the need for additional hardware such as a broadcast control console, thus improving cost-effectiveness.

[0050] Specifically, in this application, image enhancement processing was performed on each frame of the scene video, and the edges of the virtual human were processed to improve the accuracy of the virtual human video.

[0051] The specific steps are as follows:

[0052] Step S10: Based on the binocular camera, acquire scene video and process the scene video to obtain the first depth information of at least one object in the scene.

[0053] In this embodiment, the scene video is determined based on the virtual human video to be generated. For example, if the virtual human video to be generated is a video of a virtual human riding a mountain bike, then the scene video of the riding journey is obtained. If the virtual human video to be generated is a display of a skiing competition, then the video from the main perspective of the skiing competition is obtained.

[0054] The scene video is acquired by a binocular camera, which contains two sensors that can obtain depth information from two sets of images. After establishing a three-dimensional coordinate system, the depth information of the detected object can be calculated using the known distance between the sensors, so as to determine the three-dimensional coordinates of each point in the scene video.

[0055] For example, the scene video is processed to identify at least one object contained in the scene video and to determine the first depth information of that object. For instance, when capturing a scene video of a press conference, it is determined that the objects included include a table, a curtain, a microphone, etc.

[0056] For example, processing the scene video to obtain first depth information of at least one object in the scene includes:

[0057] Step S11: Based on the OpenCV algorithm library, perform edge extraction on each frame of the scene image in the scene video to determine at least one object in the scene.

[0058] In this embodiment, based on the algorithm in the OpenCV algorithm library, edge extraction is performed on each frame of the scene image in the scene video to determine at least one object in each frame of the scene image.

[0059] Step S12: Determine the first depth information of each object in each image and perform three-dimensional reconstruction to obtain the current scene.

[0060] In this embodiment, the binocular camera includes two sensors, which can obtain depth information from two sets of images respectively. After establishing a three-dimensional coordinate system, the depth information of the detected object can be calculated using the known distance between the sensors, thereby determining the three-dimensional coordinates of multiple points of each object in the scene video. Based on the three-dimensional coordinates of multiple points of each object, a three-dimensional reconstruction is performed to obtain the current scene.

[0061] For example, processing the scene video includes:

[0062] Step S13: Extract multiple consecutive frames of scene images from the acquired scene video.

[0063] In this embodiment, the scene image is a series of consecutive images extracted from the scene video acquired by binocular video.

[0064] Step S14: Determine whether the illuminance of each scene image is less than the preset illuminance.

[0065] In this embodiment, the preset illuminance is set as needed, and no specific limitation is made. For example, the preset illuminance of each second image is 0.5 lx.

[0066] Step S15: If the illuminance of the scene image is less than the preset illuminance, then mark it as an image to be processed.

[0067] In this embodiment, the image to be processed is a second image with an illuminance lower than a preset illuminance. The second image with an illuminance lower than the preset illuminance is marked to obtain the image to be processed, so that the illuminance of the image to be processed can be enhanced in the subsequent process.

[0068] For example, using an F1.2 lens, when the brightness of the subject is as low as 0.04 lx, the amplitude of the video signal output by the camera is 50% of its maximum amplitude, which is 350 mV. Therefore, the minimum illumination of this camera is said to be 0.04 lx / F1.2. If the brightness of the subject is even lower, the amplitude of the video signal output by the camera will not reach 350 mV, resulting in a dark image on the screen that is difficult to discern details.

[0069] Step S16: Perform image enhancement processing on the image to be processed to achieve a preset illumination level.

[0070] In this embodiment, the marked image to be processed is subjected to image enhancement processing to obtain a processed scene image to achieve a preset illumination.

[0071] In this embodiment, scene reconstruction requires feature extraction, digitization, and reconstruction of the acquired real-world scene. The scene images acquired by the binocular camera are affected by the ambient light intensity. If the ambient light intensity is low, the resulting scene image will have low illumination, leading to some loss of information about the real-world scene. Excessive loss of information about the real-world environment in the second image will result in low scene accuracy when constructing the scene based on the image. This application improves the accuracy of scene reconstruction by performing image enhancement processing on scene images with illumination levels below a preset threshold.

[0072] For example, the image enhancement processing of the image to be processed to achieve a preset illumination includes:

[0073] Step S161: Decompose the image to be processed into an illumination image and a reflection image.

[0074] In this embodiment, based on Retinex theory, the initial image can be decomposed into an illumination image of ambient lighting and a reflection image of the object surface reflecting the illumination light. If I(x,y) represents the image information observed by the human eye or acquired by a data acquisition device, L(x,y) represents the illumination image of ambient light, and R(x,y) represents the reflection image of the target object, then the relationship between the initial image, the illumination image, and the reflection image can be expressed as I(x,y) = L(x,y) * R(x,y). The illumination image is related to the ambient light; that is, the stronger the ambient light intensity, the stronger the illumination image intensity; and the weaker the ambient light intensity, the weaker the illumination image intensity. The reflection image is related to the object's own color attributes and is independent of the ambient light intensity.

[0075] Step S162: Perform illumination enhancement processing on the illumination image to obtain the target illumination image.

[0076] In this embodiment, the illumination image is input into the illuminance model, the illuminance of the illumination image is enhanced based on the illuminance model, and the target illumination image is output when the target illuminance is reached.

[0077] The illuminance model is a one-dimensional convolutional neural network model. It inputs labeled training data into a pre-defined training model, which is the initial illuminance model. This model processes the training data to predict the illuminance image's state. The state labels include "qualified" and "unqualified." Training samples with illuminance components reaching the pre-defined illuminance are considered "qualified," while those with illuminance components not reaching the pre-defined illuminance are considered "unqualified."

[0078] For example, the preset illuminance can be set as needed, and this embodiment does not impose any specific limitations.

[0079] For example, the number of training samples of illuminated images can be set as needed, and this embodiment does not impose a specific limitation. For instance, the number of training samples of illuminated images is 300.

[0080] In this embodiment, by using illumination image training samples, a one-dimensional convolutional neural network is used to train illumination images with different illuminance levels, resulting in two state data models: "qualified" and "unqualified." This illuminance model includes these two state data models. A preset training model is used as the initial training model. Illumination image training samples and their state labels are input into the preset training model for iterative training to obtain an illuminance model that meets the accuracy requirements.

[0081] Step S163: Denoise the reflected image to obtain the target reflected image.

[0082] In this embodiment, the reflected image is input into the denoising model, and the reflected image is denoised based on the denoising model. When the preset noise level is reached, the target reflected image is output.

[0083] For example, the preset noise can be set as needed, and this embodiment does not impose any specific limitations.

[0084] The denoising model is a one-dimensional convolutional neural network model, and its construction method is basically the same as the method for constructing the illumination model mentioned above, so it will not be described again here.

[0085] Step S164: Based on the target illumination image and the target reflection image, reconstruct the image to be processed to achieve a preset illuminance.

[0086] In this embodiment, a second image of the target is reconstructed based on the target reflection component and the target illumination component, and its clarity is higher than that of the initial second image.

[0087] Step S20: Import multiple images of virtual humans into the scene video; each image is associated with the second depth information of the virtual human.

[0088] In this embodiment, the virtual human's image consists of multiple two-dimensional images, each associated with the virtual human's second depth information, which is used to determine the virtual human's position within the scene. The virtual human's image is extracted from a video of the virtual human.

[0089] For example, before importing multiple images of virtual humans into the scene video, the process includes:

[0090] Step S21: Construct a 3D model of the virtual human.

[0091] In this embodiment, a video of a human body is captured by a binocular camera, multiple ordered images are extracted from the video, and a three-dimensional model of a virtual human body is constructed based on the changes in the posture of the human body in the images.

[0092] For example, human postures include athletes in a running posture on a track and field, skiers in a skiing posture on a ski slope, and commentators in a commentary posture during a competition.

[0093] For example, multiple parts of the human body in each image are marked as key points for monitoring. These key points are critical parts of the human body, and their determination is based on the body's state of motion. For instance, if the body is running, the head, shoulders, wrists, knees, and other joints are marked as key points; if the body is giving a presentation, the mouth, head, and hands are marked as key points. The first image captured by the binocular camera contains the positional information of each point. The positional information of the marked key points in each first image is extracted. Based on the shooting time sequence of each first image, the positional information of each key point in the first image is sequentially sorted to obtain a positional information matrix corresponding to each key point. These positional information matrices are then concatenated to obtain the posture and motion trajectory information of the person. The positional information matrix of each key point is input into a template model of the virtual human to obtain a three-dimensional motion model of the virtual human in the corresponding scene.

[0094] Step S22: Obtain multiple images of the three-dimensional model.

[0095] In this embodiment, the multiple images are two-dimensional images, which are multiple consecutive frames of images during human movement.

[0096] Step S30: Based on the first depth information and the second depth information, fuse the images of each virtual human and the scene video to obtain a virtual human fused video and transmit it to the display device so that the display device can display the virtual human fused video on the interactive interface.

[0097] In this embodiment, the first depth information is the depth information of the scene objects, and the second depth information is the depth information of the virtual human. The first depth information and the second depth information are fused in the binocular camera to obtain the virtual human fused video and transmit it to the display device so that the display device can display the virtual human fused video on the interactive interface.

[0098] For example, the step of fusing the images of each virtual human and the scene video based on the first depth information and the second depth information to obtain a virtual human fused video and transmitting it to a display device includes:

[0099] Step S31: Based on the first depth information and the second depth information, the image of the virtual human is fused with the corresponding scene image to obtain multiple fused images.

[0100] In this embodiment, the image of the virtual human is fused with the corresponding scene image based on the first depth information and the second depth information to obtain multiple fused images.

[0101] Step S32: Summarize the fused images to obtain the virtual human fused video and transmit it to the display device.

[0102] In this embodiment, the merged images are summarized to obtain a virtual human merged video, which is then transmitted to a display device.

[0103] For example, based on the first depth information and the second depth information, after fusing the images of each virtual human and the scene video, the result includes:

[0104] Step S33: Construct a mask model for the virtual human's edges;

[0105] Step S34: Based on the mask model, perform noise reduction processing on the edges of the virtual human.

[0106] In this embodiment, existing virtual human-video fusion algorithms only consider image feature changes, resulting in noise from the virtual human's image being input into the virtual scene during the fusion process. To prevent noise contamination when the virtual human is integrated into the virtual scene video, a mask model of the virtual human's edges can be constructed to suppress interference and preserve good edge information of the virtual human.

[0107] Compared to existing technologies that import background video and virtual human video together into a broadcast console for video compositing, this method loses the positional information of the video field during video compositing, failing to accurately represent the relative position of the virtual human within the video scene. This application, based on a binocular camera, acquires scene video and processes it to obtain first depth information of at least one object in the scene; it imports multiple virtual human images into the scene video; each image is associated with the virtual human's second depth information; based on the first and second depth information, it fuses the images of each virtual human with the scene video to obtain a fused virtual human video, which is then transmitted to a display device for display on an interactive interface. This application inputs virtual human images into the scene video and fuses them according to the first and second depth information, accurately representing the relative position of the virtual human within the video scene. Therefore, this application improves the accuracy of displaying the virtual human's position within the video scene.

[0108] For example, this application also provides an apparatus for virtual human video fusion, the apparatus comprising:

[0109] The acquisition module is used to acquire scene video based on a binocular camera and process the scene video to obtain the first depth information of at least one object in the scene;

[0110] The import module is used to import multiple images of virtual humans into the scene video; each image is associated with the second depth information of the virtual human.

[0111] The fusion module is used to fuse the images of each virtual human and the scene video based on the first depth information and the second depth information to obtain a virtual human fused video and transmit it to the display device so that the display device can display the virtual human fused video on the interactive interface.

[0112] For example, the acquisition module includes:

[0113] The extraction submodule is used to perform edge extraction on each frame of the scene image in the scene video based on the OpenCV algorithm library to determine at least one object in the scene.

[0114] The determination submodule is used to determine the first depth information of each object in each image and perform three-dimensional reconstruction to obtain the current scene.

[0115] For example, the fusion module includes:

[0116] The fusion submodule is used to fuse the image of the virtual human with the corresponding scene image based on the first depth information and the second depth information to obtain multiple fused images;

[0117] The aggregation submodule is used to aggregate the various fused images, obtain the virtual human fused video, and transmit it to the display device.

[0118] For example, the virtual human video fusion apparatus further includes:

[0119] Modules for building 3D models of virtual humans;

[0120] The acquisition module is used to acquire multiple images of the 3D model;

[0121] For example, the acquisition module includes:

[0122] The extraction submodule is used to extract multiple consecutive frames of scene images from the acquired scene video;

[0123] The judgment submodule is used to determine whether the illuminance of each scene image is less than the preset illuminance;

[0124] The marking submodule is used to mark an image as an image to be processed if the illuminance of the scene image is less than the preset illuminance.

[0125] The enhancement submodule is used to perform image enhancement processing on the image to be processed in order to achieve a preset illumination level.

[0126] For example, the enhancement submodule includes:

[0127] A decomposition unit is used to decompose the image to be processed into an illumination image and a reflection image;

[0128] An enhancement unit is used to perform illuminance enhancement processing on the illumination image to obtain a target illumination image;

[0129] A denoising unit is used to denoise the reflection image to obtain a target reflection image;

[0130] The reconstruction unit is used to reconstruct the image to be processed based on the target illumination image and the target reflection image to achieve a preset illuminance.

[0131] For example, the virtual human video fusion apparatus further includes:

[0132] Modules for building mask models of virtual human edges;

[0133] The denoising module is used to denoise the edges of the virtual human based on the mask model.

[0134] The specific implementation of the virtual human video fusion apparatus of this application is basically the same as the embodiments of the virtual human video fusion method described above, and will not be repeated here.

[0135] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0136] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0137] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, device, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0138] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for virtual human video fusion, characterized in that, The method includes: Based on a binocular camera, scene video and human video are acquired, and the scene video is processed to obtain the first depth information of at least one object in the scene. Multiple ordered frames of images are extracted from the video of the human body. Multiple parts of the human body in each image are marked as key points for monitoring. The key points are key parts of the human body, and the determination of the key parts is based on the movement state of the human body. The first image captured by the binocular camera contains the position information of each key point. The position information of the marked key points in each first image is extracted. According to the shooting time sequence of each first image, the position information of each key point in the first image is sorted sequentially to obtain the position information matrix corresponding to each key point. The position information matrix of each key point is spliced ​​to obtain the posture information and movement trajectory information of the person. The position information matrix of each key point is input into the template model of the virtual human to obtain the three-dimensional motion model of the virtual human in the corresponding scene. According to the change of human posture in the image, the three-dimensional model of the virtual human is constructed, and multiple images of the three-dimensional model are obtained. Multiple images of virtual humans are imported into the scene video; each image is associated with the second depth information of the virtual human. Based on the first depth information and the second depth information, the images of each virtual human and the scene video are fused to obtain a virtual human fused video, which is then transmitted to a display device for display on the interactive interface.

2. The method as described in claim 1, characterized in that, The process of processing the scene video to obtain first depth information of at least one object in the scene includes: Based on the OpenCV algorithm library, edge extraction is performed on each frame of the scene image in the scene video to determine at least one object in the scene. The first depth information of each object in each image is determined, and three-dimensional reconstruction is performed to obtain the current scene.

3. The method as described in claim 1, characterized in that, The step of fusing the images of each virtual human and the scene video based on the first depth information and the second depth information to obtain a virtual human fused video and transmitting it to the display device includes: Based on the first depth information and the second depth information, the image of the virtual human is fused with the corresponding scene image to obtain multiple fused images; The merged images are combined to obtain a virtual human merged video, which is then transmitted to a display device.

4. The method as described in claim 1, characterized in that, The processing of the scene video includes: Extract multiple consecutive frames of scene images from the acquired scene video; Determine whether the illuminance of each scene image is less than the preset illuminance; If the illuminance of the scene image is less than the preset illuminance, it is marked as an image to be processed; The image to be processed is subjected to image enhancement processing to achieve a preset illumination level.

5. The method as described in claim 4, characterized in that, The image enhancement processing of the image to be processed to achieve a preset illumination includes: The image to be processed is decomposed into an illuminated image and a reflected image; The illumination image is subjected to illuminance enhancement processing to obtain the target illumination image; The reflection image is denoised to obtain the target reflection image; Based on the target illumination image and the target reflection image, the image to be processed is reconstructed to achieve the preset illumination.

6. The method as described in claim 1, characterized in that, The process of fusing the images of each virtual human and the scene video based on the first depth information and the second depth information includes: Construct a mask model for the edges of a virtual human; Based on the mask model, the edges of the virtual human are denoised.

7. A device for virtual human video fusion, characterized in that, The device includes: The acquisition module is used to acquire scene video and human video based on a binocular camera, and process the scene video to obtain the first depth information of at least one object in the scene. The construction module is used to extract multiple ordered images from the video of the human body, mark multiple parts of the human body in each image as key points for monitoring, the key points for monitoring are key parts of the human body, and the determination of the key parts is based on the movement state of the human body. The first image captured by the binocular camera contains the position information of each key point. The position information of the marked key points in each first image is extracted. According to the shooting time sequence of each first image, the position information of each key point in the first image is sorted sequentially to obtain the position information matrix corresponding to each key point. The position information matrix of each key point is spliced ​​to obtain the posture information and movement trajectory information of the person. The position information matrix of each key point is input into the template model of the virtual human to obtain the three-dimensional motion model of the virtual human in the corresponding scene. The three-dimensional model of the virtual human is constructed according to the changes in the human body posture in the image. The acquisition module is used to acquire multiple images of the 3D model; The import module is used to import multiple images of virtual humans into the scene video; each image is associated with the second depth information of the virtual human. The fusion module is used to fuse the images of each virtual human and the scene video based on the first depth information and the second depth information to obtain a virtual human fused video and transmit it to the display device so that the display device can display the virtual human fused video on the interactive interface.

Citation Information

Patent Citations

  • Virtual-real occlusion real-time processing method based on multi-view image

    CN106803286A

  • Low-illumination video processing method, device and storage medium

    CN113824943A