Information display method based on dynamic digital human avatar, and electronic device
Generating multi-view videos through multi-camera shooting and image rendering technology solves the problems of high cost and insufficient realism in the existing technology of dynamic digital people, and realizes low-cost, high-fidelity, real-time interactive dynamic digital people display, suitable for multi-terminal devices.
Patent Information
- Application Number
- PCT/CN2024/130678
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-11-08
- Publication Date
- 2025-07-24
AI Technical Summary
The existing technology is expensive and difficult to provide a sense of reality when generating dynamic digital human images. Especially in the model show of e-commerce platforms, the digital humans generated by 3D modeling lack reality, dynamic expressions and motion capture are complex, AI-based portrait synthesis requires a large amount of training data and high-performance hardware, and deep image rendering relies on high-cost hardware and deep information quality.
Real-view videos are captured simultaneously by multi-camera, character images are extracted, intermediate viewing images are estimated using image-based rendering technology, multi-viewing videos are generated, and displayed in virtual 3D spatial scenes. Compress and transmit them in a general video encoder to provide continuous viewing angle switching effect.
It realizes dynamic real-person digital effects that are low-cost, high-fidelity, and real-time interactive. It supports multi-terminal device analysis, avoids complex modeling processes and high costs, and provides a more realistic visual experience and continuous perspective switching.
Smart Images

Figure CN2024130678_24072025_PF_FP_ABST
Abstract
Description
Method and electronic device for displaying information based on dynamic digital human image
[0001] Cross-references
[0002] This application refers to Chinese patent application No. 202410069436X, filed on January 17, 2024, entitled “Method and electronic device for information display based on dynamic digital human image”, which is incorporated into this application in its entirety by reference. Technical Field
[0003] The present application relates to the field of information processing technology, and in particular to a method and electronic device for displaying information based on a dynamic digital human image. Background Art
[0004] With the rapid development of science and technology, especially the maturation of computer vision, augmented reality (AR), and virtual reality (VR), virtual-reality fusion technologies are attracting increasing attention. This technology provides users with an immersive experience by integrating real-world elements with computer-generated virtual elements, opening up new application areas and opportunities across various industries. In particular, in industries such as entertainment, advertising, and e-commerce, the rise of these technologies is not only bringing new experiences to consumers but also opening up new markets and marketing opportunities for businesses. For example, on e-commerce platforms targeting the apparel industry, features such as "model catwalks" and "fitting rooms" can be provided to provide consumers with an immersive visual experience.
[0005] However, achieving these functions also presents new technical and design challenges. For example, in the "model catwalk" scene, it involves providing a more realistic virtual environment and "digital human," etc. Regarding "digital human," existing technologies generally offer the following implementation methods:
[0006] Method 1: Traditional 3D portrait modeling technology: A real-life model is scanned using high-precision 3D scanning equipment, and then texture mapping and detailed sculpting are performed manually or semi-automatically to create a human body model. However, this method typically requires expensive scanning equipment and highly skilled experts to process the scanned data. The time required from scanning to completed model is long, and it is mainly suitable for static modeling; dynamic expressions and motion capture require additional work.
[0007] Method 2: AI-based portrait synthesis: This method uses neural networks and a large amount of training data to directly synthesize or convert portrait perspectives and poses. However, this method requires a large amount of labeled data for training, and the generated results may not be realistic in certain complex scenes and angles. In addition, real-time applications may require high-performance hardware support and high computing requirements.
[0008] Method three, a stereo image rendering solution based on depth images: a new viewpoint image is generated using a color image and a corresponding depth image. Through depth information, this method can estimate the 3D structure of the scene and render the scene from a new viewpoint. However, the output quality of this method is heavily dependent on the quality of the depth image, and low-quality or inaccurate depth information may also cause artifacts or distortion in the rendered image. In addition, the acquisition cost of depth images is relatively high and requires some hardware, such as lidar, to obtain, which may increase cost and complexity.
[0009] Therefore, how to provide users with a more realistic visual experience at a lower cost in the process of information display based on digital humans has become a technical problem that needs to be solved by those skilled in the art.
[0010] Summary of the Invention
[0011] This application provides a method and electronic device for displaying information based on a dynamic digital human image, which can provide users with a more realistic visual experience at a lower cost.
[0012] This application provides the following solutions:
[0013] A method for displaying information based on a dynamic digital human image, comprising:
[0014] Obtaining video content from multiple real perspectives, wherein the video content from the multiple real perspectives is obtained by synchronously capturing a process in which a real person, while wearing a target garment, displays the target garment from different perspectives using multiple camera devices;
[0015] Extracting character images contained in the plurality of video frames from the video content;
[0016] Using image-based rendering technology, based on pixel offset relationship information between character images at adjacent real perspectives at the same time point, character images at corresponding time points are generated for multiple intermediate perspectives between the adjacent real perspectives;
[0017] A multi-perspective video is generated based on the character images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points, so that the multi-perspective video can be matched to a preset virtual 3D space scene model on the client to provide content in which a dynamic digital human image displays the target clothing in the virtual 3D space scene, and to provide an interactive effect for simulating continuous perspective switching.
[0018] The video contents of the multiple real perspectives are obtained by synchronously shooting a process in which a real person wearing target clothing walks in a target space to display the target clothing through multiple camera devices from different perspectives.
[0019] The step of generating character images at corresponding time points for a plurality of intermediate perspectives between the adjacent real perspectives includes:
[0020] Using image-based rendering technology, according to the pixel offset relationship information between the character images of adjacent real perspectives at the same time point, the pixel positions of multiple intermediate perspectives between the adjacent real perspectives at corresponding time points are estimated to generate the character images of the intermediate perspectives at multiple time points.
[0021] The estimating pixel positions of a plurality of intermediate perspectives between the adjacent real perspectives at corresponding time points includes:
[0022] Taking the character images corresponding to adjacent real perspectives at the same time point as input, a dense optical flow field between the character images of adjacent real perspectives is fitted through a deep learning model, and the dense optical flow field is used to estimate the pixel positions of the character images of multiple intermediate perspectives between the adjacent real perspectives at corresponding time points.
[0023] Among them, also include:
[0024] The multi-view video is compressed for transmission to a client.
[0025] The compressing of the multi-view video includes:
[0026] Splicing video frames corresponding to multiple viewpoints at a time point to obtain a frame sequence formed by multiple combined frames; wherein the multiple video frames corresponding to the multiple viewpoints at the time point are divided into multiple sets, and the multiple video frames in each set are spliced into a combined frame, and the video frames of adjacent viewpoints are located at the same position in adjacent combined frames;
[0027] A general video encoder is used to encode a frame sequence formed by the plurality of combined frames, and inter-frame compression processing is performed on the plurality of combined frames.
[0028] The resolution of each combined frame is lower than the maximum resolution supported by the terminal device.
[0029] Among them, also include:
[0030] After encoding and inter-frame compression processing are performed on a frame sequence formed by multiple combined frames, the frame sequence is also sliced so that it can be transmitted in units of fragments obtained after slicing, and independently decoded and played in units of fragments at the receiving end.
[0031] The encoding of the frame sequence formed by the plurality of combined frames further includes:
[0032] According to the number of combined frames included in each slice, the key frame interval in the inter-frame coding process is controlled so as to reduce the number of frames encoded as key frames in the same slice.
[0033] The encoding of the frame sequence formed by the plurality of combined frames further includes:
[0034] For combined frames other than key frames, the number of frames encoded as bidirectional reference frames in the same slice is increased by lowering the judgment threshold for bidirectional reference frames.
[0035] A method for displaying information based on a dynamic digital human image, comprising:
[0036] In response to a viewing request initiated by a user, a multi-perspective video is obtained, wherein the multi-perspective video is generated by: synchronously capturing a process in which a real person, while wearing a target garment, displays the target garment using multiple camera devices from different perspectives to obtain video content from multiple real perspectives; extracting character images contained in multiple video frames included in the video content from the multiple real perspectives; generating character images for multiple intermediate perspectives between adjacent real perspectives using an image-based rendering technique; and generating the multi-perspective video based on the character images from the multiple real perspectives and the multiple intermediate perspectives at multiple time points.
[0037] Decoding the multi-view video;
[0038] Matching the multi-view video to a preset virtual 3D space scene model to provide a dynamic digital human image displaying the target clothing in the virtual 3D space scene;
[0039] In response to the interactive operation of continuous perspective switching, an interactive effect simulating continuous perspective switching is provided by switching to character image video content of other perspectives.
[0040] A method for generating a dynamic digital human image, comprising:
[0041] Obtaining video content from multiple real perspectives, where the video content from the multiple real perspectives is obtained by synchronously capturing a process in which a real person performs a target action from different perspectives using multiple camera devices;
[0042] Extracting character images contained in the plurality of video frames from the video content;
[0043] Using image-based rendering technology, based on pixel offset relationship information between character images at adjacent real perspectives at the same time point, character images at corresponding time points are generated for multiple intermediate perspectives between the adjacent real perspectives;
[0044] A multi-perspective video is generated based on the multiple real perspectives and the character images of the multiple intermediate perspectives at multiple time points, so as to display the corresponding dynamic digital human image through the multi-perspective video.
[0045] A method for displaying clothing information based on a dynamic digital human image, comprising:
[0046] In response to a request to display the target garment using a dynamic digital human, a virtual 3D space scene model and a dynamic digital human image expressed in the form of a multi-view video are obtained, wherein the multi-view video is used to display the target garment while the target person is wearing the target garment from multiple perspectives;
[0047] The multi-view video is matched to the virtual 3D space scene model to provide content in which a dynamic digital human image displays the target clothing in the virtual 3D space scene, and to provide an interactive effect for simulating continuous viewpoint switching.
[0048] A device for displaying information based on a dynamic digital human image, comprising:
[0049] a video content obtaining unit, configured to obtain video content from multiple real perspectives, wherein the video content from the multiple real perspectives is obtained by synchronously shooting a process in which a real person, while wearing a target garment, displays the target garment from different perspectives using multiple camera devices;
[0050] A character image extraction unit, configured to extract character images contained in each of the plurality of video frames contained in the video content;
[0051] a perspective synthesis unit, configured to generate character images at corresponding time points for a plurality of intermediate perspectives between adjacent real perspectives using an image-based rendering technique based on pixel offset relationship information between character images at the same time point from adjacent real perspectives;
[0052] A multi-perspective video storage and unit is used to generate a multi-perspective video based on the character images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points, so that the multi-perspective video can be matched to a preset virtual 3D space scene model on the client to provide content of a dynamic digital human image displaying the target clothing in the virtual 3D space scene, and provide an interactive effect for simulating continuous perspective switching.
[0053] A device for displaying information based on a dynamic digital human image, comprising:
[0054] a multi-perspective video acquisition unit, configured to acquire a multi-perspective video in response to a viewing request initiated by a user, the multi-perspective video being generated by: synchronously capturing, from different perspectives, a process of a real person wearing a target garment and displaying the target garment, using multiple camera devices, to obtain video content from multiple real perspectives; extracting character images contained in multiple video frames included in the video content from the multiple real perspectives; generating character images for multiple intermediate perspectives between adjacent real perspectives using an image-based rendering technique; and generating the multi-perspective video based on the character images from the multiple real perspectives and the multiple intermediate perspectives at multiple time points;
[0055] A decoding unit, configured to decode the multi-view video;
[0056] An adding unit is used to match the multi-view video to a preset virtual 3D space scene model to provide a dynamic digital human image displaying the target clothing in the virtual 3D space scene;
[0057] The perspective switching interaction unit is used to respond to the interactive operation of continuous perspective switching and provide an interactive effect simulating continuous perspective switching by switching to character image video content of other perspectives.
[0058] A device for generating a dynamic digital human image, comprising:
[0059] a video content obtaining unit, configured to obtain video content from multiple real perspectives, wherein the video content from the multiple real perspectives is obtained by synchronously shooting a process in which a real person performs a target action from different perspectives using multiple camera devices;
[0060] A character image extraction unit, configured to extract character images contained in each of the plurality of video frames contained in the video content;
[0061] a perspective synthesis unit, configured to generate character images at corresponding time points for a plurality of intermediate perspectives between adjacent real perspectives using an image-based rendering technique based on pixel offset relationship information between character images at the same time point from adjacent real perspectives;
[0062] The dynamic digital human asset generation unit is used to generate a multi-perspective video based on the character images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points, so as to display the corresponding dynamic digital human image through the multi-perspective video.
[0063] A device for displaying clothing information based on a dynamic digital human image, comprising:
[0064] a request receiving unit, configured to, in response to a request for displaying a target garment through a dynamic digital human, obtain a virtual 3D space scene model and a dynamic digital human image expressed in the form of a multi-view video, wherein the multi-view video is configured to display a process of the target person wearing the target garment displaying the target garment from multiple viewpoints;
[0065] A display unit is used to match the multi-view video to the virtual 3D space scene model to provide content in which a dynamic digital human image displays the target clothing in the virtual 3D space scene, and to provide an interactive effect for simulating continuous perspective switching.
[0066] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any of the aforementioned methods.
[0067] An electronic device, comprising:
[0068] one or more processors; and
[0069] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of any of the aforementioned methods.
[0070] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0071] Through the embodiments of the present application, when it is necessary to provide users with functions such as a "model show", multiple camera devices can be used to synchronously shoot the process of a real person wearing a target garment and displaying the target garment from different perspectives, thereby obtaining video content from multiple real perspectives. Afterwards, the character images contained therein can be extracted from the multiple video frames contained in the video content. Then, using image-based rendering technology, based on the pixel offset relationship information between the character images of adjacent real perspectives at the same time point, character images at corresponding time points can be generated for multiple intermediate perspectives between the adjacent real perspectives. Then, a multi-perspective video can be generated based on the multiple real perspectives and the character images of the multiple intermediate perspectives at multiple time points, so that the character image video content of one of the perspectives can be added to a preset virtual 3D space scene model on the client for rendering and display, thereby providing content of a dynamic digital human displaying the target garment in the virtual 3D space scene and providing an interactive effect for simulating continuous perspective switching. This approach eliminates the need for explicit modeling of the human figure; instead, a 3D-like digital human effect can be achieved using images captured from multiple perspectives, avoiding complex modeling processes and high modeling costs. This results in a low-cost, high-fidelity, and real-time interactive dynamic digital human effect. Furthermore, because the digital human assets created in this embodiment are in a standard video format like HEVC, they can be parsed by most mobile devices, avoiding complex rendering processes and significant computational overhead.
[0072] In addition, for multiple video contents corresponding to multiple viewpoints, specifically during video encoding, the multiple video frames corresponding to the multiple viewpoints can be spliced together at a time point to obtain a frame sequence formed by multiple combined frames. Furthermore, the multiple video frames corresponding to the multiple viewpoints at that time point are divided into multiple sets, and the multiple video frames in each set are spliced together into a combination, such that video frames from adjacent viewpoints are located at the same position in adjacent combined frames. Subsequently, a general video encoder can be used to encode the frame sequence formed by the multiple combined frames, and inter-frame compression processing can be performed on the multiple combined frames to eliminate or reduce redundant information between video frames from adjacent viewpoints. In this way, since the video frames of multiple perspectives are grouped and spliced, the resolution of the spliced combined frames is not too high, which facilitates real-time decoding in most terminal devices. In addition, since the grouping method and arrangement method are controlled during group splicing, the video frames of adjacent perspectives are located in the same position of the adjacent combined frames, that is, the video frames of adjacent perspectives are located in different but adjacent combined frames, and the positions in different combined frames are the same, and the video frames of adjacent perspectives have a relatively high similarity. Therefore, the adjacent combined frames spliced in this way have a relatively high similarity, and then a general inter-frame compression algorithm can be used to eliminate or reduce the redundant information between the video frames of adjacent perspectives, thereby obtaining a higher compression rate. In other words, in the embodiment of the present application, an ideal compression rate can be obtained by a general video encoder, and accordingly, decoding can be completed by using a general decoder at the decoding end, so that it can be supported on more terminal devices.
[0073] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0075] FIG1 is a schematic diagram of a system architecture provided in an embodiment of the present application;
[0076] FIG2 is a flow chart of a first method provided in an embodiment of the present application;
[0077] FIG3 is a schematic diagram of a frame reordering method provided in an embodiment of the present application;
[0078] FIG4 is a flow chart of a second method provided in an embodiment of the present application;
[0079] FIG5 is a flow chart of a third method provided in an embodiment of the present application;
[0080] FIG6 is a flowchart of a fourth method provided in an embodiment of the present application;
[0081] FIG7 is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0082] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0083] First of all, it should be noted that in the embodiments of the present application, it is mainly possible to provide users with scenes such as "model shows" in application systems such as commodity information service systems. The effect to be achieved in this scene is usually: users can view through the client that a "digital human" is wearing a certain outfit, walking in a certain spatial scene, or performing certain display actions, and it is usually necessary to provide users with an interactive effect of continuous perspective switching. Among them, a digital human (Digital Human / Meta Human) is a digital character image close to a human image created using digital technology. In other words, when a user views this "model show" interface through a terminal device such as a mobile phone, he can switch to an "arbitrary" perspective for viewing by sliding on the mobile phone screen, just like the user is personally at the "show" scene and can walk to different positions to watch the model's show process.
[0084] In existing technology, some product information service systems also have products related to "model shows." However, these primarily use modeling to generate "3D digital humans," then match 3D clothing models to these "3D digital humans" to simulate the effect of a real-life "catwalk." This process can provide interactive effects such as multi-perspective viewing. However, because the "3D digital humans" are generated through 3D modeling, they have a distinctly "cartoon-like" feel and lack realism, and their movements, such as walking and posture, are not natural. Furthermore, it is difficult to fully reproduce the characteristics of the clothing through 3D clothing models.
[0085] In the embodiment of the present application, in order to provide users with a more realistic visual experience in scenes such as the "model show" and "fitting room" in the product information service system, so that users can see more realistic human images (rather than cartoon images that appear to lack relative realism), an implementation solution is provided that uses "real digital people" to provide users with functions such as "model show".
[0086] Specifically, multi-view video can replace existing 3D modeling. This means that multiple camera captures can be used to generate videos from multiple perspectives. This multi-view video can then be used to provide users with a "model show" experience, making the figures and clothing on the "model" show appear more realistic. Furthermore, because the actual assets produced are in the form of multi-view videos rather than models, they can be rendered and displayed on a wider range of devices.
[0087] In implementing the "model show" through the aforementioned multi-view video, the following issues remain: As previously mentioned, the "model show" function typically requires providing users with an interactive effect of continuous perspective switching, for example, allowing users to continuously switch perspectives by sliding the screen. However, in the embodiments of the present application, since the specific "digital human" is represented through multi-view video, and the perspectives are discrete, if perspective switching is required during the client presentation, it can only be done between these fixed perspectives, for example, switching from perspective 1 to perspective 2. If the distance between the perspectives is large, the user may experience a noticeable jump. Of course, in theory, if the number of camera devices is large enough and the distribution is dense enough to capture video content from more perspectives, then during playback on the playback end, even switching between discrete perspectives can simulate the interactive effect of continuous perspective switching to a certain extent. However, a larger number of camera devices means higher costs.
[0088] Therefore, to achieve a more cost-effective simulation of continuous perspective switching, embodiments of the present application also employ image-based rendering technology. This technology utilizes images actually captured from two adjacent perspectives to estimate pixel positions at multiple intermediate perspectives and generate images from these intermediate perspectives. In other words, without actually increasing the number of camera devices, image-based rendering technology can be used to supplement images from multiple intermediate perspectives where no camera devices are actually deployed, thereby generating image content from a greater number of perspectives and achieving a better simulation of continuous perspective switching.
[0089] Of course, in the specific implementation, it may also be necessary to provide a variety of different scenes in the "Model Show" function, such as indoor scenes, outdoor scenes, technological scenes, etc. If these scenes are built directly in the real space, the cost may still be relatively high, and the same model character will be required to complete the show again in multiple different scenes and re-shoot, etc. Therefore, the time cost will also be relatively high.
[0090] To this end, in an embodiment of the present application, the generation of a real-life digital human and the creation of a scene can be performed independently. After capturing multi-perspective video content with a camera device, the model character in the video can be "cut out." Subsequent image synthesis of intermediate perspectives can be performed based on the resulting character image. Furthermore, a variety of different 3D virtual space models can be provided. When displayed on the client, the multi-perspective character image video can be placed into the 3D virtual space scene model for display. In this way, the same multi-perspective character image video can be displayed in different 3D virtual space scene models.
[0091] Furthermore, since the embodiments of this application involve multi-view video, and in order to simulate continuous view switching, the number of viewpoints (including real viewpoints and intermediate viewpoints supplemented by the algorithm) may be very large (perhaps one viewpoint every 3 or 5 degrees, for a total of 120 or 72 viewpoints, etc.). Therefore, the compression and transmission of such multi-view video also need to be considered. To address this issue, the embodiments of this application also provide a corresponding solution, which will be described in detail later.
[0092] From a system architecture perspective, as shown in Figure 1, embodiments of the present application may involve an acquisition end, a server, and a client. On the acquisition end, multiple cameras can be used to capture a model's catwalk, generating video content from multiple real-world perspectives. On the server end, processes such as portrait cutouts, intermediate-perspective image synthesis, and multi-perspective video compression can be performed. The compressed multi-perspective video can be transmitted to the client, where it can be sliced to reduce client latency. The client can download the multi-perspective video and render a 3D spatial scene model. A "digital human" (a character image video content from a default perspective) is rendered from the multi-perspective video, and then the "digital human" is added to the 3D spatial scene model. During this process, light and shadow fusion can be performed, for example, to achieve shadow consistency across different perspectives, to enhance the presentation. During the presentation, the server can respond to user perspective switching operations by switching to adjacent video content from other perspectives, simulating the effect of continuous perspective switching. Furthermore, the server can respond to user zooming operations, etc.
[0093] The specific implementation scheme provided in the embodiments of this application is introduced in detail below.
[0094] Example 1
[0095] First, from the perspective of the server, this embodiment 1 provides a method for displaying information based on a dynamic digital human image. Referring to FIG2 , the method may specifically include:
[0096] S201: Obtain video contents from multiple real perspectives, where the video contents from multiple real perspectives are obtained by synchronously shooting a process in which a real person wears a target garment and displays the target garment from different perspectives using multiple camera devices.
[0097] Among them, the so-called real perspective is the perspective at which camera equipment is actually deployed for shooting. In specific implementation, for the "model show" function, first, multiple camera equipment can be arranged 180 / 360 degrees around real people (i.e., real-life model characters) in a real space scene such as a studio to ensure that each perspective can be clearly captured. When the real-life model characters are wearing the specific clothing to be displayed, all camera equipment are used to synchronize the shooting so that the character's movements and postures are captured at the same time. In this way, the specific model characters can walk or make certain display movements in the above-mentioned space scene while wearing the specific clothing to be displayed. During this process, the above-mentioned multiple camera equipment can be synchronized to obtain video content from multiple perspectives. In the embodiment of the present application, the specific number of camera equipment does not need to be too large, for example, it can be on the order of more than ten, and so on.
[0098] S202: Extracting character images contained in the plurality of video frames from the video content.
[0099] After acquiring video content from multiple perspectives, it is necessary to generate a "digital human" in order to facilitate integration with multiple different scene models. Therefore, the character images contained in the video frames from multiple perspectives can be extracted separately. Specifically, since each perspective corresponds to a video content, each video content includes multiple video frames, and each video frame can include a character image, the character image can be extracted from it using the "cutout" technology. In this way, multiple character images can be extracted from each perspective, corresponding to the same character. The "multiple character images" here refer to the character images extracted from the video frames at multiple different time points from each perspective. Since the video frames have time information, the extracted character images will also have corresponding time point information. These character images can be arranged according to the time point to form a character image sequence, and a character image sequence can be obtained from each perspective.
[0100] S203: Using image-based rendering technology, based on pixel offset relationship information between character images of adjacent real perspectives at the same time point, generate character images at corresponding time points for multiple intermediate perspectives between the adjacent real perspectives.
[0101] After obtaining the character image sequence corresponding to each perspective, in order to better simulate the effect of continuous perspective switching, character image sequences from more perspectives can be synthesized by interpolation. Specifically, the perspective synthesis process can be performed between every two adjacent real perspectives. For example, assuming that there are 10 real perspectives, and perspective 1 is adjacent to perspective 2, then character images can be synthesized for multiple intermediate perspectives between perspective 1 and perspective 2. That is, it is estimated what the character image would look like if a camera device was installed at the intermediate perspective. In addition, image synthesis can also be performed for intermediate perspectives between perspectives 2 and 3, between perspective 3 and perspective 4, and other adjacent perspectives. In this way, multiple images at intermediate perspectives can be estimated between every two adjacent real perspectives.
[0102] Among them, the selection of the intermediate perspective can be determined according to actual needs. For example, after testing, if a perspective is set every 3 degrees, the user will not obviously feel the jump phenomenon during the switching between perspectives. Therefore, the interval of the intermediate perspective can be set to 3 degrees; of course, if in order to better simulate continuous perspective switching, the interval of the intermediate perspective can be set smaller, or, if in order to avoid too high a transmission bit rate, the interval of the intermediate perspective can be set larger, and so on.
[0103] After determining the intervals between intermediate perspectives, the pixel positions of each intermediate perspective can be estimated based on the pixel positions of the person images corresponding to the two adjacent perspectives and the pixel offset relationship information, thereby synthesizing the person image at the specific intermediate perspective. Since each perspective contains multiple person images corresponding to different time points, the intermediate perspective image synthesis can also be performed on a time point basis. That is, multiple person images corresponding to each time point are generated for a specific intermediate perspective. In this way, each intermediate perspective can correspond to a sequence of person images.
[0104] In specific implementations, estimating the pixel positions of the intermediate perspective based on person images from two adjacent perspectives can be done in a variety of ways. For example, one approach involves using person images corresponding to adjacent real-world perspectives at the same time as input, and fitting a dense optical flow field between the person images corresponding to the adjacent perspectives using a deep learning model. This dense optical flow field can then be used to estimate the pixel positions of multiple intermediate perspectives between the adjacent real-world perspectives at corresponding time points. Optical flow is the instantaneous velocity of pixel motion of a spatially moving object on the observation imaging plane. The optical flow method uses the temporal changes in pixels in an image sequence and the correlation between adjacent frames to find the correspondence between the previous and current frames, thereby calculating the motion information of the object between adjacent frames. In space, motion can be described using a motion field. On an image plane, the motion of an object is often reflected by the different grayscale distributions of different images in the image sequence. Therefore, the transfer of the spatial motion field to the image is represented as an optical flow field. The optical flow field is a two-dimensional vector field, which reflects the trend of grayscale changes at each point on the image. It can be regarded as the instantaneous velocity field generated by the movement of grayscale pixels on the image plane. The information it contains is the instantaneous motion velocity vector information of each image point. Among them, dense optical flow is an image registration method that performs point-by-point matching on an image or a specified area. It calculates the offset of all points on the image to form a dense optical flow field. Through this dense optical flow field, pixel-level image registration can be performed. Therefore, in an embodiment of the present application, this optical flow field information can be first estimated, and then, based on the optical flow field, the pixel position of the character image of each intermediate perspective at the corresponding time point can be estimated, and then the corresponding character image can be generated for the intermediate perspective.
[0105] S204: Generate a multi-perspective video based on the character images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points, so that the multi-perspective video can be matched to a preset virtual 3D space scene model on the client for rendering and display, thereby providing content of a dynamic digital human image displaying the target clothing in the virtual 3D space scene and providing an interactive effect for simulating continuous perspective switching.
[0106] After obtaining a sequence of character images from multiple intermediate perspectives, the sequences consisting of multiple real perspectives and character images from the multiple intermediate perspectives at multiple time points can be organized into video formats to obtain multi-perspective videos. That is, in an embodiment of the present application, a multi-perspective video can include videos corresponding to multiple real perspectives and videos corresponding to multiple intermediate perspectives. This multi-perspective video can be used to express a "digital human." This "digital human" is generated by photographing a real person, cutting out images, and synthesizing intermediate perspectives, rather than through 3D modeling. Therefore, while maintaining the 3D effect, the "digital human" can also appear more realistic and less cartoon-like, and can also achieve a more realistic and natural display effect for products such as clothing.
[0107] After obtaining the multi-view video through the above method, it can be compressed and encoded so that it can be provided to the client for display. When displayed on the client, the specific multi-view video can be decoded. Because it has been processed by cutting out the image, it can have features such as a transparent background. Therefore, it can be placed in a pre-generated 3D spatial scene model, presenting the visual effect of a specific "digital human" located in the 3D spatial scene model.
[0108] Specifically, when performing compression encoding, since multi-view video is involved, embodiments of the present application may also implement special processing for the compression encoding method. This is because, compared to ordinary video, multi-view video generally has the following characteristics: In terms of resolution, multi-view video may include video data corresponding to dozens or even hundreds of viewpoints. If each viewpoint corresponds to high-definition video, the total video data volume will be enormous. Assuming there are 120 viewpoints, the overall resolution will exceed 32K or even higher, which is a heavy load for most devices. In terms of video bitrate, high resolution means higher video bitrate, which makes real-time transmission and smooth playback more difficult. The bitrate of ordinary 720P video may be 2-5Mbps, but the bitrate of multi-view video can increase by dozens or even hundreds of times. In terms of video size, multi-view video contains data from multiple viewpoints, which leads to a rapid increase in video file size. One hour of multi-view video may require tens or even hundreds of GB of storage space. These technical challenges make the storage, compression, and transmission of multi-view video particularly difficult.
[0109] In the prior art, there are some solutions for compressing and transmitting multi-view videos. For example:
[0110] Method 1 involves simply splicing multi-view videos. Specifically, all images from all viewpoints at the same time point can be combined into a single frame, which is then compressed and transmitted. However, this results in a spliced video with a resolution that is too high, making it difficult to decode and play in real time, and also difficult to transmit.
[0111] Method 2 is to transmit through streaming media. However, on the one hand, the compression efficiency is not ideal. On the other hand, in order to enable the client to switch perspectives, the multi-perspective videos need to be sliced separately before streaming. When the user needs to switch from perspective A to perspective B at a certain moment, the data of perspective B in the corresponding time slice can be pulled for playback; however, the delay of segmented streaming is relatively large, and whether smooth perspective switching is possible also depends on the size of the slice, because the previous slice must be played before switching to the next slice of the next perspective for playback.
[0112] Method three uses a codec designed for multi-view video for video encoding and decoding. This method provides better compression performance, but it increases the complexity of encoding and decoding and requires higher computing power. In addition, since MV-HEVC is a relatively new standard, a dedicated decoder is required to complete decoding. Therefore, not all devices support this format, especially ordinary mobile phones or computers and other terminal devices usually cannot support it. Even if they can support it, there will be freezes and other phenomena when opening and playing the video.
[0113] In response to the above situation, in the embodiments of the present application, corresponding solutions are also provided for the storage and compression of multi-view videos. Specifically, before encoding a specific multi-view video, the video contents corresponding to the multiple viewpoints can be spliced first, that is, the video contents corresponding to different viewpoints at the same time point can be spliced. However, it is not simply splicing all the video contents of all viewpoints into the same frame. Instead, multiple video frames corresponding to multiple viewpoints at the same time point can be grouped to obtain multiple sets, and multiple video frames in the same set can be spliced into the same frame (for ease of distinction, the frame obtained after such splicing can be called a "combined frame"). In other words, the same time point can correspond to multiple different combined frames. In this way, the resolution of the combined frame can be made not too high.
[0114] Specifically, when grouping multi-view video frames, the number of groups required and the number of viewpoints included in each group can be determined based on information such as the number of specific viewpoints, the resolution of individual video frames at each viewpoint, and the maximum resolution supported by the terminal device, so that the resolution of each combined frame is lower than the maximum resolution supported by the terminal device. For example, assuming there are 72 viewpoints, and the resolution of the video frame at each viewpoint is 720P, most terminal devices currently on the market can generally support real-time decoding of 4K resolution. In this case, the 72 viewpoints can be divided into 12 groups, and each combined frame will include video frames corresponding to 6 viewpoints at the same time point. The resolution of each combined frame is 720P×6=4320P, which is close to the resolution of a typical 4K image. Therefore, most terminal devices can achieve real-time decoding of image frames of this resolution.
[0115] After determining the specific number of groups, to achieve higher compression efficiency during the encoding process, the grouping method for the different views and their arrangement within the combined frame can also be determined. Various specific grouping and arrangement methods are possible. For example, in the simplest approach, views 1 through 6 can be grouped together, views 7 through 12 can be grouped together, and so on. Within a combined frame, the views can be divided into 3×2 (three rows and two columns) blocks, and the views can be arranged in numerical order within these blocks. However, given that video frames captured at the same time from different views often have high similarity in content, especially between adjacent views, this high similarity between adjacent views can contain a significant amount of redundant information from an information encoding perspective, which can be compressed during the encoding process. In other words, the presence of redundant information improves compression efficiency. Therefore, fully utilizing this redundant information during the encoding process can significantly improve compression efficiency.
[0116] In the video encoding process, specific information compression techniques can be divided into intra-frame compression and inter-frame compression. Intra-frame compression is performed in the spatial domain (on the X and Y axes), primarily considering the similarity between data within the frame. Inter-frame compression, on the other hand, utilizes redundancy between different frames in a video sequence, such as the similarity between previous and subsequent frames, to reduce the amount of data through prediction. Generally speaking, inter-frame compression achieves higher compression rates than intra-frame compression.
[0117] However, if the simple perspective grouping and arrangement described in the above example is followed, the video frames of adjacent perspectives with the highest redundancy are in the same combined frame. Therefore, when compressing the combined frame, this redundant information can only be used in the intra-frame compression process and cannot be fully utilized in the inter-frame coding process.
[0118] To this end, in an embodiment of the present application, a better perspective grouping and arrangement method is also provided. Specifically, the video frames of adjacent perspectives can be located at the same position of adjacent combined frames. That is, the video frames of adjacent perspectives will be divided into different but adjacent groups, and will be located at the same position of adjacent combined frames.
[0119] For example, assume there are 36 views (the number of views has been reduced for ease of description), divided into 6 groups, with each group containing 6 video frames from each view. That is, every 6 video frames from each view constitute a composite frame. As shown in Figure 3, assume each composite frame consists of 3×2 blocks, each used to hold a video frame from one view, with the positions of each block numbered 0, 1, 2, 3, 4, and 5. Furthermore, assume the 36 views are represented by A1, A2, A3, ..., and A36, respectively. As shown in Figure 3, views A1, A2, A3, A4, A5, and A6 are located at position 0 in composite frames 1 through 6, views A7, A8, A9, A10, A11, and A12 are located at position 1 in composite frames 1 through 6, and so on. That is, perspectives A1, A7, A13, A19, A25, and A31 form the first group, spliced into group frame 1; A2, A8, A14, A20, A26, and A32 form the second group, spliced into group frame 2, and so on. It can be seen that within each group frame, the perspective numbers form an arithmetic progression, and the difference between the perspective numbers is the number of groupings, which is 6 in this example. In this way, the perspectives at the same position between adjacent group frames are also adjacent. This ensures that the image content at the same position between adjacent groups of frames is highly similar. During inter-frame compression coding, the redundant information generated by the high content similarity between adjacent perspectives can be fully utilized, thereby achieving a higher compression rate.
[0120] Of course, the above example only shows the splicing of video frames from various perspectives at one time point. Video frames from various perspectives at other time points can also be grouped and arranged in the same manner. In this way, 6 combined frames can be spliced at each time point. After each time point is spliced in the above manner, the resulting combined frames can be formed into a frame sequence. For example, if each combined frame is represented as "combined frame mn", where m represents the number of the time point and n represents the number of each combined frame corresponding to the same time point, the formed frame sequence can be: (combined frame 11, combined frame 12, combined frame 13, combined frame 14, combined frame 15, combined frame 16, combined frame 21, combined frame 22, combined frame 23, combined frame 24, combined frame 25, combined frame 26, combined frame 31, combined frame 32...).
[0121] Through the above grouping and arrangement, these video frames from adjacent perspectives are dispersed into different but adjacent composite frames and located at the same position within the adjacent frames. Since video frames from adjacent perspectives typically have a high degree of similarity, the composite frames, at least at each position, can have a high degree of similarity. This means that there is a large amount of redundant information, which is the target of compression optimization during inter-frame coding. Therefore, when encoding the composite frames, inter-frame coding can fully utilize the redundant information between video frames from adjacent perspectives to achieve higher compression efficiency. Furthermore, since this high compression rate is achieved through inter-frame coding, and general video encoders have inter-frame coding capabilities, encoding can be achieved using a general video encoder without relying on an encoder dedicated to multi-view coding. Accordingly, decoding can be performed using a general video decoder on the playback end, further supporting decoding and playback on a wider range of terminal devices.
[0122] That is to say, through the embodiment of the present application, a method of grouping and splicing video frames of multiple perspectives is adopted. Compared with the method of directly splicing all perspectives into the same frame, the resolution of each combined frame can be reduced, and the decoding pressure of the terminal device can be reduced; in addition, the arrangement of different perspectives in different combined frames is specially processed so that the content similarity between different combined frames will be relatively high. In this way, the redundant information between video frames of adjacent perspectives can be eliminated or reduced by inter-frame compression of the combined frames, thereby obtaining a higher compression rate, so that a general encoder (for example, HEVC (High Efficiency Video Coding, high-efficiency video coding) etc.) can complete the encoding process. Correspondingly, a general decoder can also be used for decoding, so that a specific video can be decoded and played on most terminal devices, avoiding a complex rendering process, a large computing overhead, and dependence on a dedicated codec.
[0123] After encoding and inter-frame compression of a frame sequence formed by multiple combined frames, it can be transmitted to a client for decoding and display. In an optional implementation, the frame sequence can be sliced before transmission so that the resulting fragments are transmitted. Each fragment can then be independently decoded and played back at the receiving end. This allows the receiving end to decode and play back only the first fragment received, without having to wait for the entire frame sequence to be transmitted, thus reducing latency.
[0124] Among them, the specific slice duration can be determined according to actual needs. If the slice is smaller, the delay at the receiving end will be smaller. For example, the duration of each slice can be 1S, or it can be 0.5S, and so on. Among them, in the embodiment of the present application, since the video frames of multiple perspectives are grouped and spliced, after determining the slice duration, the number of combined frames that need to be included in each slice can be determined according to the playback frame rate of the playback end. For example, still taking 72 perspectives divided into 12 groups as an example, each time point corresponds to 12 combined frames. In addition, assuming that the playback frame rate of the playback end is 30 frames / S, the slice duration when the combined frame is sliced is 1S, then each segment needs to contain 30×72 / 6=360 combined frames. In other words, the combined frames included in each segment must meet the number of frames required to be played by the player within a 1-second duration. Specifically, during playback, the player needs to decode the combined frames, select the video frames corresponding to a specific viewpoint, and play them. Furthermore, the 30 frames played by the player within 1 second are typically 30 video frames from the same viewpoint, and playback can be performed from any specific viewpoint. Therefore, when slicing the combined frames, if each segment is 1 second, it is necessary to ensure that each viewpoint in the same segment contains 30 video frames. If the number of viewpoints is 72, the number of video frames is 30 × 72. Since these video frames are grouped and spliced into combined frames, the number of combined frames is 30 × 72 / 6 = 360. Of course, if the above assumptions remain unchanged, if the slice duration is changed to 0.5 seconds, each segment can contain 180 combined frames, and so on.
[0125] Furthermore, if segmented transmission is used, the compression ratio can be further improved during inter-frame coding by controlling the number of keyframes and bidirectional reference frames. Specifically, the encoder encodes multiple images into segments called GOPs (Group of Pictures). During playback, the decoder reads each GOP segment, decodes it, and renders it for display. A GOP is a group of consecutive pictures, consisting of an I-frame and several B / P-frames. It is the basic unit of video access by the encoder and decoder, and its order is repeated until the end of the video. An I-frame is an intra-coded frame (also known as a keyframe), a P-frame is a forward-predicted frame (forward reference frame), and a B-frame is a bidirectionally interpolated frame (bidirectional reference frame). Specifically, an I-frame is typically a complete picture, while P-frames and B-frames record changes relative to the I-frame. P-frames and B-frames do not contain complete picture data; they only contain the difference between the previous frame and the previous and next frames. B-frames require less information and therefore generally have a higher compression ratio. If a GOP includes fewer I frames and more B frames, the overall compression rate will be higher.
[0126] In practical applications, the specific frames encoded as I-frames, P-frames, B-frames, etc. are typically determined by the encoder based on an algorithm. However, in embodiments of the present application, to further control the video compression rate, the encoder can be manipulated to reduce the number of I-frames and increase the number of B-frames. Specifically, the keyframe interval during inter-frame encoding can be controlled based on the number of composite frames included in each slice to reduce the number of frames encoded as keyframes within the same slice. For example, assuming each slice contains 360 composite frames, the keyframe interval can be set to 360 or 180 frames, ensuring that only one or two frames within the same slice are encoded as I-frames. Furthermore, for composite frames other than keyframes, the number of frames encoded as bidirectional reference frames within the same slice can be increased by lowering the threshold for determining bidirectional reference frames. Specifically, for B-frames, the encoder typically determines whether the current frame can be encoded as a B-frame by calculating the similarity between the current frame and the preceding and following frames and comparing it with a certain threshold. In embodiments of the present application, lowering this threshold allows more frames to be encoded as B-frames, thereby improving the compression rate.
[0127] It should be noted here that since the decoding of P frames and B frames depends on I frames, and the decoding of B frames depends on the previous frame and the next frame, theoretically, if the number of I frames is small and the number of B frames is large, although the compression rate will be better, the image quality may be affected during decoding. However, in the embodiment of the present application, since the multi-perspective video frames are grouped and spliced, and the video frames of adjacent perspectives are located in the same position of the adjacent combined frames, each two adjacent combined frames will have a relatively high similarity. Under this premise, even if the number of I frames and B frames is controlled in the above manner, it will generally not affect the image quality at the decoding end. After testing, the solution provided by the embodiment of the present application has significantly reduced resolution and bit rate compared to the solution of encoding after simple splicing, and PSNR (Peak Signal-to-Noise Ratio, which represents the ratio of the maximum possible signal power to the destructive noise power that affects its representation accuracy, and is one of the indicators for measuring image quality) has been improved, as shown in Table 1:
[0128] Table 1
[0129] Of course, in practical applications, if higher picture quality is required, the number of I frames can be appropriately increased and the number of B frames can be reduced. For example, each slice can include 2 or more I frames, and so on.
[0130] The above provides a detailed introduction to the compression and transmission of multi-view videos. In actual applications, assuming that a user initiates access to a "model show" through a client, the specific multi-view video can be transmitted to the client, which can download it and can be decoded using a general video decoder. Afterwards, the video frames from one of the perspectives (usually one of the perspectives can be set as the default perspective) can be added to a preset 3D spatial scene model. The 3D spatial scene model can be pre-generated by 3D modeling, that is, the client can render the 3D spatial scene model and the multi-view video separately, and add the video content corresponding to one of the perspectives to the 3D spatial scene model for display.
[0131] In specific implementations, to achieve a better integration between the video content and the 3D spatial scene model, the scene perspective and the character perspective can be rendered synchronously, and the perspectives can also be switched synchronously. In addition, the light and shadow integration of the character and the scene can be achieved. For example, lighting information can be added to the 3D spatial scene, and then the character's shadow position and light position can be calculated based on the lighting information. Shadow information and light information can also be added to the 3D spatial scene, so that the character and the scene can be better integrated.
[0132] In summary, through the embodiments of the present application, when it is necessary to provide users with functions such as a "model show", multiple camera devices can be used to synchronously shoot the process of a real person wearing a target garment and displaying the target garment from different perspectives, thereby obtaining video content from multiple real perspectives. Afterwards, the character images contained therein can be extracted from the multiple video frames contained in the video content. Then, using image-based rendering technology, based on the pixel offset relationship information between the character images of adjacent real perspectives at the same time point, character images at corresponding time points can be generated for multiple intermediate perspectives between the adjacent real perspectives. Then, a multi-perspective video can be generated based on the multiple real perspectives and the character images of the multiple intermediate perspectives at multiple time points, so that the client can add the character image video content of one of the perspectives to a preset virtual 3D space scene model for rendering and display, thereby providing content of a dynamic digital human displaying the target garment in the virtual 3D space scene and providing an interactive effect for simulating continuous perspective switching. This approach eliminates the need for explicit modeling of the human figure; instead, a 3D-like digital human effect can be achieved using images captured from multiple perspectives, avoiding complex modeling processes and high modeling costs. This results in a low-cost, high-fidelity, and real-time interactive dynamic digital human effect. Furthermore, because the digital human assets created in this embodiment are in a standard video format like HEVC, they can be parsed by most mobile devices, avoiding complex rendering processes and significant computational overhead.
[0133] In addition, for multiple video contents corresponding to multiple viewpoints, specifically during video encoding, the multiple video frames corresponding to the multiple viewpoints can be spliced together at a time point to obtain a frame sequence formed by multiple combined frames. Furthermore, the multiple video frames corresponding to the multiple viewpoints at that time point are divided into multiple sets, and the multiple video frames in each set are spliced together to form a combined frame, such that video frames from adjacent viewpoints are located at the same position in adjacent combined frames. Subsequently, a general video encoder can be used to encode the frame sequence formed by the multiple combined frames, and inter-frame compression processing can be performed on the multiple combined frames to eliminate or reduce redundant information between video frames from adjacent viewpoints. In this way, since the video frames of multiple perspectives are grouped and spliced, the resolution of the spliced combined frames is not too high, which facilitates real-time decoding in most terminal devices. In addition, since the grouping method and arrangement method are controlled during group splicing, the video frames of adjacent perspectives are located in the same position of the adjacent combined frames, that is, the video frames of adjacent perspectives are located in different but adjacent combined frames, and the positions in different combined frames are the same, and the video frames of adjacent perspectives have a relatively high similarity. Therefore, the adjacent combined frames spliced in this way have a relatively high similarity, and then a general inter-frame compression algorithm can be used to eliminate or reduce the redundant information between the video frames of adjacent perspectives, thereby obtaining a higher compression rate. In other words, in the embodiment of the present application, an ideal compression rate can be obtained by a general video encoder, and accordingly, decoding can be completed by using a general decoder at the decoding end, so that it can be supported on more terminal devices.
[0134] Example 2
[0135] The second embodiment corresponds to the first embodiment and provides a method for displaying information based on a dynamic digital human image from the perspective of a client. Referring to FIG4 , the method may include:
[0136] S401: In response to a viewing request initiated by a user, a multi-perspective video is obtained, wherein the multi-perspective video is generated by: synchronously capturing a process in which a real person, while wearing a target garment, displays the target garment using multiple camera devices from different perspectives to obtain video content from multiple real perspectives; extracting character images contained in multiple video frames included in the video content from the multiple real perspectives; generating character images for multiple intermediate perspectives between adjacent real perspectives using an image-based rendering technique; and generating the multi-perspective video based on the character images from the multiple real perspectives and the multiple intermediate perspectives at multiple time points.
[0137] S402: Decoding the multi-view video;
[0138] S403: Matching the multi-view video to a preset virtual 3D space scene model to provide a dynamic digital human image displaying the target clothing in the virtual 3D space scene;
[0139] S404: In response to the interactive operation of continuous perspective switching, provide an interactive effect simulating continuous perspective switching by switching to character image video content of other perspectives.
[0140] Example 3
[0141] The above embodiments 1 and 2 mainly introduce specific implementation solutions based on the need to provide users with functions such as "model show" in the commodity information service system. In actual applications, the method of producing dynamic real-life digital humans provided in the embodiments of this application can also be used in other application scenarios. To this end, this embodiment 3 also provides a method for generating dynamic digital human images. See Figure 5. This method may include:
[0142] S501: Obtain video contents from multiple real perspectives, where the video contents from multiple real perspectives are obtained by synchronously shooting a process in which a real person performs a target action from different perspectives using multiple camera devices.
[0143] The specific target action can be determined according to the needs of the actual scene, for example, it can be performing a certain dance move, etc.
[0144] S502: Extracting character images contained in the plurality of video frames from the video content.
[0145] S503: Using image-based rendering technology, based on pixel offset relationship information between character images at the same time point from adjacent real perspectives, generate character images at corresponding time points for multiple intermediate perspectives between the adjacent real perspectives.
[0146] S504: Generate a multi-perspective video based on the multiple real perspectives and the character images of the multiple intermediate perspectives at multiple time points, so as to display the corresponding dynamic digital human image through the multi-perspective video.
[0147] The relevant data of the dynamic digital human image produced by the above method can exist in the form of multi-perspective video. Multi-perspective video is composed of files in ordinary video format corresponding to multiple perspectives. Therefore, this real-life digital human is terminal-friendly and can be displayed in various scenarios, including being integrated with a certain 3D spatial scene model, thereby presenting the perspective effect of a specific digital human performing corresponding actions in a specific 3D spatial scene. In addition, users can continuously switch or zoom in and out of the perspective, etc.
[0148] Example 4
[0149] In this fourth embodiment, the method of matching multi-view videos to a virtual 3D space scene model is mainly protected. Specifically, this fourth embodiment provides a method for displaying clothing information based on a dynamic digital human image. Referring to FIG6 , the method may include:
[0150] S601: In response to a request for displaying a target garment through a dynamic digital human, a virtual 3D space scene model and a dynamic digital human image expressed in the form of a multi-view video are obtained, wherein the multi-view video is used to display a process of the target person wearing the target garment and displaying the target garment from multiple perspectives;
[0151] S602: Matching the multi-view video to the virtual 3D space scene model to provide a dynamic digital human image displaying the target clothing in the virtual 3D space scene, and providing an interactive effect for simulating continuous viewpoint switching.
[0152] In this fourth embodiment, the multi-view video can be generated using the methods provided in the aforementioned embodiments, or other methods. For example, multi-view video can be generated directly by densely deploying more cameras. Specifically, when matching the multi-view video to the virtual 3D space scene model, this can be achieved through a 3D rendering engine. Of course, during the specific implementation, some functional customization can be performed on the basis of a general 3D rendering engine, for example, achieving shadow consistency across multiple different perspectives, etc.
[0153] For the contents not described in detail in Examples 2 to 4, please refer to the description in Example 1 and other parts of this specification, which will not be repeated here.
[0154] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described herein within the scope permitted by applicable laws and regulations, subject to the requirements of applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).
[0155] Corresponding to the first embodiment, the embodiment of the present application further provides a device for displaying information based on a dynamic digital human image, which may include:
[0156] a video content obtaining unit, configured to obtain video content from multiple real perspectives, wherein the video content from the multiple real perspectives is obtained by synchronously shooting a process in which a real person, while wearing a target garment, displays the target garment from different perspectives using multiple camera devices;
[0157] A character image extraction unit, configured to extract character images contained in each of the plurality of video frames contained in the video content;
[0158] a perspective synthesis unit, configured to generate character images at corresponding time points for a plurality of intermediate perspectives between adjacent real perspectives using an image-based rendering technique based on pixel offset relationship information between character images at the same time point from adjacent real perspectives;
[0159] A multi-perspective video storage and unit is used to generate a multi-perspective video based on the character images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points, so that the multi-perspective video can be matched to a preset virtual 3D space scene model on the client to provide content of a dynamic digital human image displaying the target clothing in the virtual 3D space scene, and provide an interactive effect for simulating continuous perspective switching.
[0160] The video contents of the multiple real perspectives are obtained by synchronously shooting a process in which a real person wearing target clothing walks in a target space to display the target clothing through multiple camera devices from different perspectives.
[0161] Specifically, the perspective synthesis unit can be used to:
[0162] Using image-based rendering technology, according to the pixel offset relationship information between the character images of adjacent real perspectives at the same time point, the pixel positions of multiple intermediate perspectives between the adjacent real perspectives at corresponding time points are estimated to generate the character images of the intermediate perspectives at multiple time points.
[0163] More specifically, the perspective synthesis unit may be used to:
[0164] Taking the character images corresponding to adjacent real perspectives at the same time point as input, a dense optical flow field between the character images of adjacent real perspectives is fitted through a deep learning model, and the dense optical flow field is used to estimate the pixel positions of the character images of multiple intermediate perspectives between the adjacent real perspectives at corresponding time points.
[0165] In addition, the device may further include:
[0166] The video compression unit is used to compress the multi-view video for transmission to the client.
[0167] The video compression unit may specifically include:
[0168] a splicing subunit, configured to splice video frames corresponding to multiple viewpoints in units of time points to obtain a frame sequence formed by multiple combined frames; wherein, for a same time point, the multiple video frames corresponding to the multiple viewpoints at that time point are divided into multiple sets, and the multiple video frames in each set are spliced into a combined frame, such that video frames of adjacent viewpoints are located at the same position in adjacent combined frames;
[0169] The inter-frame compression subunit is configured to encode the frame sequence formed by the plurality of combined frames using a universal video encoder, and perform inter-frame compression processing on the plurality of combined frames.
[0170] The resolution of each combined frame is lower than the maximum resolution supported by the terminal device.
[0171] In addition, the device may further include:
[0172] The slicing processing unit is used to slice the frame sequence formed by multiple combined frames after encoding and inter-frame compression processing, so that the frame sequence can be transmitted in units of fragments obtained after slicing, and independently decoded and played in units of fragments at the receiving end.
[0173] Furthermore, it may also include:
[0174] The key frame number control unit is used to control the key frame interval in the inter-frame coding process according to the number of combined frames included in each slice, so as to reduce the number of frames encoded as key frames in the same slice.
[0175] The bidirectional reference frame quantity control unit is used to increase the number of frames encoded as bidirectional reference frames in the same slice by lowering the judgment threshold of the bidirectional reference frames for the combined frames other than the key frames.
[0176] Corresponding to the second embodiment, the embodiment of the present application further provides a device for displaying information based on a dynamic digital human image, which may include:
[0177] a multi-perspective video acquisition unit, configured to acquire a multi-perspective video in response to a viewing request initiated by a user, the multi-perspective video being generated by: synchronously capturing, from different perspectives, a process of a real person wearing a target garment and displaying the target garment, using multiple camera devices, to obtain video content from multiple real perspectives; extracting character images contained in multiple video frames included in the video content from the multiple real perspectives; generating character images for multiple intermediate perspectives between adjacent real perspectives using an image-based rendering technique; and generating the multi-perspective video based on the character images from the multiple real perspectives and the multiple intermediate perspectives at multiple time points;
[0178] A decoding unit, configured to decode the multi-view video;
[0179] An adding unit is used to match the multi-view video to a preset virtual 3D space scene model to provide a dynamic digital human image displaying the target clothing in the virtual 3D space scene;
[0180] The perspective switching interaction unit is used to respond to the interactive operation of continuous perspective switching and provide an interactive effect simulating continuous perspective switching by switching to character image video content of other perspectives.
[0181] Corresponding to the third embodiment, the present embodiment further provides a device for generating a dynamic digital human image, which may include:
[0182] a video content obtaining unit, configured to obtain video content from multiple real perspectives, wherein the video content from the multiple real perspectives is obtained by synchronously shooting a process in which a real person performs a target action from different perspectives using multiple camera devices;
[0183] A character image extraction unit, configured to extract character images contained in each of the plurality of video frames contained in the video content;
[0184] a perspective synthesis unit, configured to generate character images at corresponding time points for a plurality of intermediate perspectives between adjacent real perspectives using an image-based rendering technique based on pixel offset relationship information between character images at the same time point from adjacent real perspectives;
[0185] The dynamic digital human asset generation unit is used to generate a multi-perspective video based on the character images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points, so as to display the corresponding dynamic digital human image through the multi-perspective video.
[0186] Corresponding to the fourth embodiment, the embodiment of the present application further provides a device for displaying clothing information based on a dynamic digital human image, which may include:
[0187] a request receiving unit, configured to, in response to a request for displaying a target garment through a dynamic digital human, obtain a virtual 3D space scene model and a dynamic digital human image expressed in the form of a multi-view video, wherein the multi-view video is configured to display a process of the target person wearing the target garment displaying the target garment from multiple viewpoints;
[0188] A display unit is used to match the multi-view video to the virtual 3D space scene model to provide content in which a dynamic digital human image displays the target clothing in the virtual 3D space scene, and to provide an interactive effect for simulating continuous perspective switching.
[0189] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0190] And an electronic device comprising:
[0191] one or more processors; and
[0192] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of any one of the method embodiments described above.
[0193] 7 exemplarily shows the architecture of an electronic device, which may include a processor 710, a video display adapter 711, a disk drive 712, an input / output interface 713, a network interface 714, and a memory 720. The processor 710, the video display adapter 711, the disk drive 712, the input / output interface 713, the network interface 714, and the memory 720 may be communicatively connected via a communication bus 730.
[0194] Among them, the processor 710 can be implemented by a general-purpose CPU (Central Processing Unit, processor), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solutions provided in this application.
[0195] The memory 720 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 720 can store an operating system 721 for controlling the operation of the electronic device 700, and a basic input and output system (BIOS) for controlling the low-level operations of the electronic device 700. In addition, a web browser 723, a data storage management system 724, and an information display processing system 725, etc. can also be stored. The above-mentioned information display processing system 725 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 720 and is called and executed by the processor 710.
[0196] The input / output interface 713 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0197] The network interface 714 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0198] The bus 730 comprises a pathway for transmitting information between the various components of the device (eg, the processor 710 , the video display adapter 711 , the disk drive 712 , the input / output interface 713 , the network interface 714 , and the memory 720 ).
[0199] It should be noted that although the above device only shows the processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, memory 720, bus 730, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.
[0200] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0201] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0202] The above describes in detail the method and electronic device for displaying information based on a dynamic digital human image provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is intended only to help understand the method and core concept of this application. At the same time, those skilled in the art will appreciate that the specific implementation methods and scope of application may vary based on the concept of this application. In summary, the contents of this specification should not be construed as limiting this application.
Claims
1. A method for information display based on a dynamic digital human image, characterized in that, Including: Obtaining video content of multiple real perspectives, where the video content of the multiple real perspectives is obtained by synchronously shooting, from different perspectives, the process of a real person showing a target clothing while wearing the target clothing by multiple camera devices; Respectively extracting the human images included in each of the multiple video frames included in the video content; Using image-based rendering technology, based on the pixel offset relationship information between the human images of adjacent real perspectives at the same time point, generating human images at the corresponding time point for multiple intermediate perspectives between the adjacent real perspectives; Generating a multi-perspective video based on the human images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points, so as to match the multi-perspective video to a preset virtual 3D space scene model on a client, to provide content for a dynamic digital human image to show the target clothing in the virtual 3D space scene, and to provide an interactive effect for simulating continuous perspective switching.
2. The method according to claim 1, wherein: The video content of the multiple real perspectives is obtained by synchronously shooting, from different perspectives, the process of a real person walking in a target space venue while wearing the target clothing to show the target clothing by multiple camera devices.
3. The method according to claim 1, wherein: Generating the human images at the corresponding time point for multiple intermediate perspectives between the adjacent real perspectives includes: Using image-based rendering technology, based on the pixel offset relationship information between the human images of adjacent real perspectives at the same time point, estimating the pixel positions of the multiple intermediate perspectives between the adjacent real perspectives at the corresponding time point, so as to generate the human images of the intermediate perspectives at multiple time points.
4. The method according to claim 3, wherein: Estimating the pixel positions of the multiple intermediate perspectives between the adjacent real perspectives at the corresponding time point includes: Taking the human images respectively corresponding to adjacent real perspectives at the same time point as inputs, fitting a dense optical flow field between the human images of the adjacent real perspectives through a deep learning model, and using the dense optical flow field to estimate the pixel positions of the human images of the multiple intermediate perspectives between the adjacent real perspectives at the corresponding time point.
5. The method according to claim 1, wherein It further includes: Performing compression processing on the multi-perspective video for transmission to a client.
6. The method according to claim 5, wherein: Performing compression processing on the multi-perspective video includes: Performing splicing processing on the video frames corresponding to multiple perspectives in units of time points to obtain a frame sequence formed by multiple combined frames; wherein, the multiple video frames corresponding to multiple perspectives at this time point are divided into multiple sets, and the multiple video frames in each set are spliced into a combined frame, and the video frames of adjacent perspectives are located at the same position in adjacent combined frames; Encoding the frame sequence formed by the multiple combined frames using a general video encoder, and performing inter-frame compression processing on the multiple combined frames.
7. The method according to claim 6, wherein: The resolution of each combined frame is lower than the maximum resolution supported by the terminal device.
8. The method according to claim 6, characterized in that, It further includes: After encoding and inter-frame compression processing of the frame sequence formed by multiple combined frames, the frame sequence is also sliced so as to be transmitted in units of the obtained segments, and independently decoded and played in units of segments at the receiving end.
9. A method for information display based on a dynamic digital human image, characterized in that, It includes: In response to a viewing request initiated by a user, a multi-view video is obtained. The multi-view video is generated in the following manner: multiple video contents of real perspectives are obtained by synchronously shooting the process of a real person showing the target clothing in the state of wearing the target clothing from different perspectives by multiple camera devices. Character images included in the multiple video frames included in the multiple video contents of real perspectives are respectively extracted. Using image-based rendering technology, character images are generated for multiple intermediate perspectives between adjacent real perspectives, and the multi-view video is generated according to the character images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points. Decode the multi-view video; Match the multi-view video to a preset virtual 3D space scene model to provide content for a dynamic digital human image to show the target clothing in the virtual 3D space scene. In response to an interactive operation of continuous perspective switching, provide an interactive effect of simulating continuous perspective switching by switching to the video content of character images of other perspectives.
10. A method for generating a dynamic digital human image, characterized in that, It includes: Obtain video contents of multiple real perspectives, where the video contents of the multiple real perspectives are obtained by synchronously shooting the process of a real person performing a target action from different perspectives by multiple camera devices; Respectively extract the character images included in the multiple video frames included in the video content; Using image-based rendering technology, according to the pixel offset relationship information between the character images of adjacent real perspectives at the same time point, generate character images of corresponding time points for multiple intermediate perspectives between the adjacent real perspectives; Generate a multi-view video according to the character images of the multiple real perspectives and the multiple intermediate perspectives at multiple time points, so as to display the corresponding dynamic digital human image through the multi-view video.
11. A method for displaying clothing information based on a dynamic digital human image, characterized in that, It includes: In response to a request to show the target clothing through a dynamic digital human, obtain a virtual 3D space scene model and a dynamic digital human image expressed in the form of a multi-view video, where the multi-view video is used to show the process of a target person showing the target clothing in the state of wearing the target clothing from multiple perspectives. Match the multi-view video to the virtual 3D space scene model to provide content for a dynamic digital human image to show the target clothing in the virtual 3D space scene, and provide an interactive effect for simulating continuous perspective switching.
12. A computer-readable storage medium, having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 11.
13. An electronic device, characterized in that, It includes: One or more processors; And A memory associated with the one or more processors, the memory being configured to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
AR virtual live broadcast method and system
CN114663633A
Multi-user free view angle video method and system based on real-time virtual view angle interpolation
CN114897681A
Multi-view synchronization method and free view system
CN117221627A
Method for displaying information based on dynamic digital human image and electronic equipment
CN117596373A
Method and apparatus for encoding and decoding multi-view video to provide uniform picture quality
US20070211796A1
Cited By
Digital human enhanced rendering method and system based on low-resolution video
CN121458812A