A method and system for generating human binocular image and disparity map
Patent Information
- Application Number
- CN202411861499.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2026-10-09
AI Technical Summary
[0004]然而,建立一个根据文本生成适应于立体视觉的双目图像和视差图仍然是一个极具挑战性的问题
[0060]本发明提出的人体双目图像与视差图生成方法,通过提取的待生成图像对应的文本语义特征和相机位姿,将提取的文本语言特征作为条件输入扩散模型,扩散模型通过去噪过程逐步生成人体参考图像,生成的该人体参考图像不仅符合文本语义,还包含特定的相机位姿信息。通过对神经辐射场的位置信息、密度信息和颜色信息进行训练,使得构建的神经辐射场具有几何位置信息,通过输入给定相机的位置信息和光线方向,可以模拟图像在三维空间中不同位置的密度和颜色,从而生成新视角下的颜色图像。将生成的新视角下的颜色图像作为左相机图像,沿水平方向平移固定距离,生成右相机图像,通过模拟双目相机的视角差异,结合双目图像和视差图之间的几何约束,利用视差和深度之间的转换关系,计算并生成新视角下的符合文本语义的高保真新视角人体双目图像和视差图,以反映左、右视角在深度上的差异。
Smart Images

Figure CN122888084A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital human integration technology, and more specifically to a method and system for generating human binocular images and disparity maps. Background Technology
[0002] Stereo vision, a key technology in the creation of virtual digital humans, has always been a challenging and important research topic in the field of computer vision. However, unlike ordinary scene or object images, human image data involves privacy and cannot be used arbitrarily for training and testing. Therefore, text-driven human image content generation technology has become a necessary approach to solve this problem.
[0003] Numerous renowned research institutions both domestically and internationally are increasingly focusing on multimodal text-to-image generation models. OpenAI embeds text and images into a shared semantic space, utilizing self-supervised contrastive learning to calculate the similarity between text and image vectors, thereby achieving cross-modal understanding. The University of Munich and Heidelberg University in Germany jointly launched the StableDiffusion model, which learns compressed representations of data through variational autoencoders and inputs text embeddings as conditional information into the denoising process of a diffusion model, achieving image generation based on text descriptions. A collaborative project between Peking University and Tencent aligned multimodal data such as images, videos, infrared images, depth maps, and speech to a unified text feature domain, generating training data for modality alignment. Therefore, it is evident that researchers both domestically and internationally have a certain theoretical and practical foundation for constructing image-to-text multimodal alignment and generation methods using deep learning.
[0004] However, generating stereoscopic images and disparity maps from text remains a highly challenging problem. While existing text-to-image models can generate images that conform to the semantics of the text, they cannot accurately understand camera pose information, making it difficult to generate images that match a specific camera pose. When generating disparity maps, relying solely on text cues cannot produce disparity maps that are geometrically consistent with stereo images. Summary of the Invention
[0005] To address the problems existing in the above-mentioned fields, this invention proposes a method and system for generating human binocular images and disparity maps, which can generate high-fidelity new-view human binocular images and disparity maps that conform to textual semantics, so as to reflect the difference in depth between the left and right perspectives.
[0006] To address the aforementioned technical problems, this invention discloses a method for generating human binocular images and disparity maps, comprising the following steps:
[0007] Extract the textual semantic features of the natural language text corresponding to the image to be generated;
[0008] Based on the extracted text semantic features and preset camera pose parameters, a human reference image with a specific camera pose that matches the text is generated through a diffusion model.
[0009] A neural radiation field is constructed by training the position, density, and color information of the neural radiation field by inputting a reference image with a given camera pose, and obtaining the trained neural radiation field. The generated human reference image with a specific camera pose is input into the trained neural radiation field, and the image is rendered under a new perspective to generate a color image rendered under the new perspective.
[0010] The color image rendered from the new perspective is used as the left camera image. The left camera image is then rendered after being horizontally shifted by a preset baseline distance to obtain the right camera image.
[0011] Based on the generated left and right camera images, the depth information from the intersection of rays from the left and right cameras to the surface of the three-dimensional object is determined. Then, the depth map is converted into a disparity map using the conversion relationship between disparity and depth, thus generating a disparity map corresponding to the binocular images.
[0012] Preferably, the step of generating a human reference image with a specific camera pose that matches the text using a diffusion model specifically includes:
[0013] Human images are generated using a diffusion model based on the semantic features Γ(P) of the generated natural language text and the preset camera pose.
[0014] The diffusion model is a probabilistic generative model that learns the data distribution by progressively denoising variables sampled from a Gaussian distribution.
[0015] Determine the initial noise map ∈ ~N(0,I) and the condition vector. Through image-to-text diffusion model Generate target image
[0016] Here, the conditional vector v consists of the textual semantic features Γ(P) and the preset camera pose. constitute;
[0017] diffusion model Training is performed on the following denoising targets:
[0018]
[0019] Where x' is the ground truth value of the image, v is the conditional vector, and α t σ t and ω t It is a function that controls sample quality with respect to the diffusion process time t~U([0,1]);
[0020] By using text semantic features and preset camera poses as conditional vectors to guide the diffusion model, a human reference image I with a specific camera pose that matches the text is generated. ref .
[0021] Preferably, acquiring the trained neural radiation field specifically includes:
[0022] By inputting a reference image with a given camera pose, the position, density, and color information of the neural radiation field are trained to construct a neural radiation field with geometric position information.
[0023] Density and color are predicted using two shallow multilayer perceptrons, respectively.
[0024] For the t-th frame image, given the position coordinates X = (x, y, z) of the reference image and the ray direction D = (δ, φ), the density model F σ Predicted density σ t (x) is:
[0025] σ t (x)=F σ (X,D)
[0026] Position and light are encoded, and the position coordinates and light direction are fed into the color model F. C Predict color c t (x) is:
[0027] c t (x)=F σ (γ x (X),γ d (D))
[0028] Where, γ x and γ d These represent the encoding functions for position and light, respectively.
[0029] Preferably, generating the color image rendered from the new perspective specifically includes:
[0030] According to density σ t and color c t Generate the rendered color image under the camera ray r(t) = o + tD:
[0031]
[0032]
[0033] Where, δ k =||x k+1 -x k ||2 represents the distance between adjacent sampling points, N kThis indicates the number of sampling points.
[0034] Preferably, generating the disparity map corresponding to the binocular images specifically includes:
[0035] Based on the generated left and right camera images, a depth map is obtained from the corresponding new perspective using the principle of light accumulation.
[0036] Using the principle of triangulation, the depth map is converted into a disparity map d based on the preset camera pose parameters baseline and focal length;
[0037] Through warp loss function Epipolar constraints are applied to the generated binocular images and disparity maps:
[0038]
[0039] Among them, F w (·) represents the anti-warping function, and H and W represent the height and width of the image, respectively;
[0040] The generated binocular images and their corresponding disparity maps are used to construct a new dataset, denoted as (I). l ,I r ,d)∈D, where the set D represents the set containing the generated left camera image, right camera image, and disparity map, I l and I r d represents the left and right views, and d represents the disparity map corresponding to the left view.
[0041] Preferably, it further includes:
[0042] Based on the pose of the reference camera, a new reference image I' is generated using neural radiation fields. ref ;
[0043] Construct an L1 loss function to measure the absolute difference between two images at the pixel level and a structural similarity loss function to measure the overall structural similarity of images;
[0044] The generated reference image is optimized at both the pixel and structural levels by minimizing the L1 loss function and the structural similarity loss function.
[0045] Preferably, the expression for the L1 loss function is:
[0046]
[0047] Among them, I ref (i) represents the reference image generated by the diffusion model, I' ref (i) represents the new reference image generated by the neural radiation field, where N is the total number of pixels in the image.
[0048] Preferably, the structural similarity loss function SSIM is calculated as follows:
[0049]
[0050] Where, μ I ,μ I' Image I ref , I' ref The average value reflecting brightness information; Image I ref , I' ref The variance, σ, reflects contrast information. II' For image I ref , I' ref The covariance reflects the structural similarity between two images; C1 and C2 are both stabilizing terms used to prevent the denominator from being zero.
[0051] Preferably, the extraction of textual semantic features from natural language text includes the following steps:
[0052] Input the natural language text corresponding to the image to be generated into the Transformer encoder;
[0053] The Transformer encoder uses self-attention and cross-attention mechanisms to encode natural language text P using a pre-trained and frozen Transformer language model, generating the textual semantic features Γ(P) of the natural language text.
[0054] Preferably, it further includes a human binocular image and disparity map generation system, comprising:
[0055] The text feature extraction module is used to extract the text semantic features of the natural language text corresponding to the image to be generated;
[0056] The reference image generation module is used to generate a human reference image with a specific camera pose that matches the text, based on the extracted text semantic features and preset camera pose parameters, through a diffusion model.
[0057] The color rendering module is used to construct a neural radiation field. By inputting a reference image with a given camera pose, the position, density, and color information of the neural radiation field are trained to obtain the trained neural radiation field. The generated human reference image with a specific camera pose is input into the trained neural radiation field, and the image is rendered from a new perspective to generate a color image rendered from the new perspective.
[0058] The disparity map generation module is used to take the color image rendered from the new perspective as the left camera image, and render the left camera image after horizontally shifting it by a preset baseline distance to obtain the right camera image. Based on the generated left and right camera images, by determining the depth information from the intersection of the rays from the left and right cameras to the surface of the three-dimensional object, the depth map is converted into a disparity map using the conversion relationship between disparity and depth, thus generating a disparity map corresponding to the binocular image.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] This invention proposes a method for generating human binocular images and disparity maps. It extracts the semantic features of the text corresponding to the image to be generated and the camera pose. The extracted textual features are used as conditional inputs to a diffusion model. The diffusion model gradually generates a human reference image through a denoising process. This generated human reference image not only conforms to the text semantics but also contains specific camera pose information. By training the neural radiation field with its position, density, and color information, the constructed neural radiation field possesses geometric position information. By inputting the given camera position information and light direction, the density and color of the image at different locations in three-dimensional space can be simulated, thereby generating a color image from a new perspective. The generated color image from the new perspective is used as the left camera image and translated a fixed distance horizontally to generate the right camera image. By simulating the perspective difference of the binocular cameras, combining the geometric constraints between the binocular image and the disparity map, and utilizing the transformation relationship between disparity and depth, a high-fidelity new perspective human binocular image and disparity map conforming to the text semantics are calculated and generated to reflect the depth difference between the left and right perspectives. Attached Figure Description
[0061] Figure 1 This is a flowchart of the method for generating human binocular images and disparity maps proposed in this invention;
[0062] Figure 2 A block diagram of a human data engine for text-driven neural rendering provided in an embodiment of the present invention;
[0063] Figure 3 A block diagram of a text-to-image diffusion model for text-driven neural rendering provided in an embodiment of the present invention. Detailed Implementation
[0064] The following will refer to the appendices in the embodiments of the present invention. Figures 1-3 The technical solutions in the embodiments of the present invention will be clearly and completely described. It should be understood that the terminology used in the present invention is only for describing particular implementation methods and is not intended to limit the present invention.
[0065] like Figure 1As shown, this invention proposes a method for generating human binocular images and disparity maps, including the following steps:
[0066] Extracting semantic features from natural language text;
[0067] Based on the extracted text semantic features and preset camera pose parameters, a human reference image with a specific camera pose that matches the text is generated through a diffusion model.
[0068] Based on human reference images, a neural radiation field is constructed to train the position, density, and color information of the 3D reconstruction model. Based on the trained 3D reconstruction model, the camera position information is input to render the model from a new perspective, generating a color image rendered from the new perspective.
[0069] The color image rendered from the new perspective is used as the left camera image. The left camera image is then rendered after being horizontally shifted by a preset baseline distance to obtain the right camera image.
[0070] Based on the generated left and right camera images, the depth information from the intersection of rays from the left and right cameras to the surface of the three-dimensional object is determined. Then, the depth map is converted into a disparity map using the conversion relationship between disparity and depth, thus generating a disparity map corresponding to the binocular images.
[0071] Specifically, extracting semantic features from natural language text includes the following steps:
[0072] Input natural language text into the Transformer encoder;
[0073] The Transformer encoder uses self-attention and cross-attention mechanisms to encode natural language text P using a pre-trained and frozen Transformer language model, generating the textual semantic features Γ(P) of the natural language text.
[0074] Generate human reference images with specific camera poses that match the text using a diffusion model, specifically including:
[0075] Human images are generated using a diffusion model based on the semantic features Γ(P) of the generated natural language text and the preset camera pose.
[0076] The diffusion model is a probabilistic generative model that learns the data distribution by progressively denoising variables sampled from a Gaussian distribution.
[0077] Determine the initial noise map ∈ ~N(0,I) and the condition vector. Through image-to-text diffusion model Generate target image
[0078] Here, the conditional vector v consists of the textual semantic features Γ(P) and the preset camera pose. constitute;
[0079] diffusion model Training is performed on the following denoising targets:
[0080]
[0081] Where x' is the ground truth value of the image, v is the conditional vector, and α t σ t and ω t It is a function that controls sample quality with respect to the diffusion process time t~U([0,1]);
[0082] By using text semantic features and preset camera poses as conditional vectors to guide the diffusion model, a human reference image I with a specific camera pose that matches the text is generated. ref .
[0083] The 3D reconstruction model is trained using its position, density, and color information, specifically including:
[0084] Based on the generated reference image I ref Training a three-dimensional reconstruction model based on neural radiation fields;
[0085] Density and color are predicted using two shallow multilayer perceptrons, respectively.
[0086] For the t-th frame image, given the position coordinates X = (x, y, z) of the reference image and the ray direction D = (δ, φ), the density model F σ Predicted density σ t (x) is:
[0087] σ t (x=F σ (X,D)
[0088] Position and light are encoded, and the position coordinates and light direction are fed into the color model F. C Predict color c t (x) is:
[0089] c t (x)=F σ (γ x (X),γ d (D))
[0090] Where, γ x and γ d These represent the encoding functions for position and light, respectively.
[0091] Generating a color image rendered from a new perspective, specifically including:
[0092] According to density σ t and color c t Generate the rendered color image under the camera ray r(t) = o + tD:
[0093]
[0094]
[0095] Where, δ k =||x k+1 -x k ||2 represents the distance between adjacent sampling points, N k This indicates the number of sampling points.
[0096] Generating multi-view binocular images and their corresponding disparity maps specifically includes:
[0097] Based on the generated left and right camera images, a depth map is obtained from the corresponding new perspective using the principle of light accumulation.
[0098] Using the principle of triangulation, the depth map is converted into a disparity map d based on the preset camera pose parameters baseline and focal length;
[0099] Through warp loss function Epipolar constraints are applied to the generated binocular images and disparity maps:
[0100]
[0101] Among them, F w (·) represents the anti-warping function, and H and W represent the height and width of the image, respectively;
[0102] The generated binocular images and their corresponding disparity maps are used to construct a new dataset, denoted as (I). l ,I r ,d)∈D, where the set D represents the set containing the generated left camera image, right camera image, and disparity map, I l and I r d represents the left and right views, and d represents the disparity map corresponding to the left view.
[0103] Furthermore, to ensure the effective generation of the new perspective image, this invention also generates a new reference image I' based on the reference camera pose using a neural radiation field. ref The reference image I generated by the diffusion model is analyzed using the L1 loss function and the structural similarity loss function. ref Reference image I' generated by neural radiation field ref The corresponding parameters are optimized to improve the reference image I generated by the diffusion model. refReference image I' generated by neural radiation field ref Maintain consistency in texture and structure.
[0104] The L1 loss function measures the absolute difference between two images at the pixel level, and its expression is:
[0105]
[0106] Among them, I ref (i) represents the reference image generated by the diffusion model, I' ref (i) represents the new reference image generated by the neural radiation field, where N is the total number of pixels in the image.
[0107] The L1 loss function directly reflects the overall difference between two images by comparing their brightness values pixel by pixel. By minimizing the L1 loss, it is ensured that the reference image generated by the diffusion model and the neural radiation field maintains consistency in brightness and texture detail.
[0108] Structural similarity loss function (SSIM) is a metric used to measure the overall structural similarity of images, and it can better capture the image quality perceived by the human eye.
[0109] The formula for calculating the Structural Similarity Loss Function (SSIM) value is as follows:
[0110]
[0111] Where, μ I ,μ I' Image I ref , I' ref The mean (reflecting brightness information), These represent the image variance (reflecting contrast information) and σ, respectively. II' Let C1 and C2 be the covariance of the two images (reflecting structural similarity), and C1 and C2 be stable terms used to prevent the denominator from being zero.
[0112] The Structural Similarity Loss Function (SSIM) focuses on the similarity of images across three levels: brightness, contrast, and structure, which is more consistent with human visual perception. By optimizing the SSIM, the structural consistency and sensory quality of the generated images can be improved. By minimizing the structural similarity loss, it ensures that the reference images generated by the diffusion model and the neural radiation field are as consistent as possible in overall texture and detailed structure.
[0113] By combining the L1 loss function and the Structural Similarity Loss Function (SSIM), the generated images can be optimized at both the pixel and structural levels, ensuring consistency in texture details and overall perception, thereby achieving high-quality image generation results.
[0114] This invention also proposes a system for generating human binocular images and disparity maps, comprising:
[0115] The text feature extraction module is used to extract the text semantic features of the natural language text corresponding to the image to be generated;
[0116] The reference image generation module is used to generate a human reference image with a specific camera pose that matches the text, based on the extracted text semantic features and preset camera pose parameters, through a diffusion model.
[0117] The color rendering module is used to construct a neural radiation field. By inputting a reference image with a given camera pose, the position, density, and color information of the neural radiation field are trained to obtain the trained neural radiation field. The generated human reference image with a specific camera pose is input into the trained neural radiation field, and the image is rendered from a new perspective to generate a color image rendered from the new perspective.
[0118] The disparity map generation module is used to take the color image rendered from the new perspective as the left camera image, and render the left camera image after horizontally shifting it by a preset baseline distance to obtain the right camera image. Based on the generated left and right camera images, by determining the depth information from the intersection of the rays from the left and right cameras to the surface of the three-dimensional object, the depth map is converted into a disparity map using the conversion relationship between disparity and depth, thus generating a disparity map corresponding to the binocular image.
[0119] The method proposed in this invention can establish a unified semantic representation of text and images, and generate new perspective human binocular images and disparity maps by combining the geometric constraints of binocular images and disparity maps through a multimodal diffusion model and neural radiation field.
[0120] Example
[0121] like Figure 2 As shown, the text-driven method for generating human binocular images and disparity maps provided in this embodiment of the invention specifically includes the following steps:
[0122] Reference image generation based on a text-to-image diffusion model
[0123] This invention extracts text semantics through a text encoder and converts it into a digital representation for encapsulation; based on text semantic information and camera pose, it uses a diffusion-based image generation model to create an image consistent with the text semantics.
[0124] Therefore, the text-to-image diffusion model can be used to generate reference images that conform to the semantics of the text, which can then be used for the generation of subsequent multi-view binocular images and their disparity maps.
[0125] Specifically, the text is input into a Transformer encoder to extract semantic information containing contextual information, and then converted into a numerical representation for encapsulation. In natural language processing, the positional relationships between words are crucial, as different positions can lead to significant semantic differences. The Transformer model utilizes self-attention and cross-attention mechanisms to achieve more accurate and in-depth language understanding. Therefore, a pre-trained and frozen Transformer language model Γ is used to encode the text P to generate semantic features Γ(P).
[0126] like Figure 3 As shown, the input is a piece of natural language text, such as "A image of the face of a Chinese man," and the semantic features of the text are extracted by the Transformer text encoder. In the reverse generation process of the diffusion model, a clear image is gradually recovered from a random noisy image, with each step using conditional probability... The calculation takes text semantic features and camera pose parameters Γ(P) as inputs and generates a human image that matches the text semantics and conforms to the specified camera pose.
[0127] Human images are generated using a diffusion model based on text semantics and camera pose.
[0128] The diffusion model is a probabilistic generative model that learns the data distribution by progressively denoising variables sampled from a Gaussian distribution. Given an initial noise map ∈ ~N(0,I) and a conditional vector... Image-to-text diffusion model It can generate target images The conditional vector v consists of the textual semantic features Γ(P) and the camera pose. constitute.
[0129] diffusion model Train on the following denoising targets:
[0130]
[0131] Where x' is the ground truth value of the image, v is the conditional vector, and α t σ t and ω t It is a function that controls sample quality with respect to the diffusion process time t~U([0,1]).
[0132] By using text semantic features and camera pose as conditional vectors to guide the diffusion model, a human reference image I with a specific camera pose that matches the text can be generated. ref .
[0133] Generation of multi-view image data based on neural radiation fields
[0134] Neural radiation fields are an advanced perspective synthesis technique that can generate lighting conditions on the surface of an object from any perspective and position.
[0135] This invention generates a reference image using a diffusion model, and then constructs a neural radiation field with geometric position information by inputting a reference image with a given camera pose. This allows the generation of a color map and a depth map at the corresponding viewpoint based on the new viewpoint camera pose.
[0136] Meanwhile, the right camera in the binocular image can be considered as the left camera horizontally translated by b units. Therefore, multi-view binocular images and their corresponding disparity maps can be generated using neural radiation fields, which can be used for subsequent stereo matching and 3D reconstruction research.
[0137] Specifically, the reference image I generated using the diffusion model ref A 3D reconstruction model based on neural radiation fields was trained. To optimize rendered image quality and convergence speed, two shallow multilayer perceptrons were used to predict density and color, respectively.
[0138] For the t-th frame image, given the position coordinates X = (x, y, z) of the reference image and the ray direction D = (δ, φ), the density model F σ Predicted density σ t (x):
[0139] σ t (x)=F σ (X,D)
[0140] To better learn high-frequency functions, position and light are encoded, and then the position coordinates and light direction are fed into the color model F. C Predict color c t (x):
[0141] c t (x)=F σ (γ x (X),γ d (D))
[0142] Where, γ x and γ d These represent the encoding functions for position and light, respectively.
[0143] Generate left and right images and their parallax maps from the new perspective based on preset new perspective camera parameters.
[0144] Using volume rendering technology, based on density σ t and color c tGenerate the rendering color under the camera ray r(t) = o + tD:
[0145]
[0146] Where, δ k =||x k+1 -x k ||2 represents the distance between adjacent sampling points, N k This indicates the number of sampling points.
[0147] To ensure the effective generation of new perspective images, the present invention adopts the following scheme:
[0148] 1) Generate a new reference image I' based on the pose of the reference camera using a trained neural radiation field model. ref ; By using the L1 loss function and the structural similarity loss function, the reference image I generated by the diffusion model is improved. ref and the reference image I' generated by NeRF ref Maintain consistency in texture and structure;
[0149] 2) Use the new perspective camera as the left camera, and translate the left camera horizontally by b units to form the pose of the right camera;
[0150] 3) Generate left and right color images based on the poses of the left and right cameras from the new perspective (I l ,I r ), and use the principle of light accumulation to obtain a depth map from the corresponding viewpoint, so as to obtain accurate depth information;
[0151] 4) Improve the accuracy of the disparity map by using the principle of triangulation. Based on the preset camera baseline and focal length, convert the depth map into a disparity map d.
[0152] 5) Using the warp loss function To ensure the generated binocular image maintains epipolar constraints with the disparity map:
[0153]
[0154] Among them, F w (·) represents the anti-warping function, and H and W represent the height and width of the image, respectively.
[0155] The generated binocular images and their corresponding disparity maps are used to construct a new dataset, denoted as (I). l ,I r ,d)∈D, the set D contains the generated left camera image, right camera image and disparity map, used to store and reference these images and their corresponding disparity information, where, I l and I r d represents the left and right views, and d represents the disparity map corresponding to the left view.
[0156] This invention utilizes a text-to-image diffusion model to generate reference images based on textual cues such as hairstyle, clothing, skin color, and posture, as well as a reference camera pose. A NeRF model is then used to render new perspective images based on the reference images and camera poses. Furthermore, a link is established between textual semantics and NeRF, forcing the diffusion model and the reference image generated by NeRF at the reference viewpoint to maintain consistent texture and structure, thereby ensuring that the generated human body data conforms to the textual semantics. Simultaneously, by constructing left and right camera poses and designing an anti-warping function, the generated binocular images and disparity maps are forced to maintain epipolar constraints. This invention aims to establish a unified semantic representation of text and images, combining the geometric constraints of binocular images and disparity maps, and generating new perspective human binocular images and disparity maps through a multimodal diffusion model and neural radiation fields.
[0157] To address the challenge of text-to-image generation models accurately generating geometrically constrained binocular images and disparity maps from multiple perspectives, this invention analyzes the potential contextual relationships between different words to extract textual semantics with global information; explores the differences and consistency in text and image feature representations to establish a text-image multimodal alignment mechanism; utilizes a diffusion model to generate reference images that conform to textual semantics; constructs a potential link between textual cues and neural radiation fields; and combines the geometric constraints between binocular images and disparity maps to generate high-fidelity novel-perspective human binocular images and disparity maps that conform to textual semantics from new perspectives.
[0158] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0159] Furthermore, unless otherwise stated, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. All references to this specification are incorporated by way of citation to disclose and describe methods relating to those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.
Claims
1. A method for generating human binocular images and disparity maps, characterized in that, Includes the following steps: Extract the textual semantic features of the natural language text corresponding to the image to be generated; Based on the extracted text semantic features and preset camera pose parameters, a human reference image with a specific camera pose that matches the text is generated through a diffusion model. A neural radiation field is constructed by training the position, density, and color information of the neural radiation field by inputting a reference image with a given camera pose, and obtaining the trained neural radiation field. The generated human reference image with a specific camera pose is input into the trained neural radiation field, and the image is rendered under a new perspective to generate a color image rendered under the new perspective. The color image rendered from the new perspective is used as the left camera image. The left camera image is then rendered after being horizontally shifted by a preset baseline distance to obtain the right camera image. Based on the generated left and right camera images, the depth information from the intersection of rays from the left and right cameras to the surface of the three-dimensional object is determined. Then, the depth map is converted into a disparity map using the conversion relationship between disparity and depth, thus generating a disparity map corresponding to the binocular images.
2. The method for generating human binocular images and disparity maps according to claim 1, characterized in that, The generation of a human reference image with a specific camera pose that matches the text through a diffusion model specifically includes: Human images are generated using a diffusion model based on the semantic features Γ(P) of the generated natural language text and the preset camera pose. The diffusion model is a probabilistic generative model that learns the data distribution by progressively denoising variables sampled from a Gaussian distribution. Determine the initial noise map ∈ ~N(0,I) and the condition vector. Through image-to-text diffusion model Generate target image Here, the conditional vector v consists of the textual semantic features Γ(P) and the preset camera pose. constitute; diffusion model Training is performed on the following denoising targets: Where x' is the ground truth value of the image, v is the conditional vector, and α t σ t and ω t It is a function that controls sample quality with respect to the diffusion process time t~U([0,1]); By using text semantic features and preset camera poses as conditional vectors to guide the diffusion model, a human reference image I with a specific camera pose that matches the text is generated. ref .
3. The method for generating human binocular images and disparity maps according to claim 2, characterized in that, The acquisition of the trained neural radiation field specifically includes: By inputting a reference image with a given camera pose, the position, density, and color information of the neural radiation field are trained to construct a neural radiation field with geometric position information. Density and color are predicted using two shallow multilayer perceptrons, respectively. For the t-th frame image, given the position coordinates X = (x, y, z) of the reference image and the ray direction D = (δ, φ), the density model F σ Predicted density σ t (x) is: σ t (x)=F σ (X,D) Position and light are encoded, and the position coordinates and light direction are fed into the color model F. C Predict color c t (x) is: c t (x)=F σ (c x (X),c d (D)) Where, γ x and γ d These represent the encoding functions for position and light, respectively.
4. The method for generating human binocular images and disparity maps according to claim 3, characterized in that, The generation of the color image rendered from the new perspective specifically includes: According to density σ t and color c t Generate the rendered color image under the camera ray r(t) = o + tD: Where, δ k =||x k+1 -x k ||2 represents the distance between adjacent sampling points, N k This indicates the number of sampling points.
5. The method for generating human binocular images and disparity maps according to claim 4, characterized in that, The generation of the disparity map corresponding to the binocular images specifically includes: Based on the generated left and right camera images, a depth map is obtained from the corresponding new perspective using the principle of light accumulation. Using the principle of triangulation, the depth map is converted into a disparity map d based on the preset camera pose parameters baseline and focal length; Through warp loss function Epipolar constraints are applied to the generated binocular images and disparity maps: Among them, F w (·) represents the anti-warping function, and H and W represent the height and width of the image, respectively; The generated binocular images and their corresponding disparity maps are used to construct a new dataset, denoted as (I). l ,I r ,d)∈D, where the set D represents the set containing the generated left camera image, right camera image, and disparity map, I l and I r d represents the left and right views, and d represents the disparity map corresponding to the left view.
6. The method for generating human binocular images and disparity maps according to claim 5, characterized in that, Also includes: Based on the pose of the reference camera, a new reference image I' is generated using neural radiation fields. ref ; Construct an L1 loss function to measure the absolute difference between two images at the pixel level and a structural similarity loss function to measure the overall structural similarity of images; The generated reference image is optimized at both the pixel and structural levels by minimizing the L1 loss function and the structural similarity loss function.
7. The method for generating human binocular images and disparity maps according to claim 6, characterized in that, The expression for the L1 loss function is: Among them, I ref (i) represents the reference image generated by the diffusion model, I' ref (i) represents the new reference image generated by the neural radiation field, where N is the total number of pixels in the image.
8. The method for generating human binocular images and disparity maps according to claim 7, characterized in that, The formula for calculating the structural similarity loss function SSIM is as follows: Where, μ I ,μ I' Image I ref , I' ref The average value reflecting brightness information; Image I ref , I' ref The variance, σ, reflects contrast information. Ⅱ' For image I ref , I' ref The covariance reflects the structural similarity between two images; C1 and C2 are both stabilizing terms used to prevent the denominator from being zero.
9. The method for generating human binocular images and disparity maps according to claim 1, characterized in that, The extraction of semantic features from natural language text includes the following steps: Input the natural language text corresponding to the image to be generated into the Transformer encoder; The Transformer encoder uses self-attention and cross-attention mechanisms to encode natural language text P using a pre-trained and frozen Transformer language model, generating the textual semantic features Γ(P) of the natural language text.
10. A system for generating human binocular images and disparity maps, characterized in that, include: The text feature extraction module is used to extract the text semantic features of the natural language text corresponding to the image to be generated; The reference image generation module is used to generate a human reference image with a specific camera pose that matches the text, based on the extracted text semantic features and preset camera pose parameters, through a diffusion model. The color rendering module is used to construct a neural radiation field. By inputting a reference image with a given camera pose, the position, density, and color information of the neural radiation field are trained to obtain the trained neural radiation field. The generated human reference image with a specific camera pose is input into the trained neural radiation field, and the image is rendered from a new perspective to generate a color image rendered from the new perspective. The disparity map generation module is used to take the color image rendered from the new perspective as the left camera image, and render the left camera image after horizontally shifting it by a preset baseline distance to obtain the right camera image. Based on the generated left and right camera images, by determining the depth information from the intersection of the rays from the left and right cameras to the surface of the three-dimensional object, the depth map is converted into a disparity map using the conversion relationship between disparity and depth, thus generating a disparity map corresponding to the binocular image.