Image generation method and device and augmented reality helmet

The XR headset captures multi-angle images to generate 3D models of users and environments, enhancing immersion by integrating them into reflective surfaces, addressing the limitations of existing XR systems.

CN120318472APending Publication Date: 2025-07-15BEIJING BOE DISPLAY TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510399652.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the existing virtual mirror technology, users have poor immersion experience during use and cannot effectively integrate virtual information with the real environment.

Method used

By expanding the reality helmet, obtaining multi-angle spatial environment images and user part images, generating spatial three-dimensional models and character three-dimensional models, and imaging them in the preset reflection surface, combining lighting and shadow processing to achieve real-time fusion of virtual information and the real environment.

Benefits of technology

It improves the realism and immersion experience of the preset reflective surface, can clearly display the user and real-time background, and enhances the user's immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318472A_ABST
    Figure CN120318472A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image generation method and device and an augmented reality helmet, and relates to the technical field of artificial intelligence. Acquiring a multi-angle first image of a space environment where the user is located, a second image of multiple parts of the user, a first position of the user in the space environment and a second position of a preset reflecting surface in the space environment by using the augmented reality helmet; generating a spatial three-dimensional model of the spatial environment based on the first images; generating a figure three-dimensional model of the user based on the second images; and generating imaging images of the spatial three-dimensional model and the figure three-dimensional model in the preset reflecting surface according to the first position and the second position. According to the invention, better immersion experience is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an image generation method, apparatus, and extended reality helmet. Background Art

[0002] Extended Reality (XR) refers to the combination of multiple technologies such as Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). Through computer technology and wearable devices, it creates an immersive experience environment that transcends the real world for users. The virtual mirror utilizes extended reality technology to fuse virtual information with real users for interactive experience, providing an experience that goes beyond the functions of traditional mirrors. In related technologies, the virtual mirror can only map preset scene images, resulting in a poor immersive experience during use. Summary of the Invention

[0003] The purpose of the embodiments of the present invention is to provide an image generation method, apparatus, and extended reality helmet to achieve a better immersive experience. The specific technical solutions are as follows:

[0004] In a first aspect, the embodiments of the present invention provide an image generation method, including:

[0005] When a user wears an extended reality helmet, use the extended reality helmet to obtain multi-angle first images of the space environment where the user is located, second images of multiple parts of the user, the first position of the user in the space environment, and the second position of a preset reflection surface in the space environment;

[0006] Generate a spatial three-dimensional model of the space environment based on each of the first images;

[0007] Generate a three-dimensional model of the user's character based on each of the second images;

[0008] Generate an imaging image of the spatial three-dimensional model and the character three-dimensional model in the preset reflection surface according to the first position and the second position.

[0009] In an embodiment of the present invention, the generating a spatial three-dimensional model of the space environment based on each of the first images includes:

[0010] Generate spatial point cloud data of the space environment based on each of the first images;

[0011] Use a trained spatial model to perform inference based on the spatial point cloud data to obtain a three-dimensional surface model of the space environment, where the trained spatial model is trained based on sample space environment data with three-dimensional model ground truth;

[0012] Perform texture mapping on the three-dimensional surface model of the spatial environment according to each of the first images to obtain the three-dimensional model of the space.

[0013] In one embodiment of the present invention, the method further includes:

[0014] Utilize the surface muscle sensor of the extended reality helmet to obtain the surface muscle contraction characteristics of the user's face.

[0015] The generating the three-dimensional model of the user based on each of the second images includes:

[0016] Generate the image features of the user based on each of the second images, where each of the second images includes the images of the user's eyes obtained by the cameras located at the eyes of the extended reality helmet, the images of the user's face, front body, and both hands obtained by the cameras located at the bottom of the extended reality helmet, and the images of the user's both hands on both sides of the body obtained by the cameras located on both sides of the extended reality helmet, and the image features include the image features of the user's eyes, face, front body, both hands, and both hands on both sides of the body.

[0017] Align the image features with the surface muscle contraction characteristics to obtain the fusion data of the user.

[0018] Utilize the trained human model to perform inference based on the fusion data to obtain the three-dimensional surface model of the user, where the trained human model is a multi-modal model including a full-body human model and a facial surface muscle model, and is trained based on the sample human data with three-dimensional model ground truth.

[0019] Perform texture mapping on the three-dimensional surface model of the user according to each of the second images to obtain the three-dimensional model of the person.

[0020] In one embodiment of the present invention, the mapping the three-dimensional model of the space and the three-dimensional model of the person to the preset reflection surface according to the first position and the second position includes:

[0021] Calculate the spatial grid data of the position to be displayed of the spatial environment in the preset reflection surface and the human grid data of the position to be displayed of the user in the preset reflection surface according to the first position and the second position.

[0022] Map the three-dimensional model of the space to the position where the spatial grid data in the preset reflection surface is located, and map the three-dimensional model of the person to the position where the human grid data in the preset reflection surface is located.

[0023] Synthesize the three-dimensional model of the person and the three-dimensional model of the space.

[0024] In one embodiment of the present invention, the synthesis of the three-dimensional human model and the three-dimensional space model includes:

[0025] Obtain the light source information in the space environment according to the first image and the second image;

[0026] Based on the light source information, calculate the lighting information and shadow information of the three-dimensional human model in the three-dimensional space model;

[0027] Perform lighting processing on the three-dimensional human model and the three-dimensional space model according to the lighting information, and perform shadow rendering in the three-dimensional space model according to the shadow information to obtain the synthesized three-dimensional human model and the three-dimensional space model.

[0028] In one embodiment of the present invention, the method further includes:

[0029] When the user wears the extended reality helmet and the three-dimensional human model and the three-dimensional space model exist in the preset reflection surface, use the extended reality helmet to obtain the third image of multiple angles of the space environment, the fourth image of multiple parts of the user, the third position of the user in the space environment, and the fourth position of the preset reflection surface in the space environment in real time;

[0030] Compare the first image with the third image, the second image with the fourth image, the first position with the third position, and the second position with the fourth position to obtain a real-time comparison result;

[0031] When the real-time comparison result indicates that the user and / or the space environment has changed, record the data to be corrected that has changed;

[0032] When the preset reflection surface is idle, based on the data to be corrected, adjust the three-dimensional human model and the three-dimensional space model in the preset reflection surface.

[0033] In a second aspect, an embodiment of the present invention provides an image generation device, including:

[0034] A first information acquisition module, configured to, when the user wears an extended reality helmet, use the extended reality helmet to obtain a first image of multiple angles of the space environment where the user is located, a second image of multiple parts of the user, a first position of the user in the space environment, and a second position of a preset reflection surface in the space environment;

[0035] A space model generation module, configured to generate a three-dimensional space model of the space environment based on each of the first images;

[0036] A character model generation module, configured to generate a three-dimensional model of the user's character based on each of the second images;

[0037] A model mapping module, configured to generate imaging images of the three-dimensional space model and the three-dimensional character model in the preset reflection plane according to the first position and the second position.

[0038] In an embodiment of the present invention, the space model generation module is specifically configured to:

[0039] Generate spatial point cloud data of the spatial environment based on each of the first images;

[0040] Use a trained space model to perform inference based on the spatial point cloud data to obtain a three-dimensional surface model of the spatial environment, where the trained space model is trained based on sample spatial environment data with three-dimensional model ground truth;

[0041] Perform texture mapping on the three-dimensional surface model of the spatial environment according to each of the first images to obtain the three-dimensional space model.

[0042] In an embodiment of the present invention, the apparatus further includes:

[0043] A facial feature acquisition module, configured to acquire surface muscle contraction features of the user's face by using a surface muscle sensor on the surface of the extended reality helmet;

[0044] The character model generation module is specifically configured to:

[0045] Generate image features of the user based on each of the second images, where each of the second images includes images of the user's eyes acquired by a camera located at the eye part of the extended reality helmet, images of the user's face, front body, and both hands acquired by a camera located at the bottom of the extended reality helmet, and images of the user's both hands located on both sides of the body acquired by cameras located on both sides of the extended reality helmet, and the image features include image features of the user's eyes, face, front body, both hands, and both hands located on both sides of the body;

[0046] Align the image features with the surface muscle contraction features to obtain fusion data of the user;

[0047] Use a trained character model to perform inference based on the fusion data to obtain a three-dimensional surface model of the user, where the trained character model is a multi-modal model including a full-body character model and a facial surface muscle model, and is trained based on sample character data with three-dimensional model ground truth;

[0048] Texture map the three-dimensional surface model of the user according to each of the second images to obtain the three-dimensional model of the person.

[0049] In one embodiment of the present invention, the model mapping module includes:

[0050] A grid data acquisition sub-module, configured to calculate the spatial grid data of the position to be displayed in the preset reflection plane of the spatial environment and the human grid data of the position to be displayed in the preset reflection plane of the user according to the first position and the second position;

[0051] A model mapping sub-module, configured to map the three-dimensional spatial model to the position where the spatial grid data is located in the preset reflection plane, and map the three-dimensional model of the person to the position where the human grid data is located in the preset reflection plane;

[0052] A model synthesis sub-module, configured to synthesize the three-dimensional model of the person and the three-dimensional spatial model.

[0053] In one embodiment of the present invention, the model synthesis sub-module is specifically configured to:

[0054] Obtain the light source information in the spatial environment according to the first image and the second image;

[0055] Based on the light source information, calculate the illumination information and shadow information of the three-dimensional model of the person in the three-dimensional spatial model;

[0056] Perform illumination processing on the three-dimensional model of the person and the three-dimensional spatial model according to the illumination information, and perform shadow rendering in the three-dimensional spatial model according to the shadow information to obtain the synthesized three-dimensional model of the person and the three-dimensional spatial model.

[0057] In one embodiment of the present invention, the device further includes:

[0058] A second information acquisition module, configured to, when the user wears the extended reality helmet and the three-dimensional model of the person and the three-dimensional spatial model exist in the preset reflection plane, use the extended reality helmet to acquire the third images of multiple angles of the spatial environment, the fourth images of multiple parts of the user, the third position of the user in the spatial environment, and the fourth position of the preset reflection plane in the spatial environment in real time;

[0059] An information comparison module, configured to compare the first image with the third image, the second image with the fourth image, the first position with the third position, and the second position with the fourth position to obtain a real-time comparison result;

[0060] A data recording module, configured to record the data to be corrected that has changed when the real-time comparison result indicates that the user and / or the spatial environment has changed;

[0061] A model adjustment module, configured to adjust the three-dimensional human model and the three-dimensional spatial model in the preset reflection surface based on the data to be corrected when the preset reflection surface is idle.

[0062] In a third aspect, an embodiment of the present invention provides an extended reality helmet, including a plurality of cameras, a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus, and each of the cameras is located in multiple orientations of the extended reality helmet;

[0063] The camera is configured to collect images of the user and the spatial environment, and obtain the position information of the user and the preset reflection surface;

[0064] The memory is configured to store a computer program;

[0065] The processor is configured to implement any of the above method steps when executing the program stored on the memory.

[0066] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, any of the above method steps is implemented.

[0067] An embodiment of the present invention further provides a computer program product including instructions, which when running on a computer, causes the computer to execute any of the above image generation methods.

[0068] Advantageous effects of the embodiments of the present invention:

[0069] The image generation method provided by the embodiments of the present invention, when the user wears an extended reality helmet, uses the extended reality helmet to obtain first images of multiple angles of the spatial environment where the user is located, second images of multiple parts of the user, the first position of the user in the spatial environment, and the second position of a preset reflecting surface in the spatial environment; generates a spatial three-dimensional model of the spatial environment based on each first image; generates a three-dimensional model of the user based on each second image; and generates an imaging image of the spatial three-dimensional model and the three-dimensional model of the user in the preset reflecting surface according to the first position and the second position. Since both the first images and the second images are real-time images of multiple angles and multiple parts collected when the user wears the extended reality helmet, a three-dimensional model of the spatial environment that is clear at each angle and a three-dimensional model of the user that is clear at each part can be generated. An imaging image of the spatial three-dimensional model and the three-dimensional model of the user is generated in the preset reflecting surface, so that the user and the real-time real background can be displayed in the preset reflecting surface, improving the realism of the preset reflecting surface and providing a better immersive experience for the user.

[0070] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.

[0072] Figure 1 It is the first flowchart of the image generation method according to the embodiments of the present invention;

[0073] Figure 2 It is a possible implementation manner of step S102 according to the embodiments of the present invention;

[0074] Figure 3 It is the second flowchart of the image generation method according to the embodiments of the present invention;

[0075] Figure 4 It is a possible implementation manner of step S104 according to the embodiments of the present invention;

[0076] Figure 5 It is a possible implementation manner of step S403 according to the embodiments of the present invention;

[0077] Figure 6 It is the third flowchart of the image generation method according to the embodiments of the present invention;

[0078] Figure 7A schematic structural diagram of an image generation device according to an embodiment of the present invention;

[0079] Figure 8-1 A schematic structural diagram of an extended reality helmet according to an embodiment of the present invention;

[0080] Figure 8-2 A front view example diagram of an extended reality helmet provided by an embodiment of the present invention;

[0081] Figure 8-3 A side view example diagram of an extended reality helmet provided by an embodiment of the present invention. Detailed implementation manners

[0082] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on this application belong to the scope of protection of the present invention.

[0083] In the related art, a virtual mirror can only map a user in a virtual background, resulting in a poor immersive experience for the user during use. To solve this problem, an embodiment of the present invention provides an image generation method and device.

[0084] Hereinafter, the image generation method of the embodiment of the present invention will be described in detail.

[0085] See Figure 1 , Figure 1 which is the first flowchart of the image generation method provided by the embodiment of the present invention, including:

[0086] Step S101, when the user wears the extended reality helmet, use the extended reality helmet to obtain multi-angle first images of the space environment where the user is located, second images of multiple parts of the user, the first position of the user in the space environment, and the second position of a preset reflection surface in the space environment.

[0087] The preset reflecting surface refers to the mirror application in extended reality, which is an application technology that simulates the functions of a real mirror through digital technology. Exemplarily, it can be a virtual mirror in extended reality. Cameras are installed around the extended reality helmet. Specifically, the cameras are installed at positions such as the eyes, bottom, front, back, and both sides of the helmet. When the user wears the extended reality helmet and activates the preset reflecting surface, the cameras can capture the 360-degree background images around the user, obtaining the first images of multiple angles of the user's spatial environment, and capture the images of the whole body and face of the person, obtaining the second images of multiple parts of the user. Both the first images and the second images are real-time images collected when the user wears the extended reality helmet.

[0088] The extended reality helmet can also obtain the first position of the user in the spatial environment and the second position of the preset reflecting surface in the spatial environment. Exemplarily, it can be achieved through the 6dof (six degrees of freedom) system carried by the extended reality helmet. Specifically, 6dof refers to the six degrees of freedom of an object in three-dimensional space, including three translational degrees of freedom (left and right movement along the X-axis, up and down movement along the Y-axis, and forward and backward movement along the Z-axis) and three rotational degrees of freedom (pitch around the X-axis such as nodding, yaw around the Y-axis such as shaking the head, and roll around the Z-axis such as tilting the head). The 6dof system can achieve the tracking of the object's movement through the cooperation of various sensors and tracking technologies possessed by the extended reality helmet, such as inertial measurement unit technology, optical tracking technology, etc., to achieve high-precision position and attitude tracking.

[0089] Based on the obtained user eye images through the 6dof system, calculate the position and orientation of the user and the preset reflecting surface in the spatial environment, as well as the coordinates of the user (where the eyes are located) and the preset reflecting surface in the spatial environment, to obtain the first position of the user in the spatial environment and the second position of the preset reflecting surface in the spatial environment.

[0090] Step S102, generate a three-dimensional spatial model of the spatial environment based on each of the first images;

[0091] Step S103, generate a three-dimensional character model of the user based on each of the second images.

[0092] Generate a three-dimensional spatial model of the spatial environment based on the first images in all directions of the spatial environment, and generate a three-dimensional human model of the user based on the second images of each part of the user. Exemplarily, various three-dimensional modeling methods such as multi-view stereo vision methods and deep learning algorithms can be used. The multi-view stereo vision method refers to analyzing the image features and texture information of the first images and the second images through vision algorithms to calculate the three-dimensional results of the spatial environment and the user. Deep learning algorithms, such as applying models like generative adversarial networks (GANs) and variational autoencoders (VAEs), calculate the mapping relationship between the two-dimensional first images and the second images and the three-dimensional model to generate the three-dimensional models of the spatial environment and the user.

[0093] Step S104: Generate an imaging image of the three-dimensional spatial model and the three-dimensional human model in a preset reflection plane according to the first position and the second position.

[0094] According to the first position of the user in the spatial environment and the second position of the preset reflection plane in the spatial environment, map the three-dimensional spatial model of the spatial environment and the three-dimensional human model of the user to the preset reflection plane to generate an imaging image of the three-dimensional spatial model and the three-dimensional human model in the preset reflection plane. Specifically, it can be to first synthesize the three-dimensional human model and the three-dimensional spatial model, and then map the synthesized three-dimensional model to the preset reflection plane, or first map the three-dimensional spatial model to the preset reflection plane, and then map the three-dimensional human model to the preset reflection plane, and synthesize the three-dimensional human model and the three-dimensional spatial model in the preset reflection plane to obtain an imaging image of the synthesized three-dimensional spatial model and three-dimensional human model in the preset reflection plane.

[0095] As can be seen from the above, in the image generation method provided by the embodiment of the present invention, when the user wears an extended reality helmet, the extended reality helmet is used to obtain the first images from multiple angles of the spatial environment where the user is located, the second images of multiple parts of the user, the first position of the user in the spatial environment, and the second position of the preset reflection plane in the spatial environment; based on each first image, generate a three-dimensional spatial model of the spatial environment; based on each second image, generate a three-dimensional human model of the user; according to the first position and the second position, generate an imaging image of the three-dimensional spatial model and the three-dimensional human model in the preset reflection plane. Since both the first images and the second images are real-time images from multiple angles and multiple parts collected when the user wears the extended reality helmet, a three-dimensional model of the spatial environment that is clear in all angles and a three-dimensional model of the user that is clear in all parts can be generated. Generate an imaging image of the three-dimensional spatial model and the three-dimensional human model in the preset reflection plane so that the user and the real-time real background can be displayed in the preset reflection plane, improving the realism of the preset reflection plane and providing a better immersive experience for the user.

[0096] In one embodiment of the present invention, as Figure 2As shown, the above step S102 generates a three-dimensional spatial model of the spatial environment based on each of the first images, including:

[0097] Step S201, generating spatial point cloud data of the spatial environment based on each of the first images;

[0098] Step S202, using the trained spatial model to perform inference based on the spatial point cloud data to obtain a three-dimensional surface model of the spatial environment;

[0099] Wherein, the trained spatial model is trained based on sample spatial environment data with three-dimensional model ground truth;

[0100] Step S203, performing texture mapping on the three-dimensional surface model of the spatial environment according to each of the first images to obtain the three-dimensional spatial model.

[0101] Perform alignment processing on the first images of the spatial environment from multiple angles to generate three-dimensional spatial point cloud data of the spatial environment. Specifically, it can be achieved through methods such as multi-view stereo vision (MVS) and structure from motion (SfM). Perform preprocessing such as noise reduction and enhancement on the spatial point cloud data, and input the processed spatial point cloud data into the trained spatial model. Specifically, the spatial model is a deep learning model, such as scene representation networks (SRNs), neural radiance fields (NeRF), etc.

[0102] The training process of the spatial model includes: inputting sample spatial environment data (images from multiple angles) with three-dimensional model ground truth into the spatial model to be trained to generate a three-dimensional model, calculating the model loss according to the difference between the generated three-dimensional model and the three-dimensional model ground truth, and adjusting the spatial model based on the model loss until the spatial model is trained.

[0103] Using the trained spatial model to perform real-time inference based on the spatial point cloud data, generating a three-dimensional surface model of the spatial environment through surface reconstruction, and then performing texture mapping on the three-dimensional surface model according to the first images to obtain a three-dimensional model of the spatial environment.

[0104] As can be seen from the above, the image generation method provided by the embodiments of the present invention uses the trained spatial model to generate a three-dimensional surface model of the spatial environment according to the point cloud data of the spatial environment, and then uses the first images to perform texture mapping to obtain a three-dimensional spatial model of the spatial environment. Generating a three-dimensional model of the spatial environment based on the first images obtained in real time realizes the real-time generation of the spatial environment model.

[0105] In one embodiment of the present invention, as Figure 3 shown, the above method further includes:

[0106] Using the surface muscle sensor of the extended reality helmet to obtain the surface muscle contraction characteristics of the user's face;

[0107] The above step S103 generates a three-dimensional model of the user's body based on each of the second images, including:

[0108] Step S301 generates image features of the user based on each of the second images;

[0109] Among them, each second image includes an image of the user's eyes obtained by a camera located at the eye part of the extended reality helmet, images of the user's face, front body, and both hands obtained by a camera located at the bottom of the extended reality helmet, and images of the user's both hands on both sides of the body obtained by cameras located on both sides of the extended reality helmet; the image features include image features of the user's eyes, face, front body, both hands, and both hands on both sides of the body;

[0110] Step S302 aligns the image features with the surface muscle contraction features to obtain the fusion data of the user;

[0111] Step S303 uses the trained body model to perform inference based on the fusion data to obtain the three-dimensional surface model of the user;

[0112] Among them, the trained body model is a multi-modal model including a full-body model and a facial surface muscle model, and is trained based on sample body data with three-dimensional models;

[0113] Step S304 performs texture mapping on the three-dimensional surface model of the user according to each of the second images to obtain the three-dimensional model of the body.

[0114] The surface muscle sensor is provided at the face-attached foam of the extended reality helmet to obtain the muscle contraction features of the user's face and obtain the surface muscle contraction features of the user's body. The camera at the eye part of the helmet captures the human eyes, the camera at the bottom of the helmet captures the face, front body, and both hands, and the cameras on both sides of the helmet capture the second images when both hands are on both sides of the body to obtain the image features of each part of the user.

[0115] Combining the image features and the surface muscle contraction features, aligning the full body and facial meshes of the person in time and space to generate the fusion data of the user. Specifically, the fusion data is the three-dimensional point cloud data of the user. Perform preprocessing such as noise reduction and enhancement on the fusion data, and input the processed fusion data into the trained body model. Specifically, the spatial model is a deep learning model, which is a multi-modal model integrating a full-body model and a facial surface muscle model, such as SMPL-X (statistical learning-based parametric three-dimensional human model), portraitGen (multi-modal editing) combined with the EMOCA model (expression control), electromyogram signal-driven multi-modal model, etc.

[0116] The training process of the human model includes: inputting sample human data (image features of multiple parts and surface muscle contraction features) with 3D model ground truth into the human model to be trained to generate a 3D model, calculating the model loss based on the difference between the generated 3D model and the 3D model ground truth, and adjusting the human model based on the model loss until the training of the human model is completed.

[0117] Using the trained human model, real-time inference is performed based on the fusion data, a 3D surface model of the user is generated through surface reconstruction, and then texture mapping is performed on the 3D surface model according to the second image to obtain a 3D model of the user's human figure.

[0118] As can be seen from the above, the image generation method provided by the embodiments of the present invention uses the trained human model to generate a 3D surface model of the human figure according to the fusion data of the user, and then uses the second image for texture mapping to obtain a 3D model of the user's human figure. By generating the image features of the user based on the second image obtained in real time, the real-time generation of the user's human model is realized, and the problem of insufficient presentation of details of multiple parts such as the user's hand is solved. Based on the surface muscle contraction features of the user obtained in real time, more detailed facial expressions of the human figure can be recognized, and the problem of insufficient presentation of facial details such as the user's face and lips is solved, so that the facial expression details of the human figure can be more clearly and accurately displayed in the preset reflection surface, further improving the user's immersive experience.

[0119] In one embodiment of the present invention, as Figure 4 shown, the above step S104 maps the spatial 3D model and the human 3D model to the preset reflection surface according to the first position and the second position, including:

[0120] Step S401, according to the first position and the second position, calculate the spatial grid data of the position to be displayed of the spatial environment in the preset reflection surface, and the human grid data of the position to be displayed of the user in the preset reflection surface;

[0121] Step S402, map the spatial 3D model to the position where the spatial grid data in the preset reflection surface is located, and map the human 3D model to the position where the human grid data in the preset reflection surface is located;

[0122] Step S403, synthesize the human 3D model and the spatial 3D model.

[0123] As mentioned above, the first position includes the orientation of the user in the spatial environment obtained through the position of the user's eyes, and the coordinates of the user in the spatial environment, and the second position includes the position orientation of the preset reflection surface in the spatial environment and the coordinates of the preset reflection surface in the spatial environment.

[0124] Based on the coordinates of the user and the coordinates of the preset reflection surface, calculate the content that the three-dimensional human model needs to display in the preset reflection surface, obtain the human mesh data in the preset reflection surface, perform scene texture mapping rendering on the human mesh data, and map the three-dimensional human model to the position where the human mesh data is located in the preset reflection surface.

[0125] Based on the coordinates of the user and the coordinates of the preset reflection surface, calculate the content that the three-dimensional space model needs to display in the preset reflection surface, obtain the space mesh data in the preset reflection surface, perform scene texture mapping rendering on the space mesh data, and map the three-dimensional space model to the position where the space mesh data is located in the preset reflection surface.

[0126] In the preset reflection surface, synthesize the three-dimensional human model and the three-dimensional space model, so that the real-time background and the user in the real-time background are displayed in the preset reflection surface. The human will cause a certain occlusion of the three-dimensional space model behind in the preset reflection surface. When the position of the human changes, the originally occluded part is exposed, and the new position where the human is located will cause a new occlusion of the three-dimensional space model.

[0127] As can be seen from the above, the image generation method provided by the embodiment of the present invention calculates the space mesh data of the space environment at the position to be displayed in the preset reflection surface and the human mesh data of the user at the position to be displayed in the preset reflection surface according to the first position and the second position, maps the three-dimensional space model to the position where the space mesh data is located, maps the three-dimensional human model to the position where the human mesh data is located, and synthesizes the three-dimensional human model and the three-dimensional space model in the preset reflection surface, realizing the display of the real-time background and the user in the real-time background in the preset reflection surface, solving the problem of insufficient presentation of the space scene behind the user, and improving the authenticity of the preset reflection surface.

[0128] In one embodiment of the present invention, as Figure 5 shown, the above step S403 of synthesizing the three-dimensional human model and the three-dimensional space model includes:

[0129] Step S501, obtain the light source information in the space environment according to the first image and the second image;

[0130] Step S502, based on the light source information, calculate the illumination information and shadow information of the three-dimensional human model in the three-dimensional space model;

[0131] Step S503, perform illumination processing on the three-dimensional human model and the three-dimensional space model according to the illumination information, and perform shadow drawing in the three-dimensional space model according to the shadow information to obtain the synthesized three-dimensional human model and three-dimensional space model.

[0132] Obtain the light source information of the spatial environment based on the first image of the spatial environment and the second image of the user. Specifically, the light source information includes the ambient light in the spatial scene, and may also include additionally set light sources, such as parallel lights, point lights, area lights, etc.

[0133] Calculate the illumination information received by the three-dimensional human model in the three-dimensional spatial model according to the light source information, and the shadow information caused by the three-dimensional human model in the three-dimensional spatial model. Specifically, it may also include the illumination reflection effect caused by the material properties (reflectivity, transparency, etc.) of the three-dimensional human model and the three-dimensional spatial model.

[0134] Perform illumination processing on the three-dimensional human model and the three-dimensional spatial model according to the illumination information, and perform shadow rendering in the three-dimensional spatial model according to the shadow information to obtain the synthesized three-dimensional human model and three-dimensional spatial model.

[0135] Exemplarily, the illumination processing can be performed by using baked lighting for static scenes in the spatial environment and real-time lighting for the user; the shadow rendering can be implemented by using methods such as Shadow Map and Distance Field Shadow.

[0136] As can be seen from the above, the image generation method provided by the embodiments of the present invention performs illumination processing and shadow rendering on the three-dimensional human model and the three-dimensional spatial model according to the real-time light source information in the spatial environment, so that the illumination and shadow in the preset reflection surface are more consistent with the actual scene and the user, improving the user's immersive experience.

[0137] In one embodiment of the present invention, as Figure 6 shown, the above method further includes:

[0138] Step S601, when the user wears an extended reality helmet and there are a three-dimensional human model and a three-dimensional spatial model in the preset reflection surface, use the extended reality helmet to obtain the third image of multiple angles of the spatial environment, the fourth image of multiple parts of the user, the third position of the user in the spatial environment, and the fourth position of the preset reflection surface in the spatial environment in real time;

[0139] Step S602, compare the first image with the third image, the second image with the fourth image, the first position with the third position, and the second position with the fourth position to obtain a real-time comparison result;

[0140] Step S603, when the real-time comparison result indicates that the user and / or the spatial environment has changed, record the data to be corrected that has changed.

[0141] Step S604. When the preset reflection surface is idle, based on the data to be corrected, adjust the 3D human model and the 3D space model in the preset reflection surface.

[0142] When the user wears the extended reality helmet and activates the preset reflection surface, first obtain the image of the current space environment and the image of the user, and determine whether the current space environment is a new space environment and whether the user is using the extended reality helmet for the first time. If so, it means that there is no 3D human model of the user and the 3D space model of the current space environment in the preset reflection surface. In this case, perform an initial pre-scan on the space environment and the user to obtain the first image of the space environment and the second image of the user, and according to the image generation method described above, obtain the synthesized 3D space model and 3D human model in the preset reflection surface.

[0143] If it is determined that the current space environment is not a new space environment and the user is not using the extended reality helmet for the first time, it means that there are a 3D human model and a 3D space model in the preset reflection surface. In this case, use the camera of the extended reality helmet to obtain the current real-time third image of multiple angles of the space environment, the current real-time fourth image of multiple parts of the user, the current real-time third position of the user in the space environment, and the current real-time fourth position of the preset reflection surface in the space environment.

[0144] Compare the first image with the third image, the second image with the fourth image, the first position with the third position, and the second position with the fourth position to obtain a real-time comparison result. The real-time comparison result indicates whether there are changes between the current states of the user and the space environment and the existing 3D space model and 3D human model in the preset reflection surface. If there are changes, record the changed parameters as the data to be corrected.

[0145] When the preset reflection surface is idle, adjust the 3D human model and the 3D space model in the preset reflection surface according to the data to be corrected, so as to realize the continuous iterative optimization and update of the model.

[0146] Exemplarily, the idle state of the preset reflection surface can be selected by the user, and real-time detection and optimization update are performed each time it is started, or optimization update is performed when the user selects the idle state. Without affecting the normal use of the user, the real-time performance of the model in the preset reflection surface is ensured.

[0147] Exemplarily, adjusting the 3D human model and the 3D space model in the preset reflection surface according to the data to be corrected can be to update and iterate the existing model, or to regenerate a new model according to the newly obtained real-time data.

[0148] In one embodiment of the present invention, the above image generation method is not only applicable to displaying a real spatial scene in a mirror, but also supports the overlay of virtual items (such as clothes, hats, accessories, etc.), thus being applied to shopping scenarios. Exemplarily, a model of a virtual item in a preset reflection surface can be established in advance. When there are a spatial three-dimensional model and a human three-dimensional model in the preset reflection surface, the virtual item is synthesized with the spatial three-dimensional model and the human three-dimensional model; it also supports the application of a multi-person conferencing system. Exemplarily, in this case, the reflection system module of the preset reflection surface can be turned off or removed, and the human three-dimensional model and the spatial three-dimensional model are directly displayed, and the participating members can conduct a real-time three-dimensional video conference.

[0149] As can be seen from the above, the image generation method provided by the embodiment of the present invention pre-scans the user and the spatial environment when a new user or a new spatial environment appears for the first time until three-dimensional models of the user and the spatial environment exist in the preset reflection surface; when three-dimensional models of the user and the spatial environment already exist in the preset reflection surface, it also obtains real-time images of the current user and the spatial environment in real time, as well as the real-time positions of the user and the preset reflection surface in the spatial environment, and continuously updates the models in the preset reflection surface according to the differences between the real-time data and the existing models, continuously iterating and optimizing to achieve real-time rendering and display, improving the real-time performance of the three-dimensional models of the user and the spatial environment.

[0150] In the technical solution of this application, operations such as the acquisition, storage, use, processing, transmission, provision, and disclosure of the user's personal information are all carried out with the user's authorization.

[0151] See Figure 7 , Figure 7 which is a schematic structural diagram of an image generation device, including:

[0152] A first information acquisition module 701, configured to, when the user wears an extended reality helmet, use the extended reality helmet to acquire multi-angle first images of the spatial environment where the user is located, second images of multiple parts of the user, the first position of the user in the spatial environment, and the second position of a preset reflection surface in the spatial environment;

[0153] A spatial model generation module 702, configured to generate a spatial three-dimensional model of the spatial environment based on each of the first images;

[0154] A human model generation module 703, configured to generate a human three-dimensional model of the user based on each of the second images;

[0155] A model mapping module 704, configured to generate imaging images of the spatial three-dimensional model and the human three-dimensional model in the preset reflection surface according to the first position and the second position.

[0156] As can be seen from the above, the image generation device provided by the embodiment of the present invention, when the user wears the extended reality helmet, uses the extended reality helmet to obtain the first images of multiple angles of the space environment where the user is located, the second images of multiple parts of the user, the first position of the user in the space environment, and the second position of the preset reflection surface in the space environment; based on each first image, generate a spatial three-dimensional model of the space environment; based on each second image, generate a three-dimensional model of the user; according to the first position and the second position, generate the imaging images of the spatial three-dimensional model and the three-dimensional model of the user in the preset reflection surface. Since both the first image and the second image are real-time images collected when the user wears the extended reality helmet. Based on this, generate the three-dimensional models of the space environment and the user, and generate the imaging images of the spatial three-dimensional model and the three-dimensional model of the user in the preset reflection surface, so that the user and the real-time real background can be displayed in the preset reflection surface, improving the realism of the preset reflection surface and providing a better immersive experience for the user.

[0157] In one embodiment of the present invention, the spatial model generation module 702 is specifically configured to:

[0158] Based on each of the first images, generate spatial point cloud data of the space environment;

[0159] Using the trained spatial model, perform inference based on the spatial point cloud data to obtain a three-dimensional surface model of the space environment, where the trained spatial model is trained based on the sample space environment data with three-dimensional models;

[0160] Perform texture mapping on the three-dimensional surface model of the space environment according to each of the first images to obtain the spatial three-dimensional model.

[0161] As can be seen from the above, the image generation device provided by the embodiment of the present invention uses the trained spatial model to generate a three-dimensional surface model of the space environment according to the point cloud data of the space environment, and then performs texture mapping using the first image to obtain the spatial three-dimensional model of the space environment. Generating the three-dimensional model of the space environment based on the real-time acquired first image realizes the real-time generation of the space environment model.

[0162] In one embodiment of the present invention, the device further includes:

[0163] A facial feature acquisition module, configured to use the surface muscle sensor of the extended reality helmet to acquire the surface muscle contraction features of the user's face;

[0164] The human model generation module 703 is specifically configured to:

[0165] Generate the user's image features based on each of the second images, where each of the second images includes an image of the user's eyes acquired by a camera located at the eye part of the extended reality helmet, an image of the user's face, front body, and both hands acquired by a camera located at the bottom of the extended reality helmet, and an image of the user's both hands on both sides of the body acquired by cameras located on both sides of the extended reality helmet, and the image features include the image features of the user's eyes, face, front body, both hands, and both hands on both sides of the body;

[0166] Perform data alignment on the image features and the surface muscle contraction features to obtain the fused data of the user;

[0167] Use the trained human model to perform inference based on the fused data to obtain the three-dimensional surface model of the user, where the trained human model is a multi-modal model including a full-body human model and a facial surface muscle model, and is trained based on sample human data with three-dimensional models;

[0168] Perform texture mapping on the three-dimensional surface model of the user according to each of the second images to obtain the three-dimensional human model.

[0169] As can be seen from the above, the image generation device provided by the embodiment of the present invention uses the trained human model to generate the three-dimensional surface model of a person according to the fused data of the user, and then uses the second image for texture mapping to obtain the three-dimensional human model of the user. Generating the user's image features based on the second images acquired in real time realizes the real-time generation of the user's human model, and solves the problem of insufficient presentation of details of multiple parts such as the human hand. Based on the surface muscle contraction features of the user acquired in real time, more detailed human facial expressions can be recognized, and the problem of insufficient presentation of facial details such as the human face and lips is solved, so that the facial expression details of the person can be more clearly and accurately displayed in the preset reflection surface, further improving the user's immersive experience.

[0170] In one embodiment of the present invention, the model mapping module 704 includes:

[0171] A grid data acquisition sub-module, configured to calculate the spatial grid data of the position to be displayed of the spatial environment in the preset reflection surface and the human grid data of the position to be displayed of the user in the preset reflection surface according to the first position and the second position;

[0172] A model mapping sub-module, configured to map the three-dimensional spatial model to the position where the spatial grid data is located in the preset reflection surface, and map the three-dimensional human model to the position where the human grid data is located in the preset reflection surface;

[0173] A model synthesis sub-module for synthesizing the three-dimensional model of the person and the three-dimensional model of the space.

[0174] As can be seen from the above, the image generation device provided by the embodiment of the present invention calculates the spatial grid data of the position to be displayed in the preset reflection plane of the spatial environment and the person grid data of the position to be displayed in the preset reflection plane of the user according to the first position and the second position, maps the three-dimensional model of the space to the position where the spatial grid data is located, maps the three-dimensional model of the person to the position where the person grid data is located, and synthesizes the three-dimensional model of the person and the three-dimensional model of the space in the preset reflection plane, realizing the display of the real-time background and the user in the real-time background in the preset reflection plane, solving the problem of insufficient presentation of the spatial scene behind the user, and improving the authenticity of the preset reflection plane.

[0175] In one embodiment of the present invention, the model synthesis sub-module is specifically used for:

[0176] Obtain the light source information in the spatial environment according to the first image and the second image;

[0177] Based on the light source information, calculate the lighting information and shadow information of the three-dimensional model of the person in the three-dimensional model of the space;

[0178] Perform lighting processing on the three-dimensional model of the person and the three-dimensional model of the space according to the lighting information, and perform shadow drawing in the three-dimensional model of the space according to the shadow information to obtain the synthesized three-dimensional model of the person and the three-dimensional model of the space.

[0179] As can be seen from the above, the image generation device provided by the embodiment of the present invention performs lighting processing and shadow drawing on the three-dimensional model of the person and the three-dimensional model of the space according to the real-time light source information in the spatial environment, so that the lighting and shadows in the preset reflection plane are more consistent with the actual scene and the user, improving the immersive experience of the user.

[0180] In one embodiment of the present invention, the device further includes:

[0181] A second information acquisition module for, when the user wears the extended reality helmet and the three-dimensional model of the person and the three-dimensional model of the space exist in the preset reflection plane, using the extended reality helmet to acquire the third image of multiple angles of the spatial environment, the fourth image of multiple parts of the user, the third position of the user in the spatial environment, and the fourth position of the preset reflection plane in the spatial environment in real time;

[0182] An information comparison module for comparing the first image with the third image, the second image with the fourth image, the first position with the third position, and the second position with the fourth position to obtain a real-time comparison result;

[0183] A data recording module, configured to record the to-be-corrected data with changes when the real-time comparison result indicates that the user and / or the space environment has changed;

[0184] A model adjustment module, configured to adjust the three-dimensional model of the person and the three-dimensional model of the space in the preset reflecting surface based on the to-be-corrected data when the preset reflecting surface is idle.

[0185] As can be seen from the above, the image generation device provided by the embodiment of the present invention pre-scans the user and the space environment when a new user or a new space environment appears for the first time until there are three-dimensional models of the user and the space environment in the preset reflecting surface; when there are already three-dimensional models of the user and the space environment in the preset reflecting surface, it also acquires the real-time images of the current user and the space environment in real time, as well as the real-time positions of the user and the preset reflecting surface in the space environment, and continuously updates the models in the preset reflecting surface according to the differences between the real-time data and the existing models, continuously iterates and optimizes, realizes real-time rendering and display, and improves the real-time performance of the three-dimensional models of the user and the space environment.

[0186] See Figure 8-1 , Figure 8-1 FIG. is a schematic structural diagram of an extended reality helmet, including a plurality of cameras 801, a processor 802, a communication interface 803, a memory 804, and a communication bus 805. Among them, the processor 802, the communication interface 803, and the memory 804 complete mutual communication through the communication bus 805, and each of the cameras 801 is located at multiple orientations of the extended reality helmet;

[0187] The camera 801 is configured to collect images of the user and the space environment and obtain the position information of the user and the preset reflecting surface;

[0188] The memory 804 is configured to store computer programs;

[0189] The processor 802 is configured to implement the image generation method described in any one of the above when executing the program stored in the memory.

[0190] Exemplarily, the embodiment of the present invention provides a front view example diagram and a side view example diagram of an extended reality helmet, as shown in Figure 8-2 and 8-3 shown.

[0191] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0192] The communication interface is used for communication between the above electronic device and other devices.

[0193] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0194] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0195] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above image generation methods are implemented.

[0196] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the image generation methods in the above embodiments.

[0197] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0198] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.

[0199] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0200] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.

Claims

1. An image generation method, characterized in that, Including: When the user wears the extended reality helmet, using the extended reality helmet to obtain multi-angle first images of the spatial environment where the user is located, second images of multiple parts of the user, the first position of the user in the spatial environment, and the second position of a preset reflecting surface in the spatial environment; Generating a spatial three-dimensional model of the spatial environment based on each of the first images; Generating a three-dimensional model of the user's character based on each of the second images; Generating an imaging image of the spatial three-dimensional model and the character three-dimensional model in the preset reflecting surface according to the first position and the second position.

2. The method according to claim 1, wherein The generating a spatial three-dimensional model of the spatial environment based on each of the first images includes: Generating spatial point cloud data of the spatial environment based on each of the first images; Using a trained spatial model to perform inference based on the spatial point cloud data to obtain a three-dimensional surface model of the spatial environment, where the trained spatial model is trained based on sample spatial environment data with three-dimensional model ground truth; Performing texture mapping on the three-dimensional surface model of the spatial environment according to each of the first images to obtain the spatial three-dimensional model.

3. The method according to claim 1, wherein The method further includes: Using the surface myoelectric sensor of the extended reality helmet to obtain the surface myoelectric contraction characteristics of the user's face; The generating a three-dimensional model of the user's character based on each of the second images includes: Generating image features of the user based on each of the second images, where each of the second images includes images of the user's eyes obtained by a camera located at the eye part of the extended reality helmet, images of the user's face, front body, and both hands obtained by a camera located at the bottom of the extended reality helmet, and images of the user's both hands on both sides of the body obtained by cameras located on both sides of the extended reality helmet, and the image features include image features of the user's eyes, face, front body, both hands, and both hands on both sides of the body; Aligning the image features with the surface myoelectric contraction characteristics to obtain the fusion data of the user; Using a trained character model to perform inference based on the fusion data to obtain a three-dimensional surface model of the user, where the trained character model is a multi-modal model including a full-body model of the character and a facial surface myoelectric model, and is trained based on sample character data with three-dimensional model ground truth; Performing texture mapping on the three-dimensional surface model of the user according to each of the second images to obtain the three-dimensional model of the user's character.

4. The method according to claim 1, wherein The mapping the spatial three-dimensional model and the character three-dimensional model to the preset reflecting surface according to the first position and the second position includes: Calculating spatial grid data of the position to be displayed of the spatial environment in the preset reflecting surface and character grid data of the position to be displayed of the user in the preset reflecting surface according to the first position and the second position; Mapping the spatial three-dimensional model to the position where the spatial grid data in the preset reflecting surface is located, and mapping the character three-dimensional model to the position where the character grid data in the preset reflecting surface is located; Synthesize the three-dimensional model of the person and the three-dimensional model of the space.

5. The method according to claim 4, wherein The synthesizing the three-dimensional model of the person and the three-dimensional model of the space includes: Obtain the light source information in the space environment according to the first image and the second image; Based on the light source information, calculate the lighting information and shadow information of the three-dimensional model of the person in the three-dimensional model of the space; Perform lighting processing on the three-dimensional model of the person and the three-dimensional model of the space according to the lighting information, and perform shadow rendering in the three-dimensional model of the space according to the shadow information to obtain the synthesized three-dimensional model of the person and the three-dimensional model of the space.

6. The method according to claim 1, wherein The method further includes: When the user wears the extended reality helmet and the three-dimensional model of the person and the three-dimensional model of the space exist in the preset reflection surface, use the extended reality helmet to obtain third images of multiple angles of the space environment, fourth images of multiple parts of the user, the third position of the user in the space environment, and the fourth position of the preset reflection surface in the space environment in real time; Compare the first image with the third image, the second image with the fourth image, the first position with the third position, and the second position with the fourth position to obtain a real-time comparison result; When the real-time comparison result indicates that the user and / or the space environment has changed, record the data to be corrected that has changed; When the preset reflection surface is idle, based on the data to be corrected, adjust the three-dimensional model of the person and the three-dimensional model of the space in the preset reflection surface.

7. An image generation device, characterized in that, It includes: A first information acquisition module, configured to, when the user wears an extended reality helmet, use the extended reality helmet to acquire first images of multiple angles of the space environment where the user is located, second images of multiple parts of the user, the first position of the user in the space environment, and the second position of the preset reflection surface in the space environment; A space model generation module, configured to generate a three-dimensional model of the space environment based on each of the first images; A person model generation module, configured to generate a three-dimensional model of the user based on each of the second images; A model mapping module, configured to generate imaging images of the three-dimensional model of the space and the three-dimensional model of the person in the preset reflection surface according to the first position and the second position.

8. The device according to claim 7, characterized in that, The space model generation module is specifically configured to: Generate spatial point cloud data of the space environment based on each of the first images; Use a trained space model to perform inference based on the spatial point cloud data to obtain a three-dimensional surface model of the space environment, where the trained space model is trained based on sample space environment data with three-dimensional model ground truth; Perform texture mapping on the three-dimensional surface model of the space environment according to each of the first images to obtain the three-dimensional model of the space.

9. The device according to claim 7, characterized in that, The device further includes: A facial feature acquisition module, configured to use the surface muscle sensor of the extended reality helmet to acquire the surface muscle contraction characteristics of the user's face; The person model generation module is specifically configured to: Generate the image features of the user based on each of the second images, where each of the second images includes images of the user's eyes obtained by a camera located at the eye part of the extended reality helmet, images of the user's face, front body, and both hands obtained by a camera located at the bottom of the extended reality helmet, and images of the user's both hands on both sides of the body obtained by cameras located on both sides of the extended reality helmet, and the image features include the image features of the user's eyes, face, front body, both hands, and both hands on both sides of the body; Align the image features with the surface muscle contraction features to obtain the fusion data of the user; Use the trained human model to perform inference based on the fusion data to obtain the three-dimensional surface model of the user, where the trained human model is a multi-modal model including a full-body human model and a facial surface muscle model, and is trained based on sample human data with three-dimensional model ground truth; Perform texture mapping on the three-dimensional surface model of the user according to each of the second images to obtain the three-dimensional human model.

10. The device according to claim 7, characterized in that, The model mapping module includes: A grid data acquisition sub-module for calculating the spatial grid data of the position to be displayed of the spatial environment in the preset reflection surface and the human grid data of the position to be displayed of the user in the preset reflection surface according to the first position and the second position; A model mapping sub-module for mapping the three-dimensional spatial model to the position where the spatial grid data is located in the preset reflection surface, and mapping the three-dimensional human model to the position where the human grid data is located in the preset reflection surface; A model synthesis sub-module for synthesizing the three-dimensional human model and the three-dimensional spatial model.

11. The device according to claim 10, characterized in that, The model synthesis sub-module is specifically used for: Obtain the light source information in the spatial environment according to the first image and the second image; Calculate the lighting information and shadow information of the three-dimensional human model in the three-dimensional spatial model based on the light source information; Perform lighting processing on the three-dimensional human model and the three-dimensional spatial model according to the lighting information, and perform shadow rendering in the three-dimensional spatial model according to the shadow information to obtain the synthesized three-dimensional human model and three-dimensional spatial model.

12. The device according to claim 7, wherein The device further includes: A second information acquisition module for, when the user wears the extended reality helmet and there are the three-dimensional human model and the three-dimensional spatial model in the preset reflection surface, using the extended reality helmet to acquire in real time third images of multiple angles of the spatial environment, fourth images of multiple parts of the user, the third position of the user in the spatial environment, and the fourth position of the preset reflection surface in the spatial environment; An information comparison module for comparing the first image with the third image, the second image with the fourth image, the first position with the third position, and the second position with the fourth position to obtain a real-time comparison result; A data recording module, configured to record the data to be corrected that has changed in the case where the real-time comparison result indicates that the user and / or the spatial environment has changed; A model adjustment module, configured to adjust the three-dimensional human model and the three-dimensional spatial model in the preset reflecting surface based on the data to be corrected when the preset reflecting surface is idle.

13. An extended reality helmet, characterized in that, It includes a plurality of cameras, a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus, and each of the cameras is located at multiple orientations of the extended reality helmet; The camera is configured to collect images of the user and the spatial environment and obtain the position information of the user and the preset reflecting surface; The memory is used to store computer programs; The processor is configured to implement the method steps described in any one of claims 1-6 when executing the programs stored on the memory.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method steps described in any one of claims 1-6.