Image synthesis method, image synthesis device, storage medium, and electronic device
By recording and utilizing camera pose change information, the background image is corrected to match the real image perspective, solving the problem of poor synthetic image quality caused by pose changes in extended reality technology and improving the user experience.
Patent Information
- Application Number
- CN202310764999.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-26
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2043-06-26
AI Technical Summary
In extended reality technology, the misalignment between the background image and the captured real image caused by changes in camera pose results in poor quality of the synthesized image, which affects the user experience.
By acquiring and recording the camera's pose change information, and using the first and second delay durations, the camera's pose at different time points is determined, the background image is corrected to match the real image perspective, and the image is composited.
It improves the visual coherence and realism of synthetic images, enhancing the user experience of extended reality technology.
Smart Images

Figure CN116645309B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of extended reality technology, and more specifically, to an image synthesis method, an image synthesis apparatus, a storage medium, and an electronic device. Background Technology
[0002] Extended reality (AR) is a technology that combines the virtual and real worlds. It overlays virtual objects onto real-world scenes, allowing users to see a blended image of the virtual and real worlds. Dynamic tracking and real-time synchronization are crucial technologies in AR applications, ensuring that virtual objects remain synchronized with the real-world environment, thereby enhancing user experience and interactivity.
[0003] When the camera moves, the poses of the real image captured by the camera and the rendered image used as the background image are different. That is, the perspectives of the real image and the rendered image are different. During compositing, the real image and the rendered image cannot be aligned, resulting in tearing of the composite image and thus a poor composite image effect. Summary of the Invention
[0004] The purpose of this disclosure is to provide an image synthesis method, an image synthesis apparatus, a storage medium, and an electronic device to improve the effect of synthesized images.
[0005] To achieve the above objectives, a first aspect of this disclosure provides an image synthesis method applied to a processing device in a virtual reality device, the virtual reality device further including a display device and a first camera, the image synthesis method comprising:
[0006] The first delay duration between the first time and the third time, and the second delay duration between the second time and the third time are obtained respectively. The first time is the time when the original background image is projected onto the screen model to obtain the screen image. The second time is the time when the first camera takes a picture of the display device displaying the screen image to obtain the first real image. The third time is the time when the processing device receives the first real image.
[0007] Based on the first delay duration, the third pose of the first camera at the third moment, and the recorded pose queue, the first pose of the first camera when the original background image is projected onto the screen model is determined. The pose queue includes the third pose of the first camera at the third moment and the pose of each historical moment before the third moment.
[0008] Based on the second delay duration, the third pose of the first camera at the third moment, and the pose queue, the second pose of the first camera when acquiring the first real image is determined;
[0009] Based on the first pose and the second pose, the original background image is modified to obtain the target background image;
[0010] The target background image, the first real image, and the screen mask image are combined to obtain a composite image.
[0011] A second aspect of this disclosure provides an image synthesis apparatus, applied to a processing device in a virtual reality device, the virtual reality device further including a display device and a first camera, the image synthesis apparatus comprising:
[0012] The first acquisition module is used to acquire the first delay duration between the first time and the third time, and the second delay duration between the second time and the third time, wherein the first time is the time when the original background image is projected onto the screen model to obtain the screen image, the second time is the time when the first camera takes a picture of the display device displaying the screen image to obtain the first real image, and the third time is the time when the processing device receives the first real image.
[0013] The first determining module is used to determine the first pose of the first camera when the original background image is projected onto the screen model based on the first delay duration, the third pose of the first camera at the third time, and the recorded pose queue. The pose queue includes the third pose of the first camera at the third time and the pose of each historical time before the third time.
[0014] The second determining module is used to determine the second pose of the first camera when acquiring the first real image based on the second delay duration, the third pose of the first camera at the third moment, and the pose queue.
[0015] The correction module is used to correct the original background image based on the first pose and the second pose to obtain the target background image;
[0016] The compositing module is used to composite the target background image, the first real image, and the screen masking image to obtain a composite image.
[0017] A third aspect of this disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the image synthesis method described in the first aspect of this disclosure.
[0018] A fourth aspect of this disclosure provides an electronic device, comprising:
[0019] A memory on which computer programs are stored;
[0020] A processor is configured to execute the computer program in the memory to implement the steps of the image synthesis method of the first restaurant of this disclosure.
[0021] By employing the above technical solution, based on the first delay duration, the second delay duration, the third pose of the first camera at the third moment, and the pose queue, the first pose of the first camera when projecting the original background image onto the screen model and the second pose of the first camera when acquiring the first real image are determined. Then, the original background image is corrected based on the first and second poses to obtain the target background image. Finally, the target background image, the first real image, and the screen mask image are used to synthesize the composite image. Thus, by considering the changes in the first camera pose and correcting the original background image to obtain the target background image, the target background image can be aligned with the first real image, making the composite image visually more coherent and realistic, improving the application experience of extended reality technology and the effect of the composite image.
[0022] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0023] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings:
[0024] Figure 1 This is a schematic diagram illustrating a virtual reality device according to an exemplary embodiment.
[0025] Figure 2 This is a schematic diagram illustrating the workflow of a virtual reality device according to an exemplary embodiment.
[0026] Figure 3 This is a flowchart illustrating an image synthesis method according to an exemplary embodiment.
[0027] Figure 4 This is a block diagram illustrating an image synthesis apparatus according to an exemplary embodiment.
[0028] Figure 5 This is a block diagram illustrating an electronic device according to an exemplary embodiment.
[0029] Figure 6 This is a block diagram illustrating another electronic device according to an exemplary embodiment. Detailed Implementation
[0030] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.
[0031] First, it should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are performed in accordance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the corresponding device. Second, it should be noted that the pose of the first camera described in this disclosure refers to the pose of the first camera relative to the screen model.
[0032] Current extended reality technologies primarily focus on the time delay between the background image and the captured real image during image synthesis, without fully considering the impact of camera pose changes on the synthesized image, especially when there are significant pose changes between the background image and the captured real image. In practical applications, due to changes in camera pose (such as rotation and translation), the alignment between the background image and the captured real image may differ significantly. This can lead to inaccurate positions of virtual objects in the background image, reducing the realism and stability of extended reality applications and resulting in a poor user experience.
[0033] To improve the user experience of extended reality technology and the effect of synthesized images, this disclosure provides an image synthesis method, an image synthesis apparatus, a storage medium, and an electronic device.
[0034] Before describing the image synthesis method provided in this disclosure in detail, the virtual reality device to which this image synthesis method is applicable will first be described.
[0035] Figure 1 This is a schematic diagram illustrating a virtual reality device according to an exemplary embodiment. Figure 1 As shown, the virtual reality device 100 may include a processing device 101, a display device 102, and a first camera 103. The processing device 101 is connected to both the display device 102 and the first camera 103. The display device 102 displays a screen image A generated based on the original background image. The first camera 103 is a real camera used to capture images of the display device 102 displaying the screen image A to obtain a first real image. The shooting range of the first camera 103 is as follows... Figure 1 The range between the two dashed lines.
[0036] In addition, virtual reality devices may also include a scene camera, an OGRE camera, and a screen model. The processing device is connected to both the scene camera and the OGRE camera. The scene camera captures and renders a frame from the virtual scene; this frame can be called the original background image. The OGRE camera projects the original background image captured and rendered by the scene camera onto the screen model in the virtual 3D space to obtain a screen image, which is then displayed on the real-world device. In other words, the OGRE camera can determine the content displayed on the display device. The first camera is a real camera used to capture real images of the real world. For example, the processing device controls the first camera to capture an image of the display device showing the screen image, thus obtaining a first real image.
[0037] It should be understood that when the scene camera acquires and renders the original background image, the processing device obtains the pose P0 of the first camera and then controls the scene camera to acquire and render the original background image while in pose P0. Similarly, when the OGRE camera projects the original background image onto the screen model, the processing device obtains the pose P1 of the first camera and then controls the OGRE camera to project the original background image onto the screen model while in pose P1 to generate a screen image. Afterwards, the processing device controls the display device to display the screen image and controls the first camera to capture a first real image on the display device. Then, the acquisition card sends the first real image back to the processing device. The processing device records the pose of the first camera when it receives the first real image, as well as the pose of the first camera at historical moments, to ensure that the pose data from historical moments can be used during the image synthesis stage.
[0038] Figure 2 This is a schematic diagram illustrating the workflow of a virtual reality device according to an exemplary embodiment. Figure 2 As shown, taking a single frame as an example, firstly, at time T0, the scene camera renders the image to obtain the original background image. Then, at time T1, the OGRE camera projects the original background image onto the screen model to obtain the screen image. After that, at time T2, the display device displays the screen image and controls the first camera to take a picture. Finally, at time T3, the acquisition card transmits the first real image captured by the first camera back to the processing device.
[0039] It should be understood that in practical applications, the scene camera, the OGRE camera, and the first camera operate in parallel. That is, during the time period from T0 to T3, the scene camera continuously performs image rendering, the OGRE camera continuously performs projection, and the first camera continuously performs shooting.
[0040] The image synthesis method provided in this disclosure will now be described in detail.
[0041] Figure 3 This is a flowchart illustrating an image synthesis method according to an exemplary embodiment. The method is applied to a processing device in a virtual reality device, which also includes a display device and a first camera. Figure 3 As shown, the image synthesis method may include the following steps.
[0042] In step S31, the first delay duration between the first time and the third time, and the second delay duration between the second time and the third time are obtained respectively.
[0043] Specifically, the first moment is the moment when the original background image is projected onto the screen model to obtain the screen image; the second moment is the moment when the first camera takes a picture of the display device showing the screen image to obtain the first real image; and the third moment is the moment when the processing device receives the first real image. Furthermore, the original background image can be a picture rendered by the scene camera at any moment prior to the first moment.
[0044] In this disclosure, the screen model provides the position and shape information of the virtual screen, ensuring that the composite image matches the position and shape of the virtual screen, thereby achieving accurate display and natural blending of the composite image.
[0045] It should be understood that the third moment can allow the processing device to perform actions such as... Figure 3 The current moment of the image synthesis method shown, that is, the processing device can begin executing the process upon receiving the first real image returned, as described above. Figure 3 The image synthesis method shown.
[0046] In step S32, the first pose of the first camera when the original background image is projected onto the screen model is determined based on the first delay duration, the third pose of the first camera at the third moment, and the recorded pose queue.
[0047] In this disclosure, the third pose of the first camera at a third moment refers to the current pose of the first camera when the processing device receives the first real image. Therefore, the processing device acquires the third pose of the first camera at a third moment when it receives the first real image.
[0048] In addition, the processing device records the pose at the current moment of receiving the first real image and the pose at historical moments. That is, the processing device records the third pose of the first camera at the third moment and the pose at every historical moment before the third moment.
[0049] For example, the processing device finds the first pose of the first camera when projecting the original background image onto the screen model in the pose queue based on the first delay duration and the third pose at the third time point. That is, the first pose of the first camera at the first time point. For example, assuming the third pose corresponding to the third time point t3 is p3 and the first delay duration is Δt1, then the pose corresponding to the time point t3-Δt1 is determined as the first pose of the first camera.
[0050] In step S33, the second pose of the first camera when acquiring the first real image is determined based on the second delay duration, the third pose of the first camera at the third moment, and the pose queue.
[0051] The method for determining the second pose is similar to that for determining the first pose, and will not be elaborated upon here.
[0052] It should be understood that steps S32 and S33 can be executed simultaneously, or step S32 can be executed first and then step S33, or step S33 can be executed first and then step S32; this disclosure does not limit this. Figure 3 The example is to execute step S32 first and then step S33.
[0053] In step S34, the original background image is modified according to the first pose and the second pose to obtain the target background image.
[0054] It should be understood that the screen image displayed on the display device is obtained by projecting the original background image onto the screen when the OGRE camera is in the first pose, while the first real image is the image captured by the first camera when it is in the second pose. Therefore, the viewing angle when the original background image is projected to obtain the screen image does not match the viewing angle when the first camera captures the first real image. When compositing the original background image and the first real image, the first real image and the original background image cannot be aligned, resulting in tearing of the composite image and thus a poor composite image effect.
[0055] In this disclosure, the original background image is modified using a first pose and a second pose to obtain the target background image. Modifying the original background image means transforming the viewing angle of the projected background image from the first pose to the second pose; that is, the visual perspective of the target background image is the second pose. Thus, the viewing angle of the target background image is consistent with the viewing angle when the first camera captures the first real image, enabling alignment between the first real image and the target background image.
[0056] In step S35, the target background image, the first real image, and the screen mask image are composited to obtain a composite image.
[0057] Step S34 can obtain a target background image that is aligned with the first real image. Then, based on the target background image, the first real image, and the screen mask image, the image is composited to obtain a composite image, avoiding tearing problems in the composite image and improving the effect of the composite image.
[0058] For example, the target background image, the first real image, and the screen mask image can be composited using the following formula: P_target i =V_mask i ·P_background i +(1-V_mask i )·P_real i ; where P_target i V_mask represents the pixel value of the i-th pixel in the synthesized image. i The pixel value representing the i-th pixel in the screen mask image. Each pixel value in the screen mask image is used to represent the weight of whether the position corresponding to that pixel value is the target background area or the first real image area, and its value range is [0,1]; P_background i P_real represents the pixel value of the i-th pixel in the target background image. i The pixel value of the i-th image in the first real image is represented, where i is an integer with a value range of [1, N], and N is the total number of pixel values in the synthesized image. The total number of pixel values of the target background image, the first real image, and the screen mask image is also N.
[0059] By employing the above technical solution, based on the first delay duration, the second delay duration, the third pose of the first camera at the third moment, and the pose queue, the first pose of the first camera when projecting the original background image onto the screen model and the second pose of the first camera when acquiring the first real image are determined. Then, the original background image is corrected based on the first and second poses to obtain the target background image. Finally, the target background image, the first real image, and the screen mask image are used to synthesize the composite image. Thus, by considering the changes in the first camera pose and correcting the original background image to obtain the target background image, the target background image can be aligned with the first real image, making the composite image visually more coherent and realistic, improving the application experience of extended reality technology and the effect of the composite image.
[0060] In the first embodiment, the first delay duration is determined as follows: First, the OGRE camera projects a completely white frame onto the screen model, and this moment is recorded as the first moment. When the OGRE camera is not projecting, a completely black frame is displayed on the screen model. Next, the projected white frame is displayed on the display device. Then, the first camera captures a real image from the display device and transmits this real image back via a capture card. Finally, the processing device records the moment the real image transmitted back from the capture card as the third moment, and the difference between the third moment and the first moment is determined as the first delay duration.
[0061] In another embodiment, the first delay duration is determined as follows: First, for each video image frame in the video stream carrying frame number information, the video image frame is sequentially projected onto the screen model to obtain a second screen imaging image. The second screen imaging image is displayed on the display device, and the first camera is controlled to take a picture of the display device displaying the second screen imaging image to obtain a second real image. Then, when the second real image is received, the frame number information of the target video image frame included in the second real image is determined, and the first delay duration is determined based on the frame number information of the target video image frame and the frame number information of the video image frame currently projected onto the screen model.
[0062] For example, assuming that the frame number information of the target video image frame included in the second real image indicates that the target video image frame is the second frame in the video stream, and the frame number information of the video image frame currently projected onto the screen model when the second real image is received indicates that the currently projected video image frame is the tenth frame in the video stream, then the sum of the frame lengths of each image frame from the second frame to the tenth frame can be determined as the first delay duration.
[0063] In one embodiment, the second delay duration is determined as follows: First, during the continuous movement of the first camera, real-time poses and timestamps are recorded. The first camera also acquires real images in real-time during the movement and transmits these images back to the processing device via a capture card. Upon receiving a real image, the processing device can record the moment when any frame of the real image is received as the third moment. Then, for each recorded pose, the screen model is converted into a screen edge mesh using that pose, resulting in multiple screen edge meshes. Finally, the multiple screen edge meshes are matched with screen regions in the real images received at the third moment. The moment corresponding to the pose of the screen edge mesh with the highest matching degree is determined as the second moment, and the difference between the third moment and the second moment is determined as the second delay duration.
[0064] In another embodiment, the second delay duration is determined as follows: First, while the first camera is in a moving state, a fourth pose of the first camera is acquired upon receiving a third real-world image captured by the first camera. The first camera can be mounted on a gimbal, and its movement is controlled by moving the gimbal. Next, based on the fourth pose, the intrinsic parameters of the first camera, and the screen model, a screen edge mesh representing the screen model region is added to the third real-world image. Then, the screen mesh delay parameter is adjusted so that the screen model region represented by the screen edge mesh matches the screen region in the third real-world image. Finally, the screen mesh delay parameter at which the screen model region represented by the screen edge mesh matches the screen region in the real-world image is determined as the second delay duration.
[0065] Because the first camera is in motion, its pose when capturing the third real image differs from its fourth pose when the processing device receives the third real image. Therefore, the screen model region represented by the screen edge mesh added to the third real image does not coincide with the screen region in the third real image. To make them coincide, the screen mesh delay parameter can be adjusted. For example, the screen mesh delay parameter can be adjusted by a fixed value. For instance, a fixed value of 0.5 is used, and the screen mesh delay parameter is increased or decreased by 0.5 each time until the screen model region represented by the screen edge mesh matches the screen region in the real image. The adjustment then stops, and the current screen mesh delay parameter is determined as the second delay duration.
[0066] Thus, the first delay duration and the second delay duration can be determined using the methods provided in the above embodiments.
[0067] It should be understood that the first delay duration and the second delay duration can be during execution. Figure 3 The image synthesis method shown is predetermined, and for a virtual reality device, the first delay duration and the second delay duration only need to be calculated once. Subsequently, when executing the image synthesis method, the first delay duration and the second delay duration can be directly obtained. Alternatively, for a virtual reality device, the first delay duration and the second delay duration can be calculated once before each image synthesis operation; this disclosure does not specify a particular method for this.
[0068] The following describes a specific implementation method for modifying the original background image to obtain the target background image.
[0069] In one embodiment, Figure 3The specific implementation of step S34, which modifies the original background image based on the first pose and the second pose to obtain the target background image, is as follows: First, the screen model is expanded so that the first OGRE camera can completely project the original background image onto the expanded screen model when it is in the first pose; then, the pose of the second OGRE camera is adjusted to the second pose, and when the second OGRE camera is in the second pose, the second OGRE camera is controlled to take a picture of the screen model with the original background image projected on it; finally, the image captured by the second OGRE camera is determined as the target background image.
[0070] In this embodiment, the virtual reality device includes at least two OGRE cameras. The first OGRE camera is used to project the original background image onto the expanded screen model in the first pose. The second OGRE camera is used to capture the screen model with the original background image projected on it in the second pose. Thus, the viewing angle of the image captured by the second OGRE camera is the same as the viewing angle of the first real image captured by the first camera. That is, the image captured by the second OGRE camera is the target background image.
[0071] Furthermore, the virtual reality device can have one or more screen models. In this embodiment, one screen model can be expanded, or all screen models can be expanded; this disclosure does not impose specific limitations in this regard. Additionally, when expanding a screen model, expansion can be performed along its edges. For example, if the screen model consists of multiple spliced screen models, expansion is performed along edges other than the splicing edges.
[0072] In another embodiment, Figure 3 The specific implementation of step S34, which modifies the original background image based on the first pose and the second pose to obtain the target background image, is as follows: First, determine the target screen model. The target screen model can be one of multiple screen models, or at least two screen models located on the same plane among multiple screen models. Specifically, when the target screen model is one of multiple screen models, the specific implementation of determining the target screen model can be: determining the area of the original background image projected onto each screen model, and determining the screen model with the largest area as the target screen model.
[0073] Next, based on the first pose, the second pose, and the intrinsic parameters of the first camera, the first image pixel coordinates and the second image pixel coordinates corresponding to the vertices of the target screen model are determined, respectively. For example, the first image pixel coordinates corresponding to the vertices of the target screen model are determined based on the first pose and the intrinsic parameters of the first camera, and the second image pixel coordinates corresponding to the vertices of the target screen model are determined based on the second pose and the intrinsic parameters of the first camera. The intrinsic parameters of the first camera can be determined through camera calibration. It should be understood that determining the image pixel coordinates corresponding to vertices based on pose and camera intrinsic parameters is a relatively mature technique, and this disclosure will not elaborate further on it.
[0074] Finally, the original background image is corrected based on the pixel coordinates of the first image and the pixel coordinates of the second image to obtain the target background image.
[0075] For example, the projection transformation relationship can be determined based on the pixel coordinates of the first image and the pixel coordinates of the second image. Then, the original background image can be corrected based on the projection transformation relationship to obtain the target background image.
[0076] In this disclosure, the original background image can be modified in the manner described in any of the above embodiments to obtain a target background image that can be aligned with the first real image.
[0077] After determining the second pose, the image synthesis method may further include:
[0078] The screen mask image generated by the first OGRE camera in the second pose is acquired. During image synthesis, the viewpoints of the target background and the first real image are both from the second pose. To further improve the quality of the synthesized image, the screen mask image from the second pose viewpoint is also used during synthesis. That is, after determining the second pose, the screen mask image generated by the first OGRE camera in the second pose is acquired.
[0079] The screen mask image is processed to obtain the processed target screen mask image. This processing can include dilation and feathering. In the screen mask image, the pixel value representing the background screen area is 0.0, and the pixel value representing the real image screen area is 1.0. Dilation reduces the size of the real image screen area, and then feathering gradually reduces the weight of the edge portions of the real image screen area from 1.0 to 0.0.
[0080] The target background image, the real image, and the target screen mask image are composited to obtain a composite image.
[0081] The above technical solution acquires a screen mask image generated when the first OGRE camera is in the second pose, and then processes the screen mask image before image synthesis. This ensures that the screen mask image, the target background image, and the first real image have the same viewpoint, further improving the visual coherence and realism of the synthesized image. Furthermore, processing the screen mask image before image synthesis makes the fusion of the real image and the background image in the synthesized image more natural and smooth, improving the quality and realism of the synthesized image.
[0082] Based on the same inventive concept, this disclosure also provides an image synthesis apparatus. Figure 4 This is a block diagram illustrating an image compositing apparatus according to an exemplary embodiment. The image compositing apparatus is applied to a processing device in a virtual reality device, which further includes a display device and a first camera. Figure 4 As shown, the image synthesis apparatus 400 may include:
[0083] The first acquisition module 401 is used to acquire the first delay duration between the first moment and the third moment, and the second delay duration between the second moment and the third moment, wherein the first moment is the moment when the original background image is projected onto the screen model to obtain the screen imaging image, the second moment is the moment when the first camera takes a picture of the display device displaying the screen imaging image to obtain the first real image, and the third moment is the moment when the processing device receives the first real image.
[0084] The first determining module 402 is used to determine the first pose of the first camera when the original background image is projected onto the screen model based on the first delay duration, the third pose of the first camera at the third time, and the recorded pose queue. The pose queue includes the third pose of the first camera at the third time and the pose of each historical time before the third time.
[0085] The second determining module 403 is used to determine the second pose of the first camera when acquiring the first real image based on the second delay duration, the third pose of the first camera at the third moment, and the pose queue.
[0086] The correction module 404 is used to correct the original background image based on the first pose and the second pose to obtain the target background image;
[0087] The compositing module 405 is used to composite the target background image, the first real image, and the screen mask image to obtain a composite image.
[0088] Optionally, the correction module 404 may include:
[0089] An extension submodule is used to extend the screen model so that the first OGRE camera can project the original background image completely onto the extended screen model when it is in the first pose.
[0090] The control submodule is used to adjust the pose of the second OGRE camera to the second pose, and when the second OGRE camera is in the second pose, control the second OGRE camera to take pictures of the screen model on which the original background image is projected.
[0091] The first determining submodule is used to determine the image captured by the second OGRE camera as the target background image.
[0092] Optionally, the correction module 404 may include:
[0093] The second determining submodule is used to determine the target screen model, wherein the target screen model is one of the multiple screen models, or at least two screen models located on the same plane among the multiple screen models;
[0094] The third determining submodule is used to determine the first image pixel coordinates and the second image pixel coordinates corresponding to the vertices of the target screen model based on the first pose, the second pose and the intrinsic parameters of the first camera.
[0095] The correction submodule is used to correct the original background image based on the pixel coordinates of the first image and the pixel coordinates of the second image to obtain the target background image.
[0096] Optionally, the target screen model is one of the multiple screen models, and the second determining submodule is used to: determine the area of the original background image projected onto each of the screen models; and determine the screen model with the largest area as the target screen model.
[0097] Optionally, after determining the second pose, the image synthesis device 400 further includes:
[0098] The second acquisition module is used to acquire the screen mask image generated by the first OGRE camera when it is in the second pose.
[0099] The processing module is used to process the screen masking image to obtain the processed target screen masking image;
[0100] The compositing module 405 is used to: composite the target background image, the real image, and the target screen mask image to obtain a composite image.
[0101] Optionally, the first delay duration is determined in the following way:
[0102] For each video image frame in the video stream carrying frame number information, the video image frame is sequentially projected onto the screen model to obtain a second screen imaging image. The second screen imaging image is displayed on the display device, and the first camera is controlled to take a picture of the display device displaying the second screen imaging image to obtain a second real image.
[0103] Upon receiving the second real image, the frame number information of the target video image frame included in the second real image is confirmed, and the first delay duration is determined based on the frame number information of the target video image frame and the frame number information of the video image frame currently projected onto the screen model.
[0104] Optionally, the second delay duration is determined in the following way:
[0105] When the first camera is in a moving state, the fourth pose of the first camera is obtained when the third real image captured by the first camera is received;
[0106] Based on the fourth pose, the intrinsic parameters of the first camera, and the screen model, a screen edge mesh is added to the third real image to represent the screen model region.
[0107] The screen grid delay parameter is adjusted so that the screen model region represented by the screen edge grid is consistent with the screen region in the third real image;
[0108] The screen mesh delay parameter when the screen model region represented by the screen edge mesh matches the screen region in the real image is determined as the second delay duration.
[0109] Regarding the image synthesis apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0110] Figure 5 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Figure 5 As shown, the electronic device 700 may include a processor 701 and a memory 702. The electronic device 700 may also include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705.
[0111] The processor 701 controls the overall operation of the electronic device 700 to complete all or part of the steps in the image synthesis method described above. The memory 702 stores various types of data to support the operation of the electronic device 700. This data may include, for example, instructions for any application or method operating on the electronic device 700, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 703 may include a screen and audio components. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 702 or transmitted via communication component 705. The audio component also includes at least one speaker for outputting audio signals. I / O interface 704 provides an interface between processor 701 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0112] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the image synthesis method described above.
[0113] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the image compositing method described above. For example, the computer-readable storage medium may be the memory 702 including the program instructions described above, which may be executed by the processor 701 of the electronic device 700 to complete the image compositing method described above.
[0114] Figure 6 This is a block diagram illustrating another electronic device according to an exemplary embodiment. For example, electronic device 1900 may be provided as a server. (Refer to...) Figure 6 The electronic device 1900 includes a processor 1922, which may be one or more, and a memory 1932 for storing computer programs executable by the processor 1922. The computer program stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processor 1922 may be configured to execute the computer program to perform the image compositing method described above.
[0115] Additionally, the electronic device 1900 may also include a power supply component 1926 and a communication component 1950. The power supply component 1926 can be configured to perform power management of the electronic device 1900, and the communication component 1950 can be configured to enable communication of the electronic device 1900, such as wired or wireless communication. Furthermore, the electronic device 1900 may also include an input / output (I / O) interface 1958. The electronic device 1900 can operate on an operating system stored in memory 1932.
[0116] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the image compositing method described above. For example, the non-transitory computer-readable storage medium may be the memory 1932 including the program instructions described above, which may be executed by the processor 1922 of the electronic device 1900 to complete the image compositing method described above.
[0117] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described image synthesis method when executed by the programmable device.
[0118] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0119] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0120] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. An image compositing method characterized by, The application discloses a processing device applied to a virtual reality device, the virtual reality device further comprising a display device and a first camera, and the image synthesis method comprises the following steps: respectively acquiring a first delay duration between a first time and a third time, and a second delay duration between a second time and the third time, wherein the first time is a time when an original background picture is projected onto a screen model to obtain a screen imaging image, the second time is a time when the first camera captures the display device displaying the screen imaging image to obtain a first real image, and the third time is a time when the processing device receives the first real image; determining a first pose of the first camera when the original background picture is projected onto the screen model according to the first delay duration, a third pose of the first camera at the third time, and a recorded pose queue, wherein the pose queue comprises the third pose of the first camera at the third time and a pose at each historical time before the third time; determining a second pose of the first camera when the first real image is collected according to the second delay duration, the third pose of the first camera at the third time, and the pose queue; correcting the original background picture according to the first pose and the second pose to obtain a target background picture; performing picture synthesis on the target background picture, the first real image and a screen mask image to obtain a synthesis image.
2. The image compositing method of claim 1, wherein, The method for correcting the original background picture according to the first pose and the second pose to obtain a target background picture comprises the following steps: extending the screen model so that a first OGRE camera can project the original background picture on the extended screen model completely when the first OGRE camera is in the first pose; adjusting a pose of a second OGRE camera to the second pose, and controlling the second OGRE camera to capture the screen model on which the original background picture is projected when the second OGRE camera is in the second pose; determining an image captured by the second OGRE camera as the target background picture.
3. The image compositing method of claim 1, wherein, The method for correcting the original background picture according to the first pose and the second pose to obtain a target background picture comprises the following steps: determining a target screen model, wherein the target screen model is one of a plurality of screen models, or is at least two screen models in the same plane in the plurality of screen models; determining first image pixel coordinates and second image pixel coordinates corresponding to vertices of the target screen model according to the first pose, the second pose and an intrinsic parameter of the first camera; correcting the original background picture according to the first image pixel coordinates and the second image pixel coordinates to obtain a target background picture.
4. The image compositing method of claim 3, wherein, The target screen model is one of a plurality of screen models, and the method for determining the target screen model comprises the following steps: determining a region of the original background picture projected onto each screen model; determining a screen model with the largest region as the target screen model.
5. The image compositing method of claim 1, wherein, After the second pose is determined, the image synthesis method further includes: obtaining a screen mask image generated by the first OGRE camera when in the second pose; processing the screen mask image to obtain a processed target screen mask image; the screen mask image is processed to obtain a processed target screen mask image; the screen mask image is processed to obtain a processed target screen mask image; 6. The image compositing method of claim 1, wherein, the screen mask image is processed to obtain a processed target screen mask image; the first delay duration is determined by: for each video image frame in the video stream carrying frame number information, sequentially project the video image frame onto the screen model to obtain a second screen imaging image, display the second screen imaging image on the display device, and control the first camera to capture the display device displaying the second screen imaging image to obtain a second real image; 7. The image compositing method of claim 1, wherein, when the second real image is received, determine the frame number information of the target video image frame included in the second real image, and determine the first delay duration according to the frame number information of the target video image frame and the frame number information of the video image frame currently projected onto the screen model. the second delay duration is determined by: when the first camera is in a moving state, when the third real image captured by the first camera is received, the fourth pose of the first camera is obtained; according to the fourth pose, the intrinsic parameter of the first camera and the screen model, a screen edge grid for representing the screen model region is added in the third real image; adjust the screen grid delay parameter so that the screen model region represented by the screen edge grid is consistent with the screen region in the third real image; 8. An image compositing apparatus characterized by comprising: the screen grid delay parameter when the screen model region represented by the screen edge grid is consistent with the screen region in the third real image is determined as the second delay duration. The processing device applied to the virtual reality device further includes a display device and a first camera, and the image synthesis device includes: a first acquisition module for acquiring a first delay duration between a first time and a third time, and a second delay duration between a second time and the third time, wherein the first time is the time when the original background picture is projected onto the screen model to obtain the screen imaging image, the second time is the time when the first camera captures the display device displaying the screen imaging image to obtain the first real image, and the third time is the time when the processing device receives the first real image; a first determination module for determining the first pose of the first camera when projecting the original background picture onto the screen model according to the first delay duration, the third pose of the first camera at the third time, and the recorded pose queue, wherein the pose queue includes the third pose of the first camera at the third time and the pose at each historical time before the third time; a second determination module, configured to determine a second pose of the first camera when the first real image is captured according to the second delay duration, a third pose of the first camera at the third time, and the pose queue; a correction module, configured to correct the original background picture according to the first pose and the second pose to obtain a target background picture; a synthesis module, configured to perform picture synthesis on the target background picture, the first real image, and a screen mask image to obtain a synthesis image.
9. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the image synthesis method in any one of claims 1-7.
10. An electronic device, comprising: comprise: a memory having a computer program stored thereon; a processor configured to execute the computer program in the memory to implement the steps of the image synthesis method in any one of claims 1-7.
Citation Information
Patent Citations
Image display method and device, readable medium and electronic equipment
CN112486318A
Distributed augmented reality image display and generation of composite images
US11138806B1