Method, device, electronic device and storage medium for rendering an image
By segmenting the external and internal images of the vehicle on the head-mounted display device and combining them with inertial data to determine the pose for low-latency rendering, the problem of insufficient image stability and positioning accuracy of head-mounted displays in mobile vehicles is solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2026-03-27
AI Technical Summary
When a head-mounted display device equipped with extended reality technology is located in a mobile vehicle, existing technologies struggle to achieve high-frequency and low-latency pose information rendering, resulting in insufficient image stability and positioning accuracy.
By setting up image sensors and inertial measurement units on a head-mounted display device, the scene image is segmented into external and internal images of the vehicle. The device's pose in the world and vehicle coordinate systems is determined by combining inertial data, and virtual objects are rendered with low latency compensation to synthesize the target virtual image.
It improves the stability and positioning accuracy of virtual objects both inside and outside the vehicle in the target virtual image, reduces the jitter and lag of virtual objects, and enhances the user experience.
Smart Images

Figure CN115100342B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of extended reality, and in particular, to a method, device, electronic device and storage medium for rendering an image. BACKGROUND
[0002] At present, extended reality technology has been widely applied to various fields, such as healthcare, retail, education, social media, entertainment, etc., to enhance user experience.
[0003] Extended reality can include augmented reality, virtual reality, mixed reality, etc., which integrates virtual content obtained by rendering with the physical world, provides an extended reality experience to the user, and allows the user to interact with a real or physical environment enhanced or augmented with virtual content.
[0004] In the related art, when a head-mounted display device equipped with extended reality technology is located in a movable carrier (such as a car), higher frequency and lower latency pose information is usually required to ensure the stability of the rendered image and the accuracy of the positioning. SUMMARY
[0005] The present disclosure provides a method, device, electronic device and storage medium for rendering an image.
[0006] In one aspect of the present disclosure, a method for rendering an image is provided, which is used for a head-mounted display device, the head-mounted display device is provided with an image sensor and an inertial measurement unit, and the head-mounted display device is located inside a movable carrier. The method comprises: segmenting a first image and a second image from a scene image collected by the image sensor, wherein the first image represents a scene located outside the carrier, and the second image represents a scene located inside the carrier; determining a first pose of the head-mounted display device in a world coordinate system based on the first image and inertial data collected by the inertial measurement unit; determining a second pose of the head-mounted display device in a carrier coordinate system based on the second image and the inertial data; rendering and low-latency compensating a virtual object outside the carrier based on the first pose to obtain a first virtual image; rendering and low-latency compensating a virtual object inside the carrier based on the second pose to obtain a second virtual image; combining the first virtual image and the second virtual image into a target virtual image; and sending the target virtual image to a display screen of the head-mounted display device to make the head-mounted display device present the target virtual image.
[0007] Yet another aspect of the embodiments of the present disclosure provides an apparatus for rendering images, for a head-mounted display device, the head-mounted display device being provided with an image sensor and an inertial measurement unit, the head-mounted display device being located inside a movable vehicle, the apparatus comprising: an image segmentation unit configured to segment a first image and a second image from an image collected by the image sensor, wherein the first image represents a scene located outside the vehicle, and the second image represents a scene located inside the vehicle; a first pose determination unit configured to determine a first pose of the head-mounted display device in a world coordinate system based on the first image and inertial data collected by the inertial measurement unit; a second pose determination unit configured to determine a second pose of the head-mounted display device in a vehicle coordinate system based on the second image and the inertial data; a first image rendering unit configured to render a virtual object outside the vehicle based on the first pose and low-latency compensation, to obtain a first virtual image; a second image rendering unit configured to render a virtual object inside the vehicle based on the second pose and low-latency compensation, to obtain a second virtual image; an image synthesis unit configured to synthesize the first virtual image and the second virtual image into a target virtual image; and an image presentation unit configured to send the target virtual image to a display screen of the head-mounted display device, so that the head-mounted display device presents the target virtual image.
[0008] Yet another aspect of the embodiments of the present disclosure provides an electronic device, comprising: a memory configured to store a computer program; and a processor configured to execute the computer program stored in the memory, and the computer program, when executed, implements the method in any of the above embodiments.
[0009] Yet another aspect of the embodiments of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implements the method in any of the above embodiments.
[0010] The technical solutions of the present disclosure are described in further detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings, which form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0012] The present disclosure can be understood moreappreciably with reference to the following detailed description when considered in connection with the accompanying drawings, in which:
[0013] FIG. 1 A scene diagram applicable to the method for rendering images of the present disclosure;
[0014] FIG. 2 A flowchart of an embodiment of the method for rendering images of the present disclosure;
[0015] FIG. 3 Flowchart of determining the second pose in one embodiment of the method for rendering an image of the present disclosure;
[0016] FIG. 4 Flowchart of segmenting the scene image in one embodiment of the method for rendering an image of the present disclosure;
[0017] FIG. 5 Flowchart of segmenting the scene image in yet another embodiment of the method for rendering an image of the present disclosure;
[0018] FIG. 6 Structural diagram of one embodiment of the apparatus for rendering an image of the present disclosure;
[0019] FIG. 7 Structural diagram of one application embodiment of the electronic device of the present disclosure. DETAILED DESCRIPTION
[0020] Various exemplary embodiments of the present disclosure will now be described in detail below with reference to the accompanying drawings. Note that the relative arrangement, numerical expressions, and numerical values of components and steps set forth in these embodiments are not limiting to the scope of the present disclosure unless specifically stated otherwise.
[0021] Those skilled in the art can understand that the terms "first", "second", and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, and neither represent any specific technical meaning nor indicate their inherent logical order.
[0022] It should also be understood that in the embodiments of the present disclosure, "multiple" can refer to two or more, and "at least one" can refer to one, two, or more.
[0023] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, unless specifically limited or given the opposite implication by the context, it can be understood as one or more in general.
[0024] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the front and rear associated objects.
[0025] It should also be understood that the description of various embodiments of the present disclosure focuses on the differences between the various embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.
[0026] It should be understood that the dimensions of the various elements shown in the figures are chosen for convenience only, and do not bear any relationship to actual proportions.
[0027] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the disclosure, its application or uses.
[0028] Techniques, methods, and apparatus known to those of ordinary skill in the relevant art can not be discussed in detail herein, but should be considered within the scope of the disclosure.
[0029] It should be noted that like reference numerals and letters refer to like items throughout the attached drawings, and thus, once certain items are defined in one drawing, further discussion of such items in subsequent drawings need not be repeated.
[0030] Embodiments of the disclosure can be applied to terminal devices, computer systems, servers, and the like electronic devices, which can operate with numerous other general purpose or special purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations that can be suitable for use with terminal devices, computer systems, servers, and the like electronic devices include, but are not limited to: personal computers, servers, thin clients, thick clients, hand-held or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputers, mainframe computers, and distributed cloud computing technology environments that include any of the above systems, and the like.
[0031] Terminal devices, computer systems, servers, and the like electronic devices can be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, which perform particular tasks or implement particular abstract data types. Computer systems / servers can be practiced in distributed cloud computing environments with remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules can be located in local or remote computer system storage media including memory storage devices.
[0032] In the method for rendering images provided by the disclosure, the head-mounted display device is located inside a movable carrier.
[0033] Exemplary system
[0034] FIG. 1 A scenario to which the method for rendering images of the disclosure is applied is shown, as FIG. 1As shown, the vehicle 100 is a carrier, and the wearer of the head-mounted display device 110 is located in the cabin. The head-mounted device is provided with an image sensor and an inertial measurement unit. The image sensor can be a camera. Through the camera, a scene image in the field of view can be collected. At the same time, through the inertial measurement unit, inertial data can be collected. When the processor of the head-mounted device 110 receives the scene image, the scene image can be segmented into a first image representing a scene outside the vehicle and a second image representing a scene inside the vehicle. In combination with the inertial data and the first image, the position and attitude of the head-mounted display device 110 in the world coordinate system 130 are determined to obtain a first pose. In combination with the inertial data and the second image, the position and attitude of the head-mounted display device 110 in the vehicle coordinate system 120 are determined to obtain a second pose. Through the rendering engine preset in the processor, a first virtual image is generated based on the first pose, and a second virtual image is generated based on the second pose. Then, the processor can combine the first virtual image and the second virtual image into a target virtual image, and present the target virtual image to the wearer by the head-mounted display device 110.
[0035] Exemplary method
[0036] The following will be described in conjunction with FIG. 2 The method for rendering images of the present disclosure is exemplarily illustrated. FIG. 2 A flowchart of one embodiment of the present disclosure for rendering images is shown as follows. FIG. 2 As shown, the flow includes the following steps:
[0037] Step 210, segmenting a first image and a second image from the scene image collected by the image sensor.
[0038] Among them, the first image represents a scene located outside the carrier, and the second image represents a scene located inside the carrier.
[0039] In this embodiment, the head-mounted display device is located inside the movable carrier, so that the scene image collected by the image sensor provided on the head-mounted display device can include the scenes inside and outside the carrier at the same time.
[0040] In one specific example, when the wearer of the head-mounted display device (which can be AR glasses, for example) is riding in a car, the camera provided on the head-mounted display device can capture the scene image inside and outside the car and send the collected scene image to the processor built-in in the head-mounted display device. Then, the scene image is segmented by the image processing algorithm preloaded in the processor. The scene image outside the car (i.e. the first image) and the scene image inside the car (i.e. the second image) are segmented from the scene image. As an example, a pre-trained deep learning model can be used to segment the scene image; the depth information of the scene image can also be estimated by using a binocular vision algorithm, and then the outside area and the inside area of the scene image are determined based on the depth information.
[0041] In step 220, a first pose of the head-mounted display device in the world coordinate system is determined based on the first image and inertial data collected by the inertial measurement unit.
[0042] Generally, the head-mounted display device is also provided with an inertial measurement unit (IMU) for acquiring inertial data of the head-mounted display device, which can be used to represent the motion state of the head-mounted display device.
[0043] For example, the inertial measurement unit can include an acceleration sensor and a gyroscope, wherein the acceleration sensor can be a three-axis acceleration sensor, or include three single-axis acceleration sensors for measuring the acceleration of the head-mounted display device along a certain coordinate axis, and the gyroscope is used to measure the attitude angle of the head-mounted display device relative to the three coordinate axes, and the obtained 6DoF (Degree of Freedom) attitude information (including the acceleration and attitude angle of the three coordinate axes) is the inertial data of the head-mounted display device.
[0044] In practice, the acquisition frequency of the inertial measurement unit is relatively high, for example, it can reach 1000Hz, so the delay between the inertial data and the true pose information of the measured object is relatively low; while the acquisition frequency of the image sensor is relatively low, usually between 10Hz-120Hz, so the delay between the image data and the true pose of the measured object is relatively high.
[0045] In this embodiment, the first pose of the head-mounted display device in the world coordinate system is determined based on the first image and the inertial data, and the inertial data and the image data are fused to obtain pose information with relatively low delay. The first pose represents the position and attitude of the head-mounted display device in the world coordinate system, for example, it can include the position and attitude of the head-mounted display device at each time.
[0046] As an example, the processor of the head-mounted display device can fuse the inertial data and the plurality of frames of first images collected continuously in a preset time period to determine the position and attitude of the head-mounted display device in the world coordinate system at each time in the preset time period, and form a pose sequence of the position and attitude at each time, thereby obtaining the first pose of the head-mounted display device.
[0047] For another example, the processor of the head-mounted display device can also construct a visual-inertial odometry (VIO) based on the first image and the inertial data, and fuse the first image and the inertial data in a tightly coupled or loosely coupled manner to determine the first pose of the head-mounted display device.
[0048] Step 230, determining a second pose of the head-mounted display device in the vehicle coordinate system based on the second image and the inertial data.
[0049] In this embodiment, the second pose represents the position and attitude of the head-mounted display device relative to the vehicle. For example, when the wearer of the head-mounted display device is located in the cabin of the car, the vehicle coordinate system can be the vehicle coordinate system of the car, and the processor of the head-mounted display device can fuse the second image and the inertial data to determine the second pose of the head-mounted display device in the vehicle coordinate system.
[0050] Step 240, rendering and low-latency compensation of the virtual object outside the vehicle based on the first pose, to obtain a first virtual image.
[0051] In this embodiment, the virtual object represents an object presented to the wearer by the head-mounted display device, which can be a virtual icon, a virtual image, or text information, and the like. The first virtual image represents an image containing one or more virtual objects outside the vehicle.
[0052] In one specific example, the head-mounted display device can process one or more virtual objects outside the vehicle through a rendering engine to generate the first virtual image. Specifically, first, based on the first pose corresponding to the current time, a pose corresponding to the on-screen time is predicted to obtain a first predicted pose, and one or more virtual objects outside the vehicle are rendered according to the first predicted pose to obtain a rendered virtual image; then, based on the first pose corresponding to the intermediate time between the current time and the on-screen time, the pose corresponding to the on-screen time is predicted to obtain a second predicted pose, and the difference between the first predicted pose and the second predicted pose is determined; thereafter, the rendered virtual image is compensated using the difference between the first predicted pose and the second predicted pose to obtain the first virtual image.
[0053] Step 250, rendering and low-latency compensation of the virtual object inside the vehicle based on the second pose, to obtain a second virtual image.
[0054] In this embodiment, the second virtual image represents an image containing one or more virtual images inside the vehicle, and its generation process is similar to that of the first virtual image in step 240, which will not be described here.
[0055] Step 260, combining the first virtual image and the second virtual image into a target virtual image.
[0056] In this embodiment, the target virtual image represents an image presented to the wearer, which contains virtual objects inside and outside the vehicle.
[0057] Step 270, sending the target virtual image to the display screen of the head-mounted display device to make the head-mounted display device present the target virtual image.
[0058] The method for rendering images in this embodiment can segment a first image representing a scene outside the vehicle and a second image representing a scene inside the vehicle from the scene image collected by the head-mounted display device; determine a first pose of the head-mounted display device in a world coordinate system based on the first image and the inertial data, and determine a second pose of the head-mounted display device in a vehicle coordinate system based on the second image and the inertial data; then, render and low-latency compensate virtual objects outside the vehicle according to the first pose to generate a first virtual image, and render and low-latency compensate virtual objects inside the vehicle according to the second pose to generate a second virtual image; and then, synthesize the first virtual image and the second virtual image into a target virtual image, and present the target virtual image by a display screen of the head-mounted display device. Rendering and low-latency compensating virtual objects outside the vehicle according to the first pose in the world coordinate system and rendering and low-latency compensating virtual objects inside the vehicle according to the second pose in the vehicle coordinate system can anchor virtual objects outside the vehicle and inside the vehicle in different areas in the target virtual image, which helps to improve the stability and positioning accuracy of virtual objects in the target virtual image, and especially when the head-mounted display device and the vehicle have relative motion, can reduce the jitter or lag of virtual objects in the target virtual image, thereby improving the user experience.
[0059] In some optional embodiments of the present embodiment, the above step 260 can further include: performing image synthesis on the first virtual image and the second virtual image to obtain the target virtual image, taking the first virtual image and the second virtual image as background and foreground respectively, wherein the depth of the first virtual image is greater than the depth of the second virtual image.
[0060] In the present embodiment, in order to avoid false occlusion information of virtual objects inside and outside the vehicle in the synthesis process, the depth information of the image is introduced in the synthesis process, a relatively large depth is set for the first virtual image, which is taken as the background; a relatively small depth is set for the second virtual image, which is taken as the foreground, so that the depth of the virtual objects inside the vehicle in the synthesized target virtual image is less than that of the virtual objects outside the vehicle, avoiding the phenomenon of false occlusion.
[0061] Reference is next made to FIG. 3 , FIG. 3 a flowchart illustrating the determination of the second pose in one embodiment of the method for rendering images of the present disclosure, as shown in FIG. 3 , the flowchart includes the following steps:
[0062] Step 310: performing failure detection on the visual-inertial odometer based on the second image and the inertial data.
[0063] In the embodiment, the relative motion state of the carrier and the head-mounted display device can be determined by performing the failure detection on the visual-inertial odometry. The visual-inertial odometry success indicates that the carrier is in a state of being at rest or close to uniform linear motion. The visual-inertial odometry failure indicates that the carrier is in a state of variable speed motion or turning.
[0064] As an example, the processor of the head-mounted display device can first perform the calculation on the second image and the inertial data by using the visual-inertial odometry, and then perform the failure detection on the visual-inertial odometry according to the intermediate data or the result data in the calculation process. For example, when the number of the calculated inliers is reduced too much compared with the number of the inliers in the calculation process, the visual-inertial odometry is determined to fail; when the pose information after the calculation is greater than the size of the carrier, the visual-inertial odometry is determined to fail.
[0065] As another example, the calculation result of the visual-inertial odometry can also be processed by using a pre-trained deep learning model to determine whether the visual-inertial odometry fails.
[0066] In another example, the head-mounted display device can also construct a visual odometry (VO) based on the second image, perform the calculation on the second image by using the visual odometry to obtain a calculation result of the visual odometry, and then determine whether the visual-inertial odometry fails by comparing the calculation result of the visual odometry with the calculation result of the visual-inertial odometry. For example, when the pose data in the two calculation results is too different, the visual-inertial odometry is determined to fail; when the number of the inliers of the visual-inertial odometry is less than a certain number compared with the number of the inliers of the visual odometry, the visual-inertial odometry is determined to fail.
[0067] In some optional embodiments of the embodiment, the visual-inertial odometry is determined to fail when any one of the following preset conditions is met; the visual-inertial odometry is determined to succeed when none of the preset conditions is met.
[0068] The preset conditions can include: the reduction of the number of inliers obtained after the visual-inertial odometry processes the second image and the inertial data relative to the preset number of inliers is greater than a first preset threshold, the preset number of inliers represents one of the following two: the number of inliers in the process of the visual-inertial odometry, the number of inliers obtained after the visual odometry processes the second image; the residual obtained after the visual-inertial odometry calculates is greater than a second preset threshold, the residual includes at least one of the following: the re-projection error of the second image, the residual of the inertial measurement unit; the change of the state parameter of the preset type of the inertial measurement unit after the visual-inertial odometry calculates exceeds a third preset threshold; the change of the pose estimated by the visual-inertial odometry exceeds the size of the carrier; the speed estimated by the visual-inertial odometry exceeds a fourth preset threshold.
[0069] The state parameters may be, for example, inertial measurement unit bias (IMU bias), IMU-camera extrinsic, IMU-camera-time sync, and the like. These state parameters usually do not change abruptly. If the change of one or more of these state parameters exceeds a reasonable range, it indicates that the visual-inertial odometry fails.
[0070] In this embodiment, the visual-inertial odometry can be detected for failure by using various preset conditions, and the visual-inertial odometry can be evaluated from multiple dimensions, which helps to improve the reliability of the failure detection of the visual-inertial odometry.
[0071] In this embodiment, when the visual-inertial odometry succeeds, the second pose can be determined by step 320; and when the visual-inertial odometry fails, the second pose can be determined by steps 330 to 360.
[0072] Step 320: In response to determining that the visual-inertial odometry succeeds, the second image and the inertial data are processed by using the visual-inertial odometry to determine the second pose.
[0073] In this embodiment, when the visual-inertial odometry succeeds, the second image and the inertial data can be processed by using the visual-inertial odometry to determine the second pose. The high-frequency characteristics of the inertial measurement unit can be used to obtain a high-frame-rate second pose, which helps to further reduce the lag of the target virtual image.
[0074] Step 330: In response to determining that the visual-inertial odometry fails, the first attitude angle of the head-mounted display device in the vehicle coordinate system is determined based on the inertial data.
[0075] In this embodiment, the first attitude angle is angle data collected by a gyroscope in the inertial measurement unit.
[0076] Step 340: The second image is processed by using the visual odometry to determine the second attitude angle and the position information of the head-mounted display device in the vehicle coordinate system.
[0077] Step 350: In response to determining that the difference between the first attitude angle and the second attitude angle is less than a preset attitude angle threshold, the second pose is determined based on the first attitude angle and the position information.
[0078] In this embodiment, when the visual-inertial odometry fails, whether the data collected by the gyroscope in the inertial measurement unit is available can be determined by comparing the first attitude angle and the second attitude angle. The attitude angle threshold can be determined by experiment or set according to experience.
[0079] Specifically, if the difference between the first attitude angle and the second attitude angle is less than the preset attitude angle threshold, it indicates that the carrier is in a variable linear motion state. In this case, the first attitude angle collected by the gyroscope and the position information determined by the visual odometry can be fused to determine the second pose.
[0080] In this embodiment, in the case where the visual-inertial odometry fails but the gyroscope is available, high-frequency attitude angle information can be introduced, and the second pose with high frame rate can be generated in combination with the position information provided by the visual odometry, and the accuracy of the second pose can be ensured.
[0081] Step 360, in response to determining that the difference between the first attitude angle and the second attitude angle is greater than or equal to the attitude angle threshold, determining the second pose by a preset manner.
[0082] The preset manner includes one of the following: increasing the acquisition frame rate of the image sensor, and processing the second image by the visual odometry to determine the second pose; processing the second image by the visual odometry to obtain an initial pose, and performing interpolation processing on the initial pose to determine the second pose.
[0083] In this embodiment, if the difference between the first attitude angle and the second attitude angle is greater than or equal to the attitude angle threshold, it indicates that the carrier is in a non-linear motion, i.e., a turning state. In this case, the first attitude angle collected by the gyroscope in the inertial measurement unit of the head-mounted display device is unavailable. At this time, the second image can be calculated by the visual odometry to determine the second pose. In order to improve the frame rate of the second pose, the acquisition frame rate of the image sensor can be increased and / or the calculation result of the visual odometry can be interpolated.
[0084] In this embodiment, in the case where the visual-inertial odometry fails and the gyroscope is unavailable, the second pose with high frame rate can be obtained by the visual odometry, so as to ensure the accuracy of the second pose.
[0085] Reference is next made to FIG. 4 , FIG. 4 shows a flowchart of segmenting a scene image in one embodiment of the method for rendering an image of the present disclosure, as shown in FIG. 4 , the flow includes the following steps:
[0086] Step 410, based on the scene image, determining the depth information of the scene image and the point cloud model corresponding to the scene image.
[0087] Step 420, based on the depth information of the scene image and the point cloud model corresponding to the scene image, segmenting the scene image to determine the image region located outside the carrier and the image region located inside the carrier in the scene image.
[0088] As an example, after the processor of the head-mounted display device receives the scene image collected by the image sensor, the scene image can be processed by using a binocular vision algorithm to obtain depth information and a point cloud model of the scene image. Then, the scene image is segmented based on the depth information, and the region with a depth greater than the size of the vehicle is determined as the region outside the vehicle, and the region with a depth not greater than the size of the vehicle is determined as the region inside the vehicle, to obtain a first segmentation result; at the same time, the point cloud model is registered with the vehicle model, and the region that matches successfully is determined as the region inside the vehicle, and the remaining part is the region outside the vehicle, to obtain a second segmentation result; then, the first segmentation result and the second segmentation result can be fused according to a preset strategy to determine the image regions inside and outside the vehicle in the scene image, for example, a first weight can be set for the image region according to the difference between the depth information of the image and the size of the vehicle, and a second weight can be set for the image region according to the difference between the point cloud model and the vehicle model. When there is a conflict in the same image region in the first segmentation result and the second segmentation result, the first weight and the second weight of the image region can be compared, and then the segmentation result with a larger weight is determined as the segmentation result of the image region.
[0089] Step 430, the image region located outside the vehicle and the image region located inside the vehicle are determined as a first image and a second image respectively.
[0090] In this embodiment, the scene image is segmented in combination with the depth information of the scene image and the point cloud model, which can improve the accuracy of image segmentation.
[0091] In some optional embodiments of this embodiment, the above step 420 can also use FIG. 5 As shown in the flowchart shown in FIG. 5 The flowchart includes the following steps:
[0092] Step 510, based on the depth information of the scene image, the scene image is determined as a first region and a second region.
[0093] The first region is an image region with a depth far exceeding the size of the vehicle, and the second region is an image region other than the first region.
[0094] As an example, the difference between the depth of the image region and the size of the vehicle can be determined first, and then when the difference exceeds a preset depth threshold, it is determined that the depth of the image region far exceeds the size of the vehicle.
[0095] Step 520, a local point cloud model corresponding to the second region is determined from the point cloud model corresponding to the scene image.
[0096] Step 530, based on the registration result of the local point cloud model and the point cloud model of the vehicle, the second region is determined as a third region and a fourth region.
[0097] The third region is an image region corresponding to a point cloud region with a successful match, and the fourth region is an image region corresponding to a point cloud region with a failed match.
[0098] In this embodiment, the point cloud model of the vehicle can be pre-generated, or can be obtained in the following manner: obtaining a plurality of point cloud models respectively corresponding to a plurality of continuous scene images, to obtain a plurality of continuous point cloud models; registering the plurality of continuous point cloud models, and determining a point cloud region with a successful match in the plurality of continuous point cloud models as the point cloud model of the vehicle. By registering the point cloud models corresponding to the plurality of continuous scene images, the point cloud model of the vehicle can be flexibly determined.
[0099] In step 540, the first region and the fourth region are determined as image regions located outside the vehicle, and the third region is determined as an image region located inside the vehicle.
[0100] In FIG. 5 In the flowchart shown, the scene image can be first segmented into the first region and the second region according to the depth information, and then the local point cloud model of the second region is registered with the point cloud model of the vehicle to segment the second region into the third region and the fourth region. On the one hand, segmenting the scene image in combination with the depth information and the point cloud model helps to improve the accuracy of image segmentation; on the other hand, the calculation amount of point cloud registration can be reduced, which helps to improve the efficiency of image segmentation.
[0101] In some optional implementations of the above embodiments, the head-mounted display device is further provided with a depth camera for collecting a depth image and obtaining a return reflectivity; and the above step 210 can further include: registering the depth image with the scene image, mapping the return reflectivity associated with each feature point in the depth image to a matching feature point in the scene image, to determine the return reflectivity corresponding to each feature point in the scene image; and determining the image region located outside the vehicle and the image region located inside the vehicle in the scene image by using the return reflectivity corresponding to each feature point in the scene image.
[0102] In this implementation, the depth camera can be, for example, a ToF (Time of Flight) lens, which can collect a depth image and obtain a return reflectivity. By registering the depth image with the image sensor, the return reflectivity corresponding to each feature point in the scene image can be determined, and then the image region located outside the vehicle and the image region located inside the vehicle in the scene image can be determined by using the return reflectivity corresponding to each feature point.
[0103] For example, a feature point can be characterized by echo reflectivity, and the distance between the feature point and the head-mounted display device can be determined by time of flight. If the distance is greater than the size of the vehicle, the feature point is located in the image region outside the vehicle, otherwise, the feature point is located in the image region inside the vehicle. For another example, when the vehicle is a vehicle, whether the near-infrared light emitted by the depth camera passes through the vehicle window glass can also be determined by the echo reflectivity corresponding to the feature point. If it passes through, it means that the feature point is located in the image region inside the vehicle, otherwise, it is located in the image region outside the vehicle.
[0104] In this embodiment, the depth map and echo reflectivity collected by the depth camera can quickly and accurately segment the scene image.
[0105] Exemplary apparatus
[0106] Reference will be made to the following FIG. 6 , FIG. 6 A structural diagram of one embodiment of an apparatus for rendering an image of the present disclosure is shown, which is used in a head-mounted display device provided with an image sensor and an inertial measurement unit, and is located inside a movable vehicle, as shown in FIG. 6 The apparatus includes: an image segmentation unit 610 configured to segment a first image and a second image from an image collected by the image sensor, wherein the first image represents a scene located outside the vehicle, and the second image represents a scene located inside the vehicle; a first pose determination unit 620 configured to determine a first pose of the head-mounted display device in a world coordinate system based on the first image and inertial data collected by the inertial measurement unit; a second pose determination unit 630 configured to determine a second pose of the head-mounted display device in a vehicle coordinate system based on the second image and the inertial data; a first image rendering unit 640 configured to render and low-latency compensate a virtual object outside the vehicle based on the first pose, to obtain a first virtual image; a second image rendering unit 650 configured to render and low-latency compensate a virtual object inside the vehicle based on the second pose, to obtain a second virtual image; an image synthesis unit 660 configured to synthesize the first virtual image and the second virtual image into a target virtual image; and an image presentation unit 670 configured to send the target virtual image to a display screen of the head-mounted display device, so that the head-mounted display device presents the target virtual image.
[0107] In one embodiment, the image synthesis unit 660 is further configured to: synthesize the first virtual image and the second virtual image as background and foreground, respectively, to obtain the target virtual image, wherein the depth of the first virtual image is greater than the depth of the second virtual image.
[0108] In one of the embodiments, the second pose determination unit 630 further comprises: a detection module configured to perform a failure detection on the visual-inertial odometry based on the second image and the inertial data; and a visual-inertial odometry module configured to perform a processing on the second image and the inertial data by the visual-inertial odometry to determine the second pose in response to a determination that the visual-inertial odometry is successful.
[0109] In one of the embodiments, the second pose determination unit 630 further comprises: a first attitude angle module configured to determine a first attitude angle of the head-mounted display device in the vehicle coordinate system based on the inertial data in response to a determination that the visual-inertial odometry is failed; a second attitude angle module configured to perform a processing on the second image by the visual odometry to determine a second attitude angle and the position information of the head-mounted display device in the vehicle coordinate system; and a pose determination module configured to determine the second pose based on the first attitude angle and the position information in response to a determination that a difference between the first attitude angle and the second attitude angle is less than a preset attitude angle threshold.
[0110] In one of the embodiments, the pose determination module is further configured to: determine the second pose by a preset manner in response to a determination that the difference between the first attitude angle and the second attitude angle is greater than or equal to the attitude angle threshold; and the preset manner comprises one of: increasing a frame rate of the image sensor and performing the processing on the second image by the visual odometry to determine the second pose; performing the processing on the second image by the visual odometry to obtain an initial pose, and performing an interpolation processing on the initial pose to determine the second pose.
[0111] In one of the embodiments, the detection module is further configured to: perform the processing on the second image and the inertial data by the visual-inertial odometry, and determine that the visual-inertial odometry is failed when any one of preset conditions is met; and determine that the visual-inertial odometry is successful when none of the preset conditions is met, wherein the preset conditions comprise: a reduction of an inlier quantity obtained after the processing on the second image and the inertial data by the visual-inertial odometry relative to a preset inlier quantity is greater than a first preset threshold, the preset inlier quantity represents one of: an inlier quantity in a processing process of the visual-inertial odometry, and an inlier quantity obtained after a processing on the second image by the visual odometry; a residual error obtained after the processing by the visual-inertial odometry is greater than a second preset threshold, the residual error comprises at least one of: a re-projection error of the second image, and a residual error of the inertial measurement unit; a variation of a state parameter of a preset type of the inertial measurement unit after the processing by the visual-inertial odometry exceeds a third preset threshold; a variation of a pose estimated by the visual-inertial odometry exceeds a size of the vehicle; and a speed estimated by the visual-inertial odometry exceeds a fourth preset threshold.
[0112] In one of the embodiments, the image segmentation unit 610 further comprises: a depth point cloud module configured to determine, based on the scene image, depth information of the scene image and a point cloud model corresponding to the scene image; a region segmentation module configured to segment the scene image based on the depth information of the scene image and the point cloud model corresponding to the scene image, to determine an image region located outside the vehicle and an image region located inside the vehicle in the scene image; and an image segmentation module configured to determine the image region located outside the vehicle and the image region located inside the vehicle as the first image and the second image respectively.
[0113] In one of the embodiments, the region segmentation module further comprises: a first segmentation sub-module configured to determine, based on the depth information of the scene image, the scene image as a first region and a second region, wherein the first region is an image region with a depth far exceeding the size of the vehicle, and the second region is an image region other than the first region; a local point cloud sub-module configured to determine, from the point cloud model corresponding to the scene image, a local point cloud model corresponding to the second region; a point cloud registration sub-module configured to determine, based on a registration result of the local point cloud model and a point cloud model of the vehicle, the second region as a third region and a fourth region, wherein the third region is an image region corresponding to a point cloud region with a successful match, and the fourth region is an image region corresponding to a point cloud region with a failed match; and a region determination sub-module configured to determine the first region and the fourth region as the image region located outside the vehicle, and to determine the third region as the image region located inside the vehicle.
[0114] In one of the embodiments, the apparatus further comprises a model generation unit configured to: obtain a plurality of continuous point cloud models corresponding to a plurality of continuous scene images respectively, to obtain the plurality of continuous point cloud models; register the plurality of continuous point cloud models, and determine, as the point cloud model of the vehicle, a point cloud region with a successful match in the plurality of continuous point cloud models.
[0115] In one of the embodiments, the head-mounted display device is further provided with a depth camera configured to collect a depth image and obtain echo reflectivity; and the image segmentation unit 610 is further configured to: register the depth image and the scene image, map the echo reflectivity associated with each feature point in the depth image to a matching feature point in the scene image, to determine the echo reflectivity corresponding to each feature point in the scene image; and determine, based on the echo reflectivity corresponding to each feature point in the scene image, the image region located outside the vehicle and the image region located inside the vehicle in the scene image.
[0116] Exemplary electronic device
[0117] In addition, the embodiments of the present disclosure further provide an electronic device, comprising:
[0118] a memory configured to store a computer program;
[0119] a processor configured to execute a computer program stored in the memory, and the computer program, when executed, implements the method for rendering an image according to any one of the embodiments of the present disclosure.
[0120] FIG. 7 FIG. 1 shows a structural schematic diagram of an electronic device according to an application embodiment of the present disclosure. Hereinafter, an electronic device according to an embodiment of the present disclosure will be described with reference to FIG. 1. FIG. 7 As shown in FIG. 1, the electronic device includes one or more processors and a memory. FIG. 7
[0121] The processor can be a central processing unit (CPU) or other form of processing unit having data processing and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.
[0122] The memory can include one or more computer program products, which can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache, and / or the like. The non-volatile memory, for example, can include read-only memory (ROM), hard disk, flash memory, and / or the like. One or more computer program instructions can be stored on the computer-readable storage media, and the processor can run the program instructions to implement the method for rendering an image according to various embodiments of the present disclosure described above and / or other desired functions.
[0123] In one example, the electronic device can further include an input device and an output device, which are interconnected through a bus system and / or other forms of connection mechanism (not shown).
[0124] In addition, the input device can further include, for example, a keyboard, a mouse, and / or the like.
[0125] The output device can output various information to the outside, including determined distance information, direction information, and / or the like. The output device can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and / or the like.
[0126] Of course, in order to simplify, FIG. 7 In FIG. 1, only some of the components in the electronic device related to the present disclosure are shown, and components such as buses, input / output interfaces, and / or the like are omitted. In addition, the electronic device can further include any other appropriate components according to specific application cases.
[0127] In addition to the above method and device, embodiments of the present disclosure can also be a computer program product, which includes computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the method for rendering an image according to various embodiments of the present disclosure described in the above parts of the specification.
[0128] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.
[0129] In addition, embodiments of the present disclosure can also be a computer readable storage medium, which stores computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the method for rendering an image according to various embodiments of the present disclosure described in the above parts of the specification.
[0130] The computer readable storage medium can take the form of one or more combinations of any type of computer readable medium. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can include, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0131] Those of ordinary skill in the art can understand that all or part of the steps of the above method embodiments can be completed by program instruction related hardware, and the aforementioned program can be stored in a computer readable storage medium, and the program executes the steps of the above method embodiments when executed; and the aforementioned storage medium includes ROM, RAM, magnetic disc or optical disc and various storage medium that can store program code.
[0132] The above describes the basic principles of the present disclosure in conjunction with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the present disclosure are merely examples and are not limiting, and these advantages, benefits, effects and the like cannot be considered as necessary for each embodiment of the present disclosure. In addition, the above specific details are only for the purpose of example and understanding, and are not limiting, and the above details do not limit the present disclosure to be necessarily implemented with the above specific details.
[0133] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between each embodiment can be mutually referred to. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0134] The block diagrams of the devices, apparatuses, equipment, systems involved in the present disclosure are only exemplary examples and are not intended to require or imply the connection, arrangement, configuration shown in the block diagram. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, which mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.
[0135] The methods and devices of the present disclosure can be implemented in many ways. For example, the methods and devices of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present disclosure are not limited to the above specific description, unless otherwise specifically described. In addition, in some embodiments, the present disclosure can also be implemented as programs recorded in recording media, which include machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers the recording media storing the programs for executing the method according to the present disclosure.
[0136] It should also be noted that in the devices, equipment and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of the present disclosure.
[0137] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0138] The above description has been presented for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the disclosure to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those of ordinary skill in the art will appreciate a variety of modifications, alternatives, permutations, additions, and sub-combinations of the described aspects and features.
Claims
1. A method for rendering an image for a head-mounted display device, the head-mounted display device having an image sensor and an inertial measurement unit, the head-mounted display device being located inside a movable vehicle, the method comprising: A first image and a second image are segmented from scene images acquired by an image sensor mounted on the head-mounted display device. The scene images include both the scene outside and inside the vehicle. The first image represents the scene outside the vehicle, and the second image represents the scene inside the vehicle. Based on the first image and the inertial data collected by the inertial measurement unit, the first pose of the head-mounted display device in the world coordinate system is determined; Based on the second image and the inertial data, the second pose of the head-mounted display device in the vehicle coordinate system is determined; Based on the first pose, the virtual objects outside the vehicle are rendered and low-latency compensation is performed to obtain the first virtual image; Based on the second pose, the virtual objects inside the vehicle are rendered and low-latency compensation is performed to obtain a second virtual image; The first virtual image and the second virtual image are combined into a target virtual image; The target virtual image is sent to the display screen of the head-mounted display device so that the head-mounted display device displays the target virtual image.
2. The method according to claim 1, wherein, Combining the first virtual image and the second virtual image into a target virtual image includes: The first virtual image and the second virtual image are used as the background and foreground, respectively. The first virtual image and the second virtual image are then combined to obtain the target virtual image, wherein the depth of the first virtual image is greater than the depth of the second virtual image.
3. The method according to claim 1 or 2, wherein, Determining the second pose of the head-mounted display device in the vehicle coordinate system based on the second image and the inertial data includes: Based on the second image and the inertial data, a failure detection is performed on the visual inertial odometry. In response to the successful determination of the visual inertial odometry, the second image and the inertial data are processed using the visual inertial odometry to determine the second pose.
4. The method according to claim 3, wherein, Determining the second pose of the head-mounted display device in the vehicle coordinate system based on the second image and the inertial data further includes: In response to determining that the visual inertial odometry has failed, a first attitude angle of the head-mounted display device in the vehicle coordinate system is determined based on the inertial data; The second image is processed using visual odometry to determine the second attitude angle and position information of the head-mounted display device in the vehicle coordinate system; In response to determining that the difference between the first attitude angle and the second attitude angle is less than a preset attitude angle threshold, the second pose is determined based on the first attitude angle and the position information.
5. The method according to claim 4, wherein, Determining the second pose of the head-mounted display device in the vehicle coordinate system based on the second image and the inertial data further includes: In response to determining that the difference between the first attitude angle and the second attitude angle is greater than or equal to the attitude angle threshold, the second pose is determined by a preset method; The preset method includes one of the following: Increase the frame rate of the image sensor and process the second image using the visual odometry to determine the second pose; The second image is processed using the visual odometry to obtain an initial pose, and the initial pose is interpolated to determine the second pose.
6. The method according to any one of claims 3 to 5, wherein, Based on the second image and the inertial data, failure detection of the visual inertial odometry is performed, including: The visual inertial odometry is used to process the second image and the inertial data. When any one of the preset conditions is met, the visual inertial odometry is determined to have failed. If none of the preset conditions are met, the visual inertial odometry is determined to be successful. The preset conditions include: the decrease in the number of inliers obtained by the visual inertial odometry (VIO) after processing the second image and the inertial data relative to a preset number of inliers is greater than a first preset threshold, where the preset number of inliers represents one of the following two: the number of inliers during the VIO processing, or the number of inliers obtained by the VIO after processing the second image; the residual obtained by the VIO calculation is greater than a second preset threshold, where the residual includes at least one of the following: the reprojection error of the second image, or the residual of the inertial measurement unit (IMU); the change in the state parameter of the IMU of a preset type after calculation by the VIO exceeds a third preset threshold; the pose change estimated by the VIO exceeds the size of the vehicle; and the velocity estimated by the VIO exceeds a fourth preset threshold.
7. The method according to any one of claims 1 to 6, wherein, Segmenting a first image and a second image from scene images acquired by the image sensor includes: Based on the scene image, determine the depth information of the scene image and the point cloud model corresponding to the scene image; Based on the depth information of the scene image and the point cloud model corresponding to the scene image, the scene image is segmented to determine the image region located outside the vehicle and the image region located inside the vehicle in the scene image; The image region located outside the vehicle and the image region located inside the vehicle are respectively designated as the first image and the second image.
8. The method according to claim 7, wherein, Based on the depth information of the scene image and the point cloud model corresponding to the scene image, the scene image is segmented to determine the image regions located outside the vehicle and the image regions located inside the vehicle, including: Based on the depth information of the scene image, the scene image is determined as a first region and a second region, wherein the first region is an image region whose depth far exceeds the size of the vehicle, and the second region is an image region outside the first region; Determine the local point cloud model corresponding to the second region from the point cloud model corresponding to the scene image; Based on the registration results of the local point cloud model and the point cloud model of the vehicle, the second region is determined as the third region and the fourth region, wherein the third region is the image region corresponding to the successfully matched point cloud region, and the fourth region is the image region corresponding to the unmatched point cloud region. The first region and the fourth region are defined as image regions located outside the vehicle, and the third region is defined as an image region located inside the vehicle.
9. The method according to claim 8, wherein, The method further includes the step of obtaining a point cloud model of the vehicle: Obtain the point cloud models corresponding to multiple consecutive scene images to obtain the multi-frame continuous point cloud model; The multi-frame continuous point cloud model is registered, and the successfully matched point cloud region in the multi-frame continuous point cloud model is determined as the point cloud model of the vehicle.
10. The method according to any one of claims 1 to 7, wherein, The head-mounted display device is also equipped with a depth camera for capturing depth images and obtaining echo reflectivity; Segmenting a first image and a second image from scene images acquired by the image sensor includes: The depth image and the scene image are registered, and the echo reflectance associated with each feature point in the depth image is mapped to the matching feature point in the scene image to determine the echo reflectance corresponding to each feature point in the scene image. By utilizing the echo reflectivity corresponding to each feature point in the scene image, the image regions located outside the vehicle and the image regions located inside the vehicle in the scene image are determined.
11. An apparatus for rendering an image for a head-mounted display device, the head-mounted display device having an image sensor and an inertial measurement unit, the head-mounted display device being located inside a movable vehicle, the apparatus comprising: An image segmentation unit is configured to segment a first image and a second image from scene images acquired by an image sensor mounted on the head-mounted display device, wherein the scene images simultaneously include scenes outside and inside the vehicle, the first image representing the scene located outside the vehicle, and the second image representing the scene located inside the vehicle; The first pose determination unit is configured to determine the first pose of the head-mounted display device in the world coordinate system based on the first image and the inertial data collected by the inertial measurement unit. The second pose determination unit is configured to determine the second pose of the head-mounted display device in the vehicle coordinate system based on the second image and the inertial data. The first image rendering unit is configured to render and perform low-latency compensation on virtual objects outside the vehicle based on the first pose to obtain a first virtual image. The second image rendering unit is configured to render and perform low-latency compensation on virtual objects inside the vehicle based on the second pose to obtain a second virtual image; An image synthesis unit is configured to synthesize the first virtual image and the second virtual image into a target virtual image; An image rendering unit is configured to send the target virtual image to the display screen of the head-mounted display device so that the head-mounted display device renders the target virtual image.
12. An electronic device, comprising: Memory, used to store computer programs; A processor for executing a computer program stored in the memory, wherein when the computer program is executed, it implements the method described in any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any one of claims 1-10.
Citation Information
Patent Citations
Display system, electronic apparatus, mobile body, and display method
JP2021110920A
Systems and methods for providing immersive extended reality experiences on moving platforms
US10767997B1