Human-centered vehicle metaverse platform for roadside AR / VR content

DE102024112593A1Pending Publication Date: 2025-09-11GM GLOBAL TECHNOLOGY OPERATIONS LLC +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024112593
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-07
Filing Date
2024-05-06
Publication Date
2025-09-11

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In an exemplary embodiment, a system is provided that includes one or more sensors and a processor. The one or more sensors are configured to receive sensor data for a vehicle. The processor is coupled to the one or more sensors and configured to at least enable selection of a physical target using the sensor data, including a physical element proximate the vehicle while the vehicle is moving; generate virtual content based on the sensor data; and deliver the virtual content to one or more users of the vehicle as an overlay for the physical target.
Need to check novelty before this filing date? Find Prior Art

Description

INTRODUCTION

[0001] The technical field relates generally to vehicles and, more specifically, to methods and systems for delivering content to users in a vehicle.

[0002] Vehicle users (e.g., passengers) can now perceive roadside content, such as billboards and the like, at the roadside. However, roadside content is not always presented to users in an optimal manner.

[0003] Accordingly, it is desirable to provide improved methods and systems for delivering content to users of a vehicle. Furthermore, other desirable features and characteristics of the present disclosure will become apparent from the following detailed description and the appended claims, taken in conjunction with the accompanying drawings and the foregoing technical field and background. DESCRIPTION

[0004] According to an exemplary embodiment, a method is disclosed comprising: obtaining sensor data from one or more sensors of a vehicle; selecting, via a processor using the sensor data, a physical target comprising a physical element proximate the vehicle while the vehicle is moving; generating, via the processor, virtual content based on the sensor data; and delivering the virtual content to one or more users of the vehicle as an overlay or in an augmented manner with respect to the physical target.

[0005] Also in an exemplary embodiment, the sensor data includes image sensor data, location sensor data, and inertial measurement unit (IMU) sensor data obtained from the vehicle; the method further includes generating, via the processor, a head pose and a vehicle pose based on the image sensor data, the location sensor data, and the IMU sensor data obtained from the vehicle; and generating the virtual content is performed using the head pose and the vehicle pose.

[0006] In an exemplary embodiment, generating the virtual content is further performed using a three-dimensional image database stored in a computer memory in combination with the head and vehicle pose.

[0007] In an exemplary embodiment, the virtual content is displayed to one or more users of the vehicle on one or more display devices in the vehicle.

[0008] In an exemplary embodiment, the physical target comprises a billboard along the roadway on which the vehicle is traveling, and the virtual content is overlaid over the billboard or displayed in an augmented manner relative to the billboard as displayed in the vehicle to one or more users.

[0009] In an exemplary embodiment, the physical target comprises a traffic control device (TCD) along the roadway on which the vehicle is traveling, and the virtual content is overlaid over, or displayed in an augmented manner with respect to, the TCD as displayed in the vehicle for one or more users.

[0010] In an exemplary embodiment, the physical target comprises a point of interest along the roadway on which the vehicle is traveling, and the virtual content is overlaid with respect to the point of interest in an augmented manner.

[0011] Also in an exemplary embodiment, the physical target comprises a road condition object along a roadway on which the vehicle is traveling; and the virtual content is overlaid over the road condition object or presented in an augmented manner relative to the road condition object as displayed in the vehicle to the one or more users.

[0012] In an exemplary embodiment, the provision of the virtual content uses a level of detail of a three-dimensional (3D) mesh provided to the vehicle from a remote system, and the level of detail is selected based on a distance between the physical target and the vehicle.

[0013] In an exemplary embodiment, the level of detail is further selected based on available network bandwidth or one or more other quality of experience (QOE) metrics.

[0014] In an exemplary embodiment, the method further comprises performing frustum and back-face culling of the 3D mesh by the processor.

[0015] In another exemplary embodiment, a system is provided that includes one or more sensors and a processor. The one or more sensors are configured to receive sensor data for a vehicle. The processor is coupled to the one or more sensors and configured to at least enable selection of a physical target including a physical element proximate the vehicle while the vehicle is moving using the sensor data; generate virtual content based on the sensor data; and provide the virtual content to one or more users of the vehicle as an overlay over the physical target or as a display in an augmented manner relative to the physical target.

[0016] Also in an exemplary embodiment, the one or more sensors comprise a plurality of sensors configured to receive the sensor data, wherein the sensor data includes image sensor data, location sensor data, and inertial measurement unit (IMU) sensor data obtained from the vehicle; and the processor is further configured to enable at least generating a head pose and a vehicle pose based on the image sensor data, the location sensor data, and the IMU sensor data obtained from the vehicle; and generating the virtual content using the head pose and the vehicle pose.

[0017] In an exemplary embodiment, the processor is further configured to enable at least the generation of the virtual content using a three-dimensional image database stored in a computer memory in combination with the head pose and the vehicle pose.

[0018] In an exemplary embodiment, the virtual content is displayed to one or more users of the vehicle on one or more display devices in the vehicle.

[0019] In an exemplary embodiment, the physical target comprises a billboard along the roadway on which the vehicle is traveling, and the virtual content is overlaid over the billboard or displayed in an augmented manner relative to the billboard as displayed in the vehicle to one or more users.

[0020] Also in an exemplary embodiment, the provision of the virtual content utilizes the level of detail of a three-dimensional (3D) mesh provided to the vehicle from a remote system; and the level of detail is selected based on a distance between the physical target and the vehicle.

[0021] In an exemplary embodiment, the level of detail is also based on available network bandwidth or one or more other quality of experience (QOE) metrics.

[0022] In an exemplary embodiment, the processor is further configured to facilitate at least frustum and back-face culling of the 3D mesh by the processor.

[0023] In another exemplary embodiment, a system is provided that includes a plurality of sensors, a processor, and a display device. The plurality of sensors is configured to receive sensor data for a vehicle, wherein the sensor data includes image sensor data, location sensor data, and inertial measurement unit (IMU) sensor data received from the vehicle.The processor is coupled to the plurality of sensors and configured to, using the sensor data, at least facilitate selection of a physical target including a physical element proximate the vehicle while the vehicle is moving; generate a head pose and a vehicle pose based on the image sensor data, the location sensor data, and the IMU sensor data obtained from the vehicle; generate virtual content using the head pose and the vehicle pose; and provide instructions for delivering the virtual content to one or more users of the vehicle as an overlay on the physical target or in an augmented manner with respect to the physical target.The device is configured to display the virtual content overlaid on the physical target or otherwise enhanced with respect to the virtual content in the vehicle to the one or more users according to instructions provided by the processor. DESCRIPTION OF THE DRAWINGS

[0024] The present specification will now be described in conjunction with the following drawings, in which like numerals refer to like elements: Fig. 1 provides an overview of a platform and process 100 for delivering content to users of a vehicle, including overlaying or augmenting virtual content with a physical element proximate the vehicle, according to an exemplary embodiment; and Fig. Figure 2 is a functional block diagram of a system for providing content to users of a vehicle used to implement the method of Fig. 1 can be used, according to exemplary embodiments; Fig. Figure 3 shows an example implementation of content generated by the system of Fig. 2 about the platform and the process of Fig. 1, according to an exemplary embodiment; Fig. Figure 4 shows an exemplary implementation of the selection of a physical target for the overlaid or enhanced display of the virtual content in conjunction with the platform and method of Fig. 1, according to an exemplary embodiment; Fig. Figure 5 shows an exemplary implementation of the selection of a level of detail for the virtual content in conjunction with the platform and method of Fig. 1, according to an exemplary embodiment; and Fig. 6 and Fig. 7 show an exemplary implementation of the culling of the virtual content with the platform and the process of Fig. 1, according to an exemplary embodiment. DETAILED DESCRIPTION

[0025] The following detailed description is merely exemplary and is not intended to limit the disclosure or its application and uses. Furthermore, there is no intention to be bound by any theory presented in the foregoing background or the following detailed description.

[0026] Fig. 1 provides an overview of a platform and process 100 for delivering content to users of a vehicle 110, including overlaying or augmenting virtual content with a physical element proximate the vehicle, according to an exemplary embodiment. According to an exemplary embodiment, the platform and process 100 are also used in conjunction with the system 200 of Fig. 2 and various implementations of the Fig. 3-7 implemented as described below.

[0027] As in Fig. 1, the platform and process 100 (collectively referred to as "process 100" for convenience) may be used in conjunction with multiple vehicles 110. In various embodiments, the process 100, for each vehicle 110, utilizes sensor data (including location, IMU, and image data) from the vehicle 110 in generating virtual content (e.g., augmented reality or virtual reality) that is overlaid on one or more physical elements (e.g., a billboard, road sign, or other traffic control device (TCD), a point of interest, road condition objects, or one or more other roadside elements) or provided in an augmented manner for display to one or more users of the vehicle 110, as described in more detail below.

[0028] In various embodiments, each vehicle 110 comprises an automobile. The vehicle 110 may be any number of different types of automobiles, such as a sedan, station wagon, truck, or sport utility vehicle (SUV), and may have two-wheel drive (2WD) (i.e., rear-wheel drive or front-wheel drive), four-wheel drive (4WD), or all-wheel drive (AWD), and / or various other types of vehicles in certain embodiments. In certain embodiments, the vehicle 110 may also include a motorcycle or other vehicle, such as an aircraft, spacecraft, watercraft, etc., and / or one or more other types of mobile platforms (e.g., a robot and / or other mobile platform).

[0029] As in Fig. 1, in various embodiments, sensor data 112 from each vehicle 110 is transmitted to a remote control center 114 for processing. In certain embodiments, the sensor data 112 includes location data (e.g., including Global Positioning System (GPS) data), IMU (Inertial Measurement Unit) sensor data, and image data (e.g., from a driver or passenger monitoring system or an image of a vehicle front view) from the vehicle 110. In various embodiments, the sensor data 112 is received from a sensor array 210 of the vehicle 110 (in Fig. 2 and described below in connection therewith) and sent to the remote center 114. In certain embodiments, the center 114 also includes one or more remote servers (e.g., the Fig. 2 and described further below in this context), which are physically remote from the vehicle 110.

[0030] As in Fig. 1, in various embodiments, the sensor data 112 is provided to a localization module 116. Also in various embodiments, the localization module 116 provides information from the sensor data 122 (e.g., in various embodiments, the raw sensor data 122 and / or filtered sensor data 122) to a cloud service 120 for localization and receives a vehicle position 124 from the cloud service 120. In various embodiments, the cloud service 120 and other remote communications are conducted via one or more wireless communication networks 206 (as in Fig. 2).

[0031] In various embodiments, the localization module 116 also generates pose data 126 based on both the vehicle pose 124 and the sensor data 122, and provides the pose data 126 to a delivery module 118 for further processing. In various embodiments, the pose data 126 provided by the localization module 116 to the delivery module 118 includes both a vehicle pose (e.g., a position and rotation of the vehicle 110) and a head pose (e.g., a position and rotation of the head of a driver and / or one or more other users of the vehicle 110). In various embodiments, the head pose also includes a gaze of the user (e.g., the driver), such as a gaze direction of a user from one or more of the user's eyes.

[0032] In various embodiments, the localization module 116 generates the pose data 126 with six degrees of freedom (DOF) in a world coordinate system (WCS). In an exemplary embodiment, the WCS corresponds to a global positioning system (GPS) coordinate system. In various embodiments, the localization module 116 uses camera images (e.g., from the sensor data 112) together with a pre-built structure-from-motion (SfM) model (e.g., as a model stored in computer memory 232 and / or 252 of Fig. 2 stored program) database images from the computer memory (e.g. from one or more stored values ​​242 and / or 262 of Fig. 2 and / or as otherwise stored in the cloud). In one example, the SfM model generates a three-dimensional (3D) point cloud detachable from two-dimensional (2D) image feature points. It should be noted that in various embodiments, the cloud-side 3D map database may be constructed using techniques other than SfM, including one or more advanced map scanning techniques that generate 3D meshes.

[0033] Furthermore, in certain embodiments, the head pose and vehicle pose of pose data 126 are generated by (i) image retrieval, (ii) feature matching, (iii) lifting, and (iv) pose estimation. In certain embodiments, image retrieval includes retrieving relevant database images based on the query image via an exhaustive search or a nearest neighbor search. In certain embodiments, matching also includes two-dimensional (2D) feature extraction and feature matching between retrieved images (e.g., from the database) and the camera image. Also in certain embodiments, lifting includes lifting two-dimensional (2D) images into three-dimensional (3D) images, including searching for 3D points corresponding to 2D features from the SfM point cloud.Additionally, also in certain embodiments, the pose estimation comprises a regression of a six degrees of freedom (6 DOF) pose from 2D-to-3D pairs via Perspective-n-Point (PnP) and a Random Sample Consensus (RANSAC) algorithm.

[0034] In various embodiments, camera calibration is also performed (e.g., with respect to one or more cameras of a driver monitoring system) to ensure that the head pose is provided in conjunction with the WCS. In certain example embodiments, a transformation is performed (e.g., via a transformation matrix) to convert the head pose from a coordinate system of the driver monitoring system (DMS) to the WCS (e.g., using a transformation matrix). In various embodiments, the vehicle pose is also converted to the WCS in a similar manner.

[0035] In certain embodiments, for example, the head position is detected by a driver monitoring system (DMS) as H d generated and the vehicle position is shown in the WCS as V w received, both as inputs. In various embodiments, calibration is performed as a relative transformation between the vehicle and the strain gauge as follows T r .

[0036] In various embodiments, the head position in the WCS is output as H w Generated. In various embodiments, the head pose in the strain gauge is converted to the head pose in the WCS via the vehicle pose in the WCS (as mentioned above) and the calibration between the strain gauge and the vehicle according to the following equation: Hw=Vw⊙(Tr⊙Hd) where "⊙" represents a matrix multiplication. In various embodiments, the head pose in the strain gauge coordinates can be considered a transformation matrix between the head position and the strain gauge, so that the head pose in the WCS is calculated by applying two transformation matrices (one between the head position and the strain gauge, and the other between the strain gauge and the vehicle). In various embodiments, the relative transformation between the vehicle and the strain gauge is also determined during the calibration step.

[0037] With further reference to Fig. 1, in various embodiments, the pose data 126 (including the head pose and the vehicle pose) is provided to the delivery module 118 as mentioned above. In various embodiments, the target selection 128 is performed using the pose data 126. More specifically, in various embodiments, the target selection 128 results in a physical target 129 of one or more physical elements along the roadside for overlaying virtual content for viewing by one or more users within the vehicle and / or for displaying the virtual content in an augmented manner relative to the one or more physical elements.

[0038] In various embodiments, the target selection uses as inputs the camera images, along with the head pose and target position in the WCS, and an aerial imagery network. Also in various embodiments, the target selection 128 selects as output one or more physical elements along the roadway on which the vehicle is traveling to overlay the virtual content over the physical element and / or to display the virtual content augmented with respect to the one or more physical elements. In one embodiment, the target comprises a billboard along the roadway. In other embodiments, the target may comprise any number of other types of physical elements along the roadway, such as one or more traffic control devices (e.g., traffic signs, traffic lights, crosswalks, etc.), one or more points of interest (e.g.,, restaurants, gas stations, service stations, hotels, shops, travel destinations and / or other points of interest, e.g., to provide augmented reality (AR)), and / or one or more road condition objects (e.g., potholes, slippery roads, traffic events, gestures, etc.) and / or other physical roadside elements.

[0039] As in Fig. 3, in an exemplary embodiment, the target 129 comprises an advertising billboard located along the roadway on which the vehicle 110 is traveling. As shown in Fig. 3, in one exemplary embodiment, a display 140 is provided to users in the vehicle 110 such that the display 140 includes an overlay of virtual content over the target 129. As mentioned above, in various embodiments, the display 140 may include an overlay of the virtual content over the physical element and / or for displaying the virtual content in an augmented manner with respect to the one or more physical elements. As mentioned previously, in various embodiments, in addition to a billboard along the roadway, the target 129 may include one or more traffic control devices (TCDs) (e.g., traffic signs, traffic lights, crosswalks, and so on), one or more points of interest (e.g., restaurants, gas stations, service stations, hotels, shops, travel destinations, and / or other points of interest, e.g., for providing augmented reality), and / or one or more road condition objects (e.g.,B. potholes, slippery roads, traffic events, gestures, etc.) and / or other physical elements at the roadside, among other possible physical elements.

[0040] As in Fig. 4, in an exemplary embodiment, the target 129 is selected from several other potential roadside targets. In various embodiments, as part of the target selection process, an image from an aerial image network is first rendered with the head pose (including position and orientation) in the WCS. In various embodiments, a 2D target 129 is selected according to the position of the target 129 in the WCS. Furthermore, in various embodiments, corresponding potential targets are detected and selected between several images, such as those shown in Fig. 4. In various embodiments, the matching is also accelerated by comparing the 2D bounding box of the target 129. In certain embodiments, a contour detection algorithm is used to describe the bounding box. Next, in various embodiments, the selected target 129 is differentiated among other potential targets 401, 402 in the second image 400.

[0041] In various embodiments, the target 129 is selected based on its proximity to the vehicle 110 and / or one or more other specific characteristics, such as whether the target 129 includes an empty space or area or is otherwise suitable for overlaying virtual content onto the target 129 and / or providing content in an enhanced manner relative to the target 129, and / or whether the target 129 includes a theme suitable for overlaying virtual content and / or providing virtual content in an enhanced manner (e.g., if the target 129 includes text corresponding to traffic rules or regulations that could be enlarged or otherwise enhanced for display to users within the vehicle, and / or if the target 129 includes content for information, advertising, or the like that could be enhanced with an overlay of virtual content, etc.).

[0042] With reference to Fig. 1, in various embodiments, a video database 130 is created. In various embodiments, the video database 130 takes as inputs information related to the vehicle and its users, such as a user's point of interest, vehicle posture, and / or various other information. In various embodiments, the video database 130 includes as output a sequence of three-dimensional (3D) images to be delivered to the user. In various embodiments, the video database 130 includes virtual 3D content to be provided to the user in the vehicle as an overlay over the selected physical target 129 and / or in another enhanced manner from the target selection 128. In various embodiments, the video database 130 also includes a sequence of 3D mesh frames.In various embodiments, each mesh has a geometry that includes a collection of vertices and faces, as well as texture information that is combined with the geometry to convolve the mesh frames.

[0043] With further reference to Fig. 1, in various embodiments, the 3D video database 130 is combined with the selected target 129 of the target selection 128 for the level of detail (LOD) selection 132. In various embodiments, the LOD selection 132 refers to a level of detail of the 3D mesh to be provided to the user, including whether a simplified version of the 3D mesh would be appropriate under certain circumstances. In various embodiments, the LOD selection 132 includes as inputs the position of the vehicle and the availability of the communications network bandwidth (including for providing virtual 3D content to the vehicle) and / or one or more other quality metrics (e.g., including potential delays, image instability, channel reliability, etc.).In various embodiments, an output of the LOD selection 132 also includes a frame of simplified 3D mesh images corresponding to a selected LOD.

[0044] In various embodiments, the selected LOD reduces the complexity and detail of the 3D mesh depending on the distance from the user. For example, in certain embodiments, when the 3D mesh is relatively close to the physical target 129, the LOD may request a 3D mesh with greater detail (e.g., with a greater number of triangles) to increase a reasonable standard of quality of the 3D mesh for user viewing. Conversely, also in certain embodiments, when the physical target 129 is relatively far from the user, the LOD may request a 3D mesh with relatively less detail (e.g., with a fewer number of triangles) to reduce the complexity of the 3D mesh for user viewing.In certain embodiments, the selected LOD may also require a 3D mesh with more details when the bandwidth is sufficiently large, and a 3D mesh with less details when the bandwidth is relatively small, etc. In various embodiments, similar adjustments may also be made with respect to one or more other QOE metrics (e.g., including possible delays, image instability, channel reliability, etc.).

[0045] Accordingly, in certain embodiments, the selection of the LOD may be based on two main rules, namely: (1) a distance-based rule that selects the LOD based on the physical distance between the user and the physical destination; and (2) a bandwidth-based rule that selects the LOD based on the network bandwidth. In various embodiments, as a result: (i) a relatively abstract LOD is selected when the user is (1) far from the destination or (2) the available network bandwidth is insufficient. In various embodiments, the system prepares 3D videos for each LOD using a mesh decimation algorithm and stores the various 3D videos in the 3D video database 130. In various embodiments, similar adjustments may also be made with respect to one or more other QOE metrics (e.g., potential delays, image instability, channel reliability, etc.).

[0046] In Fig. 5 shows a flowchart for an exemplary implementation of the LOD selection 132 of Fig. 1. As shown in Fig. 5, a vehicle pose is provided in 502 and implemented in conjunction with a distance-based rule in 504 to generate a first selected LOD “A” in 506. In addition, also as in Fig. 5, the network bandwidth 508 is implemented in conjunction with a bandwidth-based LOD rule (and / or one or more other QOE metrics) at 510 to generate a second selected LOD “B” at 512.

[0047] With further reference to Fig. 5, in various embodiments, the first selected LOD "A" of 506 and the second selected LOD "B" of 512 are considered together at 514. Specifically, in various embodiments, the minimum of the first selected LOD "A" and the second selected LOD at 514 (e.g., the one having the lowest level of detail) is selected and referred to as the selected LOD at 516. In various embodiments, the selected LOD of 516 is integrated with a 3D mesh database at 518 to generate a 3D mesh frame with the selected LOD at 520.

[0048] Back to Fig. 1: In various embodiments, frustum and back-face culling are performed at 134. In various embodiments, frustum and back-face culling is performed by removing the mesh (or a portion thereof) that is outside the viewer's perspective. In various embodiments, back-face culling is also performed by removing certain features of the mesh (e.g., including certain vertices and triangles of the 3D mesh) that are not visible to the user. The terms "viewer" and "user" refer to the vehicle occupants who are intended to view the display.

[0049] In certain embodiments, the frustum and backface extraction are performed by introducing a new data structure to efficiently calculate whether a given triangle lies within the user's frustum, and the extraction is performed with adaptive resolution. Specifically, in certain embodiments, the mesh is converted from a Cartesian coordinate system to an aspherical coordinate system where the origin is the user's position, as follows: (x,y,z) in Cartesian coordinates →(θ,π,r) in spherical coordinates

[0050] In exemplary embodiments, adaptive bin resolution is also implemented. In particular, in certain embodiments, the 3D space in spherical coordinates is divided into bins, where each bin contains at most one visible vertex. In certain embodiments, the division is based on the resolution level of the bins (Resolution bin) to account for the different distances between the 3D object and the user. In particular, in various embodiments, for equal (θ, π) angles, the 3D mesh will (a) contain a relatively smaller number of vertices and faces when the 3D object is relatively close to the origin (e.g., to the user); whereas conversely, (b) the 3D mesh will contain a relatively larger number of vertices and faces when the 3D object is relatively far from the origin (e.g., from the user). Accordingly, in various embodiments, a relatively smaller angle of resolution is implemented when the 3D object is farther from the user's position, so that the bin converges on the same or a similar part of the mesh, while a relatively larger angle of resolution is implemented when the 3D object is closer to the user's position.

[0051] With reference to Fig. 6, in various embodiments, according to a first algorithm, the frustum and backface extraction begins with the calculation of the average distance between neighboring vertices (d avg ) 602 in Ft (e.g. as in the Fig. from Fig. 6) when a frame of the 3D mesh (Ft) is present. In various embodiments, a distance between the target position and the user position is also calculated, namely R sphere 604, as in Fig. 6. Furthermore, in various embodiments, a bin resolution, namely the resolution bin 606, as in Fig. 6, calculated according to the following equation: resolutionbin=2arctan(davg2Rsphere)

[0052] In various embodiments, this method ensures that each bin has at most one vertex at a given resolution value. Furthermore, in various embodiments, after obtaining the resolution bin a corresponding bin index for each vertex in F t assign.

[0053] Furthermore, in certain embodiments, the frustum and backface selection may also be implemented in conjunction with a second algorithm. In this second algorithm, in certain embodiments, a specific value for both F t and resolution bin given, then for each vertex (v i ) in Ft a bin (π i , θ i ) -index of v igenerated using the first algorithm above. In certain embodiments, a "truncated cone read" is also performed and continued if the current bin is outside the user's "truncated cone". In various embodiments, the backface read is also performed based on whether the bin (π i , θ i ) is empty. In various embodiments, in particular, (A) if the field (π i , θ i ) is empty, v i added to the current container; and (B) otherwise, v j used to represent the vertex in the current bin, the distance to the origin is used for both v i as well as for v j compared, and the vertex is updated to be the nearest vertex of the previously visited bin.

[0054] In Fig. 7 is a flowchart for an exemplary embodiment of the truncated cone and back face readout of 134 from Fig. 1. As shown in Fig. 7, an exemplary readout process 700 in various embodiments includes the generation of a 3D mesh in 702. In various embodiments, the 3D mesh of the vehicle pose is used to calculate a bin resolution in 704. In various embodiments, the calculation of 704 results in the Resolution bin 606 for the resolution, as described above.

[0055] As also in Fig. 7, in various embodiments, the bin resolution 606 is combined with the 3D mesh 710 and the head position 712 for the readout 708, including the truncated cone readout 714 and the back face readout 716, as described above. In various embodiments, the readout 708 results in a readout 3D mesh. Furthermore, as shown in Fig. 7, in various embodiments, the readout 708 is used in conjunction with a representation of the 3D mesh in spherical coordinates 720 to generate a vertex “i” (with x, y, and z coordinates) at 722. As in Fig. 7, in various embodiments, a spherical transformation is then performed at 724, generating a bin index (θ, π) at 726.

[0056] In various embodiments, at 728, it is determined whether the bin index (θ, π) is within the user's head position. If at 728, it is determined that the bin index is not within the user's head position, the process returns to point 722 described above.

[0057] If, on the other hand, it is determined in step 728 that the bin index is within the user's head pose, the process proceeds to step 730, where it is determined whether the bin index (θ, π) already has a vertex. If, on the other hand, it is determined in step 720 that the bin index (θ, π) does not have a vertex, in various embodiments the current vertex is added to the bin at 732, after which the process returns to step 722 described above. If, on the other hand, it is determined in step 720 that the bin index (θ, π) already contains a vertex, in various embodiments the bin is updated accordingly at 734, after which the process returns to step 722 described above.

[0058] Back to Fig. 1: In various embodiments, additional compression is performed at 136. In various embodiments, after the sorting at 134, further compression is performed at 136, including encoding and decoding, to further reduce the size of the 3D mesh to be transmitted to the vehicle. In certain embodiments, the further compression at 134 comprises using a list of vertices and faces in a 3D mesh frame along with connectivity information (in certain embodiments using a Dracoper frame encoding technique), which produces an encoded output (e.g., in one embodiment, including an encoded binary output). In certain embodiments, frames of texture files are also used with one or more compression techniques to generate encoded video for the 3D mesh.

[0059] With further reference to Fig. 1, in various embodiments, the resulting compressed 3D output is rendered at 138. In particular, in various embodiments, a 3D video rendering for each vehicle 110 of Fig. 1 using one or more display devices 139 (e.g., one or more screens, head-up display devices, or the like) of the vehicle 110, resulting in the above-mentioned display 140 of the virtual 3D video content overlaid on the selected roadside destination and / or in another enhanced manner with respect to the selected roadside destination.

[0060] Accordingly, as described above, the platform and process 100 (and the various embodiments described herein) utilize images from inside the vehicle (e.g., from a driver monitoring system of the vehicle) in conjunction with a database of images and other information to deliver virtual content overlaid on one or more physical selected roadside elements and / or in another enhanced manner, as presented via one or more vehicle display devices (e.g., one or more screens, head-up display elements, or the like). In various embodiments, as described above, the platform and process 100 utilize target selection, bin-based readout, and LOD selection, also through the integration of vehicle signals such as vehicle image data, vehicle sensor data, GPS data, and other information.

[0061] As mentioned above, Fig. 2 is a functional block diagram of a system 200 for providing content to users of a vehicle, which is suitable for implementing the platform and process 100 of Fig. 1 and the implementations of the Fig. 3-7 according to exemplary embodiments. As in Fig. 2, in various embodiments, the system 200 may include a vehicle control system 202, a remote control system 204, a communications network 206, and one or more third-party services 208.

[0062] In various embodiments, the vehicle control system 202 is part of the vehicle 110 of Fig. 1. In certain embodiments, a separate vehicle control system 202 may be included in and as part of each respective vehicle 110. As in Fig. 2, the vehicle control system 202 includes, in various embodiments, a sensor assembly 210, a transceiver 214, a display 216, and a control unit 218.

[0063] In various embodiments, the sensor array 210 includes various sensors for operating the vehicle 110 and also for obtaining sensor data useful for the platform and process 100 of Fig. 1. In various embodiments, the sensor assembly 210 includes one or more location sensors 222 (e.g., as part of a GPS and / or other location system), IMU sensors 224 (e.g., as part of an inertial measurement unit of the vehicle 110), and image sensors 226 (e.g., including cameras as part of a driver monitoring system and / or otherwise receiving camera images of the driver and other users / passengers within the vehicle 110, as well as additional cameras receiving camera images of the roadway and the environment outside the vehicle 110).

[0064] In various embodiments, the transceiver 214 is used to communicate with the remote server (e.g., hosting the remote system 204) and, in certain embodiments, with one or more third-party services 208 (e.g., providing, for example, weather data, traffic data, or the like in certain embodiments). In various embodiments, the transceiver 214 is used, for example, to transmit sensor data from the vehicle and receive virtual 3D content for display to users of the vehicle 110. In various embodiments, the communication network 206 may also include any number of different types of wireless communication networks, such as cellular, satellite, internet, cloud-based networks, and the like.

[0065] In various embodiments, the display 216 of the vehicle 110 is also used to display content for users in the vehicle 110, including the virtual 3D content overlaid over the selected physical roadside destination and / or in another enhanced manner for display to the users of the vehicle 110. In certain embodiments, the display corresponds to the display described above in connection with Fig. 1. In various embodiments, the device 139 may also include a physical screen (e.g., a flip-up screen, a display element in the mirror, a display element in the windshield, e.g., as part of a rearview display, a forward-view display, a head-up display, and / or any one or more other types of displays for users of the vehicle 110).

[0066] In various embodiments, control unit 218 includes a computer system of vehicle 110 that performs certain processing functions of method 100 and the implementations described above. In certain embodiments, certain processing functions are performed within control unit 218 of vehicle 110, while other processing functions are instead performed via a control unit 248 of remote system 204 (as described further below). In certain other embodiments, the processing functions may be performed exclusively by one of control units 218 or 248.

[0067] As in Fig. 2, the control unit 218 of the vehicle 110, in certain embodiments, includes a processor 230, a memory 232, an interface 234, a storage device 236, and a bus 238. The processor 230 performs the computational and control functions of the control unit 218 and may comprise any type of processor or multiple processors, individual integrated circuits such as a microprocessor, or any suitable number of integrated devices and / or circuit boards that cooperate to perform the functions of a processing unit. During operation, the processor 230 executes one or more programs 240 contained in the memory 232 and, as such, generally controls the general operation of the control unit 218 in carrying out the processes described herein, such as the process 100 of Fig. 1 and the implementations of the Fig. 3-8 and described herein.

[0068] Memory 232 may be any suitable storage, including various types of non-transitory computer-readable storage media. In certain examples, memory 232 is located on the same computer chip as and / or co-located with processor 230. In the illustrated embodiment, memory 232 stores the aforementioned program 240 along with stored values ​​242 (e.g., stored images and / or other information).

[0069] The interface 234 enables communication with the computer system of the control unit 218, e.g., from a system driver and / or another computer system, and may be implemented using any suitable method and apparatus. In one embodiment, the interface 234 receives the various data from the sensor assembly 210, the remote system 204, and / or the third-party services 208, among other possible data sources. The interface 234 may include one or more network interfaces to communicate with other systems or components. The interface 234 may also include one or more network interfaces to communicate with technicians and / or one or more storage interfaces to connect to storage devices, such as the device 236.

[0070] The storage device 236 may be any suitable type of storage device, including various types of random access memory and / or other storage devices. In an exemplary embodiment, the device 236 includes a program product from which the memory 232 can receive a program 240 that performs one or more embodiments of one or more processes of the present description, such as the steps of process 100 of Fig. 1 and its implementations described herein. In another exemplary embodiment, the program product may be stored and / or otherwise accessed directly in memory 232 and / or a disk (e.g., disk 244), as described below.

[0071] Bus 238 is used to transfer programs, data, status, and other information or signals between the various components of the computer system of control unit 218. Bus 238 may be any suitable physical or logical means for interconnecting computer systems and components. These include, but are not limited to, direct, hard-wired connections, fiber optic, infrared, and wireless bus technologies. During operation, program 240 is stored in memory 232 and executed by processor 230.

[0072] While this exemplary embodiment is described in the context of a fully functional computer system, those skilled in the art will recognize that the mechanisms of the present description may be distributed as a program product having one or more types of non-transitory, computer-readable, signal-bearing media used to store the program and its instructions and to effect distribution, such as a non-transitory, computer-readable medium carrying the program and having computer instructions stored therein for causing a computer processor (such as processor 230) to execute the program.

[0073] As in Fig. 2, in exemplary embodiments, the remote control system 204 includes, among other possible components, a transceiver 246 and a control unit 248. In various embodiments, the transceiver 246 is similar to the transceiver 214 of the vehicle 110 and is used by the remote system 204 to communicate with the vehicles 110 and, in certain embodiments, with the third-party services 208 via the communication network 206. In certain embodiments, the transceiver 246 receives sensor data from the vehicle 110 and delivers virtual 3D content from the remote system 204 to the vehicle 110.

[0074] In various embodiments, the control unit 248 is also similar to the control unit 218 of the vehicle 110 described above. In various embodiments, the control unit 248 has a processor 250 and a memory 252 that stores one or more programs 260 and stored values ​​262, with functions similar or identical to those described above in connection with the control unit 218, the processor 230, the memory 232, the programs 240, and the stored values ​​242 with respect to the vehicle 110. In various embodiments, the control unit 248 may also include other components as described above in connection with the control unit 218, among other possible components.Furthermore, in various embodiments, the control unit 248 processes the sensor data and other information (including data from the vehicle 110 and the third-party services 208) to generate and deliver virtual 3D content to the vehicle 110, which is displayed within the vehicle 110 as an overlay on one or more physical targets and / or in another enhanced manner relative to the physical targets selected along the roadside. As mentioned above, in various embodiments, the processing steps of the method 100 and its implementations are performed via the control unit 248 of the remote system 204, while in certain embodiments, such processing steps may also be performed in whole or in part via the control unit 218 of the vehicle 110.

[0075] It should be understood that the systems, vehicles, and methods may vary from those depicted in the figures and described herein. For example, the platform and process 100 may differ from Fig. 1 of the Fig. 1. Likewise, the system 200 may differ from Fig. 2 and / or its components from the Fig. 2 and / or described herein. It is also recognized that the implementations of the Fig. 3-7 may differ from those shown in the drawings and / or described herein.

[0076] Although at least one exemplary embodiment has been presented in the foregoing detailed description, it should be understood that numerous variations exist. It should also be appreciated that the exemplary embodiment or exemplary embodiments are merely examples and are not intended to limit the scope, applicability, or configuration of the description in any way. Rather, the foregoing detailed description is intended to provide one skilled in the art with a convenient guide for implementing the exemplary embodiment or exemplary embodiments. It should be understood that various changes in the function and arrangement of elements may be made without departing from the scope of the description as set forth in the appended claims and their legal equivalents.

Claims

[1] Method comprising: Obtaining sensor data from one or more sensors of a vehicle; selecting, by a processor using the sensor data, a physical target comprising a physical element proximate the vehicle during a travel of the vehicle; Generating virtual content by the processor based on the sensor data; and providing the virtual content to one or more users of the vehicle as an overlay for the physical target or in an augmented form for the physical target. [2] The method of claim 1, wherein: the sensor data includes image sensor data, location sensor data, and inertial measurement unit (IMU) sensor data obtained from the vehicle; the method further comprises generating a head position and a vehicle position via the processor based on the image sensor data, the location sensor data, and the IMU sensor data obtained from the vehicle; and the generation of the virtual content is carried out using the head position and the vehicle position. [3] The method of claim 2, wherein generating the virtual content is further performed using a three-dimensional image database stored in a computer memory in combination with the head pose and the vehicle pose. [4] The method of claim 1, wherein the virtual content is displayed to the one or more users of the vehicle on one or more display devices within the vehicle. [5] The method of claim 1, wherein: the physical target includes an advertising billboard along a roadway on which the vehicle is traveling; and the virtual content is superimposed over, or displayed in an enhanced manner with respect to, the advertising panel as it is displayed inside the vehicle for the one or more users. [6] The method of claim 1, wherein: the physical target comprises a traffic control device along a roadway on which the vehicle is traveling; and the virtual content is superimposed over, or presented in an enhanced manner with respect to, the traffic control device as displayed inside the vehicle for the one or more users. [7] The method of claim 1, wherein: providing the virtual content uses a level of detail of a three-dimensional (3D) mesh provided to the vehicle by a remote system; and the level of detail is selected based on the distance between the physical target and the vehicle. [8] The method of claim 7, wherein the level of detail is further selected based on available network bandwidth or one or more other Quality of Experience (QOE) metrics. [9] The method of claim 7, further comprising: Performing frustum and back-face culling of the 3D meshes by the processor. [10] System comprising: one or more sensors configured to receive sensor data for a vehicle; and a processor coupled to the one or more sensors and configured to at least enable: selecting a physical target comprising a physical element proximate the vehicle during a travel of the vehicle using the sensor data; Generating virtual content based on the sensor data; and Providing the virtual content to one or more users of the vehicle as an overlay over the physical target or in an augmented form relative to the physical target.

Citation Information

Patent Citations

  • Display device and computer program

    DE112018004583T5