Dynamic multi-dimensional media content projection

By employing 3D surface modeling and ray casting to establish correspondence between 2D and 3D coordinates, the method addresses the challenge of projecting polygons onto obscured ground planes, improving object detection accuracy.

JP2025137461APending Publication Date: 2025-09-19FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025033344
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2025-03-04
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing 3D image projection techniques struggle to accurately project polygons onto the ground plane when it is partially or fully obscured by obstacles, such as trees or traffic lights, which is crucial for accurate object detection in real-world scenarios.

Method used

A method involving 3D surface modeling and ray casting to determine correspondence information between 2D pixel coordinates and 3D surface coordinates, allowing for the projection of 3D polygons onto the ground segment of a 3D surface model, even in occluded conditions.

Benefits of technology

Enables accurate projection of 3D polygons onto the ground plane, enhancing object detection capabilities in real-world scenarios by providing comprehensive 3D perspectives despite occlusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025137461000001_ABST
    Figure 2025137461000001_ABST
Patent Text Reader

Abstract

To provide a method of projecting dynamic multi-dimensional content.SOLUTION: Image data including 2D images depicting views of real-world is acquired. 3D polygons are detected corresponding to objects on a ground of the 2D images. 3D surface model is segmented into 3D segments including ground segments and non-ground segments. The ground segment of the 3D surface model is extracted by removing the one or more non-ground segments from the plurality of 3D segments. Corresponding information between 2D pixel coordinates in the 2D images and 3D surface coordinates of the ground segments is determined by applying a ray casting operation on the 2D images and the ground segment of the 3D surface model. The 3D polygons are projected onto the 3D surface model based on the correspondence information.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The embodiments discussed in this disclosure relate to dynamic multi-dimensional media content projection. [Background technology]

[0002] Based on its capabilities, 3D perspectives for object detection can provide a more complete understanding of real-world scenarios than 2D images. This can be particularly useful in situations requiring thorough investigation, such as car accidents. However, the widespread use of 2D monocular cameras, such as CCTV, due to their low cost, has prevented the full realization of the potential of 3D perspectives. Previous methods of projecting polygons onto the ground plane can encounter significant problems. While this method can work well when the entire ground plane is visible to the 2D monocular camera, it is unable to project polygons onto the invisible ground plane. This can be particularly problematic in occlusion situations, where objects such as trees, traffic lights, etc. obstruct the view.

[0003] The claimed subject matter in this disclosure is not limited to embodiments that solve any disadvantages or that operate only in such environments. Rather, this background is provided only to illustrate one example technology where some embodiments described in this disclosure may be practiced. Summary of the Invention

[0004] According to one aspect of an embodiment, a method may include a set of operations that may include acquiring image data that may include a 2D image depicting a view of a real-world location. The set of operations may further include detecting a 3D polygon corresponding to an object on a ground region of the 2D image and acquiring a 3D surface model of the real-world location. The 3D surface model may be segmented into a plurality of 3D segments including a ground segment and one or more non-ground segments. The ground segment of the 3D surface model may be extracted by removing the non-ground segments from the plurality of 3D segments. The set of operations may further include determining correspondence information between 2D pixel coordinates in the 2D image and 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment of the 3D surface model, and projecting a 3D polygon corresponding to the object onto the 3D surface model based on the correspondence information.

[0005] The object and advantages of the embodiments will be realized and attained at least by the elements, features, and combinations particularly pointed out in the claims.

[0006] Both the foregoing general description and the following detailed description are given by way of example and are explanatory only and are not limitations of the invention, as claimed. [Brief explanation of the drawings]

[0007] The exemplary embodiments will be described and explained with additional specificity and detail through the use of the accompanying drawings.

[0008] [Figure 1] FIG. 1 depicts an exemplary environment for dynamic multi-dimensional media content projection.

[0009] [Figure 2] FIG. 1 is a block diagram illustrating an exemplary electronic device for dynamic multi-dimensional media content projection.

[0010] [Figure 3A] FIG. 1 illustrates an execution pipeline for object detection and ground segmentation in dynamic multi-dimensional media content projection.

[0011] [Figure 3B] FIG. 1 illustrates an execution pipeline for determining correspondence information for dynamic multi-dimensional media content projection.

[0012] [Figure 4] FIG. 1 illustrates a flowchart of an exemplary method for estimating a 3D polygon corresponding to an object in a 2D image.

[0013] [Figure 5] FIG. 1 illustrates an exemplary scenario for 3D polygon detection and estimation.

[0014] [Figure 6] FIG. 1 illustrates an exemplary scenario for the acquisition of a 3D surface model.

[0015] [Figure 7] FIG. 10 illustrates an example flowchart for determining correspondence information for dynamic multi-dimensional media content projection.

[0016] [Figure 8] FIG. 1 illustrates an exemplary scenario for estimation of initial camera pose parameters corresponding to an initial view of a 3D surface model in 3D space.

[0017] [Figure 9] FIG. 1 illustrates an example flowchart for dynamic multi-dimensional media content projection.

[0018] All according to at least one embodiment described in this disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0019] Some embodiments described herein relate to a method and system for dynamic multidimensional media content projection. In the present disclosure, image data may be acquired. The image data may include a two-dimensional (2D) image depicting a view of a real-world location. Then, a 3D polygon (e.g., a bounding box) corresponding to an object on a ground region of the 2D image may be detected. A three-dimensional (3D) surface model of the real-world location may be acquired for segmentation into a plurality of 3D segments, including a ground segment and one or more non-ground segments. Based on the segmented 3D surface model, a ground segment of the 3D surface model may be extracted by removing one or more non-ground segments from the plurality of 3D segments. Correspondence information between 2D pixel coordinates in the 2D image and the 3D surface coordinates of the ground segment may be determined by applying a ray-casting operation to the 2D image and the ground segment. One or more 3D polygons corresponding to the object may be projected onto the 3D surface model.

[0020] According to one or more embodiments of the present disclosure, the technical field of 3D image generation and projection may be improved by configuring an electronic device to dynamically project multidimensional media content using correspondence information between 2D spatial coordinates and 3D spatial coordinates. The electronic device may acquire image / video data including a 2D image depicting a view of a real-world location. Furthermore, the electronic device may detect a 3D polygon corresponding to an object on a ground region of the 2D image. Based on the acquired 3D surface model of the real-world location, the 3D surface model may be segmented into a plurality of 3D segments including ground segments and non-ground segments. The ground segment of the 3D surface model may be extracted by removing the non-ground segments from the plurality of 3D segments. Furthermore, the electronic device may determine correspondence information between 2D pixel coordinates in the 2D image and the 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment. Based on the correspondence information, the electronic device may project a 3D polygon corresponding to the object onto the 3D surface model.

[0021] Conventional techniques for 3D image projection may offer several advantages by understanding real-world situations beyond the limitations of 2D images. 3D perspectives may assist, for example, in visualizing a vehicle accident from multiple angles to identify potential problems. It will be appreciated that generating a 3D model from 2D images may involve capturing 2D images from different perspectives of an object or environment. The 2D images may be processed using a 3D model generation component, which generates a 3D model based on the 2D images and their associated 3D data. Alternatively, laser-based 3D scanning may be used to directly acquire the 3D surface model.

[0022] Traditional 3D object detection approaches often encounter significant challenges when projecting polygons (e.g., bounding cubes) onto the ground. While this approach works well when the entire ground is visible from an imaging device (e.g., a 2D imaging device), it may fail to project them onto unseen ground. This is particularly problematic in occlusion situations where objects such as trees or traffic lights can obstruct the view. The primary challenge is projecting such polygons onto accurate ground locations, which may be required for accurate object detection. Sub-problems may arise when polygons are unavailable or are incorrectly detected, further complicating the task. As a result, while 3D perspectives may have the potential to improve object detection, these issues must be addressed before their benefits can be fully realized. In contrast, the present disclosure provides a technique for accurately projecting polygons from 2D images into 3D space by introducing a correspondence table between 2D pixel coordinates in the 2D image and 3D surface coordinates, even when obstacles are present and the object is partially obscured or invisible.

[0023] Embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0024] FIG. 1 is a diagram depicting an example environment for dynamic multi-dimensional media content projection arranged in accordance with at least one embodiment described herein. Referring to FIG. 1, environment 100 is shown. Environment 100 may include an electronic device 102, a server 104, a database 106, a communications network 108, a 2D imaging device 110, and a 3D imaging device 112. Electronic device 102, server 104, 2D imaging device 110, and 3D imaging device 112 may be communicatively coupled to each other via communications network 108. Electronic device 102 may include a display device 102A. Server 104 may be communicatively coupled to database 106. FIG. 1 also illustrates a set of images 110A acquired by 2D imaging device 110 and a 3D surface model 112A acquired by 3D imaging device 112. The set of images (e.g., 2D images 110A and 3D surface model 112A) may be stored in database 106.

[0025] The set of images (e.g., 2D image 110A) and 3D surface model 112A of Figure 1 are presented by way of example only, and the set of images may include N images without departing from the scope of this disclosure.

[0026] A 2D imaging device 110 (e.g., a camera) and a 3D imaging device 112 (e.g., a LiDAR) may be communicatively coupled to the electronic device 102. In some embodiments, the electronic device 102 may include the 2D imaging device 110 and the 3D imaging device 112.

[0027] The electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to project multidimensional media content as described herein. In some embodiments, the electronic device 102 may be configured to acquire a 2D image 110A and a 3D surface model 112A, as shown in FIG. 1. Examples of multidimensional media content may include, but are not limited to, the 2D image 110A, the 3D surface model 112A, 2D video, and 3D video. Examples of user end devices 102 may include, but are not limited to, mobile devices, desktop computers, laptops, computer workstations, action cameras, 360 cameras, computing devices, mainframe machines, gaming consoles, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality (MR) devices, mainframe machines, computer workstations, and Internet of Things (IoT) devices. In one or more embodiments, the electronic device 102 may include a user terminal device and a server communicatively coupled to the user terminal device. The electronic device 102 may be implemented using hardware, including a processor or microprocessor (e.g., to perform or control the performance of one or more operations). In some other instances, the electronic device 102 may be implemented using a combination of hardware and software.

[0028] The server 104 may include suitable logic, interfaces, and / or code that may be configured to store neural network models (e.g., object detection model, polygon estimation model 208B, 3D surface model 112A, etc.). The server 104 may be configured to retrieve historical data associated with 3D polygons corresponding to objects in the 2D image 110A or the 3D surface model 112A from the database 106. Each 3D polygon may enclose a corresponding object of the plurality of objects. The server 104 may be configured to estimate 3D polygons corresponding to one or more objects detected in the 2D image 110A based on the trained polygon estimation model 208B.

[0029] The database 106 may be stored or cached on a device such as a server or the electronic device 102. The device storing the database 106 may be configured to receive queries for data (e.g., the 2D image 110A and the 3D surface model 112A) from the electronic device 102. In response, the electronic device 102 of the database 106 may be configured to retrieve and provide the queried data to the electronic device 102 based on the received query. In some embodiments, the database 106 may be hosted on multiple servers stored at the same location or at different locations. The operations of the database 106 may be performed using hardware, including a processor, a microprocessor (e.g., to perform or control the execution of one or more operations), a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some other examples, the database 106 may be implemented using software.

[0030] The communication network 108 may include a communication medium through which the electronic device 102 may communicate with a server or device that may store the database 106 and with the imaging devices (e.g., the 2D imaging device 110 and the 3D imaging device 112). Examples of the communication network 108 may include, but are not limited to, the Internet, a cloud network, Wireless Fidelity (Wi-Fi), a Personal Area Network (PAN), a Local Area Network (LAN), a cellular network (such as a Long Term Evolution (or 4G) cellular network or a 5G cellular network), a satellite network (such as a network of low-earth orbit satellites), and / or a Metropolitan Area Network (MAN). The various devices in the environment 100 may be configured to connect to the communication network 108 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, light fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access points (APs), device-to-device communication, cellular communication protocols, and / or Bluetooth (BT) communication protocols, or combinations thereof.

[0031] The server 104, the 2D imaging device 110, and the 3D imaging device 112 may be communicatively coupled via a wired or wireless network. In some embodiments, the electronic device 102 may incorporate the functionality of the server 104, the 2D imaging device 110, and the 3D imaging device 112.

[0032] The 2D imaging device 110 may include suitable logic, circuitry, and interfaces that may be configured to capture an image or multiple images of a real-world location, including various objects (e.g., buildings, trees, roads, terrain, traffic signals, people, etc.). The 2D imaging device 110 may be further configured to capture a 2D image 110A corresponding to the real-world location. The 2D image 110A may be any digital data that may be rendered, streamed, broadcast, or stored on any electronic device 102 or storage device. Examples of the 2D imaging device 110 may include, but are not limited to, an image sensor, a wide-angle camera, an action camera, a closed-circuit television (CCTV) camera, a camcorder, a camera with an integrated depth sensor, a cinematic camera, a digital single-lens reflex (DSLR) camera, a digital single-lens mirrorless (DSLM) camera, a digital camera, a camera phone, a time-of-flight camera (ToF camera), a night-vision camera, and / or other capture devices.

[0033] The 3D imaging device 112 may include appropriate logic, circuitry, and interfaces that may be configured to capture 3D data by performing 3D scanning of real-world objects (e.g., buildings, trees, roads, terrain, traffic lights, people, etc.). The 3D data (e.g., a depth map or a point cloud) may be converted into a 3D surface model 112A. Examples of the 3D imaging device 112 may include, but are not limited to, a LiDAR, a time-of-flight camera (ToF camera), a stereoscopic camera, a structured light 3D scanner, and a CT scanner. The 3D surface model 112A of the real-world location may be obtained from the 3D imaging device 112.

[0034] 3D data may be collected from a 3D imaging device 112 to obtain a 3D surface model 112A of a real-world location (e.g., a street in a city). The 3D data may be any digital data that can be rendered, streamed, broadcast, or stored on any electronic device 102 or storage device. Examples of 3D data may include, but are not limited to, depth maps, 3D volume data, and point cloud data.

[0035] During operation, the electronic device 102 may acquire image data including a 2D image 110A depicting a view of a real-world location. The electronic device 102 may be configured to detect 3D polygons corresponding to objects on a ground region of the 2D image 110A. The electronic device 102 may further be configured to acquire a 3D surface model 112A of the real-world location from the 3D data received by the 3D imaging device 112.

[0036] In at least one embodiment, the electronic device 102 may estimate initial camera pose parameters corresponding to an initial view of the 3D surface model 112A in 3D space. The electronic device 102 may optimize the initial camera pose parameters to align the view of the 3D surface model 112A with the view of the 2D image 110A.

[0037] The electronic device 102 may be further configured to segment the 3D surface model 112A into a plurality of 3D segments including ground segments and non-ground segments. The ground segments of the 3D surface model 112A may be extracted by removing one or more non-ground segments from the plurality of 3D segments. By applying a ray-cast operation to the 2D image 110A and the ground segments of the 3D surface model 112A, the electronic device 102 may be configured to determine correspondence information between 2D pixel coordinates in the 2D image 110A and 3D surface coordinates of the ground segments. By way of example and not limitation, the correspondence information may be a correspondence table that may include 2D pixel coordinates of the 2D image 110A and 3D surface coordinates corresponding to the 2D pixel coordinates. The electronic device 102 may be further configured to project one or more 3D polygons (e.g., 3D cubes / cuboids) corresponding to the object onto the 3D surface model 112A based on the correspondence information.

[0038] The electronic device 102 may generate vertex information including a mapping of 3D surface coordinates to vertices of one or more 3D polygons. The electronic device 102 may further generate edge information including a mapping of vertices to edges of one or more 3D polygons. Projection may be performed based on the vertex information and the edge information. The vertex information may include a primary table including a vertex ID, a polygon ID, a vertex type, and 3D surface coordinates. The edge information may include a secondary table including a polygon ID, an edge ID, and an ID of the edge vertex. The vertex type may include either ground or object. Each 3D polygon of the one or more 3D polygons may enclose a corresponding object of the multiple objects in the 3D space of the 3D surface model 112A.

[0039] In at least one embodiment, the electronic device 102 may retrieve, from a database, historical data associated with 3D polygons corresponding to multiple objects in a set of images (e.g., the 2D image 110A). Each 3D polygon of the multiple 3D polygons surrounds a corresponding object among the multiple objects. The electronic device 102 may train the polygon estimation model 208B based on the historical data. During inference, the electronic device 102 may estimate a 3D polygon corresponding to an object detected in the 2D image 110A based on the trained polygon estimation model 208B. The electronic device 102 may project the detected 3D polygon onto a ground region of the 2D image 110A based on a determination that the difference between the estimated 3D polygon and the detected 3D polygon is less than a predefined threshold. Alternatively, the electronic device 102 may project the estimated 3D polygon onto a ground region of the 2D image 110A based on a determination that the difference between the estimated 3D polygon and the detected 3D polygon is greater than a predefined threshold. Details regarding historical and behavioral data are provided further in FIG. 4, for example (at 402, 412).

[0040] Modifications, additions, or omissions may be made to FIG. 1 without departing from the scope of the present disclosure. For example, environment 100 may include more or fewer elements than those shown and described in this disclosure. For example, in some embodiments, environment 100 may include electronic device 102 but may not include user device 106. Additionally, in some embodiments, the functionality of each of user device 106 may be incorporated into electronic device 102 without departing from the scope of the disclosure.

[0041] Figure 2 is a block diagram illustrating an example electronic device 102 for dynamic multi-dimensional media content projection, arranged in accordance with at least one embodiment described herein. Figure 2 is described in conjunction with elements from Figure 1. With reference to Figure 2, a block diagram 200 of the electronic device 102 is shown. The electronic device 102 may include a network interface 202, input / output (I / O) devices 204, circuitry 206, and memory 208. The I / O devices may include a display device 102A. The memory may include an object detector 208A, a polygon estimation model 208B, and a 3D surface model 208C.

[0042] The network interface 202 may include suitable logic, circuitry, interfaces, and / or code that may be configured to establish communications between the electronic device 102, the server of the database 106, the 2D imaging device 110, and the 3D imaging device 112 over the communications network 108. The network interface 202 may be implemented using various known technologies to support wired or wireless communications of the electronic device 102 over the communications network 108. The network interface 202 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, and / or a local buffer.

[0043] The I / O device(s) 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive user input. For example, the I / O device(s) 204 may display image data (e.g., 2D image 110A) and 3D surface model 112A captured using 2D imaging device 110 and 3D imaging device 112, respectively. The I / O device(s) 204 may project 3D polygons corresponding to objects onto 3D surface model 112A. The I / O device(s) 204 may include various input and output devices that may be configured to communicate with other components, such as circuitry 206 and network interface 202. Examples of input devices may include, but are not limited to, a touchscreen, a keyboard, a mouse, a joystick, and / or a microphone. Examples of output devices may include, but are not limited to, a display (e.g., display device 102A) and a speaker.

[0044] Circuitry 206 may include suitable logic, circuitry, and / or interfaces that can be configured to execute program instructions associated with different operations performed by electronic device 102. For example, some operations may include acquiring image data, detecting 3D polygons, segmenting the 3D surface model 208C, extracting ground and non-ground segments, determining correspondence information, and projecting the 3D polygons. Circuitry 206 may include any suitable special-purpose or general-purpose computer, computing entity, or processing device, including various computer hardware or software modules, and may be configured to execute instructions stored on any applicable computer-readable storage medium. For example, circuitry 206 may include a microprocessor, a microcontroller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any other digital or analog circuitry configured to interpret and / or execute program instructions and / or process data.

[0045] Although shown as a single processor in FIG. 2 , circuitry 206 may include any number of processors configured, individually or collectively, to perform or direct the performance of any number of operations of electronic device 102 described herein. Additionally, one or more processors may reside in one or more different electronic devices, such as different servers. In some embodiments, circuitry 206 may interpret and / or execute program instructions stored in memory 208 and / or process data. In some embodiments, circuitry 206 may fetch program instructions from memory 208 and load program instructions into memory 208. After the program instructions are loaded into memory 208, circuitry 206 may execute the program instructions. Some examples of circuitry 206 may be a Graphical Processing Unit (GPU), a Central Processing Unit (CPU), a Reduced Instruction Set Computer (RISC) processor, an ASIC processor, a Complex Instruction Set Computer (CISC) processor, a coprocessor, and / or combinations thereof.

[0046] Memory 208 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store program instructions executable by circuitry 206. In certain embodiments, memory 208 may be configured to store an operating system and associated application-specific information. Memory 208 may include computer-readable storage media for carrying or having computer-executable instructions or data structures stored thereon. Such computer-readable storage media may include any available media that can be accessed by a general-purpose or special-purpose computer, such as circuitry 206. By way of example, and not limitation, such computer-readable storage media may include tangible or non-transitory computer-readable storage media, including random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory devices (e.g., solid-state memory devices), or any other storage medium that may be used to carry or store specific program code in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer. Combinations of the above may also be included within the scope of computer-readable storage media. Computer-executable instructions may include, for example, instructions and data configured to cause circuitry 206 to perform a particular operation or group of operations associated with electronic device 102. Memory 208 may include, but is not limited to, various models, object detector 208A, polygon estimation model 208B, 3D surface model 208C. The various models are described in detail in FIGS. 3, 5, 8, and 9.

[0047] By way of example, and not limitation, such computer-readable storage media may include tangible or non-transitory computer-readable storage media such as compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices (e.g., hard-disk drives (HDDs)), flash memory devices (e.g., solid-state drives (SSDs), secure digital (SD) cards, other solid-state memory devices), or any other storage medium that can be used to carry or store specific program code in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer. Combinations of the above may also be included within the scope of computer-readable storage media. Computer-executable instructions may include, for example, instructions and data configured to cause circuitry 206 to perform a particular operation or group of operations associated with electronic device 102.

[0048] Object detection (using the object detector 208A) may be achieved through semantic segmentation, in which different regions in an image (e.g., the 2D image 110A) are defined and labeled based on the real-world objects (e.g., vehicles, people, buildings, etc.) they depict. The resulting object detection may then be used to generate point features that can be located on a map. In an exemplary scenario, a user walking on a sidewalk may appear in an image (e.g., the 2D image 110A). The user may be considered an object. The object may be a static object or a dynamic object. In a dynamic sequence of images, some of the images may contain occlusions that partially or completely obscure the object. Object detection may be useful for many applications, including, but not limited to, security, surveillance, facial recognition, autonomous driving, and robotics. The circuitry 206 may detect and track pedestrians, vehicles, and traffic signs on busy streets. This can help improve road safety, traffic management, and navigation. Based on object detection in the 2D image 110A, an algorithm may be enabled to distinguish the object from the background. Polygons may also handle occlusion, as they can accurately label only the visible parts of occluded objects.

[0049] The polygon estimation model 208B may be a neural network or a computational network, or may be a system arranged in multiple layers with artificial neurons as nodes. During inference, the polygon estimation model 208B may process input images to localize objects across a 2D image (e.g., 2D image 110A) by extracting behavioral data associated with the movement or appearance of such objects. The multiple layers of the neural network may include an input layer, one or more hidden layers, and an output layer. Each of the multiple layers may include one or more nodes (or artificial neurons, represented, for example, by circles). The output of every node in the input layer may be connected to at least one node in the hidden layer. Similarly, the input of each hidden layer may be connected to the output of at least one node in another layer of the neural network. The output of each hidden layer may be connected to the input of at least one node in another layer of the neural network. A node in the final layer may receive input from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyperparameters of the neural network. Such hyperparameters may be set before or after training the neural network on the training dataset.

[0050] Each node of the neural network in polygon estimation model 208B may correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) having a set of parameters that can be adjusted during training of the network. The set of parameters may include, for example, weight parameters, regularization parameters, etc. Each node may use the mathematical function to calculate an output based on one or more inputs from nodes in other layers (e.g., previous layers) of the neural network. All or some of the nodes of the neural network may correspond to the same or different mathematical functions.

[0051] In training the neural network of polygon estimation model 208B, one or more parameters of each node of the neural network may be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result based on a loss function for the neural network. The above process may be repeated for the same or different inputs until a minimum loss function is achieved and the training error is minimized. Several methods for training are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, metaheuristics, etc.

[0052] Examples of polygon estimation model 208B may be neural network-based models such as deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), CNN-recurrent neural networks (CNN-RNNs), R-CNNs, Fast R-CNNs, Faster R-CNNs, artificial neural networks (ANNs), You Only Look Once (YOLO) networks, Long Short Term Memory (LSTM) network-based RNNs, CNNs + ANNs, LSTMs + ANNs, gated recurrent unit (GRU)-based RNNs, fully connected neural networks, Connectionist Temporal Classification (CTC)-based RNNs, deep Bayesian neural networks, generative adversarial networks (GANs), and / or combinations of such networks. In particular embodiments, neural network-based polygon estimation model 208B may be based on a hybrid architecture of multiple deep neural networks (DNNs).

[0053] The 3D surface model 208C may be a digital representation of a real or hypothetical feature in three-dimensional space. Some examples of 3D surfaces may be a landscape, an urban corridor, an underground gas deposit, and a network of well depths for determining the depth of a water table. The 3D surface model 208C may be created from various data sources, such as points, lines, polygons, or images, using interpolation or triangulation methods. The 3D surface model 208C may be used for various applications, such as determining mass properties, checking for interferences, generating cross sections, and creating finite element meshes.

[0054] In some embodiments, the 3D surface model 208C may be obtained from a 3D imaging device 112 (e.g., LiDAR), or the 3D surface may be generated based on point cloud data (e.g., 602). The captured point cloud data may be converted into the 3D surface model 208C (or a 3D mesh with or without color / texture).

[0055] Display device 102A may include suitable logic, circuitry, interfaces, and / or code that may be configured to display image data (e.g., 2D image 110A) and 3D data information captured using 2D imaging device 110 and 3D imaging device 112, respectively. Display screen 212 may be configured to receive user input (e.g., for projection of 3D surface model 112A) from user 114. In such a case, display device 102A may be a touchscreen for receiving user input. Display device 102A may be implemented via several known technologies, such as, but not limited to, Liquid Crystal Display (LCD), Light Emitting Diode (LED), plasma display, and / or organic LED (OLED) display technologies, and / or other display technologies.

[0056] Modifications, additions, or omissions may be made to the exemplary electronic device 102 without departing from the scope of the present disclosure. For example, in some embodiments, the exemplary electronic device 102 may include any number of other components that may not be explicitly shown or described for the sake of brevity. The object detector 208A, polygon estimation model 208B, and 3D surface model 208C are described in FIGS. 3, 4, 5, and 9.

[0057] 3A and 3B collectively illustrate an execution pipeline for dynamic multidimensional media content projection, according to one embodiment of the present disclosure. FIGS. 3A and 3B are described in conjunction with elements from FIGS. 1 and 2. Referring to FIGS. 3A and 3B, an execution pipeline 300 is shown. The exemplary execution pipeline 300 may include a set of operations that may be performed by one or more electronic components, such as the electronic device 102 of FIG. 1. The operations may include image data (e.g., 2D image 110A) acquisition, 3D model (e.g., 3D surface model 112A) acquisition, polygon (e.g., 3D polygon) detection, 3D model surface model segmentation, ground segment extraction, correspondence information determination, and 3D polygon projection. This set of operations may be performed by the electronic device 102 for dynamic multidimensional media content projection, as described herein.

[0058] At 302, operations for 2D image acquisition may be performed. In one embodiment, circuitry 206 may be configured to receive image data including 2D image 110A. 2D image 110A may be any digital data that can be rendered, streamed, broadcast, or stored in 2D on any electronic device 102. For example, circuitry 206 may receive image data from an imaging device (e.g., 2D imaging device 110). 2D image 110A may include objects such as vehicles, ground surfaces, buildings, roads, utility poles, and people. In one embodiment, 2D image 110A may be pre-stored in memory 208, and circuitry 206 may retrieve the pre-stored image data 110A from memory 208.

[0059] At 304, operations for 3D polygon detection may be performed. The objects may be static objects, dynamic objects, or a combination thereof. As part of the operations, circuitry 206 may detect one or more 3D polygons (e.g., 3D polygon 304B) corresponding to one or more objects (e.g., cars) on the ground region of 2D image 110A. By way of example and not limitation, 3D polygon 304B may be a bounding cube / cuboid that may enclose an object in the image (such as the 2D image shown in 304A). Alternatively, 3D polygon 304B may be of any suitable shape without departing from the scope of this disclosure.

[0060] In one embodiment, detection may be performed using a neural network-based object detector (e.g., object detector 208A) that can identify and label different regions in the 2D image 304A based on the real-world objects contained in those regions. To locate the objects, a 3D polygon 304B may be placed around each object. As an example, a 2D image 110A of a street may be considered. The detected objects may be, for example, streetlights, vehicles, people, buildings, traffic lights, roads, trees, etc. The resulting object detections may be used to generate features that can be located on a map. Consider an exemplary scenario of the 2D image 304A in which a vehicle appears to be parked on the street. The vehicle may be considered as an object that can be detected by the object detector 208A. A 3D polygon 304B (e.g., a 3D cube) may be placed to surround the vehicle based on the results generated by the object detector 208A. In some examples, the 2D image 304A may also include occlusions (not shown) that may fully or partially occlude the detected objects.

[0061] At 306, operations may be performed to obtain the 3D surface model 112A. In one embodiment, the circuitry 206 may be configured to receive the 3D surface model 112A, which may include 3D mesh or point data. The 3D surface model 112A may be any digital content that can be rendered, streamed, broadcast, or stored in 3D on any electronic device. The 3D surface model 112A may be generated from 3D data such as depth maps, 3D volume data, and point cloud data.

[0062] Circuitry 206 may receive 3D data from an imaging device (e.g., 3D imaging device 112), which may correspond to an object in 2D image 110A. The object may be a vehicle, a surface, a building, a person, etc. In another embodiment, the 3D data may be collected from 3D imaging device 112 and processed to obtain a 3D surface model 112A of a real-world location (e.g., a street in a city). In another embodiment, the 3D data may be pre-stored in memory 208, and circuitry 206 may retrieve the pre-stored 3D data from memory 208.

[0063] At 308, operations for view alignment may be performed. View alignment of the acquired 3D surface model 112A may include estimating initial pose parameters corresponding to an initial view of the 3D surface model 112A in 3D space. The initial camera pose parameters may be optimized to align the view of the 3D surface model 112A with the view of the 2D image 110A. The real-world location and the 3D surface model 112A may have different coordinates. While the real-world location uses longitude and latitude, the 3D surface model 112A uses x, y, and z axes. This makes it difficult to obtain the exact location and orientation of a fixed camera. Therefore, values ​​such as position, angle, and zoom may be adjusted to align the views of the 2D image 110A and the 3D surface model 112A. The initial pose parameters corresponding to the initial view of the 3D surface model 112A may be optimized to align the view of the 3D surface model 112A in 3D space to obtain an aligned 3D surface model 308A. The aligned 3D surface model 308A may include a ground segment 308B and a non-ground segment 308C. For example, the aligned 3D surface model 308A shows a set of buildings (i.e., non-ground segment 308C) as part of the 3D surface model 112A. Camera pose estimation is further described in FIG. 8.

[0064] At 310, the aligned 3D surface model 308A may be segmented into 3D segments, including ground segments 308B and non-ground segments 308C. Segmentation may be performed on the aligned 3D surface model 308A by dividing the aligned 3D surface model 308A into regions or segments based on criteria such as color, texture, shape, or semantic meaning (semantic segmentation). In one embodiment, the circuitry 206 may identify and label objects or portions of the 2D image. For example, the object detector 208A may locate and identify objects in the image by drawing bounding boxes around the objects and assigning class labels such as "car," "person," "road," or "dog." Based on the labeled objects or portions of the image, the aligned 3D surface model 308A may be segmented as ground segments 308B and non-ground segments 308C.

[0065] Semantic segmentation of the aligned 3D surface model 308A may involve assigning a class label to each point in the aligned 3D surface model 308A, such as “car,” “road,” or “sky.” Semantic segmentation may not distinguish between instances of the same class, such as multiple cars in the same image. Instance segmentation may distinguish between instances of the same class, such as multiple automobiles in the same aligned 3D surface model 308A. Panoptic segmentation may combine semantic and instance segmentation by assigning a class label and an instance ID to each point on the aligned 3D surface model 308A, but may group points that belong to the same semantic domain, such as “road 1” or “sidewalk 1.” This type of segmentation may provide a comprehensive and consistent representation of the aligned 3D surface model 308A.

[0066] In some embodiments, segmentation may be performed using various techniques and algorithms, such as, but not limited to, thresholding, clustering, edge detection, region growing, graph-based methods, or deep learning models, such as, but not limited to, convolutional neural networks (CNNs), U-Net, Mask R-CNN, DeepLab, etc. CNN architectures consisting of an encoder-decoder structure may have an encoder that reduces the spatial resolution of the input image and extracts high-level features, while a decoder may improve the spatial resolution and generate an output segmentation map.

[0067] At 312, extraction of ground segments 308B from the aligned 3D surface model 308A may be performed by removing non-ground segments 308C from the 3D surface model 308A. For example, the segmented image 312A depicts the ground segments 312B (i.e., roads).

[0068] At 314, a ray casting operation may be performed. The ray casting operation may be performed on the 2D image 110A and the ground segment 308B of the 3D surface model 112A. Using the 3D surface model 112A, each pixel of the 2D image may be assigned a 3D coordinate in 3D space based on a perspective projection equation. As an example, the image size of the 2D image 110A may be extracted, and 2D coordinates may be selected from the 2D pixel coordinates. A ray casting operation may be applied to the selected 2D coordinates to estimate corresponding 3D surface coordinates on the ground segment 308B in 3D space. The above operation may be repeated for all pixels of the 2D image. Correspondence information may be determined based on the 2D pixel coordinates and the 3D surface coordinates.

[0069] At 316, correspondence information between 2D pixel coordinates in the 2D image 110A and 3D surface coordinates of the ground segment 308B may be determined based on applying a ray casting operation to the 2D image 110A and the ground segment 308B of the 3D surface model 112A. In one embodiment, the correspondence information may be stored in the database 106.

[0070] In an exemplary embodiment, the correspondence information may be a table 316A that includes 2D pixel coordinates and 3D surface coordinates that correspond to the 2D pixel coordinates. The correspondence table 316A may be used to project a 3D polygon (e.g., a 3D cube around a 3D object) onto the surface of the 3D surface model 112A. In the correspondence table 316A, the 2D pixel coordinates are represented by, for example, a width (W) (as shown). 2D ) and height (h 2D ) parameters. The 3D pixel coordinates may include 3D information, such as horizontal distance (X 3D ), vertical distance (Y 3D ), and depth distance (Z 3D ) parameters. The camera coordinate system or 2D pixel coordinates may be

number

number

[0071] At 318, a polygon shaping operation may be performed. A polygon 318C (e.g., a 2D polygon or a 3D polygon 318C) may include vertices and edges. The polygon shaping representation may include a type of projection having ground vertices and non-ground vertices. For example, polygon 318C may be generated based on the ground vertices. Table 1 represents a vertex table. The vertex table may include, for example, a polygon ID, a vertex ID, a vertex type, and 3D pixel coordinates. Table 2 represents an edge table. An edge may include, for example, a polygon ID, an edge ID, a vertex FROM ID, and a vertex TO ID. The vertex table may be referred to as a primary table, and the edge table may be referred to as a secondary table. [Table 1] [Table 2]

[0072] Polygon ID, vertex ID, and vertex type are some of the attributes that can be used to describe and manipulate polygons. Polygon ID may be a unique identifier that can be assigned to each polygon 318C in an image. Polygon ID may be used to access, modify, or delete a specific polygon from a collection of polygons. For example, if an image contains multiple polygons, each polygon may be given a different ID, such as 1, 2, or 3. The ID can be used to select a polygon and change its color, position, or shape. Vertex ID may be a unique identifier assigned to each vertex in a polygon. Vertex ID may be used to access, modify, or delete a specific vertex from a polygon. For example, if a polygon has four vertices, each vertex is given a different ID, such as 1, 2, 3, or 4. The ID can be used to select a vertex and change its coordinates or add or delete edges to or from it. Vertex type may be a classification that indicates the role or function of the vertex within the polygon. The vertex type may be used to determine how the polygon is drawn, filled, or textured. For example, some common vertex types include corner vertices that form a sharp angle between two edges, curve vertices that form a smooth curve between two edges, texture vertices that specify texture coordinates for the polygon, etc. 3D image 318A (obtained after polygon shaping) includes 3D polygon 318C on ground segment 318B, shaped based on correspondence table 316A.

[0073] At 320, a projection of the 3D polygon 318C may be performed. The 3D polygon 318C may be projected onto the ground segment 318B of the projected 3D surface model 320A based on correspondence information corresponding to one or more objects. Furthermore, the projection may be performed based on vertex information and edge information. Polygon projection is described in detail in Figures 4, 5, 7, and 9.

[0074] Figure 4 illustrates a flowchart of an exemplary method for estimating a 3D polygon corresponding to an object in a 2D image, according to one embodiment of the present disclosure. Figure 4 is described in conjunction with elements from Figures 1, 2, and 3. Referring to Figure 4, a flowchart 400 is shown. The method illustrated in flowchart 400 may begin at 402 and may be performed by any suitable system, apparatus, or device, such as the exemplary electronic device 102 of Figure 1 or the circuitry 206 of Figure 2. Although illustrated with discrete blocks, steps and operations associated with one or more of the blocks of flowchart 400 may be divided into additional blocks, combined into fewer blocks, or eliminated, depending on the implementation.

[0075] At block 402, historical data associated with 3D polygons corresponding to objects in 2D image 110A may be retrieved. The historical data may be stored in database 106. Each 3D polygon may enclose a corresponding object of the plurality of objects. Historical data collection may be performed through various known methods to examine historical images stored in database 106, such as objects captured in 2D image 110A or 3D images over various time frames, different angles, etc.

[0076] At block 404, the polygon estimation model 208B may be trained based on historical data.

[0077] At block 406, a 3D polygon corresponding to the detected object may be estimated. The object may be detected based on the 3D polygon on the ground region of the 2D image 110A. Predicting behavior data may involve using machine learning to analyze how the polygon's behavior data is analyzed under various scenarios, such as translation, rotation, and scaling. The polygon estimation model 208B may include a reinforcement learning (RL) framework, where an agent learns to control the polygon by interacting with the environment and receiving rewards or penalties based on the behavior of the object (corresponding to the 3D polygon).

[0078] The polygon estimation model 208B may optimize the object's behavior data for different purposes, such as, but not limited to, reaching a goal, avoiding an obstacle, or following a path. In one embodiment, a generative adversarial network (GAN) may be used to determine the object's behavior data corresponding to a 3D polygon, where a generator attempts to estimate the polygon behavior and a discriminator attempts to distinguish between real behavior and generated behavior. Based on the behavior data, 3D polygon estimation may be performed. The generator may learn to generate diverse and natural polygon behaviors that match the data distribution. The predetermined threshold may vary based on past instances of the object's behavior and represent physical laws and human movement. For example, a predefined threshold may be used to determine whether the object detector 208A correctly identified an object in an image based on intersection over union (IoU).

[0079] At block 408, a difference between estimated 3D polygons corresponding to detected objects may be estimated. When the difference between the estimated 3D polygons and the detected 3D polygons is less than a predefined threshold, control may pass to 410. When the difference between the estimated 3D polygons and the detected 3D polygons is greater than a predefined threshold, control may pass to 412. Because the detected polygons may be inaccurate due to occlusion, the polygons may be further estimated to optimize object detection. Thus, the estimated 3D polygons corresponding to the objects and the detected 3D polygons may be compared to determine differences in behavior.

[0080] The threshold may depend on the object's motion, rotation, scale, or shape changes in the previous image or video. For example, if the object is a car, the threshold may be higher when the car is moving fast or making sharp turns, and lower when the car is moving slowly or stationary. This means that the threshold may be based on rules or patterns that describe how the object behaves in the real world. For example, if the object is a person, the threshold may be based on human anatomy, posture, gestures, or facial expressions. If the object is, for example, a ball, the threshold may be based on the physics of motion, gravity, or friction. Polygon projection is exemplarily described in detail in FIGS. 5 and 9.

[0081] At block 410, based on a determination that the difference between the estimated 3D polygon and the detected 3D polygon is less than a predefined threshold, the detected 3D polygon may be projected onto a ground region of the 2D image 110A. Additionally or alternatively, the projection of the detected 2D polygon may be used to project the polygon (or 3D polygon) onto the 3D surface model 112A.

[0082] In one embodiment, the detected 3D polygon may be dynamically detected based on the 2D image 110A. A 3D polygon estimated based on historical data may be referred to as an estimated polygon. The 3D polygon may include depth in addition to other 2D parameters. Because the detected polygon may be inaccurate or missing due to occlusion, the polygon may be estimated even if the detected polygon can be obtained from the dynamic 2D image 110A. In the case of a human object, the estimated polygon may be learned from past instances and may broadly represent the laws of physics and human movement.

[0083] At block 412, based on a determination that the difference between the estimated 3D polygon and the detected 3D polygon is greater than a predefined threshold, the estimated 3D polygon may be projected onto a ground region of the 2D image 110A. Additionally or alternatively, the projection of the estimated 3D polygon may be used to project the 3D polygon (or polygons) onto the 3D surface model 112A.

[0084] Because detected 2D polygons may be inaccurate due to occlusion, estimated polygons may be estimated even if detected 3D polygons may be obtained from the dynamic 2D image 110A. Estimated 3D polygons may be learned from past instances and broadly represent the laws of physics and human motion. Therefore, they are expected to not deviate significantly compared to polygons detected based on fragmented information.

[0085] Figure 5 illustrates an exemplary scenario for estimating a 3D polygon, according to one embodiment of the present disclosure. Figure 5 is described in conjunction with elements of Figures 1, 2, 3, and 4. With reference to Figure 5, an exemplary scenario 500 is shown. The method illustrated in exemplary scenario 500 may be performed by any suitable system, apparatus, or device, such as the exemplary electronic device 102 of Figure 1 or the circuitry 206 of Figure 2.

[0086] Consider a set of image frames 504, 506, 508, 510 (e.g., 2D images) as an example for determining behavior data of object 504D across the set of image frames 504, 506, 508, 510. The number of image frames shown in FIG. 5 is provided by way of example only, and such example should not be construed as limiting the present disclosure. Image frames 504, 506, 508, and 510 may include only one image frame or more than N image frames for determining behavior data without departing from the scope of the present disclosure. For simplicity, only four image frames are shown in FIG. 5. However, in some embodiments, there may be more than four image frames without limiting the scope of the present disclosure.

[0087] In frame 504, object 504D may be detected. The number of objects shown in Figure 5 is provided by way of example only. Object 504D may include only one object or may include multiple objects without departing from the scope of this disclosure.

[0088] The first step is to determine the number of objects 504D in the set of image frames 504, 506, 508, and 510. The set of image frames 504, 506, 508, and 510 may be analyzed to determine an object of interest (e.g., object 504D) or an object to be tracked for behavior data. Dynamic objects may, in some embodiments, be considered objects of interest in image frames 504, 506, 508, and 510 for purposes of determining behavior data. For object 504D detected in image frame 504, a 3D polygon 504C may be estimated at time "t." Image frame 504 at time "t" may include multiple objects. For example, image frame 504 may include a ground segment and object 504D, which is to be tracked for behavior data. Object 504D may be at location 1 in image frame 504 at time t (the object may be surrounded by 3D polygon 504C). Consider an image frame 506 at time t+1, which may include occlusions 504A and 504B. Now, an object 504D may be detected at a second location in the second image frame 506 by considering the object's behavior data. The polygon estimation model 208B may be pre-trained using historical object data and may compensate for missing polygons due to non-detection and incorrect object locations (at the locations of occlusions 504B and 504A in the image frames 506 and 508). Polygon ID, vertex ID, and vertex type are some of the attributes that may be used to describe and manipulate polygons. In one embodiment, each polygon may include vertices and edges, and information about the vertices and edges may be stored in the form of a table (Table 1 and Table 2 in FIG. 3 ). The table may include vertex information and edge information.

[0089] In image frame 508, object 504D may be hidden behind occlusions 504A and 504B at time t+j. Image frame 510 at time t+j+1 may be considered to indicate that object 504D may be detected at time t+j+1. Object 504D, corresponding to 3D polygon 504C, may be estimated for image frames 506 and 508, where object 504D is not visible due to occlusions 504B, 504A. The estimated location of object 504D may be determined based on behavior data.

[0090] Figure 6 illustrates an exemplary scenario for obtaining a 3D surface model, according to one embodiment of the present disclosure. Figure 6 will be described in conjunction with elements of Figures 1, 2, 3, 4, and 5. With reference to Figure 6, an exemplary scenario 600 is shown. The method illustrated in exemplary scenario 600 may be performed by any suitable system, apparatus, or device, such as the exemplary electronic device 102 of Figure 1 or the circuitry 206 of Figure 2.

[0091] Scenario 600 may include a 3D imaging device 112. The 3D imaging device 112 may include suitable logic, circuitry, or interfaces that may be configured to capture a 3D view of a real-world location (e.g., a city street). According to one embodiment, the 3D imaging device 112 may include multiple image sensors (not shown) to capture 3D views of the real-world location from multiple perspectives. According to one embodiment, the 3D imaging device 112 may be integrated into the electronic device 102. Examples of the 3D imaging device 112 may include, but are not limited to, a LiDAR, a time-of-flight camera (ToF camera), a stereoscopic camera system, a structured light 3D scanner, and a CT scanner.

[0092] In one embodiment, the 3D imaging device 112 may be a LiDAR camera. A LiDAR camera may function by sending out fast pulses of laser light and detecting the reflected light with a sensor. By scanning a laser beam across a scene, the LiDAR camera may generate a 3D point cloud 602A of the object or scene 602, which may be further processed into a 3D surface model image 604. For example, the 3D point cloud 602A may be converted into a surface model or mesh by applying a meshing operation to the 3D point cloud 602A. The 3D surface model image 604 may be referred to as geometric data, which is most often composed of a bundle of connected triangles (or polygons) that explicitly describe a surface. The 3D surface model image 604 may consist of vertices, edges, and faces that define the surface of the object. Different methods and algorithms may be used to generate the 3D surface model image 604 from the 3D point cloud 602A, such as Poisson surface reconstruction, ball-pivot algorithm, Delaunay triangulation, or marching cubes.

[0093] FIG. 7 illustrates an example flowchart for determining correspondence information for dynamic multidimensional media content projection, according to one embodiment of the present disclosure. FIG. 7 will be described in conjunction with elements from FIGS. 1, 2, 3, 4, and 5. Referring to FIG. 7, an example flowchart 700 is shown. The method illustrated in example flowchart 700 may be performed by any suitable system, apparatus, or device, such as example electronic device 102 of FIG. 1 or circuitry 206 of FIG. 2. Although illustrated with discrete blocks, steps and operations associated with one or more of the blocks of flowchart 700 may be divided into additional blocks, combined into fewer blocks, or eliminated, depending on the particular implementation.

[0094] At 704, the image size of the 2D image 110A may be extracted. The 2D image 110A may be provided as input for performing image extraction. The 2D image 110A may be retrieved from the database 106, and the image size may be extracted based on the image's dimensions in terms of pixels, such as height and width. In another embodiment, the image size may be extracted by dividing the 2D image 110A into smaller patches of a fixed size and using the dimensions of each patch to extract the image size.

[0095] At 706, a 2D coordinate of a pixel may be selected from the 2D pixel coordinates of the 2D image 110A. The 2D pixel coordinate may be X 2D , Y 2D It may be considered that X 2D is the pixel position along the x-axis, and Y 2D is the pixel's position along the y-axis. As an example, the 2D coordinates of two pixels (ID1_2D and ID2_2D) may be expressed as (x,y)=(305,656) and (x,y)=(506,656).

[0096] At 708, a ray casting operation may be applied to the selected 2D coordinates to estimate corresponding 3D surface coordinates on the ground segment of the 3D surface model 112A in 3D space. The 3D surface coordinates may include 3D pixel coordinates (X3D, Y3D, Z3D). As an example, the 3D coordinates for two pixels ID1_2D and ID2_3D are given below: (x, y, z) = (0.1, 0.02, 0.21), (x, y, z) = (0.21, 0.5, 0.1); ...

[0097] At 710, correspondence information (e.g., correspondence table 316A) may be determined based on the 2D pixel coordinates and the 3D surface coordinates. 2D coordinate selection from 2D pixel coordinates may be a process of transforming pixel locations in 2D image 110A to a different coordinate system, such as world coordinates, camera coordinates, or intrinsic coordinates. This is useful for various applications, such as building a 3D surface model, image registration, or feature extraction.

[0098] At 712, it may be determined whether the selected 2D coordinate is the last coordinate in the 2D image. If the selected 2D coordinate is the last coordinate in the 2D image, control may pass to End. Otherwise, control may pass to 706, where the next 2D coordinate from the 2D image may be selected. The process from 706 through 710 may be repeated for all pixels in the 2D image.

[0099] FIG. 8 illustrates an exemplary scenario for estimating initial camera pose parameters corresponding to an initial view of a 3D surface model in 3D space, according to one embodiment of the present disclosure. FIG. 8 will be described in conjunction with elements from FIGS. 1, 2, 3, 4, 5, 6, and 7. Referring to FIG. 8, an exemplary scenario 800 is illustrated. In the exemplary scenario 800, an initial view 802 of a 3D surface model, a final view 804 of the 3D surface model (after alignment), and a base 2D image 806 are illustrated. To align the initial view 802 of the 3D surface model with the view of the base 2D image 806, the circuitry 206 may estimate initial camera pose parameters corresponding to the initial view 802 of the 3D surface model in 3D space. The circuitry 206 may then iteratively optimize the initial camera pose parameters to align the initial view 802 of the 3D surface model with the view of the base 2D image 806. After alignment, the final view 804 of the 3D surface model may match the view of the base 2D image 806 .

[0100] Figure 9 illustrates an example flowchart for dynamic multi-dimensional media content projection, according to one embodiment of the present disclosure. Figure 9 will be described in conjunction with elements from Figures 1, 2, 3A, 3B, 4, 5, 6, 7, and 8. Referring to Figure 9, a flowchart 900 is shown. Operations 902-916 may be implemented by any computing system, such as the electronic device 102 or circuitry 206 of the electronic device 102 of Figure 1. Operations may begin at 902 and may proceed to 904.

[0101] At 904, image data may be acquired, including a 2D image 110A depicting a view of a real-world location. The 2D image 110A may be any digital data that can be rendered, streamed, broadcast, or stored on any electronic device or storage device. Examples of the 2D image 110A may include an image (such as overlay graphics), a real-world image captured by an imaging device, or audio / video data.

[0102] At 906, 3D polygons corresponding to the objects may be detected on the ground region of the 2D image 110A. The polygons may be 3D polygons with straight edges that form a closed area. The polygons may include cubes, cuboids, etc. The vertices, edges, and faces of these polygons may be defined using points (Pt) having x, y, and z coordinates.

[0103] Polygon generation may be a process that involves creating a set of points to define the boundary of an object or area of ​​interest in an image or video. The points or vertices may be connected by lines to form a polygon that encloses the object or area and takes its shape. For example, vertices may be generated from bounding box coordinates by selecting four points along the perimeter of a rectangular bounding box and using such points as polygon vertices.

[0104] At 908, a 3D surface model 208C of a real-world location may be obtained. The 3D surface model 112A may be a digital representation of a feature in three-dimensional space. Some examples of 3D surfaces may be a landscape, an urban corridor, an underground gas deposit, and a network of well depths for determining the depth of a water table. The 3D surface model 112A may be created from various data sources, such as points, lines, polygons, or images, using interpolation or triangulation methods.

[0105] In one embodiment, initial camera pose parameters may be estimated corresponding to an initial view of the 3D surface model 208C in 3D space. The initial camera pose parameters may be optimized to align the view of the 3D surface model 208C with the view of the 2D image 110A.

[0106] At 910, the 3D surface model 112A may be segmented into 3D segments including a ground segment and one or more non-ground segments. In some embodiments, model segmentation may be performed using various techniques and algorithms, such as, but not limited to, thresholding, clustering, edge detection, region growing, graph-based methods, or deep learning models. Additionally or alternatively, segmentation may be performed by acquiring 2D images 110A having multiple ground segments, such as roads, and non-ground segments, such as buildings. The 2D images 110A, for example, traffic surveillance images featuring multiple objects, may be collected. The collected 2D images 110A may be annotated. This involves drawing bounding boxes around objects of interest in each 2D image 110A and labeling them. The object of interest may be, for example, the ground segment 312B. For example, a bounding box may be drawn around the ground segment and labeled “ground.” The annotated 2D images 110A may be used to train a machine learning model. The model may learn to recognize features of each object from the annotated images.

[0107] In one embodiment, acquiring the 3D surface model 112A may include capturing 3D data via an image capture device and generating a 3D point cloud based on the 3D data. The point cloud data may be converted into the 3D surface model 208C.

[0108] At 912, ground segments of the 3D surface model 112A may be extracted by removing one or more non-ground segments from the plurality of 3D segments.

[0109] At 914, correspondence information between 2D pixel coordinates in 2D image 110A and 3D surface coordinates of the ground segment may be determined by applying a ray casting operation to the ground segment of 2D image 110A and 3D surface model 112A. The ray casting operation may be applied to the selected 2D coordinates to estimate corresponding 3D surface coordinates on the ground segment in 3D space.

[0110] In an exemplary embodiment, the correspondence information may be a table 316A that includes 2D pixel coordinates and their corresponding 3D surface coordinates. The correspondence table 316A may be useful for quickly transitioning polygon data on ground segments of the 3D surface model 208C from 2D coordinates to 3D coordinates.

[0111] At 916, one or more 3D polygons corresponding to the objects may be projected onto the 3D surface model 208C based on the correspondence information. In one embodiment, each 3D polygon of the one or more 3D polygons may surround a corresponding one of the objects in the 3D space of the 3D surface model 208C. Vertex information may be generated that includes a mapping of 3D surface coordinates to vertices of the one or more 3D polygons. Edge information may be generated that includes a mapping of vertices to edges of the one or more 3D polygons. Projection may be further performed based on the vertex information and the edge information. The vertex information may include a primary table that includes a vertex ID, a polygon ID, a vertex type, and the 3D surface coordinates, and the edge information may include a secondary table that includes a polygon ID, an edge ID, and an ID of the edge vertex. The vertex type may include either ground or object.

[0112] In one embodiment, projecting the polygon may include retrieving historical data associated with 3D polygons corresponding to objects in the 2D image 110A or the 3D surface model 112A. Each 3D polygon surrounds a corresponding object among the plurality of objects. The polygon estimation model 208B may be trained based on the historical data. A 3D polygon corresponding to an object detected in the 2D image 110A may be estimated based on the trained polygon estimation model 208B. Based on a determination that a difference between the estimated 3D polygon and the detected 3D polygon is less than a predefined threshold, the detected 3D polygon may be projected onto a ground region of the 2D image 110A. Based on a determination that a difference between the estimated 3D polygon and the detected 3D polygon is greater than a predefined threshold, the estimated 3D polygon may be projected onto a ground region of the 2D image 110A.

[0113] Those skilled in the art will understand that the scope of the present disclosure is not limited to 2D / 3D images. According to one embodiment, the 2D and 3D images may be any data, such as video, moving images, etc., without departing from the scope of the present disclosure.

[0114] Various embodiments of the present disclosure may provide one or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause a system (such as the exemplary electronic device 102) to perform operations. The operations may include acquiring image data (e.g., 2D image data) including a two-dimensional (2D) image (e.g., a live recording received from a CCTV camera) depicting a view of a real-world location. The operations may include detecting 2D polygons corresponding to one or more objects on a ground region of the 2D image and acquiring a 3D surface model of the real-world location. The set of operations may further include segmenting the 3D surface model from the plurality of 3D segments into a plurality of 3D segments including a ground segment and one or more non-ground segments. The set of operations may further include determining correspondence information between 2D pixel coordinates in the 2D image and 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment of the 3D surface model, and projecting one or more 3D polygons corresponding to the one or more objects onto the 3D surface model based on the correspondence information.

[0115] As used in this disclosure, the terms “module” and “component” may refer to a specific hardware implementation and / or a software object or software routine that may be stored and / or executed on general-purpose hardware of the electronic device 102 (e.g., a computer-readable medium, a processing device, etc.) configured to perform the actions of the module or component. In some embodiments, different components, modules, engines, and services described in this disclosure may be implemented as objects or processes that run on the electronic device 102 (e.g., as separate threads). While some of the systems and methods described in this disclosure are described as being implemented in software (stored on and / or executed by general-purpose hardware), specific hardware implementations or combinations of software and specific hardware implementations are also possible and contemplated. In this description, a “computing entity” may be any electronic device 102 as defined above in this disclosure, or any module or combination of modules operating on the electronic device 102.

[0116] Additionally, where a specific number of introduced claim provisions are intended, such intention will be expressly set forth in the claim; where such a provision is absent, such intention does not exist. For example, as an aid to understanding, the following appended claims may contain the use of the introductory phrases "at least one" and "one or more" to introduce claim provisions. However, the use of such phrases should not be construed to suggest that the introduction of a claim provision with the indefinite article "a" or "an" limits any particular claim containing such introduced claim provision to embodiments containing only one such provision, even when the same claim includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"), and the same is true for any use of an indefinite article used to introduce a claim provision.

[0117] Furthermore, any disjunctive word or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both of the terms. For example, the phrase "A or B" should be understood to include the possibilities of "A" or "B" or "A and B."

[0118] All examples and conditional language set forth in this disclosure are intended for educational purposes to aid the reader in understanding the disclosure and the concepts the inventors have contributed to furthering the art, and should not be construed as being limited to the examples and conditions so specifically set forth. While embodiments of the present disclosure have been described in detail, various changes, substitutions, and alterations can be made thereto without departing from the spirit and scope of the present disclosure.

[0119] This specification also discloses the following supplementary information: (Appendix 1) 1. A method executed by a processor in an electronic device, comprising: acquiring image data including a two-dimensional (2D) image depicting a view of a real-world location; detecting one or more 3D polygons corresponding to one or more objects on a ground region of the 2D image; obtaining a three-dimensional (3D) surface model of the real-world location; segmenting the 3D surface model into a plurality of 3D segments, including a ground segment and one or more non-ground segments; extracting the ground segment of the 3D surface model by removing the one or more non-ground segments from the plurality of 3D segments; determining correspondence information between 2D pixel coordinates in the 2D image and 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment of the 3D surface model; and projecting the one or more 3D polygons corresponding to the one or more objects onto the 3D surface model based on the correspondence information. (Appendix 2) estimating initial camera pose parameters corresponding to an initial view of the 3D surface model in 3D space; 2. The method of claim 1, further comprising: optimizing the initial camera pose parameters to align a view of the 3D surface model with a view of the 2D image. (Appendix 3) 2. The method of claim 1, wherein the correspondence information includes a table including the 2D pixel coordinates and the 3D surface coordinates corresponding to the 2D pixel coordinates. (Appendix 4) generating vertex information comprising a mapping of the 3D surface coordinates to vertices of the one or more 3D polygons; generating edge information comprising a mapping of the vertices to edges of the one or more 3D polygons; 2. The method of claim 1, wherein the projection of the one or more 3D polygons is further performed based on the vertex information and the edge information. (Appendix 5) the vertex information includes a primary table containing vertex IDs, polygon IDs, vertex types, and the 3D surface coordinates; 5. The method of claim 4, wherein the edge information includes a secondary table containing the polygon ID, edge ID, and edge vertex ID. (Appendix 6) 5. The method of claim 4, wherein the vertex type includes one of ground or object. (Appendix 7) 2. The method of claim 1, wherein each 3D polygon of the one or more 3D polygons encloses a corresponding object of the plurality of objects in 3D space of the 3D surface model. (Appendix 8) retrieving, from a database, historical data associated with a plurality of 2D polygons corresponding to a plurality of objects in a set of images; each 3D polygon of the plurality of 3D polygons encircling a corresponding object of the plurality of objects; training a polygon estimation model based on the historical data; estimating one or more 3D polygons corresponding to the detected one or more objects in the 2D image based on the trained polygon estimation model; and 2. The method of claim 1, further comprising: projecting the detected one or more 3D polygons onto the ground region of the 2D image based on a determination that a difference between the estimated one or more 3D polygons and the detected one or more 3D polygons is less than a predefined threshold. (Appendix 9) 9. The method of claim 8, further comprising projecting the estimated one or more 3D polygons onto the ground region of the 2D image based on a determination that a difference between the estimated one or more 3D polygons and the detected one or more 3D polygons is greater than the predefined threshold. (Appendix 10) The determination of the correspondence information includes: extracting an image size of the 2D image; selecting a 2D coordinate from the 2D pixel coordinates; applying a ray casting operation to the selected 2D coordinates to estimate corresponding 3D surface coordinates of the 3D surface coordinates on the ground segment in 3D space; determining the correspondence information based on the 2D pixel coordinates and the 3D surface coordinates. (Appendix 11) The acquisition of the 3D surface model includes: capturing 3D data via an image capture device; generating a 3D point cloud based on the 3D data; 2. The method of claim 1, further comprising converting the point cloud data to the 3D surface model. (Appendix 12) One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause an electronic device to perform an operation, the operation including: acquiring image data including a two-dimensional (2D) image depicting a view of a real-world location; detecting one or more 3D polygons corresponding to one or more objects on a ground region of the 2D image; obtaining a three-dimensional (3D) surface model of the real-world location; segmenting the 3D surface model into a plurality of 3D segments, including a ground segment and one or more non-ground segments; extracting the ground segment of the 3D surface model by removing the one or more non-ground segments from the plurality of 3D segments; determining correspondence information between 2D pixel coordinates in the 2D image and 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment of the 3D surface model; and projecting the one or more 3D polygons corresponding to the one or more objects onto the 3D surface model based on the correspondence information. (Appendix 13) The process comprises: estimating initial camera pose parameters corresponding to an initial view of the 3D surface model in 3D space; and optimizing the initial camera pose parameters to align a view of the 3D surface model with a view of the 2D image. (Appendix 14) The process comprises: generating vertex information comprising a mapping of the 3D surface coordinates to vertices of the one or more 3D polygons; generating edge information comprising a mapping of the vertices to edges of the one or more 3D polygons; 13. The one or more non-transitory computer-readable storage media of claim 12, wherein projection is further performed based on the vertex information and the edge information. (Appendix 15) the vertex information includes a primary table containing vertex IDs, polygon IDs, vertex types, and the 3D surface coordinates; 15. The one or more non-transitory computer-readable storage media of claim 14, wherein the edge information includes a secondary table including the polygon ID, edge ID, and edge vertex ID. (Appendix 16) The process comprises: retrieving, from a database, historical data associated with a plurality of 3D polygons corresponding to a plurality of objects in a set of images; each 3D polygon of the plurality of 3D polygons encircling a corresponding object of the plurality of objects; training a polygon estimation model based on the historical data; estimating one or more 3D polygons corresponding to the detected one or more objects in the 2D image based on the trained polygon estimation model; and and projecting the detected one or more 3D polygons onto the ground region of the 2D image based on a determination that a difference between the estimated one or more 3D polygons and the detected one or more 3D polygons is less than a predefined threshold. (Appendix 17) The process comprises: extracting an image size of the 2D image; selecting a 2D coordinate from the 2D pixel coordinates; applying a ray casting operation to the selected 2D coordinates to estimate corresponding 3D surface coordinates of the 3D surface coordinates on the ground segment in 3D space; determining the correspondence information based on the 2D pixel coordinates and the 3D surface coordinates. (Appendix 18) The process comprises: capturing 3D data via an image capture device; generating a 3D point cloud based on the 3D data; and converting the point cloud data into the 3D surface model. (Appendix 19) 1. An electronic device comprising: a memory configured to store instructions; a processor coupled to the memory that executes the instructions to perform a process, the process comprising: acquiring image data including a two-dimensional (2D) image depicting a view of a real-world location; detecting one or more 3D polygons corresponding to one or more objects on a ground region of the 2D image; obtaining a three-dimensional (3D) surface model of the real-world location; segmenting the 3D surface model into a plurality of 3D segments, including a ground segment and one or more non-ground segments; extracting the ground segment of the 3D surface model by removing the one or more non-ground segments from the plurality of 3D segments; determining correspondence information between 2D pixel coordinates in the 2D image and 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment of the 3D surface model; and projecting the one or more 3D polygons corresponding to the one or more objects onto the 3D surface model based on the correspondence information.

Claims

1. 1. A method executed by a processor in an electronic device, comprising: acquiring image data including a two-dimensional (2D) image depicting a view of a real-world location; detecting one or more 3D polygons corresponding to one or more objects on a ground region of the 2D image; obtaining a three-dimensional (3D) surface model of the real-world location; segmenting the 3D surface model into a plurality of 3D segments including a ground segment and one or more non-ground segments; extracting the ground segment of the 3D surface model by removing the one or more non-ground segments from the plurality of 3D segments; determining correspondence information between 2D pixel coordinates in the 2D image and 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment of the 3D surface model; and projecting the one or more 3D polygons corresponding to the one or more objects onto the 3D surface model based on the correspondence information.

2. estimating initial camera pose parameters corresponding to an initial view of the 3D surface model in 3D space; The method of claim 1 , further comprising optimizing the initial camera pose parameters to align a view of the 3D surface model with a view of the 2D image.

3. The method of claim 1 , wherein the correspondence information comprises a table containing the 2D pixel coordinates and the 3D surface coordinates corresponding to the 2D pixel coordinates.

4. generating vertex information comprising a mapping of the 3D surface coordinates to vertices of the one or more 3D polygons; generating edge information comprising a mapping of the vertices to edges of the one or more 3D polygons; The method of claim 1 , wherein the projection of the one or more 3D polygons is further performed based on the vertex information and the edge information.

5. the vertex information includes a primary table containing vertex IDs, polygon IDs, vertex types, and the 3D surface coordinates; The method of claim 4 , wherein the edge information includes a secondary table containing the polygon ID, edge ID, and edge vertex IDs.

6. The method of claim 5 , wherein the vertex type includes one of ground or object.

7. The method of claim 1 , wherein each 3D polygon of the one or more 3D polygons encloses a corresponding object of a plurality of objects in 3D space of the 3D surface model.

8. retrieving, from a database, historical data associated with a plurality of 2D polygons corresponding to a plurality of objects in a set of images; each 3D polygon of the plurality of 3D polygons encircling a corresponding object of the plurality of objects; training a polygon estimation model based on the historical data; estimating one or more 3D polygons corresponding to the detected one or more objects in the 2D image based on the trained polygon estimation model; and 2. The method of claim 1, further comprising: projecting the detected one or more 3D polygons onto the ground region of the 2D image based on a determination that a difference between the estimated one or more 3D polygons and the detected one or more 3D polygons is less than a predefined threshold.

9. 9. The method of claim 8, further comprising: projecting the estimated one or more 3D polygons onto the ground region of the 2D image based on a determination that a difference between the estimated one or more 3D polygons and the detected one or more 3D polygons is greater than the predefined threshold.

10. The determination of the correspondence information includes: extracting an image size of the 2D image; selecting a 2D coordinate from the 2D pixel coordinates; applying a ray casting operation to the selected 2D coordinates to estimate corresponding 3D surface coordinates of the 3D surface coordinates on the ground segment in 3D space; and determining the correspondence information based on the 2D pixel coordinates and the 3D surface coordinates.

11. The acquisition of the 3D surface model comprises: Capturing 3D data via an image capture device; generating a 3D point cloud based on the 3D data; The method of claim 1 , further comprising: converting the 3D point cloud into the 3D surface model.

12. One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause an electronic device to perform an operation, the operation including: acquiring image data including a two-dimensional (2D) image depicting a view of a real-world location; detecting one or more 3D polygons corresponding to one or more objects on a ground region of the 2D image; obtaining a three-dimensional (3D) surface model of the real-world location; segmenting the 3D surface model into a plurality of 3D segments including a ground segment and one or more non-ground segments; extracting the ground segment of the 3D surface model by removing the one or more non-ground segments from the plurality of 3D segments; determining correspondence information between 2D pixel coordinates in the 2D image and 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment of the 3D surface model; and projecting the one or more 3D polygons corresponding to the one or more objects onto the 3D surface model based on the correspondence information.

13. 1. An electronic device comprising: a memory configured to store instructions; a processor coupled to the memory that executes the instructions to perform a process, the process comprising: acquiring image data including a two-dimensional (2D) image depicting a view of a real-world location; detecting one or more 3D polygons corresponding to one or more objects on a ground region of the 2D image; obtaining a three-dimensional (3D) surface model of the real-world location; segmenting the 3D surface model into a plurality of 3D segments including a ground segment and one or more non-ground segments; extracting the ground segment of the 3D surface model by removing the one or more non-ground segments from the plurality of 3D segments; determining correspondence information between 2D pixel coordinates in the 2D image and 3D surface coordinates of the ground segment by applying a ray casting operation to the 2D image and the ground segment of the 3D surface model; and projecting the one or more 3D polygons corresponding to the one or more objects onto the 3D surface model based on the correspondence information.