Natural language 3D data search

By merging multimodal embedded information with 3D spatial information, and combining instance segmentation and octree structure, the problem of low storage and processing efficiency of point cloud data is solved, enabling efficient 3D data query and privacy-preserving user interaction.

CN121889786APending Publication Date: 2026-04-17SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2024-12-06
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing 3D scanning technology generates point cloud data that suffers from problems such as being discrete, incomplete, disordered, and having a large data volume, making it difficult to store and process effectively, and resulting in low user privacy and data processing efficiency.

Method used

By merging multimodal embedded information with 3D spatial information and combining instance segmentation to achieve global and local scene understanding, an octree structure is used to optimize data storage and querying, and natural language or image references are used for querying and retrieval.

Benefits of technology

It improves the efficiency of 3D data storage and retrieval, simplifies user interaction, protects privacy, reduces data transmission volume, and enhances the flexibility and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121889786A_ABST
    Figure CN121889786A_ABST
Patent Text Reader

Abstract

A method includes merging information from a multi-modal embedding in an indexed point cloud data structure with three-dimensional (3D) spatial information for a captured scene. The method further includes performing at least one of querying the 3D point cloud data or retrieving the 3D point cloud data based on a user input, the user input including at least one of a natural language or an image reference. The method further includes using instance segmentation in combination with multi-modal embedding to implement global scene understanding and local scene understanding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to machine learning systems. More specifically, this disclosure relates to a system and method for searching three-dimensional (3D) data in natural language. Background Technology

[0002] Thanks to significant investments in robotics and autonomous driving, 3D capture technology has made tremendous progress over the past decade, with substantial reductions in the cost, size, complexity, and noise of sensing technologies. This has opened up unprecedented new areas and use cases for consumer electronics. The 3D data created by these technologies presents numerous unresolved new challenges in terms of privacy, user experience, and processing. 3D scanning is the process of capturing a physical object or environment to create a digital representation in the form of a 3D model. Many technologies exist for 3D scanning, including contact scanning technologies (such as zoomers) and optical scanning technologies (such as LiDAR, photogrammetry, and structured light).

[0003] Typically, 3D scanning produces point clouds, which are discrete sets of 3D points and potential appearance attributes (such as color and intensity). Point clouds are a convenient representation for many applications because they are easy to visualize and manipulate. However, point clouds also have some inherent limitations. For example, point clouds are discrete, meaning the number of points representing a scanned object or environment is limited. This can lead to incomplete representations because small features or details may not be captured. Point clouds are often large and difficult to compress. Point clouds are also unordered, meaning the points reflecting the location of the scanned object or environment do not have an inherent order. Summary of the Invention

[0004] Solution This disclosure relates to systems and methods for searching three-dimensional (3D) data in natural language.

[0005] In a first embodiment, a method includes: merging information from multimodal embeddings in an indexed point cloud data structure with 3D spatial information specific to the captured scene. The method further includes: performing at least one of querying or retrieving 3D point cloud data based on user input, the user input including at least one of natural language or image references. The method also includes: using instance segmentation in conjunction with multimodal embeddings to achieve global scene understanding and local scene understanding.

[0006] In a second embodiment, an apparatus includes at least one processing unit configured to merge information from multimodal embeddings in an indexed point cloud data structure with 3D spatial information specific to the captured scene. The at least one processing unit is further configured to perform a query for or retrieval of at least one of the 3D point cloud data based on user input, the user input including at least one of natural language or image references. The at least one processing unit is also configured to use instance segmentation in conjunction with multimodal embeddings to achieve both global and local scene understanding.

[0007] In a third embodiment, a non-transitory computer-readable medium includes instructions that, when executed, cause at least one processor of an electronic device to: merge information from multimodal embeddings in an indexed point cloud data structure with 3D spatial information for the captured scene. The instructions, when executed, also cause at least one processor of the electronic device to: perform a query for or retrieval of at least one of the 3D point cloud data based on user input, the user input including at least one of natural language or image references. The instructions, when executed, also cause at least one processor of the electronic device to: combine multimodal embeddings with instance segmentation to achieve global scene understanding and local scene understanding.

[0008] Other technical features will be obvious to those skilled in the art based on the following figures, description and claims.

[0009] Before proceeding with the detailed embodiments below, it may be advantageous to define the specific words and phrases used in this patent document. The terms “send,” “receive,” and “communicate,” and their derivatives, include both direct and indirect communication. The terms “comprise,” “include,” and their derivatives, indicate, but are not limited to, those included in, those included in. The term “or” is inclusive, indicating and / or. The phrase “associated with,” and its derivatives, indicate including, being included within, interconnected with, containing, contained within, connected to or connected with, combined with or combined with, capable of communicating with, cooperating with, interleaving, juxtaposing, proximate, bound to or bound with, having, possessing the nature of, having a relationship to, or a relationship with, etc.

[0010] Furthermore, the various functions described below may be implemented or supported by one or more computer programs, each of which is formed by computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium accessible by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, optical disc (CD), digital video disc (DVD), or any other type of storage. "Non-transitory" computer-readable medium excludes wired, wireless, optical, or other communication links that transmit transient electrical or other signals. Non-transitory computer-readable medium includes media that permanently store data and media that store data and can be rewritten later, such as rewritable optical discs or erasable memory devices.

[0011] As used herein, terms and phrases such as “having,” “may have,” “comprising,” or “may include” indicate the presence of a feature (such as a number, function, operation, or component, like a part) and do not exclude the presence of other features. Furthermore, as used herein, the phrases “A or B,” “at least one of A and / or B,” or “one or more of A and / or B” can include all possible combinations of A and B. For example, “A or B,” “at least one of A and B,” and “at least one of A or B” can indicate all of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. Furthermore, as used herein, the terms “first” and “second” can modify various components regardless of importance and do not limit the components. These terms are used only to distinguish one component from another. For example, a first user device and a second user device can indicate user devices that are different from each other, regardless of the order or importance of the devices. A first component can be referred to as a second component without departing from the scope of this disclosure, and vice versa.

[0012] It will be understood that when an element (such as a first element) is referred to as being (operably or communicatively) "combined" with / "attached to" another element (such as a second element) or "connected" with / "connected to" another element (such as a second element), it may be directly combined or connected to / attached to or connected to that other element, or combined or connected to / attached to / connected to that other element via a third element. Conversely, it will be understood that when an element (such as a first element) is referred to as being "directly combined" with / "directly attached to" another element (such as a second element) or "directly connected" with / "directly connected to" another element (such as a second element), no other element (such as a third element) is between that element and that other element.

[0013] As used herein, the phrase “configured (or set) to” may be used interchangeably with the phrases “suitable for,” “capable of,” “designed to,” “adapted to,” “manufactured as,” or “capable”, depending on the context. The phrase “configured (or set) to” does not inherently imply “specifically designed to” in hardware. Rather, the phrase “configured to” may indicate that a device can perform operations in conjunction with another device or component. For example, the phrase “processor configured (or set) to perform A, B, and C” may refer to a general-purpose processor (such as a CPU or application processor) that can perform operations by running one or more software programs stored in a memory device, or a dedicated processor (such as an embedded processor) for performing operations.

[0014] The terms and phrases used herein are provided only to describe some embodiments of this disclosure and are not intended to limit the scope of other embodiments of this disclosure. It should be understood that, unless the context clearly indicates otherwise, the singular form includes plural references. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this disclosure pertain. It will also be understood that terms and phrases (such as those defined in common dictionaries) should be interpreted as having the meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly defined herein. In some cases, the terms and phrases defined herein may be interpreted as excluding embodiments of this disclosure.

[0015] Examples of "electronic devices" according to embodiments of this disclosure may include at least one of the following: smartphones, tablet PCs, mobile phones, video phones, e-book readers, desktop PCs, laptop computers, netbooks, workstations, personal digital assistants (PDAs), portable multimedia players (PMPs), MP3 players, mobile medical devices, cameras, or wearable devices (such as smart glasses, head-mounted displays (HMDs), electronic clothing, electronic bracelets, electronic necklaces, electronic accessories, electronic tattoos, smart mirrors, or smartwatches). Other examples of electronic devices include smart home appliances. Examples of smart home appliances may include at least one of the following: television, digital video disc (DVD) player, audio player, refrigerator, air conditioner, vacuum cleaner, oven, microwave oven, washing machine, dryer, air purifier, set-top box, home automation control panel, security control panel, TV box (such as SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), smart speaker or speaker with integrated digital assistant (such as SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), game console (such as XBOX, PLAYSTATION, or NINTENDO), electronic dictionary, electronic key, camera, or electronic photo frame. Other examples of electronic devices include at least one of the following: various medical devices (such as various portable medical measurement devices (e.g., blood glucose measuring devices, heart rate measuring devices, or body temperature measuring devices), magnetic resonance angiography (MRA) devices, magnetic resonance imaging (MRI) devices, computed tomography (CT) devices, imaging devices, or ultrasound devices), navigation devices, global positioning system (GPS) receivers, event data loggers (EDR), flight data loggers (FDR), automotive infotainment devices, navigation electronics (such as navigation devices or gyrocompasses), avionics devices, safety devices, vehicle head units, industrial or household robots, automated teller machines (ATMs), point-of-sale (POS) devices, or Internet of Things (IoT) devices (such as light bulbs, various sensors, electricity or gas meters, sprinklers, fire alarms, thermostats, streetlights, toasters, fitness equipment, hot water tanks, heaters, or boilers). Other examples of electronic devices include at least a part of a piece of furniture or building / structure, electronic boards, electronic signature receivers, projectors, or various measuring devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). It should be noted that, according to various embodiments of this disclosure, the electronic device may be one or a combination of the devices listed above. According to some embodiments of this disclosure, the electronic device may be a flexible electronic device. The electronic devices disclosed herein are not limited to the devices listed above and may include new electronic devices as the technology develops.

[0016] In the following description, electronic devices are described with reference to the accompanying drawings according to various embodiments of the present disclosure. As used herein, the term "user" may mean a person using the electronic device or another device (such as an artificial intelligence electronic device).

[0017] Definitions for other specific words and phrases are provided throughout this patent document. It will be understood by those skilled in the art that, in many cases (and even in most cases), these definitions apply to the prior and future use of the words and phrases thus defined.

[0018] The descriptions in this application should not be construed as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the patent subject matter is defined solely by the claims. Furthermore, unless the exact phrase "method for..." is followed by a participle, none of the claims are intended to invoke 35 USC § 112(f). Any other terminology used in the claims (including, but not limited to, "mechanism," "module," "device," "unit," "component," "element," "building block," "device," "machine," "system," "processor," or "controller") is understood by the applicant to refer to structures known to a person skilled in the art and is not intended to invoke 35 USC § 112(f). Attached Figure Description

[0019] To gain a more complete understanding of this disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein similar reference numerals denote similar parts: Figure 1 An example network configuration including electronic devices is shown according to this disclosure; Figure 2 An example of point cloud segmentation is shown; Figure 3A and Figure 3B An example of point cloud completion is shown; Figure 4A and Figure 4B An example of point cloud hole filling is shown; Figure 5A and Figure 5B An example of point cloud meshing is shown; Figure 6A and Figure 6B An example octree process according to this disclosure is shown; Figure 7 An example comparison between image representation and language representation according to this disclosure is shown; Figure 8 An example process for projecting visual and linguistic features into a shared embedding space, according to this disclosure, is shown; Figure 9Aand Figure 9B An example 3D projection process according to this disclosure is shown; Figure 10A and Figure 10B An example of natural language 3D capture according to this disclosure is shown; Figure 11 An example process for natural language 3D search according to this disclosure is shown; Figure 12 The process for deriving natural language 3D search data according to this disclosure is illustrated; Figure 13 The following is shown in accordance with the present disclosure: Figure 12 Example output corresponding to the process; Figure 14 The following is shown in accordance with this disclosure: Figure 13 An example retrieval process for data points in a data structure; Figure 15 An example process for ingesting, deriving, and retrieving natural language 3D search data according to this disclosure is shown; Figure 16A and Figure 16B Example image segmentation according to this disclosure is shown; Figure 17 An example of a joint language-visual representation according to this disclosure is shown; Figure 18 An example of using cosine similarity according to an embodiment of this disclosure is shown; Figure 19 An example alternative process for ingesting, deriving natural language 3D search data, and retrieving data according to this disclosure is shown; Figure 20 The process of adding tags is shown, which adds tags to the combination. Figure 12 , Figure 13 and Figure 14 The discussion focuses on the octree of chairs and lamps, and uses it as an associated query. Figure 21 , Figure 21A and Figure 21B The component segmentation process according to this disclosure is illustrated; Figure 22 A selective filtering process according to this disclosure is illustrated; Figure 23 A selective filtering process according to this disclosure is illustrated; Figure 24 The filtering process according to this disclosure is illustrated; and Figure 25 An example method for natural language 3D data search according to this disclosure is shown. Detailed Implementation

[0020] The following discussion is described with reference to the accompanying drawings. Figures 1 to 25 Various embodiments of this disclosure are also described. However, it should be understood that this disclosure is not limited to these embodiments, and all changes and / or equivalents or substitutions thereof are also within the scope of this disclosure. Throughout the specification and drawings, the same or similar reference numerals may be used to refer to the same or similar elements.

[0021] As mentioned above, thanks to significant investments in robotics and autonomous driving, 3D capture technology has made tremendous progress over the past decade, with substantial reductions in the cost, size, complexity, and noise of sensing technologies. This has opened up unprecedented new areas and application scenarios for consumer electronics. The 3D data created by these technologies presents numerous unresolved new challenges in terms of privacy, user experience, and processing. 3D scanning is the process of capturing a physical object or environment to create a digital representation in the form of a 3D model. Many technologies exist for 3D scanning, including contact scanning technologies (such as zoomers) and optical scanning technologies (such as LiDAR, photogrammetry, and structured light).

[0022] Typically, 3D scanning produces point clouds, which are discrete sets of 3D points and potential appearance attributes (such as color and intensity). Point clouds are a convenient representation for many applications because they are easy to visualize and manipulate. However, point clouds also have some inherent limitations. For example, point clouds are discrete, meaning the number of points representing a scanned object or environment is limited. This can lead to incomplete representations because small features or details may not be captured. Point clouds are often large and difficult to compress. Point clouds are also unordered, meaning the points reflecting the location of the scanned object or environment do not have an inherent order.

[0023] Point clouds generated from 3D scanning are also susceptible to various artifacts that can affect the quality of 3D models. Due to the inherent properties of 3D capture and point clouds, they are not suitable for a wide range of tasks, such as computer-aided design (CAD), augmented reality, simulation, and aesthetic representation. Therefore, one task for applications that wish to utilize point cloud data is to transform the raw point cloud capture into a semantically equivalent continuous representation (such as a mesh or radiation field). These are not easy tasks, and performing them requires significant technical expertise as well as sufficient computational power to ingest and process the data.

[0024] Capturing point clouds is a more challenging process than two-dimensional (2D) representations because 2D representations are both continuous and ordered images, meaning that algorithms for visualizing and analyzing these images are relatively easier. As mentioned above, point cloud data can be affected by various artifacts that can impact the quality of the resulting 3D model. These artifacts can include noise, holes, and outliers. Noise can be caused by inaccuracies during the scanning process, such as sensor noise or environmental interference. Holes can occur when parts of the scanned object or environment are not visible to the scanner (such as occluded areas or poorly lit areas). Outliers can occur when points that do not belong to the scanned object or environment are present (such as debris or reflections). Similar to 2D imaging, scan resolution decreases with distance.

[0025] This disclosure improves the performance of 3D algorithms by leveraging the fact that neighboring points in 3D space are most often associated in a meaningful way, resulting in spatial data structures that allow for more efficient storage and retrieval of 3D data. Many possible approaches exist for this, among which octrees are one example. An octree is a data structure used for efficiently representing three-dimensional space. An octree divides space into increasingly smaller cubes, each of which can be empty or occupied by points or objects. The root of the tree represents the entire space, while the leaves represent the smallest possible cubes. Octrees have applications in computer graphics and computer vision for tasks such as collision detection, ray tracing, and object rendering, but can also be used in Geographic Information Systems (GIS) for spatial indexing and data compression. By using octrees, the computational load required to perform the aforementioned segmentation, completion, hole-filling, and meshing tasks can be reduced, resulting in faster and more efficient algorithms.

[0026] Figure 1 An example network configuration 100 including electronic devices according to this disclosure is shown. Figure 1 The embodiment of network configuration 100 shown is for illustrative purposes only. Other embodiments of network configuration 100 may be used without departing from the scope of this disclosure.

[0027] According to embodiments of this disclosure, electronic device 101 is included in network configuration 100. Electronic device 101 may include at least one of bus 110, processor 120, memory 130, input / output (I / O) interface 150, display 160, communication interface 170, and sensor 180. In some embodiments, electronic device 101 may not include at least one of these components, or at least one other component may be added. Bus 110 includes circuitry for connecting components 120-180 to each other and for transmitting communication (such as control messages and / or data) between components.

[0028] Processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In some embodiments, processor 120 includes one or more of a central processing unit (CPU), application processor (AP), communication processor (CP), or graphics processing unit (GPU). Processor 120 is capable of performing control and / or performing operations or data processing related to communication or other functions on at least one of the other components of electronic device 101.

[0029] Memory 130 may include volatile memory and / or non-volatile memory. For example, memory 130 may store commands or data associated with at least one other component of electronic device 101. According to embodiments of this disclosure, memory 130 may store software and / or program 140. Program 140 includes, for example, kernel 141, middleware 143, application programming interface (API) 145, and / or application program (or “application”) 147. At least a portion of kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).

[0030] Kernel 141 can control or manage system resources (such as bus 110, processor 120, or memory 130) used to perform operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). Kernel 141 provides an interface that allows middleware 143, API 145, or application 147 to access various components of electronic device 101 to control or manage system resources. These functions can be performed by a single application or by multiple applications, each performing one or more of these functions. For example, middleware 143 can act as a repeater to allow API 145 or application 147 to communicate data with kernel 141. Multiple applications 147 can be provided. Middleware 143 is able to control work requests received from applications 147, such as by prioritizing the use of system resources (such as bus 110, processor 120, or memory 130) of electronic device 101 for at least one of the multiple applications 147. API 145 is an interface that allows application 147 to control functions provided from kernel 141 or middleware 143. For example, API 145 includes at least one interface or function (such as commands) for file control, window control, image processing, or text control.

[0031] I / O interface 150 serves as an interface for transmitting, for example, commands or data input from a user or other external device to other components of electronic device 101. I / O interface 150 can also output commands or data received from other components of electronic device 101 to the user or other external device.

[0032] Display 160 includes, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot light-emitting diode (QLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. Display 160 can also be a depth-sensing display (such as a multi-focus display). Display 160 is capable of displaying various content to a user, such as text, images, videos, icons, or symbols. Display 160 may include a touchscreen and can receive input such as touch, gestures, proximity, or hover input using an electronic pen or a user's body part.

[0033] Communication interface 170, for example, enables communication between electronic device 101 and external electronic devices (such as first electronic device 102, second electronic device 104, or server 106). For instance, communication interface 170 can be connected to network 162 or 164 via wireless or wired communication to communicate with external electronic devices. Communication interface 170 can be a wired transceiver or a wireless transceiver or any other component for transmitting and receiving signals.

[0034] Wireless communication can use at least one of the following as a communication protocol: WiFi, LTE, LTE-A, 5G, millimeter wave or 60 GHz wireless communication, wireless USB, CDMA, WCDMA, UMTS, Wi-Fi, or GSM. Wired connections may include at least one of the following: USB, HDMI, RS-232, or POTS. Network 162 or 164 includes at least one communication network, such as a computer network (e.g., a local area network (LAN) or wide area network (WAN)), the Internet, or a telephone network.

[0035] Electronic device 101 also includes one or more sensors 180 that can measure physical quantities or detect the activation state of electronic device 101 and convert the measured or detected information into electrical signals. Sensor 180 may also include one or more buttons for touch input, one or more microphones, a depth sensor, a gesture sensor, a gyroscope or gyroscope sensor, an atmospheric pressure sensor, a magnetic sensor or magnetometer, an accelerometer or accelerometer, a grip sensor, a proximity sensor, a color sensor (such as an RGB sensor), a biophysical sensor, a temperature sensor, a humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyography (EMG) sensor, an electroencephalography (EEG) sensor, an electrocardiogram (ECG) sensor, an infrared (IR) sensor, an ultrasonic sensor, an iris sensor, or a fingerprint sensor. Sensor 180 may also include an inertial measurement unit, which may include one or more accelerometers, gyroscopes, and other components. Furthermore, sensor 180 may include control circuitry for controlling at least one of the sensors included herein. Any of these sensors 180 may be located within electronic device 101.

[0036] In some embodiments, the first external electronic device 102 or the second external electronic device 104 may be a wearable device or an electronically mountable wearable device (such as an HMD). When electronic device 101 is mounted in electronic device 102 (such as an HMD), electronic device 101 can communicate with electronic device 102 via communication interface 170. Electronic device 101 can be directly connected to electronic device 102 to communicate with electronic device 102 without involving a separate network. Electronic device 101 may also be an augmented reality wearable device (such as glasses) including one or more imaging sensors.

[0037] The first external electronic device 102, the second external electronic device 104, and the server 106 may be devices of the same or different types as electronic device 101. According to a specific embodiment of this disclosure, server 106 includes a combination of one or more servers. Furthermore, according to a specific embodiment of this disclosure, all or some of the operations performed on electronic device 101 may be performed on another electronic device or a plurality of other electronic devices (such as electronic devices 102 and 104 or server 106). Furthermore, according to a specific embodiment of this disclosure, when electronic device 101 is required to perform some functions or services automatically or on request, electronic device 101 may request another device (such as electronic devices 102 and 104 or server 106) to perform at least some of the functions associated with it, rather than performing the function or service alone, or electronic device 101 may request another device (such as electronic devices 102 and 104 or server 106) to perform at least some of the functions associated with it in addition to running the function or service. Another electronic device (such as electronic devices 102 and 104 or server 106) is capable of performing the requested function or additional function and transmitting the result of the performance to electronic device 101. Electronic device 101 can provide the requested function or service by processing the received result as is or additionally. For this purpose, cloud computing, distributed computing, or client-server computing technologies can be used, for example. Although Figure 1 The electronic device 101 is shown to include a communication interface 170 for communicating with an external electronic device 104 or a server 106 via a network 162 or 164, but according to some embodiments of this disclosure, the electronic device 101 may operate independently without a separate communication function.

[0038] Server 106 may include components 110-180 (or suitable subsets thereof) that are the same as or similar to those in electronic device 101. Server 106 may support driving electronic device 101 by performing at least one of the operations (or functions) implemented on electronic device 101. For example, server 106 may include a processing module or processor that can support processor 120 implemented in electronic device 101.

[0039] although Figure 1 An example of a network configuration 100 including electronic device 101 is shown, but it is possible to modify it. Figure 1 Various changes can be made. For example, network configuration 100 can include any number of each component in any suitable arrangement. Typically, computing and communication systems have a wide variety of configurations, and Figure 1 This disclosure is not intended to limit the scope to any particular configuration. Furthermore, although… Figure 1 An operating environment is shown in which the various features disclosed in this patent document can be used, but these features can be used in any other suitable system.

[0040] Figures 2 to 5B Example transformations of 3D point cloud data according to this disclosure are shown. Many transformations can be performed on 3D point cloud data before use. Outlier removal, hole filling, completion, meshing, and segmentation can all be performed as initial processing steps. Figure 2 Example 200 of segmentation is shown, in which points within a point cloud are logically segmented based on features of representation, such as street 201, cars parked on the street 202, and buildings adjacent to the street 203 in the example shown. Figure 3A and Figure 3B Example 300 of the completion is shown, where the point cloud is identified as an incomplete representation and additional points are added. For example, Figure 3A The portion of the floor lamp shown has points added based on projection symmetry to produce... Figure 3B The point cloud shown in the image. Figure 4A and Figure 4B Example 400 of void filling is shown. Based on the continuity of the projected surface, for... Figure 4A The point cloud of the rabbit shown is perceived as missing points, and these missing points are added to form... Figure 4B The point cloud shown in the image. Figure 5A and Figure 5B Example 500 of the gridded model is shown, where, Figure 5A The (complete) outer surface of the rabbit shown in the image is transformed into... Figure 5B The image shows a smaller polygonal surface shape. Once the mesh is built, traditional editing and rendering techniques can be used to render the 3D data for application use.

[0041] although Figures 2 to 5B An example of transformation of 3D point cloud data is shown, but it is possible to modify... Figures 2 to 5B Various changes can be made. For example, the 3D point cloud shown can be various objects, and Figures 2 to 5B The objects shown are merely examples. Typically, 3D point cloud data can have a wide variety of arrangements, and Figures 2 to 5B This disclosure is not intended to limit the scope of any particular arrangement of 3D point cloud data.

[0042] Figure 6A and Figure 6B An example octree process 600 according to this disclosure is shown. As described above, this disclosure improves the performance of 3D algorithms by leveraging the fact that neighboring points in 3D space are most often associated in a meaningful way, to produce spatial data structures that allow for more efficient storage and retrieval of 3D data. There are many possible approaches to this, among which the octree is one example. An octree is a data structure used for efficiently representing three-dimensional space. Figure 6AAs shown, process 600 involves using an octree to divide the 3D space into increasingly smaller cubes, each of which can be empty or occupied by points or objects. For example... Figure 6B As shown, process 600 also includes creating an octree spatial data structure corresponding to the partitioned 3D space. This octree spatial data structure is hierarchical, with each level of the tree representing a different level of detail. The root of the tree represents the entire space, while the leaves represent the smallest possible cubes. As mentioned above, octrees have applications in computer graphics and computer vision for tasks such as collision detection, ray tracing, and object rendering, but they can also be used in Geographic Information Systems (GIS) for spatial indexing and data compression. By using octrees, the computational load required to perform the aforementioned partitioning, completion, hole-filling, and meshing tasks can be reduced, resulting in faster and more efficient algorithms.

[0043] The octree concept described above can be applied to 3D search in natural language. To process text (e.g., "natural language") in a machine learning setting, a mathematical representation that preserves the meaning of the objects must be designed—this is called a language or word embedding. One approach would involve simply browsing an English dictionary and assigning an ordinal identifier to each word (aardvark=00000001, etc.), but this has many drawbacks. This low-dimensional representation only encodes the alphabetical order of words and does not indicate meaning, frequency, part of speech, context, etc. The fact that "aardvark" is assigned 0000000000000000000000000000001 and the word "final" is assigned 00000000000000010001000101110000 encodes very little useful information. To improve this, the encoding system can be extended to more dimensions. By projecting words into a higher-dimensional space (e.g., R300), the geometric relationships between points encode a rich set of information, such as part of speech, word origin, lexical semantics, etc. In this space, mathematics can be effectively "completed" using the word. For example, in an embedded space, king + woman can equal queen, etc.

[0044] While useful in many applications, such as search, word embeddings are not rich enough for many tasks. Because word embeddings are context-insensitive, a word like "bat" will map to the same point in the embedding space, regardless of whether it refers to an animal or an object used for playing baseball. To address this issue, sentence embedding models have been developed that encode entire sentences as meaningful vector representations. For example, these embeddings allow determining that "What is the capital of the United States?" is more closely associated with "Washington, D.C. [...] is the capital of the United States" than "The death penalty [...] already exists in the United States [...]".

[0045] Because these embedding models are derived from massive text corpora, the embeddings capture many nuances and variations of human language and can evolve into “common sense.” For example, the embedding for the phrase “make me dinner” would exist in a space of words like “kitchen,” “panel,” “microwave,” and “food,” as these words frequently appear alongside the word “dinner” in the training corpus. This allows for the deployment of systems with a rich yet superficial understanding of the world.

[0046] Just like natural language, image data must be encoded in a format suitable for machine learning tasks. Digital images come in many shapes, sizes, and formats, and naturally encode a very rich set of information. The task of an image encoder is to normalize image data and represent it in a simplified form that preserves the "meaning" of the data. Does this embedding represent what is actually captured, which depends entirely on the task? Embeddings can represent objects in an image, the style of an artwork, the type of camera used to capture the photo, and so on.

[0047] Figure 7 An example comparison 700 between image representation and language representation according to this disclosure is shown. For example... Figure 7 As shown, image embeddings can be jointly learned with sentence embeddings, enabling direct comparison between linguistic and image representations. This allows for the identification of an image as a dog, rather than an image of a black dog. Previous image classification methods were limited in generality by the need for researchers to manually label and model large visual concept spaces. This joint linguistic representation allows for the creation of very nuanced and sensitive image classifiers using a vast pre-annotated image corpus.

[0048] Figure 8 An example process 800 for projecting visual and linguistic features into a shared embedding space according to this disclosure is shown. Research on visual language models (VLMs) has been driven by use cases in a wide range of fields, such as search, e-commerce, customer support, robotics, and media accessibility. In short, as... Figure 8 As shown, these models work by interweaving and projecting visual and linguistic features into a shared embedding space, creating joint representations that can be used by standard machine learning models, such as multilayer perceptron (MLP) models. Due to the existence of large, readily available corpora of annotated 2D media, these models typically only work with 2D data, such as images and videos. Deep learning foundational models, such as convolutional neural networks (CNNs), are well understood and applicable to 2D images.

[0049] Figure 9A and Figure 9B An example 3D projection process 600 according to this disclosure is illustrated. This is to extend these models into 3D space, such as... Figure 9A and Figure 9BAs shown, an algorithm is applied to explore space and project 3D data onto 2D from multiple viewpoints. Sample selection is a crucial factor in both result quality and algorithm time performance. Each 2D sample is then fed into the network. However, the above method can encounter problems because 3D data is a rich data format that encodes structural and visual information over arbitrarily large spatial areas. Every facet of a room or object can be captured, measured, and communicated through this data. As with any such complex data, navigating and understanding it is a challenging task for both experts and laypeople, especially when performed on devices with 2D touch interfaces, such as phones or tablets. Even in desktop environments, experts use a “3D mouse” to improve navigation efficiency.

[0050] Furthermore, searching large sets of "non-verbal data" is challenging, especially when the data cannot be readily examined, such as being presented with 100 movie files and asked to find a car in a single frame. Additionally, users are reluctant to share complete scans of their homes with third parties: this data could expose many private details about an individual, such as medical conditions, hobbies, relationships, hygiene habits, wealth, etc. Moreover, the files generated by various 3D capture technologies are extremely large, often reaching gigabytes in size. Third parties wishing to utilize this data will find it difficult to ingest, process, understand, and store it. Furthermore, many users live in areas with poor internet access, reducing the experience and feasibility of such applications.

[0051] This disclosure provides a system that allows a user / application programming interface (API) / agent (collectively referred to herein as "user") to capture navigation and extract data from 3D using a build language or natural language interface. For example, Figure 10A and Figure 10B An example of natural language 3D capture according to this disclosure is shown. As provided in this disclosure, the user can make potentially ambiguous open requests (such as... Figure 10A and Figure 10B The query will automatically find a subset of data that matches the query (as shown in the "good places to hang paintings").

[0052] Depending on the use case, further processing can be performed to add / remove information, refine the rendered view, generate relevant metrics, etc. This feature simplifies interaction with 3D data, allows users to protect their privacy, and significantly reduces the amount of data sent or rendered to the screen over wires. In this example, the user captures 3D information with various devices and stores it on a secure (local or remote) device, where the 3D information is then preprocessed to generate a spatial database that allows retrieval of subsets of the point cloud using language or visual queries. For example, given a point cloud of an Ikea store, the user can query "long white sofa," or provide an image of a similar sofa and retrieve a related subset of that point cloud. Figure 10A In specific example 1001, the natural language expression "good places to hang paintings" is provided, and a result can be identified as follows: Figure 10A The large open-walled space shown, and another result that can be identified is as follows: Figure 10B Example 1002 shows a feature (such as a fireplace) and indicates the space above that feature (whether occupied or empty).

[0053] although Figure 10A and Figure 10B An example of natural language 3D capture according to this disclosure is shown, but it is possible to modify it further. Figure 10A and Figure 10B Various changes can be made. For example, the natural language provided, the 3D scene involved, and the determination of appropriate aspects of the 3D scene that conforms to the natural language can vary depending on the implementation method or the specific use case.

[0054] Figure 11 An example process 1100 for natural language 3D search according to this disclosure is shown. For ease of explanation, Figure 11 The process 1100 shown in the diagram is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The electronic device 101 in the network configuration 100 is supported. However, such as when process 1100 is implemented on or supported by server 106, Figure 11 The process 1100 shown can be used with any other suitable device and in any other suitable system.

[0055] like Figure 11 As shown, 3D data is captured and preprocessed via preprocessing function 1101 to generate a complete mesh with as little noise and artifacts as possible. Then, the system captures numerous 2D samples from the point cloud via scene sampling function 1102, fully exploring the space and generating a representative summary of the model. Next, each of the 2D samples is fed into segmentation function 1103, which identifies and extracts a subset of the image ((x,y,z), (w,x,y,z)). The image samples are then passed to a pre-trained image encoder that performs embedding function 1104, embedding the data into a shared latent space. Finally, the source point cloud ID, the point cloud subsets contained in the 2D samples, and the embeddings are stored in a spatial database for later retrieval.

[0056] although Figure 11 An example process 1100 for natural language 3D search is shown, but it is possible to... Figure 11 Make various changes. For example, Figure 11The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired.

[0057] Figure 12 A process 1200 for deriving natural language 3D search data according to this disclosure is shown. For ease of explanation, Figure 12 The process 1200 shown in the diagram is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The network configuration 100 is supported by electronic device 101. However, such as when process 1200 is implemented on server 106 or supported by server 106, Figure 12 The process 1200 shown can be used with any other suitable device and in any other suitable system.

[0058] What will be understood is... Figure 12 The embodiments shown are for illustrative purposes only. Other embodiments of process 1200 may be used without departing from the scope of this disclosure. Process 1200 begins with multiple views 1201, 1202 of an environment, which in the illustrated example includes an armchair near a floor lamp. Each view 1201, 1202 is then segmented into portions 1203, 1204 corresponding to one object (armchair) in the environment and portions 1205, 1206 corresponding to another object (floor lamp). Each 3D point from the segmented images 1203-1204 for the objects is then assigned an embedding vector 1207, 1208 for the entire segment containing the segmented images.

[0059] For each frame 1209, it is checked whether any point is also included in a previously processed frame. If not (e.g., for the first frame processed), all points receive the same new group identifier (ID). If there is overlap between two frames 1210, all points from the currently processed frame are assigned the previously assigned group ID. Embeddings from the overlapping images 1207 and 1208 are then unified by averaging 1211.

[0060] although Figure 12 An example process 1200 for exporting natural language 3D search data is shown, but it is possible to... Figure 12 Make various changes. For example, Figure 12 The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired.

[0061] Figure 13 The following is shown in accordance with the present disclosure: Figure 12The example output corresponding to process 1200 is 1300. For ease of explanation, Figure 13 The output 1300 shown is described as being generated by... Figure 1 The electronic device 101 in the network configuration 100 is provided. However, such as when process 1200 is implemented on or supported by server 106, Figure 13 The output 1300 shown can be used with any other suitable device and in any other suitable system.

[0062] Figure 12 Process 1200 generates a set of octree indexes and corresponding embeddings for each group ID. Following process 1200, a data structure 1302 is provided for the grouping of capture points and image embeddings. Since many neighboring points will share group embeddings, the octree data structure is used to compress this data structure 1302. Therefore, voxels can be stored instead of individual points.

[0063] although Figure 13 An example output of 1300 is shown, but it is possible to... Figure 13 Make various changes. For example, Figure 13 The various components and data within can be combined, further subdivided, copied, or rearranged as needed. Additionally, one or more additional components and data may be included if required or desired.

[0064] Figure 14 The following is shown in accordance with this disclosure: Figure 13 Example retrieval process for data points in a data structure 1400. For ease of explanation, Figure 14 The process 1400 shown in the diagram is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The network configuration 100 is supported by electronic device 101. However, such as when process 1400 is implemented on server 106 or supported by server 106, Figure 14 The output 1400 shown can be used with any other suitable device and in any other suitable system.

[0065] like Figure 14 As shown, to retrieve points, a query is first embedded in step 1401 (“chair next to the lamp” in the example shown), and then a scan is performed on the embedded entries within data structure 1302. Entries with similarity scores below an implementation-defined threshold are then discarded. However, when an entry exceeds the threshold, that entry and all other entries in the group are returned as query result 1402. The result of querying data structure 1302 is a set of split points 1404 that match the query. The final group can then be selected based on any of several metrics, such as aggregate confidence, size, proximity, etc.

[0066] although Figure 14 An example retrieval process 1400 for data points is shown, but it is possible to... Figure 14 Make various changes. For example, Figure 14 The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired.

[0067] Figure 15 An example process 1500 for ingesting, deriving, and retrieving natural language 3D search data according to this disclosure is shown. For ease of explanation, Figure 15 The process 1500 shown in the diagram is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The network configuration 100 is supported by electronic device 101. However, such as when process 1500 is implemented on server 106 or supported by server 106, Figure 15 The output 1500 shown can be used with any other suitable device and in any other suitable system.

[0068] Figure 15 The embodiments shown are for illustrative purposes only. Other embodiments of process 1500 may be used without departing from the scope of this disclosure. The complete ingestion and retrieval process 1500 is logically arranged into three blocks: ingestion 1501; processing 1502; and retrieval 1503. Ingestion 1501 and processing 1502 generally correspond to the process by... Figure 12 The process shown produces Figure 13 The data structure, and retrieval usually corresponds to the data structure of... Figure 14 The process shown is for Figure 13 The data structure is used for operations.

[0069] The acquisition 1501 begins by receiving point cloud data 1504. This process is not limited to any specific capture technique, only requiring that the point cloud data 1504 contains color data. Optionally, the process also employs the real-world camera pose used during data acquisition. During acquisition 1501, processing and mesh generation 1505 are performed on the point cloud data 1504, and scene sampling 1506 is performed on the output of this processing and mesh generation 1505.

[0070] Scene sampling 1506 projects 3D data into 2D samples 1507, which can be used in conjunction with 3D samples 1508 and camera pose 1509. During processing 1502, each 2D sample is passed to segmentation algorithm 1510, which identifies and extracts objects from the image, generating fragments 1511 corresponding to the objects. The extracted image fragments 1511 are then encoded using a pre-trained encoder 1512, which generates an embedding 1513 for each encoded image fragment. A mask 1514 based on fragment 1511 is applied to the 3D samples 1508 to generate 3D fragments 1515. For each 2D sample, a check is performed by group tracker 1516 to determine whether any of the 3D points corresponding to the 3D points have been processed. If so, any new points contained in the 2D sample are added to the existing group ID 1517; otherwise, a new group identifier 1517 is created for the 2D sample.

[0071] Embedsions 1513, point and group IDs 1517 are stored in a spatial data structure 1518 such as a K-tree (kd-tree), octree, etc. During retrieval 1503, users and applications can query this spatial data structure 1518 by passing an input language query 1519 or a visual query 1521 through a language encoder 1520 or an image encoder 1522 (operating in a manner corresponding to the image encoder 1512). Once the query is encoded, the database can be searched, and some similarity measure (such as cosine similarity) can be applied to find candidate entries. These entries are then returned along with associated metadata (image embeddings) 1523 and point cloud data segments 1524.

[0072] Obtaining representative 2D samples from the raw data during scene sampling 1506 is a crucial component of the final solution. A sample-efficient approach is needed, meaning neither too many nor too few samples are collected to capture all the relevant features of the input point cloud. Too few (or unrepresentative) samples will result in poor recall. Too many samples require significantly more computational resources. Several different approaches can be adopted.

[0073] Scene sampling can be easily achieved using real-world camera poses for capturing 3D space. Depending on the capture method, there will be instances that can be captured via Simultaneous Localization and Mapping (SLAM) or other existing photogrammetric techniques (e.g., Figures 9A to 9B Alternatively, multiple camera poses can be extracted (considering near / far clipping planes). This has the advantage of close alignment with the user's focus; the spatial region of interest to the user will likely receive more and higher-resolution images. This method also allows the algorithm to better calibrate the spatial scale of the focused object.

[0074] Scene sampling methods can be improved by tracking which parts of the point cloud have been sampled and at what resolution. Using real-world poses as a starting point and readily available spatial hashing algorithms, the space can be explored while simultaneously tracking both sampled and unsampled content.

[0075] When performing segmentation 1510 for the purpose of embedding 1513, the embedding models discussed above do not provide a method for identifying regions of interest or partitioning objects within a scene; instead, they only provide directly comparable mathematical representations of objects of different categories. Existing 2D semantic segmentation models can be used to identify and extract subsets of 3D data for indexing. These models take many forms, and some generate segments with discrete sets of class labels for segmentation (e.g., “tree,” “dog,” “chair”), while others simply produce masks without any indication of the nature of the segmented objects.

[0076] Within process 1500, no assumptions are made about the type of the segment being executed. Depending on the application, process 1500 can employ semantic segmentation, instance segmentation, component segmentation, etc. Applications can even choose to use a set of segmentation models to achieve a "higher resolution" search. For example, Figure 16A and Figure 16B An example image segmentation according to this disclosure is shown. In the segmentation... Figure 16A Image 1601 to produce Figure 16B In segment 1602, component segmentation is not used (see below). Figure 21 and Figures 21A to 21B (Discussion) You can search for "cat" or "cat tail", but you cannot search for "tail" alone.

[0077] Refer again Figure 15 The model, encoded by the image encoder 1512 / 1522 or the language encoder 1520, can be trained in various ways using a variety of data sources. The only constraint is that the model produces a joint-linguistic visual representation. Figure 17 An example joint language-visual representation 1700 according to this disclosure is shown. The generation of the joint language-visual representation can be accomplished using a corpus of text-image pairs obtained from a web crawler. These pairs are then fed through a transformer network and optimized using a contrastive loss function such as maximum interval contrastive loss. As explained in this disclosure, these models allow both text and images to be projected into a shared embedding space that makes the text and images directly contrastable. The exact choice of model may depend on a variety of factors, such as the specific use case, cost, data availability, data licensing issues, etc. Various models can be used for this task with minimal fine-tuning. These models can be trained on a very broad range of visual concepts, capturing data from satellite imagery to social media posts. Alternatively, models can be specifically trained to focus on capturing visual concepts common in home and social settings.

[0078] To retrieve embedded vectors from the spatial data structure 1518, some concept of similarity is needed. In this setting, direct element-wise comparison is practically useless because any slight change in the input (e.g., altering a single pixel) will slightly perturb the encoding, meaning that "cat" and "feline" will not map to exactly the same value. Instead, one of many similarity measures can be used, which models similarity as some distance measure in Euclidean space. For example, Figure 18 Example 1800 illustrating the use of cosine similarity according to embodiments of the present disclosure is shown. Using cosine similarity yields good results. Simply put, cosine similarity measures the angle between two normalized embedding vectors and maps the value between 1 and π using a cosine function.

[0079] Once a subset of the point cloud is identified, classic 3D modeling and rendering techniques can be used to generate new views of the data. Given knowledge of the intrinsic matrix of the virtual camera, trigonometric methods can be used to find the optimal viewing distance. Finding the correct viewpoint is somewhat complex due to several confounding factors that make the optimal choice less obvious. First, at any given viewpoint, non-target geometry may occlude the view of the region of interest. Second, there is a possibility that the mesh reconstruction method is not optimal, leading to the construction of meshes using various completion and hole-filling techniques. In some embodiments, these completed or filled areas may not be aesthetically pleasing or may not capture any real information, and therefore may not be presented to the user. In short, the chosen viewpoint should be “natural,” matching the angle at which the user is most accustomed to seeing objects daily. For example, when viewing a sofa, a view focusing on the underside of the sofa should not be chosen. Third, due to the sensitive nature of the classifier, queries on the spatial data structure 1518 can return many valid results. To find the most suitable viewpoint, heuristics collected from the data capture process can be utilized, and in the absence of such heuristics, training data can be used. Similar to 2D sampling, known camera poses that were initially used to generate 3D data can be utilized.

[0080] although Figure 15 An example process 1500 for ingesting, exporting, and retrieving natural language 3D search data is shown, but it is applicable to... Figure 15 Make various changes. For example, Figure 15 The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired. Additionally, although shown as a series of steps, Figure 15 The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0081] Figure 19 An example alternative process 1900 for ingesting, deriving natural language 3D search data, and retrieving data according to this disclosure is shown. For ease of explanation, Figure 19 The process 1900 shown in the diagram is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The network configuration 100 is supported by electronic device 101. However, such as when process 1900 is implemented on server 106 or supported by server 106, Figure 19 The process 1900 shown can be used with any other suitable device and in any other suitable system.

[0082] Figure 19 The embodiments shown are for illustrative purposes only. Other embodiments of process 1900 may be used without departing from the scope of this disclosure. Figure 15 Building upon the described functionality, some changes can be introduced to enable technologies such as Large Language Models (LLMs) to interface with 3D data. This will allow the intelligent assistant to fulfill requests (such as "turn on the red light next to the sofa").

[0083] To achieve this, processing 1902 introduces a segmentation model, which also generates class labels 1950 for each segment produced. These class labels 1950 can be "chair", "table", "lamp", etc. These class labels 1950 are then stored along with each entry in the group, allowing LLM 1951 to know exactly what the grouped objects are, which is not clear from individual image embeddings.

[0084] Once a subset is extracted from the spatial data structure 1518, structures such as oriented bounding boxes can be computed, and the scene can be represented in a plain text format (High-Level Representation 1952) that is understandable by LLM 1951. High-Level Representation 1952 can be directly read and understood by LLM 1951, or a code library interfacing with LLM 1951 can be provided. For example, Figure 20 The addition process 2000 is shown, which adds tags to the combination. Figure 12 , Figure 13 and Figure 14 The discussion focuses on the octree of chairs and lamps, and uses related queries.

[0085] although Figure 19 An example alternative process 1900 is shown for ingesting, deriving, and retrieving natural language 3D search data, but it is applicable to... Figure 19 Make various changes. For example, Figure 19The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired. Additionally, although shown as a series of steps, Figure 19 The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0086] Figure 21 , Figure 21A and Figure 21B The component segmentation process 2100 according to this disclosure is shown. For ease of explanation, Figure 21 , Figure 21A and Figure 21B The process 2100 shown in the diagram is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The network configuration 100 is supported by electronic device 101. However, such as when process 2100 is implemented on server 106 or supported by server 106, Figure 21 , Figure 21A and Figure 21B The process 2100 shown can be used with any other suitable device and in any other suitable system.

[0087] Apart from Figure 15 and Figure 19 In addition to the above process, a segmentation process 2100 can also be performed. Because each point / voxel is associated with an image embedding, the above techniques can be used to perform part segmentation. Figure 21 As shown, when retrieving points and embeddings from spatial data structure 1518 for a query ("front of the TV" in the example shown), the confidence is not uniformly distributed? This is to be expected when considering the specificity of the user query. The front of the TV (e.g., Figure 21A The image depicted should not be the same as the image on the back of the television (such as...). Figure 21B (The depicted data) is a perfect match. In fact, the confidence gradient is extracted from each group of points / voxels. This fact can be used to identify finer-grained features of the point cloud in one of several ways. First, a simple threshold filter can be applied to eliminate all data below a specified value. This might work in some applications, but it's not sensitive enough and can result in many outliers, depending on the data and query. To address this, more sophisticated methods might involve clustering. Depending on the use case, sophisticated clustering algorithms such as k-means clustering, density-based hierarchical spatial clustering (HDBScan) for noisy applications, mean-shift clustering, etc., can be used.

[0088] although Figure 21 , Figure 21A and Figure 21BAn example component segmentation process 2100 is shown, but it is possible to perform a different process. Figure 21 , Figure 21A and Figure 21B Make various changes. For example, Figure 21 , Figure 21A and Figure 21B The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired. Additionally, although shown as a series of steps, Figure 21 , Figure 21A and Figure 21B The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0089] Figure 22 A selective filtering process 2200 according to this disclosure is illustrated. For ease of explanation, Figure 22 The process 2200 shown is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The electronic device 101 in the network configuration 100 is supported. However, such as when process 2200 is implemented on or supported by server 106, Figure 22 The process 2100 shown can be used with any other suitable device and in any other suitable system.

[0090] Process 2200 can be Figure 15 or Figure 19 The following version of the process performs selective filtering of the underlying data to enable selective sharing of point cloud data with third parties based on the user's pre-specified privacy preferences. This process is similar to that of... Figure 11 The process is illustrated below. First, 3D data is captured and preprocessed at step 2201. The encrypted 3D data is then provided to service 2202, which handles third-party requests 2203. Utilizing a hypothetical sensitive natural language classifier architecture, the third party can request a highly precise view of the user data via natural language or constructed language queries. For example, an e-commerce platform wishing to generate a preview of a painting in the home could request "good places to hang paintings," and once identified, only this subset will be sent to the Visual Language Model (VLM) 2204. Users can use natural language in their user privacy preferences 2205 to ensure that specific types of content (such as "medical devices," "hobbies," etc.) are not shared. Using the constrained 3D data, the process then combines the above at step 2206. Figure 15 or Figure 19 The described method performs segmentation, mesh creation, and simplification, and generates output 2207, which in this example represents the defined location of the hanging painting.

[0091] although Figure 22 An example selective filtering process 2200 is shown, but it is possible to... Figure 22 Make various changes. For example, Figure 22 The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired. Additionally, although shown as a series of steps, Figure 22 The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0092] Figure 23 A selective filtering process 2300 according to this disclosure is illustrated. For ease of explanation, Figure 23 The process 2300 shown in the diagram is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The electronic device 101 in the network configuration 100 is supported. However, such as when process 2300 is implemented on server 106 or supported by server 106, Figure 23 The process 2300 shown can be used with any other suitable device and in any other suitable system.

[0093] Figure 23 Is with Figure 22 The flowchart corresponds to the process described above. That is, utilizing the same ingestion method 1501 and spatial database 1518 described above, a third-party query 2219 operates on 3D point cloud data 2324 and embeddings 2323, and optionally, operates on room information 2325 (such as labels or advanced representations) to add an additional box 2304 to perform filtering during retrieval 2303. This filtering block 2304 includes an embedding matching filter 2360 and a post-processing stage 2361. The embedding filter 2360 receives the embeddings 2323 retrieved from the spatial database 1518, along with any additional associated metadata, as input. Based on user preferences 2205, the system determines whether the received content can be shared with a third party. For each region of the input 3D data processed by the embedding filter 2360, points rejected by the embedding filter 2360 are discarded; otherwise, a fully segmented model is produced as output.

[0094] although Figure 23 An example selective filtering process 2300 is shown, but it is possible to... Figure 23 Make various changes. For example, Figure 23The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired. Additionally, although shown as a series of steps, Figure 23 The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0095] Figure 24 The filtering process 2400 according to this disclosure is illustrated. For ease of explanation, Figure 24 The process 2400 shown in the diagram is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The network configuration 100 is supported by electronic device 101. However, such as when process 2400 is implemented on server 106 or supported by server 106, Figure 24 The process 2400 shown can be used with any other suitable device and in any other suitable system.

[0096] Filtration process 2400 can be combined with Figure 22 and Figure 23 The process is used in combination. To filter points from the point cloud, the inherent properties of the embeddings are relied upon to match them against items in a user-defined blacklist. Users can provide natural language cues (such as "medical device," "children's items," etc.) as reference points in the embedding space. To make this filter more robust, another language model can be used to extend the user-provided preferences. Input such as "medical device" can be parsed into a series of cues, including "pulse oximeter," "CPAP machine," "blood glucose monitor," etc. Complicating matters further, users may provide irrelevant information in the cues, such as instructions or punctuation that may reduce the model's sensitivity. To address this, classic natural language processing techniques can be used to label various words as nouns or named entities and extract these words from the user input.

[0097] In another example of a use case addressing the above process, the robot agent typically has a very narrow set of capabilities explicitly designed and planned by the robot's creator. Discrete actions such as picking up and placing known objects and navigating space can only be accomplished through extensive planning and testing. Beyond these tasks, the robot has little planning ability and struggles to understand its occupied environment. A robot working in a kitchen tasked with making dinner would have to know the presence of a refrigerator, know what food is inside, know how to retrieve food, and which surfaces to place ingredients on. By using the techniques described above in conjunction with an LLM agent, the robot can quickly assess what in the environment is relevant to the task at hand and plan accordingly more effectively. As an example, when a robot in the kitchen is asked to "make me dinner," using this technique, the robot would be able to quickly locate relevant items such as appliances, food countertops, etc.

[0098] Another example of a use case for the above process involves search. Users will increasingly capture and interact with 3D data on devices, accumulating thousands of 3D scans, which will require robust search capabilities. Users have become accustomed to natural language search of 2D images because such functionality is built into devices and is prominently featured in search engines. The search capabilities described in this paper can also be built into many applications or could drive new services storing 3D data. For example, companies providing robot fleet management software collect vast amounts of 3D data from robots in operation. This data may be underutilized due to the difficulty in finding a useful subset. The techniques described in this paper allow for the sifting and automatic tagging of the data.

[0099] Users may accumulate vast amounts of 3D data, which can be difficult to navigate and search. Therefore, the technology in this application is a natural addition to existing applications such as image sets. In the long term, it is expected that 3D data will become commonplace in many industries, and end users will become accustomed to using it. E-commerce will have deeper integrations that allow users to preview goods online, and social media will regularly feature 3D content. The types of search and segmentation discussed above will be important in driving these future systems. In fact, products such as map applications have already familiarized a large number of people with the basic technology, but they do not offer the kind of search functionality proposed in this paper. For example, when virtually visiting a home, a virtual camera must "walk" through the space, rather than allowing the user to specify what to display using natural language. By combining the subject matter of this disclosure with digital architectural datasets, many industries such as architecture, real estate, interior design, construction, and property management can significantly improve efficiency and scalability. That is, leveraging existing applications, users navigate 3D space using a click-based interface, and the authors of 3D scans must manually annotate the space to allow users to quickly jump to points of interest. The process disclosed herein will enable users to navigate using natural language queries with terms such as “table,” “poster,” or “mirror,” while significantly reducing the time and effort required to keep scans up-to-date.

[0100] although Figure 24 An example filtering process 2400 is shown, but it is possible to... Figure 24 Make various changes. For example, Figure 24 The various components and functions within can be combined, further subdivided, replicated, or rearranged as needed. Furthermore, one or more additional components and functions may be included if required or desired. Additionally, although shown as a series of steps, Figure 24 The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0101] Figure 25 An example method 2500 for natural language 3D data search according to this disclosure is shown. For ease of explanation, Figure 25 The method 2500 shown is described as being in Figure 1 Implemented on or by electronic device 101 in network configuration 100 Figure 1 The network configuration 100 is supported by electronic device 101. However, such as when process 2500 is implemented on server 106 or supported by server 106, Figure 25 The method 2500 shown can be used with any other suitable device and in any other suitable system.

[0102] For ease of explanation, refer to at least... Figure 15Method 2500 is described by the process flow depicted and described above. However, Method 2500 can be used with any suitable process flow and system, and can be easily modified to accommodate changes in the underlying process flow and / or system.

[0103] Method 2500 includes, at step 2502, merging information from the multimodal embeddings in the indexed point cloud data structure with 3D spatial information specific to the captured scene. In step 2504, at least one of querying or retrieving 3D point cloud data is performed based on user input, which includes at least one of natural language or image references. In step 2506, instance segmentation is used in conjunction with the multimodal embeddings to achieve both global and local scene understanding.

[0104] In any of the foregoing embodiments, a multimodal embedding with 3D spatial information can be generated by preprocessing 3D data for the captured scene to produce a 3D point cloud. 2D samples can be extracted from the 3D point cloud via projection, forming a representative summary of the 3D point cloud. 3D samples can be extracted from the 3D point cloud. The 2D and 3D samples can be used to generate a spatial data structure for the captured scene.

[0105] In various embodiments, the operations for generating spatial data structures may include: segmenting 2D samples into 2D fragments corresponding to objects in the captured scene; each of the 2D fragments may be image-encoded based on 3D points corresponding to the respective 2D fragment using a pre-trained encoder to form part of a multimodal embedding; information associated with the 2D fragments may be used to mask 3D samples to generate 3D fragments; and group tracking information and group identifiers may be determined for the 3D fragments.

[0106] In various embodiments, the operation of using instance segmentation in conjunction with multimodal embedding to achieve global scene understanding and local scene understanding may include storing the 3D points corresponding to any one of the multimodal embeddings and 2D segments, group tracking information, and group identifiers in a spatial data structure.

[0107] In various embodiments, the operation of generating the spatial data structure may include receiving camera pose information associated with 2D and 3D samples. The camera pose information can be used to determine group tracking information and group identifiers (IDs) for 3D segments by tracking sampled portions of the 3D point cloud and the corresponding sampling resolution for each sampled portion.

[0108] In various embodiments, operations employing camera pose information may include: assigning embedding vectors for the entire fragment set to each 3D point contained throughout the fragment set; examining each frame to determine the status of a fragment set with respect to 3D points seen in previously processed, different fragment sets; assigning new group IDs to 3D points not seen in any previously processed fragment sets; assigning previously assigned group IDs to 3D points seen in at least one previously processed fragment set; and averaging the embeddings for overlapping images.

[0109] In various embodiments, the operation of using instance segmentation in conjunction with multimodal embedding to achieve global scene understanding and local scene understanding may include: using a 2D semantic segmentation model to identify and extract a subset of a 3D point cloud, wherein the 2D semantic segmentation model generates one of a discrete set of class labels or a mask.

[0110] In various embodiments, performing an operation to query or retrieve at least one of 3D point cloud data based on user input, including at least one of natural language or image references, may include: performing a scan of multimodal embeddings in a spatial database based on the user input; discarding entries in the spatial database whose similarity scores are below a defined threshold; and selecting a final group of point clouds from the set of point clouds corresponding to the remaining 2D and 3D fragments after discarding the entries, based on one or more metrics selected from aggregation confidence, size, or proximity.

[0111] although Figure 25 An example of method 2500 for natural language 3D data search is shown, but it is applicable to... Figure 25 Make various changes. For example, although shown as a series of steps, Figure 25 The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0112] Although this disclosure has been described with reference to various exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. This disclosure is intended to include such changes and modifications that fall within the scope of the appended claims.

Claims

1. A method comprising: The information from the multimodal embedding in the indexed point cloud data structure is merged with the three-dimensional 3D spatial information for the captured scene; Querying or retrieving 3D point cloud data based on user input that includes at least one of natural language or image references; as well as The multimodal embedding is combined with instance segmentation to achieve global scene understanding and local scene understanding.

2. The method of claim 1, wherein, The multimodal embedding with three-dimensional spatial information is generated through the following operations: Preprocess the 3D data for the captured scene to generate a 3D point cloud; Two-dimensional 2D samples are extracted from the 3D point cloud by projection, and the 2D samples form a representative summary of the 3D point cloud; Extract 3D samples from the 3D point cloud; as well as Using the 2D and 3D samples, a spatial data structure for the captured scene is generated.

3. The method of claim 2, wherein, The operations for generating the spatial data structure include: The 2D sample is segmented into 2D fragments corresponding to objects in the captured scene; Based on each corresponding 3D point in the 2D segment, the corresponding 2D segment is image encoded using a pre-trained encoder to form part of the multimodal embedding. Use information associated with the 2D fragment to mask the 3D sample to generate a 3D fragment; and Determine the group tracking information and group identifier for the 3D segment.

4. The method of claim 3, wherein, The operations for achieving global scene understanding and local scene understanding using instance segmentation in conjunction with the aforementioned multimodal embedding include: The multimodal embedding, the 3D point corresponding to any one of the 2D segments, the group tracking information, and the group identifier are stored in the spatial data structure.

5. The method of claim 3, wherein, The operations for generating the spatial data structure include: Receive camera pose information associated with the 2D sample and the 3D sample; and The camera pose information is used to determine the group tracking information and the group identifier ID for the 3D segment by tracking the sampled portions of the 3D point cloud and the corresponding sampling resolution for each sampled portion.

6. The method of claim 5, wherein, Operations using the camera pose information include: The embedding vector for the entire fragment set is assigned to each 3D point contained in the entire fragment set; Each frame is examined to determine the status of a segment set within the segment set relative to the 3D points seen in a different segment set previously processed within that segment set; Assign the new group ID to a 3D point that has not been seen in any previously processed fragment set; Assign the previously assigned group ID to the 3D point seen in at least one previously processed fragment set; and Average the embeddings for overlapping images.

7. The method of claim 3, wherein, The operations for achieving global scene understanding and local scene understanding using instance segmentation in conjunction with the aforementioned multimodal embedding include: A 2D semantic segmentation model is used to identify and extract a subset of the 3D point cloud, wherein the 2D semantic segmentation model generates one of a discrete set of class labels or a mask.

8. The method of claim 3, wherein, Performing an operation to query 3D point cloud data or retrieve at least one of 3D point cloud data based on user input including at least one of the natural language or the image reference includes: A scan of the multimodal embeddings in the spatial database is performed based on the user input; Discard entries in the spatial database whose similarity scores are below a defined threshold; and From the point cloud set corresponding to the remaining 2D and 3D fragments after the discarded entries, a final group of point clouds is selected based on one or more metrics chosen from aggregation confidence, size, or proximity.

9. An apparatus comprising: At least one processing device is configured to: The information from the multimodal embedding in the indexed point cloud data structure is merged with the three-dimensional 3D spatial information for the captured scene; Querying or retrieving 3D point cloud data based on user input that includes at least one of natural language or image references; as well as The multimodal embedding is combined with instance segmentation to achieve global scene understanding and local scene understanding.

10. The device as claimed in claim 9, wherein, At least one processing device is configured to generate the multimodal embedding having three-dimensional spatial information by: Preprocess the 3D data for the captured scene to generate a 3D point cloud; Two-dimensional 2D samples are extracted from the 3D point cloud by projection, and the 2D samples form a representative summary of the 3D point cloud; Extract 3D samples from the 3D point cloud; as well as Using the 2D and 3D samples, a spatial data structure for the captured scene is generated.

11. The device as claimed in claim 10, wherein, At least one processing device is configured to generate the spatial data structure by: The 2D sample is segmented into 2D fragments corresponding to objects in the captured scene; Based on each corresponding 3D point in the 2D segment, the corresponding 2D segment is image encoded using a pre-trained encoder to form part of the multimodal embedding. Use information associated with the 2D fragment to mask the 3D sample to generate a 3D fragment; as well as Determine the group tracking information and group identifier for the 3D segment.

12. The device as claimed in claim 11, wherein, At least one processing device is configured to achieve global scene understanding and local scene understanding by combining the multimodal embedding with instance segmentation: The multimodal embedding, the 3D point corresponding to any one of the 2D segments, the group tracking information, and the group identifier are stored in the spatial data structure.

13. The device as claimed in claim 11, wherein, At least one processing device is configured to generate the spatial data structure by: Receive camera pose information related to the 2D sample and the 3D sample; as well as The camera pose information is used to determine the group tracking information and the group identifier ID for the 3D segment by tracking the sampled portions of the 3D point cloud and the corresponding sampling resolution for each sampled portion.

14. The device as claimed in claim 13, wherein, At least one processing device is configured to employ the camera pose information by: The embedding vector for the entire fragment set is assigned to each 3D point contained in the entire fragment set; Each frame is examined to determine the status of a segment set within the segment set relative to the 3D points seen in a different segment set previously processed within that segment set; Assign the new group ID to a 3D point that has not been seen in any previously processed fragment set; Assign the previously assigned group ID to the 3D point seen in at least one previously processed fragment set; and Average the embeddings for overlapping images.

15. The device as claimed in claim 11, wherein, At least one processing device is configured to achieve global scene understanding and local scene understanding by combining the multimodal embedding with instance segmentation: A 2D semantic segmentation model is used to identify and extract a subset of the 3D point cloud, wherein the 2D semantic segmentation model generates one of a discrete set of class labels or a mask.