Scene representation method and device for enhancing three-dimensional space understanding of large language model

By constructing three-dimensional spatial representation layer by layer, the lack of spatial relationships and position perception in the understanding of three-dimensional scenes is solved, and a richer and more accurate three-dimensional spatial understanding is achieved.

CN119942027AActive Publication Date: 2025-05-06北京数原数字化城市研究中心

Patent Information

Application Number
CN202510017129.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing large language models lack effective perception of three-dimensional spatial relationships and precise position generation in the understanding of three-dimensional scenes, which limits the richness of three-dimensional scene representation and hinders the comprehensive perception of three-dimensional space by the large language models.

Method used

Multiple subset points are obtained by sampling and offsetting the scene point cloud characterization, multiple visual references are clustered, and global spatial distribution modeling is carried out through the message propagation mechanism, combining multi-layer attention mechanism and position fine network to construct three-dimensional spatial representation layer by layer.

Benefits of technology

Effectively capture and enhance the spatial representation of location information, and improve the spatial understanding and reasoning capabilities of large language models in three-dimensional visual language tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942027A_ABST
    Figure CN119942027A_ABST
Patent Text Reader

Abstract

The invention provides a scene representation method and device for enhancing three-dimensional space understanding of a large language model, and relates to the technical field of data processing. Constructing a plurality of visual reference objects, learning spatial information of points in a local area corresponding to the visual reference objects, and obtaining a three-dimensional space representation of a first level; global spatial distribution modeling among different visual reference objects is promoted through a message passing mechanism, so that each visual reference object not only captures local features of the visual reference object, but also can understand a global spatial relationship between the visual reference object and an adjacent reference object, and a second-level three-dimensional spatial representation is formed. Information interaction between a visual reference object and a global scene is realized through an attention mechanism, and a position fine tuning network is added to refine positioning of the visual reference object, so that three-dimensional space representation of a third layer is obtained. Therefore, the progressive three-dimensional space representation from the first hierarchy to the third hierarchy is adopted, the space representation enhancing the position information is captured, and the space understanding and reasoning ability of a large language model in processing a three-dimensional visual language task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a scene representation method and device for enhancing three-dimensional spatial understanding of a large language model. Background Art

[0002] Three-dimensional scene understanding is an important way for robots to perceive the rich semantic and spatial information of the real world. It aims to solve the task of understanding specified objects or spaces in natural language queries by using specific scenes as context. At present, large language models have made some progress in three-dimensional scene understanding, but still lack effective perception of three-dimensional spatial relationships and precise position generation. Existing models often rely on overall scene information or specific areas when understanding 3D scenes, which limits the richness of three-dimensional scene representation and hinders the comprehensive perception of three-dimensional space by large language models. Summary of the invention

[0003] In view of this, the present application provides a scene representation method and device for enhancing the three-dimensional spatial understanding of a large language model, aiming to enhance the scene representation of the three-dimensional spatial understanding of a large language model.

[0004] In a first aspect, the present application provides a scene representation method for enhancing three-dimensional spatial understanding of a large language model, including:

[0005] The scene point cloud representation is obtained by sampling and position shifting a plurality of subset points, and the scene point cloud representation is clustered based on the plurality of subset points to construct a plurality of visual reference objects, wherein the plurality of visual reference objects serve as the first-level three-dimensional representation of the corresponding local area in the scene point cloud representation;

[0006] Performing global spatial distribution modeling on the multiple visual reference objects through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects;

[0007] The second-level three-dimensional representations of the multiple visual reference objects and the scene point cloud representation are processed through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation. The position fine-tuning network is a neural network that adjusts the spatial positioning of each of the visual reference objects based on real coordinates.

[0008] Optionally, obtaining a plurality of subset points by sampling and position shifting the scene point cloud representation, and clustering the scene point cloud representation based on the plurality of subset points to construct a plurality of visual reference objects includes:

[0009] Sampling the scene point cloud representation by a farthest point sampling algorithm to obtain a first number of sampling points that are uniformly distributed in space;

[0010] Aligning each sampling point with the nearest object center through a first feedforward neural network to obtain a subset of points corresponding to the sampling point;

[0011] The scene point cloud representation is clustered and pooled based on the first number of subset points to form a first number of visual reference objects.

[0012] Optionally, the global spatial distribution modeling of the multiple visual reference objects through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects includes:

[0013] The first number of visual reference objects are input into a graph convolutional network, feature information of the visual reference objects is used to model graph nodes, and edge information is constructed using spatial distances between each of the visual reference objects to obtain a second-level three-dimensional representation of each of the visual reference objects.

[0014] Optionally, the processing of the second-level three-dimensional representations of the multiple visual references and the scene point cloud representation through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation includes:

[0015] Inputting the second-level three-dimensional representation of the visual reference object into a self-attention network, inputting the processing result output by the self-attention network and the scene point cloud representation into a cross-attention network, and updating the third-level three-dimensional representation of the visual reference object;

[0016] The third-level three-dimensional representation is input into a position fine-tuning network to update the third-level three-dimensional representation.

[0017] Optionally, the method further includes:

[0018] By means of a visual-linguistic bridge, the updated third-level three-dimensional representation is aligned to the language text space to obtain a visual prompt;

[0019] The visual prompt is input into a large language model so that the large language model can understand and process the visual prompt.

[0020] Optionally, the method further includes:

[0021] Collect the point cloud of the target scene to obtain the scene point cloud;

[0022] The scene point cloud is encoded by a point cloud encoder to obtain a scene point cloud representation.

[0023] Optionally, the method further includes:

[0024] Configuring a center loss function and an inter-pair spatial constraint loss function for the position fine-tuning network to implement feedback adjustment of a neural network that generates a third-level three-dimensional representation from the second-level three-dimensional representation, a neural network that generates a second-level three-dimensional representation from multiple visual references, and a neural network that generates multiple visual references from a scene point cloud representation;

[0025] The position fine-tuning network configures a center loss function and an inter-pair spatial constraint loss function;

[0026] The center loss function is:

[0027]

[0028] Where M represents the number of visual references, represents the coordinates of the predicted visual reference object, Represents the coordinates of the actual visual reference object;

[0029] The pairwise spatial constraint loss function:

[0030]

[0031] Where N represents the number of visual reference pairs, is the distance between the predicted visual reference pair (i, j), is the actual distance between the visual reference pair (i, j).

[0032] Optionally, the method further includes:

[0033] An overall optimization loss function is set based on the center loss function, the pairwise space constraint loss function, and the generation loss function of the large language model, and feedback optimization is performed on the large language model, the visual-language bridge, the neural network corresponding to the third-level three-dimensional representation generated from the second-level three-dimensional representation, the neural network corresponding to the second-level three-dimensional representation generated from multiple visual references, and the neural network corresponding to the multiple visual references generated from the scene point cloud representation. The formula of the overall optimization loss function is:

[0034] L total =L LLM +α1L center +α2L psc

[0035] Among them, L LLM represents the generation loss function of the large language model, α1 represents the coefficient of the center loss function, and α2 represents the coefficient of the inter-space constraint loss function.

[0036] In a second aspect, the present application provides a scene representation device for enhancing three-dimensional spatial understanding of a large language model, the device comprising:

[0037] A clustering abstraction module, configured to obtain a plurality of subset points by sampling and position shifting the scene point cloud representation, cluster the scene point cloud representation based on the plurality of subset points to construct a plurality of visual reference objects, and the plurality of visual reference objects serve as the first-level three-dimensional representation of the corresponding local area in the scene point cloud representation;

[0038] A message transmission module, used for performing global spatial distribution modeling of the multiple visual reference objects through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects;

[0039] The interaction module is used to process the second-level three-dimensional representations of the multiple visual reference objects and the scene point cloud representation through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation. The position fine-tuning network is a neural network that adjusts the spatial positioning of each of the visual reference objects based on real coordinates.

[0040] Optionally, the device further includes:

[0041] A visual-language bridge, for aligning the updated third-level three-dimensional representation to the language text space to obtain a visual prompt;

[0042] A large language model is used to understand and process the visual cues.

[0043] In a third aspect, the present application provides a device comprising a memory and a processor, wherein the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes a scene representation method for enhancing three-dimensional spatial understanding of a large language model as described in any one of the first aspects.

[0044] In a fourth aspect, the present application provides a computer storage medium having codes stored therein. When the codes are executed, the device executing the codes implements a scene characterization method for enhancing three-dimensional spatial understanding of a large language model as described in any one of the first aspects above.

[0045] The present application provides a scene representation method and device for enhancing the three-dimensional spatial understanding of a large language model. When executing the method, the scene point cloud representation is first obtained by sampling and position offsetting multiple subset points, and the scene point cloud representation is clustered based on the multiple subset points to construct multiple visual references. After that, the multiple visual references are modeled for global spatial distribution through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual references, and then the second-level three-dimensional representation of the multiple visual references and the scene point cloud representation are processed through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation. In this way, multiple visual references are first randomly constructed, and the first-level three-dimensional spatial representation is obtained by learning the spatial information of the points in the local area corresponding to the visual reference. Then, through the message passing mechanism, the global spatial distribution modeling between different visual references is promoted, so that each visual reference can not only capture its local features, but also understand the global spatial relationship with the adjacent references, forming a second-level three-dimensional spatial representation. Finally, the information interaction between the visual reference and the global scene is realized through the attention mechanism, and the position fine-tuning network is added to refine the positioning of the visual reference to obtain the third-level three-dimensional spatial representation. In this way, a progressive position-enhanced three-dimensional spatial representation is adopted from the first level to the third level to capture the spatial representation with enhanced position information, better realize spatial position perception, and help improve the spatial understanding and reasoning ability of the large language model in processing three-dimensional visual language tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 A flow chart of a scene representation method for enhancing three-dimensional spatial understanding of a large language model provided in an embodiment of the present application;

[0048] Figure 2 A schematic diagram of the structure of a scene representation device for enhancing three-dimensional spatial understanding of a large language model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] At present, the 3D vision-language model technology that combines three-dimensional (3D) scenes with large-scale language models (LLM) has explored a variety of spatial learning paradigms for perceiving and understanding the 3D world. The architecture of these technologies usually includes 3D vision encoders, vision-language bridges, and language large models. Among them, some methods are mainly limited to understanding object attributes and lack accurate 3D spatial position perception. Other methods first use scene encoders to extract scene point cloud features; then use vision-language bridges to extract instruction-specific visual information and align it with the text space; finally, input LLM to complete the 3D scene understanding task. However, the implementation of this type of technology is often limited by the ability of vision-language bridges to fit tasks, and cannot handle complex tasks in the scene, such as spatial layout understanding, distance estimation between objects, and spatial layout editing, and lacks effective perception of three-dimensional spatial relationships and precise position generation.

[0050] Based on the above problems, the present application provides a scene representation method and device for enhancing the three-dimensional spatial understanding of a large language model. The method first randomly constructs multiple visual references, and obtains the first-level three-dimensional spatial representation by learning the spatial information of points in the local area corresponding to the visual reference. Then, through the message passing mechanism, the global spatial distribution modeling between different visual references is promoted, so that each visual reference can not only capture its local features, but also understand the global spatial relationship with adjacent references, forming a second-level three-dimensional spatial representation. Finally, the information interaction between the visual reference and the global scene is realized through the attention mechanism, and the position fine-tuning network is added to refine the positioning of the visual reference to obtain the third-level three-dimensional spatial representation. In this way, the three-dimensional spatial representation is enhanced by progressive position enhancement from the first level to the third level, which effectively captures the spatial representation of the enhanced position information, better realizes spatial position perception, and is conducive to improving the spatial understanding and reasoning ability of the large language model in processing three-dimensional visual language tasks.

[0051] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0052] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0053] Unless otherwise stated, the term "plurality" means two or more.

[0054] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0055] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0056] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0057] See also Figure 1 , Figure 1 A flow chart of a scene representation method for enhancing three-dimensional spatial understanding of a large language model provided in an embodiment of the present application, a scene representation method for enhancing three-dimensional spatial understanding of a large language model, comprising:

[0058] S101. A scene point cloud representation is obtained by sampling and position shifting a plurality of subset points, and the scene point cloud representation is clustered based on the plurality of subset points to construct a plurality of visual reference objects, wherein the plurality of visual reference objects serve as a first-level three-dimensional representation of a corresponding local area in the scene point cloud representation.

[0059] The scene point cloud representation is obtained by encoding the point cloud of the target scene through a point cloud encoder.

[0060] A point cloud is a collection of discrete points in three-dimensional space. Each point carries three-dimensional coordinate information. In addition, this information can also include other attributes such as color and reflectivity.

[0061] Scene point cloud representation is to extract meaningful features of point clouds and convert them into a form that can be understood and processed by computers for subsequent tasks such as scene understanding, target recognition, and 3D reconstruction. Scene point cloud representation can include the position of points in the point cloud and the distance between points, shape and contour (which can be represented by normal vectors, curvature, etc.), spatial distribution relationship (describing the relative position relationship between objects in the scene), attribute information (color, reflectivity), etc.

[0062] The above method obtains multiple sub-points uniformly distributed in space through sampling, and aligns each sub-point with the center of the nearest object through position offset to obtain a subset of points representing the object or a part of the object. Then, the scene point cloud representation is clustered based on the multiple subset points to construct multiple visual references. Through clustering, the scene point cloud representation is divided into regions corresponding to the multiple subset points to obtain multiple sub-regions, and then the point sets of each sub-region are compressed through pooling to form visual references.

[0063] In this way, the spatial information of the internal points of the local area can be obtained, thereby obtaining the first-level three-dimensional spatial representation.

[0064] S102: Perform global spatial distribution modeling on the multiple visual reference objects through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects.

[0065] Optionally, global spatial distribution modeling can model graph nodes with feature information of multiple visual reference objects, construct edge information through spatial position distances between two visual reference objects, and promote global spatial distribution modeling between different visual reference objects through a message passing mechanism.

[0066] In this way, we can focus on the implicit relationship between visual references, capture the local features of visual references, and capture the global spatial relationship semantics of adjacent visual references. Through the message propagation mechanism, we learn the enhanced representation of the spatial relationship of each visual reference and obtain the second-level three-dimensional spatial representation.

[0067] S103. Process the second-level three-dimensional representations of the multiple visual reference objects and the scene point cloud representation through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation. The position fine-tuning network is a neural network that adjusts the spatial positioning of each of the visual reference objects based on real coordinates.

[0068] Optionally, self-attention can be used to process the second-level three-dimensional representation of the visual reference object, and then combined with the scene point cloud representation to achieve information interaction between the visual reference object and the global scene through cross-attention, to obtain a third-level three-dimensional representation, and the third-level three-dimensional representation is updated through a position fine-tuning network.

[0069] Based on the above steps S101-S103, it can be known that first, multiple visual references are randomly constructed, and the first-level three-dimensional spatial representation is obtained by learning the spatial information of the points in the local area corresponding to the visual reference. Then, through the message passing mechanism, the global spatial distribution modeling between different visual references is promoted, so that each visual reference can not only capture its local features, but also understand the global spatial relationship with the adjacent references, forming a second-level three-dimensional spatial representation. Finally, the information interaction between the visual reference and the global scene is realized through the attention mechanism, and the position fine-tuning network is added to refine the positioning of the visual reference to obtain the third-level three-dimensional spatial representation. In this way, the progressive position-enhanced three-dimensional spatial representation from the first level to the third level is adopted to effectively capture the spatial representation of the enhanced position information, better realize spatial position perception, and help improve the spatial understanding and reasoning ability of the large language model in processing three-dimensional visual language tasks.

[0070] Optionally, before the above step S101, a scene point cloud representation may be obtained, and the specific steps may be:

[0071] First, the point cloud of the target scene is collected to obtain the scene point cloud.

[0072] For example, a large number of points on the surface of an object in three-dimensional space can be acquired by using a scanner, camera or other device to form a scene point cloud.

[0073] Then, the scene point cloud is encoded through a point cloud encoder to obtain a scene point cloud representation.

[0074] The point cloud encoder mainly compresses, extracts features or converts the original scene point cloud data to form a scene point cloud representation.

[0075] Optionally, after the above step S103, the third-level three-dimensional representation that enhances the three-dimensional spatial understanding of the large language model can also be used to solve the understanding task of the specified object or space in the natural language query, such as ScanQA, ScanRefer and Scan2Cap. It can include:

[0076] First, through the visual-language bridge, the third-level three-dimensional representation output in step S103 is aligned to the language text space to obtain a visual prompt.

[0077] The vision-language bridge aligns the third-level three-dimensional representation of the enhanced scene representation to the language text space as a visual cue, so that the subsequent large language model can process three-dimensional spatial information and natural language input simultaneously, enhancing multimodal understanding capabilities and facilitating accurate scene analysis.

[0078] Then, the visual prompt is input into the large language model so that the large language model can understand and process the visual prompt to obtain the result of the understanding task about the specified object or space in the natural language query.

[0079] In the embodiment of the present application, the above Figure 1 There are possible implementations of step S101, which are described in detail below. It should be noted that the implementations described below are only exemplary and do not represent all implementations of the embodiments of the present application.

[0080] In step S101, a scene point cloud representation is obtained by sampling and position shifting a plurality of subset points, and the scene point cloud representation is clustered based on the plurality of subset points to construct a plurality of visual reference objects. The specific steps may include:

[0081] First, the scene point cloud representation is sampled using a farthest point sampling algorithm to obtain a first number of sampling points that are uniformly distributed in space.

[0082] The farthest point sampling algorithm can sample the scene point cloud representation to obtain a first number of sampling points that are uniformly distributed in space. In this way, representative sampling points are selected from a given point cloud set (scene point cloud representation), and these sampling points can reflect the overall distribution characteristics of the point cloud to a certain extent.

[0083] Then, each sampling point is aligned with the center of the nearest object through a first feedforward neural network to obtain a subset of points corresponding to the sampling point.

[0084] The offset between the current sampling point and the center of the nearest object is obtained through the first feedforward neural network, so that each sampling point is aligned with the center of the nearest object.

[0085] Finally, clustering and pooling the scene point cloud representation based on the first number of subset points to form a first number of visual reference objects.

[0086] Optionally, the above clustering may adopt a K-nearest neighbor algorithm, which is an algorithm for classification and regression tasks, and a plurality of points closest to each subset point in the scene point cloud representation are selected by KNN to form a sub-region point set.

[0087] Then, each sub-region point set is pooled, compressed, and the information of multiple points included in the sub-region point set is aggregated to obtain the point set features with rich information of the local region corresponding to the sub-region point set, thereby forming the first-level three-dimensional representation of the local region corresponding to the scene point cloud representation, that is, forming a visual reference for the sub-region point set.

[0088] In the embodiment of the present application, the above Figure 1There are possible implementations of step S102, which are described in detail below. It should be noted that the implementations described below are only for illustrative purposes and do not represent all implementations of the embodiments of the present application.

[0089] In step S102, the multiple visual reference objects are globally distributed in space by using a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects. The specific steps include:

[0090] The first number of visual reference objects are input into a graph convolutional network, feature information of the visual reference objects is used to model graph nodes, and edge information is constructed using spatial distances between each of the visual reference objects to obtain a second-level three-dimensional representation of each of the visual reference objects.

[0091] The graph convolutional network mentioned above is a deep learning model specifically designed to process graph structured data. A graph is a data structure in which nodes represent entities and edges represent the relationships between entities. A set of nodes and their relationships (edges) can be modeled.

[0092] By effectively extracting features from graph structured data through graph convolutional networks, the present application models graph nodes through feature information of multiple visual reference objects, constructs edge information based on the spatial distance between each pair of visual reference objects, and implements global spatial distribution modeling. In this way, the local features of the visual reference objects are captured, and the global spatial relationship features of adjacent visual reference objects are captured, and the enhanced representation of the spatial relationship of each visual reference object is learned to obtain a second-level three-dimensional representation.

[0093] In the embodiment of the present application, the above Figure 1 There are possible implementations of step S103, which are described in detail below. It should be noted that the implementations described below are only exemplary and do not represent all implementations of the embodiments of the present application.

[0094] In step S103, the second-level three-dimensional representations of the multiple visual reference objects and the scene point cloud representation are processed through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation. The specific steps include:

[0095] First, the second-level three-dimensional representation of the visual reference is input into the self-attention network, and the processing result output by the self-attention network and the scene point cloud representation are input into the cross-attention network to update the third-level three-dimensional representation of the visual reference.

[0096] The self-attention network can better capture the long-distance dependencies and important information in the second-level three-dimensional representation of each visual reference object. The cross-attention network can fuse the processing results output by the above self-attention network and the correlation information between the scene point cloud representation to achieve information interaction between each visual reference object and the global scene. In this way, the inherent characteristics of the visual reference object and the information of the global scene are effectively combined, so that the visual reference object has the attribute of being fully perceived in space.

[0097] The third-level three-dimensional representation is input into a position fine-tuning network to update the third-level three-dimensional representation.

[0098] Exemplarily, the position fine-tuning network may adopt a FFN including multiple layers.

[0099] The spatial positioning of the visual reference is adjusted by minimizing the relative distance between the visual reference and the real coordinates, so that the visual reference is closer to the coordinate center of the object.

[0100] In a specific implementation, to implement the above steps S101-S103, a spatial perception model may be set and trained corresponding to the neural network corresponding to the above steps S101-S103, specifically:

[0101] The spatial perception model may include a clustering abstraction module, a message passing module, and an interaction module;

[0102] The clustering abstraction module is used to execute the above step S101, and the clustering abstraction module may specifically include a farthest point sampling algorithm, a feedforward neural network layer, a clustering layer (k-nearest neighbor algorithm clustering may be used), and a pooling layer. The message passing module is used to execute the above step S102, and the message passing module may specifically include a graph convolutional network. The interaction module is used to execute the above step S103, and the message passing module is specifically configured with a self-attention layer, a cross-attention layer, and a position fine-tuning network layer.

[0103] When specifically training the spatial perception model, a multi-layer feedforward neural network can be set in the position fine-tuning network layer, and a center loss function and an inter-pair spatial constraint loss function can be designed. The above-mentioned spatial perception model is trained based on a training set containing scene point cloud representation (which can be a training set containing scene point cloud representation obtained by encoding the training set containing scene point cloud through a point cloud encoder). The Euclidean distance between the predicted coordinates and the real coordinates of the visual reference object obtained by the spatial perception model is penalized by the loss function, so as to encourage the spatial perception model to predict a more accurate position of the visual reference object.

[0104] The center loss function is:

[0105]

[0106] Where M represents the number of visual references, represents the coordinates of the predicted visual reference object, Represents the coordinates of the actual visual reference object;

[0107] The pairwise spatial constraint loss function:

[0108]

[0109] Where N represents the number of visual reference pairs, is the distance between the predicted visual reference pair (i, j), is the actual distance between the visual reference pair (i, j).

[0110] Furthermore, in order to make the third 3D spatial feature better for the large language model, the above spatial perception model can be further trained in combination with the visual-language bridge and the large language model. The specific steps are as follows:

[0111] First, a cloud representation of a scenic spot in the training set is input into the above-mentioned spatial perception model to obtain the third-level three-dimensional representation output by the spatial perception model, and the spatial perception model is trained based on the feedback of the center loss function and the pairwise spatial constraint loss function;

[0112] The training set can be composed of various visual-language understanding and visual-language localization tasks, wherein the training set also includes point cloud scene representations or point cloud scenes corresponding to various visual-language understanding and visual-language localization tasks. If it is a point cloud scene, it needs to be encoded by a point cloud encoder first to obtain the corresponding point cloud scene representation.

[0113] Secondly, the third-level three-dimensional representation is input into the visual-language bridge, and the third-level three-dimensional representation is aligned to the language text space to obtain the visual prompt output by the visual-language bridge.

[0114] Then, the visual cue is input into the large language model, so that the large language model understands and processes the visual cue to obtain the results of various visual-language understanding and visual-language localization tasks corresponding to the point cloud scene representation. The above-mentioned spatial perception model, visual-language bridge and large language model are optimized and adjusted by overall optimization loss function.

[0115] The overall optimization loss function is: L total =L LLM +α1L center +α2L psc

[0116] Among them, L LLMrepresents the generation loss function of the large language model, α1 represents the coefficient of the center loss function, and α2 represents the coefficient of the inter-space constraint loss function.

[0117] The above is some specific implementation methods of a scene representation method for enhancing the three-dimensional spatial understanding of a large language model provided by the embodiment of the present application. Based on this, the present application also provides a corresponding device. The device provided by the embodiment of the present application will be introduced from the perspective of functional modularization.

[0118] See also Figure 2 A schematic diagram of the structure of a scene representation device for enhancing three-dimensional spatial understanding of a large language model is shown, and a scene representation device for enhancing three-dimensional spatial understanding of a large language model includes:

[0119] A clustering abstraction module 201 is used to obtain a plurality of subset points by sampling and position shifting the scene point cloud representation, and to cluster the scene point cloud representation based on the plurality of subset points to construct a plurality of visual reference objects, wherein the plurality of visual reference objects serve as a first-level three-dimensional representation of a corresponding local area in the scene point cloud representation;

[0120] A message transmission module 202, configured to perform global spatial distribution modeling on the multiple visual reference objects through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects;

[0121] The interaction module 203 is used to process the second-level three-dimensional representations of the multiple visual reference objects and the scene point cloud representation through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation. The position fine-tuning network is a neural network that adjusts the spatial positioning of each of the visual reference objects based on real coordinates.

[0122] This application uses the above-mentioned device to randomly construct multiple visual references through the clustering abstraction module 201, and obtains the first-level three-dimensional spatial representation by learning the spatial information of the points in the local area corresponding to the visual reference. Then, the message passing module 202 promotes the modeling of the global spatial distribution between different visual references, so that each visual reference can not only capture its local features, but also understand the global spatial relationship with the adjacent references, forming a second-level three-dimensional spatial representation. Finally, the interaction module 203 realizes the information interaction between the visual reference and the global scene through a multi-layer attention mechanism, and adds a position fine-tuning network to refine the positioning of the visual reference to obtain the third-level three-dimensional spatial representation. In this way, the progressive position-enhanced three-dimensional spatial representation from the first level to the third level is adopted to effectively capture the spatial information containing the position information, better realize the spatial position perception, and is conducive to improving the spatial understanding and reasoning ability of the large language model in processing three-dimensional visual language tasks.

[0123] Optionally, the device further includes: a collection module, used to collect a point cloud of the target scene to obtain a scene point cloud.

[0124] Optionally, the device further includes: a point cloud encoder, used to encode the scene point cloud to obtain a scene point cloud representation.

[0125] Optionally, the device further includes:

[0126] A visual-language bridge, for aligning the updated third-level three-dimensional representation to the language text space to obtain a visual prompt;

[0127] A large language model is used to understand and process the visual cues.

[0128] In one possible implementation, the clustering abstraction module 201 is specifically used to sample the scene point cloud representation through a farthest point sampling algorithm to obtain a first number of sampling points that are evenly distributed in space; align each sampling point with the center of the nearest object through a first feedforward neural network to obtain a subset of points corresponding to the sampling point; and cluster and pool the scene point cloud representation based on the first number of subset points to form a first number of visual reference objects.

[0129] In one possible implementation, the message passing module 202 is specifically used to input the first number of visual reference objects into a graph convolutional network, use the feature information of the visual reference objects to model graph nodes, construct edge information based on the spatial distances between the visual reference objects, and obtain a second-level three-dimensional representation of each of the visual reference objects.

[0130] In one possible implementation, the interaction module 203 is specifically used to input the second-level three-dimensional representation of the visual reference object into the self-attention network, input the processing result output by the self-attention network and the scene point cloud representation into the cross-attention network, and update the third-level three-dimensional representation of the visual reference object; input the third-level three-dimensional representation into the position fine-tuning network, and update the third-level three-dimensional representation.

[0131] In a possible implementation, the clustering abstraction module, the message passing module, and the interaction module form a spatial perception model, and the point cloud scene representation is input into the spatial perception model to obtain a scene representation that can enhance the three-dimensional spatial understanding of the large language model. The training method of the spatial perception model can be:

[0132] The spatial perception model may include;

[0133] The training is performed based on a training set containing scene point cloud representations (which may be a training set containing scene point cloud representations obtained by encoding the training set containing scene point cloud with a point cloud encoder). The Euclidean distance between the predicted coordinates and the actual coordinates of the visual reference objects obtained by the spatial perception model is penalized by a loss function, so as to encourage the spatial perception model to predict more accurate positions of the visual reference objects. The loss function configures an interactive module that outputs the third-level three-dimensional representation. The loss function includes a center loss function and an inter-pair spatial constraint loss function.

[0134] The center loss function can be:

[0135]

[0136] Where M represents the number of visual references, represents the coordinates of the predicted visual reference object, Represents the coordinates of the actual visual reference object;

[0137] The inter-space constraint loss function can be:

[0138]

[0139] Where N represents the number of visual reference pairs, is the distance between the predicted visual reference pair (i, j), is the actual distance between the visual reference pair (i, j).

[0140] Furthermore, in order to make the third 3D spatial feature more suitable for the large language model, the above spatial perception model can be trained in combination with the visual-language bridge and the large language model. The specific steps are as follows:

[0141] First, a scenic spot cloud representation in the training set is input into the above-mentioned spatial perception model to obtain the third-level three-dimensional representation output by the spatial perception model, and the spatial perception model is trained based on the feedback of the center loss function and the pairwise spatial constraint loss function;

[0142] The training set can be composed of various visual-language understanding and visual-language localization tasks, wherein the training set also includes point cloud scene representations or point cloud scenes corresponding to various visual-language understanding and visual-language localization tasks. If it is a point cloud scene, it needs to be encoded by a point cloud encoder first to obtain the corresponding point cloud scene representation.

[0143] Secondly, the third-level three-dimensional representation is input into the visual-language bridge, and the third-level three-dimensional representation is aligned to the language text space to obtain the visual prompt output by the visual-language bridge.

[0144] Then, the visual cue is input into the large language model, so that the large language model understands and processes the visual cue to obtain the results of various visual-language understanding and visual-language localization tasks corresponding to the point cloud scene representation. The above-mentioned spatial perception model, visual-language bridge and large language model are optimized and adjusted by overall optimization loss function.

[0145] The overall optimization loss function is: L total =L LLM +α1L center +α2L psc

[0146] Among them, L LLM represents the generation loss function of the large language model, α1 represents the coefficient of the center loss function, and α2 represents the coefficient of the inter-space constraint loss function.

[0147] The embodiments of the present application also provide corresponding devices and computer storage media for implementing the solutions provided by the embodiments of the present application.

[0148] The device includes a memory and a processor, the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes a scene representation method for enhancing three-dimensional spatial understanding of a large language model as described in any embodiment of the present application.

[0149] The computer storage medium stores code, and when the code is executed, the device executing the code implements a scene representation method for enhancing three-dimensional spatial understanding of a large language model as described in any embodiment of the present application.

[0150] The "first" and "second" in the names such as "first" and "second" (if any) mentioned in the embodiments of the present application are only used as name identifiers and do not represent the first or second in order.

[0151] Through the description of the above implementation methods, it can be known that those skilled in the art can clearly understand that all or part of the steps in the above-mentioned embodiment method can be implemented by means of software plus a general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in each embodiment of the present application or some parts of the embodiments.

[0152] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.

[0153] The above description is merely an exemplary embodiment of the present application and is not intended to limit the protection scope of the present application.

Claims

1. A scene representation method for enhancing three-dimensional spatial understanding of a large language model, characterized in that: include: The scene point cloud representation is obtained by sampling and position shifting a plurality of subset points, and the scene point cloud representation is clustered based on the plurality of subset points to construct a plurality of visual reference objects, wherein the plurality of visual reference objects serve as the first-level three-dimensional representation of the corresponding local area in the scene point cloud representation; Performing global spatial distribution modeling on the multiple visual reference objects through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects; The second-level three-dimensional representations of the multiple visual reference objects and the scene point cloud representation are processed through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation. The position fine-tuning network is a neural network that adjusts the spatial positioning of each of the visual reference objects based on real coordinates.

2. The method according to claim 1, characterized in that The step of obtaining a plurality of subset points by sampling and position shifting the scene point cloud representation, and clustering the scene point cloud representation based on the plurality of subset points to construct a plurality of visual reference objects includes: Sampling the scene point cloud representation by a farthest point sampling algorithm to obtain a first number of sampling points that are uniformly distributed in space; Aligning each sampling point with the nearest object center through a first feedforward neural network to obtain a subset of points corresponding to the sampling point; The scene point cloud representation is clustered and pooled based on the first number of subset points to form a first number of visual reference objects.

3. The method according to claim 2, characterized in that The global spatial distribution modeling of the multiple visual reference objects is performed through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects, including: The first number of visual reference objects are input into a graph convolutional network, feature information of the visual reference objects is used to model graph nodes, and edge information is constructed using spatial distances between each of the visual reference objects to obtain a second-level three-dimensional representation of each of the visual reference objects.

4. The method according to claim 3, characterized in that The processing of the second-level three-dimensional representations of the multiple visual references and the scene point cloud representation through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation includes: Inputting the second-level three-dimensional representation of the visual reference object into a self-attention network, inputting the processing result output by the self-attention network and the scene point cloud representation into a cross-attention network, and updating the third-level three-dimensional representation of the visual reference object; The third-level three-dimensional representation is input into a position fine-tuning network to update the third-level three-dimensional representation.

5. The method according to claim 4, characterized in that The method further comprises: By means of a visual-linguistic bridge, the updated third-level three-dimensional representation is aligned to the language text space to obtain a visual prompt; The visual prompt is input into a large language model so that the large language model can understand and process the visual prompt.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Collect the point cloud of the target scene to obtain the scene point cloud; The scene point cloud is encoded by a point cloud encoder to obtain a scene point cloud representation.

7. The method according to claim 5, characterized in that The method further comprises: Configuring a center loss function and an inter-pair spatial constraint loss function for the position fine-tuning network to implement feedback adjustment of a neural network that generates a third-level three-dimensional representation from the second-level three-dimensional representation, a neural network that generates a second-level three-dimensional representation from multiple visual references, and a neural network that generates multiple visual references from a scene point cloud representation; The position fine-tuning network configures a center loss function and an inter-pair spatial constraint loss function; The center loss function is: Where M represents the number of visual references, represents the coordinates of the predicted visual reference object, Represents the coordinates of the actual visual reference object; The pairwise spatial constraint loss function: Where N represents the number of visual reference pairs, is the distance between the predicted visual reference pair (i, j), is the actual distance between the visual reference pair (i, j).

8. The method according to claim 7, characterized in that The method further comprises: An overall optimization loss function is set based on the center loss function, the pairwise space constraint loss function, and the generation loss function of the large language model, and feedback optimization is performed on the large language model, the visual-language bridge, the neural network corresponding to the third-level three-dimensional representation generated from the second-level three-dimensional representation, the neural network corresponding to the second-level three-dimensional representation generated from multiple visual references, and the neural network corresponding to the multiple visual references generated from the scene point cloud representation. The formula of the overall optimization loss function is: L total =L LLM +α1L center +α2L psc Among them, L LLM represents the generation loss function of the large language model, α1 represents the coefficient of the center loss function, and α2 represents the coefficient of the inter-space constraint loss function.

9. A scene representation device for enhancing three-dimensional spatial understanding of a large language model, characterized in that: The device comprises: A clustering abstraction module, configured to obtain a plurality of subset points by sampling and position shifting the scene point cloud representation, cluster the scene point cloud representation based on the plurality of subset points to construct a plurality of visual reference objects, and the plurality of visual reference objects serve as the first-level three-dimensional representation of the corresponding local area in the scene point cloud representation; A message transmission module, used for performing global spatial distribution modeling of the multiple visual reference objects through a message propagation mechanism to obtain a second-level three-dimensional representation of each of the visual reference objects; The interaction module is used to process the second-level three-dimensional representations of the multiple visual reference objects and the scene point cloud representation through a multi-layer attention mechanism and a position fine-tuning network to obtain a third-level three-dimensional representation. The position fine-tuning network is a neural network that adjusts the spatial positioning of each of the visual reference objects based on real coordinates.

10. The device according to claim 9, characterized in that The device further comprises: A visual-language bridge, for aligning the updated third-level three-dimensional representation to the language text space to obtain a visual prompt; A large language model is used to understand and process the visual cues.

Citation Information

Patent Citations

  • Industrial robot auxiliary programming method based on natural language

    CN111267097A

  • Scene analysis method based on point-by-point spatial attention mechanism

    CN111611879A

  • Video abstract generation method fusing local target features and global features

    CN113139468A

  • Method for monitoring movement of reference object through laser radar and robot positioning method and device

    CN115755070A

  • Visual question and answer method based on scene graph relation information enhancement

    CN116187349A

Cited By

  • Virtual-real fusion scene automatic construction method based on AIGC script generation

    CN121861245A