Three-dimensional dense description method based on knowledge guide feature enhancement and multi-modal fusion

By building a three-dimensional intensive description network based on knowledge-guided feature enhancement and multimodal fusion, combined with ConceptNet knowledge graph and multimodal fusion technology, the shortcomings of three-dimensional data description in the existing technology are solved, and a three-dimensional description with higher accuracy and robustness are achieved.

CN120354344APending Publication Date: 2025-07-22SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415249.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing three-dimensional data description methods have shortcomings in generalization ability, interpretability and feature expression, and have failed to effectively utilize prior knowledge for feature enhancement and multimodal fusion, resulting in insufficient accuracy and robustness of three-dimensional description.

Method used

A three-dimensional intensive description method based on knowledge-guided feature enhancement and multimodal fusion is adopted, and a detailed natural language description is generated through the end-to-end network of the three-dimensional object detection module, the knowledge-guided feature enhancement module, the multimodal fusion module and the text generation module, combined with the ConceptNet knowledge graph and multimodal fusion technology.

Benefits of technology

It significantly enhances the semantic expression ability of the target, provides a new way of information fusion, and generates descriptions that are more accurate and detailed, improving the accuracy and robustness of the three-dimensional description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354344A_ABST
    Figure CN120354344A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional dense description method based on knowledge guide feature enhancement and multi-modal fusion. The method comprises the following steps: firstly, constructing and training an end-to-end three-dimensional dense description network comprising a three-dimensional target detection module, a knowledge guide feature enhancement module, a multi-modal fusion module and a text generation module; in a description process, an input scene point cloud is firstly converted into visual features through a three-dimensional target detection module, a knowledge guide feature enhancement module is utilized to retrieve and encode related knowledge features from a knowledge graph by taking a target category as a query condition, and the knowledge features and the visual features are fused into a target feature vector through a multi-modal fusion module; and inputting the target feature vector into a text generation module, and generating a natural language text of each target in the scene and the mutual spatial relationship thereof. Compared with the prior art, the method has the advantages that by introducing the prior semantic knowledge, the generated description text is more accurate and detailed in the aspects of target details, spatial relations and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and artificial intelligence, and particularly relates to a three-dimensional dense description method based on knowledge-guided feature enhancement and multi-modal fusion, which can be used for tasks such as three-dimensional object recognition, scene understanding, and point cloud analysis. Background Art

[0002] With the continuous development of three-dimensional vision technology, three-dimensional data has been widely used in fields such as autonomous driving, robotics, and medical imaging. However, due to problems such as sensor noise, perspective occlusion, and data loss, the representation of three-dimensional data is often incomplete or has large inter-modal differences. Therefore, how to effectively use prior knowledge for feature enhancement and at the same time combine multiple modal information to improve the accuracy and robustness of three-dimensional dense description has become the focus of current research.

[0003] Currently, research on three-dimensional data description mainly focuses on deep learning methods: with the rise of convolutional neural networks and graph neural networks, more and more research uses deep learning models to extract global or local features from three-dimensional data. Such methods can automatically learn the implicit patterns and high-level semantic information in the data. However, due to the small amount of three-dimensional data and the complex data distribution, deep learning models still have deficiencies in generalization ability, interpretability, and feature expression. To make up for the deficiencies of a single visual modality, in recent years, some methods have begun to attempt to combine two-dimensional images, text descriptions, and even other sensor data for multi-modal fusion to improve the accuracy and robustness of three-dimensional description. Multi-modal fusion methods can make full use of the complementarity of each modal data in information expression. For example, two-dimensional images provide rich texture and color information, while three-dimensional point clouds provide accurate geometric shapes.

[0004] Existing methods have ignored the important role of prior knowledge in scene understanding. Knowledge graphs contain a large amount of information about object attributes, functions, and relationships between concepts, which is of great significance for supplementing the missing semantics in scene point clouds and improving the completeness of descriptions. However, effectively integrating external knowledge with visual features requires both efficient encoding of prior knowledge and the design of a mechanism that can perform cross-modal fusion and dynamically adjust feature weights to achieve a comprehensive description of complex scenes. Summary of the Invention

[0005] The main purpose of the present invention is to overcome the deficiencies of the prior art and provide a three-dimensional dense description method based on knowledge-guided feature enhancement and multi-modal fusion.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A 3D dense description method based on knowledge-guided feature enhancement and multimodal fusion, the 3D dense description method comprising the following steps: S1. Construct a 3D dense description network based on knowledge-guided feature enhancement and multimodal fusion, the network comprising a 3D object detection module, a knowledge-guided feature enhancement module, a multimodal fusion module, and a text generation module connected in sequence; S2. Input the 3D indoor scene point cloud obtained by lidar scanning of a real indoor scene , and perform feature encoding and object detection on through the 3D object detection module, and extract the 3D detection box features of each target object in , and map the 3D detection box features to obtain the target category ; ; S3. Use the knowledge-guided feature enhancement module, with the target category as the query condition, retrieve the knowledge nodes related to each target from the ConceptNet knowledge graph, and use a pre-trained word embedding model to vectorize the knowledge nodes, and obtain knowledge features through a graph convolutional network ; S4. Use the multimodal fusion module to perform multimodal fusion on the knowledge features and the 3D detection box features to generate knowledge-enhanced target features ; S5. Input the target features into the text generation module, the text generation module adopts an autoregressive decoding structure, and in each decoding process, based on the currently generated text sequence and the fusion features, predicts the next decoding process, and gradually generates a natural language description of each target and its mutual relationship in the scene ; S6. Use the 3D dense description dataset to perform end-to-end training on the 3D dense description network. During the training process, by minimizing the joint objective function of the 3D object detection loss and the text generation loss, use the backpropagation algorithm to update the parameters of the 3D dense description network until the network converges; S7. Input the 3D indoor scene point cloud to be described into the trained 3D dense description network to obtain the natural language description of each target object in the scene.

[0007] Further, the working process of the 3D object detection module is as follows: S21. Perform spatial partitioning on each point of a given 3D indoor scene point cloud to obtain a sparse voxel grid , to solve the problem of sparse and irregular point cloud data, facilitating subsequent sparse convolution operations, and performing sparse convolution on the sparse voxel grid to obtain a high-resolution feature map through sparse convolution , for the high-resolution feature map to obtain a low-resolution feature map through sparse convolution , for the low-resolution feature map to perform transposed convolution to obtain a low-resolution feature map with the same size as the high-resolution feature map , for the low-resolution feature map and the high-resolution feature map to add them together to obtain a global multi-scale feature map ; S22. Randomly sample the feature space of the global multi-scale feature map to obtain an initial 3D content query . Perform a fully connected mapping on the coordinates of the 3D content query to obtain a 3D position query . Add the 3D content query and the 3D position query to obtain a 3D object query ; S23. Perform self-attention operation on the 3D object query to obtain context features . Perform cross-attention operation on the context features and the global multi-scale feature map to obtain a 3D detection box ; Among them, the self-attention operation process is as follows: Generate a query matrix , a key matrix , and a value matrix through three groups of different linear layers for the 3D object query , where , , are the first, second, and third learnable weight matrices; Calculate the similarity between the query matrix and the key matrix to obtain a similarity score ; Scale the similarity score to obtain a scaled score , is the feature dimension; Perform probability normalization on the scaled score to obtain an attention matrix , represents the dimension of the scaled score, represents the th item of the scaled score, ​Indicates exponentiation; uses the attention matrix For the value matrix Perform weighted summation to obtain the context feature representation , Indicates the th row of the attention matrix Indicates the th column of the value matrix; Self-attention calculation can efficiently capture the global dependencies at any position in the sequence, and has the ability of parallel computing and better feature expression ability; Among them, the cross-attention operation process is as follows. The context feature and the global multi-scale feature map Generate the query matrix , the key matrix and the value matrix through three groups of different linear layers, where , , is a learnable weight matrix; Perform dot product similarity calculation on the query matrix and the key matrix to obtain the similarity score ; Perform scaling processing on the similarity score to obtain the scaled score ; Perform probability normalization on the scaled score to obtain the attention matrix ; Use the attention matrix to perform weighted summation on the value matrix to obtain the 3D detection box , Indicates the th row of the attention matrix Indicates the th column of the value matrix; Cross-attention is an extended attention mechanism that can capture the information correlation between different modalities or different levels, and is often used in multi-modal learning and feature alignment.

[0008] Furthermore, the working process of the knowledge-guided feature enhancement module is as follows: Input the category name as a keyword into the query interface of the ConceptNet knowledge graph to retrieve the nodes that are similar to or exactly match the target category name , query the neighbor nodes of the node to obtain the candidate knowledge subgraph , represents the node set, represents the edge set, perform graph convolution operation on the candidate knowledge subgraph to obtain the knowledge feature ; The ConceptNet knowledge graph was proposed by Speer et al. in the paper "ConceptNet 5.5: An Open Multilingual Graph of General Knowledge", which contains a large amount of general common sense knowledge and provides support for natural language understanding and reasoning; Among them, the graph convolution operation process is as follows: For the node set each node in is input into Word2Vec to obtain the initial node feature as , Word2Vec is a technology that maps words into continuous vector representations. For the initial node feature a message propagation operation is performed to obtain the second node feature = , represents the union of the neighbor set of node and node , is a normalization constant, is the th layer of the learnable weight matrix, The function represents taking the larger value between x and 0; The above message propagation operation is repeated times to obtain the final node feature . For the final node feature of each node, global average pooling is performed to obtain the knowledge feature ; Word2Vec was proposed by Mikolov et al. in the paper "Efficient Estimation of Word Representations in Vector Space" and can convert words into vector representations to capture semantic similarities and context relationships; Graph convolution was proposed by Kipf and Welling in 2016 in the paper "Semi-Supervised Classification with Graph Convolutional Networks", and its advantage lies in being able to efficiently extract node features on the graph structure, capture local and global structure information, and at the same time have good scalability and generalization ability.

[0009] Furthermore, the working process of the multi-modal fusion module is as follows: For the three-dimensional detection box feature and the knowledge feature query matrices = , key matrices and value matrices are respectively generated through three groups of different linear layers , where , , are the first, second, and third learnable weight matrices; similarity calculation is performed on the query matrix and the key matrix to obtain a similarity score ; the similarity score is scaled to obtain a scaled score , is the feature dimension; probability normalization is performed on all scaled scores to obtain an attention matrix ; for the attention matrix weighted summation is performed on the value matrix to obtain a knowledge-enhanced target feature .

[0010] Furthermore, the working process of the text generation module is as follows: A linear transformation is performed on the knowledge-enhanced target feature to obtain an initial query embedding , is the fourth learnable weight matrix, and self-attention operation is performed on the initial query embedding to obtain a context relationship , and cross-attention operation is performed on the context relationship and the knowledge-enhanced target feature to obtain an enhanced sequence representation , and a multi-layer perceptron is used to operate on the enhanced sequence representation to obtain an output , probability normalization is performed on the output to obtain a word probability distribution , the maximum value is selected from the word probability distribution to obtain the output word at the current time step , and the output word is used as a new query embedding, and the above process is repeated until the end-of-sequence token END is generated or the maximum length is reached; the multi-layer perceptron was proposed by Chen et al. in the paper "End-to-End 3D Dense Captioning with Vote2Cap-DETR", has a simple structure, is easy to train, and can capture complex feature relationships and patterns through non-linear activation functions; Among them, the self-attention operation process is as follows: the query embedding is respectively passed through three different linear layers to generate a query matrix , a key matrix and a value matrix , where , , are the first, second, and third learnable weight matrices; perform dot product similarity calculation on the query matrix and the key matrix ; perform scaling processing on the similarity scores to obtain scaled scores ; , is the feature dimension; perform probability normalization on the scaled scores to obtain the attention matrix ; perform weighted summation on the attention matrix on the value matrix to obtain the context relationship represents the th row of the attention matrix, represents the th column of the value matrix; Among them, the cross-attention operation process is as follows: generate the query matrix and the knowledge-enhanced target feature through three different linear layers to generate the query matrix , the key matrix , and the value matrix , where , , are learnable weight matrices; perform dot product similarity calculation on the query matrix and the key matrix to obtain the similarity scores ; perform scaling processing on the similarity scores to obtain the scaled scores ; is the feature dimension; perform probability normalization operation on all scaled scores to obtain the attention matrix ; use the attention matrix to perform weighted summation on the value matrix to obtain the enhanced sequence representation .

[0011] Furthermore, the joint objective function is defined as follows:

[0012] Among them, is the center point regression loss term, is the size regression loss term, is the text description generation loss term; The center point regression loss term is calculated by the following formula:

[0013] Among them, represents the predicted center coordinates, represents the actual center coordinates, represents the smooth loss, which is calculated by the following formula:

[0014] The dimension regression loss term is calculated by the following formula:

[0015] where, represents the predicted dimension, represents the actual dimension; The text description generation loss term is calculated by the following formula:

[0016] where represents the probability distribution corresponding to the word output at each time step of the output.

[0017] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. The prior knowledge and visual information are effectively fused. By using the ConceptNet knowledge graph to retrieve relevant knowledge nodes and performing multimodal fusion with the detection boxes extracted by the 3D object detection module, the present invention can significantly enhance the semantic expression ability of the object.

[0018] 2. By mapping the 3D object detection box and knowledge features to a unified embedding space and adopting a cross-attention mechanism for adaptive alignment, the present invention proposes a novel multimodal information fusion method, which not only retains the detailed information of the object but also fully introduces external semantic knowledge, providing a new idea for information fusion. Brief Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a flowchart of the 3D dense description method disclosed in the present invention, including a training process and a testing process; Figure 2 is a schematic flowchart of knowledge-guided feature enhancement in the 3D dense description method disclosed in the present invention; Figure 3It is a schematic diagram of the composition of the multimodal fusion module in the three-dimensional dense description method disclosed in the present invention; Figure 4 It is a comparison result diagram of the three-dimensional dense description method disclosed in the present invention and the existing method when the number of layers of the graph convolutional network is 2; Figure 5 It is a comparison result diagram of the three-dimensional dense description method disclosed in the present invention and the existing method when the number of layers of the graph convolutional network is 4. Detailed implementation manners

[0021] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.

[0022] Referring to "embodiment" in this application means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.

[0023] Embodiment 1 This embodiment discloses a three-dimensional dense description method based on knowledge-guided feature enhancement and multimodal fusion, which specifically includes the following steps: S1. Construct a three-dimensional dense description network based on knowledge-guided feature enhancement and multimodal fusion, which includes a three-dimensional object detection module, a knowledge-guided feature enhancement module, a multimodal fusion module, and a text generation module connected in sequence; S2. Input the three-dimensional indoor scene point cloud obtained by lidar scanning of a real indoor scene , and perform feature encoding and object detection on through the three-dimensional object detection module, and extract the three-dimensional detection box features of each target object in , and map the three-dimensional detection box features to obtain the target category ; ; S3. Use the knowledge-guided feature enhancement module with the target category Taking [the query condition], retrieve the knowledge nodes related to each target from the ConceptNet knowledge graph, and use a pre-trained word embedding model to vectorize the knowledge nodes, and obtain knowledge features through a graph convolutional network ; Among them, the number of layers of the graph convolutional network is 2; S4. Use the multimodal fusion module to fuse the knowledge features with the 3D detection box features to perform multimodal fusion and generate target features enhanced by knowledge ; S5. Input the target features into the text generation module. The text generation module adopts an autoregressive decoding structure, and in each decoding process, based on the currently generated text sequence and fusion features, predicts the next decoding process, and gradually generates a natural language description of each target and its mutual relationship in the scene ; S6. Use the 3D dense description dataset to perform end-to-end training on the 3D dense description network. During the training process, by minimizing the joint objective function of the 3D object detection loss and the text generation loss, use the backpropagation algorithm to update the parameters of the 3D dense description network until the network converges; Among them, the 3D dense description dataset uses the ScanRefer dataset proposed by Chen et al. in the paper "ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language". This dataset contains 8,060 real indoor scenes, 11,046 target objects, 51,583 descriptions of objects, and on average each target object corresponds to 3 to 4 descriptions under different references; S7. Input the point cloud of the scene to be described into the trained 3D dense description network to obtain the natural language description of each target object in the scene.

[0024] Joint objective function is defined as follows:

[0025] Among them, is the center point regression loss term, is the size regression loss term, is the text description generation loss term; Center point regression loss term is calculated by the following formula:

[0026] Among them, represents the predicted center coordinates, represents the actual center coordinates represents the smooth loss, which is calculated by the following formula:

[0027] The size regression loss term is calculated by the following formula:

[0028] where represents the predicted size represents the actual size; The text description generation loss term is calculated by the following formula:

[0029] where represents each time step the probability distribution corresponding to the output word.

[0030] During the training process, the initial learning rate is set to 0.0001, and the learning rate decays to 0.7 times the original every 20 epochs, and a total of 200 epochs are trained.

[0031] The training process of the entire network is as Figure 1 shown, the hardware parameters of the used training platform are shown in Table 1, and the software parameters are shown in Table 2. After the network training is completed, referring to the Figure 1 test process in, the scene point cloud is given as input for three-dimensional dense description to obtain natural language text.

[0032] Table 1. Table of Hardware Environment Parameters

[0033] Table 2. Table of Software Environment Parameters

[0034] Table 3 shows the quantitative comparison results of the current popular methods Scan2Cap, D3Net, Vote2Cap-DETR and the present invention on the test set of the ScanRefer dataset. The evaluation metrics are CIDEr, BLEU, ROUGE, and METEOR, and the larger the value, the better. It can be seen from the table that the method disclosed in the present invention exceeds the existing methods in all four metrics, proving the vividness of the natural language generated by the present invention.

[0035] Table 3. Comparison Table of Quantitative Results of the Three-Dimensional Dense Method Disclosed in the Present Invention and Existing Methods ( )

[0036] As Figure 4 shown, on the left is a part of the objects in the scene point cloud, the first row on the right is the ground truth description, the second row is the description obtained in step S7, and the third row is the description generated by the method Vote2cap-DETR. As can be seen from Figure 4 this, the description generated by the present invention correctly outputs the position of the round table in the room, while the description generated by the method Vote2cap-DETR misjudges the position of the round table as the center.

[0037] Embodiment 2 This embodiment discloses a three-dimensional dense description method based on knowledge-guided feature enhancement and multimodal fusion, which specifically includes the following steps: S1. Construct a three-dimensional dense description network based on knowledge-guided feature enhancement and multimodal fusion. This network includes a three-dimensional object detection module, a knowledge-guided feature enhancement module, a multimodal fusion module, and a text generation module connected in sequence; S2. Input the three-dimensional indoor scene point cloud obtained by lidar scanning of the real indoor scene , and perform feature encoding and object detection on it through the three-dimensional object detection module, and extract the three-dimensional detection box features of each target object in , and map the three-dimensional detection box features to obtain the target category ; S3. Use the knowledge-guided feature enhancement module to retrieve the knowledge nodes related to each target from the ConceptNet knowledge graph with the target category as the query condition, and use the pre-trained word embedding model to vectorize the knowledge nodes, and obtain the knowledge features through the graph convolutional network ; Among them, the number of layers of the graph convolutional network is 4; S4. Use the multimodal fusion module to perform multimodal fusion on the knowledge features and the three-dimensional detection box features to generate the target features enhanced by knowledge ; S5. Input the target features into the text generation module. This text generation module adopts an autoregressive decoding structure, and in each decoding process, it predicts the next decoding process based on the currently generated text sequence and the fusion features, and gradually generates a natural language description of each target and its mutual relationship in the scene ; ; ; S6. Use the 3D dense description dataset to perform end-to-end training on the 3D dense description network. During the training process, update the parameters of the 3D dense description network using the backpropagation algorithm by minimizing the joint objective function of the 3D object detection loss and the text generation loss until the network converges; S7. Input the scene point cloud to be described into the trained 3D dense description network to obtain the natural language descriptions of each target object in the scene.

[0038] Joint objective function is defined as follows:

[0039] where is the center point regression loss term, is the size regression loss term, is the text description generation loss term; Center point regression loss term is calculated by the following formula:

[0040] where represents the predicted center coordinates, represents the actual center coordinates, represents the smooth loss, which is calculated by the following formula:

[0041] Size regression loss term is calculated by the following formula:

[0042] where represents the predicted size, represents the actual size; Text description generation loss term is calculated by the following formula:

[0043] where represents the probability distribution corresponding to the word output at each time step During the training process, the initial learning rate is set to 0.0001, and the learning rate decays to 0.7 times the original every 20 epochs, and a total of 100 epochs are trained.

[0044] Table 4 shows the quantitative comparison results of the current popular methods Scan2Cap, D3Net, Vote2Cap-DETR and the present invention on the test set of the ScanRefer dataset. The evaluation metrics are CIDEr, BLEU, ROUGE, and METEOR, and the larger the value, the better. As can be seen from the table, the method disclosed in the present invention exceeds the existing methods in all four metrics, demonstrating the vividness of the natural language generated by the present invention.

[0045] Table 4. Comparison table of quantitative results between the three-dimensional dense method disclosed in the present invention and existing methods ( )

[0046] As Figure 5 shown, on the left is a part of the objects in the scene point cloud, the first row on the right is the ground truth description, the second row is the description obtained in step S7, and the third row is the description generated by the method Vote2cap-DETR. As can be seen from the figure, the description generated by the present invention correctly outputs the relative positional relationship between the door and the fireplace, while the description generated by the method Vote2cap-DETR misjudges the door to be on the left of the chair.

[0047] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.

[0048] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered as the scope described in this specification.

[0049] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A three-dimensional dense description method based on knowledge-guided feature enhancement and multimodal fusion, characterized in that The three-dimensional dense description method includes the following steps: S1. Construct a three-dimensional dense description network based on knowledge-guided feature enhancement and multi-modal fusion. The network includes a three-dimensional object detection module, a knowledge-guided feature enhancement module, a multi-modal fusion module, and a text generation module that are sequentially connected in order; S2. Input the three-dimensional indoor scene point cloud obtained by lidar scanning of the real indoor scene , and perform feature encoding and object detection on through the three-dimensional object detection module, and extract the three-dimensional detection box features of each target object in , and map the three-dimensional detection box features to obtain the target category ; S3. Use the knowledge-guided feature enhancement module with the target category as the query condition to retrieve knowledge nodes related to each target from the ConceptNet knowledge graph, and use a pre-trained word embedding model to vectorize the knowledge nodes, and obtain knowledge features through a graph convolutional network ; S4. Use the multi-modal fusion module to fuse the knowledge features with the 3D detection box features for multi-modal fusion to generate target features enhanced with knowledge ; S5. Input the target feature into the text generation module, which adopts an autoregressive decoding structure. In each decoding process, it predicts the next decoding process based on the currently generated text sequence and the fused feature, and gradually generates a natural language description of each target in the scene and their mutual relationships ; S6. Use the three-dimensional dense description dataset to perform end-to-end training on the three-dimensional dense description network. During the training process, update the parameters of the three-dimensional dense description network using the backpropagation algorithm by minimizing the joint objective function of the three-dimensional object detection loss and the text generation loss until the network converges; S7. Input the point cloud of the three-dimensional indoor scene to be described into the trained three-dimensional dense description network to obtain the natural language descriptions of each target object in the scene.

2. The three-dimensional dense description method based on knowledge-guided enhancement and multimodal fusion according to claim 1, wherein The working process of the three-dimensional object detection module is as follows: S21. Perform spatial partitioning on each point of a given three-dimensional indoor scene point cloud to obtain a sparse voxel grid . For the sparse voxel grid , perform sparse convolution to obtain a high-resolution feature map . For the high-resolution feature map , perform sparse convolution to obtain a low-resolution feature map . For the low-resolution feature map , perform transposed convolution to obtain a low-resolution feature map with the same size as the high-resolution feature map . For the low-resolution feature map and the high-resolution feature map , add them together to obtain a global multi-scale feature map ; ; S22. Randomly sample the feature space of the global multi-scale feature map to obtain an initial three-dimensional content query . For the three-dimensional content query , perform a fully connected mapping on its coordinates to obtain a three-dimensional position query . Add the three-dimensional content query and the three-dimensional position query to obtain a three-dimensional object query ; S23. Perform a 3D object query to obtain context features through self-attention operation , and perform cross-attention operation on the context features and the global multi-scale feature map to obtain 3D detection bounding boxes ; Among them, the self-attention operation process is as follows. The three-dimensional object query generates query matrices , key matrices , and value matrices respectively through three groups of different linear layers, where , , are the first, second, and third learnable weight matrices; perform similarity calculation on the query matrix and the key matrix to obtain a similarity score ; perform scaling processing on the similarity score to obtain a scaled score , is the feature dimension; perform probability normalization on the scaled score to obtain an attention matrix , represents the dimension of the scaled score, represents the -th term of the scaled score, represents the exponential operation; use the attention matrix to perform weighted summation on the value matrix to obtain a context feature representation , represents the -th row of the attention matrix, represents the -th column of the value matrix; Among them, the cross-attention operation process is as follows. The context features and the global multi-scale feature map generate a query matrix , a key matrix , and a value matrix through three groups of different linear layers respectively, where , , are learnable weight matrices; perform dot-product similarity calculation on the query matrix and the key matrix to obtain a similarity score ; perform scaling processing on the similarity score to obtain a scaled score ; perform probability normalization on the scaled score to obtain an attention matrix ; use the attention matrix to perform weighted summation on the value matrix to obtain a 3D detection box , represents the th row of the attention matrix, and represents the th column of the value matrix.

3. The three-dimensional dense description method based on knowledge-guided enhancement and multi-modal fusion according to claim 1, characterized in that The working process of the knowledge-guided feature enhancement module is as follows: Input the category name as a keyword into the query interface of the ConceptNet knowledge graph to retrieve nodes that are similar to or exactly match the target category name , query the neighbor nodes of the node to obtain a candidate knowledge subgraph , represents a set of nodes, represents a set of edges, perform a graph convolution operation on the candidate knowledge subgraph to obtain knowledge features ; Among them, the graph convolution operation process is as follows: For the node set each node in is input into Word2Vec to obtain the initial node feature as For the initial node feature as perform a message propagation operation to obtain the second node feature = , represents the union of the neighbor set of node and node , is the normalization constant, is the -th layer of the learnable weight matrix, The function means taking the larger value between x and 0; Repeat the above message propagation operation times to obtain the final node feature For the final node feature of each node perform global average pooling to obtain the knowledge feature .

4. The three-dimensional dense description method based on knowledge-guided enhancement and multimodal fusion according to claim 1, wherein The working process of the multi-modal fusion module is as follows: For the three-dimensional detection box features and the knowledge features Three different groups of linear layers are respectively used to generate query matrices = , key matrices and value matrices , where , , are the first, second, and third learnable weight matrices; the query matrix and the key matrix are used to calculate the similarity to obtain the similarity score ; The similarity score is scaled to obtain a scaled score , is the feature dimension; Perform a probability normalization operation on all scaling fractions to obtain an attention matrix ; For the attention matrix Perform a weighted sum on the value matrix to obtain the target feature enhanced by knowledge .

5. The three-dimensional dense description method based on knowledge-guided enhancement and multimodal fusion according to claim 1, characterized in that, The working process of the text generation module is as follows: Perform a linear transformation on the knowledge-enhanced target feature to obtain an initial query embedding , where is the fourth learnable weight matrix, and perform a self-attention operation on the initial query embedding to obtain a context relationship . Then, perform a cross-attention operation on the context relationship and the knowledge-enhanced target feature to obtain an enhanced sequence representation . Use a multi-layer perceptron to perform an operation on the enhanced sequence representation to obtain an output . Perform probability normalization on the output to obtain a word probability distribution . Select the maximum value from the word probability distribution to obtain the output word at the current time step . Use the output word as the new query embedding and repeat the above process until the end-of-sequence token END is generated or the maximum length is reached; Among them, the self-attention operation process is as follows: Embed the query Generate query matrix , key matrix and value matrix through three groups of different linear layers, where , , are the first, second, and third learnable weight matrices; perform dot product similarity calculation on the query matrix and the key matrix ; perform scaling processing on the similarity scores to obtain scaled scores , is the feature dimension; perform probability normalization on the scaled scores to obtain the attention matrix ; perform weighted summation on the attention matrix on the value matrix to obtain the context relationship , represents the th row of the attention matrix, represents the th column of the value matrix; Among them, the cross-attention operation process is as follows: the context relationship and the target features enhanced by knowledge generate a query matrix , a key matrix and a value matrix through three different linear layers respectively, where , , are learnable weight matrices; perform dot product similarity calculation on the query matrix and the key matrix to obtain a similarity score ; perform scaling processing on the similarity score to obtain a scaled score , is the feature dimension; perform probability normalization operation on all scaled scores to obtain an attention matrix ; use the attention matrix to perform weighted summation on the value matrix to obtain an enhanced sequence representation .

6. The three-dimensional dense description method based on knowledge-guided enhancement and multimodal fusion according to claim 1, wherein The combined objective function is defined as follows: Among them, is the center point regression loss term, is the size regression loss term, is the text description generation loss term; Center point regression loss term Calculated by the following formula: Among them, represents the predicted center coordinates, represents the actual center coordinates, represents the smoothing loss, which is calculated by the following formula: Size regression loss term It is calculated by the following formula: Among them, represents the predicted size, represents the actual size; Text description generation loss term Calculated by the following formula: Among them represents the probability distribution corresponding to the output word at each time step ​