Three-dimensional point cloud recognition method based on visual angle specific cue word and related equipment
Through the design of multi-view projection and perspective-specific prompt words, combined with the image text alignment capabilities of the multi-source prompt thesaurus and the CLIP model, the data dependence and cross-modal alignment problems in three-dimensional point cloud recognition are solved, and high-precision zero-sample or few-sample recognition is achieved.
Patent Information
- Application Number
- CN202510475322.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-19
AI Technical Summary
The existing three-dimensional point cloud recognition technology has problems such as high data dependence, insufficient cross-modal alignment and insufficient utilization of multi-view information, resulting in unsatisfactory recognition accuracy.
Through the design of multi-view projection and perspective-specific prompt words, combined with a multi-source prompt thesaurus that can be learned, language model generation and hand-designed multi-source prompt words, dynamically optimize the view-specific prompt words, and use the image text alignment ability of the CLIP model to perform similarity calculations to achieve high-precision recognition.
Achieve high-precision three-dimensional point cloud recognition under zero or few samples conditions, improving the generalization ability and efficiency of recognition.
Smart Images

Figure CN120509381A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a three-dimensional point cloud recognition technology, and in particular to a three-dimensional point cloud recognition method based on perspective-specific prompt words and related equipment. Background Art
[0002] With the continuous advancement of spatial sensor technology, 3D point cloud data has attracted widespread attention in both academia and industry, with numerous applications in augmented reality (AR), virtual reality (VR), autonomous driving, and robotics. The following section will gradually introduce point cloud recognition technology, explaining its background and developments.
[0003] First, single-modal deep learning point cloud recognition methods focus on directly extracting and identifying target objects from 3D point cloud data, without relying on information from other modalities (such as images or language). Although this recognition method can achieve point cloud recognition, it has the following drawbacks: high dependence on data, poor performance in untrained environments, and poor performance for open-world tasks.
[0004] To address the aforementioned issues, many works have utilized multimodal pre-training methods and developed them into 3D point cloud recognition technology. The following briefly introduces the main work and shortcomings in this field. PointCLIP is a pioneering model in this field and a powerful 3D open-world learning framework. It combines CLIP (Contrastive Language Image Pretraining) and GPT (Generative Pretrained Transformer) for zero-shot 3D classification, segmentation, and detection tasks. It can significantly outperform single-modal recognition models without any training, demonstrating strong generalization capabilities and scalable application to a variety of 3D open-world tasks. PointCLIP V2 is a further study of the PointCLIP work. This method prompts CLIP to generate more realistic depth maps through a shape projection module, narrowing the domain gap between projected point clouds and natural images. It also uses the GPT model to generate 3D-specific text as input to the CLIP text encoder, thereby better aligning 3D data with pre-trained language knowledge. While the PointCLIP series of works has made significant progress in 3D open-world learning, several limitations remain. These works fail to fully exploit multimodal information from multiple viewpoints. The same cue words are used for text encoding and feature computation for each viewpoint, resulting in a loss of view-specific geometric information. ULIP 2 is a simple yet effective trimodal pre-training framework designed to address the limitations of existing frameworks in collecting linguistic descriptions of 3D shapes, which are not scalable and diverse enough. This framework creates scalable multimodal 3D data by rendering 2D images of 3D shapes from a fixed set of holistic viewpoints and leveraging a large multimodal model to generate detailed descriptions for each rendered image. ULIP 2 achieves significant results on downstream tasks such as zero-shot 3D classification, fine-tuning standard 3D classification, and 3D-to-language generation. However, this work leverages an external 3D encoder to improve feature expressiveness and does not fully explore multimodal pre-training methods. This results in a limited improvement in the recognition capabilities of the multimodal pre-trained model.
[0005] In summary, existing technologies in the field of 3D point cloud recognition suffer from the following major flaws and shortcomings: 1) Strong data dependence: Single-modal deep learning point cloud recognition methods require a large amount of labeled point cloud data for training, and the high cost of manual labeling limits the model's performance in data-scarce scenarios. 2) Insufficient cross-modal alignment: While point cloud recognition methods based on contrastive language-image pre-training can reduce reliance on labeled point cloud data, their text prompts are relatively simple in design, and they do not fully utilize multi-view information. They fail to consider semantic differences across different viewpoints and fail to fully capture the diversity and geometric characteristics of point cloud categories, resulting in suboptimal modal alignment and impacting recognition accuracy. Summary of the Invention
[0006] In order to at least solve one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a three-dimensional point cloud recognition method and related equipment based on perspective-specific prompt words of a comparative language-image pre-training model.
[0007] The first technical solution adopted by the present invention is:
[0008] A three-dimensional point cloud recognition method based on perspective-specific prompt words includes the following steps:
[0009] Acquire point cloud data, project the point cloud through multiple preset perspectives, and obtain multi-perspective images;
[0010] Input the multi-view images into the contrastive language-image pre-trained image encoder to obtain multi-view features;
[0011] Obtaining perspective-specific prompt words from a preset multi-source perspective-specific prompt word library, wherein each perspective corresponds to a set of perspective-specific prompt words;
[0012] Input the viewpoint-specific cue words into the contrastive language-image pre-trained text encoder to obtain text features;
[0013] Similarity is calculated based on multi-view features and text features, and recognition results are obtained based on the calculated similarity.
[0014] Furthermore, the multi-source perspective specific prompt word library is constructed in the following way:
[0015] Based on the category of the object represented by the point cloud to be identified, dedicated prompt words are assigned to different viewpoints; the viewpoint-specific prompt words are generated as follows:
[0016] 1) Learnable prompt words: A lightweight neural network generates trainable prompt word vectors, automatically optimizing their matching with point cloud features. The prompt words for each view are combined to form a complete prompt word combination.
[0017] 2) Language model generation prompts: Using a large language model to generate diverse category description text based on the categories of objects represented by the point cloud, taking into account the differences in geometric features from different perspectives;
[0018] 3) Manually designed prompt words: Design structured templates based on domain knowledge to meet the differentiation requirements of different perspectives.
[0019] Furthermore, the three-dimensional point cloud recognition method also includes a dynamic optimization mechanism for the combination of multi-source perspective specific prompt words:
[0020] Similarity calculation and screening: For each viewpoint, a greedy algorithm is used to calculate the similarity between the prompt word and the multi-view features of the point cloud through the CLIP text encoder. In each calculation, the prompt word that gives the highest similarity index is screened in real time.
[0021] Multiple rounds of iterative updates: In each round of iteration, the prompt word library is updated based on the current similarity calculation results;
[0022] Elimination of low-relevance descriptions: During the iteration process, we identify and eliminate those prompt words that have low similarity with the multi-view features of the point cloud, thereby improving the quality and effectiveness of the entire prompt word library.
[0023] Through the dynamic optimization mechanism, the multi-source perspective-specific prompt word library can be optimized in real time and efficiently to ensure that prompt words that highly match the multi-view features of the point cloud can be generated under different perspectives, thereby improving the overall recognition accuracy and efficiency.
[0024] Furthermore, the step of acquiring point cloud data and projecting the point cloud through a plurality of preset viewing angles to obtain a multi-view image includes:
[0025] The point cloud data is sampled to meet the point number standard of the input model, and the number of sampling points is 1024 points;
[0026] The point cloud is gridded to convert the data format into a grid, and then the noise is reduced through a convolutional network.
[0027] The point cloud data is rotated by a rotation matrix that stores multi-directional parameters so that the point cloud presents geometric features in different directions, and then the point cloud is projected into a two-dimensional 224*224 grayscale image.
[0028] Furthermore, the plurality of preset viewing angles include: a top view, a side view, and a front view;
[0029] Dedicated prompt words are assigned to each perspective. The prompt words for the top view emphasize the outline features of the object, while the prompt words for the side view highlight the structural features of the object. This perspective-specific prompt word design can help the model better understand the geometric information under different perspectives, thereby improving recognition accuracy.
[0030] Furthermore, the step of obtaining the perspective-specific prompt words from the preset multi-source perspective-specific prompt word library includes:
[0031] Select a set of prompt words from a multi-source viewpoint-specific prompt word library, or choose to preset a new set of prompt words. After the model inference is completed, the new set of prompt words can be added to the vocabulary. The number of prompt words is the number of viewpoints multiplied by the number of categories, where the number of viewpoints is the number of viewpoints selected when converting the point cloud data into a multi-view projection image, and the number of categories is the total number of categories of objects represented by the point cloud in the dataset used.
[0032] After model inference is completed, a dynamic optimization mechanism for the multi-source perspective-specific prompt word combination is initiated. For each perspective of each category, new prompt words are iteratively replaced, and the recognition result indicators are calculated. The prompt word combination is updated with the results of the indicator calculation as feedback and with the purpose of optimizing the indicators.
[0033] Furthermore, the three-dimensional point cloud recognition method further includes the steps of constructing and training a model:
[0034] Build a contrastive language-image pre-training model, which includes a text encoder and an image encoder;
[0035] Among them, the text encoder adopts the Transformer architecture, and the input text sequence is encoded into a high-dimensional vector; the image encoder is implemented using a convolutional neural network or Vision Transformer;
[0036] During the forward reasoning of the model, the multi-view features of the point cloud are combined with the text features of the view-specific prompt word combination to obtain the logical child, and the recognition result is obtained based on the logical child;
[0037] During model training, the parameters of the text encoder and image encoder are frozen, and the domain knowledge of downstream tasks is learned by adding a learnable network (MLP).
[0038] The second technical solution adopted by the present invention is:
[0039] A three-dimensional point cloud recognition system based on viewpoint-specific prompt words, comprising:
[0040] Point cloud acquisition module, used to acquire point cloud data, project the point cloud through multiple preset perspectives, and obtain multi-perspective images;
[0041] A view feature extraction module is used to input multi-view images into a comparative language-image pre-trained image encoder to obtain multi-view features;
[0042] A prompt word acquisition module is used to obtain perspective-specific prompt words from a preset multi-source perspective-specific prompt word library, where each perspective corresponds to a set of perspective-specific prompt words;
[0043] A text feature extraction module is used to input viewpoint-specific prompt words into the contrastive language-image pre-trained text encoder to obtain text features;
[0044] The recognition module is used to calculate similarity based on multi-view features and text features, and obtain recognition results based on the calculated similarity.
[0045] The third technical solution adopted by the present invention is:
[0046] An electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a three-dimensional point cloud recognition method based on perspective-specific prompt words as described above.
[0047] The fourth technical solution adopted by the present invention is:
[0048] A computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a three-dimensional point cloud recognition method based on perspective-specific prompt words as described above.
[0049] The fifth technical solution adopted by the present invention is:
[0050] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method for three-dimensional point cloud recognition based on viewpoint-specific prompt words.
[0051] The beneficial effects of the present invention are as follows: the present invention retains geometric information by projecting the point cloud into a multi-view image; assigns a dedicated prompt word to each view (such as emphasizing the outline in the top view and highlighting the structure in the side view), and integrates multi-view semantic features. By utilizing the image-text alignment capability of the CLIP model, the multi-view features of the point cloud and the multi-view semantic features are similarly calculated, thereby achieving high-precision recognition under zero-sample or few-sample conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present invention or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.
[0053] Figure 1 Schematic diagram of building a multi-source perspective specific prompt vocabulary in an embodiment of the present invention;
[0054] Figure 2 Schematic diagram of a dynamic optimization mechanism for a combination of multi-source perspective specific prompt words in an embodiment of the present invention;
[0055] Figure 3 This is an overall framework diagram of the multi-view projection feature and view-specific prompt word calculation strategy in an embodiment of the present invention;
[0056] Figure 4 is a schematic diagram of a three-dimensional point cloud recognition method based on viewpoint-specific prompt words in an embodiment of the present invention;
[0057] Figure 5 This is a flowchart of the steps of a three-dimensional point cloud recognition method based on perspective-specific prompt words in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application. For the step numbers in the following embodiments, they are provided only for the convenience of explanation and are not intended to limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0059] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present application. The singular forms of "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise clearly defined, words such as setting, installing, and connecting should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0060] In the description of this application, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on this application.
[0061] In the description of this application, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The terms "first" and "second" are used solely to distinguish technical features and are not to be construed as indicating or implying relative importance, or as implicitly specifying the number or order of the technical features indicated.
[0062] In the description of this application, "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0063] To overcome the problems of high data dependence, insufficient cross-modal alignment, and insufficient utilization of multi-view information in existing technologies, this paper proposes a 3D point cloud recognition solution based on dynamic multi-prompt words and multi-view fusion, which includes the following contents:
[0064] 1) Multi-view projection features and view-specific prompt word calculation strategy:
[0065] By projecting the point cloud into multi-view images, geometric information is preserved. Each view is assigned a dedicated cue word (e.g., top view emphasizes contours, side view highlights structures). Multi-view semantic features are integrated, and the CLIP model's image-text alignment capability is leveraged to calculate similarity between the multi-view features of the point cloud and the multi-view semantic features, achieving high-precision recognition in zero- or few-shot conditions.
[0066] 2) Multi-source perspective specific prompt word generation and optimization technology:
[0067] Combining learnable prompt words (optimized by neural networks), language model-generated prompt words (such as diverse text descriptions generated by GPT), and manually designed prompt words (injected with domain knowledge), we build a prompt word library covering category diversity.
[0068] For each perspective’s prompt word combination, a greedy algorithm is used to dynamically select the prompt word that best matches the point cloud features, thereby improving the accuracy of text-point cloud alignment.
[0069] Example 1
[0070] like Figure 5 As shown, this embodiment provides a three-dimensional point cloud recognition method based on perspective-specific prompt words, including the following steps:
[0071] S1. Obtain point cloud data and project the point cloud through multiple preset perspectives to obtain multi-perspective images;
[0072] S2. Input the multi-view image into the contrastive language-image pre-trained image encoder to obtain multi-view features;
[0073] S3, obtaining perspective-specific prompt words from a preset multi-source perspective-specific prompt word library, wherein each perspective corresponds to a set of perspective-specific prompt words;
[0074] S4, inputting the viewpoint-specific cue words into the contrastive language-image pre-trained text encoder to obtain text features;
[0075] S5. Calculate similarity based on the multi-view features and the text features, and obtain a recognition result based on the calculated similarity.
[0076] The present embodiment is explained in detail below with reference to the accompanying drawings and specific implementation methods.
[0077] (1) Multi-source perspective specific prompt word generation and optimization technology
[0078] (1.1) Multi-source prompt word generation
[0079] See also Figure 1 , use a variety of technologies to generate prompt words, and organically integrate the prompt words generated by different means into a database. The following introduces the technologies used.
[0080] 1) Learnable prompt words: Generate trainable prompt word vectors through lightweight neural networks (such as MLP) and automatically optimize their matching degree with point cloud features.
[0081] 2) Language model generation prompt words: Use large language models (such as Deepseek, GPT, etc.) to generate diverse category description texts (such as "wooden chair with armrests") to expand semantic coverage.
[0082] 3) Manually design prompt words: Design structured templates based on domain knowledge (such as "point cloud projection of [category] with [geometric features]") to ensure that key attributes are not missing.
[0083] (1.2) Viewpoint-specific prompt words:
[0084] See also Figure 1,Depending on the category of the object represented by the point cloud to be identified, ,special prompt words are assigned to different perspectives, which also include ,multiple sources. The following introduces the generation method of perspective-specific ,prompt words based on different sources.
[0085] 1) Learnable cue words: Similar to non-viewpoint-specific cue words, learnable cue word vectors are generated for the channels of a single view through a lightweight neural network (such as an MLP) to automatically optimize their matching with point cloud features. The cue words for each view are combined to form a complete cue word combination.
[0086] 2) Language model generation prompts: Use large language models (such as Deepseek, GPT, etc.) to generate diverse category description text (such as "circular base" from a top view and "vertical pillar" from a side view) based on the category of the object represented by the point cloud, taking into account the differences in its geometric features from different perspectives.
[0087] 3) Manually designed prompt words: Based on domain knowledge, structured templates are designed to meet the discrimination of different perspectives (such as "a point cloud projection of [category] observed from [perspective]").
[0088] (1.3) Dynamic Optimization Mechanism for Combining Specific Prompt Words from Multiple Sources
[0089] See also Figure 2 After establishing a multi-source perspective-specific prompt word library, the dynamic optimization mechanism optimizes the prompt words in the following ways:
[0090] 1) Real-time Similarity Calculation and Selection: For each viewpoint, the CLIP text encoder uses a greedy algorithm to calculate the similarity between the prompt word and the multi-view features of the point cloud. During each calculation, the prompt word that maximizes the similarity metric is selected in real time. This process is similar to the simple and efficient heuristic greedy algorithm, where each decision point (i.e., each viewpoint) optimizes overall performance based on the current best choice.
[0091] 2) Multiple Iterative Updates: The dynamic optimization mechanism is not completed in one go, but rather gradually optimizes the prompt word library through multiple iterations. In each iteration, the system updates the prompt word library based on the current similarity calculation results.
[0092] 3) Elimination of Low-Relevance Descriptions: During the iteration process, the system identifies and eliminates prompt words with low similarity to point cloud features. This is similar to environmental adaptability assessment in dynamic optimization. By testing the degree of match between prompt words and point cloud features, prompt words that are not suitable for the current environment (i.e., point cloud features) are removed, thereby improving the quality and effectiveness of the entire prompt word library.
[0093] Through the above-mentioned dynamic optimization mechanism, the system can optimize the multi-source perspective-specific prompt word library in real time and efficiently, ensuring that prompt words that are highly matched with point cloud features can be generated under different perspectives, thereby improving overall recognition accuracy and efficiency.
[0094] (2) Multi-view projection features and view-specific prompt word calculation strategy
[0095] (2.1) Multi-view projection technology
[0096] See also Figure 3 Projecting point clouds into multi-view images: This embodiment first projects 3D point cloud data into multi-view 2D images. This process uses multiple preset perspectives (such as top view, side view, and front view) to ensure that the geometric information of the point cloud is preserved from different angles. Compared with traditional single-view projection, multi-view projection can more comprehensively capture the shape and structural features of the object, thereby providing richer information for subsequent feature extraction and semantic alignment.
[0097] (2.2) Special prompt word allocation
[0098] Dedicated prompt words are assigned to each view. The previous section on generating multi-source view-specific prompt words explained how to optimize prompt words based on view characteristics. For example, prompt words for top views emphasize the object's outline, while prompt words for side views highlight its structural features. This view-specific prompt word design helps the model better understand geometric information from different viewpoints, thereby improving recognition accuracy.
[0099] (2.3) Multi-view semantic feature fusion
[0100] See also Figure 3 , this embodiment fuses the semantic features of multiple perspectives to form a comprehensive semantic feature vector. This process is achieved through an advanced feature fusion algorithm, which ensures that the semantic features of different perspectives can complement each other, thereby providing a more comprehensive semantic description. The fused semantic features can better reflect the overall characteristics of the object and provide more accurate input for subsequent similarity calculations. Then, the powerful image-text alignment capability of the CLIP model is used to calculate the similarity between the multi-perspective features of the point cloud and the multi-perspective semantic features. Through the pre-trained image-text alignment capability, the CLIP model can automatically learn the mapping relationship between the multi-perspective projection image features of the point cloud and the perspective-specific prompt word text features, thereby achieving efficient similarity calculation. This process can not only make full use of the pre-trained knowledge of the CLIP model, but also achieve high-precision recognition under zero-sample or few-sample conditions.
[0101] Through the above-mentioned technical means and measures, the present invention can achieve high-precision point cloud recognition under zero-sample or few-sample conditions, significantly improving the generalization ability and recognition accuracy of point cloud recognition, and providing a new solution for the field of point cloud recognition.
[0102] As some optional implementations, Figure 4 As shown, this embodiment discloses a three-dimensional point cloud recognition method based on perspective-specific prompt words, comprising the following steps:
[0103] Step 1: Data preprocessing.
[0104] Point cloud data can be obtained from public datasets, synthetic datasets, and sensor-collected datasets. The point cloud data is organized into 3D coordinates and sampled to meet the input model's point count standard, with a sampling point count of 1024 points. The point cloud is gridded to create a grid-like data format, and then processed through a convolutional network to reduce noise. Next, the 3D data is rotated using a rotation matrix that stores multi-directional parameters to reflect geometric features in different directions. The data is then projected into a 2D 224x224 grayscale image.
[0105] Step 2: Prepare perspective-specific prompt words.
[0106] You can select a set of prompt words from a multi-source perspective-specific library, or you can preset a new set of prompt words. After model inference is complete, you can choose to add the new set of prompt words to the library. The number of prompt words is the number of perspectives multiplied by the number of categories, where the number of perspectives is the number of perspectives selected when converting the point cloud data into a multi-perspective projection image, and the number of categories is the total number of categories of objects represented by the point cloud in the dataset used. After model inference is complete, you can choose to activate a dynamic optimization mechanism for the combination of multi-source perspective-specific prompt words. For each perspective in each category, new prompt words are iteratively replaced and recognition results are calculated. The prompt word combination is updated based on the results of the indicator calculation and the goal of indicator optimization.
[0107] Step 3: Build a deep learning model.
[0108] A language-image pre-training model that compares multi-view projection features with a view-specific cue word calculation strategy is constructed. The model consists of two encoders: text and image. Optionally, the text encoder uses a Transformer architecture, such as the BERT model. The input text sequence is encoded into a high-dimensional vector. The image encoder can use a convolutional neural network (CNN), such as ResNet, or a Vision Transformer (ViT). Taking ViT as an example, it first divides the multi-view projection image of the point cloud into multiple small patches, and then inputs these patches as a sequence into the Transformer for feature extraction.
[0109] During the forward reasoning of the model, the multi-view projection map features of the point cloud are combined with the view-specific prompt word combination features to obtain logical sub-segments, and the recognition results are obtained based on the logical sub-segments.
[0110] Step 4: Model training.
[0111] By freezing the parameters of the text encoder and image encoder, the learnable network (MLP) is augmented with domain knowledge for downstream tasks.
[0112] Step 5: Zero-shot recognition.
[0113] Powerful zero-shot recognition is a major advantage shared by the research field of this invention. The model can achieve excellent recognition results without any training, demonstrating its generalization capabilities. Therefore, conducting experiments on the zero-shot recognition capabilities of this embodiment is a key focus.
[0114] Step 6: Model evaluation.
[0115] Evaluate the trained model using an independent test dataset. Calculate metrics such as accuracy, recall, and F1 value to assess model performance.
[0116] Step 7: Point cloud recognition application.
[0117] The trained model is deployed in an actual point cloud recognition system to perform real-time recognition on input point clouds. The system receives the object-level point cloud output from the pre-processing process, quickly recognizes it using the model, and returns the recognition results for further visual understanding tasks.
[0118] In summary, the method of the present invention has at least the following advantages and beneficial effects compared to the prior art:
[0119] 1) The present invention uses multi-source prompt words to ensure sufficient representation capabilities for point clouds. In the prior art, when designing prompt word combinations to represent point clouds, there is a tendency to look for the optimal combination obtained from a single source, which results in insufficient robustness and generalization performance. In different scenarios or with different recognition objects, due to the single source of prompt words, it may be difficult for the comparative language-image pre-training model to achieve ideal performance. In view of this, the present invention has designed a strategy for integrating prompt words from multiple sources, organically combining prompt words obtained by different means, and forming a multi-source prompt word library with database functions and characteristics. Due to the diversity of its sources and content, the model can have a wider range of choices and higher flexibility when it is put into use.
[0120] 2) Optimize search strategies to ensure rapid application. This technique, based on the premise of establishing a multi-source cue vocabulary, utilizes the principles of a greedy algorithm to iteratively select optimal combinations from the library and eliminate suboptimal combinations. This technique, leveraging the powerful representation capabilities of the multi-source cue vocabulary, allows us to dynamically adjust the cue word combinations used for recognition when new datasets or new usage scenarios are introduced, achieving the ideal recognition result as quickly as possible.
[0121] 3) Multi-perspective specific strategy to ensure minimal loss of geometric information. The existing technology lacks in-depth exploration of the modal alignment of multiple perspective projection feature maps and prompt words of point clouds, which results in the loss of geometric information of point clouds from different perspectives. In response to this, the present invention designs a multi-perspective projection feature and perspective-specific prompt word calculation strategy, adds new channels to the geometric features of point clouds from different perspectives, and uses features fused from multi-perspective semantic features to perform feature calculations to obtain recognition results. In addition, by relying on domain knowledge, language models and other means, through analysis, understanding and experimentation, perspective-specific prompt words suitable for the new calculation strategy are designed. In addition, by utilizing the prompt word library and optimized search strategy, a recognition effect as close to the ideal situation as possible is achieved. This solves the lack of prior art in utilizing geometric information of point clouds from different perspectives, and achieves the best possible recognition effect.
[0122] Example 2
[0123] This embodiment provides a three-dimensional point cloud recognition system based on perspective-specific prompt words, including:
[0124] Point cloud acquisition module, used to acquire point cloud data, project the point cloud through multiple preset perspectives, and obtain multi-perspective images;
[0125] A view feature extraction module is used to input multi-view images into a comparative language-image pre-trained image encoder to obtain multi-view features;
[0126] A prompt word acquisition module is used to obtain perspective-specific prompt words from a preset multi-source perspective-specific prompt word library, where each perspective corresponds to a set of perspective-specific prompt words;
[0127] A text feature extraction module is used to input viewpoint-specific prompt words into the contrastive language-image pre-trained text encoder to obtain text features;
[0128] The recognition module is used to calculate similarity based on multi-view features and text features, and obtain recognition results based on the calculated similarity.
[0129] As an optional embodiment, the system implementation scheme is as follows:
[0130] 1) Hardware Architecture: Deployed on a GPU server, it supports real-time processing of point cloud streaming data (e.g., in industrial inspection scenarios). It integrates point cloud acquisition devices (e.g., LiDAR) and rendering computing modules.
[0131] 2) Software modules: Multi-source view-specific prompt word management library (text generation, storage, and retrieval interface); point cloud preprocessing and multi-view projection module; image feature and text extraction module; multi-view projection feature and view-specific prompt word calculation module; classifier, segmenter (supporting different recognition tasks); downstream task adapter network (supporting cross-dataset migration).
[0132] Since the system corresponds to the method of embodiment 1 of the present invention, and the principle of solving the problem by the system is similar to that of the method, the implementation of the system can refer to the implementation process of the above method embodiment, and the repeated parts will not be repeated.
[0133] Example 3
[0134] An embodiment of the present invention further provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the following Figure 5 A three-dimensional point cloud recognition method based on view-specific prompt words is shown.
[0135] It is understood that the memory may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created based on the use of the server, etc.
[0136] The processor may include one or more processing cores. The processor utilizes various interfaces and circuits to connect various components within the server. It executes various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory, as well as accessing data stored in memory. Optionally, the processor may be implemented using at least one of the following hardware forms: digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor may integrate one or a combination of a central processing unit (CPU) and a modem. The CPU primarily processes the operating system and application programs, while the modem handles wireless communications. It is understood that the modem may not be integrated into the processor and may be implemented separately via a single chip.
[0137] Since the electronic device is an electronic device corresponding to a three-dimensional point cloud recognition method based on perspective-specific prompt words in an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0138] Example 4
[0139] An embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 5 A three-dimensional point cloud recognition method based on view-specific prompt words is shown.
[0140] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0141] Since the storage medium is a storage medium corresponding to a three-dimensional point cloud recognition method based on perspective-specific prompt words in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0142] Example 5
[0143] In some possible implementations, various aspects of the methods of the embodiments of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is executed on a computer device, the program code is used to cause the computer device to perform the steps of the method for three-dimensional point cloud recognition based on viewpoint-specific prompt words according to various exemplary embodiments of the present application described above in this specification. The executable computer program code or "code" used to perform the various embodiments may be written in a high-level programming language such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0144] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0145] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0146] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made based on the essence of the present invention are intended to be covered by the scope of protection of the present invention.
Claims
1. A three-dimensional point cloud recognition method based on perspective-specific prompt words, characterized in that: The following steps are involved: Acquire point cloud data, project the point cloud through multiple preset perspectives, and obtain multi-perspective images; Input the multi-view images into the contrastive language-image pre-trained image encoder to obtain multi-view features; Obtaining perspective-specific prompt words from a preset multi-source perspective-specific prompt word library, wherein each perspective corresponds to a set of perspective-specific prompt words; Input the viewpoint-specific cue words into the contrastive language-image pre-trained text encoder to obtain text features; Similarity is calculated based on multi-view features and text features, and recognition results are obtained based on the calculated similarity.
2. The three-dimensional point cloud recognition method based on perspective-specific prompt words according to claim 1, characterized in that: The multi-source perspective specific prompt word library is constructed in the following way: Based on the category of the object represented by the point cloud to be identified, dedicated prompt words are assigned to different viewpoints; the viewpoint-specific prompt words are generated as follows: 1) Learnable prompt words: A lightweight neural network generates trainable prompt word vectors, automatically optimizing their matching with point cloud features. The prompt words for each view are combined to form a complete prompt word combination. 2) Language model generation prompts: Using a large language model to generate diverse category description text based on the categories of objects represented by the point cloud, taking into account the differences in geometric features from different perspectives; 3) Manually designed prompt words: Design structured templates based on domain knowledge to meet the differentiation requirements of different perspectives.
3. The three-dimensional point cloud recognition method based on perspective-specific prompt words according to claim 1, characterized in that: The three-dimensional point cloud recognition method also includes a dynamic optimization mechanism for the combination of multi-source perspective specific prompt words: Similarity calculation and screening: For each viewpoint, a greedy algorithm is used to calculate the similarity between the prompt word and the multi-view features of the point cloud through the text encoder. In each calculation, the prompt word that can make the similarity index the highest is selected; Multiple rounds of iterative updates: In each round of iteration, the prompt word library is updated based on the current similarity calculation results; Elimination of low-relevance descriptions: During the iteration process, we identify and eliminate those prompt words that have low similarity with the multi-view features of the point cloud; To improve the quality and effectiveness of the entire prompt vocabulary; Through the dynamic optimization mechanism, the multi-source perspective-specific prompt word library can be optimized in real time and efficiently to ensure that prompt words that highly match the multi-view features of the point cloud can be generated under different perspectives, thereby improving the overall recognition accuracy and efficiency.
4. The three-dimensional point cloud recognition method based on perspective-specific prompt words according to claim 1, characterized in that: The step of acquiring point cloud data and projecting the point cloud through a plurality of preset viewing angles to obtain a multi-view image includes: Sampling the point cloud data to meet the point count standard of the input model; The point cloud is gridded to convert the data format into a grid, and then the noise is reduced through a convolutional network. The point cloud data is rotated by a rotation matrix that stores multi-directional parameters so that the point cloud presents geometric features in different directions, and then the point cloud is projected into a two-dimensional grayscale image.
5. The three-dimensional point cloud recognition method based on perspective-specific prompt words according to claim 1, characterized in that: The multiple preset perspectives include: top view, side view, and front view; Dedicated prompt words are assigned to each perspective. The prompt words for the top view emphasize the outline features of the object, while the prompt words for the side view highlight the structural features of the object. This perspective-specific prompt word design can help the model better understand the geometric information under different perspectives, thereby improving recognition accuracy.
6. The three-dimensional point cloud recognition method based on perspective-specific prompt words according to claim 1, characterized in that: The step of obtaining the perspective-specific prompt words from a preset multi-source perspective-specific prompt word library includes: Select a set of prompt words from a multi-source viewpoint-specific prompt word library, or choose to preset a new set of prompt words. After the model inference is completed, the new set of prompt words can be added to the vocabulary. The number of prompt words is the number of viewpoints multiplied by the number of categories, where the number of viewpoints is the number of viewpoints selected when converting the point cloud data into a multi-view projection image, and the number of categories is the total number of categories of objects represented by the point cloud in the dataset used. After model inference is completed, a dynamic optimization mechanism for the multi-source perspective-specific prompt word combination is initiated. For each perspective of each category, new prompt words are iteratively replaced, and the recognition result indicators are calculated. The prompt word combination is updated with the results of the indicator calculation as feedback and with the purpose of optimizing the indicators.
7. The three-dimensional point cloud recognition method based on perspective-specific prompt words according to claim 1, characterized in that: The three-dimensional point cloud recognition method further includes the steps of constructing and training a model: Build a contrastive language-image pre-training model, which includes a text encoder and an image encoder; Among them, the text encoder adopts the Transformer architecture, and the input text sequence is encoded into a high-dimensional vector; the image encoder is implemented using a convolutional neural network or Vision Transformer; During the forward reasoning of the model, the multi-view features of the point cloud are combined with the text features of the view-specific prompt word combination to obtain the logical child, and the recognition result is obtained based on the logical child; During model training, the parameters of the text encoder and image encoder are frozen, and domain knowledge of downstream tasks is learned by adding a learnable network.
8. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The computer program product comprises computer instructions, which are used to perform the method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Cited By
LLM-based zero sample precast concrete reinforcement point cloud segmentation method and device
CN121504962A
A Zero-Sample Precast Concrete Reinforcing Steel Point Cloud Segmentation Method and Device Based on LLM
CN121504962B