Image recognition method based on point cloud-oriented local semantics and cross-level dependence

By designing multiple feature extraction modules and intermediary attention modules in the point cloud model, combined with the scaling residual block SRB, the problem of insufficient generalization and feature extraction capabilities of point cloud models in the existing technology is solved, and more efficient point cloud analysis and recognition effects are achieved.

CN120147655AActive Publication Date: 2025-06-13SOUTHWEST UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510206592.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Due to the limitations of model scale, insufficient generalization and feature extraction capabilities, the prior art is difficult to meet the needs of the ever-expanding point cloud analysis tasks.

Method used

A method of image recognition based on local semantics and cross-level dependence based on point cloud-oriented local semantics and cross-level dependence is designed. By constructing multiple feature extraction modules and mediating attention modules, context features across levels and long distances are captured, and the model is expanded by scaling the residual block SRB.

Benefits of technology

It effectively improves the generalization ability and feature extraction ability of the point cloud model, can capture information in the point cloud more deeply, and improves the accuracy and robustness of image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147655A_ABST
    Figure CN120147655A_ABST
Patent Text Reader

Abstract

An image recognition method based on point cloud-oriented local semantics and cross-level dependence is characterized by comprising the following steps: 1, constructing an image recognition system based on point cloud-oriented local semantics and cross-level dependence; 2, the image acquisition module acquires original image data; 3, a preprocessing module preprocesses the original image data to obtain standard data; 4, an embedding layer in the LCHNet network obtains standard data, the feature dimension of the standard data is improved, and embedded data is obtained; 5-8, performing feature extraction on the input data by each feature extraction module to obtain feature data; 9, performing intermediary attention mechanism calculation by an intermediary attention module to obtain multi-level feature data; 10, an addition unit adds the fourth feature data and the multi-level feature data to obtain comprehensive feature data; and 11, the image recognition module performs image recognition operation on the comprehensive feature data and outputs an image recognition result. The method has the effect that the generalization ability and the feature extraction ability of the point cloud model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to an image recognition method based on local semantics and cross-level dependencies for point clouds. Background Art

[0002] The analysis of 3D point cloud data is widely used in healthcare, autonomous driving, and embodied artificial intelligence. Different from traditional planar images, point cloud entities consist of a series of discrete points, showing the characteristics of spatial irregularity and randomness. These unstructured properties pose great challenges to the design of models and algorithms. Currently, there are mainly two methods for processing point cloud data.

[0003] One common strategy is to convert the point cloud into formatted data for further processing. According to the different forms of the converted data, these methods can be mainly divided into two categories. Another is the technology based on multi-view projection, which projects the point cloud data into multiple 2D views and then uses a traditional convolutional neural network for feature learning; another technology is to quantize the point cloud into a voxel grid and then use 3D convolution to process the quantized data, so as to achieve the understanding of the point cloud. PointNet introduced in 2017 marks the emergence of a new method for processing point clouds, which directly processes the original point cloud, thus saving unnecessary conversion costs.

[0004] Since the proposal of PointNet, the research on model expansion has been continuously deepened. However, most of the research only improves the neighborhood feature pooling module and does not expand the model itself, resulting in a small existing network scale and failing to fully explore its performance potential. Some studies have shown that increasing the depth of the network by adding feature extraction layers can slightly improve the performance. However, as the network gets deeper, effective feature information is often lost during the propagation process, indicating that this method may not be sustainable in the long term. In addition, simply deepening the network does not ensure the improvement of generalization ability.

[0005] Disadvantages of the prior art: With the continuous expansion of the point cloud scale, due to scale limitations, the existing models have insufficient generalization ability and feature extraction ability, and can no longer meet the requirements of point cloud analysis tasks. Summary of the Invention

[0006] An image recognition method based on local semantics and cross-level dependencies for point clouds provided by the present invention effectively improves the generalization ability and feature extraction ability of the point cloud model.

[0007] To achieve the above object, a key aspect of an image recognition method based on local semantics and cross-level dependencies for point clouds provided by the present invention includes the following steps:

[0008] Step 1: Construct an image recognition system based on local semantics and cross-level dependencies of point clouds. The image recognition system is provided with an image acquisition module, a preprocessing module, and a local semantics and cross-level dependency network LCHNet of point clouds connected in sequence;

[0009] The local semantics and cross-level dependency network LCHNet of point clouds is provided with an embedding layer, a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, an addition unit, and an image recognition module connected in sequence at the head and tail. The first feature extraction module, the second feature extraction module, the third feature extraction module, the fourth feature extraction module, and the addition unit are also connected with a mediation attention module;

[0010] Step 2: The image acquisition module collects the original image data a in real time and transmits it to the preprocessing module;

[0011] Step 3: The preprocessing module performs preprocessing operations on the original image data a to obtain standard data b and transmits it to the local semantics and cross-level dependency network LCHNet of point clouds;

[0012] Step 4: The embedding layer in the local semantics and cross-level dependency network LCHNet of point clouds obtains the standard data b, enhances the feature dimension of the standard data b to obtain embedded data c, and then transmits it to the first feature extraction module;

[0013] Step 5: The first feature extraction module performs feature extraction operations on the embedded data c to obtain first feature data d and transmits it to the second feature module and the mediation attention module;

[0014] Step 6: The second feature extraction module performs feature extraction operations on the first feature data d to obtain second feature data e and transmits it to the third feature module and the mediation attention module;

[0015] Step 7: The third feature extraction module performs feature extraction operations on the second feature data e to obtain third feature data f and transmits it to the fourth feature module and the mediation attention module;

[0016] Step 8: The fourth feature extraction module performs feature extraction operations on the third feature data f to obtain fourth feature data g and transmits it to the addition unit and the mediation attention module;

[0017] Step 9: The mediation attention module performs mediation attention mechanism calculations on the first feature data d, the second feature data e, the third feature data f, and the fourth feature data g, captures the long-range semantic information in each feature data, obtains multi-level feature data h, and transmits it to the addition unit;

[0018] Step 10: The addition unit adds the fourth feature data g and the multi-level feature data h to obtain the comprehensive feature data k, and transmits it to the image recognition module;

[0019] Step 11: The image recognition module performs an image recognition operation on the comprehensive feature data k and outputs an image recognition result.

[0020] Through the above design, multiple feature extraction modules are designed to obtain cross-level feature information, and then the mediation attention module is used to calculate and enhance the interaction between different granularity levels, delving deeper into deeper information, making up for the information loss brought by the backbone network directly processing the global point cloud, and effectively improving the feature extraction ability of the model.

[0021] In addition, a scaling residual block SRB is designed in the feature extraction module to expand the model. The scaling residual block SRB uses a residual structure to reduce information loss and enhance the ability to extract fine-grained features in the network, making the model have strong generalization performance.

[0022] Preferably: in the step 1, the first feature extraction module, the second feature extraction module and the fourth feature extraction module have the same structure, and are all provided with a first downsampling layer, a first scaling residual block, a first aggregation layer and a first post-MLP module connected in sequence;

[0023] The third feature extraction module is provided with a second downsampling layer, a second scaling residual block, a third scaling residual block, a second aggregation layer, a second post-MLP module and a third post-MLP module connected in sequence.

[0024] A total of four key downsampling layers are included in the four feature extraction modules. This continuous downsampling process gradually reduces the number of points of the object in the point cloud and increases the dimension of the features; the aggregation layer is used to aggregate local features into global features; the post-MLP module is used to perform feature processing on the received features.

[0025] Preferably: the first scaling residual block, the second scaling residual block and the third scaling residual block have the same structure, and are all provided with a first multi-layer perceptron layer, a second multi-layer perceptron layer, a batch normalization layer, a third multi-layer perceptron layer and a residual connection layer, and a ReLU activation function is built in the output end of the residual connection layer.

[0026] The scaling residual block SRB first processes the point cloud feature input using the first multi-layer perceptron layer, increasing the number of channels to four times the original and mapping the information to a high-dimensional space. Then, the second MLP layer continues to process these features, with the input and output channel counts remaining unchanged, both being 4 times the input feature channel count. This helps to more effectively understand different features and their relationships. After batch normalization, the features are passed to the final multi-layer perceptron layer, where the channel count is reduced to the required output count. Then, the processed features are concatenated with the original features residually and activated through the ReLU activation function to generate new feature information.

[0027] The scaling residual block SRB can capture finer feature details and enhance the generalization ability of the model. Its unique residual connection design not only retains the original feature information but also effectively prevents the problem of gradient disappearance, thereby reducing the common information loss in deep networks.

[0028] Preferably, the output expression of the scaling residual block is as follows:

[0029]

[0030] Where, represents the ReLU activation function, F i represents the i-th input data of the scaling residual block, φ 1 represents the first multi-layer perceptron layer, φ 2 represents the second multi-layer perceptron layer, φ 3 represents the third multi-layer perceptron layer, Υ represents the batch normalization layer, represents the output data of the scaling residual block.

[0031] By designing the scaling residual block SRB to expand the backbone model, the scaling residual block SRB only utilizes the multi-layer perceptron layer MLP to effectively capture finer-grained feature information, making the model have strong generalization performance. In addition, the scaling residual block SRB adopts a residual connection structure, which not only effectively prevents the problem of gradient disappearance during training but also significantly retains the original information during the feature extraction stage, minimizing the degradation of valuable data, achieving more comprehensive information acquisition, and reducing the loss of effective information.

[0032] Preferably, in the step 9, the mediation attention module is provided with a first multiplication unit, a second multiplication unit, a third multiplication unit, and a fourth multiplication unit, and Softmax activation functions are built in at the output ends of the first multiplication unit and the second multiplication unit;

[0033] The process by which the mediation attention module obtains the multi-level feature data h based on the first feature data d, the second feature data e, the third feature data f, and the fourth feature data g is as follows:

[0034] S1: The first multiplication unit obtains the first feature data d and the second feature data e, multiplies the two, and then outputs the first multiplication data x1 to the fourth multiplication unit through the Softmax activation function;

[0035] S2: The second multiplication unit obtains the first feature data d and the third feature data f, multiplies the two, and then outputs the second multiplication data x2 to the third multiplication unit through the Softmax activation function;

[0036] S3: The third multiplication unit obtains the second multiplication data x2 and the fourth feature data g, multiplies the two, and then outputs the third multiplication data x3 to the fourth multiplication unit;

[0037] S4: The fourth multiplication unit obtains the first multiplication data x1 and the third multiplication data x3, multiplies the two, and then outputs the multi-level feature data h to the addition unit.

[0038] The traditional triple attention module adopted by the existing network has limitations in the representation ability of point clouds. Simply stacking triple attention not only introduces additional computational overhead, but also fails to significantly enhance its representation ability and reduces its applicability in the scene.

[0039] The proposed mediation attention mechanism in the present invention can effectively promote the communication between different information levels and the extraction of long-distance features through four cross-level inputs, enhancing the model's ability to capture long-distance features and multi-level feature information.

[0040] Preferably: the output expression of the mediation attention module is as follows:

[0041] O = At S (Q, I, At φ (I, K, V)) = σ(QI T )σ(IK T )V

[0042] Wherein, O represents the multi-level feature data h; Q represents the first feature data d, i.e., the query matrix; I represents the second feature data e, i.e., the mediation matrix; K represents the third feature data f, i.e., the key matrix; V represents the fourth feature data g, i.e., the value matrix; At S represents the softmax attention mechanism algorithm, At φ represents the linear attention mechanism algorithm, σ represents the softmax activation function, and the superscript T represents the matrix transpose.

[0043] The intermediate attention module not only integrates the Softmax attention mechanism and the linear attention mechanism, but also innovates its input source. The outputs of the four feature extraction modules correspond to the four parameters Q, I, K, and V in sequence. The intermediate attention module accepts the outputs of four different levels as inputs, utilizing information at different scales. The intermediate matrix I transforms between the key matrix K and the query matrix Q, thereby facilitating communication across all functional levels. This process realizes the interaction of cross-level feature information and captures the long-range semantic information from the first feature extraction module to the fourth feature extraction module.

[0044] Preferably: in the step 3, the preprocessing operation includes but is not limited to moving, rotating, resizing the image, removing the background noise of the image, and normalization processing.

[0045] The preprocessing module is used to perform preprocessing operations on the original image data to make it meet the input requirements of the LCHNet network, reducing the interference of factors such as image background noise and occlusion on the image recognition task.

[0046] Preferably: in the step 11, the image recognition module is provided with an image classification module and an image partial segmentation module. The image classification module is used to perform image classification tasks, and the image partial segmentation module is used to perform image partial segmentation tasks;

[0047] The image recognition result is either an image classification recognition result; or an image partial segmentation recognition result.

[0048] In subsequent image recognition processing, different processing can be performed according to different tasks to meet various task requirements.

[0049] Preferably: the image classification module is provided with a first fully connected layer, a first batch normalization layer, a first dropout layer, a second fully connected layer, a second batch normalization layer, a second dropout layer, a third fully connected layer, and a classification output layer. The output ends of the first batch normalization layer and the second batch normalization layer are both built-in with ReLU activation functions;

[0050] The third fully connected layer outputs the probability distribution values of various image categories to the classification output layer, and the classification output layer outputs the image category corresponding to the maximum probability distribution as the image classification recognition result.

[0051] For the image classification task, the obtained comprehensive feature data k is directly input into the image classification module for classification, and the final target category is output.

[0052] Preferably: the image partial segmentation module is provided with a first upsampling layer, a second upsampling layer, a third upsampling layer, a fourth upsampling layer, a concat splicing layer, an information processing unit, and a partial segmentation classification head;

[0053] The input end of the first upsampling layer is connected to the output ends of the third feature extraction module and the fourth feature extraction module. The output end of the first upsampling layer is connected to the input end of the second upsampling layer. The input end of the second upsampling layer is also connected to the output end of the second feature extraction module. The output end of the second upsampling layer is connected to the input end of the third upsampling layer. The input end of the third upsampling layer is also connected to the output end of the first feature extraction module. The output end of the third upsampling layer is connected to the input end of the fourth upsampling layer. The input end of the fourth upsampling layer is also connected to the output end of the embedding layer. The output end of the fourth upsampling layer is connected to the input end of the concat splicing layer. The input end of the concat splicing layer is also connected to the output end of the addition unit. The output end of the concat splicing layer is connected to the input end of the partial segmentation classification head through an information processing unit, and the partial segmentation classification head outputs the recognition result of the image partial segmentation;

[0054] The partial segmentation classification head is provided with a fourth fully connected layer, a third batch normalization layer, a third dropout layer, and a fifth fully connected layer connected in sequence. The output end of the third batch normalization layer is built-in with a ReLU activation function.

[0055] For the partial segmentation task, more refined classification is required. Feature propagation in four stages is performed on the obtained features to gradually restore the point cloud of the original object and classify each point.

[0056] The information processing unit is used to process global context information and class label information, and then uses this information for the final partial segmentation recognition task.

[0057] Advantages of the present invention:

[0058] 1. The model is extended by adding a scaling residual block SRB to the backbone network. The scaling residual block SRB uses a residual structure to reduce information loss and enhance the ability of the network to extract fine-grained features;

[0059] 2. A mediating attention mechanism for point clouds is proposed. This mediating attention mechanism enables the model to capture cross-level and long-distance context features. The mediating attention mechanism effectively promotes communication between different information levels and the extraction of long-distance features through four cross-level inputs.

[0060] 3. The local semantic and cross-level dependency network LEHNet for point clouds achieves efficient feature extraction and can capture deep information between points. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0062] Figure 2 It is a schematic structural diagram of an image recognition system;

[0063] Figure 3 It is a schematic structural diagram of a local semantics and cross-level dependency network LEHNet for point clouds;

[0064] Figure 4 It is a schematic structural diagram of a scaling residual block SRB;

[0065] Figure 5 It is a schematic structural diagram of an image classification module;

[0066] Figure 6 It is a schematic diagram of the local semantics and cross-level dependency network LEHNet in the embodiment;

[0067] Figure 7 It is a visualization diagram of the ground truth and prediction results in a partial segmentation task in the embodiment;

[0068] Figure 8 It is a result curve graph of the robustness test in the embodiment. Detailed implementation manners

[0069] The present invention will be further described in detail below with reference to the accompanying drawings and specific examples. The following examples or drawings are used to illustrate the present invention, but not to limit the scope of the present invention.

[0070] As Figure 1 shown: An image recognition method based on local semantics and cross-level dependencies for point clouds includes the following steps:

[0071] Step 1: Construct an image recognition system based on local semantics and cross-level dependencies for point clouds. The image recognition system is provided with an image acquisition module, a preprocessing module, and a local semantics and cross-level dependency network LCHNet for point clouds that are connected in sequence;

[0072] As Figure 2 shown, the local semantics and cross-level dependency network LCHNet for point clouds is provided with an embedding layer, a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, an addition unit, and an image recognition module that are connected in sequence at the head and tail. The first feature extraction module, the second feature extraction module, the third feature extraction module, the fourth feature extraction module, and the addition unit are also connected to a mediation attention module;

[0073] Step 2: The image acquisition module real-time collects the original image data a and transmits it to the preprocessing module;

[0074] Step 3: The preprocessing module performs preprocessing operations on the original image data a to obtain standard data b, and transfers it to the Local Semantic and Cross-level Dependence Network for Point Cloud (LCHNet).

[0075] Step 4: The embedding layer in the Local Semantic and Cross-level Dependence Network for Point Cloud (LCHNet) obtains the standard data b, enhances the feature dimension of the standard data b to obtain embedded data c, and then transfers it to the first feature extraction module.

[0076] Step 5: The first feature extraction module performs feature extraction operations on the embedded data c to obtain first feature data d, and transfers it to the second feature module and the mediation attention module.

[0077] Step 6: The second feature extraction module performs feature extraction operations on the first feature data d to obtain second feature data e, and transfers it to the third feature module and the mediation attention module.

[0078] Step 7: The third feature extraction module performs feature extraction operations on the second feature data e to obtain third feature data f, and transfers it to the fourth feature module and the mediation attention module.

[0079] Step 8: The fourth feature extraction module performs feature extraction operations on the third feature data f to obtain fourth feature data g, and transfers it to the addition unit and the mediation attention module.

[0080] Step 9: The mediation attention module performs mediation attention mechanism calculations on the first feature data d, second feature data e, third feature data f, and fourth feature data g, captures the long-range semantic information in each feature data, obtains multi-level feature data h, and transfers it to the addition unit.

[0081] Step 10: The addition unit adds the fourth feature data g and the multi-level feature data h to obtain comprehensive feature data k, and transfers it to the image recognition module.

[0082] Step 11: The image recognition module performs image recognition operations on the comprehensive feature data k and outputs the image recognition result.

[0083] As Figure 3 shown, the first feature extraction module, the second feature extraction module, and the fourth feature extraction module have the same structure, and are all provided with a first downsampling layer, a first scaling residual block, a first aggregation layer, and a first post-MLP module connected in sequence.

[0084] The third feature extraction module is provided with a second downsampling layer, a second scaling residual block, a third scaling residual block, a second aggregation layer, a second post-MLP module, and a third post-MLP module connected in sequence.

[0085] As shown Figure 4 in the figure, the first scaling residual block, the second scaling residual block, and the third scaling residual block have the same structure, and are all provided with a first multi-layer perceptron layer, a second multi-layer perceptron layer, a batch normalization layer, a third multi-layer perceptron layer, and a residual connection layer connected in sequence. The output end of the residual connection layer is built-in with a ReLU activation function.

[0086] The output expression of the scaling residual block is as follows:

[0087]

[0088] where represents the ReLU activation function, F i represents the i-th input data of the scaling residual block, φ 1 represents the first multi-layer perceptron layer, φ 2 represents the second multi-layer perceptron layer, φ 3 represents the third multi-layer perceptron layer, Υ represents the batch normalization layer, represents the output data of the scaling residual block.

[0089] The intermediate attention module is provided with a first multiplication unit, a second multiplication unit, a third multiplication unit, and a fourth multiplication unit. The output ends of the first multiplication unit and the second multiplication unit are both built-in with a Softmax activation function;

[0090] The process by which the intermediate attention module obtains the multi-level feature data h according to the first feature data d, the second feature data e, the third feature data f, and the fourth feature data g is as follows:

[0091] S1: The first multiplication unit obtains the first feature data d and the second feature data e, multiplies the two, and then outputs the first multiplication data x1 to the fourth multiplication unit through the Softmax activation function;

[0092] S2: The second multiplication unit obtains the first feature data d and the third feature data f, multiplies the two, and then outputs the second multiplication data x2 to the third multiplication unit through the Softmax activation function;

[0093] S3: The third multiplication unit obtains the second multiplication data x2 and the fourth feature data g, multiplies the two, and then outputs the third multiplication data x3 to the fourth multiplication unit;

[0094] S4: The fourth multiplication unit obtains the first multiplication data x1 and the third multiplication data x3, multiplies the two, and then outputs the multi-level feature data h to the addition unit.

[0095] The output expression of the intermediate attention module is as follows:

[0096] O = At S (Q, I, At φ (I, K, V)) = σ(QI T )σ(IK T )V

[0097] where O represents the multi-level feature data h, Q represents the first feature data d, I represents the second feature data e, K represents the third feature data f, V represents the fourth feature data g, At S represents the softmax attention mechanism algorithm, Atφ represents the linear attention mechanism algorithm, σ represents the softmax activation function, and the superscript T represents matrix transpose.

[0098] In the step 3, the preprocessing operation includes moving, rotating, resizing the image, removing the background noise of the image, and normalization processing.

[0099] In the step 11, the image recognition module is provided with an image classification module and an image partial segmentation module. The image classification module is used for performing the image classification task, and the image partial segmentation module is used for performing the image partial segmentation task;

[0100] The image recognition result is either the image classification recognition result or the image partial segmentation recognition result.

[0101] As Figure 5 shown, the image classification module is provided with a first fully connected layer, a first batch normalization layer, a first dropout layer, a second fully connected layer, a second batch normalization layer, a second dropout layer, a third fully connected layer, and a classification output layer. The output ends of the first batch normalization layer and the second batch normalization layer are both built-in with ReLU activation functions;

[0102] The third fully connected layer outputs the probability distribution values of various image categories to the classification output layer, and the classification output layer outputs the image category corresponding to the maximum probability distribution as the image classification recognition result.

[0103] As Figure 3 、 Figure 6 shown, the image partial segmentation module is provided with a first upsampling layer, a second upsampling layer, a third upsampling layer, a fourth upsampling layer, a concat splicing layer, an information processing unit, and a partial segmentation classification head;

[0104] The input end of the first upsampling layer is connected to the output ends of the third feature extraction module and the fourth feature extraction module. The output end of the first upsampling layer is connected to the input end of the second upsampling layer. The input end of the second upsampling layer is also connected to the output end of the second feature extraction module. The output end of this second upsampling layer is connected to the input end of the third upsampling layer. The input end of the third upsampling layer is also connected to the output end of the first feature extraction module. The output end of this third upsampling layer is connected to the input end of the fourth upsampling layer. The input end of the fourth upsampling layer is also connected to the output end of the embedding layer. The output end of this fourth upsampling layer is connected to the input end of the concat splicing layer. The input end of the concat splicing layer is also connected to the output end of the addition unit. The output end of the concat splicing layer is connected to the input end of the partial segmentation classification head through an information processing unit, and the partial segmentation classification head outputs the recognition result of the image partial segmentation;

[0105] The partial segmentation classification head is provided with a fourth fully connected layer, a third batch normalization layer, a third dropout layer, and a fifth fully connected layer connected in sequence. The output end of the third batch normalization layer is built-in with a ReLU activation function.

[0106] Next, a comprehensive performance evaluation of the local semantic and cross-level dependency network LCHNet for point clouds will be carried out on multiple datasets. In the test part of each dataset, detailed information on network configuration will be supplemented, including but not limited to key factors such as the depth of the network and parameter settings, to ensure the integrity and reproducibility of the evaluation.

[0107] In this embodiment, the local semantic and cross-level dependency network LCHNet for point clouds will be comprehensively evaluated on 3D point cloud object shape classification and per-point part segmentation tasks to verify its performance and effectiveness, and three datasets are used: ScanObjectNN, ModelNet40, and ShapeNet-Part. The reason for choosing these two tasks for evaluation is that in the research on downstream tasks of point cloud analysis, classifying object shapes and segmenting parts represent two extremes. The classification task requires the network to learn and master the overall context of the point cloud and capture its overall features. In contrast, the part segmentation task requires the network to learn the local information of each point for fine-grained depiction. For the classification task, the number of input points for each instance is set to 1024, and for the segmentation task, the number of points is set to 2048.

[0108] The backbone network of the point cloud-oriented local semantics and cross-level dependency network LCHNet consists of four carefully designed downsampling feature extraction stages. Each stage contains two key components: the scaling residual block SRB and the post-MLP module Pos-MLP. Through a series of experiments, the optimal configuration of each stage is determined: except for the third stage which includes two scaling residual blocks SRB and two post-MLP modules Pos-MLP, the remaining stages include one scaling residual block SRB and one post-MLP module Pos-MLP, and the position of the scaling residual block SRB in each stage is before the post-MLP module Pos-MLP. Finally, the results of the mediation attention module are fused with the original features in a certain proportion and adjusted according to the dataset.

[0109] First, the LCHNet network was evaluated on the ScanObjectNN dataset. Released in 2019, ScanObjectNN is a point cloud dataset containing 15,000 instances of 15 categories. Among them, 2,902 instances were obtained through real-world scanning, which poses a severe challenge to existing point cloud understanding technologies due to factors such as background interference, noise, and occlusion. In the embodiment, a variant of the ScanObjectNN dataset, called PB_T50_RS, was used. This variant performs geometric transformations on the instances, such as moving, rotating, and resizing, further increasing the complexity of the dataset and is considered a highly challenging version of the point cloud classification task. To verify the effectiveness of the network, extensive evaluations were carried out on the PB_T50_RS dataset, and the proportion between the original features and the features processed by the intermediate attention mechanism was adjusted. Specifically, different proportion configurations were tested, such as 1:1, 3:1, 1:3, 4:1, and 1:4. The results show that when the proportion is 4:1, the network performance is optimal, and the classification accuracy reaches 85.7%. After multiple tests, the average accuracy can be stably maintained at 85.6%. As shown in Table 1, LCHNet has achieved significant improvements in both the average class accuracy mAcc and the overall accuracy OA, exceeding most existing methods. In this pioneering work of deploying the mediation attention mechanism in the field of point cloud analysis, although there is an admitted gap with the cutting-edge algorithms, the results demonstrate its promising capabilities.

[0110] Table 1

[0111] Classification results on the ScanObjectNN dataset.

[0112]

[0113] Then, the LCHNet network was evaluated on the ModelNet40 dataset. This dataset contains 9,843 training instances and 2,468 test instances, including 40 classes of mesh CAD models. Following the common evaluation criteria, the average class accuracy mAcc and the overall accuracy OA on the test set were reported. In the experiment, the batch size of the training input was set to 24, and all models were trained with a comprehensive training scheme of 280 epochs. A detailed test was conducted on the connection rate between the features processed by the mediation attention mechanism and the unprocessed features. Specifically, different feature connection ratios were tried in this embodiment, including 1:1, 1:3, and 3:1. In multiple experiments, the splicing ratio of 1:1 showed the highest average accuracy rate, reaching 93.6%. The correct splicing rate of 1:3 reached 93.22%, while the correct splicing rate of 3:1 did not exceed that of the equal-ratio splicing. Based on these results, it can be determined that on the ModelNet40 dataset, the features processed by the mediation attention module should be connected to the unprocessed features at a ratio of 1:1.

[0114] Table 2 summarizes the experimental results of various methods on ModelNet40. In these comparisons, the LCHNet network is superior to some existing networks in some aspects, but it can be observed that there is still a gap in accuracy compared with the existing networks. This paper believes that the fundamental reason for the accuracy difference may be related to the characteristics of the ModelNet40 dataset. The ModelNet40 dataset consists of sampled points of computer-aided design (CAD) models, which have relatively simple scene simulations and limited available training samples. These factors may limit the network's ability to learn more complex scenes, thereby affecting the aggregation ability of the model.

[0115] Table 2

[0116] Classification Results on ModelNet40

[0117]

[0118] Finally, the part segmentation performance of the LCHNet model on ShapeNet-Part was evaluated. ShapeNet-Part is a comprehensive dataset containing 16,881 pre-aligned 3D instances, divided into 16 shape categories and further subdivided into 50 unique part categories. By evaluating on this comprehensive and diverse dataset, the effectiveness and robustness of the network in handling complex shape part segmentation tasks in the field of point cloud deep learning were fully verified. According to the settings of Qi et al., 2048 points were randomly selected from each object as the network input. In Table 3, the LCHNet network was compared with several recent works, including the Kd network, the self-organizing network SO-Net, and Point-PlaneNet, and the results showed that the LCHNet network achieved formidable results. In addition, Figure 7 The ground truth and prediction results in the part segmentation task are visualized. Intuitively, the prediction results of the LCHNet network are very close to the ground truth, intuitively illustrating the rationality and effectiveness of the LCHNet network design.

[0119] Table 3

[0120] Part segmentation results on the ShapeNet-Part dataset

[0121]

[0122] Next, to verify the contribution of the proposed mediation attention mechanism to the model performance, ablation experiments were conducted on the ScanObjectNN dataset. This experiment not only verified the implementation details of the input of the mediation attention module but also verified the effectiveness of the LCHNet model. The LCHNet consists of 4 basic feature extraction stages. To explore the impact of the mapping relationship between the outputs of different stages and the quadruple parameters (Q, I, K, V) on the model performance, three different experimental attempts were made. The results are shown in Table 4. When the mapping of the stage to the parameters is set to (Q, I, K, V), the model has the best training effect, with an OA of 85.26%, significantly better than other configuration schemes, which are 84.52%, 84.34%, and 81.92% respectively.

[0123] Table 4

[0124] Ablation experiment results of the mediation attention mechanism on ScanObjectNN

[0125]

[0126] To enhance the performance of the enhanced intermediate attention module, a scaled residual block (SRB) is introduced, aiming to reduce information loss during processing and effectively control the overfitting phenomenon. Therefore, this study conducts component combination exploration experiments to investigate how to maximize the benefits using the SRB. In the experiments, the Pre-MLPBlock and Pos-MLPBlock in the original network architecture are respectively replaced with the SRB to seek the optimal component combination. While making this change, other configurations of the network remain unchanged to ensure the comparability of experimental results. Table 5 details the experimental results on the ScanObjectNN dataset, where the model is optimal when only the Pre-MLP module in LCHNet is replaced with the SRB. Compared with the original network without the SRB, the introduction of the SRB improves the model performance to a certain extent. Specifically, in terms of the key metric of overall accuracy (OA), the model with the SRB is improved by 1.492 compared to the model without the SRB.

[0127] Table 5

[0128] Results of component ablation experiments on ScanObjectNN (crosses indicate the absence of the component, and blanks indicate its presence)

[0129]

[0130] It has been previously confirmed that the combination of the SRB and Pos-MLP block for feature extraction is effective and demonstrates its superiority in performance. However, the adjustment of the network depth, i.e., the number of organizational components, also plays a decisive role in the overall performance of the network model. To gain an in-depth understanding of the specific impact of changing the number of network layers on the model performance, a series of ablation experiments are conducted on the ScanObjectNN dataset for the number configurations of the SRB and Pos-MLP modules. The experimental results are shown in Table 6 in detail. In the initial experiment, referring to the settings of PointMLP, the number of SRBs in each contact feature extraction stage is set to 2, represented by the array [2, 2, 2, 2]. At the same time, to reduce the loss of original information, the list of the number of Pos-MLP blocks is set to

[0131] [1, 1, 1, 1]. Under this configuration, after two independent trainings, the results are 84.39% and 83.91% respectively, and the average overall accuracy OA reaches 84.15%. Then, the number of SRBs is reduced to [1, 1, 1, 1]. Under this configuration, the average OA of the trained model reaches 83.91%, and the specific values are 84.11% and 83.72% respectively, which is lower than the initial configuration. Therefore, the number of Pos-MLP blocks is increased to [2, 2, 2, 2], hoping to improve the performance through more complex feature extraction. However, the experimental results show that the model is not improved under this setting, which indicates that in the current network structure, too many Pos-MLP blocks may lead to performance saturation and their contributions become insignificant. According to previous experience, the numbers of SRBs and Pos-MLP blocks are readjusted and their configuration is changed to [1, 1, 2, 1]. After two trainings, this adjustment increases the average OA of the model to 85.58%, showing the positive impact of configuration optimization on performance. In addition, several other configurations such as [1, 1, 1, 2] and

[0132] [1, 2, 1, 1] are tested, but the experimental results of these configurations do not reach the effect of the [1, 1, 2, 1] configuration. Due to the limitation of computing resources, more in-depth parameter fine-tuning cannot be carried out. Nevertheless, there is reason to believe that with the further optimization of the network and the adjustment of parameters, the performance of the model is expected to be further improved.

[0133] Table 6

[0134] Module-level configuration and corresponding results

[0135]

[0136] To evaluate the robustness of the proposed deep learning network model in processing point cloud data, a series of tests are carried out by gradually reducing the number of points in the input point cloud to check the performance of the model under different data sparsities. In the initial test, a complete point cloud dataset is used as a benchmark, and each input contains 1024 points. Subsequently, three decreasing tests are carried out, each reducing the number of points by 25%, corresponding to 768 points, 576 points, and 432 points respectively. After each reduction, the performance of the model on the ScanObjectNN dataset is re-evaluated. Figure 8The results of these robustness tests are presented, where the classification accuracy of the model corresponds to the decreasing point density. The experimental results show that the network model can maintain a high accuracy and stability even when the number of points is significantly reduced. Specifically, in the first test, the model achieved an accuracy of 85.46% in the classification task on the ScanObjectNN dataset. As the number of points decreased, the performance decreased significantly; however, in the third test, as the number of points decreased to 432, the accuracy of the model in the classification task on the dataset still remained at 83.03%. The results are as Figure 8 shown. These findings indicate that the proposed network model shows good adaptability to the sparsity of point cloud data and can provide reliable classification results even when the number of data points is significantly reduced. This finding has important implications for deploying the model in practical applications with resource constraints or incomplete data.

[0137] The present invention revisits the existing 3D point cloud models and proposes the LCHNet network, a cross-level semantic dependency network for point cloud analysis. Adding SRB effectively expands the model, reduces information loss while deepening the network hierarchy, and prevents the phenomenon of gradient disappearance. In addition, the proposed mediating attention mechanism also has a powerful feature capture ability without stacking, which can achieve cross-level information interaction and long-distance semantic capture. LCHNet has shown excellent performance in point cloud classification and part segmentation. The experiments have demonstrated the effectiveness and importance of the method.

[0138] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image recognition method based on local semantics and cross-level dependencies for point clouds, characterized in that: The following steps are involved: Step 1: Construct an image recognition system based on local semantics and cross-level dependencies for point clouds, which includes an image acquisition module, a preprocessing module, and a local semantics and cross-level dependency network LCHNet for point clouds connected in sequence; The local semantic and cross-level dependency network LCHNet for point cloud is provided with an embedding layer, a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, an addition unit and an image recognition module connected in sequence from beginning to end, wherein the first feature extraction module, the second feature extraction module, the third feature extraction module, the fourth feature extraction module and the addition unit are also connected with an intermediate attention module; Step 2: The image acquisition module acquires the original image data a in real time and transmits it to the preprocessing module; Step 3: The preprocessing module performs a preprocessing operation on the original image data a to obtain standard data b, and passes it to the point cloud-oriented local semantic and cross-level dependency network LCHNet; Step 4: The embedding layer in the point cloud-oriented local semantics and cross-level dependency network LCHNet obtains the standard data b, and enhances the feature dimension of the standard data b to obtain the embedded data c, which is then passed to the first feature extraction module; Step 5: The first feature extraction module performs a feature extraction operation on the embedded data c to obtain first feature data d, and passes it to the second feature module and the intermediate attention module; Step 6: The second feature extraction module performs a feature extraction operation on the first feature data d to obtain second feature data e, and passes it to the third feature module and the intermediate attention module; Step 7: The third feature extraction module performs a feature extraction operation on the second feature data e to obtain third feature data f, and passes it to the fourth feature module and the intermediate attention module; Step 8: The fourth feature extraction module performs a feature extraction operation on the third feature data f to obtain fourth feature data g, and passes it to the addition unit and the intermediate attention module; Step 9: The intermediate attention module performs intermediate attention mechanism calculation on the first feature data d, the second feature data e, the third feature data f and the fourth feature data g, captures the long-range semantic information in each feature data, obtains multi-level feature data h, and passes it to the addition unit; Step 10: the adding unit adds the fourth feature data g and the multi-level feature data h to obtain comprehensive feature data k, and transmits it to the image recognition module; Step 11: The image recognition module performs an image recognition operation on the comprehensive feature data k and outputs an image recognition result.

2. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 1, characterized in that: In the step 1, the first feature extraction module, the second feature extraction module and the fourth feature extraction module have the same structure, and are all provided with a first downsampling layer, a first scaling residual block, a first aggregation layer and a first post-MLP module connected in sequence; The third feature extraction module is provided with a second downsampling layer, a second scaling residual block, a third scaling residual block, a second aggregation layer, a second post-MLP module and a third post-module which are connected in sequence.

3. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 2, characterized in that: The first scaled residual block, the second scaled residual block and the third scaled residual block have the same structure, and are all provided with a first multi-layer perceptron layer, a second multi-layer perceptron layer, a batch normalization layer, a third multi-layer perceptron layer and a residual connection layer connected in sequence, and a ReLU activation function is built in the output end of the residual connection layer.

4. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 3, characterized in that: The output expression of the scaled residual block is as follows: in, ReLU activation function, F i represents the i-th input data of the scaled residual block, φ1 represents the first multi-layer perceptron layer, φ2 represents the second multi-layer perceptron layer, φ3 represents the third multi-layer perceptron layer, Υ represents the batch normalization layer, Represents the output data of the scaled residual block.

5. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 1, characterized in that: In step 9, the intermediate attention module is provided with a first multiplication unit, a second multiplication unit, a third multiplication unit and a fourth multiplication unit, and the output ends of the first multiplication unit and the second multiplication unit are both built-in with a Softmax activation function; The process of obtaining the multi-level feature data h by the intermediate attention module according to the first feature data d, the second feature data e, the third feature data f and the fourth feature data g is as follows: S1: The first multiplication unit obtains the first feature data d and the second feature data e, multiplies the two, and then outputs the first multiplication data x1 to the fourth multiplication unit through the Softmax activation function; S2: the second multiplication unit obtains the first feature data d and the third feature data f, multiplies the two, and then outputs the second multiplication data x2 to the third multiplication unit through the Softmax activation function; S3: the third multiplication unit obtains the second multiplication data x2 and the fourth feature data g, multiplies the two, and then outputs the third multiplication data x3 to the fourth multiplication unit; S4: The fourth multiplication unit obtains the first multiplication data x1 and the third multiplication data x3, multiplies the two, and then outputs the multi-level feature data h to the addition unit.

6. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 5, characterized in that: The output expression of the intermediate attention module is as follows: O=At s (Q,I,At φ (I,K,V))=σ(QI T )σ(IK T )V Wherein, O represents multi-level feature data h, Q represents the first feature data d, I represents the second feature data e, K represents the third feature data f, V represents the fourth feature data g, At S represents the softmax attention mechanism algorithm, Atφ represents the linear attention mechanism algorithm, σ represents the softmax activation function, and the superscript T represents the matrix transpose.

7. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 1, characterized in that: In step 3, the preprocessing operations include but are not limited to moving, rotating, adjusting image size, removing image background noise and normalizing.

8. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 1, characterized in that: In the step 11, the image recognition module is provided with an image classification module and an image part segmentation module, the image classification module is used to perform an image classification task, and the image part segmentation module is used to perform an image part segmentation task; The image recognition result may be an image classification recognition result or an image partial segmentation recognition result.

9. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 8, characterized in that: The image classification module is provided with a first fully connected layer, a first batch of normalized layers, a first discard layer, a second fully connected layer, a second batch of normalized layers, a second discard layer, a third fully connected layer and a classification output layer which are connected in sequence, and the output ends of the first batch of normalized layers and the second batch of normalized layers are both built-in with a ReLU activation function; The third fully connected layer outputs probability distribution values ​​of various image categories to the classification output layer, and the classification output layer outputs the image category corresponding to the maximum value of the probability distribution as the image classification recognition result.

10. The image recognition method based on point cloud-oriented local semantics and cross-level dependencies according to claim 8, characterized in that: The image part segmentation module is provided with a first upsampling layer, a second upsampling layer, a third upsampling layer, a fourth upsampling layer, a concat splicing layer, an information processing unit and a part segmentation classification head; The input end of the first upsampling layer is connected to the output ends of the third feature extraction module and the fourth feature extraction module, the output end of the first upsampling layer is connected to the input end of the second upsampling layer, the input end of the second upsampling layer is also connected to the output end of the second feature extraction module, the output end of the second upsampling layer is connected to the input end of the third upsampling layer, the input end of the third upsampling layer is also connected to the output end of the first feature extraction module, the output end of the third upsampling layer is connected to the input end of the fourth upsampling layer, the input end of the fourth upsampling layer is also connected to the output end of the embedding layer, the output end of the fourth upsampling layer is connected to the input end of the concat splicing layer, the input end of the concat splicing layer is also connected to the output end of the addition unit, the output end of the concat splicing layer is connected to the input end of the partial segmentation classification head via the information processing unit, and the partial segmentation classification head outputs the partial segmentation recognition result of the image; The partial segmentation classification head is provided with a fourth fully connected layer, a third batch normalization layer, a third discard layer and a fifth fully connected layer which are connected in sequence, and a ReLU activation function is built in the output end of the third batch normalization layer.

Citation Information

Patent Citations

  • Target processing method and device based on cross-level, cross-scale and cross-attention mechanism

    CN113177555A

  • Image classification identification method and device based on adaptive dynamic convolutional network, and computer equipment

    CN114445664A

  • Image recognition method and device based on cross-layer feature mining and electronic equipment

    CN117911755A

  • Point cloud semantic segmentation method based on global feature enhancement

    CN118247511A

  • Ship detection method and system combining spatial explicit vision and improved big nuclear attention

    CN118968017A