Deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes

By combining SHAP attribution analysis and deep-sea prior knowledge modeling network modulation methods, the problem of insufficient accuracy of point cloud semantic understanding in deep-sea environments is solved, and a more efficient and robust semantic segmentation effect is achieved.

CN119538073BActive Publication Date: 2025-05-13ZHEJIANG LAB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510099896.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

In deep-sea environments, traditional point cloud semantic understanding methods are difficult to effectively process sparse and incomplete sonar point cloud data, resulting in insufficient accuracy of semantic understanding results.

Method used

Using SHAP-based attribution analysis, deep-sea prior knowledge modeling and network modulation methods, the intermediate features of the deep learning network are dynamically adjusted through feature interpreters and formal modeling representations to improve the accuracy of semantic classification results.

Benefits of technology

It improves the performance of deep-sea point cloud data in semantic segmentation, reduces the model's illusion output, realizes more robust point cloud semantic understanding, reduces the need for large-scale training data, and thus reduces computational cost and time overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538073B_ABST
    Figure CN119538073B_ABST
Patent Text Reader

Abstract

The present invention discloses a deep learning network modulation method for semantic understanding of sonar point clouds for deep-sea scenes, the method comprising: performing feature statistical analysis on a pre-trained semantic understanding network to determine feature vectors of different target categories; and determining key feature components based on SHAP attribution analysis to establish a feature interpreter; at the same time, formalizing human prior knowledge of semantic understanding of deep-sea point clouds into logical expressions, and performing modulation and correction of target intermediate features in combination with the pre-classification results of the feature interpreter during the reasoning process to achieve better semantic classification results. The present invention realizes knowledge-enhanced deep learning through dual-drive of data and knowledge, which helps to improve the adaptability and accuracy of the sonar point cloud semantic understanding network in complex deep-sea environments, optimize the semantic classification results of the model, and avoid unreasonable hallucination outputs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep sea environment perception technology, and in particular to a sonar point cloud semantic understanding deep learning network modulation method for deep sea scenes. Background Art

[0002] The deep-sea environment is complex and highly complex and uncertain. Traditional point cloud semantic understanding methods based on sonar, lidar or visual sensors face many challenges in practical applications. Especially in deep-sea scenes, due to the scarcity of light, unclear texture and the sparsity of sonar point cloud data, conventional semantic segmentation methods often fail to achieve ideal results. Therefore, how to achieve accurate point cloud semantic understanding in deep-sea environments has become a key issue in the current research of unmanned submersibles and underwater robot systems.

[0003] Existing deep-sea point cloud semantic understanding and environmental perception methods mainly rely on traditional feature processing technology. Patent document CN109035224A discloses a submarine pipeline detection and 3D reconstruction method based on multi-beam point cloud, which includes the following steps: first, based on the underwater sonar image obtained by multi-beam bathymetric sonar detection of the pipeline, the image pixels are classified and extracted using the threshold method and the Canny edge detection method to obtain the 3D point cloud data of the pipeline; then, a point cloud denoising filtering method based on density analysis is used to set different initial radii R and minimum number of neighbors k to obtain the 3D point cloud data of the pipeline after filtering and denoising; then, a histogram-based statistical method and a spatial linear fitting method are used to perform circle fitting on the point cloud data of each section of the pipeline to obtain the radius of the fitting circle and the linearly changing center point of the circle; finally, the AlphaShape algorithm is used to reconstruct the pipeline in 3D based on the radius of the fitting circle and the linearly changing center point of the circle. This method directly extracts point cloud data from sonar images, has low computational complexity, and is suitable for the detection and three-dimensional reconstruction of various underwater pipelines. However, when processing pipelines with a smaller radius, this method has less point cloud data, which may affect the accuracy of reconstruction. Patent document CN110517193A discloses a method for processing submarine sonar point cloud data, which specifically includes the following steps: first, the original point cloud data is fused to form four-dimensional point cloud data; then, a filtering algorithm based on the average distance of neighboring points is used to remove high-frequency noise; then, the point cloud normal vector is corrected for the point cloud after high-frequency denoising, and the normal vector is smoothed using a Gaussian weight function; then, the sonar point cloud is low-frequency denoised, and the low-frequency noise is filtered using an anisotropic smoothing denoising algorithm based on bilateral filtering; then, the denoised point cloud is globally simplified, kd-tree is used to search the k-neighborhood of the point, the point cloud area is segmented based on the region growing method, and different point cloud simplification methods are applied to flat areas and non-flat areas; finally, the four-dimensional submarine sonar point cloud holes are repaired using a point cloud hole repair algorithm based on local expansion of concentric circles. The point cloud data processed by this method facilitates the filtering algorithm to identify and detect artificial targets in the sonar point cloud data, which not only saves the time consumed by manual interpretation of sonar image data, but also improves the accuracy of the artificial target recognition algorithm. However, when processing large-scale point cloud data, this method has a large amount of calculation, which may affect the processing efficiency.

[0004] With the continuous maturity of artificial intelligence technology, point cloud semantic understanding technology based on deep learning has been developed, such as convolutional neural networks (CNN) and recurrent neural networks (RNN) and Tranformer architecture based on point cloud data. These methods usually extract target features in deep-sea scenes and perform target semantic segmentation by directly training point cloud data. However, these methods face two major problems: first, the lack of training data for deep-sea scenes leads to unstable performance of the model in complex environments; second, in deep-sea environments, the sparsity and incompleteness of point cloud data make it more difficult to train deep learning models, resulting in insufficient accuracy of semantic understanding results.

[0005] In order to overcome these challenges, knowledge-enhanced deep learning methods based on network modulation have gradually become a research hotspot in recent years. By combining the prior knowledge of domain experts with deep learning models, the performance of the model in specific tasks can be effectively improved, especially when data is scarce and samples are unbalanced. However, most of the current knowledge enhancement methods focus on optimizing the model through simple knowledge injection, and lack an effective mechanism to integrate prior knowledge and point cloud data in deep-sea environments. This makes it impossible for traditional methods to fully utilize the advantages of prior knowledge when facing complex deep-sea point cloud data, resulting in reduced model detection accuracy and even unconventional hallucination outputs. Therefore, how to fuse deep-sea specific prior knowledge and point cloud data through innovative network modulation methods to achieve efficient semantic segmentation has become one of the current research focuses in the field of deep-sea point cloud semantic understanding. Summary of the invention

[0006] The purpose of the present invention is to provide a deep learning network modulation method for sonar point cloud semantic understanding for deep-sea scenes in view of the deficiencies of the prior art. The present invention can effectively improve the performance of deep-sea point cloud data in semantic segmentation, avoid the model's hallucination output, and achieve more robust point cloud semantic understanding by combining SHAP-based attribution analysis, deep-sea prior knowledge modeling, and network modulation.

[0007] The objective of the present invention is achieved through the following technical solutions: In a first aspect, an embodiment of the present invention provides a sonar point cloud semantic understanding deep learning network modulation method for deep-sea scenes, comprising:

[0008] Based on the semantic understanding network that has been pre-trained on the point cloud dataset, the classification results are statistically analyzed to determine the feature vectors of different target categories;

[0009] The SHAP attribution analysis method is used to perform attribution analysis on the classifiers in the semantic understanding network, and the key feature components of different target categories are determined by setting a floating SHAP importance threshold.

[0010] According to the feature vectors and key feature components of different target categories, a feature interpreter is established for pre-classification;

[0011] Formal modeling of prior knowledge on semantic understanding of deep-sea point clouds to convert it into logical expressions;

[0012] During the forward reasoning process of the semantic understanding network model, pre-classification is performed through the feature interpreter. At the same time, the pre-classification result is corrected by combining the logical expression corresponding to the prior knowledge of semantic understanding of deep-sea point clouds represented by the above-mentioned formal modeling. On this basis, network modulation is performed to achieve better semantic classification results.

[0013] Furthermore, the feature vectors of different target categories are expressed as:

[0014]

[0015] in, Represents the feature vector of target category i, i represents the classification result of the i-th category, represents the classifier in the semantic understanding network, Indicates the calculation The nth feature classified as category i, N represents the number of features used to calculate The total number of features classified into category i.

[0016] Furthermore, the SHAP attribution analysis method is used to perform attribution analysis on the classifier in the semantic understanding network, specifically including:

[0017] The SHAP attribution analysis method is used to perform attribution analysis on the classifier in the semantic understanding network to obtain the SHAP value of each component of the intermediate feature vector of different target categories, which is expressed as:

[0018]

[0019] in, represents the intermediate eigenvector The kth component in SHAP value of represents an intermediate feature vector of the target classification result i, and K represents the intermediate feature vector The total number of components in is a non-existent A subset of features, Indicates discharge from F , In the feature subset and The results of training the classifier and making predictions on the joint dataset of In the feature subset The results of training the classifier and making predictions.

[0020] Furthermore, the key feature components of different target categories are determined by setting a floating SHAP importance threshold, specifically including:

[0021] Based on the SHAP values ​​of each component of the intermediate feature vector of different target categories and the set floating SHAP importance threshold, the following conditions are used to determine the key feature components of different target categories:

[0022]

[0023] in, Represents the feature vector The kth eigencomponent in , represents the floating SHAP importance threshold, which is determined by the following formula:

[0024]

[0025] in, is the sensitivity threshold.

[0026] Furthermore, the feature interpreter is expressed as:

[0027]

[0028] in, represents a feature interpreter, represents the intermediate feature, I represents the total number of target categories, Represents intermediate features The feature vector with target category i The difference between them is calculated as:

[0029]

[0030] in, Indicates Key feature components, The feature vector representing the target category i The components, and G represents the total number of key feature components.

[0031] Furthermore, the pre-classification is performed by the feature interpreter, and the pre-classification result is corrected by combining the logical expression corresponding to the prior knowledge of the semantic understanding of deep-sea point clouds represented by the above-mentioned formal modeling, including the judgment criteria for whether a single target needs to be corrected and modulated, which is expressed as:

[0032]

[0033] in, represents the classification category given by the logical expression converted from the prior knowledge of the semantic understanding of deep-sea point clouds, and P represents the original attribute features output by the sensor corresponding to the target; that is, when the pre-classification result of the feature interpreter is different from the classification result given by the logical expression corresponding to the prior knowledge of the semantic understanding of deep-sea point clouds represented by formal modeling, the pre-classification result needs to be corrected.

[0034] Furthermore, the performing network modulation specifically includes:

[0035] The network modulation of the intermediate features of the target is expressed as:

[0036]

[0037] Among them, i is the ideal category of modulation, that is, the classification category given by the logical expression corresponding to the prior knowledge; For the intermediate features The amount of modulation added is expressed as = - .

[0038] A second aspect of an embodiment of the present invention provides a sonar point cloud semantic understanding deep learning network modulation device for deep-sea scenes, which is used to implement the above-mentioned sonar point cloud semantic understanding deep learning network modulation method for deep-sea scenes, and the device includes:

[0039] The SHAP analysis unit is used to perform attribution analysis on the classifiers in the semantic understanding network, and calculate the SHAP value of the feature component of each intermediate feature by using the SHAP attribution analysis method to identify the key feature components of different target categories;

[0040] The knowledge modeling representation unit is used to formally model and represent the prior knowledge of the deep-sea point cloud scene by symbolization, regularization or graphing, so as to convert the prior knowledge into a logical expression;

[0041] A knowledge mapping unit, which is used to pre-classify the intermediate features and verify the pre-classification results using formalized prior knowledge, and to assist the semantic understanding network in dynamically adjusting the intermediate features of the semantic understanding network during the model reasoning process; and

[0042] The network modulation unit is used to modify the intermediate features of the semantic understanding network in combination with the output of the knowledge mapping unit during the model reasoning process, so that the semantic understanding network can integrate information from prior knowledge when processing deep-sea point cloud data and optimize the recognition of target categories.

[0043] A third aspect of an embodiment of the present invention provides an electronic device, comprising one or more processors and a memory, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned sonar point cloud semantic understanding deep learning network modulation method for deep-sea scenes.

[0044] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, is used to implement the above-mentioned sonar point cloud semantic understanding deep learning network modulation method for deep-sea scenes.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] (1) By combining SHAP attribution analysis with formal modeling of prior knowledge, the present invention makes it possible to identify and utilize key feature components and modulate the model during the reasoning process; this helps to improve the model's adaptability to complex deep-sea environments, reduce misjudgments caused by data sparsity or noise, and thus improve the accuracy and robustness of semantic segmentation.

[0047] (2) The present invention adopts a data and knowledge dual-driven strategy, extracting the key features of the model through SHAP analysis, and dynamically correcting the model by using the deep-sea prior knowledge of domain experts; through knowledge enhancement, the model can obtain better generalization capabilities in an environment with scarce data and unbalanced samples.

[0048] (3) The present invention formalizes humans’ prior knowledge about the semantic understanding of deep-sea point clouds and integrates it into the model, allowing the model to refer to these empirical rules when dealing with unknown or edge cases, thereby making judgments that are more in line with the actual situation; this method not only improves the generalization ability of the model, but also avoids the model’s hallucination output.

[0049] (4) The present invention combines pre-classification with correction based on prior knowledge to reduce the demand for large-scale training data while ensuring or even improving the quality of semantic segmentation, thereby reducing computing costs and time overhead. In addition, by analyzing the importance of features, we can selectively focus on those features that have a greater impact on the prediction results, further optimizing computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flow chart of the deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes of the present invention;

[0051] Figure 2 It is a SHAP attribution analysis result diagram of each component of the three categories of intermediate feature vectors of the present invention;

[0052] Figure 3 is a distribution diagram of three categories of key characteristic components of the present invention;

[0053] Figure 4 It is a schematic diagram of the network modulation principle of the present invention;

[0054] Figure 5 It is a structural schematic diagram of a sonar point cloud semantic understanding deep learning network modulation device for deep-sea scenes of the present invention;

[0055] Figure 6 It is a structural schematic diagram of an electronic device of the present invention. DETAILED DESCRIPTION

[0056] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0057] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0058] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0059] The present invention is described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the features of the following embodiments and implementations can be combined with each other.

[0060] The deep learning network modulation method for sonar point cloud semantic understanding for deep-sea scenes of the present invention realizes knowledge enhancement of deep-sea sonar point cloud semantic understanding network by combining SHAP (SHapley Additive exPlanations) attribution analysis and deep learning network modulation. The method adopts a data and knowledge dual-driven strategy. First, the SHAP attribution method is used to perform feature statistics and attribution analysis on the classification results of the point cloud semantic understanding network, determine the key feature components and construct a feature interpreter; at the same time, the prior knowledge of the deep-sea scene is converted into a decidable logical expression through formal modeling representation; the prior knowledge represented by the formal modeling is combined with the pre-classification results generated by the feature interpreter, and the intermediate features are dynamically modulated and corrected during the reasoning process, thereby achieving a more accurate result of the point cloud semantic understanding and avoiding the hallucination output of the model.

[0061] See also Figure 1 The deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes of the present invention specifically includes the following steps:

[0062] S110. Based on the semantic understanding network that has been pre-trained on the point cloud dataset, feature statistical analysis is performed on the classification results to determine feature vectors of different target categories.

[0063] It should be understood that the semantic understanding network is already pre-trained, and its input is a frame of original point cloud data, and it outputs the category information of each point. The semantic understanding network mainly includes a point cloud encoding module, a multi-scale feature extraction module, a feature fusion module and a classifier module.

[0064] In this embodiment, the point cloud in an unstructured seabed scene is divided into a traversable area category (category i=1), an obstacle category (category i=2), and a neutral area category between the two (category i=3), so as to provide obstacle avoidance and navigation capabilities for the seabed unmanned platform. Therefore, the feature vectors of different target categories are expressed as:

[0065]

[0066] in, Represents the feature vector of target category i, i represents the classification result of the i-th category, represents the classifier in the semantic understanding network, Indicates the calculation The nth feature classified as category i, N represents the number of features used to calculate The total number of features classified as category i, represents an intermediate feature vector of the target classification result i, and K represents the intermediate feature vector In this embodiment, the N values ​​corresponding to categories 1, 2, and 3 are 716, 214, and 69 respectively. The total number of components K=1536.

[0067] S120, using the SHAP attribution analysis method to perform attribution analysis on the classifier in the semantic understanding network in step S110, and determining the key feature components of different target categories by setting a floating SHAP importance threshold.

[0068] In this embodiment, the classifier in the semantic understanding network is a multi-layer perceptron module. The SHAP attribution analysis method is used to perform attribution analysis on the classifier in the semantic understanding network to obtain the SHAP value of each component of the intermediate feature vector of different target categories, which is expressed as:

[0069]

[0070] in, represents the intermediate eigenvector The kth component in SHAP value of represents an intermediate feature vector of the target classification result i, and K represents the intermediate feature vector The total number of components in is a non-existent A subset of features, Indicates discharge from F , In the feature subset and The results of training the classifier and making predictions on the joint dataset of In the feature subset In this embodiment, the SHAP values ​​of the components of the intermediate feature vectors of the three categories are as follows: Figure 2 shown.

[0071] Then, based on the SHAP values ​​of each component of the intermediate feature vector of different target categories and the set floating SHAP importance threshold, the following conditions are used to determine the key feature components of different target categories:

[0072]

[0073] in, Represents the feature vector The kth eigencomponent in , is the feature vector of different target categories, is determined by the average value of multiple intermediate feature vectors F of this category, that is, an intermediate feature vector of target classification result i is , the average value of multiple intermediate eigenvectors F is , for the convenience of description, it is described as ; represents the floating SHAP importance threshold, which is determined by the following formula:

[0074]

[0075] in, is the sensitivity threshold. In this embodiment, , correspondingly, the key feature components of the three categories are as follows Figure 3 shown.

[0076] It should be noted that when determining the key feature components of different target categories, it is necessary to judge Is it greater than or equal to ,like Greater than or equal to , then As the key feature component of the target category; otherwise, It is not the key feature component of the target category. The key feature components of different target categories are expressed as .

[0077] S130: Establish a feature interpreter according to feature vectors and key feature components of different target categories for pre-classification.

[0078] Furthermore, the feature interpreter is expressed as:

[0079]

[0080] in, represents a feature interpreter, represents the intermediate feature, I represents the total number of target categories, Represents intermediate features The feature vector with target category i The difference between them is calculated as:

[0081]

[0082] in, Indicates Key feature components, The feature vector representing the target category i The components, G represents the total number of key feature components. Specifically, for the intermediate feature , calculate its feature vector with all target categories The difference between , the category corresponding to the feature with the smallest difference is regarded as The category corresponding to the target.

[0083] S140. Formalize and model the prior knowledge related to human semantic understanding of deep-sea point clouds to convert the prior knowledge into a decidable logical expression. Prior knowledge includes important information such as the geometric shape, physical characteristics, and behavior patterns of the target.

[0084] In this embodiment, the following two kinds of prior knowledge related to human semantic understanding of deep-sea point clouds are given, and formal modeling is performed to represent them, so as to convert the prior knowledge into a decidable logical expression. The first kind of prior knowledge is that targets with a height difference greater than a certain value should be classified as obstacles, and the corresponding logical expression is The second prior knowledge is that non-obstacle targets close to obstacles should be classified as neutral areas, and the corresponding logical expression is .in, It represents the classification category given by the determinate logical expression converted from the prior knowledge of human beings’ semantic understanding of deep-sea point clouds. P represents the original attribute features corresponding to the target output by the sensor, such as the three-dimensional coordinates of the target, etc. h represents the target height. represents the target height threshold, d represents the shortest distance between the target and other obstacles, Indicates the distance threshold between the target and other obstacles.

[0085] S150. In the forward reasoning process of the semantic understanding network model, pre-classification is performed through the feature interpreter. At the same time, the pre-classification result is corrected by combining the logical expression corresponding to the relevant prior knowledge of human beings on the semantic understanding of deep-sea point clouds represented by the above-mentioned formal modeling. On this basis, network modulation is performed to achieve better semantic classification results.

[0086] In this embodiment, the modulation principle of the deep learning network for point cloud semantic understanding of deep sea scenes is as follows: Figure 4 As shown in the figure, the first prior knowledge is applied to the target. Specifically, if the height of the target is greater than the threshold, the target can be judged as category 2 (obstacle). In layman's terms, higher objects should be regarded as obstacles. That is, according to the logical expression corresponding to the first prior knowledge, the target is judged to belong to the obstacle category, but the feature interpreter judges the target to be a non-obstacle category, that is:

[0087]

[0088] At this time, the pre-classification results need to be corrected, which is also the judgment criterion for whether a single target needs to be corrected and modulated. That is, when the pre-classification results of the feature interpreter are different from the classification results given by the logical expression corresponding to the prior knowledge of the semantic understanding of deep-sea point clouds represented by formal modeling, the pre-classification results need to be corrected.

[0089] Furthermore, network modulation is performed, specifically including: network modulation of the intermediate features of the target, which is expressed as:

[0090]

[0091] Wherein, i is the ideal category of modulation, that is, the classification category given by the logical expression corresponding to the prior knowledge. In this embodiment, the ideal category of modulation is the obstacle category, that is, i=2; For the intermediate features The amount of modulation added is expressed as = - .

[0092] It is worth mentioning that, based on the same inventive concept, the embodiment of the present invention also provides a sonar point cloud semantic understanding deep learning network modulation device 500 for deep-sea scenes, and its structure is as follows: Figure 5 The network modulation device 500 includes a SHAP analysis unit 510 , a knowledge modeling representation unit 520 , a knowledge mapping unit 530 and a network modulation unit 540 .

[0093] In this embodiment, the SHAP analysis unit 510 is used to perform attribution analysis on the classifier in the semantic understanding network, and calculates the SHAP value of the feature component of each intermediate feature as a contribution to the classifier prediction result by using the SHAP attribution analysis method to identify the key feature components of different target categories. This analysis can quantify the importance of each feature component in the model decision and identify the key feature components that have a greater impact on the classification decision.

[0094] In this embodiment, the knowledge modeling and representation unit 520 is used to formally model and represent the prior knowledge of the deep-sea point cloud scene in a symbolic, regularized or graphical manner, so as to convert the prior knowledge into a logical expression for decision reasoning.

[0095] In this embodiment, the knowledge mapping unit 530 is used to pre-classify the intermediate features and verify the pre-classification results using formalized prior knowledge, and assist the semantic understanding network to dynamically adjust the intermediate features of the semantic understanding network during the model reasoning process. This unit is a bridge between the knowledge modeling representation unit 520 and the network modulation unit 540.

[0096] In this embodiment, the network modulation unit 540 is used to modify the intermediate features of the semantic understanding network in combination with the output of the knowledge mapping unit 530 during the model reasoning process, so that the semantic understanding network can more effectively integrate information from prior knowledge when processing deep-sea point cloud data and optimize the recognition of target categories.

[0097] Corresponding to the aforementioned embodiment of the sonar point cloud semantic understanding deep learning network modulation method for deep-sea scenes, the present invention also provides an embodiment of an electronic device.

[0098] See also Figure 6 An electronic device provided by an embodiment of the present invention includes one or more processors and a memory, and the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the sonar point cloud semantic understanding deep learning network modulation method for deep-sea scenes in the above embodiment.

[0099] The embodiments of the electronic device of the present invention can be applied to any device with data processing capability, and the device with data processing capability can be a device or apparatus such as a computer. The electronic device embodiments can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as an electronic device in a logical sense, it is formed by the processor of any device with data processing capability in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, if Figure 6 As shown, it is a hardware structure diagram of any device with data processing capability where the electronic device of the present invention is located, except Figure 6 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the electronic device in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0100] The implementation process of the functions and effects of each unit in the above electronic device is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here.

[0101] For the electronic device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The electronic device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0102] An embodiment of the present invention also provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the deep learning network modulation method for sonar point cloud semantic understanding for deep-sea scenes in the above-mentioned embodiment is implemented.

[0103] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capability and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.

[0104] The above is only a preferred implementation case of the present invention and does not limit the present invention in any form. Although the implementation process of the present invention is described in detail above, for those familiar with the art, they can still modify the technical solutions recorded in the above examples, or replace some of the technical features therein with equivalents. All modifications, equivalent replacements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes, characterized in that: include: Based on the semantic understanding network that has been pre-trained on the point cloud dataset, the classification results are statistically analyzed to determine the feature vectors of different target categories; The SHAP attribution analysis method is used to perform attribution analysis on the classifiers in the semantic understanding network, and the key feature components of different target categories are determined by setting a floating SHAP importance threshold. According to the feature vectors and key feature components of different target categories, a feature interpreter is established for pre-classification; the feature interpreter is expressed as: Among them, E feature () represents the feature interpreter, F represents the intermediate feature, I represents the total number of target categories, D(F,Γ i ) represents the feature vector Γ of the intermediate feature F and the target category i i The difference between them is calculated as: Among them, f g represents the g-th key feature component, The feature vector Γ represents the target category i i The g-th component in , G represents the total number of key characteristic components; Formal modeling of prior knowledge on semantic understanding of deep-sea point clouds to convert it into logical expressions; During the forward reasoning process of the semantic understanding network model, pre-classification is performed through the feature interpreter. At the same time, the pre-classification result is corrected by combining the logical expression corresponding to the prior knowledge of semantic understanding of deep-sea point clouds represented by the above-mentioned formal modeling. On this basis, network modulation is performed to achieve better semantic classification results.

2. The deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes according to claim 1 is characterized in that: The feature vectors of different target categories are expressed as: Among them, Γ i represents the feature vector of target category i, i represents the classification result of the i-th category, M(·) represents the classifier in the semantic understanding network, and F n It is used to calculate Γ i The nth feature classified as category i, N represents the feature used to calculate Γ i The total number of features classified into category i.

3. The deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes according to claim 1 is characterized in that: The SHAP attribution analysis method is used to perform attribution analysis on the classifier in the semantic understanding network, specifically including: The SHAP attribution analysis method is used to perform attribution analysis on the classifier in the semantic understanding network to obtain the SHAP value of each component of the intermediate feature vector of different target categories, which is expressed as: Among them, Θ SHAP (f k ) represents the kth component f in the intermediate eigenvector F k SHAP value of F = [f1,f2,…,f k ,…,f K ] represents an intermediate feature vector of the target classification result i, K represents the total number of components in the intermediate feature vector F, and S is a vector that does not include f k The feature subset of F\{f k } means to discharge f from F k , M(S∪{f k }) is in the feature subset S and f k M(S) is the result of training a classifier on the joint dataset of and making predictions, and M(S) is the result of training a classifier on the feature subset S and making predictions.

4. The deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes according to claim 3 is characterized in that: The key feature components of different target categories are determined by setting the floating SHAP importance threshold, specifically including: Based on the SHAP values ​​of each component of the intermediate feature vector of different target categories and the set floating SHAP importance threshold, the following conditions are used to determine the key feature components of different target categories: I SHAP (f k )≥V shap Among them, f k represents the eigenvector Γ=[f1,f2,…,f k ,…,f K ], V shap represents the floating SHAP importance threshold, which is determined by the following formula: Among them, λ is the sensitivity threshold.

5. The deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes according to claim 1 is characterized in that: The pre-classification is performed by the feature interpreter, and the pre-classification result is corrected by combining the logical expression corresponding to the prior knowledge of the semantic understanding of the deep-sea point cloud represented by the above formal modeling, including the judgment criteria for whether a single target needs to be corrected and modulated, which is expressed as: E feature (F)≠E knowledge (P) Among them, E knowledge () represents the classification category given by the logical expression converted from the prior knowledge of the semantic understanding of deep-sea point clouds, and P represents the original attribute features output by the sensor corresponding to the target; that is, when the pre-classification result of the feature interpreter is different from the classification result given by the logical expression corresponding to the prior knowledge of the semantic understanding of deep-sea point clouds represented by formal modeling, the pre-classification result needs to be corrected.

6. The deep learning network modulation method for sonar point cloud semantic understanding for deep sea scenes according to claim 1 is characterized in that: The performing network modulation specifically includes: The network modulation of the intermediate features of the target is expressed as: T NM (F,i)=F+ΔF Where i is the ideal category of modulation, that is, the classification category given by the logical expression corresponding to the prior knowledge; ΔF is the modulation amount added to the intermediate feature F, expressed as ΔF = Γ i -F.

7. A sonar point cloud semantic understanding deep learning network modulation device for deep sea scenes, used to implement the sonar point cloud semantic understanding deep learning network modulation method for deep sea scenes described in any one of claims 1-6, characterized in that: The device comprises: The SHAP analysis unit is used to perform attribution analysis on the classifiers in the semantic understanding network, and calculate the SHAP value of the feature component of each intermediate feature by using the SHAP attribution analysis method to identify the key feature components of different target categories; The knowledge modeling representation unit is used to formally model and represent the prior knowledge of the deep-sea point cloud scene by symbolization, regularization or graphing, so as to convert the prior knowledge into a logical expression; A knowledge mapping unit, which is used to pre-classify the intermediate features and verify the pre-classification results using formalized prior knowledge, and to assist the semantic understanding network in dynamically adjusting the intermediate features of the semantic understanding network during the model reasoning process; and The network modulation unit is used to modify the intermediate features of the semantic understanding network in combination with the output of the knowledge mapping unit during the model reasoning process, so that the semantic understanding network can integrate information from prior knowledge when processing deep-sea point cloud data and optimize the recognition of target categories.

8. An electronic device comprising one or more processors and a memory, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the sonar point cloud semantic understanding deep learning network modulation method for deep-sea scenes as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that: A program is stored thereon, which, when executed by a processor, is used to implement the sonar point cloud semantic understanding deep learning network modulation method for deep-sea scenes as described in any one of claims 1-6.

Citation Information

Patent Citations

  • A subsea pipeline detection and three-dimensional reconstruction method based on multi-beam point cloud

    CN109035224A

  • Seabed sonar point cloud data processing method

    CN110517193A

  • Risk control method, device and equipment based on Shapril additive interpretation and medium

    CN115953248A

  • Slope deformation prediction and interpretation method based on fuzzy echo state network and SHAP

    CN118410903A