Semi-supervised three-dimensional semantic segmentation class distribution mismatched weight sensing distillation method

By obtaining the point mutual information between the labeled data and the unlabeled data, predicting and evaluating the pseudo-label weights of the unlabeled data, and performing knowledge distillation, the class distribution mismatch problem in semi-supervised three-dimensional semantic segmentation is solved, and the optimization quality and effect of the classifier are improved.

CN120388239AInactive Publication Date: 2025-07-29CHINA COAL RES INST +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510856561.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, semi-supervised three-dimensional semantic segmentation has class distribution mismatch problems, resulting in pseudo-label errors and intrusion errors of unknown categories, affecting classifier performance.

Method used

By obtaining the point mutual information between the labeled data and the unlabeled data, predicting the pseudo-labels of the unlabeled data and evaluating their weights, selectively performing knowledge distillation, and optimizing the classification model.

Benefits of technology

Effectively reduce the negative impact of pseudo-label errors and unknown categories, and improve the optimization quality and classification effect of classifiers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388239A_ABST
    Figure CN120388239A_ABST
Patent Text Reader

Abstract

The invention provides a semi-supervised three-dimensional semantic segmentation class distribution mismatched weight sensing distillation method, which comprises the following steps of: obtaining a first data pair containing first data and a first label of the first data and second data, and determining a point mutual information queue corresponding to any second data and all the first data, determining a second label of the second data and a weight of the second label according to the size of each point mutual information in the point mutual information queue and the corresponding first label, and calculating a loss value of the initial classification model based on the first data pair, the second data, the second label of the second data and the weight of the second label; the parameters of the initial classification model are updated, and optimization of the target classification model is completed. The self-adaptive weight and the high-quality pseudo tag of the target instance are captured by exploring the point mutual information in the representation space, so that the function of unmarked data can be played to the maximum extent, unknown categories can be filtered, and the efficiency and the quality of classification model optimization are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of semi-supervised learning, knowledge distillation, 3D semantic segmentation, and class distribution mismatch, and in particular, to a weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch. Background Art

[0002] 3D semantic segmentation involves various different types of objects or regions (such as buildings, roads, vegetation, etc.). Optimizing the classifier can enable it to more accurately learn the feature representations of various classes, so as to more precisely assign each 3D point or voxel to the correct class during the segmentation process, reduce misclassification phenomena, and significantly improve the overall segmentation accuracy. Summary of the Invention

[0003] The present application aims to at least partly solve one of the technical problems in the related art.

[0004] To this end, the first object of the present application is to propose a weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch, so as to predict the pseudo-labels of unlabeled data through the point mutual information between labeled data and unlabeled data, evaluate the weights of the pseudo-labels, and then selectively perform knowledge distillation on the classification model according to the weights to optimize the classification effect, avoid the negative impacts caused by incorrect labels and unknown classes, and improve the quality of classifier optimization.

[0005] The second object of the present application is to propose a weight-aware distillation device for semi-supervised 3D semantic segmentation with class distribution mismatch.

[0006] The third object of the present application is to propose an electronic device.

[0007] The fourth object of the present application is to propose a computer-readable storage medium.

[0008] The fifth object of the present application is to propose a computer program product.

[0009] To achieve the above object, the first aspect embodiment of the present application proposes a weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch, including: Obtain a plurality of first data pairs and a plurality of second data, wherein each first data pair includes first data and a first label of the first data; Extract features from the first data and the second data respectively to obtain a first feature of the first data and a second feature of the second data; Determine the point mutual information queue corresponding to any second data based on the relationship between the second feature of any second data and the first feature of each first data, where the point mutual information queue contains multiple point mutual informations, and each point mutual information corresponds to a first data; Determine the second label of any second data and the weight of the second label according to the magnitudes of the respective point mutual informations in the point mutual information queue and the first label of the first data corresponding thereto; Calculate the loss value of the initial classification model based on the multiple first data pairs, the multiple second data, the second label of each second data, and the weight of each second label; Update the parameters of the initial classification model according to the loss value to obtain an optimized target classification model.

[0010] To achieve the above object, an embodiment of the second aspect of the present application proposes a weight-aware distillation device for semi-supervised 3D semantic segmentation class distribution mismatch, including: An acquisition module, configured to acquire a plurality of first data pairs and a plurality of second data, where each first data pair includes a first data and a first label of the first data; A processing module, configured to perform feature extraction on the first data and the second data respectively to obtain a first feature of the first data and a second feature of the second data; A first calculation module, configured to determine the point mutual information queue corresponding to any second data based on the relationship between the second feature of any second data and the first feature of each first data, where the point mutual information queue contains multiple point mutual informations, and each point mutual information corresponds to a first data; A second calculation module, configured to determine the second label of any second data and the weight of the second label according to the magnitudes of the respective point mutual informations in the point mutual information queue and the first label of the first data corresponding thereto; A third calculation module, based on the multiple first data pairs, the multiple second data, the second label of each second data, and the weight of each second label, calculates the loss value of the initial classification model; An update module, configured to update the parameters of the initial classification model according to the loss value to obtain an optimized target classification model.

[0011] To achieve the above object, an embodiment of the third aspect of the present application proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer-executable instructions stored in the memory to implement the weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch as described in the embodiments of the first aspect of this application.

[0012] To achieve the above object, an embodiment of the fourth aspect of this application proposes a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch as described in the embodiments of the first aspect of this application.

[0013] To achieve the above object, an embodiment of the fifth aspect of this application proposes a computer program product, including a computer program, which when executed by a processor implements the weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch as described in the embodiments of the first aspect of this application.

[0014] The weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch provided by this application predicts the pseudo-labels of unlabeled data through the point mutual information between labeled data and unlabeled data, evaluates the weights of the pseudo-labels, and then selectively obtains the knowledge beneficial to the classification task according to the weights, performs knowledge distillation on the classification model, optimizes the classification effect, avoids the negative impacts caused by incorrect labels and unknown categories, and improves the quality of classifier optimization.

[0015] The additional aspects and advantages of this application will be partly given in the following description, partly become obvious from the following description, or be understood through the practice of this application. Description of the Drawings

[0016] The above and / or additional aspects and advantages of this application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where: Figure 1 It is a schematic flowchart of a weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch provided by an embodiment of this application; Figure 2 It is a schematic diagram of the process of feature extraction for data provided by an embodiment of this application; Figure 3 It is a schematic flowchart of another weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch provided by an embodiment of this application; Figure 4 It is a schematic flowchart of another weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch provided by an embodiment of this application; Figure 5A flow chart of another weight-aware distillation method for semi-supervised 3D semantic segmentation with mismatched class distribution provided by an embodiment of the present application; Figure 6 A flow chart of another weight-aware distillation method for semi-supervised 3D semantic segmentation with mismatched class distribution provided by an embodiment of the present application; Figure 7 A schematic diagram of a system based on a weight-aware distillation optimization classification model provided in an embodiment of the present application; Figure 8 A structural schematic diagram of a weight-aware distillation device for semi-supervised three-dimensional semantic segmentation with mismatched class distribution provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0018] The following describes the weight-aware distillation method and device for semi-supervised three-dimensional semantic segmentation with mismatched class distribution with reference to the accompanying drawings, according to an embodiment of the present application.

[0019] Figure 1 A flow chart of a weight-aware distillation method for semi-supervised 3D semantic segmentation with mismatched class distribution provided in an embodiment of the present application.

[0020] Existing research in semi-supervised learning (SSL) primarily focuses on image classification and semantic image segmentation. Semi-supervised 3D semantic segmentation utilizes a small number of densely labeled point cloud frames and a large number of unlabeled point cloud frames for model training, which can alleviate the annotation burden to a certain extent. Class distribution mismatch is a common problem in semi-supervised learning. In addition to incorrectly annotated data of the target class, data of unknown classes can also intrude into the classifier's training as outliers, leading to classification errors in semi-supervised learning. Therefore, semi-supervised learning errors can be divided into pseudo-labeling errors and intrusion errors.

[0021] Current approaches to combating class distribution mismatch primarily utilize pseudo-labeling or consistency regularization to expand the label pool and weighted filtering of instances with unknown categories. Semi-supervised learning methods with consistency regularization and pseudo-labeling rely heavily on the performance of the classifier and are generally less effective. Furthermore, weighted filtering of instances with unknown categories can result in many instances with unknown categories having the same output and receiving high weights, impacting the classifier and impairing its performance.

[0022] To address this issue, the embodiment of the present application provides a weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch, which can decouple the semi-supervised learning error under 3D scene class distribution mismatch into pseudo-labeling error and intrusion error. By using the point mutual information between labeled data and unlabeled data, the pseudo-labels of unlabeled data are predicted, and the weights of these pseudo-labels are evaluated. Then, according to the weights, knowledge distillation is selectively performed on the classification model to optimize the classification effect, avoid the negative impacts caused by incorrect labels and unknown classes, and improve the quality of classifier optimization.

[0023] As Figure 1 shown, the weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch includes the following steps: Step 101, obtain a plurality of first data pairs and a plurality of second data.

[0024] Each first data pair includes first data and a first label of the first data.

[0025] The first data is labeled data, and the first label is the classification to which the first data is labeled. According to actual needs, the first labels of multiple first data can be of multiple types. For example, there can be k classifications, where k is a positive integer.

[0026] The second data is unlabeled data.

[0027] In the embodiment of the present application, the first data and the second data can be 3D point cloud data. Each first data or second data is a point, which can include values in three dimensions, expressed as ( ). When studying the K-classification problem in a semi-supervised learning manner, the first data can be denoted as , the second data can be denoted as , , is the class set, and the set contains K classifications. There are first data, and second data.

[0028] It should be noted that since labeled data requires manual annotation, which is costly and time-consuming, while unlabeled data can be automatically collected through sensors or logs, is easy to obtain and low-cost, and unlabeled data can provide distribution information, alleviate overfitting, and improve model performance. Therefore, in the present application, the number of second data should be much larger than the number of first data, that is, , and the specific numbers of the first data and the second data can be adjusted according to actual needs.

[0029] Step 102, perform feature extraction on the first data and the second data respectively to obtain the first feature of the first data and the second feature of the second data.

[0030] In the embodiment of the present application, unsupervised contrastive coding can be used as a teacher model to learn representations from labeled and unlabeled data. Using the pre-trained teacher model, feature extraction is performed on the first data and the second data respectively to obtain the first feature of the first data and the second feature of the second data. The process of the teacher model extracting features from the first data or the second data can be as follows: Figure 2 As shown, Figure 2 This is a schematic diagram of a feature extraction process provided in an embodiment of the present application.

[0031] exist Figure 2 In the feature extraction process shown, the direction and shape of the first data or the second data input to the teacher model can be normalized through transformation matrix and matrix multiplication, the data direction difference can be eliminated, the data can be aligned and normalized, and the robustness of feature extraction can be improved. Figure 2 In the step of transforming the matrix, the first data or the second data (i.e. Figure 2 The output is an affine transformation matrix, which maps the input data to a standard coordinate system, eliminating any directional differences. The matrix multiplication step then applies this affine transformation to the input data, normalizing its orientation and shape.

[0032] Then, if Figure 2 As shown in Figure 2, after obtaining the three-dimensional data with normalized direction and shape, the three-dimensional data is passed through a multi-layer perceptron with two hidden layers and 64 neurons in each layer. The multi-layer perceptron can map the three-dimensional data to a high-dimensional feature space and obtain a vector with a feature dimension of 64 (i.e. Figure 2 (n*64 in [1]). This 64-dimensional feature vector can then be normalized in direction and shape using a transformation matrix and matrix multiplication. The transformation matrix and matrix multiplication for data of different dimensions are similar, except that the affine transformation has different dimensions. The affine transformation dimension matches the data dimension to be mapped. The multiplication result can then be passed through a multilayer perceptron with three hidden layers, each containing 64, 128, and 1024 neurons, to obtain a feature extraction result (i.e., the first or second feature) with a feature dimension of 1024.

[0033] It should be noted that in Figure 2 The feature extraction process shown in is only a schematic illustration. In different specific implementation scenarios, the parameters of the multilayer perceptron can be adjusted as needed, and this application does not limit this.

[0034] In an embodiment of the present application, during the training of the teacher model, first data may be input as input data into the teacher model to obtain a predicted label output by the teacher model for the first data. The predicted label may be a soft label, i.e., a set of K scores corresponding to K categories. A correction gradient for the teacher model may then be determined based on the distance between the predicted label and the first label. The teacher model is then iteratively updated until the gap between the predicted label and the true labeled first label is less than a threshold, completing the training of the teacher model.

[0035] Step 103 : determining a point mutual information queue corresponding to any second data based on a relationship between a second feature of any second data and a first feature of each first data.

[0036] Among them, pointwise mutual information (PMI) is used to measure the correlation strength between two specific events or data. The point mutual information queue can contain multiple point mutual information, and each point mutual information corresponds to a first data.

[0037] In an embodiment of the present application, the marginal distribution of each first feature and the conditional distribution of the first feature based on any second feature can be calculated first, and then the point mutual information between the corresponding first data and the second data can be determined based on the marginal distribution and the conditional distribution. After that, all the point mutual information corresponding to the same second data can be arranged in order from large to small to obtain the point mutual information queue corresponding to the second data.

[0038] Alternatively, in some embodiments, the conditional and marginal distributions may not be directly calculated, but the point mutual information is proportional to the inner product. The larger the inner product, the larger the point mutual information. Therefore, the inner product between each first feature and the second feature can be calculated, and the result of the inner product calculation can be used as the point mutual information between the corresponding first data and the second data. Then, all the point mutual information corresponding to the same second data can be arranged in order from large to small to obtain the point mutual information queue corresponding to the second data.

[0039] Step 104 : Determine a second label and a weight of any second data according to the magnitude of each point mutual information in the point mutual information queue and the first label of the corresponding first data.

[0040] The second label is the predicted category of the second data, which can also be called a pseudo label.

[0041] The weight, or confidence level of the label, is used to measure the teacher model's predicted labels for unlabeled data and its role in training the student model's classifier. The larger the weight, the greater the role. In this application, the weights of each data label are used to transfer the teacher model's knowledge to the student model, completing knowledge distillation to address the class distribution mismatch problem in semi-supervised 3D semantic segmentation.

[0042] In this embodiment of the present application, to reduce intrusion errors, instances of unknown class should be filtered out. After determining the point mutual information queue corresponding to each second data item, a higher point mutual information in the point mutual information queue indicates a stronger correlation or similarity between the corresponding first data item and the second data item, further indicating a higher likelihood that the second data item and the first data item belong to the same class distribution. Therefore, the first label of the first data item corresponding to the maximum point mutual information in the point mutual information queue can be determined as the second label of the second data item corresponding to that point mutual information queue.

[0043] It should be noted that in some possible situations, some second data may have similar point-wise mutual information with two first data with different first labels. In this case, the classification to which the second data belongs is on the decision boundary between the two categories, which may introduce incorrect pseudo-labels and impair the performance of the classifier. Therefore, in this application, the ratio between the first maximum point-wise mutual information and the second maximum point-wise mutual information can be used to evaluate the weight of the pseudo-label (i.e., the second label) of the second data to avoid the negative impact caused by incorrect labels and unknown categories.

[0044] In an embodiment of the present application, the first point mutual information with the largest point mutual information and the second point mutual information with the second largest point mutual information in the point mutual information queue can be obtained first, and then the ratio of the second point mutual information to the first point mutual information is calculated. The closer the ratio is to 1, the more similar the first point mutual information is to the second point mutual information. The first labels corresponding to the first point mutual information and the second point mutual information may be the category of the second data, and the label is not clear enough. At this time, if the first label corresponding to the first point mutual information is used as the second label, there may be a classification error problem, and the weight of the second label will be lower.

[0045] It should be noted that when weighing weights, the maximum mutual information value in the point mutual information queue must also be considered. This is because when the maximum mutual information is relatively small, the correlation strength between the second data and the first data corresponding to the maximum mutual information is also relatively small, and the probability of belonging to the same category is relatively small. In this case, the first label of the first data corresponding to the maximum mutual information is used as the second label of the second data, and the weight is not high. Therefore, a monotonically increasing function can be constructed with the largest mutual information value in the point mutual information queue as the independent variable to estimate the probability that the second data belongs to the category of the first data corresponding to the first largest mutual information value. The higher the function value, the more likely the second data is to belong to the category of the first data corresponding to the first largest mutual information value, and the higher the weight.

[0046] In this application, the weight of the second label is determined to filter out instances of unknown categories and incorrectly labeled instances in the target category (i.e., the K categories for classifying the data), so that when performing knowledge distillation and optimization on the classification model, the use of instances with high-quality pseudo-labels from the target category is encouraged.

[0047] In the embodiment of the present application, in order to make full use of the unlabeled second data, a pseudo label (i.e., second label) can be assigned to each second data, which is expressed as Pseudo-labels can be obtained by classification from a teacher model, which is learned from all labeled and unlabeled data in an unsupervised manner through contrastive learning. Knowledge distillation is then achieved by transferring the beneficial knowledge of the teacher model to the classification model (i.e., the student model or classifier). The teacher model can produce closely aligned representations for instances from the same category and maximize the mutual information between them.

[0048] Step 105 : Calculate the loss value of the initial classification model based on the plurality of first data pairs, the plurality of second data, the second label of each second data, and the weight of each second label.

[0049] Among them, the loss value of the initial classification model can be a weight-aware knowledge distillation loss, which includes the traditional supervision loss in the first data, i.e., the labeled data, and the weight-aware supervision loss in the second data, i.e., the unlabeled data.

[0050] In an embodiment of the present application, each first data in the first data pair can be input into the initial classification model to obtain the predicted label output by the initial classification model. Then, the distance between the predicted label and the first label of the first data in the first data pair is calculated. The distances corresponding to all first data are added and then divided by the number of first data to obtain the traditional supervised loss in the first data, i.e., the labeled data. The traditional supervised loss aims to minimize the distance between the predicted probability and the ground truth label.

[0051] In an embodiment of the present application, each second data can be input into the initial classification model to obtain the predicted label output by the initial classification model, and then the distance between the predicted label and the second label of the second data is calculated. The weight corresponding to the second label is used to perform a weighted sum of the distances corresponding to all the second data and then divided by the number of second data to obtain the weight-aware supervision loss of the second data, i.e., the unlabeled data. The weight-aware supervision loss of the second data, i.e., the unlabeled data, is mainly to selectively transfer the useful knowledge in the teacher model to the student model through weights to reduce the negative impact of unknown categories and improve the classification model. Afterwards, the traditional supervision loss in the first data, i.e., the labeled data, and the weight-aware supervision loss of the second data, i.e., the unlabeled data, are added together to determine the loss value of the initial classification model.

[0052] Step 106: Update the parameters of the initial classification model according to the loss value to obtain an optimized target classification model.

[0053] In the embodiments of the present application, an optimization algorithm can be used to update the parameters of the student model, minimize the loss value, and obtain an optimized target classification model. For example, a gradient descent method can be used to calculate the gradient of the loss value with respect to the parameters of the initial classification model, and then the parameters of the initial classification model can be updated based on the actual required learning rate and gradient.

[0054] It should be noted that in this application, pseudo labels and weights are used to reduce pseudo labeling and intrusion errors, and then reduce semi-supervised learning errors, to update the initial classification model, and perform knowledge distillation. Since the labeled data is limited, it may cause the knowledge of pseudo labels and weights to be biased. Therefore, after several forward iterations, some reliable instances can be gradually added to the labeled data. Multiple data with higher reliability can be determined in the second data, used as new first data, and deleted from the original second data. The updated first data and second data are then used to update the point mutual information, determine the second label, and calculate the loss, etc., to achieve further distillation of the classification model. Labeled data can be added multiple times, and polynomial decay can be used to dynamically adjust the amount of data added each time to prevent the negative impact of unknown categories that gradually increase with iteration.

[0055] In this embodiment, multiple first data pairs and multiple second data are first obtained, where each first data pair includes the first data and the first label of the first data. Feature extraction is then performed on the first and second data to obtain the first features of the first data and the second features of the second data. Based on the relationship between the second features of any second data and the first features of each first data, a point mutual information queue corresponding to any second data is determined. Next, based on the magnitude of each point mutual information in the point mutual information queue and the first label of the corresponding first data, the second label and the weight of the second label are determined. The loss value of the initial classification model is then calculated based on the multiple first data pairs, the multiple second data, the second labels of each second data, and the weight of each second label. The parameters of the initial classification model are then updated based on the loss value to obtain an optimized target classification model. Pseudo-labels for the unlabeled data are predicted using the point mutual information between the labeled and unlabeled data, and the weights of the pseudo-labels are evaluated. Based on the weights, knowledge that is beneficial to the classification task is selectively extracted to perform knowledge distillation on the classification model, optimizing the classification performance, avoiding the negative impact caused by incorrect labels and unknown categories, and improving the quality of classifier optimization.

[0056] This embodiment provides another weight-aware distillation method for semi-supervised 3D semantic segmentation with mismatched class distribution. Figure 3 A flowchart of another weight-aware distillation method for semi-supervised 3D semantic segmentation with mismatched class distribution is provided for this embodiment.

[0057] like Figure 3 As shown, the classification model optimization method may include the following steps: Step 301: Acquire multiple first data pairs and multiple second data.

[0058] Step 302 : extract features from the first data and the second data respectively to obtain a first feature of the first data and a second feature of the second data.

[0059] For detailed description of the above steps 301 and 302, please refer to the above embodiments of the present application and will not be repeated here.

[0060] Step 303: Determine a first probability distribution of the first data with respect to any second data based on the second feature of any second data and the first feature of the first data.

[0061] The first probability distribution is a conditional distribution of the first data with the second data as a condition, and is the ratio between the probability of the first data and the second data occurring simultaneously and the probability of the second data.

[0062] In the embodiments of the present application, after feature extraction is performed on the first data and the second data, the potential structure of the data can be determined through the extracted first feature and second feature, and the conditional dependence relationship between the features can be analyzed to determine the distribution of a certain random feature under the condition of given other features, so as to obtain a first probability distribution of a first data for any second data.

[0063] In the embodiments of the present application, the first probability distribution can be expressed as , where is a second data, takes values from 1 to n, n is the number of second data, u represents that the data belongs to the set of unlabeled data, is a first data, takes values from 1 to m, m is the number of first data, represents that the data belongs to the set of labeled data, and k is the classification to which the first data belongs.

[0064] Step 304: Determine the second probability distribution of the first data based on the first feature of the first data.

[0065] Among them, the second probability distribution is the marginal distribution of the first data, which describes the distribution of the first feature of a single first data without considering the influence of other variables.

[0066] In the embodiments of the present application, after feature extraction is performed on the first data, the potential structure of the first data can be determined through the extracted first feature, so that the marginal distribution of each first feature can be estimated, and the second probability distribution of the first data can be obtained.

[0067] In the embodiments of the present application, the second probability distribution can be expressed as , where is a first data, takes values from 1 to m, m is the number of first data, represents that the data belongs to the set of labeled data, and k is the classification to which the first data belongs.

[0068] Step 305: Calculate the point mutual information between the first data and any second data according to the first probability distribution and the second probability distribution.

[0069] In the embodiments of the present application, the point mutual information can be calculated based on the following formula (1): (1) In formula (1), is the first data, is the second data, is and the point mutual information between, is the first probability distribution, is the second probability distribution.

[0070] Step 306 , arranging all point mutual information between any second data and all first data from largest to smallest, to obtain a point mutual information queue corresponding to any second data.

[0071] In the embodiment of the present application, the above formula (1) can be used to obtain all the point mutual information of each second data with respect to all the first data, and then these point mutual information are sorted in descending order of value to obtain the point mutual information queue corresponding to the second data, thereby obtaining the point mutual information queues corresponding to all the second data.

[0072] In this embodiment, by calculating the point mutual information between each labeled data and the unlabeled data based on the calculated conditional distribution and marginal distribution, the degree of association between each unlabeled data and all labeled data can be determined, and a point mutual information queue can be obtained, which provides conditions for determining the labeled data with the highest degree of association with the unlabeled data and improving the reliability of obtaining pseudo labels for the unlabeled data.

[0073] Step 307 : Determine a second label and a weight of any second data according to the magnitude of each point mutual information in the point mutual information queue and the first label of the corresponding first data.

[0074] In step 308 , a loss value of the initial classification model is calculated based on the plurality of first data pairs, the plurality of second data, the second label of each second data, and the weight of each second label.

[0075] Step 309: Update the parameters of the initial classification model according to the loss value to obtain an optimized target classification model.

[0076] For detailed descriptions of steps 307 to 309, please refer to the above embodiments of the present application and will not be repeated here.

[0077] It should be noted that in some embodiments, the conditional distribution and marginal distribution may not be directly calculated. In this case, the point mutual information between the first data and the second data cannot be calculated using formula (1). However, the point mutual information is proportional to the inner product. The larger the inner product, the larger the point mutual information. Therefore, the inner product between each first feature and the second feature can be calculated, and the result of the inner product calculation is used as the point mutual information between the corresponding first data and the second data. Then, all the point mutual information corresponding to the same second data are arranged in order from large to small, and the point mutual information queue corresponding to the second data can be obtained.

[0078] Optionally, the inner product between the second feature of any second data and the first feature of the first data can be calculated first to determine the point mutual information between the first data and any second data. Then, all the point mutual information between any second data and all the first data is arranged in descending order to obtain the point mutual information queue corresponding to any second data.

[0079] In the embodiments of the present application, the relationship between the point mutual information and the inner product is shown by the following formula (2): (2) Wherein, represents the second data and the first data the inner product between them, that is, the cosine of the features of the first data and the second data is calculated, and the first data and the second data can be normalized when calculating the inner product. represents "proportional to".

[0080] In the present application, in the case of determining the point mutual information between the first data and the second data based on the inner product between the first data and the second data, the second label of the second data can be represented by the following formula (3): (3) Wherein, is the second label of the second data and is the first label of the first data corresponding to the maximum value among the inner products between the second data and all the first data.

[0081] In the embodiments of the present application, the class label of the labeled instance with the maximum point mutual information can be assigned to the unlabeled instance, so as to accurately capture the point mutual information from the representation space to generate high-quality pseudo-labels and reduce pseudo-label errors.

[0082] The embodiments of the present application also provide another weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch. Figure 4 It is a schematic flow chart of another weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch provided by this embodiment.

[0083] As Figure 4 shown, this weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch may include the following steps: Step 401, obtain a plurality of first data pairs and a plurality of second data.

[0084] Step 402, respectively perform feature extraction on the first data and the second data to obtain the first feature of the first data and the second feature of the second data.

[0085] Step 403: Determine the point mutual information queue corresponding to any second data based on the relationship between the second feature of any second data and the first feature of each first data.

[0086] For the detailed descriptions of the above steps 401 to 403, reference can be made to the above embodiments of this application, which will not be elaborated here.

[0087] Step 404: Determine the first point mutual information with the maximum point mutual information in the point mutual information queue.

[0088] In the embodiment of this application, the first point mutual information can be expressed by the following formula (4): (4) Wherein, represents the first point mutual information corresponding to the second data ; is the maximum value of the point mutual information between the second data and all the first data (i.e., the maximum value of the inner product); takes values from 1 to m, where m is the number of the first data.

[0089] Step 405: Determine the first label of the first data corresponding to the first point mutual information as the second label of any second data.

[0090] In the embodiment of this application, after determining the point mutual information queue corresponding to each second data, the higher the point mutual information in the point mutual information queue, the stronger the association or similarity between the corresponding first data and the second data, and the more likely it is that the first data and the second data belong to the same class distribution. Therefore, the first label of the first data corresponding to the maximum first point mutual information in the point mutual information queue can be determined as the second label of the second data corresponding to this point mutual information queue.

[0091] Step 406: Determine the second point mutual information with the maximum point mutual information except the first point mutual information in the point mutual information queue.

[0092] In the embodiment of this application, the second point mutual information can be expressed by the following formula (5): (5) Wherein, represents the second point mutual information corresponding to the second data ; is the maximum value of the point mutual information between the second data and the other first data except the first data (i.e., the maximum value of the inner product).

[0093] Step 407: Calculate the weight of the second label based on the first point mutual information and the second point mutual information.

[0094] In the embodiment of the present application, the weight of the second tag can be calculated based on the following formula (6): (6) in, For the second data The weight of the second label, , It can be any monotonically increasing function.

[0095] From the above formula (6), we can see that the monotonically increasing function The first point mutual information is used as the independent variable to estimate the probability that the second data belongs to the first label corresponding to the first point mutual information. The higher the first point mutual information, The higher the function value, the higher the probability that the second data belongs to the category of the first label corresponding to the mutual information of the first point, and the higher the weight of the first label as the second label. by is the independent variable, that is, the mutual information at the second point Mutual information with the first point The closer the ratio is to 1 (i.e. and the closer), The lower the function value is, the lower the weight of the first label as the second label is, and the second data has unclear labels between the two categories corresponding to the first point mutual information and the second point mutual information.

[0096] In step 408 , a loss value of the initial classification model is calculated based on the plurality of first data pairs, the plurality of second data, the second label of each second data, and the weight of each second label.

[0097] Step 409: Update the parameters of the initial classification model according to the loss value to obtain an optimized target classification model.

[0098] For detailed descriptions of steps 408 to 409, please refer to the above embodiments of the present application and will not be repeated here.

[0099] In this embodiment, for the second data, the ratio between the two corresponding maximum point mutual information is determined to measure the first label of the first data corresponding to the maximum point mutual information. This is used as the weight of the pseudo-label of the second data. This can filter out instances of unknown categories and incorrectly labeled instances in the target category, reduce intrusion errors, and help improve the efficiency and quality of knowledge distillation of the classification model.

[0100] This embodiment provides another weight-aware distillation method for semi-supervised 3D semantic segmentation with mismatched class distribution. Figure 5 A flowchart of another weight-aware distillation method for semi-supervised 3D semantic segmentation with mismatched class distribution is provided for this embodiment.

[0101] like Figure 5 As shown, the weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch may include the following steps: Step 501: Acquire multiple first data pairs and multiple second data.

[0102] Step 502 : extract features from the first data and the second data respectively to obtain a first feature of the first data and a second feature of the second data.

[0103] Step 503: Determine a point mutual information queue corresponding to any second data based on a relationship between the second feature of any second data and the first feature of each first data.

[0104] Step 504 : Determine a second label and a weight of any second data according to the magnitude of each point mutual information in the point mutual information queue and the first label of the corresponding first data.

[0105] For detailed description of the above steps 501 to 504, please refer to the above embodiments of the present application and will not be repeated here.

[0106] Step 505: input the first data into the initial classification model to obtain a first prediction label output by the initial classification model.

[0107] The initial classification model is a classifier, also known as a student model. It can be a lightweight and compact model with high computational efficiency.

[0108] It is understandable that compared with large and complex deep neural networks (i.e., teacher models), the initial classification model is smaller and its classification accuracy is not as high as that of the teacher model. Directly training the initial classification model may be difficult to achieve the performance of the teacher model. Therefore, in this application, the adaptive weights and high-quality pseudo-labels captured in the teacher model can be used to transfer knowledge from the teacher model to the initial classification model for knowledge distillation, thereby completing the optimization of the initial classification model.

[0109] In the embodiment of the present application, the first data can be input into the initial classification model to obtain the first predicted label output by the initial classification model. The first predicted label can be expressed as ,in, are the parameters of the initial classification model.

[0110] Step 506: Determine a first difference between the first predicted label and the first label corresponding to each first data.

[0111] In the embodiment of the present application, any loss function can be used to calculate the first difference between the first predicted label corresponding to each first data and its corresponding first label. The first difference can be expressed as ,in First Data The first label.

[0112] Step 507: sum up all the first differences and take the average value to obtain a first loss value.

[0113] Among them, the first loss value is the traditional supervision loss in the labeled data.

[0114] In the embodiment of the present application, the first loss value can be calculated by the following formula (7): (7) in, is the first loss value, is the set of first data, is the number of the first data.

[0115] Step 508: Input the second data into the initial classification model to obtain a second predicted label output by the initial classification model.

[0116] In the embodiment of the present application, the second data can be input into the initial classification model to obtain the second predicted label output by the initial classification model. The second predicted label can be expressed as ,in, are the parameters of the initial classification model.

[0117] Step 509 : Determine a second difference between the second predicted label and the second label corresponding to each second data.

[0118] In this embodiment of the present application, any loss function can be used to calculate the second difference between the second predicted label corresponding to each second data and its corresponding second label. The second difference can be expressed as ,in For the second data The second label.

[0119] In step 510 , all second differences are weightedly summed using weights and the average value is taken to obtain a second loss value.

[0120] Among them, the second loss value is the weight-aware supervision loss in unlabeled data.

[0121] In the embodiment of the present application, the first loss value can be calculated by the following formula (8): (8) in, is the second loss value, is the set of second data, is the number of the second data, For the second data The weight of the second label can be calculated by formula (6) in the above embodiment.

[0122] Step 511: Add the first loss value and the second loss value to obtain the loss value of the initial classification model.

[0123] In the embodiment of the present application, the loss value of the initial classification model can be expressed as 。

[0124] In this embodiment, by calculating the traditional supervision loss in labeled data and the weight-aware supervision loss in unlabeled data, the weight-aware knowledge distillation loss of the initial classification model is obtained, and then the update and optimization of the initial classification model is achieved, which can reduce the negative impact of unknown categories on the knowledge distillation of the classification model, and is conducive to improving the reliability of the knowledge distillation of the classification model and improving the robustness of the classification model.

[0125] Step 512: Update the parameters of the initial classification model according to the loss value to obtain an optimized target classification model.

[0126] For a detailed description of the above step 512, please refer to the above embodiment of the present application, which will not be repeated here.

[0127] It should be noted that in this application, pseudo labels and weights are used to reduce pseudo labeling and intrusion errors, and then reduce semi-supervised learning errors, to update the initial classification model and perform knowledge distillation. Since the labeled data is limited, it may cause the knowledge of pseudo labels and weights to be biased. Therefore, after several forward iterations, some reliable instances can be gradually added to the labeled data to further optimize the classification model.

[0128] The present application also provides another weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch. Figure 6 A flowchart of another weight-aware distillation method for semi-supervised 3D semantic segmentation with mismatched class distribution is provided for this embodiment.

[0129] like Figure 6 As shown, the weight-aware distillation method for semi-supervised 3D semantic segmentation class distribution mismatch may include the following steps: Step 601: Input the second data into the optimized target classification model to obtain the third predicted label output by the target classification model.

[0130] In the embodiments of the present application, after updating the parameters of the initial classification model to obtain the optimized target classification model, since the parameters have changed, the label predicted by the classification model for the second data may also have changed. In order to measure the classification quality of the optimized target classification model and further distill and update the target classification model, the third predicted label output by the target classification model for the second data can be obtained, and then based on the difference between the third predicted label and the pseudo-label predicted by the teacher model, the reliability of the classification result of the classification model for the second data at this time can be determined.

[0131] Step 602: Calculate the difference between the third predicted label and the second label to determine the credibility of the second data.

[0132] In the embodiments of the present application, the credibility of the second data can be determined by the following formula (9). The smaller the result obtained by formula (9), the higher the credibility of the second data.

[0133] (9) Wherein, is the cross-entropy function, is the parameter of the classification model at the t-th iteration.

[0134] Step 603: Sort the multiple second data in descending order of credibility, and obtain one or more target second data whose order is before the order threshold.

[0135] Among them, the order threshold can be the first few percentages in the order, and its value can be determined according to needs.

[0136] For example, if there are 100 second data, after sorting them in descending order of credibility, when the order threshold is 10%, the 10 second data with the highest credibility in the sorting can be determined as the target data.

[0137] It should be noted that in the present application, the target second data can be obtained multiple times as the labeled data to update the point mutual information between the labeled data and the unlabeled data, and then the classification model is updated again. In order to prevent the negative impact of the unknown classes that gradually increase with the iteration, the order threshold can be dynamically adjusted using polynomial decay, and when the order threshold decays to the minimum value, the iterative update of the classification model is stopped.

[0138] Step 604: Update the original first data pair and the second data using the target second data and the second label corresponding to the target second data.

[0139] In an embodiment of the present application, after determining the target second data, the target second data and the second label corresponding to the target second data can be placed as a new first data pair into the marked data, and the target second data can be deleted from the unmarked data, so that an updated first data pair and second data can be obtained.

[0140] Step 605, based on the updated first data pair and the second data, calculate the loss of the target classification model, update the target classification model, and return to the step of using the updated target classification model to predict the label of the second data, and perform multiple iterations until the iteration stop condition is met to obtain the final optimized target classification model.

[0141] Among them, the iteration termination condition can be that the difference between the classification prediction of the classification model for all unlabeled data and the pseudo label of the teacher model is less than the threshold, or it can be that the sequence threshold decays to the minimum value at a certain iteration when the polynomial decay is used to dynamically adjust the sequence threshold.

[0142] In this embodiment of the present application, the target classification model can be updated by recalculating the pseudo labels, point mutual information, and target classification model loss based on the updated first data pair and second data according to the steps described in the above embodiment. The updated classification model can then be used to return to step 601 to once again obtain the predicted label for the current second data, and iterate multiple times until the iteration stop condition is met, thereby obtaining the final optimized target classification model.

[0143] In this embodiment, multiple data with higher reliability are determined in the second data, used as new first data, and deleted from the original second data. The updated first data and second data are then used to update the point mutual information, determine the second label, and calculate the loss, etc., to achieve further distillation and update of the classification model, which can further improve the quality of the classification model and improve the reliability and robustness of the classification model.

[0144] The following is based on Figure 7 To describe the process of optimizing classification models through knowledge distillation, Figure 7 A system diagram of a weight-aware distillation optimization classification model provided in an embodiment of the present application.

[0145] Depend on Figure 7It can be seen that by using the encoder of unsupervised contrastive coding as the teacher model and inputting unlabeled instances (i.e., the second data in the above embodiments) into the teacher model, the point mutual information between the unlabeled instances and all labeled instances (i.e., the first data in the above embodiments) can be obtained, and then the maximum point mutual information and the second largest point mutual information can be determined. After that, the pseudo-label (i.e., the second label in the above embodiments) and weight (i.e., the weight) of the unlabeled instance can be determined from the maximum point mutual information and the second largest point mutual information. Then, the weight-aware supervision loss in the unlabeled data can be calculated by using the weight information and pseudo-label information and the predicted label obtained after inputting the unlabeled data into the student model (i.e., the classifier, classification model). And input the labeled data into the student model to determine the traditional supervision loss between its predicted label and the true labeled label. Then use and the sum of to update the mutual information between the labeled instances and the unlabeled instances, so as to perform knowledge distillation on the student model and optimize the classifier.

[0146] To implement the above embodiments, the present application also proposes a weight-aware distillation device for semi-supervised 3D semantic segmentation with class distribution mismatch.

[0147] Figure 8 FIG. is a schematic structural diagram of a weight-aware distillation device for semi-supervised 3D semantic segmentation with class distribution mismatch provided by an embodiment of the present application.

[0148] As Figure 8 shown, the weight-aware distillation device 80 for semi-supervised 3D semantic segmentation with class distribution mismatch includes: An acquisition module 801, configured to acquire a plurality of first data pairs and a plurality of second data, wherein each first data pair includes first data and a first label of the first data; A processing module 802, configured to perform feature extraction on the first data and the second data respectively to obtain a first feature of the first data and a second feature of the second data; A first calculation module 803, configured to determine a point mutual information queue corresponding to any second data based on the relationship between the second feature of any second data and the first feature of each first data, wherein the point mutual information queue includes a plurality of point mutual informations, and each point mutual information corresponds to a first data; A second calculation module 804, configured to determine a second label of any second data and the weight of the second label according to the magnitudes of the point mutual informations in the point mutual information queue and the first label of the first data corresponding thereto; A third calculation module 805, configured to calculate a loss value of the initial classification model based on a plurality of first data pairs, a plurality of second data, the second label of each second data, and the weight of each second label; The updating module 806 is used to update the parameters of the initial classification model according to the loss value to obtain an optimized target classification model.

[0149] Furthermore, in a possible implementation of the embodiment of the present application, the first calculation module 803 may be specifically configured to: Determine a first probability distribution of the first data with respect to any second data based on the second feature of any second data and the first feature of the first data; determining a second probability distribution of the first data based on the first feature of the first data; According to the first probability distribution and the second probability distribution, the point mutual information between the first data and any second data is calculated. Arrange all point mutual information between any second data and all first data from largest to smallest to obtain a point mutual information queue corresponding to any second data; Among them, the point mutual information is calculated based on the following formula,

[0150] In the formula, is the first data, is the second data, for and The point mutual information between is the first probability distribution, is the second probability distribution.

[0151] Furthermore, in a possible implementation of the embodiment of the present application, the first calculation module 803 may be specifically configured to: Calculate the inner product between the second feature of any second data and the first feature of the first data to determine the point mutual information between the first data and any second data; Arrange all point mutual information between any second data and all first data from largest to smallest to obtain a point mutual information queue corresponding to any second data.

[0152] Furthermore, in a possible implementation of the embodiment of the present application, the second calculation module 804 may be specifically configured to: Determine the first point mutual information with the largest point mutual information in the point mutual information queue; Determine the first label of the first data corresponding to the first point mutual information as the second label of any second data; Determine the second point mutual information with the largest mutual information in the point mutual information queue except the first point mutual information; Based on the mutual information of the first point and the mutual information of the second point, the weight of the second label is calculated. The weight is calculated based on the following formula:

[0153] in, For the second data The weight of the second label, and is a monotonically increasing function, is the first point mutual information, is the second mutual information.

[0154] Furthermore, in a possible implementation of the embodiment of the present application, the third calculation module 805 may be specifically configured to: Inputting the first data into the initial classification model to obtain a first predicted label output by the initial classification model; Determine a first difference between a first predicted label and the first label corresponding to each first data; Sum all the first differences and take the average to obtain the first loss value; Inputting the second data into the initial classification model to obtain a second predicted label output by the initial classification model; Determine a second difference between the second predicted label and the second label corresponding to each second data; Using the weights, all the second differences are weighted and summed up, and the average is taken to obtain the second loss value; Add the first loss value to the second loss value to get the loss value of the initial classification model.

[0155] Furthermore, in a possible implementation of the embodiment of the present application, the updating module 806 may also be configured to: Inputting the second data into the optimized target classification model to obtain a third predicted label output by the target classification model; Calculating the difference between the third predicted label and the second label to determine the credibility of the second data; Sorting the plurality of second data in descending order of credibility, and obtaining one or more target second data whose order in the sorting is before the order threshold; Updating the original first data pair and the second data using the target second data and the second label corresponding to the target second data; Based on the updated first data pair and the second data, the loss of the target classification model is calculated, the target classification model is updated, and the step of using the updated target classification model to predict the label of the second data is returned to. Multiple iterations are performed until the iteration stop condition is met to obtain the final optimized target classification model.

[0156] It should be noted that the above explanation of the weight-aware distillation method embodiment for the class distribution mismatch in semi-supervised 3D semantic segmentation also applies to the weight-aware distillation device for the class distribution mismatch in semi-supervised 3D semantic segmentation of this embodiment, and will not be elaborated here.

[0157] In the embodiments of the present application, the point mutual information between the labeled data and the unlabeled data is used to predict the pseudo-labels of the unlabeled data, and the weights of the pseudo-labels are evaluated. Then, according to the weights, the knowledge beneficial to the classification task is selectively obtained to perform knowledge distillation on the classification model, optimize the classification effect, avoid the negative impacts caused by wrong labels and unknown classes, and improve the quality of classifier optimization.

[0158] To implement the above embodiments, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments. To implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method provided in the foregoing embodiments.

[0159] To implement the above embodiments, the present application also proposes a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method provided in the foregoing embodiments.

[0160] In the present application, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0161] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and signing an agreement / authorization including authorizing the relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to the personal information data comply with the relevant policies and processes.

[0162] The present application anticipates providing an implementation plan for users to selectively block the use or access of personal information data. That is, the present disclosure anticipates providing hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to ensure user information security.

[0163] In the descriptions of the foregoing embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.

[0164] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0165] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0166] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0167] It should be understood that various parts of the present application can be implemented in the form of hardware, software, firmware, or a combination. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented in the form of any one or a combination of the following techniques well known in the art: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0168] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0169] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0170] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A weight-aware distillation method for class distribution mismatch in semi-supervised 3D semantic segmentation, characterized in that Including the following steps: Obtain a plurality of first data pairs and a plurality of second data, wherein each of the first data pairs includes first data and a first label of the first data; Perform feature extraction on the first data and the second data respectively to obtain a first feature of the first data and a second feature of the second data; Based on the relationship between the second feature of any one of the second data and the first feature of each of the first data, determine the point mutual information queue corresponding to the any one of the second data, wherein the point mutual information queue includes a plurality of point mutual informations, and each of the point mutual informations corresponds to a first data; According to the magnitudes of the respective point mutual informations in the point mutual information queue and the first label of the first data corresponding thereto, determine the second label of the any one of the second data and the weight of the second label; Based on the plurality of first data pairs, the plurality of second data, the second label of each of the second data, and the weight of each of the second labels, calculate the loss value of the initial classification model; According to the loss value, update the parameters of the initial classification model to obtain an optimized target classification model.

2. The method according to claim 1, wherein The determining the point mutual information queue corresponding to any one of the second data based on the relationship between the second feature of the any one of the second data and the first feature of each of the first data includes: Based on the second feature of any one of the second data and the first feature of the first data, determine the first probability distribution of the first data with respect to the any one of the second data; Based on the first feature of the first data, determine the second probability distribution of the first data; According to the first probability distribution and the second probability distribution, calculate the point mutual information between the first data and the any one of the second data, Arrange all the point mutual informations between the any one of the second data and all the first data in descending order to obtain the point mutual information queue corresponding to the any one of the second data; wherein, the point mutual information is calculated based on the following formula, In the said formula, is the first data, is the second data, is the point mutual information between and is the first probability distribution, is the second probability distribution.

3. The method according to claim 1, wherein The determining the point mutual information queue corresponding to any one of the second data based on the relationship between the second feature of the any one of the second data and the first feature of each of the first data includes: Calculate the inner product between the second feature of any one of the second data and the first feature of the first data, and determine the point mutual information between the first data and the any one of the second data; Arrange all the point mutual informations between the any one of the second data and all the first data in descending order to obtain the point mutual information queue corresponding to the any one of the second data.

4. The method according to any one of claims 1 to 3, characterized in that The determining the second label of the any one of the second data and the weight of the second label according to the magnitudes of the respective point mutual informations in the point mutual information queue and the first label of the first data corresponding thereto includes: Determine the first point mutual information with the largest point mutual information in the point mutual information queue; Determine the first label of the first data corresponding to the first point mutual information as the second label of the any one of the second data; Determine the second point mutual information with the largest point mutual information except the first point mutual information in the point mutual information queue; Based on the first point mutual information and the second point mutual information, calculate the weight of the second label, and the weight is calculated based on the following formula, Among them, is the weight of the second tag of the second data , and is a monotonically increasing function is the first pointwise mutual information is the second pointwise mutual information 5. The method according to claim 1, wherein Calculating the loss value of the initial classification model based on the multiple first data pairs, the multiple second data, the second labels of each second data, and the weights of each second label, includes: Inputting the first data into the initial classification model to obtain a first predicted label output by the initial classification model; Determining a first difference between the first predicted label corresponding to each first data and the first label; Summing up all the first differences and then taking the average to obtain a first loss value; Inputting the second data into the initial classification model to obtain a second predicted label output by the initial classification model; Determining a second difference between the second predicted label corresponding to each second data and the second label; Using the weights to perform a weighted sum of all the second differences and then taking the average to obtain a second loss value; Adding the first loss value and the second loss value to obtain the loss value of the initial classification model.

6. The method according to claim 1, wherein After updating the parameters of the initial classification model according to the loss to obtain an optimized target classification model, it further includes: Inputting the second data into the optimized target classification model to obtain a third predicted label output by the target classification model; Calculating the difference between the third predicted label and the second label to determine the credibility of the second data; Sorting the multiple second data in descending order of credibility, and obtaining one or more target second data whose order is before the order threshold in the sorting; Updating the original first data pairs and second data by using the target second data and the second labels corresponding to the target second data; Based on the updated first data pairs and second data, calculating the loss of the target classification model, updating the target classification model, and returning to the step of predicting the label of the second data by using the updated target classification model, and performing multiple iterations until the iteration stop condition is met to obtain a finally optimized target classification model.

7. A weight-aware distillation device for semi-supervised 3D semantic segmentation with class distribution mismatch, characterized in that, Including the following steps: An acquisition module, configured to acquire multiple first data pairs and multiple second data, wherein each first data pair includes first data and a first label of the first data; A processing module, configured to perform feature extraction on the first data and the second data respectively to obtain a first feature of the first data and a second feature of the second data; A first calculation module, configured to determine a point mutual information queue corresponding to any one of the second data based on the relationship between the second feature of any one of the second data and the first features of each first data, wherein the point mutual information queue includes multiple point mutual informations, and each point mutual information corresponds to one first data; A second calculation module, configured to determine the second label of any one of the second data and the weight of the second label according to the magnitudes of the point mutual informations in the point mutual information queue and the first labels of the first data corresponding thereto; A third calculation module, calculating the loss value of the initial classification model based on the multiple first data pairs, the multiple second data, the second labels of each second data, and the weights of each second label; An update module, configured to update parameters of the initial classification model according to the loss value to obtain an optimized target classification model.

8. An electronic device, characterized in that, Comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch as described in any one of claims 1-6.

10. A computer program product, characterized in that, Comprising a computer program, which when executed by a processor implements the weight-aware distillation method for semi-supervised 3D semantic segmentation with class distribution mismatch as described in any one of claims 1-6.

Citation Information

Patent Citations

  • National defense science and technology text structured semantic analysis method suitable for few-sample scene

    CN115994204A