A face entity matching method based on ensemble learning

Through the integrated learning CatBoost model and Bayesian optimization, hyperparameters are adjusted, combined with shape, area, direction and position similarity indicators, the surface entity matching is transformed into classification problems, solving the problem of threshold and weight setting, and improving matching accuracy and model transparency.

CN116955682BActive Publication Date: 2025-07-08Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210958616.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-07-08
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

The threshold or index weights in existing surface entity matching methods are difficult to determine, resulting in the problem of low matching accuracy.

Method used

Using an integrated learning-based approach, the CatBoost model is used to combine Bayesian optimization to adjust hyperparameters, transforming the face entity matching into classification problems through shape, area, direction and position similarity indicators, and introducing a Shapley addition interpretation framework to improve model transparency.

Benefits of technology

It effectively avoids the problem of threshold and weight setting, improves matching accuracy and model training efficiency, and enhances the generalization ability and interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116955682B_ABST
    Figure CN116955682B_ABST
Patent Text Reader

Abstract

The present invention relates to a surface entity matching method based on ensemble learning, belonging to the technical field of image matching. The present invention constructs an ensemble learning algorithm, the CatBoost model, as the surface entity matching model, and transforms the surface entity matching problem into a classification problem. The CatBoost model can use the similarity of each index as the feature of the surface entity pair to be matched, and based on this feature, classify the entity pairs to be matched into two categories: "matched" and "unmatched", so as to achieve the matching of surface entities. In this matching process, there is no need to set thresholds or index weights, effectively avoiding the problems of threshold and weight setting in conventional matching methods and improving the matching accuracy. At the same time, the Bayesian optimization method is used to adjust the hyperparameters of the CatBoost model, improving the model training efficiency and accuracy, and further improving the matching accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for matching surface entities based on ensemble learning, belonging to the technical field of image matching. Background Art

[0002] As an important type of geospatial data, the scale of vector spatial data is also constantly expanding. However, there are many inconsistencies among multi-source vector spatial data with different production methods and different currency, which brings great difficulties to the interconnection, interoperability and interoperation of data. Integrating and fusing multi-source heterogeneous vector spatial data is an important means to solve data inconsistency. As a key technology for integrating and fusing multi-source heterogeneous vector spatial data, homonymous entity matching aims to identify the same entity in different datasets and has important application value in aspects such as spatial database update, spatial query and change detection.

[0003] Surface entities account for a large proportion in vector spatial data and are one of the important carriers for expressing spatial information. Considering the differences in matching index types, surface entity matching methods are mainly divided into geometric matching, topological matching and semantic matching. Topological matching is more sensitive to the differences in topological relationships and is prone to matching failures. Semantic matching highly depends on the completeness of spatial data attribute information, so it can only be used as an auxiliary matching method. Geometric matching identifies homonymous entities by comparing the geometric feature similarities between entities and is the most commonly used surface entity matching algorithm. There are many main application methods, such as the surface object matching method based on spatial direction similarity and the multi-scale surface entity matching method using the area overlap degree as an evaluation index. However, the geometric matching results based on a single index often cannot meet the actual requirements in terms of reliability. Therefore, people more often consider comprehensively using multiple similarity indexes for surface entity matching. For example, an algorithm for calculating the total similarity of entities based on position, shape and area indexes; a multi-operator weighted matching method: based on the three indexes of area, position and geometric shape, the surface entity matching is completed by using the analytic hierarchy process. Comprehensively utilize the position, rotation angle and weighted similarity of associated edges between boundary points; a surface entity matching method based on boundary points of homonymous surface entities. A multi-scale surface entity matching method: starting from the edges and corners of the composed surface entities, four matching factors of displacement, stretching, rotation and generalization degree are extracted. Homonymous entity matching is realized based on two metric indexes of the Euclidean distance and the overlapping area between the center points of two surface entities. However, whether based on a single or multiple geometric features, the setting of thresholds or index weights in the matching process is always a difficult problem, and inappropriate setting of thresholds or index weights will affect the final matching result and reduce the matching accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for matching surface entities based on ensemble learning to solve the problem of low matching accuracy caused by the difficulty in determining thresholds or index weights in the current surface entity matching process.

[0005] The present invention provides a surface entity matching method based on ensemble learning to solve the above technical problems. The matching method includes the following steps:

[0006] 1) Obtain a surface entity dataset with known pairing information, select surface entity pairs from it as training entity pairs, calculate the index similarity of each entity pair as the feature of each entity pair, and construct training samples according to the features of each entity pair. The training samples include samples with successful matches and samples with unsuccessful matches;

[0007] 2) Input the constructed training samples into the CatBoost model, and use the Bayesian optimization method to adjust the hyperparameters of the CatBoost model to achieve the training of the CatBoost model;

[0008] 3) Obtain a set of surface entities to be matched, select surface entity pairs to be matched from it, calculate the index similarity of the selected surface entity pairs to be matched as the feature of the surface entity pairs, and input the obtained features into the trained CatBoost model to achieve the matching of surface entities.

[0009] The present invention constructs an ensemble learning algorithm CatBoost model as a surface entity matching model, transforms the surface entity matching problem into a classification problem. The CatBoost model can use each index similarity as the feature of the surface entity pair to be matched, and based on this feature, classify the surface entity pairs to be matched into two categories: "matched" and "unmatched", without setting thresholds or index weights, effectively avoiding the problems of threshold and weight setting in conventional matching methods, and improving the matching accuracy. At the same time, the Bayesian optimization method is used to adjust the hyperparameters of the CatBoost model, improving the model training efficiency and accuracy, and further improving the matching accuracy.

[0010] Further, the training samples constructed in step 1) are samples after unbalanced sample processing. Unbalanced sample processing refers to oversampling the samples with successful matches and randomly sampling the samples with unsuccessful matches.

[0011] The present invention adjusts the proportion of different category samples in the training samples by oversampling or random sampling, avoids the phenomenon of category imbalance caused by a large difference in the number of unmatched samples and matched samples in the training samples, and thus improves the generalization ability of the model.

[0012] Further, the index similarities calculated in step 1) and step 3) include at least two of the shape similarity, area similarity, direction similarity, and position similarity between at least two surface entities.

[0013] Further, the calculation process of the shape similarity between two surface entities is as follows:

[0014] a. Sample the outer contours of the two surface entities respectively. N sampling points are obtained for each surface entity, where N = 2 T , and T is a positive integer;

[0015] b. Take one of the sampling points P i as the starting point and divide the outer contour of the surface entity into several equal parts to obtain H equally divided points distributed on both sides of the starting point from far to near, realizing the multi-scale division of the outer contour of the surface entity;

[0016] c. Take the angles corresponding to the starting point at different scales as geometric features to obtain H geometric features of each surface entity. Perform discrete Fourier transform on each geometric feature, retain the first M Fourier low-frequency coefficients, and normalize using the modulus of the first low-frequency coefficient to obtain a normalized shape description matrix;

[0017] d. Calculate the shape difference degree between the two surface entities according to the shape description matrices of the two surface entities, and calculate the shape similarity degree between the two surface entities according to the obtained shape difference degree. The formula used is as follows:

[0018] sim shape = 1 - dis shape (A, B)

[0019]

[0020] where sim shape is the shape similarity degree between the two surface entities, and dis shape (A, B) is the shape difference degree between the two surface entities, are the elements in the normalized shape description matrices Z A and Z B of the two surface entities respectively.

[0021] Furthermore, the calculation formula for the area similarity degree between the two surface entities is:

[0022] sim area = min(S A , S B ) / max(S A , S B )

[0023] where sim area represents the area similarity degree between the two surface entities, and S A , S B are the areas of the two surface entities respectively.

[0024] Furthermore, the calculation formula for the direction similarity degree between the two surface entities is:

[0025]

[0026] where sim direc represents the directional similarity between two surface entities, and θ A , θ B are respectively the principal axis direction angles of the two surface entities. The principal axis direction angle refers to the included angle between the principal axis and the horizontal direction, and the principal axis refers to the major axis of the circumscribed rectangle with the minimum area of the surface entity.

[0027] Furthermore, the calculation formula for the position similarity between two surface entities is as follows:

[0028]

[0029] where sim posi represents the position similarity between two surface entities, centroid(A,B) represents the centroid distance between two surface entities, and dist max (A,B) represents the maximum distance value between any boundary points of the two surface entities.

[0030] Furthermore, in step 3), the minimum bounding rectangle (MBR) algorithm is used to select the pairs of surface entities to be matched from the set of surface entities to be matched. The MBR is the minimum bounding rectangle algorithm.

[0031] In order to avoid matching most unmatched surface entities and improve the matching efficiency, the present invention determines a matching candidate set from the set of surface entities to be matched by using the MBR algorithm (minimum bounding rectangle), and only uses the pairs of surface entities in the matching candidate set as the pairs of surface entities to be matched.

[0032] Furthermore, the process of adjusting the hyperparameters of the CatBoost model by the Bayesian optimization method in step 2) is as follows:

[0033] Use a Gaussian process as a surrogate model to approximately approximate the model evaluation function with an unknown structure; determine the sampling points most likely to obtain extreme values from the search space through a sampling function, and update the surrogate model according to the sampling point information; iteratively update to obtain a set of hyperparameters that satisfy the global optimum from the search space.

[0034] Furthermore, the method further includes the step of interpreting the CatBoost model using the Shapley additive explanation framework.

[0035] The present invention introduces the Shapley Additive exPlanation (SHAP) framework to quantify the influence of each input feature on the model classification result, improve the transparency of the CatBoost machine learning black box model, and solve the problems of complex structure and poor interpretability existing in the ensemble learning algorithm. Description of the Drawings

[0036] Figure 1 is a flowchart of the method for matching surface entities based on ensemble learning according to the present invention;

[0037] Figure 2-a is a schematic diagram of multi-scale division in the process of calculating shape similarity in an embodiment of the present invention;

[0038] Figure 2-b is a schematic diagram of constructing geometric features at a single scale in the process of calculating shape similarity in an embodiment of the present invention;

[0039] Figure 3-a is a schematic diagram of training samples for one-to-one matching in the experimental process of the present invention;

[0040] Figure 3-b is a schematic diagram of training samples for one-to-two matching in the experimental process of the present invention;

[0041] Figure 3-c is a schematic diagram of the first type of unmatched training samples in the experimental process of the present invention;

[0042] Figure 3-d is a schematic diagram of the second type of unmatched training samples in the experimental process of the present invention;

[0043] Figure 4 is the learning curve of the CatBoost model after hyperparameter optimization in the experimental process of the present invention;

[0044] Figure 5 is a schematic diagram of the test confusion matrix of the CatBoost model in the experimental process of the present invention;

[0045] Figure 6 is the SHAP summary plot of the CatBoost model in the experimental process of the present invention;

[0046] Figure 7 is the feature importance ranking based on the average absolute value of SHAP in the experimental process of the present invention;

[0047] Figure 8-a is the SHAP dependence graph of the direction similarity obtained in the experimental process of the present invention;

[0048] Figure 8-b is the SHAP dependence graph of the area similarity obtained in the experimental process of the present invention;

[0049] Figure 8-c is the SHAP dependence graph of the position similarity obtained in the experimental process of the present invention;

[0050] Figure 8-dIt is the SHAP dependence graph of the shape similarity obtained during the experiment of the present invention;

[0051] Figure 9 It is the prediction explanation graph of the negative example samples obtained during the experiment of the present invention;

[0052] Figure 10 It is the prediction explanation graph of the positive example samples obtained during the experiment of the present invention. Specific embodiments

[0053] The following further illustrates the specific embodiments of the present invention with reference to the accompanying drawings.

[0054] Method embodiment

[0055] The present invention transforms the face entity matching problem into a classification problem. The CatBoost classifier classifies the entity pairs to be matched into two categories of "matched" and "unmatched" according to four geometric similarity indicators of shape, area, direction, and position, effectively avoiding the problem of difficult threshold and weight setting in conventional matching methods. At the same time, aiming at the problems of complex structure and poor interpretability existing in the ensemble learning algorithm, the Shapley Additive exPlanation (SHAP) framework is introduced to quantify the influence of each input feature on the model classification result, improving the transparency of the CatBoost machine learning model. The implementation process of this method is as Figure 1 shown. The following details the specific steps of the entire implementation process.

[0056] 1. Establish a CatBoost model as the matching model.

[0057] CatBoost is an improved algorithm based on the Gradient Boosting Decision Tree (GBDT). GBDT uses the CART regression decision tree as the base learner, and serially trains a series of learners through the Boosting method, and accumulates the outputs of all learners as the final result. First, CatBoost introduces the Ordered boosting technique to solve the prediction shift phenomenon existing in GBDT. Second, CatBoost borrows the idea of online learning and introduces the Ordered Target Statistics (OrderedTS), for example, directly converting categorical features into numerical statistics, increasing the direct support for categorical features. Finally, CatBoost selects a symmetric binary tree as the base learner, and executes the same splitting criterion at each layer, so it is not easy to overfit, and the stability and prediction speed of the model are also significantly improved.

[0058] 2. Obtain a face entity dataset with known pairing information, select face entity pairs from it as training entity pairs, calculate the index similarity of each entity pair as the feature of each entity pair, and construct training samples based on the features of each entity pair.

[0059] Suppose the face entity dataset with known pairing information obtained in this embodiment is dataset D 1 and dataset D 2 , in dataset D 1 part of the face entities match the face entities in dataset D 2 . Before constructing training data, it is necessary to preprocess the face entity data in dataset D 1 and dataset D 2 by unifying the coordinate system, projection, and topological inspection. Select the matching face entity pairs from the preprocessed dataset D 1 and dataset D 2 as training entity pairs, and extract features from the training entity pairs. The present invention uses the index similarity between two face entities as the feature of the entity pair. The index similarity of the present invention includes shape similarity, area similarity, direction similarity, and position similarity. At least two of the selected similarities can be used as the extracted features. The calculation of each similarity is described below:

[0060] 1) Calculation of shape similarity

[0061] The present invention calculates the shape similarity of face entities based on the leaf recognition method. The implementation principle of this method is as Figure 2-a and Figure 2-b shown. First, obtain the external contour of the face entity. The external contour is sampled at equal distances to obtain N sampling points and satisfy N = 2 T (T is a positive integer, generally taken as 8), Figure 2-a represents a resampled face entity contour, which contains 256 sampling points in total; taking one side of any sampling point P i (represented by a five-pointed star) as an example, starting from P i , the entire contour is divided into three equal parts to obtain the first equal division point S6; then the sub-contour segment is divided into two equal parts to obtain the second equal division point S5; divide into two equal parts to obtain the third equal division point S4, and so on. After 1 time of three equal divisions and 5 times of two equal divisions, a total of 6 equal division points S6 to S1 can be obtained, and circles, stars, diamonds, crosses, triangles, and squares are used for highlighting respectively. Here, the distance between P i and each equal division point is called the scale; the equal division points are distributed on both sides of P i from far to near, and the distance from P iTogether, they achieve a multi-scale division of the contour. In this embodiment, 6 equal divisions are performed to obtain 6 equal division points. As other implementation manners, the number of equal divisions can also be adjusted according to actual situations. For example, it can be divided into 5 times, 7 times, 8 times, etc.

[0062] Then calculate P i The corresponding angle value θ at different scales is used as the geometric feature of each scale. As Figure 2-b shown, the angle between P i and the second equal division point S5 is used as the geometric feature of the second scale. Through the construction of such geometric features, the geometric features at 6 scales under each contour point can be calculated to obtain the shape description matrix Θ N×6 , where each column corresponds to one scale.

[0063] To eliminate the influence of the starting point P i position on the matrix Θ, a discrete Fourier transform is performed on each scale, the first M Fourier low-frequency coefficients are retained and normalized using the modulus of the first low-frequency coefficient to obtain the shape description matrix Z that satisfies rotation, translation, and scaling invariance M×6 .

[0064] Finally, according to the obtained shape description matrix Z M×6 calculate the shape difference degree dis shape (A, B) of the surface entities A and B, and calculate the shape similarity according to the shape difference degree. The specific calculation formula is as follows:

[0065] sim shape = 1 - dis shape (A, B)

[0066]

[0067] where sim shape is the shape similarity between two surface entities, dis shape (A, B) is the shape difference degree between two surface entities, are the elements in the normalized shape description matrices Z A and Z B of the two surface entities respectively.

[0068] 2) Calculation of area similarity

[0069] Area is one of the indicators for judging the similarity degree of two surface entities. The smaller the difference degree between entities, the closer their areas are. Let the areas of the surface entities A and B be represented by S A and S B respectively. Then the calculation formula for the area similarity of the two surface entities is:

[0070] sim area= min(S A , S B ) / max(S A , S B )

[0071] 3) Calculation of direction similarity

[0072] The direction similarity is represented by the similarity degree of the axial directions of the main axes of two face entities. First, calculate the main axis direction angle of the entity. In this application, the long axis of the minimum-area circumscribed rectangle of the face entity is set as the main axis, and the angle between the main axis and the horizontal direction is the main axis direction angle. Let θ A and θ B represent the main axis direction angles of face entities A and B respectively, and θ A ∈ [0, π), θ B ∈ [0, π). Then the calculation formula for the direction similarity sim direc is as follows:

[0073]

[0074] 4) Calculation of position similarity

[0075] The similarity degree of positions is calculated by the Euclidean distance between the centroids of two face entities. Use centroid(A, B) to represent the centroid distance between face entities A and B, and dist max (A, B) represents the maximum value of the distances between any boundary points of face entities A and B. Then the calculation formula for the position similarity is:

[0076]

[0077] Through the above process, the shape similarity, area similarity, direction similarity, and position similarity between two face entities can be obtained. Taking these four similarity values as the features between two face entities, the features of the training entity pairs can be obtained, and the training samples can be obtained.

[0078] Let two datasets D 1 , D 2 contain m and n face entities respectively. If the generation of candidate matching sets is not considered, one face entity in D 1 needs to be traversed and matched with each entity in D 2 . Ideally, D 2There is an entity that matches it. Therefore, the class imbalance between the non-matching and matching results is n - 1:1. According to the Probably Approximately Correct (PAC) learning theory, machine learning model training requires a large number of samples to ensure good generalization ability. In actual situations, due to the limitation of the scale of the dataset to be matched, while ensuring the total number of training samples, the number of non-matching samples and matching samples selected in the training samples will differ greatly, resulting in class imbalance. Unbalanced training samples often cause the classification model to focus too much on the majority class and ignore the minority class, thereby affecting the generalization ability of the model.

[0079] SMOTE over-sampling (Synthetic Minority Over-sampling Technique) generates new samples by randomly interpolating linearly between two minority class samples. Random under-sampling balances the class distribution by randomly deleting majority class samples until the desired class ratio between the majority class samples and the minority class samples is reached. The present invention uses a hybrid resampling method that combines under-sampling and over-sampling to process the training samples with class imbalance. Perform SMOTE over-sampling on the minority class (matching) samples, and then perform random under-sampling on the majority class (non-matching) samples to reduce the class imbalance of the original training samples, in order to ensure the model performance when the scale of the dataset to be matched is small.

[0080] 3. Input the obtained training samples into the constructed CatBoost model, and use the Bayesian optimization method to adjust the hyperparameters of the CatBoost model to achieve the training of the CatBoost model.

[0081] For a given training set D = (x k , y k ) k=1,2,...,n , where is an m-dimensional input feature, and y k is a corresponding numerical label. If the strong learner generated after the previous step of training is F t-1 , then the training objective for this round is to obtain a tree h t from the set of CART trees H such that the expectation E(·) of the loss function L(·) is minimized, that is:

[0082] h t = argmin E L(y, F t-1 (x) + h t (x)) (1)

[0083] where: (x, y) is a test sample independent of the training set. GBDT uses the negative gradient of the loss function to fit the CART decision tree h for this round of training t , and after T rounds of iteration, the final model F shown in Equation (2) is obtained, where F0 represents the initial weak learner, and a t represents the step size of the t-th round of training.

[0084]

[0085] Hyperparameter optimization is an important step in machine learning tasks. The CatBoost model has a complex structure and a large hyperparameter search space. The efficiency of using manual tuning methods, grid search methods, and random search methods is relatively low. This application selects the Bayesian optimization method for CatBoost hyperparameter optimization. The Bayesian optimization method is based on Bayes' theorem and uses the prior information of the previous search results to determine the next search point, so the optimization efficiency is higher. Bayesian optimization first uses Gaussian Processes (GP) as a surrogate model to approximately approximate the model evaluation function with an unknown structure, and then determines the next sampling point most likely to obtain an extreme value from the search space through a sampling function, and updates the surrogate model according to the sampling point information. After continuous iterative updates, a set of hyperparameters that satisfy the global optimum are finally obtained from the search space.

[0086] 4. Use the trained CatBoost model to match the surface entities to be matched.

[0087] Obtain the set of surface entities to be matched. Since there are generally many surface entities in the set of surface entities to be matched, and the number of unmatched surface entity pairs is much larger than the number of matched surface entity pairs, that is, if any two surface entities are selected for matching, the probability of successful matching is relatively low. To avoid matching most of the unmatched surface entities and improve the matching efficiency, the present invention first uses the MBR algorithm (minimum bounding rectangle) to determine the matching candidate set from the set of surface entities to be matched, and only performs subsequent matching on the surface entity pairs in the matching candidate set.

[0088] First, calculate the index similarity of the surface entity pairs in the matching candidate set as the feature of the surface entity pair. The index similarity includes shape similarity, area similarity, direction similarity, and position similarity. The specific calculation process is the same as the calculation methods of each similarity in step 2, and will not be specifically described here. Then, input the obtained features into the trained CatBoost model to achieve the matching of surface entities.

[0089] When using the trained CatBoost model for face entity matching, the present invention also uses the SHAP explanation framework to interpret and analyze the CatBoost model. The SHAP explanation framework is the Shapley Additive exPlanation (SHAP) framework. As an ex-post explanation framework, this explanation framework draws on the idea of cooperative game theory and interprets the prediction value of a machine learning model as the sum of the contributions of each input feature. SHAP fits and explains a certain prediction model f through a simple additive model g, that is

[0090]

[0091] To further illustrate the effect of the present invention, a simulation experiment is first conducted on the face entity matching method based on ensemble learning of the present invention. The data selected for this experiment comes from the Qinghai-Tibet Plateau lake dataset and OSM data respectively, containing 230 and 992 face entities respectively. The Qinghai-Tibet Plateau lake dataset is provided by the National Tibetan Plateau Data Center of China and is obtained by vectorizing satellite images.

[0092] This experiment is implemented using Python 3.7. First, the two groups of data are preprocessed. Next, 50 pairs of matching entities and 500 pairs of non-matching entities are randomly selected from the two datasets as the original training samples. The training samples are only used for model training and do not overlap with the subsequent samples to be matched. Some training samples are as Figure 3-a 、 Figure 3-b 、 Figure 3-c and Figure 3-d shown. Figure 3-a is a one-to-one matching sample in Figure 3-b is a one-to-two matching sample Figure 3-c and Figure 3-d are both non-matching samples. The class imbalance of the original training samples reaches 10. After processing the original training samples using the hybrid resampling technique, the number of matching entities and non-matching entities in the samples is 200 pairs and 400 pairs respectively, and the class imbalance is significantly reduced. The processed training samples are divided into two parts: 80% of the samples are used as the training set, and the remaining 20% are used as the test set.

[0093] The CatBoost model has many hyperparameters. In actual operation, it is extremely difficult to optimize all hyperparameters simultaneously. Therefore, this experiment selects 6 hyperparameters that have a greater impact on the model performance to generate a search space, selects the AUC value, that is, the area under the ROC curve, as the model evaluation function, and realizes the Bayesian optimization of each hyperparameter in the search space based on stratified 10-fold cross-validation. Based on the hyperparameter optimization results, the learning curve of the model is drawn (such as Figure 4As shown, it can be seen that as the number of model training samples increases, the training set score and the validation set score gradually converge. The trained model is tested on the test set, Figure 5 The confusion matrix of

[0094] indicates that the model has achieved good test results, with good fitting effect, strong learning ability and excellent generalization ability. CatBoost is a tree ensemble model. Therefore, in this experiment, the tree model explainer TreeExplainer in SHAP is selected to globally and locally explain the CatBoost matching model to analyze how the model makes predictions.

[0095] First, use SHAP Summary Plots to globally explain how each input feature affects model prediction. Figure 6 Each row in the vertical axis direction in Figure 7 represents a feature, which is sorted in descending order of feature importance from top to bottom. The horizontal axis represents the SHAP value. Each point in the figure represents a feature of a sample. The darker the color, the larger the feature value of the point. When the points coincide, they will be stacked vertically.

[0096] It shows that the feature that has the greatest impact on model prediction is position, followed by area and shape in turn, and the feature with the least impact is direction. Figure 7 Quantify the feature importance in Figure 7 The order of the quantified feature importance in Figure 6 is consistent with

[0097] . The feature ranked first is position, and its importance is about 1.65 times that of area, about 3.63 times that of shape, and about 4.06 times that of direction. The importance of the two features of shape and direction ranks third and fourth respectively, and their influence on the model prediction value is relatively close. In addition, it can also be seen from the SHAP summary plot that there may be a correlation between the size of the feature value and its contribution to model prediction. For example, for the position feature, from the color distribution of the points, it can be seen that the smaller the feature value, the smaller the corresponding SHAP value, and the larger the feature value, the larger the corresponding SHAP value. Next, with the help of the SHAP dependence plot, the relationship between each feature and its corresponding SHAP value is analyzed in more detail. The vertical axis of the SHAP dependence plot represents the SHAP value, and the horizontal axis represents the size of the feature value. As Figure 8-a , Figure 8-b , Figure 8-c and Figure 8-dAs shown, it can be seen that when the direction similarity is in the range of about 0.1 to 0.9, the SHAP value is less than 0, showing a negative contribution to the model prediction value. When the similarity is greater than 0.8, the similarity value is positively correlated with the SHAP value. When the area similarity is less than 0.1, it mainly shows a negative contribution to the model prediction value. When the similarity value is greater than 0.1, it shows a positive contribution to the model prediction value. When the similarity value is less than 0.3, the positive correlation between the similarity value and the SHAP value is strong. After the similarity value is greater than 0.3, the monotonicity of the SHAP value is not obvious. The position similarity is positively correlated with the SHAP value. Generally speaking, the SHAP value increases monotonically. When the position similarity is less than 0.5, it shows a negative contribution to the model prediction value. When the position similarity is greater than 0.5, it shows a positive contribution to the model prediction value. When the shape similarity is less than about 0.84, its contribution to the model prediction value is negative. When the similarity is greater than about 0.84, its contribution to the model prediction value is positive. After the similarity value reaches about 0.9, the SHAP value no longer continues to increase and remains relatively stable.

[0098] Next, from the perspective of a single sample, it is explained how the CatBoost model makes predictions. A positive example sample and a negative example sample are selected, corresponding to the two situations of "matched" and "unmatched" respectively. Figure 9 The prediction process of the negative example sample is explained. In the figure, f(x) represents the SHAP value output by the model, and the base value represents the benchmark value, which is determined by the model itself. Red represents that this feature value will increase the probability that the sample is predicted as a positive example sample, and blue represents that this feature value will decrease the probability that the sample is predicted as a positive example sample. Figure 9 In it, the value of the direction similarity is 0.943, which plays a positive role in the model prediction value. However, its positive promoting effect is much smaller than the total negative inhibiting effect of the other three features on the model. The finally output SHAP value is -8.86, which is less than the benchmark value of -4.076. This sample is predicted as a negative example sample, that is, "unmatched". Figure 10 In it, all four features play a positive promoting role in the model prediction. The finally output SHAP value is 8.80, which is much greater than the benchmark value of -4.076. This sample is predicted as a positive example sample, that is, "matched".

[0099] As can be seen from the above process, the present invention can transform the surface entity matching problem into a classification problem by constructing a CatBoost model. The CatBoost model classifies the entity pairs to be matched into two categories of "matched" and "unmatched" according to four geometric similarity indexes of shape, area, direction and position, effectively avoiding the problems of threshold and weight setting in conventional matching methods. At the same time, aiming at the problems of complex structure and poor interpretability existing in the ensemble learning algorithm, the Shapley Additive exPlanation (SHAP) framework is introduced to quantify the influence of each input feature on the model classification result, improving the transparency of the CatBoost machine learning black box model.

Claims

1. A surface entity matching method based on ensemble learning, characterized in that The matching method includes the following steps: 1) Obtain a face entity dataset with known pairing information, select face entity pairs from it as training entity pairs, calculate the index similarity of each entity pair as the feature of each entity pair, and construct training samples according to the features of each entity pair. The training samples include samples with successful matches and samples with unsuccessful matches; 2) Input the constructed training samples into the CatBoost model, and use the Bayesian optimization method to adjust the hyperparameters of the CatBoost model to achieve the training of the CatBoost model; 3) Obtain a set of face entities to be matched, select face entity pairs to be matched from it, calculate the index similarity of the selected face entity pairs to be matched as the feature of the face entity pair, and input the obtained features into the trained CatBoost model to achieve the matching of face entities.

2. The method for matching surface entities based on ensemble learning according to claim 1, wherein The training samples constructed in step 1) are samples after imbalanced sample processing. Imbalanced sample processing refers to oversampling the samples with successful matches and randomly sampling the samples with unsuccessful matches.

3. The method for face entity matching based on ensemble learning according to claim 1, wherein The index similarities calculated in step 1) and step 3) include at least two of the shape similarity, area similarity, direction similarity, and position similarity between at least two face entities.

4. The method for face entity matching based on ensemble learning according to claim 3, wherein The calculation process of the shape similarity between two face entities is as follows: a. Sample the outer contours of the two surface entities respectively, and N sampling points are obtained for each surface entity, where N = 2 T , where T is a positive integer; b. Starting from one of the sampling points P i perform several equal division processes on the outer contour of the surface entity to obtain H equally divided points distributed on both sides of the starting point from far to near, thereby realizing the multi-scale division of the outer contour of the surface entity; c. Use the angles corresponding to the starting points at different scales as geometric features to obtain H geometric features of each face entity. Perform discrete Fourier transform on each geometric feature, retain the first M Fourier low-frequency coefficients, and normalize using the modulus of the first low-frequency coefficient to obtain a normalized shape description matrix; d. Calculate the shape difference degree between two face entities according to the shape description matrices of the two face entities, and calculate the shape similarity between the two face entities according to the obtained shape difference degree. The formula used is as follows: sim shape = 1 - dis shape (A,B) where sim shape is the shape similarity between two surface entities, and dis shape (A, B) is the shape difference degree between two surface entities, are the normalized shape description matrices Z A and Z B in the elements respectively.

5. The method for face entity matching based on ensemble learning according to claim 3, wherein The calculation formula for the area similarity between two face entities is: sim area = min(S A , S B ) / max(S A , S B ) where sim area represents the area similarity between two surface entities, S A , S B and S are the areas of the two surface entities respectively.

6. The method for matching surface entities based on ensemble learning according to claim 3, wherein The calculation formula for the direction similarity between two face entities is: where sim direc represents the directional similarity between two surface entities, and θ A , θ B are respectively the principal axis direction angles of the two surface entities. The principal axis direction angle refers to the angle between the principal axis and the horizontal direction, and the principal axis refers to the major axis of the circumscribed rectangle with the smallest area of the surface entity.

7. The face entity matching method based on ensemble learning according to claim 3, wherein The calculation formula for the position similarity between two face entities is: where sim posi represents the positional similarity between two surface entities, centroid(A,B) represents the centroid distance between two surface entities, and dist max (A,B) represents the maximum distance value between any boundary points of two surface entities.

8. The method for surface entity matching based on ensemble learning according to any one of claims 1-7, characterized in that In step 3), the MBR (Minimum Bounding Rectangle algorithm) is used to select face entity pairs to be matched from the set of face entities to be matched.

9. The method for matching surface entities based on ensemble learning according to claim 8, wherein The process of adjusting the hyperparameters of the CatBoost model by the Bayesian optimization method in step 2) is as follows: Use the Gaussian process as a surrogate model to approximately approximate the model evaluation function with unknown structure; determine the sampling points most likely to obtain extreme values from the search space through the sampling function, and update the surrogate model according to the sampling point information; iterate and update to obtain a set of hyperparameters that satisfy the global optimum from the search space.

10. The method for matching surface entities based on ensemble learning according to claim 8, wherein The method also includes the step of interpreting the CatBoost model using the Shapley additive explanation framework.

Citation Information

Patent Citations

  • Different-period spatial entity hierarchical matching method and system based on multi-source information

    CN104317793A

  • Hybrid entity matching to drive program execution

    US20200210466A1