A POI data classification method based on semantic enhancement
By transforming the POI data classification problem into a binary classification problem and capturing semantic relationships using RoBERTa's cross-coding method, the problem of uneven distribution of labels in POI data categories is solved, and the effect of improving classification accuracy under a small amount of data is achieved.
Patent Information
- Application Number
- CN202310972101.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-03
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-08-03
AI Technical Summary
The existing POI data classification methods ignore the explicit semantic information provided by the POI data category label, resulting in low classification accuracy and cannot be applied to the learning needs of a small number of samples.
By introducing POI data category tags, the disassembly method is used to transform multi-classification problems into binary classification problems, and the semantic relationship between POI data category tags and names is captured using RoBERTa's cross-coding method, expand the data set and adjust the model hyperparameters, and train the POI data classifier.
It effectively solves the problem of uneven distribution of labels in POI data categories, improves classification accuracy, and can train new categories of POI data classifiers under a small number of data samples.
Smart Images

Figure CN117033632B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a POI data classification method based on semantic enhancement. Background Art
[0002] Points of Interest (POI) are a term used in geographic information systems (GIS) to refer to any geographical object that can be abstracted as a point and are closely related to people's lives. For example, people often use POI data categories to search for businesses that meet their needs on apps like Meituan and Ele.me. However, with the rapid development of cities, a large amount of new POI data emerges daily. Meeting the demand for personalized POI recommendation services, alleviating the information overload faced by users, and achieving automated classification of POI thematic categories have become key and challenging areas in POI data analysis and mining research.
[0003] Currently, POI data classification methods fall into two main categories: content-based classification and feature-based classification. The former combines review information and POI names for classification, but its scope of application is very limited. The latter requires word segmentation and feature-based extraction of POI names, making it impossible to classify POI names that lack feature-based classification. Both of these methods ignore the explicit semantic information provided by the POI data's category labels, rely too heavily on keywords in the POI data, and ignore the semantic information of the POI names themselves. Consequently, they have low accuracy and are unsuitable for continued learning on a small number of samples.
[0004] Based on this, a POI data classification method based on semantic enhancement is proposed, which provides a new technical solution to solve the classification problem of POI data. Summary of the Invention
[0005] The purpose of the present invention is to provide a POI data classification method based on semantic enhancement to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: a POI data classification method based on semantic enhancement, comprising the following steps:
[0007] Step 1: Collect POI data and clean it;
[0008] Step 2: Filter POI data and build a POI dataset;
[0009] Step 3: Introduce POI category labels and construct a binary form of POI name and category, transforming the POI data classification problem into the relationship between POI name and category, thereby transforming the multi-classification problem into a binary classification problem;
[0010] Step 4: Expand the dataset and input POI names and their categories into RoBERTa. Using RoBERTa's cross-coding method, we capture the word-to-word interaction between labels and POI names and convert the relationship between them into a vector representation.
[0011] Step 5: The relationship vector between the two is mapped into a one-dimensional vector containing two values through a linear layer, and normalized with the Softmax function. The true or false prediction result is output according to the probability.
[0012] Step 6: Adjust the optimal hyperparameters of the model to obtain the optimal model for POI data classification, input the POI data into the trained optimal model, and classify the POI data.
[0013] Furthermore, the step 1 of collecting POI data and cleaning it includes the following steps:
[0014] Use crawler tools to crawl POI data from Baidu Maps;
[0015] Delete special symbols contained in the captured POI data and further delete duplicate data.
[0016] Furthermore, the screening of POI data in step 2 and the construction of a POI dataset include the following steps:
[0017] The collected POI data were analyzed and sorted, and M different categories of POI data were screened out;
[0018] The POI data of each different category are put into the training set, validation set and test set in a certain proportion to construct the POI dataset.
[0019] Furthermore, the POI category label is introduced in step 3 to construct a binary form of the POI name and its category, which is defined as:
[0020] <T i , L i >
[0021] Where T i , L i Respectively represent the i-th POI name and its corresponding category;
[0022] In step 3, the POI data classification problem is converted into the relationship between its name and its category, thereby converting the multi-classification problem into a binary classification problem, which includes the following steps:
[0023] For any POI name T k , category label L i , the function is defined as follows:
[0024]
[0025] Where i represents the i-th category label, k represents the k-th POI name in the training set; if T k Belongs to category label L i , the value of f is 1, otherwise it is 0.
[0026] Furthermore, expanding the data set in step 4 includes the following steps:
[0027] For a classification task with M categories, fill N bigrams for each sample. If the i-th input T i The label is the i-th category L i , then there is ( <T i ,L i >, True) as a training positive sample, there are (N-1) negative samples ( <T k , L i >, False), where k≠i;
[0028] In step 4, the POI name and its category are input into RoBERTa as follows:
[0029] [[CLS],T i ,[SEP],L i ,[SEP]]
[0030] "[CLS]" indicates the beginning of a sentence, and "[SEP]" indicates the end of a sentence or the division of two sentences;
[0031] In step 4, the cross-coding method based on RoBERTa captures the word-to-word interaction information between the label and the POI name, and converts the relationship between the two into a vector representation:
[0032] R=[β1,β2,…,β n ]
[0033] Where n is the number of embedded tokens, β n The calculation process is as follows:
[0034] In order to obtain the embedding vector of the nth token, it is necessary to derive it from the initialization vectors x1, x2, and x3;
[0035] Use x1, x2, x3 and three transformation matrices W respectively q 、W k 、W v Multiply them to get q, k, v, where
[0036] Multiply the query matrix q by the transpose of the matching matrix k to get α 11 , α 12 , α 13 ;
[0037] To α 11 , α 12 , α 13 Do Softmax normalization and get The calculation formula is as follows:
[0038]
[0039] use Multiplying by vector v gives β n , the calculation formula is as follows:
[0040]
[0041] Furthermore, in step 5, the relationship vector between the two is mapped into a one-dimensional vector containing two values through a linear layer:
[0042] γ=RW=[γ1,γ2]
[0043] Where E is the weight matrix;
[0044] In step 5, the normalization is performed in combination with the Softmax function:
[0045]
[0046] Furthermore, the optimal hyperparameters for adjusting the model in step 6 are:
[0047] Set the learning rate range to [10 -7 ,1], the batch values are set to 8, 16 and 32 respectively, and the number of negative samples is set to 5, 10, 15 and 20 respectively for discussion;
[0048] The screening of the optimal model in step 6 uses macro-P, macro-R and macro-F1 to measure the overall performance of the multi-classification task model. i , then the corresponding accuracy recall Rate and The value is calculated as follows:
[0049]
[0050] In the formula Indicates tag i The number of correct classifications, Misclassify other types of tags as tags i The number of Indicates that the tag i The number of labels that are incorrectly identified as other categories is macro-P, macro-R, and macro-F1, and the calculation formula is as follows:
[0051]
[0052] Furthermore, the special symbols include “ / ”, “#” and space.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] The present invention introduces POI data category labels and adopts a disassembly method to introduce several negative classes for each input positive class, thereby converting the multi-classification problem into a binary classification problem. Then, the semantic relationship between the POI data category labels and their names is captured based on the cross-coding method of RoBERTa. This effectively solves the problem of uneven distribution of POI data category labels and allows new classes of POI data classifiers to be trained using only a small number of data samples without retraining. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Schematic diagram of the POI data classification model structure of the present invention;
[0056] Figure 2 This is an example display diagram of positive and negative samples of the present invention. DETAILED DESCRIPTION
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0058] Example 1:
[0059] A POI data classification method based on semantic enhancement, using PyCharm as the development platform and Python as the development language, the method structure is as follows Figure 1 As shown; including the following steps:
[0060] Step 1: Collect POI data and clean it;
[0061] In this embodiment, the following steps are specifically included:
[0062] Use crawler tools to crawl POI data from Baidu Maps;
[0063] Delete special symbols contained in the captured POI data and further delete duplicate data;
[0064] Special symbols include " / ", "#" and space.
[0065] Step 2: Filter POI data and build a POI dataset;
[0066] Here, the specific steps include:
[0067] The collected POI data were analyzed and sorted, and M different categories of POI data were screened out;
[0068] The POI data of each different category are put into the training set, validation set and test set in a certain proportion to construct the POI dataset.
[0069] Step 3: Introduce POI category labels and construct a binary form of POI name and category, transforming the POI data classification problem into the relationship between POI name and category, thereby transforming the multi-classification problem into a binary classification problem;
[0070] Among them, the POI category label is introduced, and the binary form of the POI name and its category is defined as:
[0071] <T i , L i >
[0072] Where T i , L i Respectively represent the i-th POI name and its corresponding category;
[0073] Here, converting the POI data classification problem into the relationship between its name and its category, thereby converting the multi-classification problem into a binary classification problem, includes the following steps:
[0074] For any POI name T k , category label L i , the function is defined as follows:
[0075]
[0076] Where i represents the i-th category label, k represents the k-th POI name in the training set; if T k Belongs to category label L i , the value of f is 1, otherwise it is 0.
[0077] Step 4: Expand the dataset and input POI names and their categories into RoBERTa. Using RoBERTa's cross-coding method, we capture the word-to-word interaction between labels and POI names and convert the relationship between them into a vector representation.
[0078] In this embodiment, expanding the data set specifically includes the following steps:
[0079] For a classification task with M categories, fill N bigrams for each sample. If the i-th input T i The label is the i-th category L i , then there is ( <T i , L i >, True) as a training positive sample, there are (N-1) negative samples ( <T k , L i >, False), where k≠i;
[0080] Among them, the POI name and its category are input into RoBERTa as follows:
[0081] [[CLS],T i ,[SEP],L i ,[SEP]]
[0082] "[CLS]" indicates the beginning of a sentence, and "[SEP]" indicates the end of a sentence or the division of two sentences;
[0083] The cross-coding method based on RoBERTa captures the word-to-word interaction information between the label and the POI name, and converts the relationship between the two into a vector representation:
[0084] R=[β1,β2,…,β n ]
[0085] Where n is the number of embedded tokens, β n The calculation process is as follows:
[0086] In order to obtain the embedding vector of the nth token, it is necessary to derive it from the initialization vectors x1, x2, and x3;
[0087] Use x1, x2, x3 and three transformation matrices W respectively q 、W k 、W v Multiply them to get q, k, v, where
[0088] Multiply the query matrix q by the transpose of the matching matrix k to get α 11 , α 12 , α 13 ;
[0089] To α 11 , α 12 , α 13 Do Softmax normalization and get The calculation formula is as follows:
[0090]
[0091] use Multiplying by vector v gives β n , the calculation formula is as follows:
[0092]
[0093] Step 5: The relationship vector between the two is mapped into a one-dimensional vector containing two values through a linear layer, and normalized with the Softmax function. The true or false prediction result is output according to the probability.
[0094] Here, the relationship vector between the two is mapped into a one-dimensional vector containing two values through a linear layer:
[0095] γ=RW=[γ1,γ2]
[0096] Where W is the weight matrix;
[0097] Among them, combined with the Softmax function for normalization:
[0098]
[0099] Step 6: Adjust the optimal hyperparameters of the model to obtain the optimal model for POI data classification, input the POI data into the trained optimal model, and classify the POI data;
[0100] In this example, the optimal hyperparameters of the adjustment model are:
[0101] Set the learning rate range to [10 -7 ,1], the batch values are set to 8, 16 and 32 respectively, and the number of negative samples is set to 5, 10, 15 and 20 respectively for discussion;
[0102] Among them, the screening of the optimal model uses macro-P, macro-R and macro-F1 to measure the overall performance of the multi-classification task model. For the category label tag in the multi-classification task i , then the corresponding accuracy recall Rate and The value is calculated as follows:
[0103]
[0104] In the formula Indicates tag i The number of correct classifications, Misclassify other types of tags as tags i The number of Indicates that the tag i The number of labels that are incorrectly identified as other categories is macro-P, macro-R, and macro-F1, and the calculation formula is as follows:
[0105]
[0106] In summary, the present invention provides a POI data classification method based on semantic enhancement. By introducing POI data category labels, a decomposition method is used to introduce several negative classes for each input positive class, thereby converting the multi-classification problem into a binary classification problem. Then, the semantic relationship between the POI data category labels and their names is captured based on the cross-coding method of RoBERTa. This can effectively solve the problem of uneven distribution of POI data category labels. Thus, without retraining, it is allowed to use only a small number of data samples to train new classes of POI data classifiers.
[0107] Example 2:
[0108] Please refer to Figure 1 and Figure 2 , a POI data classification method based on semantic enhancement, using PyCharm as the development platform and Python as the development language, the method structure is as follows Figure 1 As shown;
[0109] The following steps are involved:
[0110] Step 1: Collect POI data and clean it.
[0111] First, the captured area is segmented using rectangles. Then, the segmented rectangular areas are traversed in the web crawler, and requests and web page access are constructed based on the Baidu Map API, and the captured POI data is saved to the local computer.
[0112] Furthermore, duplicate POI data are deleted, and special symbols contained therein, including “ / ”, “#”, spaces, etc., are deleted to complete the cleaning of POI data.
[0113] Step 2: Filter POI data and build a POI dataset.
[0114] The present invention performs semantic recognition based on the POI classification of Baidu Map. Due to the complexity of POI semantic recognition, it will not perform very detailed semantic classification like Amap, Baidu, etc. The selection of POI categories follows the following two principles: (1) POI names belonging to multiple categories are not subdivided, and only the primary category to which they belong is identified. (2) The phenomenon of multi-semantic intersection is treated specifically. Here, the food category in Baidu Map is used as an example to illustrate. For the categories such as "Chinese restaurant" and "fast food restaurant" in the secondary classification of food, since the keywords "Chinese restaurant" and "fast food restaurant" rarely appear in their names, they are uniformly classified as the primary category of food. For "cafe" and "cake dessert" that contain obvious keywords, they are subdivided. By analyzing and sorting the collected POI data in Hangzhou, a total of 75 different categories of POI data were screened out, including 139,471 POI names. The specific classification is shown in Table 1. The screened data are divided into three parts: training set, test set and validation set, with the proportions of 60%, 20% and 20% respectively.
[0115] Table 1 POI data categories
[0116]
[0117]
[0118] Step 3: Introduce POI category labels and construct a binary form of POI name and its category, converting the POI data classification problem into the relationship between POI name and its category, thereby converting the multi-classification problem into a binary classification problem.
[0119] Specifically, the POI data category label is introduced to provide display semantic information and construct a tuple of POI name and its category. <T i , L i >, thereby converting the multi-classification problem of POI data into a binary classification problem between its category labels and their names.
[0120] Furthermore, for this problem, a function approximation method is defined. For any POI name T k , category label L i , the function is defined as follows:
[0121]
[0122] Where i represents the i-th category label, and k represents the k-th POI name in the training set. k Belongs to category label L i , the value of f is 1, otherwise it is 0.
[0123] Step 4: Expand the dataset and input POI names and their categories into RoBERTa. Based on RoBERTa's cross-coding method, the interaction information between labels and POI names, such as words and characters, is captured and the relationship between the two is converted into a vector representation.
[0124] First, expand the dataset, enhance the generalization ability of the model and learn the classification principles of POI data. For a classification task containing M categories, fill N bigrams for each sample text. If the i-th input T i The label is the i-th category L i , then there is ( <T i , L i >, True) as a training positive sample, there are (N-1) negative samples ( <T k , L i >, False), where k≠i. Figure 2 As shown in the figure, taking two negative samples as an example, the label of the positive sample is true, otherwise it is false.
[0125] Secondly, construct the input format of RoBERTa, use the sentence start symbol "[CLS]", the end symbol or the separator symbol "[SEP]" of two sentences, and concatenate the POI name and its category into [[CLS], T i ,[SEP],L i , [SEP]].
[0126] Finally, the cross-coding method based on RoBERTa captures the interactive information between the label and the POI name, such as words and characters, and converts the relationship between the two into a vector representation:
[0127] R=[β1,β2,…,β n ]
[0128] Where n is the number of embedded tokens, β n The calculation process is as follows:
[0129] (1) To obtain the embedding vector of the nth token, it is necessary to derive it starting from the initialization vectors x1, x2, and x3;
[0130] (2) Use x1, x2, and x3 to transform the three matrices W q 、W k 、W v Multiply them to get q, k, v, where
[0131] (3) Multiply the query matrix q by the transpose of the matching matrix k to obtain α 11 , α 12 , α 13 ;
[0132] (4) To α 11 , α 12 , α 13 Do Softmax normalization and get The calculation formula is as follows:
[0133]
[0134] (5) Use Multiplying by vector v gives β n , the calculation formula is as follows:
[0135]
[0136] Step 5: Pass the relationship vector between the two through a linear layer to map it into a one-dimensional vector containing two values. At the same time, it is normalized with the Softmax function and the true or false prediction result is output according to the probability.
[0137] First, use the relationship vector R obtained in step 4 to multiply the weight vector W of the linear layer to obtain a two-dimensional vector:
[0138] A=RW=[A1,A2]
[0139] Then, the Softmax function is used for normalization:
[0140]
[0141] Step 6: Adjust the optimal hyperparameters of the model to obtain the optimal model for POI data classification, input the POI data into the trained optimal model, and classify the POI data.
[0142] First, the range of learning rate is set to [10 -7 ,1], the batch values are set to 8, 16 and 32 respectively, and the number of negative samples is set to 5, 10, 15 and 20 respectively for discussion. Through the change curves of learning rate and loss function value under different parameters, it is found that when the number of negative samples is 5, the batch value and learning rate should be set to 32 and 0.001 respectively; when the number of negative samples is 10, the batch value and learning rate should be set to 16 and 0.001 respectively; when the number of negative samples is 15, the batch value and learning rate should be set to 16 and 0.01 respectively; when the number of negative samples is 20, the batch value and learning rate should be set to 8 and 0.001 respectively.
[0143] After determining the optimal hyperparameters of the model under different negative sample values, start training the model. Then, use macro-P, macro-R and macro-F1 to measure the overall performance of the multi-classification task model. i , then the corresponding accuracy recall Rate and The value is calculated as follows:
[0144]
[0145] In the formula Indicates tag i The number of correct classifications, Misclassify other types of tags as tags i The number of Indicates that the tag i The number of labels that are incorrectly identified as other categories is macro-P, macro-R, and macro-F1, and the calculation formula is as follows:
[0146]
[0147] When the number of negative samples is 5, 10, 15, and 20, the corresponding maximum macro-F1 values on the test set are 0.8317, 0.8334, 0.84, and 0.8389, respectively. These results indicate that setting the number of negative samples to 15 is a good choice. Finally, the optimal model is used to classify the POI data.
[0148] In order to illustrate the POI data classification effect of the present invention, under the same conditions, the multi-classification method of RoBERTa-RNN was compared with the binary classification method of the present invention using the same training set, test set and validation set. The predicted results are as follows: (1) The macro-F1 value of the RoBERTa-RNN multi-classification method without introducing POI data category labels on the test set is 0.827, while the macro-F1 value of the method proposed in the present invention on the test set is 0.840. The results show that the macro-F1 value of the present method is improved by 1.3%; (2) For POI data containing 95 entrances and exits, the F1 value of the RoBERTa-RNN method for entrances and exits is 0, and the characteristics of this type of POI data are not learned at all. The F1 value of the method proposed in the present invention for entrances and exits is 0.984. The results show that the present method can well capture the explicit semantic information provided by the natural language category labels. At the same time, converting the binary classification problem into a multi-classification problem can effectively solve the problem of data category imbalance.
[0149] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0150] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A POI data classification method based on semantic enhancement, characterized in that: The following steps are involved: Step 1: Collect POI data and clean it; Step 2: Filter POI data and build a POI dataset; Step 3: Introduce POI category labels and construct a binary form of POI name and category, transforming the POI data classification problem into the relationship between POI name and category, thereby transforming the multi-classification problem into a binary classification problem; Step 4: Expand the dataset and input POI names and their categories into RoBERTa. Using RoBERTa's cross-coding method, we capture the word-to-word interaction between labels and POI names and convert the relationship between them into a vector representation. Step 5: The relationship vector between the two is mapped into a one-dimensional vector containing two values through a linear layer, and normalized with the Softmax function. The true or false prediction result is output according to the probability. Step 6: Adjust the optimal hyperparameters of the model to obtain the optimal model for POI data classification, input the POI data into the trained optimal model, and classify the POI data; Expanding the data set in step 4 includes the following steps: For a classification task with M categories, fill N bigrams for each sample. If the i-th input T i The label is the i-th category L i , then there is ( <T i ,L i >, True) as a training positive sample, there are N-1 negative samples ( <T k ,L i >, False), where k≠i; In step 4, the POI name and its category are input into RoBERTa as follows: [[CLS],T i ,[SEP],L i ,[SEP]] "[CLS]" indicates the beginning of a sentence, and "[SEP]" indicates the end of a sentence or the division of two sentences; In step 4, the cross-coding method based on RoBERTa captures the word-to-word interaction information between the label and the POI name, and converts the relationship between the two into a vector representation: R=[β1,β2,…,β n ] Where n is the number of embedded tokens, β n The calculation process is as follows: In order to obtain the embedding vector of the nth token, it is necessary to derive it from the initialization vectors x1, x2, and x3; Use x1, x2, x3 and three transformation matrices W respectively q 、W k 、W v Multiply them to get q, k, v, where Multiply the query matrix q by the transpose of the matching matrix k to get α 11 , α 12 , α 13 ; To α 11 , α 12 , α 13 Do Softmax normalization and get The calculation formula is as follows: use Multiplying by vector v gives β n , the calculation formula is as follows:
2. A POI data classification method based on semantic enhancement according to claim 1, characterized in that: Collecting POI data and cleaning it in step 1 includes the following steps: Use crawler tools to crawl POI data from Baidu Maps; Delete special symbols contained in the captured POI data and further delete duplicate data.
3. The POI data classification method based on semantic enhancement according to claim 1, characterized in that: The screening of POI data in step 2 and the construction of a POI dataset include the following steps: The collected POI data were analyzed and sorted, and M different categories of POI data were screened out; The POI data of each different category are put into the training set, validation set and test set in a certain proportion to construct the POI dataset.
4. The POI data classification method based on semantic enhancement according to claim 1, characterized in that: The POI category label is introduced in step 3, and the binary form of the POI name and its category is defined as: <T i ,L i > Where T i ,L i Respectively represent the i-th POI name and its corresponding category; In step 3, the POI data classification problem is converted into the relationship between its name and its category, thereby converting the multi-classification problem into a binary classification problem, which includes the following steps: For any POI name T k , category label L i , the function is defined as follows: Where i represents the i-th category label, k represents the k-th POI name in the training set; if T k Belongs to category label L i , the value of f is 1, otherwise it is 0.
5. The POI data classification method based on semantic enhancement according to claim 1, characterized in that: In step 5, the relationship vector between the two is passed through a linear layer and mapped into a one-dimensional vector containing two values: γ=RW=[γ1,γ2] Where W is the weight matrix; In step 5, the Softmax function is combined to perform normalization:
6. The POI data classification method based on semantic enhancement according to claim 1, characterized in that: The optimal hyperparameters for adjusting the model in step 6 are: Set the learning rate range to [10 -7 ,1], the batch values are set to 8, 16 and 32 respectively, and the number of negative samples is set to 5, 10, 15 and 20 respectively for discussion; The screening of the optimal model in step 6 uses macro-P, macro-R and macro-F1 to measure the overall performance of the multi-classification task model. i , then the corresponding accuracy recall Rate and The value is calculated as follows: In the formula Indicates tag i The number of correct classifications, Misclassify other types of tags as tags i The number of Indicates that the tag i The number of labels that are incorrectly identified as other categories is macro-P, macro-R, and macro-F1, and the calculation formula is as follows:
7. The POI data classification method based on semantic enhancement according to claim 2, characterized in that: The special symbols include " / ", "#" and space.
Citation Information
Patent Citations
Interest point classification method, device and equipment and storage medium
CN111767359A