Chinese short text-oriented semi-supervised place name data labeling method
Through the semi-supervised place name data annotation method for short Chinese texts, using large models and manual review, the problems of insufficient classification and low labeling quality of Chinese place name data sets in the prior art are solved, and efficient and high-quality place name data annotation is achieved.
Patent Information
- Application Number
- CN202411895010.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the Chinese place name data set lacks detailed classification, and the manual labeling data consumes a large resource and is of low quality, which cannot meet the labeling needs of multi-grained place names in social media text.
The semi-supervised place name data annotation method for short Chinese texts is adopted. By acquiring and preprocessing social media text data, a semantic annotation system and a priori knowledge are established, and a large model is used for fine-tuning and incremental fine-tuning, and combined with manual review, high-quality annotation data are screened out.
It improves the quality and efficiency of Chinese place name data labeling, reduces the direct participation of manual labeling, and meets the labeling needs of multi-grained place names in social media texts.
Smart Images

Figure CN120068866A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a semi-supervised place name data annotation method for Chinese short texts. Background Art
[0002] Chinese place name datasets and text data annotation methods are the data support and foundation for many research tasks in the cross-field of natural language processing and geographic information systems. Existing publicly available datasets are all general named entity recognition datasets, and there is no detailed classification for place name entities. They are only represented and annotated with a single label, which has certain limitations for Chinese place name recognition tasks.
[0003] In the prior art, publicly available annotation datasets for Chinese place names and geographical entities are relatively lacking. Manually annotating data consumes a large amount of resources, and the existing datasets assign a single label type to place name data, which cannot meet the annotation data requirements for multi-granularity place names in place name recognition and parsing tasks for social media texts, resulting in low-quality annotation data.
[0004] Therefore, there is an urgent need to provide a solution to improve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a semi-supervised place name data annotation method for Chinese short texts to improve the problem of low-quality annotation data in the prior art.
[0006] A semi-supervised place name data annotation method for Chinese short texts provided by the present invention includes the following steps:
[0007] Obtain and preprocess the text data of social media to generate multiple place name classification types, establish a semantic annotation system based on the place name classification types, and construct multiple annotation prior knowledges based on the semantic annotation system and place name characteristics;
[0008] Fine-tune and initialize a large model based on the annotation prior knowledge and manually annotated data to obtain an initialized large model;
[0009] Incrementally fine-tune the initialized large model based on unannotated data to obtain multiple fine-tuned large models, calculate the comprehensive metric values for each of the fine-tuned large models, and select the fine-tuned large model corresponding to the largest comprehensive metric value as the annotation model;
[0010] Input the unannotated data that has not participated in fine-tuning into the annotation model to obtain multiple candidate annotation data, and group the multiple candidate annotation data to obtain multiple grouped candidate verification sets;
[0011] Randomly select the to-be-screened labeled data in each group of the to-be-verified set for manual review. Discard the groups of the to-be-verified set that fail the review based on the review accuracy threshold, and add the groups of the to-be-verified set that pass the review to the labeled result set until the number of groups of the to-be-verified set in the labeled result set reaches the task requirement.
[0012] The beneficial effects of a semi-supervised place name data annotation method for Chinese short texts provided by the present invention are as follows: The present invention analyzes the data characteristics and place name features of Chinese translation texts on social media, refers to previous classification standards and annotation systems, and reclassifies the geographical elements of place names therein. According to this classification and combined with the actual task requirements, a Chinese place name semantic annotation system including Chinese place name entity semantic annotation and Chinese place name ambiguous entity annotation is constructed. On this basis, annotation prior knowledge is constructed in combination with the expression characteristics of Chinese place names in the text. Based on the Chinese place name semantic annotation system and annotation prior knowledge, a small amount of data is manually annotated as a hint to fine-tune the large model, and the fine-tuned large model is used to annotate the social media text data. Finally, manual review is carried out by means of random sampling in batches and groups, which reduces the direct participation of manual annotation while ensuring the quality of the annotated data.
[0013] Optionally, the process of obtaining and preprocessing the text data of social media includes:
[0014] Remove irrelevant characters from the text data, and the irrelevant characters include: emoticons, "#" symbols, URL links, and "@" mentions;
[0015] Based on the text length threshold and delimiter punctuation, segment the text data to obtain the preprocessed text data.
[0016] Optionally, when constructing multiple annotation prior knowledge, the annotation prior knowledge includes: four types of annotations: POI type annotation, LocForModification type annotation, OtherObject type annotation, and annotation in special cases. The POI type annotation consists of an annotation judged by an annotation window and an annotation of a geographical common noun that refers to a certain place in the text semantics. The LocForModification type annotation consists of an LFM type semantic annotation and an LFM type institutional name prefix place name annotation. The annotation in special cases includes: the to-be-annotated word is in a symbol, the to-be-annotated name has the same name but different levels of place name types, the same type of annotated words appear in combination, and the to-be-annotated word is expressed by a combination of multiple places.
[0017] Optionally, when the annotation determined by the annotation window occurs in a to-be-annotated word with a composition structure of "continent, country or administrative division place name + suspected POI location" in the text content, the to-be-annotated word needs to be wholly or partially annotated through a judgment criterion, and the judgment criterion is: based on a preposition, the first half of the continent, country or administrative division place name in the to-be-annotated word is annotated as a place name semantic annotation type, and the suspected POI location is annotated as a POI type.
[0018] Optionally, the LFM type semantic annotation includes: if the to-be-annotated word and the location or geographical location referred to in the semantics are in the same region or scope, annotate according to the location type to which the to-be-annotated word belongs; if the to-be-annotated word and the location or geographical location referred to in the semantics are not in the same region or scope, annotate the to-be-annotated word as the LFM type; if there is no referred location or geographical location in the semantics, directly annotate the to-be-annotated word as the LFM type.
[0019] Optionally, the LFM type prefix place name annotation for an organization name includes:
[0020] Obtain the organization name in the text data. When the prefix of the organization name contains the to-be-annotated place name and there is no referred location and geographical location in the text semantics, annotate the to-be-annotated word as the LFM type.
[0021] Optionally, the annotation of the OtherObject type includes: if the to-be-annotated word has the meaning of a location or geographical location, annotate the to-be-annotated word as the corresponding location type; if the to-be-annotated word does not have the meaning of a location or geographical location and exists in the form of an entity object in the semantics, annotate it as the OO type.
[0022] Optionally, when calculating the comprehensive metric value for each of the fine-tuned large models, the expression of the comprehensive metric value is:
[0023]
[0024] where TP represents the number of correctly annotated place names, FP represents the number of wrongly annotated place names, FN represents the number of non-missed annotated place names, P represents precision, R represents recall, and F 1 is the comprehensive metric value.
[0025] Optionally, when initializing the fine-tuning of the large model based on the annotation prior knowledge and the manually annotated data, it includes the process of adjusting the text output format of the large model, where the text output format is defined as: "{loc_name: the place name of the text, class: the specific type in the Chinese place name semantic annotation type corresponding to the place name}".
[0026] Optionally, the process of obtaining the annotation model includes:
[0027] Manually annotate multiple text data to construct an evaluation test set for the large model's performance, and determine the accuracy rate, recall rate, and comprehensive metric value as the performance evaluation indicators for the large model;
[0028] Group the unannotated data to construct an unannotated data set. Use the initialized large model to annotate a group of data in the unannotated data set, and add the annotation results to the initialized fine-tuning data to form the fine-tuning data set 1;
[0029] Use the fine-tuning data set 1 as the input to fine-tune the initial large model. Calculate and record the performance evaluation indicators using the large model weights and model files. At the same time, annotate a group of data in the unannotated data set, and add the annotation results to the fine-tuning data set 1 to form the fine-tuning data set 2;
[0030] Repeat the above steps to calculate the performance evaluation indicators for the initial large model after each fine-tuning. Subsequently, take a group of unannotated data from the unannotated data set for annotation, use the currently fine-tuned large model for annotation, and add the annotation results to the fine-tuning data set i (i = 1, 2, 3,... n) to form the fine-tuning data set i+1 for the next fine-tuning;
[0031] Comprehensively evaluate the performance indicators of the fine-tuned large model under different data volumes obtained from the above steps, and select the large model with the highest comprehensive metric value as the annotation model. Description of the Drawings
[0032] Figure 1 It is a flowchart of a semi-supervised place name data annotation method for Chinese short texts provided by the present invention;
[0033] Figure 2 It is a flowchart of the fine-tuned large model provided by the present invention;
[0034] Figure 3 It is a flowchart of text data annotation and manual random review provided by the present invention;
[0035] Figure 4 It is the specified data format diagram provided by the present invention. Detailed Embodiments
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention pertains. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.
[0037] An embodiment of the present invention provides a semi-supervised place name data annotation method for Chinese short texts. Refer to Figure 1 , which includes the following steps:
[0038] S1. Obtain and preprocess the text data of social media to generate multiple place name classification types, establish a semantic annotation system based on the place name classification types, and construct multiple annotation prior knowledges based on the semantic annotation system and place name characteristics;
[0039] S2. Fine-tune and initialize a large model based on the annotation prior knowledge and manually annotated data to obtain an initialized large model;
[0040] S3. Perform incremental fine-tuning on the initialized large model based on unannotated data to obtain multiple fine-tuned large models, calculate the comprehensive metric values for each of the fine-tuned large models, and select the fine-tuned large model corresponding to the largest comprehensive metric value as the annotation model;
[0041] S4. Input the unannotated data that has not participated in the fine-tuning into the annotation model to obtain multiple to-be-screened annotation data, and group the multiple to-be-screened annotation data to obtain multiple grouped to-be-verified sets;
[0042] S5. Randomly select the to-be-screened annotation data in each grouped to-be-verified set for manual review, discard the grouped to-be-verified sets that fail the review based on the review accuracy threshold, and add the grouped to-be-verified sets that pass the review to the annotation result set until the number of grouped to-be-verified sets in the annotation result set reaches the task requirement.
[0043] In some embodiments, in the process of performing step S1 to obtain and preprocess the text data of social media, it includes:
[0044] S1-1. Remove the irrelevant characters in the text data, where the irrelevant characters include: emoticons, "#" symbols, URL links, and "@" mentions;
[0045] S1-2. Segment the text data based on the text length threshold and delimiter punctuation marks to obtain the preprocessed text data.
[0046] Specifically, in the process of performing step S1-1 to remove the "#" symbol, the "#" symbol in the text is used to refer to the topic label, and the content following the "#" is the content of the topic label, which has a certain connection with the text semantics. Therefore, when processing, different from before, instead of directly filtering it out, the "#" symbol is coarsened and the content of the topic label is retained.
[0047] Furthermore, when performing step S1-2, the text data is short text data, the text length threshold is 128. If it exceeds this value, the text is segmented from the second-to-last delimiter punctuation mark (including the period at the end of the sentence); the delimiter punctuation marks include: ",", ";", ".", "……", etc.
[0048] In some embodiments, when performing step S1 to construct multiple annotation prior knowledges, the annotation prior knowledges include: annotations of POI types, annotations of LocForModification types, annotations of OtherObject types, and annotations in special cases. The annotations of POI types are composed of annotations judged by the annotation window and annotations of geographical common nouns that refer to a certain location in the text semantics. The annotations of LocForModification types are composed of LFM type semantic annotations and LFM type institutional name prefix place name annotations. The annotations in special cases include: the word to be annotated is in a symbol, the to-be-annotated name has the same name but different levels of place name types, the same type of annotated words appear in combination, and the word to be annotated is expressed by a combination of multiple locations.
[0049] Specifically, the annotations of geographical common nouns that refer to a certain location in the text semantics include geographical common nouns such as schools, bars, parks, houses, etc., and their characteristic is to refer to a certain POI location through the text semantics.
[0050] For further explanation, the following are examples:
[0051] Such as "a bar in Los Angeles", "a dock near the port", "a lawn was laid in the school". By analyzing the text semantics, it can be seen that "bar", "port", "dock", and "school" are all used to refer to a certain POI location in the text semantics. Therefore, the above 4 words are all annotated as POI types.
[0052] Further, when the annotation determined by the annotation window occurs in a to-be-annotated word with a composition structure of "continent, country or administrative division place name + suspected POI place" in the text content, the to-be-annotated word needs to be wholly or partially annotated through a judgment criterion, and the judgment criterion is: based on a preposition, the first half of the continent, country or administrative division place name in the to-be-annotated word is annotated as a place name semantic annotation type, and the suspected POI place is annotated as a POI type.
[0053] For further explanation, the following are examples:
[0054] For example, for "Catherine Palace in Russia", through analyzing the word composition, it can be judged that the to-be-annotated word has a composition pattern of "continent, country or administrative division place name + suspected POI place name", where "Russia" is a country place name and "Catherine Palace" is a suspected POI place name. According to the above prior knowledge, the result after adding the preposition "of" between "Russia" and "Catherine Palace" is "Catherine Palace of Russia". At this time, "Catherine Palace" can still accurately refer to the POI place. Therefore, "Russia" in the to-be-annotated word is annotated as its corresponding Country (countries and regions at the national level) type, and at the same time, "Catherine Palace" is annotated as a POI type.
[0055] Specifically, the LFM type semantic annotation includes: if the to-be-annotated word and the place or geographical location referred to in the semantics are in the same region or scope, it is annotated according to the place type to which the to-be-annotated word belongs; if the to-be-annotated word and the place or geographical location referred to in the semantics are not in the same region or scope, the to-be-annotated word is annotated as the LFM type; if there is no place or geographical location referred to in the semantics, the to-be-annotated word is directly annotated as the LFM type.
[0056] For further explanation, the following are examples:
[0057] Example 1: "On the ring road in the UK, Indian protesters carried out large-scale protests". Through text semantic analysis, it can be known that the place referred to in the semantics is "the ring road in the UK". The to-be-annotated words "UK" and "ring road" are consistent with the place mentioned in the semantics, so they are respectively annotated as their corresponding Country type and S&R (pipeline route) type. However, for the to-be-annotated word "India", "Indian protesters" are located on "the ring road in the UK" rather than in "India", so it is not in the same region or scope as the geographical location "ring road" referred to in the semantics, and thus is annotated as the LFM type;
[0058] Example 2: "Qatar is one of the richest Gulf countries". Through text semantic analysis, it can be seen that there is no referred location or geographical location in the semantics. Therefore, the to-be-annotated word "Gulf" is directly annotated as the LFM type. At the same time, another to-be-annotated word "Qatar" exists in the semantics in the form of a national entity and does not have the meaning of a location or geographical location. Therefore, it is annotated as the OO type.
[0059] Specifically, the annotation of the prefix geographical name of the LFM type institution name includes:
[0060] Obtain the institution name in the text data. When the prefix of the institution name contains the to-be-annotated geographical name and there is no referred location and geographical location in the text semantics, annotate the to-be-annotated word as the LFM type.
[0061] For further explanation, the following examples are given:
[0062] For example, "Understand how the National Academy of Artificial Intelligence of the United States has promoted the development of artificial intelligence for more than 50 years". Through text semantic analysis, it can be seen that "National Academy of Artificial Intelligence of the United States" is the institution name. Therefore, only analyze the leading to-be-annotated geographical name "United States" among them. At the same time, there is no referred location or geographical location in the semantics. Therefore, directly annotate the to-be-annotated word "United States" as the LFM type.
[0063] Specifically, the annotation of the OtherObject type includes: If the to-be-annotated word has the meaning of a location or geographical location, annotate the to-be-annotated word as the corresponding location type; if the to-be-annotated word does not have the meaning of a location or geographical location and exists in the semantics in the form of an entity object, annotate it as the OO type.
[0064] For further explanation, the following examples are given:
[0065] Example 1: "Emergency food aid supplies provided by China to Sudan". Through text semantic analysis, it can be seen that the to-be-annotated words "China" and "Sudan" exist in the semantics in the form of national entity objects and do not have the meaning of a location or geographical location. Therefore, both of them are annotated as the OO type.
[0066] Example 2: "András Simor, the former central bank governor who guided Hungary through the financial crisis, said that Hungary will vigorously develop the real economy". Through text semantic analysis, it can be seen that the to-be-annotated words "Hungary" and "Hungary" exist in the semantics in the form of national entity objects and do not have the meaning of a location or geographical location. Therefore, both of them are annotated as the OO type.
[0067] The annotation in special cases includes:
[0068] 1) The to-be-annotated word is in "()".
[0069] Specifically, in such cases, it is necessary to analyze the semantics of the text to determine whether the word to be annotated is a supplementary explanation of the place name or geographical location before the parentheses. If the word to be annotated is used as a supplementary explanation of the place name or geographical location before the parentheses, it shall be annotated together with the word to be annotated and the whole parentheses according to the type of place where the place name or geographical location before the parentheses belongs; otherwise, the word to be annotated within the parentheses shall be annotated separately.
[0070] For further explanation, the following examples are given:
[0071] For example: "Antonov Airlines / BE Brave Like Antonov A Flight 124100 ADB5502, at an altitude of 8250 feet, is flying south over Erdmansheim (Germany) in Saxony."
[0072] Through text semantic analysis, it can be seen that the word to be annotated "Germany" is a supplementary explanation of the place "Erdmansheim" before the parentheses. Therefore, the place name "Erdmansheim" before the parentheses and the word to be annotated are annotated as a whole "Erdmansheim (Germany)" according to the type of place where "Erdmansheim" belongs, which is of the City (regions at the level of municipalities directly under the central government) type.
[0073] 2) The word to be annotated is within "《》"
[0074] Specifically, in such cases, directly ignore the word to be annotated within "《》" and do not annotate it. For example, words such as "slaughterhouse", "shelter", and "Yokosuka" in "From the Slaughterhouse to the Shelter" and "Yokosuka Story" are not annotated.
[0075] 3) There are place name types with the same name but different levels for the word to be annotated
[0076] Specifically, for the annotation of such cases, by default, annotate at the highest level of the word to be annotated. For example: "Okinawa" in Japan has both "Okinawa Prefecture" and "Okinawa City", but in Japan, the administrative division level of "prefecture" is higher than that of "city". Therefore, if "Okinawa" appears, it is by default annotated as the County type of administrative division at the county level.
[0077] 4) Combinations of annotation words of the same type appear
[0078] Specifically, for the annotation of such cases, it generally occurs when words to be annotated of the same place type are connected by conjunctions or punctuation marks. In such cases, these words to be annotated and the conjunctions or prepositions shall be annotated as a whole as the place type they belong to. For further explanation, the following examples are given:
[0079] "In China and Russia", where the words to be annotated are "China" and "Russia", and both belong to the same place type and the conjunction connecting them is "and", so "China and Russia" is annotated as a whole as the Country type.
[0080] 5) The word to be annotated is expressed by a combination of multiple locations of different types.
[0081] Specifically, for the annotation of this type of situation, the location to be annotated is composed of multiple locations of multiple types in a parallel relationship. It is necessary to perform semantic analysis on the text to determine whether the above combined expression is used to represent locations of the same location type. If it represents locations of the same type, it is annotated as the location type to which the location belongs; otherwise, multiple locations of multiple types are respectively annotated according to their respective location types. For further explanation, the following examples are given:
[0082] For example, "international waters off the coasts of Mexico, Central America, and South America". Through semantic analysis of the text, it can be seen that the above text contains a location of the Country type, "Mexico", locations of the Continent (continent, region at the continent level) type, "Central America and South America", and a location of the Nature (physical geography) type, "international waters". The multiple locations of multiple types above are combined and expressed in a parallel relationship as locations of the same type, that is, the "international waters" of the Nature type. Therefore, "international waters off the coasts of Mexico, Central America, and South America" is overall annotated as the Nature type.
[0083] Furthermore, based on the constructed Chinese place name semantic annotation system and annotation prior knowledge, fine-tune the prior knowledge. Clearly define the place name semantic annotation types, precautions for related type annotations, and related type annotation examples, and combine them together in the form of questions (connecting questions and answers with "***").
[0084] For further explanation, the following examples are given:
[0085] 1) For clarifying the entity category of place names: "What are the Chinese place name semantic annotation types? ***Continent: continent, region at the continent level; Country: country, region at the country level; Province: first-level administrative division, province, state, autonomous region, etc."
[0086] 2) For the precautions for related type annotations: "What are the specifications and precautions during annotation? ***The first specification for annotation is as follows: Annotation of the OO type: The annotation of the OO type requires judging whether the word to be annotated has the meaning of a location or geographical position according to semantics. If the word to be annotated has the meaning of a location or geographical position, then the word to be annotated is annotated as the corresponding location type; otherwise, it is annotated as OO.
[0087] Furthermore, the process of constructing the place name semantic annotation system includes: classifying the common place name types in social media text through analysis of text data to obtain place name classifications, as shown in Table 1.
[0088] Table 1 Local Name Classification Table of Social Media Short Texts
[0089]
[0090]
[0091] Based on the classification of place names, a Chinese place name semantic annotation system including Chinese place name entity semantic annotation and Chinese ambiguous entity annotation is constructed, as shown in Table 2 and Table 3.
[0092] Table 2 Chinese Place Name Entity Semantic Annotation Types
[0093]
[0094]
[0097] 3) Annotation examples for relevant entity types: "Are there examples, supplementary explanations, or other specifications? *** For example, 'Atami Mi Jamal is possible. It does not have the meaning of a location or geographical position, so both are labeled as OO.'
[0098] In some embodiments, when performing step S2 to fine-tune and initialize the large model based on the annotation prior knowledge and manual annotation data, it includes a process of adjusting the text output format of the large model, where the text output format is defined as: "{loc_name: the place name of the text, class: the specific type in the Chinese place name semantic annotation type corresponding to the place name}".
[0099] Furthermore, the example prompt enables the large model to learn the output format of the current task through the output format definition and relevant examples, as follows:[[]]
[0100] "Are there examples, supplementary explanations, or other specifications? *** For example: {"text": "In response to the request of the Kingdom of Tonga for volcanic disaster relief, the Chinese military dispatched the Wuzhishan Ship and the Chaganhu Ship to form a maritime transport formation to Tonga to carry out the task of transporting relief supplies. The Chinese naval ship formation transporting relief supplies to Tonga recently arrived in Nuku'alofa, the capital of Tonga.", "label": "{loc_name: Tonga, class: OO}, {loc_name: China, class: OO}, {loc_name: Tonga, class: Country}, {loc_name: Tonga, class: OO}, {loc_name: Tonga, class: Country}, {loc_name: Nuku'alofa, class: City}"}."
[0101] In some embodiments, when performing step S2 to initialize the fine-tuning of the large model, it is necessary to preprocess the input data according to the format specified by the large model, and the specified data format is as Figure 4 shown.
[0102] Figure 4 The input format shown in is for a single piece of data. Here, id is usually used to represent the round number of the current conversation, conversations represents the content of the conversation, and at the same time is the main content of the fine-tuning data. from and value respectively represent the object of the conversation and the specific content of the conversation. It can be understood that the first value is the question and the second value is the answer. The fine-tuning operation of the large model is implemented in this specific conversation form.
[0103] The prepared fine-tuning prior knowledge is split by "***", and the results before and after the split are assigned to the two values in turn, and other content is filled in according to the above-specified format to realize the construction of the fine-tuning prior knowledge, as shown in Appendix 1.
[0104] In some embodiments, in the process of performing step S3 to screen out the fine-tuned large model with the largest comprehensive metric value as the annotation model, it includes:
[0105] S3-1. Manually annotate multiple pieces of text data to construct a performance test set for evaluating the large model, and determine accuracy, recall, and the comprehensive metric value as the performance evaluation indicators of the large model;
[0106] S3-2. Group the unannotated data to construct an unannotated data set, use the initialized large model to annotate a group of data in the unannotated data set, and add the annotation results to the initialized fine-tuning data to form the fine-tuning data set 1;
[0107] S3-3. Use the fine-tuning data set 1 as the input to fine-tune the initial large model, calculate and record the performance evaluation indicators using the large model weights and model files, and at the same time annotate a group of data in the unannotated data set, and add the annotation results to the fine-tuning data set 1 to form the fine-tuning data set 2;
[0108] S3-4. Repeat the above steps S3-1, S3-2, and S3-3 to calculate the performance evaluation indicators for the initial large model after each fine-tuning, and then take a group of unannotated data from the unannotated data set for annotation, use the currently fine-tuned large model for annotation, and add the annotation results to the fine-tuning data set i (i = 1, 2, 3,... n) to form the fine-tuning data set i+1 for the next fine-tuning;
[0109] S3-5. Comprehensively evaluate the performance indicators of the fine-tuned large models under different data volumes obtained in the above steps, and select the large model with the highest comprehensive metric value as the annotation model.
[0110] Refer to Figure 2 , after starting, first initialize and construct a fine-tuning dataset, and fine-tune the large model to obtain a fine-tuned large model. On the one hand, first construct an unlabeled dataset, take a set of unlabeled data for data annotation to obtain a set of labeled data. On the other hand, evaluate the performance of the fine-tuned large model based on the test set, record the performance metrics of the fine-tuned large model with the current data volume, and determine whether the fine-tuning data volume reaches the maximum value. If so, select the large model with the best performance metrics and end the program. Otherwise, integrate a set of labeled data with the current fine-tuning dataset, reconstruct the fine-tuning dataset, and fine-tune the large model again.
[0111] Specifically, when fine-tuning the initial large model, the object of each fine-tuning is the original large model, and what changes continuously is the number of fine-tuning.
[0112] In some embodiments, when performing step S3 to calculate the comprehensive metric value for each of the fine-tuned large models, the expression of the comprehensive metric value is:
[0113]
[0114] Among them, FP represents the number of correctly labeled place names, FP represents the number of wrongly labeled place names, FN represents the number of place names not missed in labeling, P represents precision, R represents recall rate, and F 1 is the comprehensive metric value.
[0115] Specifically, precision is the ratio of the number of correctly labeled place names to the total number of place names labeled by the model, recall rate is the ratio of the number of correctly labeled place names to the actual number of all place names in the text, and the comprehensive metric value is the harmonic mean of precision and recall rate, which is used to comprehensively reflect the performance of the model.
[0116] In some embodiments, when performing step S4, the grouping is performed in groups of every 1000 data.
[0117] In some embodiments, when performing step S5, refer to Figure 3 , the process of manual review is as follows:
[0118] S5-1. Determine the data volume of the labeled result set required for the task;
[0119] S5-2. Take 1000 data from the text dataset, use the annotation model in step 3 for data annotation, and construct a set of data to be verified for annotation;
[0120] S5-3. Divide the set of data to be verified for annotation into 10 groups, and then randomly select 10% of the data in each group for manual review;
[0121] S5-4. If the manual review accuracy rate of this group of labeled data is greater than or equal to 80%, add it to the labeled result set; otherwise, discard it.
[0122] S5-5. Determine whether the quantity in the labeled result set is consistent with the required quantity for the determined task. If it is consistent, end the process and obtain the labeled result set with the required quantity for the task; if it is not consistent, repeat steps S5-2, S5-3, and S5-4 until the quantity of the labeled result set reaches the required quantity for the task.
[0123] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes are all within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein may have other embodiments and can be implemented or realized in various ways.
Claims
1. A semi-supervised place name data annotation method for Chinese short texts, characterized in that: include: Acquire and pre-process text data from social media to generate multiple place name classification types, establish a semantic annotation system based on the place name classification types, and construct multiple annotation prior knowledge based on the semantic annotation system and place name characteristics; Fine-tune and initialize the large model based on the annotated prior knowledge and the manually annotated data to obtain an initialized large model; Incrementally fine-tune the initialized large model based on the unlabeled data to obtain multiple fine-tuned large models, and calculate a comprehensive metric value for each of the fine-tuned large models, and select the fine-tuned large model corresponding to the maximum comprehensive metric value as the labeled model; Inputting unlabeled data that has not participated in fine-tuning into the labeling model to obtain a plurality of labeled data to be screened, and grouping the plurality of labeled data to be screened to obtain a plurality of grouped sets to be verified; Randomly select the labeled data to be screened in each group to be verified set for manual review. Based on the review accuracy threshold, the group to be verified sets that fail the review are discarded, and the group to be verified sets that pass the review are added to the annotation result set until the number of group to be verified sets in the annotation result set reaches the task requirement.
2. According to the semi-supervised place name data labeling method for Chinese short texts according to claim 1, it is characterized in that: The process of obtaining and preprocessing text data from social media includes: Remove irrelevant characters from the text data, including emojis, "#" symbols, URL links, and "@" mentions; The text data is segmented based on the text length threshold and the separation punctuation to obtain the preprocessed text data.
3. According to the semi-supervised place name data labeling method for Chinese short texts according to claim 1, it is characterized in that: When constructing multiple annotation prior knowledge, the annotation prior knowledge includes four categories: POI type annotation, LocForModification type annotation, OtherObject type annotation, and annotation in special cases. The POI type annotation consists of annotations judged by the annotation window and geographical general nouns and annotations that refer to a certain place in the text semantics. The LocForModification type annotation consists of LFM type semantic annotations and LFM type institution name prefix place name annotations. The annotations in special cases include: the word to be annotated is in the symbol, the name to be annotated is the same but the place name type of different levels, the same type of annotation words appear in combination, and the word to be annotated is expressed by a combination of multiple places.
4. According to the semi-supervised place name data labeling method for Chinese short texts according to claim 3, it is characterized in that: The annotation determined by the annotation window occurs in the text content of the word to be annotated whose composition structure is "continent, country or administrative division name + suspected POI location". The word to be annotated needs to be annotated as a whole or partially according to the judgment standard. The judgment standard is: based on the preposition, the first half of the continent, country or administrative division name in the word to be annotated is marked as the place name semantic annotation type, and the suspected POI location is marked as the POI type.
5. According to the semi-supervised place name data labeling method for Chinese short texts according to claim 3, it is characterized in that: The LFM type semantic annotation includes: if the word to be annotated and the place or geographical location referred to in the semantics are in the same area or range, annotating according to the place type to which the word to be annotated belongs; if the word to be annotated and the place or geographical location referred to in the semantics are not in the same area or range, annotating the word to be annotated as the LFM type; if there is no place or geographical location referred to in the semantics, directly annotating the word to be annotated as the LFM type.
6. According to the semi-supervised place name data labeling method for Chinese short texts according to claim 3, it is characterized in that: The LFM type institution name prefix place name annotation includes: The organization name in the text data is obtained. When the prefix of the organization name contains the place name to be annotated and there is no place or geographical location referred to in the text semantics, the word to be annotated is annotated as the LFM type.
7. A semi-supervised place name data labeling method for Chinese short text according to claim 3, characterized in that: The annotation of the OtherObject type includes: if the word to be annotated has the meaning of place or geographical location, annotating the word to be annotated as the corresponding place type; if the word to be annotated does not have the meaning of place or geographical location and exists in the form of an entity object in semantics, annotating it as OO type.
8. According to the semi-supervised place name data labeling method for Chinese short texts according to claim 1, it is characterized in that: When calculating the comprehensive metric value for each fine-tuned large model, the expression of the comprehensive metric value is: Among them, the TP table represents the number of correctly labeled place names, the FP table represents the number of incorrectly labeled place names, the FN table represents the number of place names that are not missed, P represents precision, R represents recall, and F1 is a comprehensive measurement value.
9. A semi-supervised place name data labeling method for Chinese short text according to claim 1, characterized in that: When fine-tuning and initializing the large model based on the annotation prior knowledge and manually annotated data, it includes a process of adjusting the text output format of the large model, wherein the text output format is defined as: "{loc_name: place name in the text, class: the specific type of the place name corresponding to the Chinese place name semantic annotation type}".
10. The semi-supervised place name data labeling method for Chinese short text according to claim 1, characterized in that: The process of obtaining the annotation model includes: Manually annotate multiple text data to build a test set for evaluating the performance of the large model, and determine the accuracy, recall rate and comprehensive measurement values as the performance evaluation indicators of the large model; The unlabeled data are grouped to construct an unlabeled data set, a group of data in the unlabeled data set is labeled using the initialization large model, and the labeled results are added to the initialization fine-tuning data to form a fine-tuning data set 1; Use fine-tuning dataset 1 as input to fine-tune the initial large model, use the large model weights and model files to calculate and record performance evaluation indicators, and label a set of data in the unlabeled dataset. Add the labeling results to fine-tuning dataset 1 to form fine-tuning dataset 2; Repeat the above steps, calculate the performance evaluation index of the initial large model after each fine-tuning, then take a group of unlabeled data from the unlabeled dataset for labeling, use the current fine-tuned large model for labeling, and add the labeling results to the fine-tuning dataset i (i = 1, 2, 3, ..., n) table to form the fine-tuning dataset i+1 for the next fine-tuning; Comprehensively evaluate the performance indicators of the fine-tuned large model under different data amounts obtained in the above steps, and select the large model with the highest comprehensive metric value as the annotation model.
Citation Information
Cited By
Place name address translation system and matching method based on artificial intelligence and special name library
CN120278165A