An automatic association processing method based on open-source geographical name data
By applying deep learning technology to semantic coding and interactive fusion of place names in place name data processing, the problems of low efficiency and poor flexibility of traditional methods are solved, and more efficient and accurate place name data association is achieved.
Patent Information
- Application Number
- CN202510128179.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-05
AI Technical Summary
Traditional place name data processing methods rely on manual annotation or rule-based matching algorithms, which are inefficient and difficult to adapt to complex and changeable place name expressions, especially when dealing with natural language texts.
Semantic coding technology based on deep learning is used to perform semantic coding and compensatory interaction fusion of candidate place names and their supplementary contents, and the semantic feature expression of candidate place names is optimized by using supplementary content as the context background, and the place name alternative list is constructed by querying the geographic database to realize automatic association of place name data.
It improves the accuracy of place name data association, reduces dependence on manual annotation, and improves data processing efficiency.
Smart Images

Figure CN119558319B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of geographical name data processing, and more specifically, to an automatic association processing method based on open-source geographical name data. Background Art
[0002] Today, with the rapid development of informatization and digitalization, the application of Geographic Information System (GIS) has penetrated into every corner of social life. As an important part of the geographic information system, the processing and management of geographical name data play an irreplaceable role in many fields such as spatial data analysis, location-based services, navigation systems, and intelligent transportation. With the development of the Internet, the number of natural language text contents containing geographical name information (such as news reports, social media, travel blogs, etc.) has increased sharply. By accurately identifying geographical names and performing geographical association processing on these open-source geographical name data, the accuracy of positioning can be effectively improved. For example, an online map platform can provide more accurate destination guidance through the association of geographical name data; emergency rescue services can quickly locate the specific location corresponding to a geographical name through the association of geographical name data, improving rescue efficiency.
[0003] However, traditional geographical name data processing methods mainly rely on manual annotation or rule-based matching algorithms. Among them, although manual annotation has high accuracy, its efficiency is extremely low. When faced with a large amount of text data, the cost is huge and the time consumption is long. The rule-based matching algorithm is relatively efficient, but its flexibility is poor, and it is difficult to adapt to complex and changeable geographical name expression forms. Especially when processing natural language texts, since geographical names may appear in various forms, such as abbreviations, aliases, dialect expressions, etc., the rule-based matching algorithm often has difficulty in accurately identifying and associating with the correct geographical location.
[0004] Therefore, an optimized automatic association processing method based on open-source geographical name data is expected. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide an automatic association processing method based on open-source geographical name data. First, entity detection is performed on natural language text content containing geographical names to extract candidate geographical names, and other content in the text is used as a supplement. A semantic encoding technology based on deep learning is used to perform semantic encoding and compensatory interactive fusion on the candidate geographical names and their supplementary content, so as to use the supplementary content as the context background to optimize the semantic feature expression of the candidate geographical names. Then, by querying the associated entity data of the candidate geographical names in the geographical database, a geographical name alternative list is constructed, and the automatic association of geographical name data is realized based on the semantic similarity between each alternative geographical name in the list and the candidate geographical name, which can effectively improve the accuracy of geographical name data association, reduce the dependence on manual annotation at the same time, and improve data processing efficiency.
[0006] According to one aspect of the present application, an automatic association processing method based on open-source place name data is provided, which includes:
[0007] Obtain the natural language text content containing place names;
[0008] Perform entity detection on the natural language text content containing place names to extract candidate place names, and define the content other than the candidate place names in the natural language text content containing place names as place name supplementary context content;
[0009] Based on the place name supplementary context content, perform semantic compensation optimization on the candidate place names based on principal component analysis to obtain an optimized candidate place name semantic embedding coding vector;
[0010] Query the associated entity data of the candidate place names in the geographical database to obtain a list of alternative place names;
[0011] Perform semantic embedding coding on each alternative place name in the list of alternative place names to obtain a sequence of alternative place name semantic embedding coding vectors;
[0012] Based on the semantic similarity between the optimized candidate place name semantic embedding coding vector and each alternative place name semantic embedding coding vector in the sequence of alternative place name semantic embedding coding vectors, establish an association between the alternative place names and the candidate place names.
[0013] Preferably, based on the place name supplementary context content, performing semantic compensation optimization on the candidate place names based on principal component analysis to obtain an optimized candidate place name semantic embedding coding vector includes:
[0014] Perform semantic embedding coding on the candidate place names to obtain a candidate place name semantic embedding coding vector;
[0015] Perform context semantic coding on the place name supplementary context content to obtain a place name supplementary content context semantic coding vector;
[0016] Perform feature principal component compensation-based interactive optimization on the place name supplementary content context semantic coding vector and the candidate place name semantic embedding coding vector to obtain the optimized candidate place name semantic embedding coding vector.
[0017] Preferably, performing context semantic coding on the place name supplementary context content to obtain a place name supplementary content context semantic coding vector includes:
[0018] Use a context semantic encoder based on the BERT model to perform context semantic coding on the place name supplementary context content to obtain the place name supplementary content context semantic coding vector.
[0019] Preferably, performing feature principal component compensatory interaction optimization on the context semantic encoding vector of the supplementary content of the place name and the semantic embedding encoding vector of the candidate place name to obtain the optimized semantic embedding encoding vector of the candidate place name, including:
[0020] Performing principal component extraction on the context semantic encoding vector of the supplementary content of the place name and the semantic embedding encoding vector of the candidate place name to obtain a set of semantic feature principal component encoding vectors of the supplementary content of the place name and a set of semantic feature principal component encoding vectors of the candidate place name;
[0021] Performing semantic difference significance measurement on the set of semantic feature principal component encoding vectors of the supplementary content of the place name and the set of semantic feature principal component encoding vectors of the candidate place name to obtain a candidate place name - supplementary content semantic difference embedding compensation encoding weight vector;
[0022] Based on the candidate place name - supplementary content semantic difference embedding compensation encoding weight vector, performing compensatory aggregation interaction encoding on the set of semantic feature principal component encoding vectors of the supplementary content of the place name and the set of semantic feature principal component encoding vectors of the candidate place name to obtain the optimized semantic embedding encoding vector of the candidate place name.
[0023] Preferably, performing semantic difference significance measurement on the set of semantic feature principal component encoding vectors of the supplementary content of the place name and the set of semantic feature principal component encoding vectors of the candidate place name to obtain a candidate place name - supplementary content semantic difference embedding compensation encoding weight vector, including:
[0024] Constructing the set of semantic feature principal component encoding vectors of the supplementary content of the place name and the set of semantic feature principal component encoding vectors of the candidate place name into a place name supplementary content semantic feature principal component aggregation encoding feature map and a candidate place name semantic feature principal component aggregation encoding feature map;
[0025] Inputting the place name supplementary content semantic feature principal component aggregation encoding feature map and the candidate place name semantic feature principal component aggregation encoding feature map into a feature embedding unit respectively to obtain a place name supplementary content semantic feature weight vector and a candidate place name semantic feature weight vector;
[0026] Calculating the candidate place name - supplementary content semantic difference embedding compensation encoding weight vector based on the place name supplementary content semantic feature weight vector and the candidate place name semantic feature weight vector.
[0027] Preferably, calculating the candidate place name - supplementary content semantic difference embedding compensation encoding weight vector based on the place name supplementary content semantic feature weight vector and the candidate place name semantic feature weight vector, including:
[0028] Calculate the position-wise difference vector between the semantic feature weight vector of the geographical name supplementary content and the semantic feature weight vector of the candidate geographical name, and take the absolute value of the position-wise difference vector to obtain the candidate geographical name-supplementary content semantic difference embedding compensation coding weight vector.
[0029] Preferably, based on the candidate geographical name-supplementary content semantic difference embedding compensation coding weight vector, perform compensatory aggregation interaction coding on the set of semantic feature principal component coding vectors of the geographical name supplementary content and the set of semantic feature principal component coding vectors of the candidate geographical name to obtain the optimized candidate geographical name semantic embedding coding vector, including:
[0030] Calculate the position-wise mean vector of the set of semantic feature principal component coding vectors of the geographical name supplementary content and the set of semantic feature principal component coding vectors of the candidate geographical name to obtain the semantic feature principal component representation coding vector of the geographical name supplementary content and the semantic feature principal component representation coding vector of the candidate geographical name;
[0031] Based on the candidate geographical name-supplementary content semantic difference embedding compensation coding weight vector, perform aggregation interaction coding on the semantic feature principal component representation coding vector of the geographical name supplementary content and the semantic feature principal component representation coding vector of the candidate geographical name to obtain the optimized candidate geographical name semantic embedding coding vector.
[0032] Preferably, based on the semantic similarity between the optimized candidate geographical name semantic embedding coding vector and each alternative geographical name semantic embedding coding vector in the sequence of alternative geographical name semantic embedding coding vectors, establish the association between the alternative geographical name and the candidate geographical name, including:
[0033] Calculate the semantic matching degree between the optimized candidate geographical name semantic embedding coding vector and each alternative geographical name semantic embedding coding vector in the sequence of alternative geographical name semantic embedding coding vectors to obtain a sequence of semantic matching degrees;
[0034] Extract the maximum value in the sequence of semantic matching degrees and determine whether the maximum value exceeds a preset threshold. If so, establish the association between the alternative geographical name corresponding to the maximum value and the candidate geographical name.
[0035] Preferably, calculate the semantic matching degree between the optimized candidate geographical name semantic embedding coding vector and each alternative geographical name semantic embedding coding vector in the sequence of alternative geographical name semantic embedding coding vectors to obtain a sequence of semantic matching degrees, including:
[0036] Calculate the cosine similarity between the optimized candidate geographical name semantic embedding coding vector and the alternative geographical name semantic embedding coding vector as the semantic matching degree.
[0037] The present application has at least the following technical effects:
[0038] Compared with the prior art, the automatic association processing method based on open-source place name data provided by the present application first performs entity detection on the natural language text content containing place names to extract candidate place names, and uses other content in the text as a supplement. It adopts a deep learning-based semantic encoding technology to perform semantic encoding and compensatory interactive fusion on the candidate place names and their supplementary content, so as to use the supplementary content as the context background to optimize the semantic feature expression of the candidate place names. Furthermore, it constructs a list of alternative place names by querying the associated entity data of the candidate place names in the geographical database, and realizes the automatic association of place name data based on the semantic similarity between each alternative place name and the candidate place name in the list. In this way, the accuracy of place name data association can be effectively improved, while reducing the dependence on manual annotation and improving the data processing efficiency. Description of the Drawings
[0039] By describing the embodiments of the present application in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of the present application will become more apparent. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. They are used together with the embodiments of the present application to explain the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0040] Figure 1 It is a flowchart of the automatic association processing method based on open-source place name data according to an embodiment of the present application.
[0041] Figure 2 It is a schematic diagram of data flow of the automatic association processing method based on open-source place name data according to an embodiment of the present application.
[0042] Figure 3 It is a flowchart of sub-step S3 of the automatic association processing method based on open-source place name data according to an embodiment of the present application.
[0043] Figure 4 It is a flowchart of sub-step S33 of the automatic association processing method based on open-source place name data according to an embodiment of the present application.
[0044] Figure 5 It is a flowchart of sub-step S332 of the automatic association processing method based on open-source place name data according to an embodiment of the present application.
[0045] Figure 6 It is a flowchart of sub-step S6 of the automatic association processing method based on open-source place name data according to an embodiment of the present application. Detailed Embodiments
[0046] As shown in this application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0047] Although this application makes various references to certain modules in the system according to the embodiments of this application, however, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0048] Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the operations before or below do not necessarily need to be executed precisely in order. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or more steps can be removed from these processes.
[0049] Next, example embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the example embodiments described here.
[0050] It should be noted that all data acquisition and processing in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.
[0051] Specifically, Figure 1 is a flowchart of an automatic association processing method based on open-source place name data according to an embodiment of this application. Figure 2 is a schematic diagram of data flow of an automatic association processing method based on open-source place name data according to an embodiment of this application. As Figure 1 and Figure 2As shown, the automatic association processing method based on open-source place name data includes the steps: S1, obtaining natural language text content containing place names; S2, performing entity detection on the natural language text content containing place names to extract candidate place names, and defining the content other than the candidate place names in the natural language text content containing place names as place name supplementary context content; S3, based on the place name supplementary context content, performing semantic compensation optimization on the candidate place names based on principal component analysis to obtain an optimized candidate place name semantic embedding coding vector; S4, querying the associated entity data of the candidate place names in the geographical database to obtain a place name alternative list; S5, performing semantic embedding coding on each alternative place name in the place name alternative list to obtain a sequence of alternative place name semantic embedding coding vectors; S6, based on the semantic similarity between the optimized candidate place name semantic embedding coding vector and each alternative place name semantic embedding coding vector in the sequence of alternative place name semantic embedding coding vectors, establishing an association between the alternative place names and the candidate place names.
[0052] In the above automatic association processing method based on open-source place name data, in step S1, natural language text content containing place names is obtained. In a specific example of the present application, the natural language text content containing place names can be sourced from various open-source channels, such as news reports, social media, user comments, travel blogs, etc. In order to extract text fragments containing place name information from various data sources, considering the data characteristics of different sources, methods such as web crawling, API interface calls, social media platform cooperation, and user-generated content (UGC) collection can be adopted.
[0053] Specifically, first, in terms of web crawling, it mainly relies on advanced web crawler algorithms. These algorithms can automatically access web page resources on the Internet according to pre-set rules, screen and capture text content containing place name information. The web crawler will update the crawling action regularly at a specific frequency to ensure that the collected information remains up-to-date. In addition, the crawler also has the ability to process dynamically loaded content and can effectively capture place name-related information asynchronously loaded through scripting languages such as JavaScript.
[0054] Secondly, using API interfaces is another efficient way to obtain text content related to place names. Many online map service providers, news media websites and social platforms have opened official API interfaces, allowing third-party developers to legally obtain specified types of data. By establishing partnerships with these service providers and following the corresponding authorization and certification processes, a large amount of high-quality original data can be easily obtained. For example, some map APIs provide functions such as searching for locations and route planning, which can directly query text descriptions with geographic location tags; while news or social platform APIs help capture mentions of place names in real-time event reports or user comments.
[0055] Furthermore, cooperation with social media platforms constitutes an indispensable data source channel. With the popularity of social applications such as Weibo and WeChat Moments, more and more people tend to share their personal experiences and insights on the Internet, including content involving specific geographical locations. Therefore, reaching cooperation agreements with major social platforms has become one of the key ways to obtain such valuable information. In this process, special attention should be paid to user privacy protection issues to ensure that all operations comply with laws and regulations and data collection work is carried out under the premise of respecting the wishes of users.
[0056] Finally, user-generated content (UGC), as an important source of information that has emerged in recent years, is also highly valued. UGC covers a wide range, from travel blogs, forum posts to text descriptions under photos and videos, which can all be regarded as forms of UGC. For this part of unstructured data rich in place name elements, more flexible and diverse collection methods are often required. On the one hand, special theme communities can be set up to attract the majority of netizens to actively participate in contributing materials; on the other hand, natural language processing technology can be used to automatically screen and classify massive UGC content. In this way, it can not only ensure that the amount of data is large enough, but also improve the quality of information to a certain extent, laying a solid foundation for subsequent in-depth analysis.
[0057] In the above-mentioned automatic association processing method based on open-source geographical name data, in step S2, entity detection is performed on the natural language text content containing geographical names to extract candidate geographical names, and the content in the natural language text content containing geographical names except the candidate geographical names is defined as geographical name supplementary context content. Specifically, in order to accurately locate the geographical name entities in the text, the present application further uses entity detection (NER, Named Entity Recognition) technology to process the natural language text content containing geographical names, so as to clarify the geographical name part in the text, and use the surrounding text as geographical name supplementary information to provide a context basis for the semantic optimization and understanding of geographical name information. In the embodiment of the present application, a conditional random field (CRF) model is used to perform entity detection on the natural language text content containing geographical names to identify and extract the geographical name entities in the text as candidate geographical names. Among them, the conditional random field model is a statistical-based discriminative model, which can accurately identify and extract geographical name entities by comprehensively considering the context information in the input natural language text content.
[0058] Specifically, first, a training data set specifically for geographical name recognition needs to be constructed. This data set should include a large number of natural language text samples annotated with geographical name entities, covering a wide range of language environments and application scenarios, such as news reports, social media, travel blogs, etc. The geographical names in each sample need to be accurately annotated manually or semi-automatically, marking the start position and end position, as well as the corresponding geographical name type (such as city name, country name, street name, etc.), as the basis for training the conditional random field model, so that the model can learn the representation forms and characteristic patterns of geographical names in different contexts.
[0059] Next, in the model training stage, the reason for choosing the conditional random field model is its strong context awareness ability and flexible feature engineering support. The conditional random field allows multiple feature functions to be considered simultaneously, not limited to word-level attributes, but also including advanced features such as part-of-speech tagging, dependency syntactic relations, named entity boundary markers, etc. In addition, the conditional random field can handle the label dependence problem in sequence data, that is, the choice of the previous label will affect the probability distribution of the subsequent label, which is crucial for correctly capturing geographical name entities in continuous text. The model parameters are adjusted through an iterative optimization algorithm to make the prediction results as close as possible to the actual annotation, and finally a geographical name recognition model with high accuracy and recall rate is obtained.
[0060] After the model training is completed and when entering the actual application stage, in the face of a new natural language text input, the conditional random field model will scan the entire sentence structure word by word and calculate the possibility of a place name appearing at each position according to the preset feature template. When encountering a suspected place name segment, the model will make a judgment by combining the clues provided by the context before and after to decide whether it is indeed recognized as a real place name entity. For those words with multiple meanings or uncertainties, such as certain words that are both common nouns and proper nouns, the model can also refer to the context information in a larger scope to assist in the decision-making. Once it is determined that there is a place name entity at a certain position, the exact boundary will be recorded, and the rest will be regarded as supplementary context for the place name.
[0061] Specifically, during the processing, special attention also needs to be paid to some special situations, such as various forms of place name variants like abbreviations, aliases, and dialect expressions. For this reason, in addition to relying on the capabilities of the conditional random field model itself, external resource libraries can also be introduced for auxiliary matching. For example, a dictionary containing common place names and their various variant forms can be established to expand the model's recognition scope; or a rule engine can be used to achieve fast searching based on specific patterns, such as the city names corresponding to postal codes. This can not only improve the recognition efficiency but also enhance the flexibility of the system in dealing with complex and diverse place name expression forms.
[0062] In the above automatic association processing method based on open-source place name data, in step S3, based on the supplementary context of the place name, semantic compensation optimization based on principal component analysis is performed on the candidate place name to obtain an optimized candidate place name semantic embedding coding vector. Among them, Figure 3 It is a flowchart of sub-step S3 of the automatic association processing method based on open-source place name data according to an embodiment of the present application. As Figure 3 shown, step S3 includes steps: S31, performing semantic embedding coding on the candidate place name to obtain a candidate place name semantic embedding coding vector; S32, performing context semantic coding on the supplementary context of the place name to obtain a place name supplementary content context semantic coding vector; S33, performing feature principal component compensation-based interactive optimization on the place name supplementary content context semantic coding vector and the candidate place name semantic embedding coding vector to obtain the optimized candidate place name semantic embedding coding vector.
[0063] Specifically, in step S31, semantic embedding encoding is performed on the candidate place names to obtain candidate place name semantic embedding encoding vectors. Specifically, this application takes into account that traditional place name association methods often rely on character-level matching operations, which have significant limitations when dealing with complex and diverse place name expression forms. Therefore, to improve the accuracy of place name association, this application further uses a deep learning-based semantic embedding encoding technology to process the candidate place names, mapping the candidate place names into a high-dimensional semantic space to obtain candidate place name semantic embedding encoding vectors. In an embodiment of this application, the Word2Vec model is used as the semantic embedding encoder to perform semantic encoding on the candidate place names. The Word2Vec model is a commonly used word embedding technology that, by mapping candidate place names into a high-dimensional vector space, can, while preserving their original semantic information, make the distances between semantically similar place name information in the vector space relatively close, thus providing a basis for subsequent place name association processing.
[0064] Specifically, in a specific example of this application, step S32 includes: using a context semantic encoder based on the BERT model to perform context semantic encoding on the context content supplemented for the place names to obtain the context semantic encoding vectors of the place name supplementary content. Specifically, since the context content supplemented for the place names contains the context environment information of the place name entities, it plays an important supplementary role in the semantic understanding of the candidate place names. Therefore, this application further uses the BERT model to perform context semantic encoding on the context content supplemented for the place names. Among them, the BERT (Bidirectional Encoder Representations from Transformers) model is a pre-trained language model based on a bidirectional Transformer architecture. Based on the self-attention mechanism, by performing bidirectional context understanding and modeling on the context content supplemented for the place names, it can effectively capture the semantic dependency relationships between various words in the text, thus providing a more comprehensive and accurate context background for the semantic understanding of place name entities and generating context semantic encoding vectors of the place name supplementary content.
[0065] Specifically, in step S33, feature principal component compensatory interaction optimization is performed on the context semantic encoding vectors of the place name supplementary content and the candidate place name semantic embedding encoding vectors to obtain the optimized candidate place name semantic embedding encoding vectors. Specifically, in order to make full use of the context semantic background information of the place name supplementary content, this application uses a compensatory interaction fusion method based on feature principal components. By performing principal component feature extraction and information complementary analysis on the context semantic encoding vectors of the place name supplementary content and the candidate place name semantic embedding encoding vectors, it realizes the deep semantic interaction and fusion between the two, thus obtaining a more accurate and comprehensive candidate place name semantic expression.Figure 4 It is a flowchart of sub-step S33 of the automatic association processing method based on open-source place name data according to an embodiment of the present application. As Figure 4 shown, the step S33 includes steps: S331, performing principal component extraction on the context semantic encoding vector of the place name supplementary content and the semantic embedding encoding vector of the candidate place name to obtain a set of principal component encoding vectors of the place name supplementary content semantic features and a set of principal component encoding vectors of the candidate place name semantic features; S332, performing semantic difference significance measurement on the set of principal component encoding vectors of the place name supplementary content semantic features and the set of principal component encoding vectors of the candidate place name semantic features to obtain a candidate place name-supplementary content semantic difference embedding compensation encoding weight vector; S333, based on the candidate place name-supplementary content semantic difference embedding compensation encoding weight vector, performing compensatory aggregation interaction encoding on the set of principal component encoding vectors of the place name supplementary content semantic features and the set of principal component encoding vectors of the candidate place name semantic features to obtain the optimized candidate place name semantic embedding encoding vector.
[0066] More specifically, the step S331 is expressed by the formula:
[0067] ;
[0068] ;
[0069] Among them, represents the context semantic encoding vector of the place name supplementary content, represents the semantic embedding encoding vector of the candidate place name, represents the principal component analysis function, represents the transpose of the vector, is the and the feature scale value, and respectively represent the and the covariance matrix, represents the matrix composed of the set of principal component encoding vectors of the place name supplementary content semantic features obtained by performing eigenvalue decomposition on the , represents the diagonal matrix composed of the set of principal component eigenvalues of the place name supplementary content semantic features obtained by performing eigenvalue decomposition on the , represents each principal component eigenvalue of the place name supplementary content semantic features, is the number of principal component eigenvalues of the place name supplementary content semantic features, Each semantic feature principal component encoding vector in the set of semantic feature principal component encoding vectors of the supplementary content of the place name represents the matrix composed of the set of candidate place name semantic feature principal component encoding vectors obtained by performing eigenvalue decomposition on the represents the diagonal matrix composed of the set of candidate place name semantic principal component eigenvalues obtained by performing eigenvalue decomposition on the represents each candidate place name semantic principal component eigenvalue Each candidate place name semantic feature principal component encoding vector in the set of candidate place name semantic feature principal component encoding vectors
[0070] Specifically, the principal component analysis (PCA) method is used to mine the main semantic feature components in the place name entity and the supplementary content of the place name, and multiple main feature components are extracted from the context semantic encoding vector of the supplementary content of the place name and the candidate place name semantic embedding encoding vector to form the set of semantic feature principal component encoding vectors of the supplementary content of the place name and the set of candidate place name semantic feature principal component encoding vectors, so as to retain the core semantic information in the place name entity and the supplementary content of the place name, while removing redundancy and noise, making the subsequent semantic interaction processing more efficient and accurate.
[0071] Figure 5 is the flowchart of sub-step S332 of the automatic association processing method based on open-source place name data according to the embodiment of the present application. As Figure 5 shown, the step S332 includes the steps: S3321, constructing the set of semantic feature principal component encoding vectors of the supplementary content of the place name and the set of candidate place name semantic feature principal component encoding vectors into a semantic feature principal component aggregation encoding feature map of the supplementary content of the place name and a semantic feature principal component aggregation encoding feature map of the candidate place name; S3322, respectively inputting the semantic feature principal component aggregation encoding feature map of the supplementary content of the place name and the semantic feature principal component aggregation encoding feature map of the candidate place name into a feature embedding unit to obtain a semantic feature weight vector of the supplementary content of the place name and a semantic feature weight vector of the candidate place name; S3323, calculating the semantic difference embedding compensation encoding weight vector of the candidate place name - supplementary content based on the semantic feature weight vector of the supplementary content of the place name and the semantic feature weight vector of the candidate place name.
[0072] In a specific example of the present application, the step S3321 is represented by the formula:
[0073] ;
[0074] ;
[0075] where and respectively represent the aggregated coding feature map of the semantic feature principal components of the supplementary content of the place name and the aggregated coding feature map of the semantic feature principal components of the candidate place name, indicating the reshaping of the feature shape.
[0076] Specifically, the set of the semantic feature principal component coding vectors of the supplementary content of the place name and the set of the semantic feature principal component coding vectors of the candidate place name are further reconstructed into the form of feature maps respectively, so as to maintain the structural relationship between the place name entity and the supplementary content of the place name in the semantic space, and form the aggregated coding feature map of the semantic feature principal components of the supplementary content of the place name and the aggregated coding feature map of the semantic feature principal components of the candidate place name.
[0077] In a specific example of the present application, the step S3322 is expressed by the formula:
[0078] ;
[0079] wherein, represents the mean pooling function, and respectively represent two learnable weight parameter matrices of the , represents the Sigmoid activation function, and respectively represent two learnable weight parameter matrices of the , represents the normalization function, represents the semantic feature weight vector of the supplementary content of the place name, represents the semantic feature weight vector of the candidate place name.
[0080] Specifically, the feature embedding unit is used to perform low-dimensional embedding coding on the aggregated coding feature map of the semantic feature principal components of the supplementary content of the place name and the aggregated coding feature map of the semantic feature principal components of the candidate place name, so as to quantify the feature importance of the semantic principal component features of the supplementary content of the place name and the semantic principal component features of the candidate place name through the processing of the multi-layer neural network, and generate the semantic feature weight vector of the supplementary content of the place name and the semantic feature weight vector of the candidate place name.
[0081] In a specific example of the present application, the step S3323 includes: calculating the position-wise difference vector between the semantic feature weight vector of the supplementary content of the place name and the semantic feature weight vector of the candidate place name and taking the absolute value of the position-wise difference vector to obtain the candidate place name-supplementary content semantic difference embedding compensation coding weight vector, which is expressed by the formula:
[0082] ;
[0083] Among them, represents a differential operation, represents the candidate place name - supplementary content semantic difference embedding compensation coding weight vector, represents taking the absolute value.
[0084] Among them, by performing a differential operation on the semantic feature weight vector of the place name supplementary content and the semantic feature weight vector of the candidate place name, the information complementarity between the place name entity and the place name supplementary content can be revealed, and a candidate place name - supplementary content semantic difference embedding compensation coding weight vector can be generated to guide subsequent semantic fusion processing.
[0085] More specifically, in a specific example of the present application, the step S333 includes: First, calculate the position - wise mean vectors of the set of semantic feature principal component coding vectors of the place name supplementary content and the set of semantic feature principal component coding vectors of the candidate place name to obtain the semantic feature principal component representation coding vector of the place name supplementary content and the semantic feature principal component representation coding vector of the candidate place name. It is expressed by the formula:
[0086] ;
[0087] ;
[0088] Among them, represents the th semantic feature principal component coding vector of the place name supplementary content in the set of semantic feature principal component coding vectors of the place name supplementary content, represents the semantic feature principal component representation coding vector of the place name supplementary content, represents the th semantic feature principal component coding vector of the candidate place name in the set of semantic feature principal component coding vectors of the candidate place name, represents the semantic feature principal component representation coding vector of the candidate place name.
[0089] Specifically, in subsequent semantic information compensation interaction, in order to comprehensively consider the semantic feature principal components of the candidate place name and the semantic feature principal components of the place name supplementary content as a whole, first, through position - wise mean calculation, the integration of the internal principal component features of the set of semantic feature principal component coding vectors of the place name supplementary content and the set of semantic feature principal component coding vectors of the candidate place name is performed.
[0090] Then, based on the candidate place name - supplementary content semantic difference embedding compensation coding weight vector, an aggregation interaction coding is performed on the semantic feature principal component representation coding vector of the place name supplementary content and the semantic feature principal component representation coding vector of the candidate place name to obtain the optimized candidate place name semantic embedding coding vector. It is expressed by the formula:
[0091] ;
[0092] Among them, represents dot product, represents dot convolution operation, represents optimizing the semantic embedding coding vector of candidate place names.
[0093] Among them, taking the semantic difference embedding compensation coding weight vector of the candidate place name - supplementary content as the compensation guidance, the main components of the semantic features of the integrated candidate place name and the main components of the semantic features of the place name supplementary content are weighted and fused, and through the processing of dot convolution coding and activation function, the key information in the fused features is further extracted and emphasized to obtain the optimized semantic embedding coding vector of the candidate place name. In this process, the semantic difference information between the place name supplementary content and the place name entity is effectively identified and compensated by means of weighted fusion, thus effectively enhancing the integrity and accuracy of the semantic expression of the candidate place name, and helping to provide a more reliable and accurate semantic information basis for subsequent place name association processing.
[0094] In the above automatic association processing method based on open-source place name data, in step S4, query the associated entity data of the candidate place name in the geographic database to obtain a list of alternative place names. Specifically, a large amount of place name entity data and its related information are stored in the geographic database. By querying the candidate place name in the geographic database, it is possible to quickly obtain the place name entity data that is literally similar or related to the candidate place name, forming a list of alternative place names, thus providing rich reference information for subsequent place name association processing.
[0095] Specifically, first at the level of geographic database design, constructing an efficient data index structure is the basis for ensuring fast query response. The geographic database usually establishes a full-text index or an inverted index for the place name field, enabling the system to quickly locate entities highly relevant to the input place name among a large number of records. In addition, considering that place names may have various variant forms, such as different language versions, historical name changes, etc., additional auxiliary indexes need to be created to support matching operations in a wider range. For example, for some ancient cities, the database not only includes the modern standard name but also preserves various names recorded in ancient documents to facilitate handling various complex query requirements.
[0096] To further improve the matching accuracy, the fuzzy matching algorithm plays an important role. Traditional string comparison-based methods are difficult to handle the diverse problems of place name expressions commonly existing in the real world. Therefore, advanced algorithms such as edit distance and phonetic and shape code conversion are introduced to address such challenges. The edit distance algorithm measures the similarity by calculating the minimum number of transformations (inserting, deleting, replacing characters) between two strings; while phonetic and shape code conversion converts text into pronunciation or shape feature codes, so as to effectively identify homophonic characters, similar characters, etc. These algorithms work together to enable the corresponding geographical entities to be found even in the face of abbreviations, aliases, and even place names expressed in dialects.
[0097] In the above automatic association processing method based on open-source place name data, in step S5, semantic embedding encoding is performed on each alternative place name in the alternative place name list to obtain a sequence of alternative place name semantic embedding encoding vectors. Specifically, in order to achieve the precise association at the semantic level between each alternative place name in the alternative place name list and the candidate place name, the present application further processes each alternative place name in the alternative place name list by using semantic embedding encoding technology (such as the Word2Vec model), maps each alternative place name to a high-dimensional vector space, and generates a sequence of alternative place name semantic embedding encoding vectors, so as to calculate and compare the semantic similarity based on the distance relationship of each place name entity data in the vector space, thereby establishing a precise semantic association between the candidate place name and the alternative place name.
[0098] In the above automatic association processing method based on open-source place name data, in step S6, based on the semantic similarity between the optimized candidate place name semantic embedding encoding vector and each alternative place name semantic embedding encoding vector in the sequence of alternative place name semantic embedding encoding vectors, an association between the alternative place name and the candidate place name is established. Among them, Figure 6 is a flowchart of sub-step S6 of the automatic association processing method based on open-source place name data according to an embodiment of the present application. As Figure 6 shown, step S6 includes steps: S61, calculating the semantic matching degree between the optimized candidate place name semantic embedding encoding vector and each alternative place name semantic embedding encoding vector in the sequence of alternative place name semantic embedding encoding vectors to obtain a sequence of semantic matching degrees; S62, extracting the maximum value in the sequence of semantic matching degrees and determining whether the maximum value exceeds a preset threshold. If so, an association between the alternative place name corresponding to the maximum value and the candidate place name is established.
[0099] Specifically, in a specific example of the present application, step S61 includes: calculating the cosine similarity between the optimized candidate place name semantic embedding coding vector and the alternative place name semantic embedding coding vector as the semantic matching degree. Specifically, the cosine similarity is an index for measuring the direction similarity between two vectors, and its value range is [-1, 1]. The closer the value is to 1, the closer the directions of the two vectors are, that is, the higher the semantic similarity. By using the cosine similarity as the measurement index for the semantic matching degree between the candidate place name and each alternative place name, the tightness of the semantic association between the candidate place name and each alternative place name can be intuitively reflected.
[0100] In a preferred example of the present application, since the optimized candidate place name semantic embedding coding vector represents the semantic coding features of the candidate place name with semantic context optimization based on the supplemented context content of the place name, and the alternative place name semantic embedding coding vector represents the semantic coding features of the alternative place name, when calculating the semantic matching degree therebetween for mapping into a common semantic matching degree space, there will be insufficient correlation correspondence between the semantic coding features of different orders under different source semantic representations, thereby causing a sparse association mapping of the optimized candidate place name semantic embedding coding vector and the alternative place name semantic embedding coding vector towards the common semantic matching degree space, and thus reducing the accuracy of the calculated semantic matching degree due to the lack of common mapping inference degree.
[0101] Based on this, before calculating the semantic matching degree between the optimized candidate place name semantic embedding coding vector and the alternative place name semantic embedding coding vector, first perform common association optimization on the optimized candidate place name semantic embedding coding vector and the alternative place name semantic embedding coding vector, which specifically includes the steps:
[0102] First, concatenate the optimized candidate place name semantic embedding coding vector and the alternative place name semantic embedding coding vector into a candidate place name - alternative place name semantic information common coding vector;
[0103] Secondly, based on the distance and distance between the eigenvalues of the candidate place name - alternative place name semantic information common coding vector, obtain a candidate place name - alternative place name semantic information common first distance matrix and a candidate place name - alternative place name semantic information common second distance matrix ;
[0104] Then, calculate the weighted sum of the candidate place name - alternative place name semantic information common first distance matrix and the candidate place name - alternative place name semantic information common second distance matrix to obtain a candidate place name - alternative place name semantic information common joint distance matrix, which is expressed by the formula:
[0105] ;
[0106] Among them, represents the candidate place name - alternative place name semantic information common one distance matrix, represents the candidate place name - alternative place name semantic information common two distance matrix, represents dot product by position, represents matrix addition, and respectively represent different weight parameters, represents the candidate place name - alternative place name semantic information common joint distance matrix;
[0107] Next, determine each eigenvalue of the candidate place name - alternative place name semantic information common joint distance matrix, and form the candidate place name - alternative place name semantic information common joint distance eigenvector ;
[0108] Then, multiply the candidate place name - alternative place name semantic information common coding vector as a row vector with the candidate place name - alternative place name semantic information common one distance matrix to obtain the candidate place name - alternative place name semantic information common one distance query vector, which is expressed by the formula:
[0109] ;
[0110] Among them, represents the candidate place name - alternative place name semantic information common coding vector, represents matrix multiplication, represents the candidate place name - alternative place name semantic information common one distance query vector;
[0111] Next, multiply the candidate place name - alternative place name semantic information common two distance matrix with the self - association matrix of the candidate place name - alternative place name semantic information common coding vector to obtain the candidate place name - alternative place name semantic information common two distance association matrix, which is expressed by the formula:
[0112] ;
[0113] Among them, represents the transpose of the vector, represents the candidate place name - alternative place name semantic information common two distance association matrix;
[0114] Then, after multiplying the candidate place name - alternative place name semantic information common distance query vector with the candidate place name - alternative place name semantic information common distance correlation matrix, and further performing a dot product with the candidate place name - alternative place name semantic information common combined distance eigenvector to obtain an optimized candidate place name - alternative place name semantic information common coding vector, which is expressed by the formula:
[0115] ;
[0116] Among them, represents the candidate place name - alternative place name semantic information common combined distance eigenvector, represents the optimized candidate place name - alternative place name semantic information common coding vector;
[0117] Finally, the optimized candidate place name - alternative place name semantic information common coding vector is split into an optimized candidate place name semantic embedding coding vector and an optimized alternative place name semantic embedding coding vector.
[0118] Specifically, for the first - distance matrix and the second - distance matrix of the candidate place name - alternative place name semantic information common coding vector obtained by concatenating the optimized candidate place name semantic embedding coding vector and the alternative place name semantic embedding coding vector as the fine - grained metric correlation cluster representation of the candidate place name - alternative place name semantic information common coding vector, the dynamic programming of the relationship between different correlation clusters of the self - correlation representation of the candidate place name - alternative place name semantic information common coding vector and the candidate place name - alternative place name semantic information common coding vector is respectively performed to simulate the sparse activation based on neuron clusters of the correlation system, and the eigen - representation of the metric correlation cluster of the first - distance matrix and the second - distance matrix of the candidate place name - alternative place name semantic information common coding vector is used to cooperate with the fine - grained predictable sparsity of the candidate place name - alternative place name semantic information common coding vector, so as to avoid the lack of correlation caused by sparsity affecting the missing of the common mapping inference degree, and improve the calculation accuracy of the semantic matching degree between the optimized candidate place name semantic embedding coding vector and the alternative place name semantic embedding coding vector.
[0119] Specifically, in step S62, the maximum value in the sequence of semantic matching degrees is extracted, and it is determined whether the maximum value exceeds a preset threshold. If so, an association is established between the alternative place name corresponding to the maximum value and the candidate place name. Specifically, the alternative place name corresponding to the maximum value in the sequence of semantic matching degrees represents the place name entity that is semantically closest or most similar to the candidate place name. When its semantic matching degree exceeds the preset threshold, it can be considered that there is a significant semantic association between this alternative place name and the candidate place name, so that it is determined as the associated place name of the candidate place name, and an association between the two is established, thereby providing an important reference basis for subsequent tasks such as place name parsing, place name standardization, or place name association analysis. If the maximum value does not exceed the preset threshold, it indicates that the degree of semantic association between the candidate place name and all alternative place names is not significant enough. At this time, it can be considered that the place name parsing or association task fails, or further processing measures need to be taken, such as expanding the scope of the place name alternative list, further correcting the candidate place name, or manual review, etc., to improve the accuracy and reliability of the place name parsing and association tasks.
[0120] In summary, the automatic association processing method based on open-source place name data according to the embodiments of the present application is elucidated. First, entity detection is performed on the natural language text content containing place names to extract candidate place names, and other content in the text is used as a supplement. The semantic encoding technology based on deep learning is used to perform semantic encoding and compensatory interactive fusion on the candidate place names and their supplementary content, so as to use the supplementary content as the context background to optimize the semantic feature expression of the candidate place names. Furthermore, the associated entity data of the candidate place names in the geographic database is queried to construct a place name alternative list, and the automatic association of place name data is realized based on the semantic similarity between each alternative place name in the list and the candidate place name. In this way, the accuracy of place name data association can be effectively improved, while reducing the dependence on manual annotation and improving the data processing efficiency.
[0121] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, advantages, effects, etc. mentioned in the present invention are only examples and not limitations. It cannot be considered that these advantages, advantages, effects, etc. are essential for each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and easy understanding, rather than limitations. The above details do not limit the present invention to necessarily adopt the above specific details to implement.
[0122] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0123] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
[0124] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.
[0125] Finally, it should be noted that the above description has been given for purposes of illustration and description. In addition, the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An automatic association processing method based on open source place name data, characterized in that: include: Get natural language text content containing place names; Performing entity detection on the natural language text content containing place names to extract candidate place names, and defining the content of the natural language text content containing place names other than the candidate place names as place name supplementary context content; Based on the supplementary context content of the place name, the candidate place name is subjected to semantic compensation optimization based on principal component analysis to obtain an optimized semantic embedding coding vector of the candidate place name; Querying the associated entity data of the candidate place name in the geographic database to obtain a candidate list of place names; Performing semantic embedding coding on each candidate place name in the candidate place name list to obtain a sequence of semantic embedding coding vectors of the candidate place names; Based on the semantic similarity between the optimized candidate place name semantic embedding coding vector and each alternative place name semantic embedding coding vector in the sequence of alternative place name semantic embedding coding vectors, an association between the alternative place name and the candidate place name is established.
2. The automatic association processing method based on open source place name data according to claim 1 is characterized in that: Based on the supplementary context content of the place name, the candidate place name is subjected to semantic compensation optimization based on principal component analysis to obtain an optimized semantic embedding coding vector of the candidate place name, including: Performing semantic embedding coding on the candidate place name to obtain a semantic embedding coding vector of the candidate place name; Performing context semantic coding on the place name supplementary context content to obtain a place name supplementary content context semantic coding vector; The place name supplementary content context semantic coding vector and the candidate place name semantic embedding coding vector are interactively optimized using principal component compensation to obtain the optimized candidate place name semantic embedding coding vector.
3. The automatic association processing method based on open source place name data according to claim 2 is characterized in that: Performing context semantic coding on the place name supplementary context content to obtain a place name supplementary content context semantic coding vector includes: The place name supplementary context content is contextually semantically encoded using a contextual semantic encoder based on the BERT model to obtain the place name supplementary content contextual semantic encoding vector.
4. The automatic association processing method based on open source place name data according to claim 3 is characterized in that: Performing feature principal component compensation interactive optimization on the place name supplementary content context semantic coding vector and the candidate place name semantic embedding coding vector to obtain the optimized candidate place name semantic embedding coding vector, including: Performing principal component extraction on the place name supplementary content context semantic coding vector and the candidate place name semantic embedding coding vector to obtain a set of place name supplementary content semantic feature principal component coding vectors and a set of candidate place name semantic feature principal component coding vectors; Performing semantic difference significance measurement on the set of principal component coding vectors of semantic features of supplementary content of place names and the set of principal component coding vectors of semantic features of candidate place names to obtain a candidate place name-supplementary content semantic difference embedding compensation coding weight vector; Based on the candidate place name-supplementary content semantic difference embedding compensation coding weight vector, the set of principal component coding vectors of the place name supplementary content semantic features and the set of principal component coding vectors of the candidate place name semantic features are subjected to compensatory aggregation interactive coding to obtain the optimized candidate place name semantic embedding coding vector.
5. The automatic association processing method based on open source place name data according to claim 4 is characterized in that: The semantic difference significance measurement is performed on the set of principal component coding vectors of the semantic features of the supplementary content of the place name and the set of principal component coding vectors of the semantic features of the candidate place name to obtain the candidate place name-supplementary content semantic difference embedding compensation coding weight vector, including: The set of principal component coding vectors of the semantic features of the supplementary content of place names and the set of principal component coding vectors of the semantic features of the candidate place names are constructed into a principal component aggregation coding feature map of the semantic features of the supplementary content of place names and a principal component aggregation coding feature map of the semantic features of the candidate place names; Inputting the principal component aggregation coding feature map of semantic features of supplementary content of place names and the principal component aggregation coding feature map of semantic features of candidate place names into a feature embedding unit respectively to obtain a semantic feature weight vector of supplementary content of place names and a semantic feature weight vector of candidate place names; Based on the place name supplementary content semantic feature weight vector and the candidate place name semantic feature weight vector, the candidate place name-supplementary content semantic difference embedding compensation coding weight vector is calculated.
6. The automatic association processing method based on open source place name data according to claim 5 is characterized in that: Calculating the candidate place name-supplementary content semantic difference embedding compensation coding weight vector based on the place name supplementary content semantic feature weight vector and the candidate place name semantic feature weight vector, including: The position difference vector between the place name supplementary content semantic feature weight vector and the candidate place name semantic feature weight vector is calculated and the absolute value of the position difference vector is taken to obtain the candidate place name-supplementary content semantic difference embedding compensation coding weight vector.
7. The automatic association processing method based on open source place name data according to claim 6 is characterized in that: Based on the candidate place name-supplementary content semantic difference embedding compensation coding weight vector, the set of the place name supplementary content semantic feature principal component coding vectors and the set of the candidate place name semantic feature principal component coding vectors are subjected to compensatory aggregation interactive coding to obtain the optimized candidate place name semantic embedding coding vector, including: Calculating the positional mean vector of the set of principal component coding vectors of the semantic features of the supplementary content of the place name and the set of principal component coding vectors of the semantic features of the candidate place names to obtain principal component representation coding vectors of the semantic features of the supplementary content of the place name and principal component representation coding vectors of the semantic features of the candidate place names; Based on the candidate place name-supplementary content semantic difference embedding compensation coding weight vector, the place name supplementary content semantic feature principal component representation coding vector and the candidate place name semantic feature principal component representation coding vector are aggregated and interactively encoded to obtain the optimized candidate place name semantic embedding coding vector.
8. The automatic association processing method based on open source place name data according to claim 7 is characterized in that: Based on the semantic similarity between the optimized candidate place name semantic embedding coding vector and each candidate place name semantic embedding coding vector in the sequence of the candidate place name semantic embedding coding vectors, establishing an association between the candidate place name and the candidate place name, including: Calculating the semantic matching degree between the optimized candidate place name semantic embedding coding vector and each candidate place name semantic embedding coding vector in the sequence of candidate place name semantic embedding coding vectors to obtain a sequence of semantic matching degrees; A maximum value in the sequence of semantic matching degrees is extracted and it is determined whether the maximum value exceeds a preset threshold. If so, an association is established between the alternative place name corresponding to the maximum value and the candidate place name.
9. The automatic association processing method based on open source place name data according to claim 8 is characterized in that: Calculating the semantic matching degree between the optimized candidate place name semantic embedding coding vector and each candidate place name semantic embedding coding vector in the sequence of candidate place name semantic embedding coding vectors to obtain a sequence of semantic matching degrees, including: The cosine similarity between the optimized candidate place name semantic embedding coding vector and the alternative place name semantic embedding coding vector is calculated as the semantic matching degree.
Citation Information
Patent Citations
Intelligent matching system and method based on natural language model
CN117521652A
Place name identification method for Chinese short text
CN119337884A