Place name address translation method based on intelligent splitting and matching of proper name library and general name library
By constructing a proper name library and a common name library for multiple rounds of word segmentation annotation and optimal segmentation scheme screening, combined with context correlation analysis and large language model generation, the accuracy and rule-making cost problems in Chinese place name address translation are solved, and high-quality place name address translation is achieved.
Patent Information
- Application Number
- CN202510779192.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
When handling Chinese place name address, the existing methods of place name address translation have low accuracy, high cost of rule formulation and maintenance, and it is difficult to deal with complex and changeable place names. The traditional word segmentation model lacks targeted recognition capabilities, resulting in uneven translation quality.
A proper name library and a common name library are built, and through multiple rounds of word segmentation annotation and optimal slicing scheme screening, combined with context association constraints and large language model generation, intelligent splitting and matching translation of place name addresses is performed.
It improves the accuracy and standardization of place name address translation, solves the problem of cleavage ambiguity, and ensures the identification and translation quality of core components.
Smart Images

Figure CN120297294A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of place name and address translation, and more specifically, to a method for translating place names and addresses by intelligent splitting and matching based on a proper name library and a common name library. Background Art
[0002] With the continuous deepening of the globalization process and the increasing frequency of international exchanges, place names and addresses, as key elements carrying geographical location information, their accurate and efficient translation is of crucial significance in many fields such as cross-language communication, international trade, cultural exchanges, navigation and positioning, and emergency response. Especially in the scenario of translating Chinese place names and addresses into foreign languages, due to the significant differences in language structures, expression habits, and place name composition rules between Chinese and Western languages, accurate and standardized translation faces many challenges, often causing communication barriers, logistics delays, and even economic losses due to improper translation.
[0003] Existing methods for translating place names and addresses either rely on general-purpose machine translation systems or adopt simple rule matching and dictionary lookup. Although general-purpose machine translation systems perform well in general text translation, when dealing with place names and addresses with highly structured and specific terms, they often result in inaccurate and non-standard translation results due to the lack of domain knowledge. For example, there are problems such as confusion between transliteration and free translation of proper name parts, improper handling of common name suffixes, and unclear expression of the hierarchical relationship of place name elements. Traditional rule- and dictionary-based methods, although able to ensure the translation accuracy of specific words to a certain extent, have limited coverage, are difficult to cope with complex, ever-changing new place names, and have high costs for rule formulation and maintenance, and weak generalization ability. Especially for Chinese place names and addresses, due to the lack of clear space separation and unified naming rules, accurately segmenting each component and accurately identifying the attributes of each component have become bottlenecks in translation. Traditional word segmentation models are trained for general corpora and lack targeted recognition ability for the unique combination of proper names and common names in place names and addresses, easily resulting in unconventional segmentation errors and being difficult to effectively solve the problems of translation ambiguity in address segmentation and out-of-vocabulary (OOV) words, thereby leading to uneven overall translation quality.
[0004] Therefore, a method for translating place names and addresses by intelligent splitting and matching based on a proper name library and a common name library is needed to solve the above technical problems. Summary of the Invention
[0005] In order to solve the above technical problems, this application is proposed.
[0006] According to one aspect of this application, a method for translating place names and addresses by intelligent splitting and matching based on a proper name library and a common name library is provided, which includes: Construct a proper name library and a common name library. Among them, the common name library stores common suffixes, type words in geographical names and addresses, and their corresponding standard translations, and the proper name library stores common proper nouns in geographical names and addresses, and their corresponding transliteration results or standard translations; Obtain the Chinese geographical name and address to be translated input by the user; Based on the proper name library and the common name library, perform multiple rounds of word segmentation annotation and optimal segmentation scheme screening on the Chinese geographical name and address to be translated to obtain an optimal splitting and matching result; Perform attribute speculation and translation generation based on context association constraints on the unannotated word segments of the Chinese geographical name and address to be translated in the optimal splitting and matching result to obtain the translation result of the unannotated geographical name and address word segments; Assemble the translation result of the unannotated geographical name and address word segments and the translation result of the annotated word segments of the Chinese geographical name and address to be translated in the optimal splitting and matching result, and perform normalization processing based on the rule engine to obtain the complete translation result of the Chinese geographical name and address.
[0007] Beneficial effects: Compared with the prior art, the intelligent splitting and matching geographical name and address translation method based on the proper name library and the common name library provided by the present application first constructs a common name library containing standard translations and a proper name library of transliterated proper names. After multiple rounds of segmentation of the geographical name and address to be translated, the word segments in each segmentation scheme are type-annotated through dual-library matching, and the optimal segmentation scheme is determined based on the annotation character coverage. On this basis, a context association inference mechanism is further introduced to perform context semantic association analysis on the unannotated word segments in the optimal segmentation scheme, speculate the type of the word segment attributes, and then call the large language model to generate an adapted translation. Finally, the rule engine performs a structured reorganization on the dual-library annotated translations and the inferred translations to form a standardized standard address translation. Through multiple rounds of word segmentation annotation and optimal segmentation scheme screening, the present application can effectively solve the segmentation ambiguity problem, accurately identify the core components in the address, and thus improve the translation quality. Description of the Drawings
[0008] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0009] Figure 1 It is a flowchart of the intelligent splitting and matching geographical name and address translation method based on the proper name library and the common name library according to the embodiments of the present application.
[0010] Figure 2A schematic diagram of data flow for a geographical name and address translation method based on intelligent splitting and matching of a proper name library and a common name library according to an embodiment of the present application.
[0011] Figure 3 A flowchart of sub-step S3 of a geographical name and address translation method based on intelligent splitting and matching of a proper name library and a common name library according to an embodiment of the present application.
[0012] Figure 4 A flowchart of sub-step S4 of a geographical name and address translation method based on intelligent splitting and matching of a proper name library and a common name library according to an embodiment of the present application.
[0013] Figure 5 A flowchart of sub-step S43 of a geographical name and address translation method based on intelligent splitting and matching of a proper name library and a common name library according to an embodiment of the present application. Detailed implementation manners
[0014] As shown in the present application and the claims, unless the context clearly indicates an exceptional situation, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0015] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0016] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. On the contrary, according to needs, various steps can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several steps of operations can be removed from these processes.
[0017] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0018] It should be noted in advance that all relevant processing of the data in the present application is carried out on the premise of complying with the corresponding data protection regulations and policies of the place where it is located and obtaining authorization from the corresponding authority manager.
[0019] Figure 1Flowchart of a place name and address translation method based on intelligent splitting and matching of proper name library and common name library according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of a place name and address translation method based on intelligent splitting and matching of proper name library and common name library according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the place name and address translation method based on intelligent splitting and matching of proper name library and common name library includes the steps of: S1, constructing a proper name library and a common name library, wherein the common name library stores common suffixes and type words in place names and addresses and their corresponding standard translations, and the proper name library stores common proper nouns in place names and addresses and their corresponding transliteration results or standard translations; S2, obtaining a Chinese place name and address to be translated input by a user; S3, based on the proper name library and the common name library, performing multiple rounds of word segmentation annotation and optimal segmentation scheme screening on the Chinese place name and address to be translated to obtain an optimal splitting and matching result; S4, performing attribute speculation and translation generation based on context association constraints on the unannotated word segments of the Chinese place name and address to be translated in the optimal splitting and matching result to obtain a translation result of the unannotated place name and address word segments; S5, assembling the translation result of the unannotated place name and address word segments and the translation result of the annotated word segments of the Chinese place name and address to be translated in the optimal splitting and matching result and performing normalization processing based on a rule engine to obtain a complete Chinese place name and address translation result.
[0020] In the above-mentioned method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library, in step S1, a proper name library and a common name library are constructed. Among them, the common name library stores common suffixes and type words in geographical names and addresses and their corresponding standard translations, and the proper name library stores common proper nouns in geographical names and addresses and their corresponding transliteration results or standard translations. It should be understood that Chinese geographical names and addresses have highly structured features and are usually composed of a proper name part (such as "Zhongguancun", "Xidan") that refers to a specific entity and a common name part (such as "district", "street", "building", "road", "hutong") that indicates its administrative division, road type, or building type. Moreover, there are often established or officially stipulated standard translation methods for these constituent elements during translation. However, due to the lack of explicit delimiters in Chinese geographical names, it leads to difficulties in segmentation and recognition during the translation process. Therefore, to lay the foundation for the accurate translation of geographical names and addresses and provide high-quality prior knowledge for subsequent intelligent segmentation and matching, this application pre-collects and organizes the core constituent elements in geographical names and addresses and their standard translations, and constructs a proper name library and a common name library with wide coverage and authoritative content to achieve the accurate recognition and translation of known geographical name elements. In the specific implementation process, the common name library will collect various types of administrative division names (such as province, city, district, county, town, township, village), road types (such as road, street, avenue, lane, hutong), point-of-interest types (such as square, park, school, hospital, shopping mall, residential area, building, unit, floor, house number), etc., as well as common suffixes and type words, and equip them with internationally common or officially recommended standardized translations. For example, "road" corresponds to "Road", "district" corresponds to "District", "building" corresponds to "Building", and "number" corresponds to "No.". The construction of the proper name library focuses on collecting common and representative proper nouns, such as city names (Beijing, Shanghai), famous streets (Chang'an Avenue, Nanjing Road), well-known landmarks (Oriental Pearl Tower, The Forbidden City), etc., and stores their widely accepted transliteration results (such as "Beijing" and "Nanjing Lu" based on Chinese pinyin) or existing standard translations (such as "The Forbidden City"). The establishment of the proper name library and the common name library can rely on official geographical information data, map service provider data, gazetteers, and various methods such as corpus mining and manual proofreading to ensure that the frequently occurring and standard-translated components in geographical names and addresses can be accurately and normatively translated, thereby significantly improving the lower limit of the overall translation quality and providing key clues for solving segmentation ambiguity problems.
[0021] In the above method for translating place names and addresses by intelligent splitting and matching based on a proper name library and a common name library, in step S2, the Chinese place name and address to be translated input by the user is obtained. Specifically, in actual application scenarios, the sources of place names and addresses are diverse (such as government forms, logistics waybills, navigation inputs), and there are problems such as chaotic formats, omission of common names, and colloquial expressions. For example, the user may input "No. 175, Inner Street, Chaoyangmen, Beijing" or abbreviate it as "Inner Chaoyangmen 175". Therefore, in order to be compatible with multi-scenario inputs and adapt to the subsequent structured parsing process, this application provides a standardized input for the subsequent translation process through a unified coding conversion and noise filtering mechanism. In the specific implementation process, the system needs to design a multi-modal input interface (supporting text, voice, and image OCR recognition), perform unified coding (such as full-width to half-width conversion) on the input content, and complete the address elements (such as supplementing the missing provincial and municipal prefixes). At the same time, an address length threshold control is established to avoid the decrease in processing efficiency caused by overly long addresses. In this way, the unstructured user input can be converted into a standardized string suitable for machine processing, laying a foundation for subsequent address segmentation and matching.
[0022] In the specific implementation process, the original inputs from different channels are received through the multi-modal interface, and preprocessing operations such as unified coding, noise filtering, and address element completion are performed on them, so as to convert the unstructured user input into a structured and standardized string format. Among them, the design of the multi-modal input interface is crucial. It not only supports the traditional text input method but also is compatible with voice recognition and image OCR (Optical Character Recognition) technologies, enabling users to input addresses in multiple ways. For example, in the logistics industry, couriers may enter the recipient address by voice; in mobile navigation applications, users may upload pictures containing address information by taking pictures. In order to adapt to these different input methods, the system needs to build corresponding parsing modules to perform semantic extraction and format normalization on the received voice text or OCR recognition results respectively, ensuring that all inputs can be uniformly processed.
[0023] After supporting various input methods, the next step is to perform standard conversion on the original input content. This process mainly includes converting full-width characters to half-width characters, standardizing punctuation marks, removing duplicate words, etc. For example, when a user enters "Jianguomenwai Street, Chaoyang District, Beijing", they may use the full-width number "175" or have extra spaces and line breaks in the address. Without processing, these formatting differences will directly affect the accuracy of subsequent segmentation and matching. Therefore, the system needs to introduce a complete encoding conversion rule library to identify and replace various non-standard characters with standard formats. In addition, for some colloquial expressions in user input, such as "near Jianguomen" or "over there by Dongzhimen", the system also needs to combine the context to determine the possible formal address names they refer to and try to map them to standardized expressions so that the subsequent processes can accurately identify the general and specific name components.
[0024] Meanwhile, to improve the integrity and consistency of address recognition, this step also needs to perform address element completion operations. In practical applications, users often omit some hierarchical information in the address for convenience. For example, they may only enter "No. 175, Chaoyangmennei Street" and omit the provincial prefix "Beijing", or only write "Chaoyangmennei 175" without indicating the street level. Although this abbreviated form does not cause much ambiguity in human understanding, it may lead to recognition failures or misjudgments for machine processing. Therefore, after obtaining the user input, the system should combine the geographical location knowledge graph and historical data to automatically supplement the missing administrative division level information. For example, when detecting the word segment "Chaoyangmennei Street", the system can infer the district it belongs to based on the existing geographical database and automatically add "Dongcheng District, Beijing" as the prefix to construct a complete address structure. Such operations not only help improve the accuracy of subsequent segmentation and matching but also ensure the standardization of the final translation result.
[0025] In addition, during the address input process, it is inevitable to encounter some noise information, such as spelling mistakes, redundant words, and irrelevant symbols. For example, users may include explanatory statements such as "Please deliver as soon as possible" or "Pay attention to the house number" in the address, or due to typing errors, there may be non-standard expressions such as "Chaoyangmen" or "Jianguomenwai Street". To address these issues, the system needs to integrate a natural language processing module to perform semantic analysis and anomaly detection on the input content. For words that are clearly not part of the address category, such as greetings or directive phrases, the system should automatically remove them; for suspected misspelled place names, they can be compared with a candidate word library through a fuzzy matching algorithm to try to correct them to the correct form. For example, if the user enters "Chaoyangmen", the system can recognize its high similarity to "Chaoyangmen" and determine that it should be "Chaoyangmen" based on the context, and then correct it to the correct form before proceeding to the next step.
[0026] Considering the complexity and diversity of address input, this step also needs to set a certain length control mechanism to avoid the decline in processing efficiency caused by overly long addresses. In some special scenarios, users may input detailed addresses with multiple levels, such as "Intelligent System Laboratory, Building 3, 2nd Floor, Area A, Institute of Automation, Chinese Academy of Sciences, No. 6A, Zhongguancun Street, Haidian District, Beijing". Although such addresses are complete in information, their lengths far exceed the range of regular addresses, which may affect the subsequent processing speed and resource consumption. Therefore, the system needs to set a reasonable address length threshold and trigger an automatic truncation or prompt mechanism when the threshold is exceeded to guide the user to input a more concise version. At the same time, through address compression technology, the lengthy descriptive content can be simplified into a standard format. For example, "Intelligent System Laboratory, Building 3, 2nd Floor, Area A, Institute of Automation, Chinese Academy of Sciences" can be simplified to "Institute of Automation, Building 3, Area A", thus improving the processing efficiency while retaining the key information.
[0027] In the above method for translating geographical name addresses by intelligent splitting and matching based on a proper name library and a common name library, in step S3, based on the proper name library and the common name library, multiple rounds of word segmentation annotation and optimal segmentation scheme screening are performed on the Chinese geographical name address to be translated to obtain the optimal splitting and matching result. Specifically, since Chinese geographical name addresses are usually continuous strings without natural delimiters such as spaces, it is easy to cause segmentation ambiguity problems, which directly affect the recognition and translation of subsequent components. At the same time, different segmentation methods will correspond to different knowledge base matching results, thus affecting the accuracy of translation. Therefore, to solve the segmentation ambiguity of Chinese geographical name addresses and make the best use of the high-quality knowledge in the constructed proper name library and common name library, this application, based on the principle of maximum coverage of the knowledge base, performs multiple rounds of word segmentation attempts with different granularities on the Chinese geographical name address to be translated, combines with the knowledge base for annotation, and then screens out the segmentation scheme that can best reflect the address structure and is most supported by the knowledge base. Among them, Figure 3 is a flowchart of sub-step S3 of the method for translating geographical name addresses by intelligent splitting and matching based on a proper name library and a common name library according to an embodiment of the present application. As Figure 3As shown, step S3 includes the steps of: S31, performing multi-round word segmentation processing on the Chinese place name and address to be translated to obtain a candidate set of sequences of Chinese place name and address word segments to be translated; S32, based on the proper name library and the common name library, performing word segment query and type annotation on each sequence of Chinese place name and address word segments in the candidate set of sequences of Chinese place name and address word segments to be translated to obtain a candidate set of sequences of Chinese place name and address word segments to be translated with annotations; S33, calculating the proportion value of the annotated characters in the total characters in each sequence of Chinese place name and address word segments with annotations in the candidate set of sequences of Chinese place name and address word segments with annotations, and extracting the sequence of Chinese place name and address word segments with annotations with the highest proportion value as the optimal splitting and matching result.
[0028] Specifically, in step S31, multi-round word segmentation processing is performed on the Chinese place name and address to be translated to obtain a candidate set of sequences of Chinese place name and address word segments to be translated. That is, for the obtained Chinese place name and address to be translated (such as "No. 1 Courtyard, Zhongguancun Street, Haidian District, Beijing"), various word segmentation strategies or algorithms (for example, a word segmenter based on a statistical language model, word segmentation under different dictionary combinations, etc.) are used for processing to generate a candidate set containing various possible segmentation results. For example, it may be possible to obtain candidate 1: ["Beijing", "City", "Haidian", "District", "Zhongguancun", "Street", "1", "No.", "Courtyard"], candidate 2: ["Beijing City", "Haidian District", "Zhongguancun Street", "No. 1 Courtyard"], candidate 3: ["Beijing City", "Haidian District", "Zhongguancun", "Street", "No. 1", "Courtyard"], etc.
[0029] Specifically, in step S32, based on the proper name library and the common name library, word segment query and type annotation are performed on each sequence of Chinese place name and address word segments in the candidate set of sequences of Chinese place name and address word segments to be translated to obtain a candidate set of sequences of Chinese place name and address word segments to be translated with annotations. Specifically, for each sequence of Chinese place name and address word segments in the candidate set, the proper name library and the common name library are queried word segment by word segment: If the word segment exists in the common name library (such as "City", "District", "Street", "No.", "Courtyard"), it is annotated as "common name" and its standard translation in the target translation language is attached; if it exists in the proper name library (such as "Beijing", "Haidian", "Zhongguancun"), it is annotated as "proper name" and its transliteration or standard translation is attached.
[0030] Specifically, in step S33, calculate the ratio of the number of labeled characters to the total number of characters in each labeled Chinese toponym address word segment sequence in the candidate set of the labeled Chinese toponym address word segment sequence, and extract the labeled Chinese toponym address word segment sequence with the highest ratio value as the optimal split matching result. That is, after the annotation is completed, calculate the total number of characters contained in all the successfully annotated (i.e., matched in the proper name library or the common name library) word segments in each Chinese toponym address word segment sequence to be translated, and then divide it by the total number of characters of the original input Chinese toponym address to be translated to obtain the ratio of the number of labeled characters to the total number of characters, that is, the knowledge base coverage. Then, from all the candidate sequences, extract and select the labeled Chinese toponym address word segment sequence with the highest ratio value as the optimal split matching result. In this way, the limitations that may be brought by a single word segmentation algorithm can be effectively overcome. Through multi-scheme comparison and knowledge base-based evaluation, the segmentation scheme that best conforms to the structural characteristics of toponym addresses and can be best explained by known knowledge is selected, thereby greatly improving the recognition accuracy and translation quality of the core components of toponym addresses and laying a solid foundation for subsequent processing of unannotated segments.
[0031] In the above-mentioned toponym address translation method based on intelligent split matching of proper name libraries and common name libraries, in step S4, perform attribute speculation and translation generation based on context association constraints on the unannotated Chinese toponym address word segments in the optimal split matching result to obtain the translation result of the unannotated toponym address word segments. It should be understood that due to the limited coverage of proper name libraries and common name libraries, they cannot cover all emerging toponyms, specific store names, uncommon building names, or descriptive attributives, etc. As a result, in the optimal split matching result, there will always be word segments that cannot be successfully annotated by the proper name library or the common name library (i.e., out-of-vocabulary words or OOV segments). If a general translation engine is simply used for translation processing, due to the lack of understanding of the specific context of toponym addresses by the general engine, it is impossible to accurately judge whether these segments should be transliterated, translated literally, or used as descriptive adjuncts, and the translation effect is often not good. Therefore, in order to break through the dictionary coverage limit and achieve semantically coherent translation, the present application further uses advanced natural language processing technology to perform context semantic association analysis on the optimal split matching result, and uses the position of the unannotated word segment in the overall address, the semantic relationship between the context address word segments, and the attribute annotation information to speculate on the possible semantic categories and attributes of the unannotated word segment, and thus perform translation on the unannotated Chinese toponym address word segment on this basis. Among them, Figure 4 It is a flowchart of sub-step S4 of the toponym address translation method based on intelligent split matching of proper name libraries and common name libraries according to an embodiment of the present application. As Figure 4As shown, step S4 includes the steps of: S41, performing word-level semantic embedding encoding on the optimal split matching result to obtain a sequence of semantic embedding encoding vectors of the to-be-translated place name and address word segments; S42, extracting the semantic embedding encoding vectors of the to-be-translated Chinese place name and address word segments corresponding to the unlabeled to-be-translated place name and address word segments from the sequence of semantic embedding encoding vectors of the to-be-translated place name and address word segments as the semantic embedding encoding vectors of the attribute query place name and address word segments; S43, performing semantic query response encoding based on context association constraints on the semantic embedding encoding vectors of the attribute query place name and address word segments and the sequence of semantic embedding encoding vectors of the to-be-translated place name and address word segments to obtain the context semantic query response encoding vectors of the attribute query place name and address word segments; S44, determining the type speculation result of the unlabeled to-be-translated Chinese place name and address word segments based on the context semantic query response encoding vectors of the attribute query place name and address word segments; S45, inputting the target translation language, the unlabeled to-be-translated Chinese place name and address word segments, and their type speculation results into a large language model to obtain the translation results of the unlabeled place name and address word segments.
[0032] Specifically, in a specific example of the present application, the step S41 includes: inputting the optimal split matching result into a BERT model containing positional encoding to obtain a sequence of semantic embedding encoding vectors of the to-be-translated geographical name and address word segments. Specifically, considering that the original Chinese geographical name and address word segments are discrete text symbols, it is difficult for a computer to directly understand their deep semantics and the complex relationships between segments. Therefore, in order to accurately capture the meanings of each to-be-translated geographical name and address word segment in the specific address context, and thus provide high-quality feature representations for subsequent attribute speculation and translation generation, based on the powerful semantic understanding ability principle of the pre-trained language model, the present application inputs the optimal split matching result (i.e., the sequence of to-be-translated Chinese geographical name and address word segments with annotations) into a BERT model containing positional encoding to achieve context awareness and semantic representation of each to-be-translated Chinese geographical name and address word segment, and obtain a sequence of semantic embedding encoding vectors of the to-be-translated geographical name and address word segments. In the specific implementation process, the BERT model first converts each to-be-translated geographical name and address word segment and its annotated attribute information (if any) in the optimal split matching result into a low-dimensional dense word vector representation through a word embedding layer, and adds positional encoding to each word vector through a positional embedding layer, so that the model can distinguish the relative positions of different word segments in the address. Subsequently, the word vectors fused with positional information are input into the Transformer structure of the BERT model for multi-layer self-attention mechanism processing. In each layer, the model dynamically adjusts the representation of each word segment by calculating the attention weights between different word segments, so that it can capture the semantic associations with other word segments. After processing through multiple Transformer structures, the semantic embedding representation of each to-be-translated geographical name and address word segment incorporates not only its own semantics but also its position in the overall address and the semantic relationships with surrounding word segments, providing a rich feature basis for subsequent attribute speculation and translation generation.
[0033] Specifically, in the step S42, the semantic embedding encoding vectors of the to-be-translated Chinese geographical name and address word segments corresponding to the unannotated to-be-translated Chinese geographical name and address word segments are extracted from the sequence of semantic embedding encoding vectors of the to-be-translated geographical name and address word segments as the semantic embedding encoding vectors of the attribute query geographical name and address word segments. That is, in order to focus on the unannotated segments that need to be subject to attribute speculation and translation generation, based on the principle of target-driven data screening, the present application extracts the semantic embedding encoding vectors of the to-be-translated geographical name and address word segments corresponding to the unannotated segments from the complete sequence of semantic embedding encoding vectors of the to-be-translated geographical name and address word segments to achieve targeted capture of target features, facilitating subsequent targeted processing.
[0034] Specifically, the step S43 performs semantic query response encoding based on context association constraints on the sequence of the semantic embedding encoding vector of the attribute query place name and address word fragment and the semantic embedding encoding vector of the place name and address word fragment to be translated to obtain the attribute query place name and address word fragment context semantic query response encoding vector. It should be understood that since the attribute inference of the unlabeled place name and address word fragment needs to integrate the synergy of its own semantics and the context semantic field (such as when the "county" in "Saga County" is a common name, "Saga" is more likely to be an overall proper name). Therefore, in order to construct a quantitative propagation channel of the semantic influence between place name and address word fragments to model the semantic transmission relationship of place name and address, the present application further models the semantic association structure in the sequence of semantic embedding coding vectors of the place name and address word fragments to be translated based on graph coding technology, and queries the place name and address global semantic association graph by using the attribute query place name and address word fragment semantic embedding coding vector to mine the potential semantic connection between the unlabeled place name and address word fragments and the surrounding labeled place name and address word fragments, thereby utilizing the semantic transmission mechanism of graph coding to enhance the semantic representation of the unlabeled place name and address word fragments, and obtain the attribute query place name and address word fragment context semantic query response coding vector. Among them, Figure 5 Flow chart of sub-step S43 of the place name and address translation method based on intelligent splitting and matching of the specific name library and the common name library according to the embodiment of the present application. Figure 5 As shown, the step S43 includes the steps of: S431, performing context association analysis between word fragments on the sequence of semantic embedding coding vectors of the place name and address word fragments to be translated to construct a topological matrix of association between semantic features of the place name and address word fragments to be translated; S432, based on the topological matrix of association between semantic features of the place name and address word fragments to be translated, performing graph structure context association perception on the sequence of semantic feature vectors of the place name and address word fragments to be translated to obtain a graph-like association coding matrix of semantic features of the place name and address word fragments to be translated; S433, performing semantic query response coding on the semantic embedding coding vector of the attribute query place name and address word fragment and the graph-like association coding matrix of the semantic features of the place name and address word fragments to be translated to obtain the context semantic query response coding vector of the attribute query place name and address word fragment.
[0035] More specifically, the step S431 includes: first, performing semantic condensation coding on each semantic feature vector of the place name and address word fragment to be translated in the sequence of semantic feature vectors of the place name and address word fragment to be translated to obtain a sequence of semantic feature condensation coding vectors of the place name and address word fragment to be translated, which is expressed by the formula: ; ; ; in, A sequence representing the semantic feature vectors of the to-be-translated geographical name and address word segments, , , and respectively represent the 1st, 2nd, th, and th semantic feature vectors of the to-be-translated geographical name and address word segments in the sequence representing the semantic feature vectors of the to-be-translated geographical name and address word segments, is the number of vectors in the sequence representing the semantic feature vectors of the to-be-translated geographical name and address word segments, represents calculating the 2-norm of the vector, represents the ReLu activation function, and respectively represent the weight matrix and bias term of the semantic condensation network, represents the sequence of semantic feature condensation encoding vectors of the to-be-translated geographical name and address word segments, , , and respectively represent the 1st, 2nd, th, and th semantic feature condensation encoding vectors of the to-be-translated geographical name and address word segments in the sequence representing the semantic feature condensation encoding vectors of the to-be-translated geographical name and address word segments.
[0036] That is, by semantic condensation encoding, irrelevant information is removed, the semantic feature vectors of each to-be-translated geographical name and address word segment are projected and refined into a low-dimensional semantic space, and its most representative core semantics are extracted, thereby improving the accuracy of subsequent semantic relatedness calculation. The sequence of semantic feature condensation encoding vectors of the to-be-translated geographical name and address word segments obtained after semantic condensation encoding realizes low-dimensionalization and semantic refinement, retains the core semantic information of the word segments in a more concise and efficient form, and provides a higher-quality feature basis for subsequent semantic relatedness analysis.
[0037] Then, calculate the semantic relatedness between any two semantic feature condensation encoding vectors of the to-be-translated geographical name and address word segments in the sequence of semantic feature condensation encoding vectors of the to-be-translated geographical name and address word segments to obtain the associated topological matrix of the semantic features of the to-be-translated geographical name and address word segments composed of multiple semantic relatednesses, which is expressed by the formula: ; where represents the th semantic feature condensation encoding vector of the to-be-translated geographical name and address word segments in the sequence of semantic feature condensation encoding vectors of the to-be-translated geographical name and address word segments, and respectively represent the semantic relatedness feature weight matrix and semantic relatedness feature bias term, represents and the semantic association coding vector between represents the semantic association coding vector of the characteristic scale value represents the semantic association coding vector the index of each position feature value in represents and the semantic association degree between, that is, the element value at the position in the association topology matrix of the semantic features of the to-be-translated place name and address word segment
[0038] That is, by quantifying the semantic association degree between the semantic feature condensed coding vectors of the to-be-translated place name and address word segment, an initial relationship graph is explicitly constructed. The association topology matrix of the semantic features of the to-be-translated place name and address word segment generated based on this can capture the deep semantic relationship between word segments, improve the recognition ability of the specific proper name and common name combination in the address, provide a structured semantic relationship basis for the subsequent graph structure context association perception, and solve the problem that the traditional word segmentation model lacks targeted recognition ability for the specific combination of place names and addresses
[0039] More specifically, the step S432 includes: First, input the association topology matrix of the semantic features of the to-be-translated place name and address word segment into a gated mask network to obtain a sparse association topology matrix between the semantic feature nodes of the to-be-translated place name and address word segment, which is expressed by the formula: ; ; wherein represents the sigmoid activation function represents the transpose of the vector represents element-wise division and respectively represent the mask weight vector and the mask bias weight parameter represents and the semantic association mask weight between the to-be-translated place name and address word segment represents the element value at the position in the sparse association topology matrix between the semantic feature nodes of the to-be-translated place name and address word segment
[0040] That is, by introducing a gating mask network, an adaptive sparsification process is performed on the correlation topology matrix of the semantic features of the to-be-translated place name and address word segments, so as to extract the most critical semantic connections and optimize the graph structure. The sparse correlation topology matrix between the semantic feature nodes of the to-be-translated place name and address word segments generated in this way significantly reduces the computational complexity of subsequent graph convolution, improves the robustness of the model to noise, effectively prevents overpropagation of information (the over-smoothing problem), and highlights the truly important semantic relationships, making the topological structure of the graph clearer, more robust, and closer to the real underlying feature manifold.
[0041] Then, a graph-spectrum encoding based on the graph convolutional network model is performed on the sequence of the semantic feature condensation encoding vectors of the to-be-translated place name and address word segments and the sparse correlation topology matrix between the semantic feature nodes of the to-be-translated place name and address word segments to obtain the semantic feature graph-spectrum correlation encoding matrix of the to-be-translated place name and address word segments, which is expressed by the formula: ; Where represents the sparse correlation topology matrix between the semantic feature nodes of the to-be-translated place name and address word segments, represents the graph convolutional network, represents the semantic feature graph-spectrum correlation encoding matrix of the to-be-translated place name and address word segments.
[0042] That is, by using the iterative neighborhood information aggregation and transformation mechanism of the graph convolutional network model, a structured depth embedding is performed on the sparse correlation topology matrix between the semantic feature nodes of the to-be-translated place name and address word segments and the semantic feature condensation encoding vectors of the to-be-translated place name and address word segments, so as to capture the high-order semantic relationships and global context correlations that go beyond simple similarity between word segments, enabling the representation of each word segment to fuse the multi-hop neighborhood structure information in the feature manifold. In this way, the generated semantic feature graph-spectrum correlation encoding matrix of the to-be-translated place name and address word segments realizes the holistic and structured depth representation of the semantic features of the to-be-translated place name and address word segments. Each row vector in the matrix contains the position and semantic role of the corresponding word segment under context awareness, effectively capturing the high-order semantic dependencies and global structure information between word segments, and providing a richer context semantic basis for subsequent semantic query response encoding.
[0043] Specifically, considering that the graph neural network model has achieved multi-order association and extraction of overall architecture information under the graph structure for the semantic feature imitation graph association coding matrix of the to-be-translated geographical name and address word segments, but due to the low-density characteristics introduced by the gated mask network in the graph topology, for each row vector in the semantic feature imitation graph association coding matrix of the to-be-translated geographical name and address word segments, it is still expected that they show convergent statistical characteristics as much as possible on the high-dimensional macroscopic distribution scale outside the graph structure, that is, it is expected to achieve the macroscopic comprehensive association of each row vector. Based on this, in a preferred example of the present application, the step S422123 includes: First, perform coupling optimization based on maintaining the global semantic field balance for each row vector in the semantic feature imitation graph association coding matrix of the to-be-translated geographical name and address word segments to obtain an optimized semantic feature imitation graph association coding matrix of the to-be-translated geographical name and address word segments.
[0044] Specifically, first, use the measure semi-ring space to calculate the neighborhood rigidity mapping function of each row vector in the semantic feature imitation graph association coding matrix of the to-be-translated geographical name and address word segments to capture complex non-linear associations that quasi-surpass covariance characterization: ; Wherein, represents the th row vector in the semantic feature imitation graph association coding matrix of the to-be-translated geographical name and address word segments, represents calculating the absolute value, represents the mean vector obtained by averaging each row vector in the semantic feature imitation graph association coding matrix of the to-be-translated geographical name and address word segments by position, represents corresponding neighborhood rigidity mapping vector.
[0045] Then, calculate the time-varying cumulative index of the neighborhood rigidity mapping vector : ; Wherein, represents 's position eigenvalue, represents the natural logarithm function with base e, represents 's time-varying cumulative index.
[0046] Finally, use the time-varying cumulative index to perform weighted optimization on the row vector : ; Wherein, It represents the th row vector in the spectrum association coding matrix of the semantic feature imitation map for optimizing the semantic feature segment of the to-be-translated place name and address term.
[0047] That is, considering that the macroscopic statistical distribution law is dominated by high-order cumulants, in order to achieve the consistency of the high-dimensional macroscopic distribution scale among different row vectors, it is necessary to constrain the global symmetric dynamics through the dynamic accumulation measure under the rigid neighborhood framework, so as to ensure the uniform transmission of macroscopic statistical characteristics under the graph structure and avoid effects such as statistical disconnection under the graph structure.
[0048] Then, perform feature interaction attention aggregation based on the feature query response intensity on the semantic embedding coding vector of the attribute query place name and address term segment and each row vector in the spectrum association coding matrix of the semantic feature imitation map for optimizing the semantic feature segment of the to-be-translated place name and address term to obtain the context semantic query response coding vector of the attribute query place name and address term segment, which is expressed by the formula: ; ; Among them, represents the semantic embedding coding vector of the attribute query place name and address term segment, represents the th row vector in the spectrum association coding matrix of the semantic feature imitation map for optimizing the semantic feature segment of the to-be-translated place name and address term, represents and the semantic interaction response attention weight between them, represents element-wise multiplication by position, represents the context semantic query response coding vector of the attribute query place name and address term segment.
[0049] That is, introduce the attention mechanism to realize the interaction between the semantic embedding coding vector of the attribute query place name and address term segment and the spectrum association coding matrix of the semantic feature imitation map for optimizing the semantic feature segment of the to-be-translated place name and address term, that is, weighted aggregation of relevant node embeddings based on the feature query response intensity, integrating the query information and the semantic information of the to-be-translated word segment after deep relationship modeling, so as to accurately locate and express the association between the query intention and the to-be-translated content in the complex semantic space. Based on this, the generated context semantic query response coding vector of the attribute query place name and address term segment can accurately reflect the location and relationship of the attribute query word segment in the structured semantic space defined by the to-be-translated word segment, dynamically capture the context dependence between word segments, and provide accurate semantic basis for the attribute speculation and translation generation of unlabeled word segments.
[0050] Specifically, in a specific example of the present application, the step S44 includes: inputting the context semantic query response encoding vector of the attribute query geographical name address word segment into the attribute speculation module based on a classifier to obtain the type speculation result. It should be understood that the context semantic query response encoding vector of the attribute query geographical name address word segment contains the semantic features of the unlabeled geographical name address word segment in the overall address context and its association information with surrounding word segments. Based on this, in order to accurately speculate on the attributes of the unlabeled geographical name address word segment, the present application designs a classifier based on a neural network architecture to perform feature learning and attribute speculation on the context semantic query response encoding vector of the attribute query geographical name address word segment. The classifier is trained on a large amount of geographical name address data, can effectively learn the semantic features of the geographical name address word segment, and establish a mapping relationship between the semantic information of the geographical name address word segment and the attribute label. During the attribute speculation process, the classifier extracts features and performs pattern recognition on the input context semantic query response encoding vector of the attribute query geographical name address word segment based on the prior knowledge learned through training, so as to accurately determine the type of the unlabeled geographical name address word segment, such as proper name / general name and road name / area name / building type, etc., thereby providing key guiding information for subsequent translation processing and ensuring the accuracy of translation.
[0051] Specifically, in the step S45, the target translation language, the unlabeled Chinese geographical name address word segment to be translated, and its type speculation result are input into a large language model to obtain the translation result of the unlabeled geographical name address word segment. Specifically, based on the prompting learning and text generation principle of a large language model (LLM), the present application drives the LLM to generate the final translation result by jointly using the original text of the unlabeled segment, its speculated attributes, and the target translation language type as inputs. In the specific implementation process, assume that the unlabeled segment is "Qinghe", its speculated attribute is "proper name - road name", and the target translation language is English. Then it is filled into a preset template prompt to construct a prompt for the large language model: "Please translate the Chinese geographical name segment 'Qinghe' into English. The attribute of this segment is 'proper name - road name'. The large language model (such as the GPT series) translates the geographical name segment 'Qinghe' according to the input prompt information, utilizes its own language understanding and generation capabilities, and generates a more idiomatic English expression considering its attribute as 'proper name - road name', such as 'Qinghe', thereby ensuring the accuracy and idiomaticity of the translation. In this way, guided by the attribute information of the word segment and combined with the powerful language generation capabilities of the large language model, the accuracy of translation can be effectively improved, and the problem of ambiguity caused by different attributes of the same geographical name in different contexts can be avoided. For example, "Xingfu" is translated as "Xingfu" when it is used as a road name, while it may be translated as "Happiness" when it is used as a community name or scenic spot name.
[0052] In the above-mentioned method for translating geographical names and addresses with intelligent splitting and matching based on a proper name library and a common name library, in step S5, the translation results of the unannotated geographical name and address word segments and the translation results of the annotated Chinese geographical name and address word segments in the optimal splitting and matching results are assembled and normalized based on a rule engine to obtain the complete Chinese geographical name and address translation result. Specifically, in order to meet the requirements of the Geographic Information System (GIS) for a standardized address format, the present application further performs a structural reorganization of the translated name components to calibrate the grammar and format of the translation results through a rule engine. In the specific implementation process, the rule engine needs to build in three types of specifications: 1) Word order rules (e.g., "X Road, No. Y" in Chinese is translated as "Y Road No. X"); 2) Component connector rules (e.g., add a space between a proper name and a common name, and use "No." for numbers); 3) Capitalization rules (e.g., capitalize the first letter of a common name). For example, for the segmentation result "Haidian District (Haidian District), North West 3rd Ring Road (North West 3rd Ring Road), No. A2 (A2)", the engine will reorganize it into "A2, North West 3rd Ring Road, Haidian District, Beijing" according to the English expression habit of "from large range to small range". At the same time, the engine will detect and repair the problem of inconsistent abbreviations between "Road" and "Rd.". In this way, it can be ensured that the finally output translation result not only conforms to the grammar norms of the target language (e.g., the English address is arranged from small to large), but also meets the requirements of the GIS system for standardized naming (e.g., facilitating database indexing and spatial matching), significantly improving machine readability.
[0053] In summary, the method for translating geographical names and addresses with intelligent splitting and matching based on a proper name library and a common name library according to the embodiments of the present application is clarified. First, a common name library containing standard translated names and a proper name library of transliterated proper names are constructed. After multiple rounds of segmentation of the geographical name and address to be translated, the word segments in each segmentation scheme are type-annotated through double-library matching, and the optimal segmentation scheme is determined based on the annotation character coverage. On this basis, a context correlation inference mechanism is further introduced to perform context semantic correlation analysis on the unannotated word segments in the optimal segmentation scheme. After inferring the word segment attribute type, a large language model is called to generate an adapted translation. Finally, the rule engine performs a structural reorganization on the double-library annotated translated names and the inferred translations to form a standardized standard address translation. Through multiple rounds of word segmentation annotation and screening of the optimal segmentation scheme, this method can effectively solve the problem of segmentation ambiguity, accurately identify the core components in the address, and thus improve the translation quality.
[0054] The basic principles of the present invention have been described above in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present invention. Additionally, the specific details of the above embodiments are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present invention to necessarily adopting the above specific details for implementation.
[0055] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the unit division is only a logical function division, and there can be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0056] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any associated drawing reference signs in the claims should not be regarded as limiting the claimed rights.
[0057] In addition, obviously the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.
[0058] Finally, it should be noted that the above description has been given for the purposes of illustration and description. In addition, the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library, characterized in that, Including: Construct a proper noun library and a common noun library. Among them, the common noun library stores common suffixes, type words in geographical names and addresses, and their corresponding standard translations, and the proper noun library stores common proper nouns in geographical names and addresses, and their corresponding transliteration results or standard translations; Obtain the Chinese geographical name and address to be translated input by the user; Based on the proper noun library and the common noun library, perform multiple rounds of word segmentation annotation and optimal segmentation scheme screening on the Chinese geographical name and address to be translated to obtain an optimal splitting and matching result; Perform attribute speculation and translation generation based on context association constraints on the unannotated Chinese geographical name and address word fragments in the optimal splitting and matching result to obtain the translation result of the unannotated geographical name and address word fragments; Assemble the translation result of the unannotated geographical name and address word fragments and the translation result of the annotated Chinese geographical name and address word fragments in the optimal splitting and matching result, and perform normalization processing based on a rule engine to obtain the complete translation result of the Chinese geographical name and address.
2. The method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library according to claim 1, wherein Based on the proper noun library and the common noun library, perform multiple rounds of word segmentation annotation and optimal segmentation scheme screening on the Chinese geographical name and address to be translated to obtain an optimal splitting and matching result, including: Perform multiple rounds of word segmentation on the Chinese geographical name and address to be translated to obtain a candidate set of sequences of Chinese geographical name and address word fragments; Based on the proper noun library and the common noun library, perform word fragment query and type annotation on each sequence of Chinese geographical name and address word fragments in the candidate set of sequences of Chinese geographical name and address word fragments to obtain a candidate set of sequences of annotated Chinese geographical name and address word fragments; Calculate the proportion value of the annotated characters in each sequence of annotated Chinese geographical name and address word fragments in the candidate set of sequences of annotated Chinese geographical name and address word fragments, and extract the sequence of annotated Chinese geographical name and address word fragments with the highest proportion value as the optimal splitting and matching result.
3. The method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library according to claim 1, characterized in that, Perform attribute speculation and translation generation based on context association constraints on the unannotated Chinese geographical name and address word fragments in the optimal splitting and matching result to obtain the translation result of the unannotated geographical name and address word fragments, including: Perform word-level semantic embedding encoding on the optimal splitting and matching result to obtain a sequence of semantic embedding encoding vectors of geographical name and address word fragments to be translated; Extract the semantic embedding encoding vector of the geographical name and address word fragment to be translated corresponding to the unannotated Chinese geographical name and address word fragment from the sequence of semantic embedding encoding vectors of geographical name and address word fragments to be translated as the semantic embedding encoding vector of the attribute query geographical name and address word fragment; Perform semantic query response encoding based on context association constraints on the semantic embedding encoding vector of the attribute query geographical name and address word fragment and the sequence of semantic embedding encoding vectors of geographical name and address word fragments to be translated to obtain the context semantic query response encoding vector of the attribute query geographical name and address word fragment; Based on the context semantic query response encoding vector of the attribute query geographical name and address word fragment, determine the type speculation result of the unannotated Chinese geographical name and address word fragment; Input the target translation language, the unannotated Chinese place name and address word segments to be translated, and their type speculation results into a large language model to obtain the translation results of the unannotated place name and address word segments.
4. The method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library according to claim 3, wherein Perform word-level semantic embedding encoding on the optimal split and match results to obtain a sequence of semantic embedding encoding vectors for the place name and address word segments to be translated, including: Input the optimal split and match results into a BERT model with positional encoding to obtain a sequence of semantic embedding encoding vectors for the place name and address word segments to be translated.
5. The method for translating place names and addresses by intelligent splitting and matching based on a proper name library and a common name library according to claim 4, characterized in that Perform semantic query response encoding based on context association constraints on the semantic embedding encoding vectors of the attribute query place name and address word segments and the sequence of semantic embedding encoding vectors of the place name and address word segments to be translated to obtain context semantic query response encoding vectors for the attribute query place name and address word segments, including: Perform inter-word segment context association analysis on the sequence of semantic embedding encoding vectors of the place name and address word segments to be translated to construct an association topology matrix between the semantic features of the place name and address word segments to be translated; Based on the association topology matrix between the semantic features of the place name and address word segments to be translated, perform graph-structured context association perception on the sequence of semantic feature vectors of the place name and address word segments to be translated to obtain an association encoding matrix of the semantic features of the place name and address word segments to be translated that mimics a graph spectrum; Perform semantic query response encoding on the semantic embedding encoding vectors of the attribute query place name and address word segments and the association encoding matrix of the semantic features of the place name and address word segments to be translated that mimics a graph spectrum to obtain the context semantic query response encoding vectors for the attribute query place name and address word segments.
6. The method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library according to claim 5, characterized in that, Perform inter-word segment context association analysis on the sequence of semantic embedding encoding vectors of the place name and address word segments to be translated to construct an association topology matrix between the semantic features of the place name and address word segments to be translated, including: Perform semantic condensation encoding on each semantic feature vector of the place name and address word segments to be translated in the sequence of semantic feature vectors of the place name and address word segments to be translated to obtain a sequence of semantic feature condensation encoding vectors for the place name and address word segments to be translated; Calculate the semantic association degree between any two semantic feature condensation encoding vectors of the place name and address word segments to be translated in the sequence of semantic feature condensation encoding vectors of the place name and address word segments to be translated to obtain the association topology matrix between the semantic features of the place name and address word segments to be translated composed of multiple semantic association degrees.
7. The method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library according to claim 6, wherein Based on the association topology matrix between the semantic features of the place name and address word segments to be translated, perform graph-structured context association perception on the sequence of semantic feature vectors of the place name and address word segments to be translated to obtain an association encoding matrix of the semantic features of the place name and address word segments to be translated that mimics a graph spectrum, including: Input the association topology matrix between the semantic features of the place name and address word segments to be translated into a gated mask network to obtain a sparse association topology matrix between the semantic feature nodes of the place name and address word segments to be translated; Perform graph spectrum encoding based on the graph convolutional network model on the sequence of semantic feature condensation encoding vectors of the place name and address word segments to be translated and the sparse association topology matrix between the semantic feature nodes of the place name and address word segments to be translated to obtain the association encoding matrix of the semantic features of the place name and address word segments to be translated that mimics a graph spectrum.
8. The method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library according to claim 7, characterized in that Semantic query response encoding is performed on the semantic embedding encoding vector of the place name and address word segment for the attribute query and the semantic feature imitation spectrum correlation encoding matrix of the to-be-translated place name and address word segment to obtain the context semantic query response encoding vector of the place name and address word segment for the attribute query, including: Performing coupling optimization based on maintaining the global semantic field balance on each row vector in the semantic feature imitation spectrum correlation encoding matrix of the to-be-translated place name and address word segment to obtain an optimized semantic feature imitation spectrum correlation encoding matrix of the to-be-translated place name and address word segment; Performing feature interaction attention aggregation based on the feature query response intensity on the semantic embedding encoding vector of the place name and address word segment for the attribute query and each row vector in the optimized semantic feature imitation spectrum correlation encoding matrix of the to-be-translated place name and address word segment to obtain the context semantic query response encoding vector of the place name and address word segment for the attribute query.
9. The method for translating geographical names and addresses by intelligent splitting and matching based on a proper name library and a common name library according to claim 3, characterized in that Based on the context semantic query response encoding vector of the place name and address word segment for the attribute query, determining the type speculation result of the unlabeled to-be-translated Chinese place name and address word segment, including: Inputting the context semantic query response encoding vector of the place name and address word segment for the attribute query into an attribute speculation module based on a classifier to obtain the type speculation result.
Citation Information
Patent Citations
Database query method and system based on self-attention syntactic perception
CN116910086A
Two-channel feature extraction named entity recognition method fusing attention mechanism in financial field
CN116956928A
LSTM (Long Short Term Memory)-based Asian geographical name and proper name automatic Chinese translation model and method
CN118070819A
Method for enhancing translation precision of place name address by using context awareness and geographic library
CN119990160A
Transformer deep learning model-based method for translating multilingual place name root into chinese
WO2022057116A1
Cited By
Intelligent image comparison method and system based on machine learning
CN120599295A