High-precision Place Name Translation Method Integrating Artificial Intelligence and Multilingual Syllable Segmentation
By integrating artificial intelligence and multilingual syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable syllable sy
Patent Information
- Application Number
- CN202510494841.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-21
AI Technical Summary
When facing multilingual place name names, traditional place name translation methods are difficult to achieve high-precision translation due to significant differences in syllable structure, morphological rules and writing systems, especially in the processing of place names in low- and medium-resource languages and mixed languages.
The method of integrating artificial intelligence and multilingual syllable segmentation is adopted to detect the source language through the language recognition model, dynamically adapt the syllable segmentation strategy, and the deep neural network model is used to perform syllable segmentation, and combined with word embedding and context encoding models, the context semantics of place name Tokens are reconstructed, the recessive syllable associations in compound words and adhesions are identified, and the local morphological characteristics are enhanced.
It realizes end-to-end effective conversion from original place names to target translation names, improves the accuracy and robustness of place name translation, especially in the processing of place name in medium and low resource languages and mixed languages, which significantly improves the translation effect.
Smart Images

Figure CN120012790B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of place name translation, and more specifically, to a high-precision place name translation method integrating artificial intelligence and multi-language syllable segmentation. Background Art
[0002] Place name translation is a key task in cross-language information processing, and its accuracy directly affects the application effects in fields such as geographic information systems, multi-language map navigation, and cross-border logistics. However, due to the significant differences in syllable structures, morphological rules, and writing systems among the world's languages, traditional translation methods face the following core challenges when dealing with multi-language place names: The syllable composition rules of different language families vary greatly. For example, Chinese is based on single-syllable Chinese characters with clear syllable boundaries; English relies on complex vowel-consonant combinations (such as "strengths" containing 7 phonemes) and has a large number of irregular spellings; agglutinative languages such as Finnish and Turkish form words by syllable superposition, with variable lengths and structures.
[0003] In addition, there are mature syllable segmentation rule libraries for high-resource languages (such as English and French), but low-resource languages (such as Vietnamese and Tibetan) lack labeled data, and the generalization ability of statistical models is limited. Moreover, in the context of globalization, place names often contain mixed language components (such as the mixture of English and Hindi in "New Delhi"), and a single translation strategy is difficult to accurately segment. These limitations in translation pose challenges to place name translation understanding and geographic information applications.
[0004] Therefore, it is necessary to provide a high-precision place name translation method integrating artificial intelligence and multi-language syllable segmentation to solve the above technical problems. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed.
[0006] According to one aspect of the present application, there is provided a high-precision place name translation method integrating artificial intelligence and multi-language syllable segmentation, which includes:
[0007] Receiving the place name string to be translated input by the user;
[0008] Using a language recognition model to detect the source language of the place name string to be translated;
[0009] Based on the source language of the place name string to be translated, determining a syllable segmentation strategy;
[0010] In response to the syllable segmentation strategy being a syllable segmentation method based on a deep neural network model, based on the syllable segmentation strategy, performing syllable segmentation on the place name string to be translated to obtain a syllable-segmented place name string;
[0011] Input the segmented geographical name string into the geographical name translation model to obtain the translated text of the geographical name in the target language.
[0012] This application has at least the following technical effects: Compared with the prior art, a high-precision geographical name translation method integrating artificial intelligence and multi-language syllable segmentation provided by this application first splits the input geographical name into a sub-word Token sequence, and uses the context encoding model of the word embedding model to realize the semantic context association encoding of the geographical name to be translated; Subsequently, by means of strengthening the reconstruction of local semantic relevance, the internal structured information and dependency relationships in the Token context semantics of the geographical name to be translated are modeled, the implicit syllable associations in compound words and agglutinative languages are identified, and the local morphological features (such as prefixes / suffixes, consonant clusters) near the syllable boundaries are enhanced to solve the problems of irregular spelling and cross-language interference, and then decode and output the segmented geographical name string for geographical name translation, generate the corresponding translated text of the geographical name in the target language, and realize the end-to-end effective conversion from the original geographical name to the target translated name. Brief Description of the Drawings
[0013] By describing the embodiments of the present application in more detail with reference to the accompanying drawings, the above and other objects, features and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0014] Figure 1 It is a flowchart of a high-precision geographical name translation method integrating artificial intelligence and multi-language syllable segmentation according to an embodiment of the present application;
[0015] Figure 2 It is a schematic diagram of data flow of a high-precision geographical name translation method integrating artificial intelligence and multi-language syllable segmentation according to an embodiment of the present application;
[0016] Figure 3 It is a flowchart of sub-step S4 of a high-precision geographical name translation method integrating artificial intelligence and multi-language syllable segmentation according to an embodiment of the present application. Detailed Description of the Embodiments
[0017] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0018] As shown in this application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. A method or device may also include other steps or elements.
[0019] Although this application makes various references to certain modules in the system according to the embodiments of this application, however, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0020] Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the operations before or below do not necessarily need to be performed precisely in sequence. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or several steps can be removed from these processes.
[0021] Next, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the exemplary embodiments described here.
[0022] It should be noted in advance that all the acquisition and processing of information or data in this application are carried out on the premise of complying with the corresponding national data protection regulations and policies and obtaining the authorization of the rights management.
[0023] Aiming at the problems in multilingual place name translation such as large differences in syllable rules, unbalanced data resources, and mixed language interference, in the technical solution of this application, a high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation is proposed. This solution adopts the design idea of "hierarchical decision-making - semantic-driven", dynamically adapts the syllable segmentation strategies of different language types, and optimizes the translation accuracy by combining context semantic understanding.
[0024] Based on this, in the technical solution of this application, a high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation is proposed. Figure 1 It is a flowchart of the high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to the embodiments of this application. Figure 2 It is a schematic diagram of the data flow of the high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to the embodiments of this application. As Figure 1 and Figure 2As shown, a high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to an embodiment of the present application includes the steps of: S1, receiving a place name string to be translated input by a user; S2, using a language recognition model to detect the source language of the place name string to be translated; S3, determining a syllable segmentation strategy based on the source language of the place name string to be translated; S4, in response to the syllable segmentation strategy being a syllable segmentation method based on a deep neural network model, performing syllable segmentation on the place name string to be translated based on the syllable segmentation strategy to obtain a syllable-segmented place name string; S5, inputting the syllable-segmented place name string into a place name translation model to obtain a place name target language translation text.
[0025] Specifically, in S1, a place name string to be translated input by a user is received. Among them, the place name string to be translated is usually composed of a series of characters representing the name of a specific geographical location. These characters can be letters, numbers, symbols, or a combination thereof. The place name string to be translated may contain single-language components or mixed-language components (for example, "New Delhi" combines English and Hindi elements). For a translation system, it is important to accurately identify the source language of the string and determine subsequent processing steps based on this. Specifically, in this process, a user-friendly interface can be designed to allow users to input the place name they wish to translate through various devices such as computers and mobile phones. This interface can be a form on a web page, an input box in a mobile application, etc.
[0026] Specifically, in S2, a language recognition model is used to detect the source language of the to-be-translated name string. Among them, the language recognition model is usually constructed based on deep learning technology and can automatically learn and identify the feature patterns of different languages. When the user inputs the to-be-translated name string, the string will be fed into a pre-trained language recognition model. The model will comprehensively analyze the input character sequence and extract the key information that can represent the specific language features. For example, for some languages with unique character sets (such as Chinese using Chinese characters and Russian using Cyrillic letters), the model can directly judge quickly according to the character type; while for those languages that use the same character set but have significant differences in grammar structure or vocabulary distribution (such as English and French), more complex statistical features and context information will be relied on for differentiation. Among them, the language recognition model can select advanced deep learning frameworks such as convolutional neural network (CNN), recurrent neural network (RNN) and its variants (such as long short-term memory network LSTM, gated recurrent unit GRU) or Transformer architecture. These models can effectively capture the local and global features in the input string and map them to a probability distribution space representing different language possibilities through multi-layer non-linear transformation. Finally, what the model outputs is a probability score for each known language, and the one with the highest score is considered to be the source language most likely corresponding to the input string.
[0027] Specifically, in step S3, based on the source language of the name string to be translated, a syllable segmentation strategy is determined. That is, according to the source language of the name string to be translated, the characteristics and rules of the language are evaluated. For example, since high-resource languages already have mature syllable segmentation rule libraries, established rules can be relatively directly applied to handle syllable boundary problems. Therefore, if a high-resource language such as English, French, or Spanish is detected, a rule-based syllable segmentation method is usually selected; among them, the rules may include how to handle special cases such as consonant-vowel combinations and irregular spellings, so as to ensure that the syllable structure within each word can be accurately identified; for those low- and medium-resource languages, such as Chinese, Japanese, Korean, and Vietnamese, etc., there may be insufficient labeled data to support rule-based methods. In this case, the system tends to adopt a syllable segmentation strategy based on a statistical model; for example, conditional random fields (CRFs), hidden Markov models (HMMs), or other models suitable for sequence labeling tasks can be used to automatically discover and apply syllable segmentation rules applicable to specific languages; when encountering place names with mixed language components, such as an example like "New Delhi" that contains both English and Hindi elements, a single language processing strategy often fails to meet the requirements. At this time, the system will select a syllable segmentation method based on a deep neural network model to be able to dynamically adapt to the interaction between different languages and capture more complex language phenomena through a deep architecture. Specifically, the deep neural network can learn the commonalities and differences between multiple languages through training, so as to maintain a high accuracy when facing cross-language interference. For example, when dealing with compound words or agglutinative morphemes, the deep neural network can identify the implicit syllable associations between roots and suffixes and enhance the local morphological features near syllable boundaries, effectively solving problems of irregular spellings and cross-language interference.
[0028] Specifically, in step S4, in response to the syllable segmentation strategy being a syllable segmentation method based on a deep neural network model, based on the syllable segmentation strategy, the name string to be translated is syllable-segmented to obtain a syllable-segmented name string. In a specific example of this application, as Figure 3 shown, step S4 includes: S41, performing context semantic association encoding based on the place name Token unit on the name string to be translated to obtain a semantic context association encoding representation of the name to be translated; S42, performing local semantic relevance reconstruction and enhancement on the semantic context association encoding representation of the name to be translated to obtain a semantic context association enhanced encoding representation of the name to be translated; S43, based on the semantic context association enhanced encoding representation of the name to be translated, determining the syllable-segmented name string.
[0029] Specifically, in S41, context semantic association encoding based on place name Token units is performed on the name string to be translated to obtain a context semantic association encoding representation of the place name to be translated. That is, in the embodiments of the present application, first, word segmentation processing is performed on the name string to be translated to obtain a sequence distribution of place name Token units. It should be understood that when the input place name character stream contains a mixed writing system (such as the simultaneous presence of Latin letters, tone marks, or agglutinative morphemes), the continuity and unstructured characteristics of the original string will hinder the model's recognition of language attribution and internal rules. For example, in the face of a long string composed of agglutinative morphemes, which may contain multiple combinations of roots and suffixes nested inside, traditional space-based coarse-grained segmentation cannot capture implicit syllable boundaries (such as consonant clusters or vowel liaison rules). At this time, by using a multilingual sub-word segmentation algorithm to decompose the character sequence into Token units with independent semantic or phonological meanings (such as decomposing agglutinative morphemes into "root + suffix" combinations), the internal word-formation logic of the language can be explicitly revealed. For example, in compound words, the core semantic unit and grammatical markers can be separated, or in non-Latin scripts, the visual adhesion of conjunct characters can be correctly processed. By generating a Token sequence distribution, the model can transform heterogeneous input character streams into a set of units with discretized semantic boundaries. For example, in languages containing tone marks, tone markers and base letters are combined into independent Tokens to retain phonological features (such as binding the tone mark "́" to the vowel letter), or in spellings with dense consonant clusters, sub-units that conform to the phoneme rules of the target language are segmented (such as decomposing consonant clusters into legal syllable combinations). This discretization processing provides a parsable basic structure for subsequent semantic encoding. Especially when dealing with non-standard spellings (such as archaic variants in historical place names), the Token units generated by the word segmentation module through an adaptive strategy can effectively distinguish core morphemes and interference noises in language variants. For example, in a mixed language component, character blocks of different language families are identified and isolated, thus avoiding segmentation errors caused by cross-language rule conflicts.
[0030] Next, semantic embedding encoding is performed on each to-be-translated place name Token unit in the sequence distribution of to-be-translated place name Token units to obtain a sequence distribution of semantic embedding encoding vectors of to-be-translated place name Token units. That is, in the technical solution of the present application, each to-be-translated place name Token unit in the sequence distribution of to-be-translated place name Token units is respectively passed through a word embedding encoder based on Word2Vec to obtain a sequence distribution of semantic embedding encoding vectors of to-be-translated place name Token units. It should be understood that in traditional natural language processing, Token units, as discrete symbols (such as sub-words, roots, or character combinations), cannot directly express semantic relevance and internal language rules. For example, in agglutinative languages, a suffix Token representing the locative case and another suffix Token representing the possessive relationship, if only existing in the form of discrete symbols, it is difficult for the model to automatically infer the similarity of their grammatical functions; in cross-language scenarios, Token symbols of synonymous roots in different language families (such as "mountain" and "mountain") lack more computable relationship expressions. The word embedding encoder based on Word2Vec can map each discrete Token to a low-dimensional continuous vector space through unsupervised training, using the statistical law of Token co-occurrence within the context window, so that Tokens with similar semantics or functions have geometric proximity in the vector space (such as the locative case suffix vectors forming a clustering distribution in the vector space). This continuous vector representation provides a quantifiable semantic basis for subsequent models.
[0031] Furthermore, the sequence distribution of the semantic embedding encoding vectors of the to-be-translated place name Token units is subjected to place name token context semantic association encoding to obtain the to-be-translated place name semantic context association encoding vector, which is used as the to-be-translated place name semantic context association encoding representation. That is, in the technical solution of this application, the sequence distribution of the semantic embedding encoding vectors of the to-be-translated place name Token units is passed through a context semantic association encoder based on BiLSTM to obtain the to-be-translated place name semantic context association encoding vector. It should be understood that the semantic embedding encoding vector of the to-be-translated place name Token unit only expresses the static semantics of the Token (such as the independent meaning of the root of an agglutinative language), and cannot model the interaction rules between adjacent Tokens (such as the constraint effect of a suffix on the syllable structure of a stem). For example, the semantic function of a suffix Token vector representing the locative case needs to be combined with the vector of the previous stem Token to fully express the grammatical meaning of "at... position", while a unidirectional LSTM can only capture forward dependencies and is difficult to capture the reverse influence of the suffix on the previous stem (such as the change of the syllable stress position of the stem by the separable verb prefix in German). BiLSTM models the temporal relationship of the Token sequence in two time dimensions, forward and backward, through a bidirectional gating mechanism, so that the hidden state at each position simultaneously integrates the context information of the past and the future (such as the grammatical function of a prefix affecting subsequent syllable segmentation, and the morphological features of a suffix modifying the pronunciation rules of the previous stem), thereby constructing a global semantic association field. Specifically, the gating structure of BiLSTM (input gate, forget gate, output gate) controls the transmission and forgetting of information flow in a parameterized manner. For example, when processing a long Token sequence of an agglutinative language, the forget gate can dynamically determine the long-term memory to be retained (such as the core semantics of the root), and the input gate filters the local features of the current Token (such as the grammatical attributes of the suffix). The superposition of the bidirectional mechanism enables the model to gradually accumulate the modification effect of the prefix on the stem (such as "mega-" meaning "huge") when encoding the "prefix + root + suffix" structure, and the reverse LSTM captures the morphological constraint of the suffix on the root in the reverse direction (such as "-polis" meaning "city"), and finally forms a context representation containing bidirectional dependencies through the concatenation of the hidden states to obtain the to-be-translated place name semantic context association encoding vector.It is worth mentioning that the bidirectional context encoding breaks through the limitation of the local window, enabling the model to handle long-distance dependencies (such as consonant assimilation phenomena spanning multiple Tokens), for example, identifying the core syllable boundaries separated by multiple modifiers in compound words; in addition, the dynamic gating mechanism enhances the model's adaptability to irregular language phenomena, such as suppressing interference noise in spelling variants (such as redundant characters in historical place names) through the forget gate, while strengthening key morphological features (such as the indication of glottal stop symbols for syllable segmentation) using the input gate; finally, the bidirectionally fused hidden state provides an ambiguity-resistant semantic representation for subsequent modules, for example, distinguishing homonymous suffixes in agglutinative languages, and its context association vectors will activate specific features in different dimensions.
[0032] Specifically, in S42, local semantic relevance reconstruction and enhancement are performed on the semantic context-related encoded representation of the to-be-translated place name to obtain a semantic context-related enhanced encoded representation of the to-be-translated place name. It should be understood that after the to-be-translated place name undergoes word embedding and context semantic encoding, although the global semantic dependency relationship has been partially modeled, the long compound words formed by syllable superposition in agglutinative languages, the spelling patterns of consonant clusters and vowel alternation in inflectional languages, and the cross-linguistic interference components of mixed-language place names often show non-linear coupling of local semantic structures. Therefore, in order to achieve feature rectification and dynamic enhancement of the semantic context-related features of the to-be-translated place name, in the technical solution of this application, local semantic relevance reconstruction and enhancement are performed on the semantic context-related encoded vector of the to-be-translated place name to obtain a semantic context-related enhanced encoded vector of the to-be-translated place name. In this process, first, based on the feature phase reconstruction of one-dimensional convolutional encoding through a sliding window mechanism, multi-scale local structural patterns are extracted from the encoded vector sequence (for example, a convolutional kernel of length 3 captures the morphological interaction features of three adjacent Tokens), and the local phase information (that is, the relative relationship between adjacent dimensions) hidden in the continuous vector space is explicitly modeled; then, since the initial phase encoding set may contain redundant information (such as the smooth transition features in non-boundary regions) and noise interference (such as irregular spelling fluctuations caused by cross-language mixing), in the information squeezing stage, feature dimension screening is performed on the local phase encoded vector through a learnable attention mechanism or sparse constraints. For example, in regions with dense consonant clusters, the model will strengthen the dimension representing the steepness of consonant transitions while suppressing the dimension representing vowel stability; in agglutinative language segments with dense suffixes, the dimension representing grammatical functions is retained; then, by statistically analyzing the effective components of the features, the information density index (that is, the number of effective component statistics) of each local phase encoded vector is calculated, and then a dynamic gain operator is generated. For example, when it is detected that the statistics of a certain local phase encoding are significantly higher than the threshold, the gain operator will non-linearly amplify the features of its corresponding dimension (such as generating an enhancement coefficient between 0 and 1 through the Sigmoid function), so as to highlight the mutation signal at the syllable boundary (such as the break point of consonant clusters or the turning point of vowel weakening) in the feature reshaping stage; finally, local structure and global context are synergistically optimized through phase significance reshaping. For example, when processing mixed-language place names, the global context encoding may be difficult to accurately locate the syllable boundary due to cross-language rule conflicts, but the local phase enhancement module suppresses the interference signals of conflicting language families (such as the long consonant sequences in agglutinative languages) by amplifying the local features that conform to the phonological rules of the target language (such as the preference for open syllables in Romance languages), so that the enhanced semantic context-related enhanced encoded vector of the to-be-translated place name can still maintain stable discriminative power in cross-language interference scenarios. This adaptive enhancement mechanism essentially constructs a multi-level feature interaction system, providing an input representation with both global robustness and local sensitivity for the deep decoder.
[0033] Specifically, first, the semantic context correlation encoding vector of the to-be-translated place name is reconstructed for semantic units based on one-dimensional convolutional encoding to obtain a set of local phase encoding vectors of the semantic features of the to-be-translated place name. It should be understood that the semantic context correlation encoding vector of the to-be-translated place name output by the BiLSTM contains global semantic dependencies (such as the cross-position grammatical association between the root and the suffix in an agglutinative language), which may weaken the microscopic interaction rules between adjacent feature dimensions (such as the pronunciation transition features in the consonant cluster area or the acoustic continuity in the vowel weakening area). For example, in a semantic segment with dense consonant clusters, although the global encoding vector can express the overall meaning of the morpheme, it is difficult to capture the steep changes within the consonant cluster (such as the alternation pattern between stop consonants and fricatives), and this local mutation is precisely the key signal of the syllable boundary. Therefore, in the technical solution of this application, the semantic context correlation encoding vector of the to-be-translated place name is reconstructed for semantic units based on one-dimensional convolutional encoding to obtain a set of local phase encoding vectors of the semantic features of the to-be-translated place name. Here, by using multiple groups of convolutional kernels with different sizes (such as window lengths of 3 / 5 / 7), the model can capture local phase patterns from different receptive fields: short windows focus on the microscopic interaction of adjacent dimensions (such as the breaking features of double consonant clusters), and long windows model the gradual change rules across dimensions (such as the tone continuity in vowel harmony). For example, in the area with dense agglutinative language suffixes, the multi-scale convolutional kernels can simultaneously capture the consonant-vowel combination rules within the suffix (short window) and the morphological superposition trend of the suffix sequence (long window), forming a set of local phase encoding vectors of the semantic features of the to-be-translated place name covering various possible forms of the syllable boundary. This multi-angle feature extraction provides a structured intermediate representation for the subsequent module, enabling the model to distinguish real syllable boundaries (such as the break of consonant clusters) from pseudo-boundaries (such as stable consonant combinations within the root). In addition, the feature phase reconstruction significantly enhances the model's sensitivity to implicit local rules. For example, when dealing with mixed language interference, the global encoding may be confused about the syllable boundary due to cross-language rule conflicts, but the local phase encoding effectively suppresses the irrelevant features of the interfering language system (such as the long consonant sequence in the agglutinative language) by strengthening the phonological patterns unique to the target language (such as the open syllable preference in the Romance language family). This fine-grained structured decoding lays an interpretable physical foundation for subsequent feature rectification and reshaping. In a specific example of this application, the semantic context correlation encoding vector of the to-be-translated place name is reconstructed for semantic units based on one-dimensional convolutional encoding with the following semantic unit reconstruction formula to obtain a set of local phase encoding vectors of the semantic features of the to-be-translated place name; where the semantic unit reconstruction formula is:
[0034] ;
[0035] Where is the semantic context correlation encoding vector of the to-be-translated place name, is the one-dimensional convolutional encoding process, is the characteristic phase reconstruction step size, are the 1st, 2nd, th, and th local phase-encoded vectors of the semantic features of the place names to be translated in the set of local phase-encoded vectors of the semantic features of the place names to be translated, respectively.
[0036] Next, semantic purification is performed on each local phase-encoded vector of the semantic features of the place names to be translated in the set of local phase-encoded vectors of the semantic features of the place names to be translated to obtain a set of local phase-encoded vectors for extracting the semantic features of the place names to be translated. It should be understood that although the set of local phase-encoded vectors of the semantic features of the place names to be translated generated by one-dimensional convolutional encoding contains multi-scale structural information (such as the steep change pattern of consonant clusters or the gradual change trend of vowel harmony), there may be dimensional redundancy (such as similar edge features extracted by adjacent convolutional kernels) and information mixing (such as the superposition of non-target phonological features caused by cross-lingual interference) inside it. For example, in the local phase encoding with dense consonant clusters, multiple convolutional kernels may simultaneously activate the response to the consonant transition features, resulting in collinearity between feature dimensions; in a mixed language scenario, some convolutional kernels may capture the phonological rules of non-target languages (such as the interference of long consonant sequences in agglutinative languages on syllable division in Romance languages), forming semantically irrelevant noise dimensions. Semantic purification projects the high-dimensional space of the local phase-encoded vectors of the semantic features of the place names to be translated through a learnable attention mechanism or sparse constraint, selects the key dimensions strongly related to syllable boundary determination (such as the mutation signal of consonant breakpoints or the turning feature of vowel weakening), and at the same time suppresses the low-information dimensions (such as the gentle fluctuations in the stable vowel area) and cross-lingual interference noise. During this process, by introducing task-driven feature importance evaluation, the model can quantify the contribution degree of each local phase-encoded vector of the semantic features of the place names to be translated to the final syllable segmentation target, so as to enhance the generalization ability by strengthening the high-value features of cross-lingual commonalities (such as the statistical law of consonant cluster breaks) and weakening language-specific noise. It is worth mentioning that semantic purification is not simply dimension reduction, but reconstructs the feature space through non-linear transformation. For example, a gating mechanism is used to dynamically adjust the activation intensity of each dimension, so that the set of local phase-encoded vectors for extracting the semantic features of the place names to be translated after extraction not only retains the diversity of multi-scale structural information, but also has discriminability suitable for the task. This purification and reconstruction of the feature space essentially constructs a noise-resistant and highly discriminative local feature base for syllable boundary detection, providing an underlying guarantee for the reliability of the end-to-end translation system. In a specific example of this application, the following semantic purification formula is used to perform semantic purification on each local phase-encoded vector of the semantic features of the place names to be translated in the set of local phase-encoded vectors of the semantic features of the place names to be translated to obtain a set of local phase-encoded vectors for extracting the semantic features of the place names to be translated; where the semantic purification formula is:
[0037] ;
[0038] Among them, is the first norm of the vector, is the th local phase encoding vector for semantic feature extraction of the place name to be translated in the set of local phase encoding vectors for semantic feature extraction of the place name to be translated.
[0039] Subsequently, calculate the number of effective components of the semantic features of the place name to be translated for each local phase encoding vector for semantic feature extraction of the place name to be translated in the set of local phase encoding vectors for semantic feature extraction of the place name to be translated. It should be understood that although the set of local phase encoding vectors after semantic purification has initially filtered redundant dimensions (such as low-variance fluctuation features in consonant cluster regions) and noise interference (such as non-target phonological patterns caused by cross-language mixing), there is still heterogeneity in the contribution degrees of feature dimensions within it - some dimensions may carry strong indication signals of syllable boundaries (such as steep gradient changes at consonant break points), while other dimensions may only carry weakly correlated or poorly generalized local information (such as rare spelling variants unique to a specific language). For example, when dealing with agglutinative morphemes, the dimension representing the suffix morphological rule may have global universality for syllable segmentation, while the dimension representing the consonant cluster inside the root may only be effective in a specific language. In the technical solution of this application, calculate the number of effective components of the semantic features of the place name to be translated for each local phase encoding vector for semantic feature extraction of the place name to be translated in the set of local phase encoding vectors for semantic feature extraction of the place name to be translated. Among them, the number of effective components is obtained through quantitative analysis (such as calculating the variance, entropy of the dimension activation value, or the mutual information with the task label) to identify highly discriminative feature dimensions (such as cross-language stable consonant transition patterns) and low-value dimensions (such as language-specific decorative character combinations), thereby providing an interpretable regulation basis for the subsequent gain operator. Through the calculation of the number of effective components of the semantic features of the place name to be translated, the model can convert the abstract local phase encoding vector for semantic feature extraction of the place name to be translated into an operable numerical index. For example, in the consonant assimilation region, the dimension with a higher number may correspond to the physical feature of the change in the consonant articulation position (such as the transition slope from alveolar to velar), while the dimension with a lower number may reflect irrelevant environmental noise (such as redundant symbols in spelling variants). This quantization control mechanism based on the number provides a task-oriented feature importance ranking for the local phase encoding vector for semantic feature extraction of the place name to be translated, enabling subsequent saliency reshaping to achieve adaptive enhancement in a data-driven manner and providing a highly robust intermediate representation for the end-to-end translation system. In a specific example of this application, calculate the number of effective components of the semantic features of the place name to be translated for each local phase encoding vector for semantic feature extraction of the place name to be translated in the set of local phase encoding vectors for semantic feature extraction of the place name to be translated with the following statistical formula; among them, the statistical formula is:
[0040] ;
[0041] Among them, is the th eigenvalue at the th position in the local phase encoding vector for extracting semantic features of the place names to be translated, represents the active ingredient count, is the trainable preset threshold, is the corresponding count of active ingredients in the semantic features of the place names to be translated.
[0042] Then, based on the statistical number of effective components of the semantic features of the to-be-translated place names that extract the local phase encoding vectors from the semantic features of each to-be-translated place name, calculate the phase reshaping gain operator of the semantic features of the local phase encoding vectors of the to-be-translated place names. It should be understood that after word embedding encoding and context semantic modeling, there are still noise interferences in the local semantic representation of the to-be-translated place names (such as consonant clusters with irregular spellings, morphological conflicts of cross-lingual mixed components), and these interferences will obscure the deep feature patterns that truly determine the syllable boundaries. Through the calculation of the statistical number of effective components of the features, the system can quantify the density of effective information strongly related to the syllable segmentation task in each local phase encoding vector. Further, in order to construct a dynamic and adaptive feature enhancement mechanism, in the technical solution of this application, calculate the initial phase reshaping gain operator of the semantic features of the local phase encoding vectors of the to-be-translated place names. Different from the traditional feature enhancement method with fixed weights, the phase reshaping gain operator integrates linguistic rules (such as the collocation rules of vowels and consonants) with data-driven features to form an interpretable feature regulation tool. Specifically, this operator will dynamically adjust the activation thresholds of different local phase vectors according to the statistical number of effective components of the features: for the encoding vectors containing high-frequency effective features (such as the syllable boundary signals of iconic suffixes in agglutinative languages), the gain operator will amplify the dimensions related to syllable segmentation in its representation space; while for the low-efficiency feature regions severely interfered by cross-languages (such as the spelling conflict areas in mixed place names), the propagation of noise signals will be suppressed through non-linear scaling. In this way, the model is enabled to distinguish between "effective syllable boundary features" and "cross-lingual interference noise". For example, when processing mixed place names containing German compound words and Finnish suffixes, it can effectively strengthen the temporal correlation features of consonant conversion nodes while weakening the morphological distortion caused by the differences in writing systems. In addition, this adaptive feature reshaping significantly improves the robustness of syllable segmentation, enabling the translation system to still perform phase alignment on the initial consonant and vowel combination patterns through the gain operator and accurately identify potential segmentation points affected by vowel length and consonant clusters when facing languages with blurred syllable boundaries such as Tibetan. In a specific example of this application, the phase reshaping gain operator of the semantic features of the local phase encoding vectors of the to-be-translated place names can be calculated through the following steps: based on the statistical number of effective components of the semantic features of the to-be-translated place names that extract the local phase encoding vectors, determine the suppression factors corresponding to the local phase encoding vectors of the semantic features of the to-be-translated place names; based on the suppression factors corresponding to the local phase encoding vectors of the semantic features of the to-be-translated place names, calculate the initial phase reshaping gain operator of the semantic features of the local phase encoding vectors of the to-be-translated place names.In a specific example of the present application, the semantic feature phase reshaping gain operator of the semantic features of each place name to be translated can be calculated through the following steps: Based on the effective component statistics of the semantic features of each place name to be translated for extracting the local phase encoding vector, determine the suppression factor corresponding to the local phase encoding vector of the semantic features of each place name to be translated; Based on the suppression factor corresponding to the local phase encoding vector of the semantic features of each place name to be translated, calculate the initial semantic feature phase reshaping gain operator of the semantic features of each place name to be translated; Perform feature phase dispersion missing correction on the initial semantic feature phase reshaping gain operator to obtain the semantic feature phase reshaping gain operator of the place name to be translated.
[0043] Specifically, when the initial semantic feature phase reshaping gain operator of the place name to be translated non-linearly amplifies the local phase encoding based on the statistics (such as enhancing the dimension representing the steep gradient of the consonant breakpoint by 1.5 times), over-focusing on the significance of local features may disrupt the overall distribution structure of the feature space. For example, in the processing of long suffix sequences of agglutinative morphemes, if the gain operators of multiple suffix Tokens independently enhance their grammatical function dimensions, it may cause the distribution of feature vectors in the space to show non-uniform scattering (such as excessive separation of some vector clusters), thereby weakening the global context relevance. This phenomenon is mathematically manifested as the lack of phase direction symmetry, that is, the geometric direction distribution of the enhanced feature vectors in the space deviates from the inherent consistency of linguistic rules. Therefore, in the preferred example of the present application, feature phase dispersion missing correction is performed on the initial semantic feature phase reshaping gain operator of the place name to be translated to obtain the semantic feature phase reshaping gain operator of the place name to be translated. That is, by introducing a flatness decomposition metric, the action of the gain operator is constrained within the scope of maintaining the holomorphic flatness of the space. For example, when processing consonant clusters in mixed-language place names, the corrected gain operator will ensure the direction compatibility between the consonant cluster features of the Romance language family and the long consonant sequences of agglutinative morphemes in the vector space, avoiding spatial distortion caused by conflicts in language family rules. This mathematical mechanism transforms the initial semantic feature phase reshaping gain operator with single-mode coupling into a canonical operation that conforms to spatial geometric constraints through gauge field construction, enabling the feature enhancement process to maintain the overall stability of the feature distribution while enhancing local significance. The corrected initial semantic feature phase reshaping gain operator enables the model to stably generalize the phonological rules (such as the separation pattern of Tibetan conjunct characters) learned from scarce labeled data to unseen language variants by maintaining the translational invariance of the feature vectors. This mathematical correction based on the holomorphic structure essentially constructs a deep mapping between linguistic rules and the geometry of the feature space, providing a feature enhancement paradigm with both local sensitivity and global consistency for the end-to-end translation system.
[0044] In this example, the statistical count of the effective components of the semantic features of the to-be-translated place names for extracting the local phase encoding vectors based on the semantic features of each to-be-translated place name is used to calculate the phase reshaping gain operator of the semantic features of the local phase encoding vectors of each to-be-translated place name by the following calculation formula; wherein, the calculation formula is:
[0045] ;
[0046] ;
[0047] ;
[0048] ;
[0049] Wherein, is the number of vectors in the set of local phase encoding vectors of the semantic features of the to-be-translated place names, is the corresponding polar angle, is the corresponding suppression factor, represents pi, represents the arctangent function, is the phase reshaping gain operator of the semantic features of the to-be-translated place names, is the corresponding initial phase reshaping gain operator of the semantic features of the to-be-translated place names, is the flatness factor of the semantic space of the to-be-translated place names, is the metric representation factor of the semantic flatness decomposition of the to-be-translated place names, is the phase reshaping gain operator of the semantic features of the to-be-translated place names.
[0050] Furthermore, a semantic feature phase reshaping gain operator for the to-be-translated place name semantic feature local phase encoding vectors is used to perform feature phase saliency reshaping on the set of to-be-translated place name semantic feature local phase encoding vectors to obtain to-be-translated place name semantic context correlation enhanced encoding vectors. Here, after the semantic feature phase reshaping gain operator for the to-be-translated place name semantic feature local phase encoding vectors is dynamically regulated by statistical number drive (for example, a 1.5-fold gain coefficient is given to the consonant break point feature), its essence is to transform the phonological rules of linguistics (such as the morphological constraints of agglutinative language suffixes) into geometric operations in the feature space. For example, when processing compound words, the local phase encoding vector corresponding to the consonant assimilation phenomenon at the junction of the root and the suffix, after the non-linear scaling of the gain operator, will form a direction-specific enhancement in the vector space (such as a specific dimension of the feature vector being stretched to amplify the steepness of the consonant transition), thereby explicitly marking the potential position of the syllable boundary. This feature enhancement mechanism based on geometric space reconstruction essentially establishes an interpretable mapping between linguistic rules and machine-understandable feature distributions, providing core technical support for the end-to-end translation system with both domain knowledge embedding and data-driven adaptability. In a specific example of this application, the following feature phase reshaping formula is used to perform feature phase saliency reshaping on the set of to-be-translated place name semantic feature local phase encoding vectors to obtain to-be-translated place name semantic context correlation enhanced encoding vectors; where, the feature phase reshaping formula is:
[0051] ;
[0052] ;
[0053] where, is the natural exponential function value with as the base, is the semantic feature phase reshaping gain weight for the to-be-translated place name, is the to-be-translated place name semantic context correlation enhanced encoding vector.
[0054] Specifically, in step S43, the syllable-segmented string of geographical names is determined based on the enhanced encoding representation of the semantic context of the geographical names to be translated. That is, in the technical solution of the present application, the enhanced encoding vector of the semantic context of the geographical names to be translated is input into the syllable segmentation decoder based on the deep neural network model to obtain the syllable-segmented string of geographical names. It should be understood that the enhanced encoding vector of the semantic context of the geographical names to be translated contains dual information of global context association and local morphological features. Essentially, it constructs a cross-lingual abstract phonological space, which not only contains the historical trajectory of root evolution but also includes the dynamic distribution pattern of syllable boundaries. Traditional decoders based on finite state machines are difficult to capture the complex syllable superposition logic inside; and due to the cross-linguistic component interaction in mixed-language geographical names, the decoder has the flexible ability to dynamically adjust the phonological parsing strategy. The deep neural network decoder can establish an inverse mapping from the abstract feature space to the specific syllable sequence. Specifically, it first analyzes the burst energy distribution of consonant clusters through multi-layer non-linear transformation (such as the plosive sound combination starting with "pf-" in German), and then dynamically focuses on the feature mutation region at the splicing point of agglutinative morphemes in combination with the attention mechanism (such as the connection boundary between the suffix and the stem in Turkish). In this way, the system can achieve a dynamic balance between phonological rule reasoning and morphological feature parsing, and break through the syllable segmentation accuracy across language barriers. This decoding paradigm based on the holomorphic structure constraint essentially constructs a mathematical bridge between linguistic prior knowledge and data-driven models, providing an end-to-end solution with both rule rigor and adaptability for geographical name translation in a globalized scenario. "suffix and the stem).
[0055] Specifically, in step S5, the syllable-segmented place name string is input into a place name translation model to obtain a place name translation text in the target language. Among them, the place name translation model is a deep learning model, which can understand and transform the complex relationships between different languages. Specifically, the model internally contains multiple layers of neural network layers, and each layer is responsible for capturing different aspects of the input data features, from the basic language symbol representation to the more advanced semantic understanding and cross-language mapping. In this process, first, the model will perform a preliminary analysis on the input place name string, identifying the meaning of each character or Token (sub-word unit) and its role in the sentence; then, using deep neural network architectures such as bidirectional long short-term memory networks (BiLSTMs), Transformers, etc., the model starts a comprehensive analysis of the input string; it not only needs to understand the meaning of each individual word, but also grasp the overall meaning of the entire place name string and the mutual relationship between its parts. For example, when processing compound words or place names containing agglutinative morphemes, the model needs to identify the implicit relationship between the root and the suffix and adjust the translation strategy accordingly; subsequently, the place name translation model will reorganize and optimize the parsed information according to the characteristics of the target language. This includes but is not limited to adjusting the grammatical structure, selecting appropriate synonyms or near-synonyms, considering cultural background differences and other factors to generate a translation result that is both faithful to the original text and conforms to the target language habits. It should be noted that this stage may also involve the processing of some special morphological features, such as prefixes / suffixes, consonant clusters, etc., to ensure that the finally output place name translation text is both natural and accurate. In this way, the system realizes an end-to-end effective conversion from the original place name to the target translation, greatly improving the efficiency and accuracy of cross-language information processing, especially showing important value in fields such as geographic information systems, multilingual map navigation, and cross-border logistics.
[0056] It is worth mentioning that those of ordinary skill in the art should be aware that, in addition to the above-disclosed technical solution of "inputting the syllable-segmented place name string into a place name translation model to obtain a place name translation text in the target language", other existing technologies can also be used to implement this technical process. For example, in another specific implementation, the place name translation solution in the paper "Neural Machine Translation for Low-Resource Named Entities" can be used to translate the syllable-segmented place name string to obtain a place name translation text in the target language. It should be understood that "Neural Machine Translation for Low-Resource Named Entities" is an existing technology, and to avoid repetition, no content expansion will be made here.
[0057] In summary, a high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to an embodiment of the present application is elucidated. First, the input place name is split into a sub-word Token sequence, and the context encoding model of the word embedding model is used to implement the semantic context association encoding of the place name to be translated. Subsequently, the internal structured information and dependency relationships in the Token context semantics of the place name to be translated are modeled by means of strengthening local semantic relevance reconstruction, the implicit syllable associations in compound words and agglutinative languages are identified, and the local morphological features (such as prefixes / suffixes, consonant clusters) near the syllable boundaries are enhanced to solve the problems of irregular spelling and cross-language interference. Furthermore, the syllable-segmented place name string is decoded and output for place name translation, and the corresponding place name target language translation text is generated to achieve an end-to-end effective conversion from the original place name to the target translation name.
[0058] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation, characterized in that: include: Receive a place name character string to be translated input by a user; Using the language identification model to detect the source language of the place name string to be translated; Determine the syllable segmentation strategy based on the source language of the place name string to be translated; In response to the syllable segmentation strategy being a syllable segmentation method based on a deep neural network model, based on the syllable segmentation strategy, the place name character string to be translated is syllable segmented to obtain the place name character string after syllable segmentation, including: performing context semantic association encoding based on the place name Token unit on the place name character string to be translated to obtain the semantic context association encoding representation of the place name to be translated; performing local semantic association reconstruction and enhancement on the semantic context association encoding representation of the place name to be translated to obtain the semantic context association enhanced encoding representation of the place name to be translated; based on the semantic context association enhanced encoding representation of the place name to be translated, determining the place name character string after syllable segmentation; The syllable-segmented place name character string is input into the place name translation model to obtain the place name target language translation text.
2. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 1 is characterized in that: Based on the source language of the place name string to be translated, determine the syllable segmentation strategy, including: If the source language of the place name string to be translated is a high-resource language, choose the rule-based syllable segmentation method; If the source language of the place name string to be translated is a medium- or low-resource language, select the syllable segmentation method based on the statistical model; If the source language of the place name string to be translated is a mixed language, select the syllable segmentation method based on the deep neural network model.
3. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 2 is characterized in that: High-resource languages include English, French, and Spanish; medium- and low-resource languages are Chinese, Japanese, Korean, and Vietnamese.
4. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 2 is characterized in that: The place name string to be translated is encoded based on the contextual semantic association of the place name Token unit to obtain the semantic contextual association encoding representation of the place name to be translated, including: Perform word segmentation on the place name string to be translated to obtain the sequence distribution of the place name Token unit to be translated; Perform semantic embedding coding on each place name Token unit to be translated in the sequence distribution of the place name Token unit to be translated to obtain a sequence distribution of the semantic embedding coding vector of the place name Token unit to be translated; The sequence distribution of the semantic embedding coding vector of the place name Token unit to be translated is encoded with the place name token context semantic association to obtain the place name semantic context association coding vector to be translated, and used as the place name semantic context association coding representation to be translated.
5. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 4 is characterized in that: Semantic embedding coding is performed on each place name Token unit to be translated in the sequence distribution of the place name Token unit to be translated to obtain a sequence distribution of the semantic embedding coding vector of the place name Token unit to be translated, including: Each place name Token unit to be translated in the sequence distribution of the place name Token unit to be translated is passed through a word embedding encoder based on Word2Vec to obtain a sequence distribution of the semantic embedding coding vector of the place name Token unit to be translated.
6. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 5 is characterized in that: Performing place name token context semantic association coding on the sequence distribution of the semantic embedding coding vector of the place name to be translated to obtain the place name semantic context association coding vector to be translated, and using it as the place name semantic context association coding representation to be translated, including: The sequence distribution of the semantic embedding coding vector of the place name Token unit to be translated is passed through a BiLSTM-based contextual semantic association encoder to obtain the semantic contextual association coding vector of the place name to be translated.
7. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 6 is characterized in that: The semantic context-related coding representation of the place name to be translated is reconstructed and enhanced in terms of local semantic relevance to obtain the semantic context-related enhanced coding representation of the place name to be translated, including: The semantic unit reconstruction and semantic purification processing are performed on the semantic context associated coding vector of the place name to be translated to obtain a set of local phase coding vectors for extracting the semantic features of the place name to be translated; Calculate the statistics of effective components of semantic features of place names to be translated of each semantic feature extraction local phase encoding vector of place names to be translated in the set of semantic feature extraction local phase encoding vectors of place names to be translated; Based on the semantic features of each place name to be translated, the statistics of the effective components of the semantic features of the place name to be translated of the local phase encoding vector are extracted, and the phase reshaping gain operator of the semantic features of the place name to be translated of the local phase encoding vector of the semantic features of each place name to be translated is calculated; Based on the phase reshaping gain operator of the semantic features of each place name to be translated, the set of local phase coding vectors of the semantic features of the place names to be translated is subjected to feature phase saliency reshaping to obtain the semantic context-related enhanced coding vector of the place names to be translated, which is used as the semantic context-related enhanced coding representation of the place names to be translated.
8. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 7 is characterized in that: The semantic unit reconstruction and semantic purification processing are performed on the semantic context associated coding vector of the place name to be translated to obtain a set of local phase coding vectors for extracting the semantic features of the place name to be translated, including: The semantic context associated coding vector of the place name to be translated is reconstructed into a semantic unit based on one-dimensional convolutional coding to obtain a set of local phase coding vectors of the semantic features of the place name to be translated; Semantic purification is performed on each local phase coding vector of semantic features of place names to be translated in the set of local phase coding vectors of semantic features of place names to be translated to obtain a set of local phase coding vectors of semantic features of place names to be translated.
9. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 8 is characterized in that: Based on the semantic features of each place name to be translated, the effective component statistics of the semantic features of the place name to be translated of the local phase encoding vector are extracted, and the phase reshaping gain operator of the semantic features of the place name to be translated of the local phase encoding vector of the semantic features of each place name to be translated is calculated, including: Based on the statistics of effective components of semantic features of place names to be translated extracted from local phase coding vectors of semantic features of place names to be translated, the inhibition factors corresponding to the local phase coding vectors extracted from the semantic features of place names to be translated are determined; Based on the semantic features of each place name to be translated, the suppression factor corresponding to the local phase encoding vector is extracted, and the initial semantic feature phase reshaping gain operator of the semantic feature of each place name to be translated is calculated; The feature phase dispersion missing correction is performed on the initial semantic feature phase reshaping gain operator of the place name to be translated to obtain the semantic feature phase reshaping gain operator of the place name to be translated.
10. The high-precision place name translation method integrating artificial intelligence and multilingual syllable segmentation according to claim 9 is characterized in that: Based on the semantic context-related enhanced coding representation of the place name to be translated, the place name string after syllable segmentation is determined, including: The semantic context-related enhanced coding vector of the place name to be translated is input into the syllable segmentation decoder based on the deep neural network model to obtain a syllable-segmented place name character string.
Citation Information
Patent Citations
Multilingual place name and root Chinese translation method based on Transformer deep learning model
CN112084796A
Translation model training method and text translation method and device
CN119089914A