Cross-language double-space alignment and pedigree guiding translation system

By building a translation system with cross-language dual spatial alignment and pedigree guidance, the shortcomings of the existing multilingual translation model in semantic alignment and structural restoration are solved, high-precision semantic mapping and structural alignment are achieved, and the adaptability and translation quality of the translation system in complex language environments are improved.

CN120493950APending Publication Date: 2025-08-15INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510496281.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When existing multilingual translation models deal with different semantic structures, language historical relationships and complex language structures, it is difficult to achieve accurate semantic alignment and structural restoration, resulting in insufficient translation quality and accuracy, especially in language pairs with significant differences in low-resource language and word order.

Method used

A translation system with cross-language dual-space alignment and pedigree guidance is adopted. By constructing a double-layer modeling of basic semantic space and language-specific semantic space, combining the language pedigree map and structural perception attention mechanism, high-precision semantic mapping and structural alignment of the source language and the target language are achieved, and translated texts that conform to the target language structure and cultural habits are generated.

Benefits of technology

It significantly improves the adaptability and accuracy of multilingual translation systems in complex language environments, and is especially suitable for international office and cross-cultural collaboration scenarios, solving the problems of inaccurate semantic alignment, lack of structural guidance of language transfer mechanisms, and insufficient translation structure restoration capabilities, and improving the translation quality and cross-cultural communication experience of low-resource languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493950A_ABST
    Figure CN120493950A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-language double-space alignment and pedigree guidance translation system, and relates to the technical field of natural language processing, and the system comprises a semantic double-space modeling module which is used for constructing a basic semantic space and a language specific semantic space based on a source language text, and forming a unified double-space semantic representation; the semantic mapping and alignment module is used for performing tensor projection and alignment of a source language and a target language based on the double-space semantic representation to generate a semantic representation adaptive to the target language; the translated text generation module is used for receiving the semantic representation and generating a translated text by means of a structure-perceived target language decoder and a language style control mechanism in combination with target language structure features and cultural habits; the structure perception attention module is used for introducing language structure features and improving perception and processing capabilities for different language structures; and the language pedigree guiding module is used for constructing a language pedigree graph and improving the translation effect of the low-resource language. The method is suitable for a multi-language communication scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a cross-language dual-space alignment and lineage-guided translation system. Background Art

[0002] With the deepening integration of global economic integration and cultural exchange, the demand for cross-language communication has exploded. From the daily work of multinational corporations and contract negotiations in international trade to international academic collaboration and the transnational dissemination of cultural industries, language barriers have become a key constraint to efficient communication. Machine translation systems based on large-scale pre-trained language models, with their automation and efficiency, have become a core technology for breaking down language barriers. They provide crucial technical support for scenarios such as document processing in multilingual offices, contract translation in international business, and information transfer in cross-cultural collaboration. Current mainstream multilingual translation models, such as mBERT (Multilingual BERT), XLM-R (Cross-Lingual Language Model RoBERTa), and mT5 (Multilingual T5), generally adopt a shared semantic representation space approach. By jointly training on massive corpora spanning multiple languages, they attempt to construct a common semantic space, enabling semantic alignment of texts in different languages within this space and enabling translation from the source language to the target language. This approach, to a certain extent, makes multilingual translation feasible and convenient.

[0003] However, these existing methods have exposed many technical bottlenecks in practical applications. From the perspective of semantic differences, there is a huge gap between different languages in the world in terms of semantic structure, grammatical rules and cultural expression. For example, English tends to express logical relationships through clear grammatical structures, while Chinese often relies on context and context to convey semantics; the direct expression habits in Western culture are completely different from the implicit and euphemistic expressions in Eastern culture. A single shared semantic space is difficult to fully and accurately capture these deep semantic differences, resulting in the inability to accurately convey some semantic details during the translation process, resulting in translation ambiguity or semantic loss, affecting the translation quality and the accuracy of information transmission.

[0004] Existing models generally have defects in terms of the historical evolution of languages. Languages, like organisms, have unique historical evolutionary contexts and kinship relationships. Languages from different language families have specific similarities and differences in vocabulary, grammar, and semantics. However, when training, existing models treat languages from different language families equally, without fully considering the kinship and migration paths of languages. This means that when dealing with low-resource languages (such as some minority languages or languages with a small number of speakers), or language pairs with significant differences in word order (such as the subject-verb-object postposition structure in Japanese and the subject-verb-object preposition structure in English), the model cannot use the inherent connections between languages to improve translation capabilities, resulting in a significant decline in translation performance and difficulty meeting actual application needs.

[0005] Current models also face challenges when processing complex language structures. For example, German, Chinese, and Japanese, for example, present significant challenges for machine translation: German's long sentence structure and complex case system, Chinese's highly concise semantic expression and flexible word order, and Japanese's honorific system and rich use of particles. The attention mechanisms in existing models lack the ability to perceive these complex language structures. During the translation process, they are unable to accurately identify the structural features of the source language and appropriately convert them into the target language, making it difficult to achieve precise structural alignment and information retention. Consequently, the translation results lack grammatical structure and semantic expression that conform to the target language's conventions, reducing readability and usability.

[0006] In summary, in the real-world context of multilingualism, coexistence of multiple language families, and complex language structures, existing machine translation technology struggles to simultaneously achieve a balance between accuracy, semantic consistency, and structural restoration. Whether it's accurate semantic communication, the rational utilization of historical language relationships, or the effective handling of complex structures, these technologies face significant shortcomings. There is an urgent need to develop a translation system with greater language adaptability and structural alignment capabilities to overcome existing technological bottlenecks and meet the growing demand for cross-language communication. Summary of the Invention

[0007] The present invention addresses technical issues existing in existing multilingual translation models, such as inaccurate semantic alignment, lack of structural guidance in language transfer mechanisms, and insufficient translation structure restoration capabilities, and provides a translation system with cross-language dual-space alignment and lineage guidance.

[0008] The present invention provides a cross-language dual-space alignment and lineage-guided translation system that solves the above-mentioned technical problems using the following technical solutions:

[0009] A cross-language bi-space alignment and lineage-guided translation system, comprising:

[0010] The semantic dual-space modeling module is used to feed the source language text into a basic semantic space encoder and a language-specific semantic modeler for parallel processing. It constructs a basic semantic space to represent cross-language common semantic information and a language-specific semantic space to represent semantic features related to the source language structure and expression habits. Finally, a weighted fusion strategy is used to form a unified dual-space semantic representation, providing an accurate and interpretable foundation for semantic mapping and alignment.

[0011] The semantic mapping and alignment module is used to implement tensor projection and alignment between the source and target languages based on the dual-space semantic representation provided by the semantic dual-space modeling module, using a semantic projection matrix, loss function training, and a context-fusion attention mechanism. This generates a semantic representation adapted to the target language and provides input for the translation generation module.

[0012] The translation generation module receives the target language semantic representation generated by the semantic mapping and alignment module, generates the translation text based on the target language structural features and cultural habits with the help of the structure-aware target language decoder and language style control mechanism, and outputs the translation text and intermediate results to support user review and system feedback optimization.

[0013] The structure-aware attention module is used to introduce language structure features into the encoding and decoding processes of the translation system. It assists the semantic dual-space modeling module in building semantic tensors, helps the semantic mapping and alignment module achieve precise alignment, and supports the translation generation module in generating text that meets structural requirements, thereby improving the translation system's ability to perceive and process different language structures.

[0014] The language genealogy guidance module is used to construct a language genealogy graph, encode the kinship, language family classification and evolutionary path between languages as structured prior knowledge, and dynamically adjust the training sample ratio and gradient propagation weight according to the graph during the training phase. In the inference phase, a bridging language mechanism is introduced for cross-language translation. At the same time, auxiliary information related to the language genealogy is provided to the semantic dual-space modeling module, the semantic mapping and alignment module and the translation generation module, thereby optimizing the translation effect of low-resource and word order-different languages.

[0015] Optionally, the basic semantic space encoder uses the Transformer structure as the underlying network framework, introduces a cross-language shared multilingual token embedding based on the SentencePiece structure, and constructs a unified basic semantic representation tensor by learning common semantic units across languages.

[0016] Among them, the semantic units common across languages include basic actions, common entities and general relations.

[0017] Optionally, the language-specific semantic modeler involved models the specific expression habits of the source language, using the language type label, grammatical structure, and pragmatic preference information of the source language to load the structural modeling sub-network of the corresponding language to accurately capture the structural characteristics of the source language. At the same time, the constructed semantic tensor is strictly aligned with the basic semantic tensor in terms of dimensional design, ensuring the accuracy and uniqueness of the semantic modeling, laying a solid foundation for subsequent fusion and projection operations;

[0018] Among them, the structural modeling sub-network is a grammatical structure graph network based on Adapter, Prefix Tuning or lightweight customization, which is used to accurately capture the structural characteristics of the source language.

[0019] Further optionally, the semantic mapping and alignment module involved adopts a dual-space mapping mechanism based on a semantic projection matrix, specifically including:

[0020] The loss function unit is used to introduce the semantic preservation loss function and the structural equivalence loss function during the training phase. The semantic preservation loss function ensures that key information is not lost during the semantic mapping process between the source language and the target language, maintaining semantic accuracy. The structural equivalence loss function ensures that the structural features of the source and target languages are preserved. This allows the translation system to take into account both semantics and structure when learning the mapping relationship, laying the foundation for accurate semantic mapping.

[0021] The context-fusion attention unit is used to introduce a context-fusion attention mechanism when performing semantic mapping and alignment tasks. It dynamically aggregates the context window information of the source language text, breaking through the limitation of single sentences and taking into account the context at the paragraph or dialogue level.

[0022] The decoupling projection unit is used to decouple the basic semantics and the language-specific semantics and then project them separately, so that the semantic tensor of the target language can integrate the semantic framework and language features of the source language in a unified space, generate a semantic representation suitable for the target language, and complete the semantic mapping and alignment tasks between the source language and the target language.

[0023] Optionally, the translation generation module may include:

[0024] The parsing unit is used to call the target language decoder with structure-aware function during the translation generation stage to parse the semantic tensor processed by the semantic mapping and alignment module;

[0025] The text generation unit is embedded in the target language decoder and is used to generate a translation text that conforms to the grammatical and structural requirements of the target language and incorporates the cultural conventions of the target language based on the structural features of the target language.

[0026] Language style unit, used to dynamically adjust the style of the translated text according to different usage scenarios and target audiences;

[0027] The result output unit is used to output the original text and the translated text together, along with intermediate results including semantic tensors, language tags, and structural alignment information for user review, helping users understand the translation process and basis;

[0028] The feedback optimization unit is used to transmit the translated text manually modified by the user as a training sample back to the training end, and realize system optimization by learning the training samples.

[0029] Furthermore, in the encoding process of the translation system, the structure-aware attention mechanism module deeply analyzes the source language text and integrates various structural features such as grammatical structure, word order, and sentence component relationships into the encoding network in the form of structured data. By analyzing and encoding the structural features of the source language, the module accurately understands the intrinsic connection between the semantics and structure of the source language. Therefore, when constructing the semantic tensor, the module can not only retain the semantic information but also record the structural characteristics of the source language, providing more accurate and comprehensive input for the semantic dual-space modeling module, helping it to construct a basic semantic space and a language-specific semantic space that are more in line with the actual situation of the source language.

[0030] During the decoding process of the translation system, the structure-aware attention mechanism module accurately identifies the grammatical rules, common sentence structures, and various structural features of language habits of the target language, and introduces these features into the decoding network. By guiding the decoding network to refer to the structural features of the target language, the auxiliary translation system follows the grammatical norms and expression habits of the target language when generating translated texts, and generates sentence structures that meet the structural requirements of the target language.

[0031] Optionally, after the semantic mapping and alignment module completes the semantic processing, the structure-aware attention mechanism module further ensures that the structure of the translation is highly consistent with the target language, achieving dual-precision conversion of semantics and structure, and providing a guarantee for the translation generation module to output high-quality translation text that conforms to the target language specifications.

[0032] Further optionally, the language lineage guidance module involved specifically includes:

[0033] A graph construction unit is used to obtain a language genealogy graph covering the evolutionary paths, language family classification, and similarity indicators between languages through two methods: static knowledge graph construction or self-supervised learning based on language features;

[0034] A training optimization unit, which dynamically adjusts the proportion of training samples and gradient propagation weights based on the "migration path distance" and "language family similarity weight" between language pairs in the graph;

[0035] The inference optimization unit is used to introduce a bridging language mechanism. When it is detected that the migration distance between the source language and the target language in the graph exceeds a set threshold, a language whose migration distance with the source language and the target language is less than the set threshold is automatically selected from the graph as a bridging language.

[0036] The cross-language dual-space alignment and lineage-guided translation system of the present invention has the following beneficial effects compared with the prior art:

[0037] 1. This invention achieves high-precision alignment and transfer of semantic information between different languages by constructing a two-layer semantic modeling structure, introducing a language genealogy graph for training guidance, and combining it with a structure-aware attention mechanism. This effectively improves the adaptability and accuracy of multilingual translation systems in complex language environments. It is particularly suitable for multilingual communication scenarios such as international offices and cross-cultural collaboration. It solves technical problems existing in existing multilingual translation models, such as inaccurate semantic alignment, lack of structural guidance in language transfer mechanisms, and insufficient translation structure restoration capabilities.

[0038] 2. By introducing a dual-layer semantic modeling mechanism of basic semantic space and language-specific semantic space, this paper effectively solves the problem of inaccurate semantic mapping caused by differences in language structures and expression habits in traditional translation systems, significantly improving the accuracy and stability of cross-language semantic alignment;

[0039] 3. This invention constructs a language family tree composed of evolutionary paths, language family classifications, and similarity indicators between multiple languages. It introduces the evolutionary relationships and language family structures between languages into the training process, forming a migration path-aware language transfer mechanism. This can significantly improve the performance of low-resource and unpopular languages in the translation system, solving the problems of poor translation quality and weak generalization ability of low-resource languages in existing technologies.

[0040] 4. Through the structure-aware attention mechanism, the present invention can identify and adapt to the differences in word order, syntactic rules, and language type characteristics between different languages, achieving structural preservation and accurate reproduction during the translation process. It is particularly suitable for translating languages with complex syntactic structures and flexible word order.

[0041] 5. The present invention introduces a language style control mechanism during the translation generation stage, which can fine-tune the output results according to the target language and cultural context and usage scenarios (such as formal / informal, professional / general), enhance the practicality and naturalness of the translated text, and help improve the cross-cultural communication experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Attachment Figure 1 It is a system module connection diagram of the present invention. DETAILED DESCRIPTION

[0043] In order to make the technical solution, the technical problems solved and the technical effects of the present invention more clear, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.

[0044] Example 1:

[0045] Reference Attachment Figure 1 This embodiment proposes a cross-language dual-space alignment and lineage-guided translation system, which includes:

[0046] The semantic dual-space modeling module is used to feed the source language text into a basic semantic space encoder and a language-specific semantic modeler for parallel processing. It constructs a basic semantic space to represent cross-language common semantic information and a language-specific semantic space to represent semantic features related to the source language structure and expression habits. Finally, a weighted fusion strategy is used to form a unified dual-space semantic representation, providing an accurate and interpretable foundation for semantic mapping and alignment.

[0047] The semantic mapping and alignment module is used to implement tensor projection and alignment between the source and target languages based on the dual-space semantic representation provided by the semantic dual-space modeling module, using a semantic projection matrix, loss function training, and a context-fusion attention mechanism. This generates a semantic representation adapted to the target language and provides input for the translation generation module.

[0048] The translation generation module receives the target language semantic representation generated by the semantic mapping and alignment module, generates the translation text based on the target language structural features and cultural habits with the help of the structure-aware target language decoder and language style control mechanism, and outputs the translation text and intermediate results to support user review and system feedback optimization.

[0049] The structure-aware attention module is used to introduce language structure features into the encoding and decoding processes of the translation system. It assists the semantic dual-space modeling module in building semantic tensors, helps the semantic mapping and alignment module achieve precise alignment, and supports the translation generation module in generating text that meets structural requirements, thereby improving the translation system's ability to perceive and process different language structures.

[0050] The language genealogy guidance module is used to construct a language genealogy graph, encode the kinship, language family classification and evolutionary path between languages as structured prior knowledge, and dynamically adjust the training sample ratio and gradient propagation weight according to the graph during the training phase. In the inference phase, a bridging language mechanism is introduced for cross-language translation. At the same time, auxiliary information related to the language genealogy is provided to the semantic dual-space modeling module, the semantic mapping and alignment module and the translation generation module, thereby optimizing the translation effect of low-resource and word order-different languages.

[0051] In this example, the basic semantic space encoder uses the Transformer structure as the underlying network framework, introduces a cross-language shared multilingual token embedding based on the SentencePiece structure, and constructs a unified basic semantic representation tensor by learning common semantic units across languages. These common semantic units include basic actions, common entities, and general relationships.

[0052] In this embodiment, the language-specific semantic modeler models the specific expression habits of the source language. It uses the language type label, grammatical structure, and pragmatic preference information of the source language to load the structural modeling subnetwork of the corresponding language, accurately capturing the structural characteristics of the source language. At the same time, the constructed semantic tensor is strictly aligned with the basic semantic tensor in terms of dimensional design, ensuring the accuracy and uniqueness of the semantic modeling and laying a solid foundation for subsequent fusion and projection operations. Among them, the structural modeling subnetwork is a grammatical structure graph network based on adapter, prefix tuning, or lightweight customization, which is used to accurately capture the structural characteristics of the source language.

[0053] In this embodiment, the semantic mapping and alignment module uses a dual-space mapping mechanism based on a semantic projection matrix, specifically including:

[0054] The loss function unit is used to introduce the semantic preservation loss function and the structural equivalence loss function during the training phase. The semantic preservation loss function ensures that key information is not lost during the semantic mapping process between the source language and the target language, maintaining semantic accuracy. The structural equivalence loss function ensures that the structural features of the source and target languages are preserved. This allows the translation system to take into account both semantics and structure when learning the mapping relationship, laying the foundation for accurate semantic mapping.

[0055] The context-fusion attention unit is used to introduce a context-fusion attention mechanism when performing semantic mapping and alignment tasks. It dynamically aggregates the context window information of the source language text, breaking through the limitation of single sentences and taking into account the context at the paragraph or dialogue level.

[0056] The decoupling projection unit is used to decouple the basic semantics and the language-specific semantics and then project them separately, so that the semantic tensor of the target language can integrate the semantic framework and language features of the source language in a unified space, generate a semantic representation suitable for the target language, and complete the semantic mapping and alignment tasks between the source language and the target language.

[0057] In this embodiment, the translation generation module involved specifically includes:

[0058] The parsing unit is used to call the target language decoder with structure-aware function during the translation generation stage to parse the semantic tensor processed by the semantic mapping and alignment module;

[0059] The text generation unit is embedded in the target language decoder and is used to generate a translation text that conforms to the grammatical and structural requirements of the target language and incorporates the cultural conventions of the target language based on the structural features of the target language.

[0060] Language style unit, used to dynamically adjust the style of the translated text according to different usage scenarios and target audiences;

[0061] The result output unit is used to output the original text and the translated text together, along with intermediate results including semantic tensors, language tags, and structural alignment information for user review, helping users understand the translation process and basis;

[0062] The feedback optimization unit is used to transmit the translated text manually modified by the user as a training sample back to the training end, and realize system optimization by learning the training samples.

[0063] In this embodiment, during the encoding process of the translation system, the structure-aware attention mechanism module deeply analyzes the source language text and integrates various structural features, such as grammatical structure, word order, and sentence component relationships, into the encoding network in the form of structured data. For example, taking the Chinese sentence "I went to the library to read in the afternoon" and the German sentence "Ich gehe am Nachmittag indie Bibliothek, um Bücher zu lesen," Chinese has a typical subject-verb-object structure, while the German sentence "am Nachmittag" has a flexible position, and the order of the verb "gehe" and the purpose adverbial "um Bücher zu lesen" is also different from that of Chinese. By analyzing and encoding the structural features of the source language, the structure-aware attention mechanism module accurately understands the intrinsic connection between the semantics and structure of the source language. Therefore, when constructing the semantic tensor, it not only retains the semantic information but also records the structural characteristics of the source language. This provides more accurate and comprehensive input for the semantic dual-space modeling module, helping it to construct a basic semantic space and a language-specific semantic space that better conforms to the actual situation of the source language.

[0064] During the translation system's decoding process, the structure-aware attention mechanism module accurately identifies the target language's grammatical rules, common sentence structures, and linguistic conventions, incorporating these features into the decoding network. For example, when translating from the source language into Japanese, where the object often precedes the predicate and contains numerous honorifics and special expressions, the structure-aware attention mechanism guides the decoding network to reference the target language's structural features. This assists the translation system in generating translated text, adhering to the target language's grammatical conventions and linguistic conventions, and generating sentence structures that meet the target language's structural requirements.

[0065] After the semantic mapping and alignment module completes the semantic processing, the structure-aware attention mechanism module further ensures that the structure of the translation is highly consistent with the target language, achieving dual-precision conversion of semantics and structure, and providing guarantees for the translation generation module to output high-quality translation text that conforms to the target language specifications, ultimately effectively improving the accuracy and fluency of the translation system in processing different language structures.

[0066] In this embodiment, the language lineage guidance module specifically includes:

[0067] A graph construction unit is used to obtain a language genealogy graph covering the evolutionary paths, language family classification, and similarity indicators between languages through two methods: static knowledge graph construction or self-supervised learning based on language features;

[0068] The training optimization unit dynamically adjusts the proportion of training samples and gradient propagation weights based on the "migration path distance" and "language family similarity weight" between language pairs in the graph. For closely related language pairs, such as Hindi and Urdu, the system increases their training weight and sample frequency, leveraging the similarities between closely related languages to achieve more efficient semantic transfer and improve the translation model's ability to handle such languages.

[0069] The inference optimization unit is used to introduce a bridging language mechanism. When it detects that the migration distance between the source language and the target language in the graph exceeds a set threshold, it automatically selects a language from the graph whose migration distance to the source and target languages is less than the set threshold as the bridging language. Specifically, a language in the graph that is close to both the source and target languages (such as English or French) is selected as an intermediary, splitting the semantic representation migration process into two stages: first migrating from the source language to the bridging language, and then migrating from the bridging language to the target language. This method mitigates the precision loss that may be caused by direct mapping and improves the accuracy of cross-language translation.

[0070] At the same time, the language genealogy map constructed by the language genealogy guidance module can provide assistance to the semantic dual-space modeling module, semantic mapping and alignment module, and translation generation module, and comprehensively optimize the translation effect of low-resource and word order-different languages.

[0071] In summary, the cross-language dual-space alignment and lineage-guided translation system of the present invention, by constructing a two-layer semantic modeling mechanism of basic semantic space and language-specific semantic space, combined with the guided training strategy of language evolution path map, and introducing structure-aware language attention mechanism, improves the semantic mapping accuracy between different languages while enhancing the translation system's modeling ability of language structure and language family evolution characteristics, thereby significantly improving the translation quality and stability of the translation system in scenarios with multiple languages, low-resource languages and large differences in language structures.

[0072] The above specific examples are used to illustrate the principles and implementation methods of the present invention in detail. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by those skilled in the art without departing from the principles of the present invention should fall within the scope of patent protection of the present invention.

Claims

1. A cross-language dual-space alignment and lineage-guided translation system, characterized by: It includes: The semantic dual-space modeling module is used to feed the source language text into a basic semantic space encoder and a language-specific semantic modeler for parallel processing. It constructs a basic semantic space to represent cross-language common semantic information and a language-specific semantic space to represent semantic features related to the source language structure and expression habits. Finally, a weighted fusion strategy is used to form a unified dual-space semantic representation, providing an accurate and interpretable foundation for semantic mapping and alignment. The semantic mapping and alignment module is used to implement tensor projection and alignment between the source and target languages based on the dual-space semantic representation provided by the semantic dual-space modeling module, using a semantic projection matrix, loss function training, and a context-fusion attention mechanism. This generates a semantic representation adapted to the target language and provides input for the translation generation module. The translation generation module receives the target language semantic representation generated by the semantic mapping and alignment module, generates the translation text based on the target language structural features and cultural habits with the help of the structure-aware target language decoder and language style control mechanism, and outputs the translation text and intermediate results to support user review and system feedback optimization. The structure-aware attention module is used to introduce language structure features into the encoding and decoding processes of the translation system. It assists the semantic dual-space modeling module in building semantic tensors, helps the semantic mapping and alignment module achieve precise alignment, and supports the translation generation module in generating text that meets structural requirements, thereby improving the translation system's ability to perceive and process different language structures. The language genealogy guidance module is used to construct a language genealogy graph, encode the kinship, language family classification and evolutionary path between languages as structured prior knowledge, and dynamically adjust the training sample ratio and gradient propagation weight according to the graph during the training phase. In the inference phase, a bridging language mechanism is introduced for cross-language translation. At the same time, auxiliary information related to the language genealogy is provided to the semantic dual-space modeling module, the semantic mapping and alignment module and the translation generation module, thereby optimizing the translation effect of low-resource and word order-different languages.

2. A cross-language dual-space alignment and lineage-guided translation system according to claim 1, characterized in that: The basic semantic space encoder uses the Transformer structure as the underlying network framework, introduces cross-language shared multilingual token embeddings based on SentencePiece construction, and constructs a unified basic semantic representation tensor by learning cross-language common semantic units. Among them, the semantic units common across languages include basic actions, common entities and general relations.

3. A cross-language dual-space alignment and lineage-guided translation system according to claim 2, characterized in that: The language-specific semantic modeler models the specific expression habits of the source language. It uses the language type label, grammatical structure, and pragmatic preference information of the source language to load the structural modeling sub-network of the corresponding language, accurately capturing the structural characteristics of the source language. At the same time, the constructed semantic tensor is strictly aligned with the basic semantic tensor in terms of dimensional design, ensuring the accuracy and uniqueness of the semantic modeling, laying a solid foundation for subsequent fusion and projection operations. The structural modeling sub-network is a grammatical structure graph network based on Adapter, Prefix Tuning or lightweight customization, which is used to accurately capture the structural characteristics of the source language.

4. A cross-language dual-space alignment and lineage-guided translation system according to claim 3, characterized in that: The semantic mapping and alignment module adopts a dual-space mapping mechanism based on the semantic projection matrix, which specifically includes: The loss function unit is used to introduce the semantic preservation loss function and the structural equivalence loss function during the training phase. The semantic preservation loss function ensures that key information is not lost during the semantic mapping process between the source language and the target language, maintaining semantic accuracy. The structural equivalence loss function ensures that the structural features of the source and target languages are preserved. This allows the translation system to take into account both semantics and structure when learning the mapping relationship, laying the foundation for accurate semantic mapping. The context-fusion attention unit is used to introduce a context-fusion attention mechanism when performing semantic mapping and alignment tasks. It dynamically aggregates the context window information of the source language text, breaking through the limitation of single sentences and taking into account the context at the paragraph or dialogue level. The decoupling projection unit is used to decouple the basic semantics and the language-specific semantics and then project them separately, so that the semantic tensor of the target language can integrate the semantic framework and language features of the source language in a unified space, generate a semantic representation suitable for the target language, and complete the semantic mapping and alignment tasks between the source language and the target language.

5. A cross-language dual-space alignment and lineage-guided translation system according to claim 4, characterized in that: The translation generation module specifically includes: The parsing unit is used to call the target language decoder with structure-aware function during the translation generation stage to parse the semantic tensor processed by the semantic mapping and alignment module; The text generation unit is embedded in the target language decoder and is used to generate a translation text that conforms to the grammatical and structural requirements of the target language and incorporates the cultural conventions of the target language based on the structural features of the target language. Language style unit, used to dynamically adjust the style of the translated text according to different usage scenarios and target audiences; The result output unit is used to output the original text and the translated text together, along with intermediate results including semantic tensors, language tags, and structural alignment information for user review, helping users understand the translation process and basis; The feedback optimization unit is used to transmit the translated text manually modified by the user as a training sample back to the training end, and realize system optimization by learning the training sample.

6. A cross-language dual-space alignment and lineage-guided translation system according to claim 4, characterized in that: During the encoding process of the translation system, the structure-aware attention mechanism module deeply analyzes the source language text and integrates various structural features such as grammatical structure, word order, and sentence component relationships into the encoding network in the form of structured data. By analyzing and encoding the structural features of the source language, it accurately understands the intrinsic connection between the semantics and structure of the source language. Therefore, when constructing the semantic tensor, it can not only retain semantic information but also record the structural characteristics of the source language, providing more accurate and comprehensive input for the semantic dual-space modeling module, helping it to construct a basic semantic space and language-specific semantic space that are more consistent with the actual situation of the source language. During the decoding process of the translation system, the structure-aware attention mechanism module accurately identifies the grammatical rules, common sentence structures, and various structural features of language habits of the target language, and introduces these features into the decoding network. By guiding the decoding network to refer to the structural features of the target language, the auxiliary translation system follows the grammatical norms and expression habits of the target language when generating translated texts, and generates sentence structures that meet the structural requirements of the target language.

7. A cross-language dual-space alignment and lineage-guided translation system according to claim 6, characterized in that: After the semantic mapping and alignment module completes the semantic processing, the structure-aware attention mechanism module further ensures that the translation is highly consistent with the target language in structure, achieving dual-precision conversion of semantics and structure, and providing a guarantee for the translation generation module to output high-quality translation text that conforms to the target language specifications.

8. A cross-language dual-space alignment and lineage-guided translation system according to claim 4, characterized in that: The language pedigree guidance module specifically includes: A graph construction unit is used to obtain a language genealogy graph covering the evolutionary paths, language family classification, and similarity indicators between languages through two methods: static knowledge graph construction or self-supervised learning based on language features; The training optimization unit dynamically adjusts the proportion of training samples and gradient propagation weights based on the "migration path distance" and "language family similarity weight" between language pairs in the graph; The inference optimization unit is used to introduce a bridging language mechanism. When it is detected that the migration distance between the source language and the target language in the graph exceeds a set threshold, a language whose migration distance with the source language and the target language is less than the set threshold is automatically selected from the graph as a bridging language.

Citation Information

Cited By

  • Semantic analysis and recognition method based on artificial intelligence

    CN121257548A