Ancient Chinese language translation method, system and equipment based on large model and storage medium

By performing higher-level semantic standardization and diachronic semantic modeling on classical Chinese passages, and combining this with a knowledge graph of classical Chinese, multiple modern Chinese translation tracks and annotation tracks are generated. This solves the problems of accuracy and stylistic consistency in classical Chinese translation, and achieves efficient and verifiable classical Chinese translation.

CN120996057AActive Publication Date: 2025-11-21COMMUNICATION UNIVERSITY OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511411391.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-11-21
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing methods for translating classical Chinese texts lack cross-documentary and cross-disciplinary knowledge integration, making it difficult to accurately convey the semantic features and cultural background of classical Chinese texts, resulting in insufficient accuracy and inconsistent style in the translations.

Method used

By performing higher-level semantic standardization on classical Chinese passages, combined with diachronic semantic modeling and classical Chinese knowledge graphs, multiple modern Chinese translation tracks and annotation tracks are generated. Evidence coverage screening and textual rearrangement are then performed to ensure the accuracy and stylistic consistency of the translations.

Benefits of technology

It achieves accurate transmission of the meaning of classical Chinese sentences in modern Chinese, provides verifiable academic translations and elegant translations, and meets the needs of different readers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996057A_ABST
    Figure CN120996057A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing, and particularly relates to an ancient language translation method, system and device based on a large model and a storage medium. The method comprises the following steps: performing context coding processing according to an ancient Chinese language processing result to fuse font information and training information, and combining with duration semantic modeling to obtain context semantic representation corresponding to an ancient Chinese language paragraph; inputting the context semantic representation and the evidence set into a translation double-track decoder to generate a plurality of modern Chinese translation tracks and modern Chinese annotation tracks; performing constraint screening on the modern Chinese translation track and the modern Chinese annotation track according to the evidence set to obtain candidate translation results meeting evidence coverage requirements; and performing text-text rearrangement according to the candidate translation result, performing executable processing on the traceability identifier in the annotation track, and outputting a translation result containing modern translations and annotations. The translation obtained through the method is graceful in language feeling, clear in annotation and operable, the reading experience is met, and the academic tracking requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of natural language processing, and particularly relates to a large model-based ancient text translation method, system, device and storage medium. BACKGROUND

[0002] With the development of artificial intelligence and natural language processing technology, a large language model (LLM) based translation technology has emerged. Through a neural network structure trained on a large amount of corpus, it has strong context modeling and semantic reasoning capabilities, and can achieve smooth and natural translation between modern languages. Compared with traditional rule-based or statistical translation methods, large model translation shows significant advantages in semantic understanding, contextual coherence and natural expression.

[0003] In existing ancient text translation research, the common method is to rely on artificially compiled dictionaries, rule systems and limited bilingual parallel corpus for translation. That is, by constructing ancient and modern dictionaries and rule templates, ancient Chinese vocabulary is mapped to modern Chinese interpretation, and then sentence-level combination and reconstruction are completed through syntactic analysis; or statistical translation or neural network translation framework is introduced, ancient text is treated as a low-resource language, and modeling and translation are performed through a small amount of aligned corpus.

[0004] However, the above method, the ancient text corpus is scarce and lacks systematic annotation, making it difficult for statistical models and neural models to fully learn the semantic features unique to ancient text; ancient language has strong diachronic differences, and the same vocabulary may have different meanings in different historical periods; existing translation methods mostly rely on single corpus or dictionary resources, lack of cross-document and cross-domain knowledge fusion, and are prone to cause insufficient accuracy of translation; traditional neural network translation methods pay more attention to formal syntactic conversion, but lack the ability to model the tone, style and cultural background implied by ancient text, resulting in a lack of overall style consistency and cultural explanatory power in the translation. SUMMARY

[0005] Therefore, it is necessary to provide a large model-based ancient text translation method, system, device and storage medium that can meet the accuracy, context consistency and knowledge coverage.

[0006] In a first aspect, the application provides a large model-based ancient text translation method, comprising:

[0007] An ancient text paragraph to be translated is obtained, and a superordinate semantic standardization process is performed on the ancient text paragraph to obtain an ancient text processing result with structured information; the ancient text processing result includes word-level information, construction recognition results and dynasty markers; the superordinate semantic standardization process corresponds to sentence segmentation, construction recognition and dynasty marking;

[0008] According to the ancient text processing result, context coding processing is performed to fuse the character form information and the annotation information, and diachronic semantic modeling is combined to obtain a context semantic representation corresponding to the ancient text paragraph;

[0009] According to the context semantic representation, similarity retrieval is performed in an ancient Chinese knowledge graph to obtain an evidence set corresponding to the ancient text paragraph;

[0010] The context semantic representation and the evidence set are input into a translation and annotation dual-track decoder to generate a plurality of modern Chinese translation tracks and modern Chinese annotation tracks;

[0011] According to the evidence set, constraint screening is performed on the modern Chinese translation tracks and the modern Chinese annotation tracks to obtain a candidate translation and annotation result satisfying an evidence coverage requirement;

[0012] According to the candidate translation and annotation result, text qi rearrangement is performed, and traceable identification in the annotation track is executable, and a translation result containing modern translation and annotation is output.

[0013] In one of the embodiments, according to the ancient text processing result, context coding processing is performed to fuse the character form information and the annotation information, and diachronic semantic modeling is combined to obtain a context semantic representation corresponding to the ancient text paragraph, including:

[0014] According to the character-level information, stroke sequences and component structures of the characters of the ancient text paragraph are extracted, and the stroke sequences and the component structures are coded to obtain character form vectors;

[0015] According to the character-level information, the diachronic annotations and the example texts corresponding to the words of the ancient text paragraph are coded to obtain semantic item vectors;

[0016] The character form vectors and the semantic item vectors are fused to obtain a corresponding fusion representation;

[0017] According to the dynasty mark, a dynasty time mark vector is obtained, and diachronic semantic modeling is performed on the fusion representation according to the dynasty time mark vector to obtain a diachronic enhanced representation;

[0018] According to the diachronic enhanced representation and the construction form recognition result, joint modeling is performed to obtain the context semantic representation.

[0019] In one of the embodiments, according to the context semantic representation, similarity retrieval is performed in an ancient Chinese knowledge graph to obtain an evidence set corresponding to the ancient text paragraph, including:

[0020] Based on the ancient Chinese knowledge graph, similarity retrieval is performed on the context semantic representation and the node vector representation to obtain a candidate node;

[0021] According to the candidate node, a classical allusion node, an annotation item, and a quotation segment corresponding to the ancient text paragraph are obtained to form an initial evidence set;

[0022] The evidences in the initial evidence set are weighted according to the similarity score, dynasty time scale consistency, style matching degree and entity alignment relationship, and are sorted according to the weighting results to obtain an evidence set corresponding to the ancient text paragraph.

[0023] In one of the embodiments, the ancient Chinese knowledge graph is constructed by the following method:

[0024] The ancient book corpus, annotation data, variant character mapping data, phonetic data and personage, place name and official position data are obtained and standardized to obtain an ancient text related knowledge index library;

[0025] The nodes and relationship edges are extracted from the ancient text related knowledge index library; the nodes include characters, words, semantic items, constructions, allusions, quotations, personages, place names, official positions, dynasties and annotation items; the relationship edges include diachronic evolution relationship, homonym relationship, allusion relationship, use case relationship and chapter relationship;

[0026] The node vector representation is generated for the nodes, and the ancient Chinese knowledge graph is formed according to the nodes, corresponding node attribute node vector representation and relationship edges.

[0027] In one of the embodiments, the context semantic representation and the evidence set are input into a translation and annotation dual-track decoder to generate a plurality of modern Chinese translation tracks and modern Chinese annotation tracks, including:

[0028] Based on the context semantic representation, the modern Chinese translation tracks are generated according to different translation styles;

[0029] According to the context semantics, the annotation tracks are generated by performing word meaning and syntax segmentation on the basis of the context semantics, and the traceability identifiers corresponding to the evidence set are marked in the annotation tracks to obtain the modern Chinese annotation tracks.

[0030] In one of the embodiments, the modern Chinese translation tracks and the modern Chinese annotation tracks are constrained and screened according to the evidence set to obtain candidate translation and annotation results that meet the evidence coverage requirement, including:

[0031] The constraint lattice is generated according to the evidence set; the constraint lattice contains constraint conditions of must cover, optional cover and prohibited cover;

[0032] Based on the constraint lattice, the modern Chinese translation tracks and the modern Chinese annotation tracks are decoded, and the translation and annotation that violates the consistency penalty term is removed to obtain a plurality of candidate translation and annotation results; the consistency penalty term includes a penalty weight when the time and person relationship between the translation and the constraint lattice are contradictory;

[0033] The candidate translation and annotation results are screened based on the evidence coverage rate to obtain the candidate translation and annotation results that meet the evidence coverage requirement.

[0034] In one of the embodiments, text atmosphere rearrangement is performed according to the candidate translation result, and the traceable identifier in the annotation track is executable, and the translation result containing the modern translation and the annotation is output, including:

[0035] The candidate translation result is scored for text atmosphere, and a text atmosphere scoring result is obtained; the text atmosphere scoring includes calculation of parallelism, comparison, rhythm, rhyme and rhetorical retention;

[0036] The candidate translation result is rearranged according to the text atmosphere scoring result, and a text atmosphere optimized translation result is obtained;

[0037] The traceable identifier is corresponded to the knowledge graph node, and the annotation path template is rendered according to the preset user view, and the translation result containing the modern translation and the annotation is obtained.

[0038] In a second aspect, the present application also provides an ancient text translation system based on a large model, including:

[0039] The standardization processing module is configured to obtain an ancient text paragraph to be translated, and perform superordinate semantic standardization processing on the ancient text paragraph to obtain an ancient text processing result with structured information;

[0040] The semantic extraction module is configured to fuse character form information and textual criticism information according to context coding processing of the ancient text processing result, and obtain a context semantic representation corresponding to the ancient text paragraph by combining diachronic semantic modeling;

[0041] The evidence module is configured to perform similarity retrieval in an ancient Chinese knowledge graph according to the context semantic representation to obtain an evidence set corresponding to the ancient text paragraph;

[0042] The translation and annotation dual-track module is configured to input the context semantic representation and the evidence set into a translation and annotation dual-track decoder to generate a plurality of modern Chinese translation tracks and modern Chinese annotation tracks;

[0043] The evidence constraint module is configured to constrain and screen the modern Chinese translation tracks and the modern Chinese annotation tracks according to the evidence set to obtain a candidate translation result meeting evidence coverage requirements;

[0044] The text atmosphere rearrangement module is configured to perform text atmosphere rearrangement according to the candidate translation result, and perform executable processing on the traceable identifier in the annotation track to output a translation result containing the modern translation and the annotation.

[0045] In a third aspect, the present application also provides a computer device including a memory and a processor, the memory stores a computer program, and the processor implements the steps of any of the above-mentioned ancient text translation methods based on a large model when executing the computer program.

[0046] In a fourth aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of any of the above-mentioned ancient Chinese translation methods based on large models.

[0047] The above-mentioned ancient Chinese translation method, system, device and storage medium based on large models can reduce ambiguous mistranslation and ensure accurate transmission of ancient Chinese sentence meaning in modern Chinese through context semantic representation and knowledge graph evidence driving. The annotation track and the traceable identifier are closely combined, so that each translation or explanation can be traced back to the original text source or related notes, and the verifiability of academic translation is realized. The translation and annotation dual-track decoder can generate multiple translation tracks according to different styles, and the annotation track provides detailed explanations, which can meet the needs of ordinary readers for easy-to-read translation and the needs of academic research for detailed interpretation of the original text. Through the text sentiment scoring and rearrangement mechanism, the couplet, parallelism, rhythm, rhyme and rhetorical features of the ancient Chinese can be maintained to the greatest extent while ensuring accurate semantics, so that the translation result is more natural and beautiful. From text analysis, semantic modeling, evidence retrieval to translation and annotation generation and optimization, the workload of manual translation and annotation is significantly reduced, and the efficiency of ancient Chinese translation is improved. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the embodiment or related art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0049] Figure 1 The flowchart of the ancient Chinese translation method based on large models of the present application is shown in the figure.

[0050] Figure 2 The flowchart of the sub-step of step S102 is shown in the figure.

[0051] Figure 3 The flowchart of the sub-step of step S105 is shown in the figure.

[0052] Figure 4 The composition structure diagram of the ancient Chinese translation system based on large models of the present application is shown in the figure. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0054] In one embodiment, as shown in Figure 1As shown, a large model-based ancient text translation method is provided. In this embodiment, the method is applied to a terminal. It should be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and can be implemented through the interaction of the terminal and the server. In this embodiment, the method includes the following steps:

[0055] In S101, an ancient text paragraph to be translated is obtained, and the ancient text paragraph is subjected to superordinate semantic standardization processing to obtain an ancient text processing result with structured information. The ancient text processing result includes word-level information, construction recognition result, and dynasty marking. The superordinate semantic standardization processing corresponds to sentence breaking, construction recognition, and dynasty marking.

[0056] By way of illustration, the original ancient text paragraph is preprocessed by using a sentence breaking algorithm, a construction recognition mechanism, and a dynasty marking method. Specifically, sentence breaking is completed by a joint model based on statistical learning and semantic constraints. The model can automatically determine the sentence boundaries in the ancient text by comprehensively considering the sentence reading symbol, common word collocation rules, and semantic rationality. Construction recognition is completed by constructing a special syntactic template library for ancient Chinese to detect and mark typical sentence structures in the ancient text paragraph, such as passive-active inversion, omitted predicate, and anaphora structure. Dynasty marking relies on a diachronic dictionary and corpus statistics to automatically identify the time background of the paragraph, so as to provide time and space constraints for subsequent semantic interpretation. After preprocessing, the ancient text paragraph is converted into a processing result with structured information, including word-level information such as the part of speech, radical, and character shape features of a single character, construction recognition results such as syntactic skeletons and semantic relationship labels, and dynasty marking, i.e., the historical period label corresponding to the paragraph.

[0057] In S102, context coding processing is performed according to the ancient text processing result to fuse character shape information and exegesis information, and diachronic semantic modeling is combined to obtain a context semantic representation corresponding to the ancient text paragraph.

[0058] Context coding is performed based on the ancient text processing result to obtain a deep semantic representation suitable for inputting a large model. Specifically, the character shape information and the exegesis information are fused, and diachronic semantic modeling is combined to improve the accuracy of ancient text understanding. By way of illustration, a character shape-exegesis encoder is based on consistency loss h c (x,d) is a context-aware character representation, s c,d is exegesis information, g cwherein, is the weight coefficient. Among them, the character information refers to the visual and structural characteristics of ancient characters in the writing system, such as the stroke configuration of oracle bone inscriptions, seal script or cursive script and its change rule in the evolution process. The above characteristics are usually modeled by convolutional coding or graph convolutional network; and the interpretation information refers to the specific meaning and usage of words or phrases in ancient books, which often has polysemy and temporal differences, and therefore needs to be represented by embedded dictionary vectors. Diachronic semantic modeling is to capture the semantic evolution trajectory of the same word in different dynasties on the time axis. Exemplarily, a time-aware embedding model is usually used to model the semantic distribution of the same word in different periods, while combining multi-task training to improve the perception of ancient Chinese grammar phenomena.

[0059] Through the fusion mechanism of multiple channels, the above information is jointly encoded into the context semantic representation, so as to retain the shape information of ancient characters and integrate the knowledge of interpretation and semantic evolution.

[0060] S103, similarity retrieval in the ancient Chinese knowledge graph according to the context semantic representation, to obtain an evidence set corresponding to the ancient text paragraph.

[0061] The context semantic representation will be compared with the external ancient Chinese knowledge graph to obtain a reliable evidence set. The ancient Chinese knowledge graph is a semantic network, whose nodes usually include ancient characters, books, characters, place names, historical events, interpretation, etc., and the edges record the semantic or diachronic relationship between them. Exemplarily, the corresponding meaning items of the word "Ren" in "Analects of Confucius" and "Mencius" may differ, and the ancient Chinese knowledge graph can distinguish them by time-meaning relationship. Specifically, the similarity between the context representation and the knowledge graph nodes is calculated by the method of semantic vector retrieval, which can use vector matching based on cosine similarity or interactive matching based on deep retrieval network. The final evidence set usually contains glossary explanations, synonymous replacements, context examples and diachronic corresponding relationships related to the ancient text paragraph, thereby providing external knowledge constraints for subsequent translation.

[0062] S104, input the context semantic representation and the evidence set into the translation and annotation dual-track decoder to generate multiple modern Chinese translation tracks and modern Chinese annotation tracks.

[0063] The contextual semantic representation and the evidence set are jointly input into a translation-annotation dual-track decoder to simultaneously generate a modern Chinese translation track and an annotation track. The translation-annotation dual-track decoder is an improved large model decoding architecture that internally includes two interrelated generation paths, one for generating a natural and fluent modern Chinese translation track and the other for generating an annotation track with explanatory annotation information. During training, the translation track is aligned and trained using a large-scale ancient and modern parallel corpus, while the annotation track is learned from dictionaries, classical book annotations, and expert annotation data, enabling it to generate explanations, source prompts, and diachronic information related to the translation. In the decoding process, the contextual semantic representation mainly provides semantic support for the text to be translated, while the evidence set is used as an external knowledge input to guide the model to maintain accuracy and consistency when selecting translation words and generating annotations.

[0064] S105, according to the evidence set, the modern Chinese translation track and the modern Chinese annotation track are constrained and screened to obtain a candidate translation-annotation result that meets the evidence coverage requirement.

[0065] Further, the generated translation track and annotation track are constrained and screened to ensure the reliability and coverage of the results. Specifically, the candidate translation-annotation result is compared with the evidence set, and if the translation words or annotations fail to cover the core information required by the evidence, they will be excluded or reweighted. At the same time, different candidate translation-annotation results are given weighted scores according to the confidence of evidence matching, so as to preferentially output the result with the highest consistency with the knowledge graph, effectively avoiding the problem of large model random translation, and ensuring the accuracy of the translation in semantic explanation and historical research.

[0066] S106, according to the candidate translation-annotation result, the text is rearranged, and the traceable identification in the annotation track is processed to output a translation result containing modern translation and annotations.

[0067] Illustratively, the candidate translation-annotation result is rearranged and the traceable identification of the annotation is processed. Text rearrangement refers to the optimization of style between multiple candidate translations, so that it is more in line with the expression habits of modern Chinese as a whole, while maintaining the original tone and style characteristics of ancient texts. Illustratively, when translating a passage from "Shiji", not only the accuracy of the meaning should be ensured, but also the brevity and solemnity of the historical record should be maintained in expression. The executable processing of traceable identification is to convert the source information generated in the annotation track into clickable or callable links, so that users can directly trace the source of the translation or the knowledge graph node, enhancing the explainability and academic reliability of the translation result.

[0068] In the above large model-based ancient text translation method, the ancient text paragraph is subjected to superordinate semantic standardization processing, including sentence breaking, construction identification and dynasty marking, and a structured ancient text processing result with word-level information, construction identification result and dynasty marking is generated, so that the hierarchical structure and language characteristics of the ancient text can be accurately understood, thereby providing a reliable foundation for subsequent encoding and translation. By fusing the word form information and the textual information, and combining the diachronic semantic modeling, a diachronically enhanced context semantic representation can be obtained, effectively capturing the polysemy, diachronic evolution and semantic nuances of the ancient text, thereby improving the accuracy and depth of translation. Similarity retrieval is performed using the ancient Chinese knowledge graph, the context semantic representation is matched with the knowledge graph nodes to form an evidence set, so that the translation generation not only relies on statistical or deep learning inference, but also combines literature and historical evidence, thereby ensuring that the translation content conforms to the academic and historical logic. By generating a modern Chinese translation track and an annotation track using a translation and annotation dual-track decoder, and combining the evidence set to build a constraint lattice for screening, high consistency and traceability of the translation and annotation are achieved, avoiding arbitrary interpretation or deviation from the evidence. On the basis of the candidate translation and annotation results, text sentiment scoring and rearrangement are performed, and the traceable identifiers in the annotation track are matched with the knowledge graph nodes for executable rendering, so that the translation is not only elegant in language, but also clear and operable in annotation, so that the final output meets the reading experience and academic tracking needs.

[0069] In one embodiment, as shown in Figure 2 According to the ancient text processing result, the context encoding processing fuses the word form information and the textual information, and combines the diachronic semantic modeling to obtain the context semantic representation corresponding to the ancient text paragraph, including:

[0070] S201, extract the stroke sequence and component structure of the word of the ancient text paragraph according to the word-level information, and encode the stroke sequence and component structure to obtain a word form vector.

[0071] The same word in ancient text often has multiple writing methods, and the evolution of word forms between different dynasties is also relatively complex. If only relying on dictionary-based word mapping, it is often difficult to capture deep rules. Illustratively, two types of structural features are used to depict word form information, one is stroke sequence, and the other is component structure. The stroke sequence can be extracted from the existing word form database to obtain the stroke order of each Chinese character, and then converted into a symbol sequence, which is input into a sequence encoding model. Illustratively, a bidirectional long short-term memory network (BiLSTM) or a convolutional neural network is used to obtain the encoding result at the stroke level; the component structure is obtained by splitting the character components, such as radical, component, and phonetic component, and the components are regarded as a structural graph, which is modeled using a graph neural network, thereby obtaining a representation reflecting the construction rules of Chinese characters. The stroke encoding and component encoding are spliced or fused to obtain a word form vector.

[0072] Exemplarily, when processing the character "義", extract its stroke sequence, i.e., the order of the left-falling stroke, horizontal stroke, vertical stroke, hook, etc., and identify its component structure as a combination of "羊" and "我". After encoding, a glyph vector that contains both physical features and structural features is obtained.

[0073] S202. Encode the ancient annotations and example texts corresponding to the words and phrases in the ancient Chinese paragraph according to the word-level information to obtain an item vector.

[0074] The meanings of words in ancient Chinese may vary significantly in different dynasties and contexts. Therefore, it is difficult to support high-quality translation based solely on modern interpretations.示意性的,可以基于字级别信息,检索预先构建的历代注疏与例证文本数据库。示例性的,若输入“仁”字,在《论语注疏》《孟子集注》以及历代典籍用例中提取其释义与例句,将语料通过Transformer编码器或基于预训练语言大模型的表示层转化为向量表示,即义项向量,不仅包含字词在古文中的义项信息,还融入了历史语料中的上下文使用规律。示例性的,“仁”在先秦语境下往往对应于德性之核心,而在后世佛教典籍中可能还会带有慈悲之义,语义细微差别可以通过义项向量加以刻画。

[0075] S203. Fuse the glyph vector and the item vector to obtain the corresponding fused representation.

[0076] Fuse the glyph vector and the item vector to form a more complete symbolic semantic representation. The fusion method can be diverse, such as vector concatenation, gating mechanism weighting, or interaction modeling based on the attention mechanism. The model can not only retain the external morphological features of the text but also combine the exegetical explanations in the semantic dimension to achieve cross-modal representation enhancement. Exemplarily, the glyph vector of the character "義" reflects its structure composed of "羊" and "我", while the item vector records its semantics of "justice, fairness". After the two are fused, both the form and semantics can be taken into account.

[0077] S204. Obtain a dynasty time-stamp vector according to the dynasty label, and perform diachronic semantic modeling on the fused representation according to the dynasty time-stamp vector to obtain a diachronically enhanced representation.

[0078] Further, the diachronic evolution of semantics is considered. Ancient Chinese semantics relies heavily on historical context, and the same vocabulary often has different usage in different dynasties. Illustratively, according to the dynasty markers in the ancient Chinese processing result, a dynasty time vector is generated, and embedding technology can be used to convert dynasty information into vector space representation. The dynasty time vector is combined with the fusion representation, and diachronic semantic modeling is used to enhance the time sensitivity of the semantic representation. Diachronic semantic modeling can use time-aware language modeling methods, such as introducing a time dimension bias in the semantic space, or building a time series Transformer, which trains segmented corpus of different periods, so that the model can capture the evolution trajectory of semantics. Illustratively, “doctor” in the Han Dynasty refers to a scholar in charge of classics, while in modern times it is an academic degree. After diachronic semantic modeling, the model will be more inclined to generate semantic interpretations consistent with the Han Dynasty context. Illustratively, a learning vector t is set for each dynasty d d , and time continuity regularization |t d -t d-1 is introduced. The semantic drift of the same word in adjacent dynasties is supervised by known annotations, and contrastive learning is performed to make correct semantic vectors close and incorrect semantic vectors far apart, i.e., through diachronic contrastive loss Contrastive learning is performed, where w is a specific word or word unit, is a positive sample representation, which refers to the semantic representation of the word or word unit in adjacent or similar historical periods d; is a negative sample representation, which refers to a different other semantic unit, or a different period without semantic correspondence.

[0079] S205, diachronic enhanced representation and construction recognition result are jointly modeled according to the context semantic representation.

[0080] Construction recognition result refers to the recognition of specific grammatical constructions, such as “with……for……” “if……then……” and other fixed sentence patterns, by analyzing the structure of ancient Chinese sentences. Constructions often have specific functions and semantics in ancient Chinese, and therefore are indispensable in context modeling. By combining diachronic enhanced representation with construction recognition result, a multi-channel attention mechanism can be used to enable the large model to understand both the diachronic semantics of individual words or words and the meaning of the overall grammatical construction. Illustratively, in the sentence “with virtue to serve people”, the construction “with……to serve people” will prompt the model that this is a cause-and-effect expression, which will work together with the diachronic enhanced representation to ultimately obtain a context semantic representation that can more accurately reflect the deep meaning of the ancient Chinese passage.

[0081] In the above method, the in-depth integration of glyph information, exegetical information, and diachronic information, combined with the results of construction recognition, finally generates a context semantic representation with rich semantic levels. This representation not only has multiple dimensions of glyphs and semantic items at the single-character level but also has context constraints of historical evolution and construction rules at the sentence level, thus providing a solid semantic foundation for subsequent translation and annotation generation.

[0082] In one embodiment, similarity retrieval is performed in the ancient Chinese knowledge graph according to the context semantic representation to obtain an evidence set corresponding to the ancient Chinese paragraph, including:

[0083] S11. Based on the ancient Chinese knowledge graph, similarity retrieval is performed between the context semantic representation and the node vector representation to obtain candidate nodes.

[0084] Schematically, the context semantic representation is mapped into the same semantic space as the nodes of the knowledge graph. Among them, the nodes of the knowledge graph usually obtain their vectorized expressions through pre-trained graph representation learning methods, such as graph neural networks, TransE, RotatE and other embedding models, so that each allusion, entry, or citation fragment can be represented in vector form. The similarity measurement method can use cosine similarity, dot product, or distance-based measurement functions. When the similarity between the context semantic representation and some node vectors exceeds a preset threshold or ranks among the top several, the corresponding nodes are selected as candidate nodes. Exemplarily, when the input ancient Chinese is "鸿鹄高飞" (swans flying high), it may retrieve in the knowledge graph the allusion nodes related to "鸿鹄" (swans) and the poem citation nodes related to "高飞" (flying high), thus becoming part of the candidate set.

[0085] S12. According to the candidate nodes, obtain the allusion nodes, annotation entries, and citation fragments corresponding to the ancient Chinese paragraph to form an initial evidence set.

[0086] It is necessary to map the knowledge content associated with the candidate nodes into directly usable evidence. Since the nodes in the knowledge graph are often connected to various types of text resources, such as allusion entries, historical annotations, poem fragments, or historical document citations, the materials directly associated with them are obtained according to the link relationships and adjacent edges of the candidate nodes. The formed initial evidence set can be either an explanatory entry for single-character exegesis or a cross-sentence citation from ancient classics.

[0087] S13. Weigh each evidence in the initial evidence set according to the similarity score, dynasty time scale consistency, genre matching degree, and entity alignment relationship, and sort according to the weighted results to obtain an evidence set corresponding to the ancient Chinese paragraph.

[0088] The final evidence set is screened by multi-dimensional constraints to ensure its accuracy and relevance. Specifically, the similarity score is the most basic indicator, which measures the semantic closeness between the context semantic representation and the candidate node. The dynasty time consistency constraint is used to ensure that the semantic interpretation is consistent with the time background of the target ancient text. For example, the interpretation of Han Dynasty texts should not be directly replaced by Song Dynasty annotations, so the dynasty time of the evidence and the dynasty mark of the input text need to be matched. The genre matching degree refers to whether the text style attribute of the candidate evidence is consistent with the genre of the text to be translated. For example, prose should refer more to historical prose annotations, while poetry tends to quote poetry corpus. The entity alignment relationship is used to determine whether the core entity in the evidence corresponds to the entity in the original text, such as “Honghu”, which should be aligned to the entity related to birds, and cannot be incorrectly mapped to unrelated characters or place names.

[0089] In the weighting process, the weight of each indicator can be automatically learned by large model training, or can be set according to empirical rules, such as increasing the weight of genre matching degree in poetry texts and increasing the weight of dynasty consistency in interpretation texts. The weighted score will comprehensively evaluate the initial evidence set, and select the optimal evidence set through the sorting mechanism. For example, for the sentence “Hongtu Gaofei”, the final evidence set may prefer to retain relevant poetic explanations in “Chu Ci·Jiuge” and the allusion items in “Zhuangzi·Xiaoyao You”, and filter out materials that do not conform to other genres or recent literature.

[0090] The above method, through the multi-dimensional screening mechanism, the final evidence set not only can ensure that it is highly relevant to the original text in terms of semantics, but also has strong consistency in terms of time, genre and entity, thereby providing reliable knowledge support for translation and annotation generation.

[0091] In one of the embodiments, the ancient Chinese knowledge graph is constructed by the following method:

[0092] S21, obtain ancient book corpus, annotation data, variant character mapping data, phonology data, and person place name official position data, and perform standardization processing to obtain an ancient text related knowledge index library.

[0093] Schematically, five types of information sources are utilized, namely ancient book corpora, annotation data, variant character mapping data, phonetic data, and data on people, place names, and official positions. Ancient book corpora refer to text materials from ancient classics, historical records, literary collections, and exegetical dictionaries. They often contain mixed use of traditional Chinese characters and variant characters, as well as differences in language habits of different dynasties. Therefore, text cleaning and unified encoding are required. Annotation data includes the explanations and commentaries of ancient scholars on classical ancient texts, usually sourced from electronic ancient book collation projects or large digital humanities databases. Variant character mapping data mainly addresses the problem of glyph divergence caused by different writings of the same morpheme, such as the correspondence between "爲" and "为". The mapping is completed by constructing a variant character comparison table and combining it with the "List of Chinese Character Standards". Phonetic data covers the results of ancient rhyme books and phonology research. It not only records the changing trajectory of pronunciation but also reveals the context of semantic evolution. Data on people, place names, and official positions mainly comes from historical geography databases and biographical materials of figures. After entity recognition and disambiguation processing, a standardized historical entity library is formed. In the standardization process, all data needs to be unified into the same character set, such as Unicode, and a unified time marking system, such as indexing by dynasty or specific years, and then word segmentation, duplicate removal, and polysemy splitting are carried out to form a clean and structured knowledge index library.

[0094] S22. Nodes and relationship edges are extracted according to the relevant knowledge index library of ancient texts; the nodes include characters, words, semantic items, constructions, allusions, citations, people, place names, official positions, dynasties, and annotation entries; the relationship edges include diachronic evolution relationships, cognate variant relationships, allusion relationships, usage example relationships, and rhetorical relationship.

[0095] Furthermore, nodes and relationship edges are extracted based on the knowledge index library. The design of the nodes covers key elements in the semantic system of ancient texts, including characters, words, semantic items, constructions, allusions, citations, people, place names, official positions, dynasties, and annotation entries. It not only ensures a refined representation at the language level but also connects language with culture and historical context. The setting of the relationship edges reflects the multi-dimensional connections of ancient text knowledge. Among them, the diachronic evolution relationship is used to describe the change of the meaning of a character or word over time; the cognate variant relationship is used to represent the consistency of different glyphs or word forms in terms of etymology; the allusion relationship is used to mark the connection between an allusion and its source text; the usage example relationship is used to connect specific citations with the semantic items they explain; the rhetorical relationship reveals the logical connection within the text in terms of rhetoric or text structure.

[0096] S23. A node vector representation is generated for the nodes, and an ancient Chinese knowledge graph is formed based on the nodes, the node vector representations corresponding to the node attributes, and the relationship edges.

[0097] Illustratively, the generation of node vector representation usually adopts a graph neural network or a context-based language model method, which incorporates the text description, attribute information of the node and its neighbor structure in the relationship network into modeling, so that the obtained node vector not only reflects the semantic attributes of a single entity, but also captures the position characteristics of the entity in the knowledge network. Illustratively, the vector of the "Confucius" node is not only from its biography description, but also is influenced by its relationship with nodes such as "Spring and Autumn Period", "State of Lu", "Analects", etc., so as to have a richer semantic representation. After obtaining the node vector, all nodes are coded and organized into a graph database in the form of ancient Chinese knowledge graph, which can support efficient retrieval and reasoning, and can also provide evidence-based support in subsequent semantic matching and translation inference.

[0098] In one embodiment, the context semantic representation and the evidence set are input into a translation and annotation dual-track decoder to generate a plurality of modern Chinese translation tracks and modern Chinese annotation tracks, including:

[0099] S31, based on the context semantic representation, generating modern Chinese translation tracks in different translation styles.

[0100] Based on the context semantic representation, translation tracks with different translation styles are generated. Illustratively, the translation style usually includes "faithful", "smooth" and "elegant", among which "faithful" emphasizes accurate conveying of the original meaning, "smooth" emphasizes the smoothness of the translation, and "elegant" pursues the beauty of the language and the improvement of the style. In order to realize the generation of diversified translation tracks, the training of the translation and annotation dual-track decoder often introduces style labels, which are annotated in the training corpus to enable the decoder to automatically favor the corresponding expression mode according to the combination of the context semantic representation and the style label during translation.

[0101] S32, according to the context semantic, performing word meaning and syntax segmentation to generate annotation tracks, and labeling the traceable identification corresponding to the evidence set in the annotation tracks to obtain modern Chinese annotation tracks.

[0102] The generation of annotation tracks depends on the combination of context semantic representation and evidence set. Specifically, the decoder performs sentence segmentation on the ancient text paragraph, and further divides the word meaning and syntax units based on the segmentation. The segmentation process can be based on rules, such as using common sentence reading marks or virtual word structure of ancient Chinese, or based on statistical models or deep learning methods, such as using syntax dependency analysis or sequence labeling model to identify sentence boundaries and word functions. After obtaining the segmented units, the decoder generates modern Chinese annotation tracks for each unit to explain the word meaning and syntax function.

[0103] Further, the annotation track will introduce the provenance identifier provided by the evidence set. The evidence set may contain allusion nodes, historical annotation items, or quotation fragments corresponding to the word or sentence. Embedding provenance information in the form of an identifier in the annotation track makes each explanation have a clear source.

[0104] Illustratively, the training process of the translation and annotation dual-track decoder will introduce artificially annotated historical annotations and knowledge graph evidence as training signals in addition to the conventional ancient and modern translation corpus, so that the model not only learns to generate translation, but also synchronously generates annotation tracks and evidence annotations. Specifically, the translation and annotation dual-track output can be presented as two columns of content, with the left side being the parallel modern Chinese translation and the right side being the sentence-by-sentence and word-by-word annotations and provenance. Illustratively, the joint loss function of the training of the translation and annotation dual-track decoder is wherein, is the joint loss function, which is the overall optimization goal, taking into account the translation track, the annotation track, and the training requirements of the alignment of the two. is the cross-entropy loss for the modern translation track, wherein Y trans is the distribution difference between the translation track generated by the model and the reference translation, which ensures that the model can output a semantically accurate and fluent modern Chinese translation. is the cross-entropy loss for the annotation track, multiplied by the weight coefficient a, wherein Y note is the distribution difference between the annotation track generated by the model and the reference annotation. a controls the importance of the annotation track training, avoiding excessive bias towards translation and ignoring annotations. is the alignment loss, multiplied by the weight coefficient b, to ensure the consistency of the translation track and the annotation track in terms of semantics and evidence set, for example, forcing the terms and explanations in the annotation track to be associated with the translation track. b controls the weight of the alignment constraint in the overall optimization.

[0105] In one embodiment, as shown in Figure 3 , the modern Chinese translation track and the modern Chinese annotation track are constrained and filtered according to the evidence set to obtain a candidate translation and annotation result that meets the evidence coverage requirement, including:

[0106] S301, generating a constraint lattice according to the evidence set; the constraint lattice contains constraint conditions that must be covered, can be covered, and must not be covered.

[0107] The constraint lattice is a multi-dimensional constraint structure, in which each evidence entry is mapped as a type of coverage condition. The core goal of the constraint lattice is to provide a followable boundary for the translation process, so that the translation is not only the result of semantic inference, but also the evidence-driven deduction. For example, when the evidence set contains the entry of “Shiji·Xiaoyao Yin” quotation, a “must cover” constraint condition is generated, that is, the candidate translation result must explicitly or implicitly mention the allusion or the corresponding semantic scene. On the contrary, if the evidence contains explicit negative content, such as some annotation has proved that the meaning of a word is irrelevant to the paragraph, a “forbidden cover” constraint condition is generated to avoid producing false trace pointing in the translation track. There is a type of evidence that does not affect the overall correctness but helps to enrich the explanation, such as the use case of the same period and the analogy expression in the parallel text, which is marked as “optional cover”. Whether it appears in the candidate result or not will not directly affect the validity of the result, but will affect the priority.

[0108] S302, based on the constraint lattice, decoding the modern Chinese translation track and the modern Chinese annotation track, and combining a consistency penalty term to remove the violation of translation and annotation, to obtain a plurality of candidate translation results; the consistency penalty term includes a penalty weight when the time and the relationship between the characters are contradictory.

[0109] Further, the decoder will synchronously check whether the output meets the coverage requirements of the constraint lattice when generating each word or phrase. If a translation selection omits the “must cover” evidence, its candidate translation track will be marked with a negative score; if the output contains “forbidden cover” components, the removal mechanism will be triggered directly. Optionally, a consistency penalty term is introduced to solve the potential logical conflict between the translation and the evidence. For example, when the translation track attributes an event to “Hanwu Emperor”, and the timestamp and the relationship between the character entities in the evidence set indicate that the event should be attributed to “Qinshihuang”, it is determined that the time and the relationship between the characters are contradictory, and the translation result will be given a higher penalty weight, resulting in a significant drop in the ranking of the result in the candidate set. The size of the penalty weight is usually determined according to the severity of the contradiction. If the contradiction involves key constraints such as time and dynasty, the penalty value is higher; if it is only a slight inconsistency in style or tone, the penalty value is lower. In this way, the decoder not only generates the translation in language form, but also ensures its consistency in historical facts, knowledge logic and literature evidence. For example, wherein, represents the candidate translation result set , and the selected one is the one that maximizes the value of the objective function Y. The candidate set is generated based on the translation and annotation dual-track decoder and is screened by preliminary constraints. log P(Y|X) is the conditional probability logarithm of the translation and annotation result Y given the input ancient text paragraph X. This item ensures the semantic matching degree of the translation and annotation result and the input ancient text, and can be understood as the confidence score of the language model itself. γ·Coverage(Y,ε) is the coverage constraint term, and γ is the adjustment weight. ε is the evidence set, that is, the explanation, use case or historical material basis in the knowledge graph corresponding to the ancient text paragraph. Coverage(Y,ε) is the coverage degree of the translation and annotation result Y to the evidence set, which is usually measured by how many evidence points are involved in the translation and annotation, the completeness or relevance of evidence utilization, which ensures that the final translation and annotation is not only semantically reasonable, but also has sufficient “source” or “support”. γ is a weight factor used to balance the importance of language model confidence and evidence coverage.

[0110] S303, screening the candidate translation and annotation result based on the evidence coverage rate to obtain a candidate translation and annotation result meeting the evidence coverage requirement.

[0111] Further, the candidate translation and annotation result after constraint and penalty processing is screened by coverage rate. The evidence coverage rate is an index for measuring the matching degree of the candidate translation and annotation result with the evidence set, which can be obtained by calculating the ratio of the number of evidence items involved in the candidate result to the total number of evidence. For example, when a candidate translation and annotation covers 6 out of 8 evidences, the coverage rate is 75%. Optionally, a coverage rate threshold is set, and only the candidate translation and annotation that meets or exceeds the threshold is retained. For candidate results with the same coverage rate, further sorting is performed in combination with additional indexes such as language fluency, style diversity and annotation clarity.

[0112] The above method controls the candidate translation and annotation result through evidence-driven constraint grid generation, multi-dimensional penalty mechanism and coverage rate screening, so that the final output candidate translation and annotation result is not only natural and fluent in language, but also reliable in knowledge logic and evidence consistency, thereby effectively solving the problems of arbitrariness, evidence loss and contradiction in traditional ancient text translation and annotation.

[0113] In one embodiment, the candidate translation and annotation result is rearranged according to the style, and the traceable identification in the annotation track is executable, and the translation result containing the modern translation and the annotation is output, including:

[0114] S41, scoring the candidate translation and annotation result according to the style to obtain a style scoring result. The style scoring includes calculation of parallelism, comparison, rhythm, rhyme and rhetorical retention.

[0115] The text style score is a comprehensive language evaluation index for measuring the quality of the translated text in terms of form and artistic feeling. Specifically, the text style score includes multiple dimensions of evaluation, including semantic fidelity S, text style fidelity Q, and readability R. Semantic fidelity S is used to measure the degree of semantic consistency between the translated text Y and the original ancient text input. It can be calculated by the similarity of context encoding vectors. Text style fidelity Q is used to measure the naturalness and expression quality of the translated text Y in modern Chinese, usually given by the perplexity or fluency score of the large model. Readability R is used to measure whether the syntax of the translated text Y is correct. Among them, the maintenance of parallelism and parallelism reflects the consistency of the translated text in sentence balance and inter-sentence structure echo; the rhythm and rhyme maintenance measures the coordination of the translated text in word and syllable and rhyme, especially for ancient poetry or verse translation; the rhetoric maintenance pays attention to whether the rhetorical devices such as metaphor, personification, contrast, and truth are preserved in the translated text. In order to realize the evaluation, a rule-based method can be used, such as statistical short sentence length, detection of sentence repetition, analysis of syllable pattern, and identification of rhetorical features by large models, to obtain the quantitative score of each candidate translation result in each dimension. The text style score result forms a vectorized index, which not only reflects the overall text style level, but also can distinguish the subtle differences between candidate translations.

[0116] S42, rearranging the candidate translation results according to the text style score results to obtain the text style optimized translation results.

[0117] The goal of the rearrangement process is to prioritize the translation with higher text style score, i.e. Y λ1S+λ2Q+λ3R, where Y is the candidate translation result, so that the final presented translation is not only semantically accurate, but also more natural and beautiful in terms of language feeling and rhetorical style. The specific implementation can include local adjustment of word order, optimization of inter-sentence conjunction words, adjustment of rhetorical structure and rhythm and rhythm. Through this process, multiple candidate translations can be sorted, and the translation with the highest comprehensive score is selected as the final output.

[0118] S43, corresponding the traceability identifier with the knowledge graph node, and rendering the annotation path template according to the preset user view, combining the translation result, to obtain the translation result containing the modern translation and the annotation.

[0119] The traceable identifier in the annotation track is executable processed to ensure that the correspondence between the annotation information and the knowledge graph node is clear and operable. Specifically, the traceable identifier of each annotation track is mapped to the node in the knowledge graph, and is presented in combination with the preset view of the user. The user view can include an academic research view to display detailed anecdotes and past annotations, or a teaching view to mainly display easy-to-read annotations and examples, or a general reader view to highlight word meaning interpretation and translation understanding, and different views will adjust the level, detail and presentation method of annotation display. According to the preset view, the annotation content is rendered into the final translation page or document through template rendering technology, such as displaying floating annotations, footnotes or side note items beside the text, while retaining the interaction function with the knowledge graph node, so that the user can click the traceable identifier to consult the original classic or related items. Through this processing, the final output translation result contains not only the modern Chinese translation track, but also the annotation track processed by visualization and interactivity, realizing high-quality fusion of translation and annotation.

[0120] In summary, through the optimization of text style scoring and rearrangement, and the executable processing of traceable identifiers, the technical solution realizes high-quality fusion of translation track and annotation track, so that the translation result achieves a balance in accuracy, readability and knowledge traceability, thereby significantly improving the overall practical value of the ancient text translation system.

[0121] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.

[0122] Based on the same inventive concept, the embodiments of the present application also provide a large model-based ancient text translation system for implementing the large model-based ancient text translation method described above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme described in the above method, so the specific limitations in one or more large model-based ancient text translation system embodiments provided below can refer to the limitations of the large model-based ancient text translation method described above, which will not be repeated here.

[0123] In one exemplary embodiment, as shown in Figure 4 a large model-based ancient text translation system is provided, comprising:

[0124] The standardization processing module 401 is configured to obtain an ancient text paragraph to be translated, and perform upper semantic standardization processing on the ancient text paragraph to obtain an ancient text processing result with structured information.

[0125] The semantic extraction module 402 is configured to perform context coding processing on the ancient text processing result, fuse glyph information and textual research information, and combine with diachronic semantic modeling to obtain a context semantic representation corresponding to the ancient text paragraph.

[0126] The evidence module 403 is configured to perform similarity retrieval on the context semantic representation in an ancient Chinese knowledge graph to obtain an evidence set corresponding to the ancient text paragraph.

[0127] The translation and annotation dual-track module 404 is configured to input the context semantic representation and the evidence set into a translation and annotation dual-track decoder to generate a plurality of modern Chinese translation tracks and modern Chinese annotation tracks.

[0128] The evidence constraint module 405 is configured to constrain and screen the modern Chinese translation tracks and the modern Chinese annotation tracks according to the evidence set to obtain a candidate translation and annotation result meeting an evidence coverage requirement.

[0129] The text atmosphere rearrangement module 406 is configured to perform text atmosphere rearrangement according to the candidate translation and annotation result, and perform executable processing on traceability identifiers in the annotation tracks to output a translation result containing a modern translation and annotations.

[0130] In one of the embodiments, the semantic extraction module 402 is further configured to:

[0131] extract stroke sequences and component structures of the words of the ancient text paragraph according to the word-level information, and encode the stroke sequences and the component structures to obtain glyph vectors;

[0132] encode the ancient commentaries and example texts corresponding to the words of the ancient text paragraph according to the word-level information to obtain meaning item vectors;

[0133] fuse the glyph vectors and the meaning item vectors to obtain a corresponding fusion representation;

[0134] obtain a dynasty time mark vector according to the dynasty mark, and perform diachronic semantic modeling on the fusion representation according to the dynasty time mark vector to obtain a diachronic enhanced representation;

[0135] perform joint modeling on the diachronic enhanced representation and the construction pattern recognition result to obtain the context semantic representation.

[0136] In one of the embodiments, the evidence module 403 is further configured to:

[0137] perform similarity retrieval on the context semantic representation and the node vector representation based on the ancient Chinese knowledge graph to obtain a candidate node.

[0138] According to the candidate node, allusion nodes, annotation items and quotation fragments corresponding to the ancient text paragraph are obtained to form an initial evidence set;

[0139] The evidences in the initial evidence set are weighted according to the similarity score, the dynasty time consistency, the style matching degree and the entity alignment relationship, and the evidences are sorted according to the weighted results to obtain an evidence set corresponding to the ancient text paragraph.

[0140] In one of the embodiments, the knowledge graph construction module is further used for:

[0141] The ancient book corpus, the annotation data, the variant character mapping data, the phonetic data and the person place official data are obtained and standardized to obtain an ancient text related knowledge index library;

[0142] According to the ancient text related knowledge index library, nodes and relationship edges are extracted; the nodes include characters, words, semantic items, constructions, allusions, quotations, persons, places, official positions, dynasties and annotation items; the relationship edges include diachronic evolution relationship, homonym relationship, allusion and quotation relationship, use case relationship and chapter relationship;

[0143] Node vector representations are generated for the nodes, and an ancient Chinese knowledge graph is formed according to the nodes, the node vector representations of corresponding node attributes and the relationship edges.

[0144] In one of the embodiments, the translation and annotation dual-track module 404 is further used for:

[0145] Based on the context semantic representation, a modern Chinese translation track is generated according to different translation styles;

[0146] According to the context semantics, a word meaning and syntax segmentation is performed to generate an annotation track, and a traceability identifier corresponding to the evidence set is marked in the annotation track to obtain a modern Chinese annotation track.

[0147] In one of the embodiments, the evidence constraint module 405 is further used for:

[0148] A constraint lattice is generated according to the evidence set; the constraint lattice contains constraint conditions of must cover, optional cover and prohibited cover;

[0149] Based on the constraint lattice, the modern Chinese translation track and the modern Chinese annotation track are decoded, and a violation of translation and annotation is removed in combination with a consistency penalty term to obtain a plurality of candidate translation and annotation results; the consistency penalty term includes a penalty weight when the time and the person relationship between the translation and the constraint lattice are contradictory;

[0150] The candidate translation and annotation results are screened based on the evidence coverage rate to obtain a candidate translation and annotation result that meets the evidence coverage requirement.

[0151] In one of the embodiments, the text atmosphere rearrangement module 406 is further configured to:

[0152] perform text atmosphere scoring on the candidate translation result to obtain a text atmosphere scoring result, wherein the text atmosphere scoring includes calculation of parallelism, anaphora, rhythm, rhyme and rhetoric retention;

[0153] rearrange the candidate translation result according to the text atmosphere scoring result to obtain a text atmosphere-optimized translation result;

[0154] correspond the provenance identifier with the knowledge graph node, and perform annotation path template rendering according to a preset user view to obtain a translation result including a modern translation and annotations.

[0155] In one embodiment, a computer device is provided, including a memory and a processor, the memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0156] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.

[0157] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts are described in the method embodiments. The above described device embodiments are only illustrative, and the components described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present disclosure according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0158] The above described embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the patent scope of the application. It should be noted that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application.

Claims

1. A method for translating classical Chinese texts based on a large model, characterized in that, The method includes: The ancient Chinese text to be translated is obtained, and the ancient Chinese text is subjected to higher-level semantic standardization processing to obtain the ancient Chinese text processing result with structured information; the ancient Chinese text processing result includes character-level information, structure recognition result and dynasty mark; the higher-level semantic standardization processing corresponds to sentence segmentation, structure recognition and dynasty mark; Based on the ancient text processing results, context encoding is performed to fuse glyph information and exegesis information, and combined with diachronic semantic modeling to obtain the context semantic representation of the corresponding ancient text paragraph; Based on the contextual semantic representation, a similarity search is performed in the ancient Chinese knowledge graph to obtain a set of evidence corresponding to the ancient text passage; The contextual semantic representation and the evidence set are input into the translation and annotation dual-track decoder to generate multiple modern Chinese translation tracks and modern Chinese annotation tracks. Based on the evidence set, the modern Chinese translation track and the modern Chinese annotation track are constrained and screened to obtain candidate translation and annotation results that meet the evidence coverage requirements; Based on the candidate translation and annotation results, the text is rearranged in terms of style and tone, and the source identifiers in the annotation track are made executable, outputting a translation result that includes the modern translation and annotations.

2. The method according to claim 1, characterized in that, The step of performing context encoding processing based on the ancient text processing results, fusing glyphic information and exegesis information, and combining it with diachronic semantic modeling to obtain the contextual semantic representation of the corresponding ancient text paragraph includes: Based on the character-level information, the stroke sequence and component structure of the characters in the ancient text paragraph are extracted, and the stroke sequence and component structure are encoded to obtain character shape vectors; The corresponding historical commentaries and illustrative texts of the ancient text paragraphs are encoded based on the character-level information to obtain semantic vectors; The glyph vector and the semantic vector are fused to obtain the corresponding fused representation; Based on the dynasty marker, a dynasty time-stamped vector is obtained, and based on the dynasty time-stamped vector, a diachronic semantic model is performed on the fused representation to obtain a diachronic enhanced representation; The contextual semantic representation is obtained by jointly modeling the diachronic augmentation representation and the configuration recognition result.

3. The method according to claim 2, characterized in that, The step of performing a similarity search in the Classical Chinese knowledge graph based on the contextual semantic representation to obtain a set of evidence corresponding to the Classical Chinese passage includes: Based on the ancient Chinese knowledge graph, a similarity search is performed on the context semantic representation and the node vector representation to obtain candidate nodes; Based on the candidate nodes, obtain the allusion nodes, annotation entries and quotations corresponding to the ancient text paragraphs to form an initial set of evidence; The evidence in the initial evidence set is weighted according to similarity score, consistency of dynasty timescale, genre matching degree and entity alignment relationship, and sorted according to the weighting results to obtain the evidence set corresponding to the ancient text passage.

4. The method according to claim 3, characterized in that, The ancient Chinese knowledge graph was constructed using the following methods: We acquire ancient text corpora, annotation data, variant character mapping data, phonological data, and data on people, places, and official titles, and perform standardization processing to obtain an index database of knowledge related to ancient texts. Nodes and relational edges are extracted from the ancient Chinese knowledge index database. The nodes include characters, words, meanings, structures, allusions, quotations, figures, place names, official titles, dynasties, and annotation entries. The relational edges include diachronic evolution relationships, cognate relationships, classical reference relationships, example relationship relationships, and compositional relationships. Generate the node vector representation for the node, and form the Ancient Chinese knowledge graph based on the node, the node vector representation of the corresponding node attribute, and the relation edges.

5. The method according to claim 1, characterized in that, The process involves inputting the contextual semantic representation and the evidence set into a dual-track decoder to generate multiple modern Chinese translation tracks and modern Chinese annotation tracks, including: Based on the aforementioned contextual semantic representation, modern Chinese translation tracks are generated according to different translation styles; Based on the contextual semantics, the sentences are segmented into semantic and syntactic parts to generate annotation tracks, and the source identification corresponding to the evidence set is marked in the annotation tracks to obtain the modern Chinese annotation tracks.

6. The method according to claim 1, characterized in that, The constraint screening of the modern Chinese translation track and the modern Chinese annotation track based on the evidence set to obtain candidate translation and annotation results that meet the evidence coverage requirements includes: A constraint grid is generated based on the evidence set; the constraint grid contains constraints that must be covered, can be covered, and cannot be covered. Based on the constraint grid, the modern Chinese translation track and the modern Chinese annotation track are decoded, and the translation annotations that violate the consistency penalty are eliminated to obtain multiple candidate translation annotation results; the consistency penalty includes the penalty weight when the translation annotations contradict the time and character relationships between them. The candidate translation annotations are filtered based on the evidence coverage rate to obtain candidate translation annotations that meet the evidence coverage requirements.

7. The method according to claim 1, characterized in that, The process of rearranging the text based on the candidate translation and annotation results, and making the source identification in the annotation track executable, outputs a translation result containing the modern translation and annotations, including: The candidate translation and annotation results are scored for literary style to obtain the literary style score; the literary style score includes the calculation of parallelism, repetition, rhythm, phonetics and rhetoric preservation. The candidate translation and annotation results are rearranged based on the textual tone scoring results to obtain the textual tone optimized translation and annotation results; The source identifier is mapped to a knowledge graph node, and the annotation path template is rendered according to the preset user view. The translation result containing the modern translation and annotation is obtained by combining the translation and annotation results.

8. A classical Chinese translation system based on a large model, characterized in that, The system includes: The standardization processing module is used to obtain the ancient Chinese text paragraphs to be translated and to perform higher-level semantic standardization processing on the ancient Chinese text paragraphs to obtain ancient Chinese text processing results with structured information. The semantic extraction module is used to perform context encoding processing based on the ancient text processing results, integrate glyph information and exegesis information, and combine it with diachronic semantic modeling to obtain the context semantic representation of the corresponding ancient text paragraph. The evidence module is used to perform similarity retrieval in the ancient Chinese knowledge graph based on the contextual semantic representation to obtain a set of evidence corresponding to the ancient text passage; The translation and annotation dual-track module is used to input the contextual semantic representation and the evidence set into the translation and annotation dual-track decoder to generate multiple modern Chinese translation tracks and modern Chinese annotation tracks. The evidence constraint module is used to constrain and filter the modern Chinese translation track and the modern Chinese annotation track according to the evidence set, so as to obtain candidate translation and annotation results that meet the evidence coverage requirements. The text rearrangement module is used to rearrange the text based on the candidate translation and annotation results, and to perform executable processing on the source identification in the annotation track, and output the translation result containing the modern translation and annotation.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Ancient Chinese understanding method based on large model

    CN119378561A