Construction method of Chinese-English corpus based on qualification identification
By building the Zizhi Tongjian Chinese-English corpus, the gap in the historical classical Chinese-English translation query platform has been solved, convenient English translation query and research support has been achieved, and the dissemination of historical culture has been promoted.
Patent Information
- Application Number
- CN202510316700.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the construction of bilingual English-Chinese corpus, the lack of parallel corpus for classical Chinese in history books has led to inconvenience in the English translation query platform of "Zizhi Tongjian" and cannot meet the needs of domestic and foreign enthusiasts and researchers.
Build a Chinese-English corpus based on Zizhitongjian, including corpus collection, processing and construction, using Tmxmall for sentence-level alignment, using Corpus Word Parser and Stanford POS Tagger for word segmentation and marking, development interface and back-end security guarantee, providing advanced query tools and continuous maintenance.
It fills the gap in the historical classical Chinese corpus, provides a convenient English translation query platform, promotes the research on English translation of "Zizhi Tongjian", enriches learning and discussion scenarios, and promotes the dissemination of traditional Chinese historiography culture.
Smart Images

Figure CN120258008A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software technology, and specifically to a construction method of a Chinese-English corpus based on Zizhi Tongjian. Background Art
[0002] Since the 21st century, corpus-based translation research at home and abroad has gradually emerged, and specialized corpora covering various fields have been created. On this basis, corpora have gradually been divided into three categories: translation corpora, analogy corpora, and bilingual parallel corpora. Theory and practice have proved that corpus-based related research has had a great impact on the field of translation research and produced rich and positive research results.
[0003] Parallel or corresponding corpora (parallel corpora) are bilingual or multilingual corpora composed of original text and its parallel corresponding translation text, and their alignment levels are divided into several types such as word level, sentence level, paragraph level, and text level. The construction of bilingual parallel corpora has emerged in combination with computer technology, providing a platform for language research, translation research, foreign language teaching, etc., and having broad prospects.
[0004] The development of bilingual parallel corpora in China has also been relatively rapid in the past decade. Chinese corpora mainly start from two types: (non-)literary texts and a certain specific type of quasi-texts. For example, the parallel corpus of the translation of A Dream of Red Mansions by Yanshan University with literature as the theme, the English-Chinese parallel corpus of Shakespeare's plays by Shanghai Jiao Tong University, or the parallel corpus of model finance English-Japanese-Chinese by Fu Jen Catholic University in Taiwan Province with a specific type as the theme, etc.
[0005] In the field of the construction of English-Chinese bilingual corpora today, there are almost no English-Chinese parallel corpora of historical books in classical Chinese, and there are also few popular, authoritative and convenient English translation query platforms for Zizhi Tongjian on the market, which causes many inconveniences to domestic and foreign enthusiasts and researchers. Summary of the Invention
[0006] The purpose of the present invention is to provide a construction method of a Chinese-English corpus based on Zizhi Tongjian, so as to solve the problem proposed in the above background art that in the field of the construction of English-Chinese bilingual corpora today, there are almost no English-Chinese parallel corpora of historical books in classical Chinese, and there are also few popular, authoritative and convenient English translation query platforms for Zizhi Tongjian on the market, which causes many inconveniences to domestic and foreign enthusiasts and researchers.
[0007] To achieve the above purpose, the present invention provides the following technical solution: A construction method of a Chinese-English corpus based on Zizhi Tongjian, including the following steps:
[0008] S1: Construct a corpus, collect corpus, process the corpus and build the corpus;
[0009] S2: Conduct website design and development, the content of which includes interface design, front-end development, back-end development, and security assurance;
[0010] S3: Conduct platform promotion and maintenance, and carry out user support and promotion plans during the continuous maintenance process.
[0011] Preferably, in the S1 step, the collection of corpus refers to integrating different translations of "Zizhi Tongjian" and "Records of the Grand Historian" literature, obtaining authorization from cooperative publishing houses and academic institutions, and providing content support for translation queries.
[0012] Preferably, in the S1 step, the method of corpus processing is to sort out the collected corpus, remove pictures and mis-translated content, and save it in TXT format first, which is convenient for segmentation, annotation and marking, and alignment procedures.
[0013] Preferably, in the S1 step, the construction of the corpus uses Tmxmall for online left-right sentence-level alignment, with the Chinese source text on the left and the corresponding English translation on the right; uses the corpus word segmentation software Corpus Word Parser to segment and mark the part-of-speech of the corpus information; uses the corpus part-of-speech tagging software Stanford POS Tagger for coding, and all processes are manually proofread twice to ensure the accuracy and objectivity of the data and the reasonable and error-free construction of the corpus.
[0014] Preferably, in the S2 step, the interface design includes a home page, a search function, a personal note area, and a forum module, ensuring that the design is responsive and adaptable to access from different devices.
[0015] Preferably, in the S2 step, the front-end development uses modern front-end technologies React or Vue.js to build the front-end part of the website and achieve dynamic interaction of the user interface.
[0016] Preferably, in the S2 step, the back-end development builds a stable server-side, uses back-end technologies such as Node.js or Python to process data requests, user management, and content updates, and the security assurance ensures the data security of the website and the protection of user privacy, and implements SSL encryption, data backup and recovery mechanisms.
[0017] Preferably, the promotion plan publicizes the website through academic conferences, cooperative institutions, and social media channels to attract the initial user group. User support: Set up online help and customer support, answer users' questions, collect user feedback, regularly update the FAQ, continuously maintain, monitor the running status of the website, and regularly perform technical upgrades and content updates to ensure the long-term stable operation of the website.
[0018] Compared with the prior art, the beneficial effects of the present invention are:
[0019] 1. Fill the gap in the construction of the classical Chinese corpus of historical books in the English-Chinese parallel corpus, providing convenience for domestic and foreign enthusiasts and researchers.
[0020] 2. Boost the research on the English translation of "Zizhi Tongjian" through the construction of the corpus, reducing the workload of discourse analysis of classical Chinese texts in the field of history from more dimensions.
[0021] 3. Contribute to enriching the learning and discussion scenarios of classical Chinese in historical books at home and abroad, and promoting the dissemination of traditional Chinese historical culture to the public. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flowchart of the present invention;
[0023] Figure 2 is a schematic diagram of the corpus word segmentation technology of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0026] Please refer to Figure 1-2 , the present invention provides a construction method based on the Chinese-English corpus of "Zizhi Tongjian", including the following steps:
[0027] S1: Construct a corpus, and conduct corpus collection, corpus processing and corpus construction;
[0028] a. Multi-source text integration: Collect and integrate various English translations of classic historical texts such as "Zizhi Tongjian" and "Records of the Grand Historian", including but not limited to academic editions, popular editions and expert annotation editions, to enrich the content and perspective of the corpus.
[0029] b. Version difference analysis: Conduct a systematic analysis of the differences between different translation versions, clarify the translation deviations and differences in translation strategies between versions, and provide more comprehensive research materials for people's learning and research.
[0030] c. Corpus metadata management: Create detailed metadata records for each text and its translation, including translator information, publication year, translation style and target audience, etc., to facilitate users to screen and compare according to different research needs.
[0031] d. Dynamic update and expansion: Establish a reasonable operation and inspection mechanism, regularly update and expand the corpus, add new translation works and revised versions, and ensure the timeliness and forefront of the corpus content.
[0032] (2) Advanced query tool
[0033] a. Develop an advanced search function that allows users to perform precise queries through various filtering conditions such as keywords, authors, text types, or translation years.
[0034] b. We divide the non - heritage Chinese - English parallel corpus of "Zizhi Tongjian" into sixteen major sub - corpora according to the recorded dynasty categories, namely "Chronicle of Zhou", "Chronicle of Qin", "Chronicle of Han", "Chronicle of Wei", "Chronicle of Jin", "Chronicle of Song", "Chronicle of Qi", "Chronicle of Liang", "Chronicle of Chen", "Chronicle of Sui", "Chronicle of Tang", "Chronicle of Later Liang", "Chronicle of Later Tang", "Chronicle of Later Jin", "Chronicle of Later Han", and "Chronicle of Later Zhou". Designing a bilingual corpus by sub - corpora is conducive to classifying and studying the language characteristics, rhetoric, and Chinese - English language comparison of the "Zizhi Tongjian" corpus; it can also conduct detailed and in - depth analysis and research on a certain type of corpus, propose targeted translation strategies and methods, which is beneficial to improving the translation quality of such non - heritage texts.
[0035] S2: Conduct website design and development, the content of which includes interface design, front - end development, back - end development, and security guarantee;
[0036] S3: Conduct platform promotion and maintenance, and carry out user support and promotion plans during the continuous maintenance process.
[0037] Furthermore, the collection of corpus in step S1 refers to integrating different translations of the "Zizhi Tongjian" and "Records of the Grand Historian" literature, obtaining authorization from cooperative publishing houses and academic institutions to provide content support for translation queries.
[0038] Furthermore, the method of corpus processing in step S1 is to sort out the collected corpus, remove pictures and mis - translated content, and save it in TXT format first, which is convenient for segmentation, annotation and marking, and alignment procedures.
[0039] Furthermore, for the construction of the corpus in step S1, Tmxmall is used for online left - right sentence - level alignment, with the Chinese source text on the left and the corresponding English translation on the right; the corpus word - segmentation software Corpus Word Parser is used to segment and mark the part - of - speech of the corpus information; the corpus part - of - speech tagging software Stanford POS Tagger is used for coding, and the entire process is manually proofread twice to ensure the accuracy and objectivity of the data and the reasonable and error - free construction of the corpus.
[0040] Furthermore, the interface design in step S2 includes a home page, a search function, a personal note area, and a forum module, ensuring that the design is responsive and adaptable to access from different devices.
[0041] Furthermore, in step S2, modern front - end technologies such as React or Vue.js are used for front - end development of the website to achieve dynamic interaction of the user interface.
[0042] Further, in step S2, a stable server - side is built in back - end development. Back - end technologies such as Node.js or Python are used to process data requests, user management, and content updates. Security guarantees ensure the data security of the website and the protection of user privacy, implementing SSL encryption, data backup, and recovery mechanisms.
[0043] Further, the promotion plan publicizes the website through academic conferences, cooperative institutions, and social media channels to attract an initial user group. User support: Set up online help and customer support to answer users' questions, collect user feedback, regularly update the FAQ, and continuously maintain. Monitor the running status of the website, and conduct regular technical upgrades and content updates to ensure the long - term stable operation of the website.
[0044] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A construction method based on the Chinese-English corpus of Zizhi Tongjian, characterized in that: It includes the following steps: S1: Build a corpus, collect corpus, process the corpus, and build the corpus; S2: Conduct website design and development. The content of website design and development includes interface design, front-end development, back-end development, and security guarantee; S3: Conduct platform promotion and maintenance, and carry out user support and promotion plans during continuous maintenance.
2. The construction method of a Chinese-English corpus based on Zizhi Tongjian according to claim 1, characterized in that: In step S1, the collection of the corpus refers to integrating different translations of the historical records "Zizhi Tongjian" and "Records of the Grand Historian", obtaining authorization from cooperative publishing houses and academic institutions, and providing content support for translation queries.
3. A construction method of a Chinese-English corpus based on Zizhi Tongjian, as described in claim 1, wherein: The method of corpus processing in step S1 is to sort out the collected corpus, remove pictures and mis-translated content, and save it in TXT format first, which is convenient for segmentation, annotation and marking, and alignment procedures.
4. A construction method of a Chinese-English corpus based on Zizhi Tongjian, characterized in that: In step S1, the construction of the corpus uses Tmxmall for online left-right sentence-level alignment, with the Chinese source text on the left and the corresponding English translation on the right; uses the corpus word segmentation software Corpus Word Parser to segment and mark the part-of-speech of the corpus information; uses the corpus part-of-speech tagging software Stanford POS Tagger for coding. All processes are manually proofread twice to ensure the accuracy and objectivity of the data and the reasonable and error-free construction of the corpus.
5. A construction method of a Chinese-English corpus based on Zizhi Tongjian, characterized in that: In step S2, the interface design includes a home page, a search function, a personal note area, and a forum module, ensuring that the design is responsive and adaptable to access from different devices.
6. A construction method of a Chinese-English corpus based on Zizhi Tongjian, characterized in that: In step S2, the front-end development uses modern front-end technologies such as React or Vue.js to build the front-end part of the website and achieve dynamic interaction of the user interface.
7. A construction method of a Chinese-English corpus based on Zizhi Tongjian, characterized in that: In step S2, the back-end development builds a stable server-side, uses back-end technologies such as Node.js or Python to process data requests, user management, and content updates. The security guarantee ensures the data security of the website and the protection of user privacy, and implements SSL encryption, data backup, and recovery mechanisms.
8. A construction method of a Chinese-English corpus based on Zizhi Tongjian, characterized in that: The promotion plan promotes the website through academic conferences, cooperative institutions, and social media channels to attract the initial user group. User support: Set up online help and customer support, answer users' questions, collect user feedback, regularly update the FAQ, continuously maintain, monitor the running status of the website, and regularly perform technical upgrades and content updates to ensure the long-term stable operation of the website.