Construction method and device for historical word meaning corpus of ancient Chinese and storage medium

By collecting and annotating ancient Chinese corpora, and utilizing large language models and dictionary alignment algorithms, a diachronic semantic corpus of ancient Chinese was constructed. This solved the problem of ancient Chinese semantic research in existing technologies and achieved efficient semantic corpus retrieval and analysis functions.

CN122045324APending Publication Date: 2026-05-15BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING NORMAL UNIVERSITY
Filing Date
2025-12-31
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies lack methods for constructing large-scale diachronic semantic corpora of Classical Chinese. Existing corpora mostly focus on Modern Chinese, with relatively little research on semantic annotation of Classical Chinese, making it difficult to study the semantic evolution of Classical Chinese vocabulary.

Method used

We collect classical Chinese corpora, use a pre-trained classical Chinese language model to automatically annotate word meanings, and align them with the meanings in the "Chinese Dictionary". Through word meaning alignment and clustering, we form a multi-level language resource library, enabling meaning-level corpus retrieval and visualization.

Benefits of technology

A high-precision diachronic semantic corpus of Classical Chinese was constructed, which improved the efficiency of Classical Chinese semantic research, reduced the difficulty of teaching, and provided a multi-level language resource database and intelligent lexical semantic analysis function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045324A_ABST
    Figure CN122045324A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method and device of an ancient Chinese calendar word meaning corpus and a storage medium. According to the method, firstly, ancient Chinese corpora containing time information are collected, then, a pre-trained language model is used for carrying out word meaning automatic labeling on words in the preprocessed ancient Chinese corpora, and word meaning alignment is carried out on the words subjected to automatic labeling. And finally, summarizing annotation alignment results, and carrying out statistics on word meaning duration proportion information to obtain an ancient Chinese duration word meaning corpus. According to the historical word meaning corpus of the ancient Chinese constructed by the method, functions of semantic item-level corpus retrieval, word semantic evolution analysis, word meaning clustering analysis and the like can be realized, so that the rule and the mechanism of the semantic change of the Chinese vocabularies can be disclosed, and important data resources are provided for the fields of classical Chinese teaching, dictionary compilation and related humanity research; and the teaching and research efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and storage medium for constructing a diachronic semantic corpus of classical Chinese. Background Technology

[0002] Lexical semantic evolution is an important dimension of language evolution, reflecting the development of human society and the changes in civilization. The semantic system of Chinese vocabulary has been passed down for thousands of years, and its processes and mechanisms of change are particularly complex. This is not only a key and difficult issue in the study of Chinese language history, but also brings many challenges to ancient Chinese information processing, language teaching, lexicography, and humanities research. Therefore, the construction of a large-scale semantic evolution resource database is of great significance.

[0003] However, current technologies for constructing large-scale corpora primarily focus on modern Chinese, with only a small number covering classical Chinese. Semantic annotation research also centers on modern Chinese and fails to reflect the evolution of classical Chinese vocabulary meanings over time. Therefore, constructing a large-scale diachronic corpus of classical Chinese semantics is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] The embodiments of this disclosure provide a method, apparatus, and storage medium for constructing a diachronic semantic corpus of Classical Chinese, so as to at least solve the technical problem of how to construct a large-scale diachronic semantic corpus of Classical Chinese in the prior art.

[0005] According to one aspect of the present disclosure, a method for constructing a classical Chinese diachronic semantic corpus is provided, comprising: collecting classical Chinese corpus containing time information, and preprocessing the classical Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The words after automatic annotation are aligned with their meanings, and then the aligned words are clustered according to their meanings. The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the ancient Chinese historical semantic corpus.

[0006] According to another aspect of the present disclosure, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0007] According to another aspect of the present disclosure, an apparatus for constructing a classical Chinese diachronic semantic corpus is also provided, comprising: a preprocessing module for collecting classical Chinese corpus containing time information and preprocessing the classical Chinese corpus; The annotation module is used to automatically annotate the meanings of words in the pre-processed Classical Chinese corpus using a pre-trained language model. The alignment and clustering module is used to align the meanings of automatically labeled words and then perform word meaning clustering on the aligned words; and The sorting module is used to summarize the results of word meaning alignment and word meaning clustering, and to statistically analyze the historical proportion of word meanings to obtain the ancient Chinese historical word meaning corpus.

[0008] According to another aspect of the present disclosure, an apparatus for constructing a classical Chinese diachronic semantic corpus is also provided, comprising: a processor; and A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: Collect ancient Chinese corpus containing time information, and preprocess the ancient Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The words after automatic annotation are aligned with their meanings, and then the aligned words are clustered according to their meanings. The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the ancient Chinese historical semantic corpus.

[0009] In this embodiment, firstly, ancient Chinese corpus containing time information is collected. Then, considering the characteristics of ancient Chinese texts, a high-precision semantic annotation and alignment algorithm is developed using large language models and vector representation technology. This algorithm automatically annotates all texts in the database and aligns them with the definitions in the *Comprehensive Dictionary of Chinese Characters*. Finally, the definitions are automatically clustered, forming a multi-level language resource database of "concept-word form-word meaning-corpus." This enables definition-level corpus retrieval, visualization of the diachronic evolution of semantic terms, semantic clustering queries, and concept name evolution analysis, improving research efficiency and reducing teaching difficulty. Attached Figure Description

[0010] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure. In the drawings: Figure 1 This is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of this disclosure; Figure 2 This is a schematic diagram of a classical Chinese diachronic semantic corpus system according to an embodiment of the present disclosure; Figure 3 This is a flowchart illustrating the method for constructing a classical Chinese diachronic semantic corpus according to an embodiment of this disclosure; Figure 4 is a schematic flowchart of a semantic alignment method according to an embodiment of the present disclosure; Figure 5 is an exemplary schematic diagram of a semantic alignment method according to an embodiment of the present disclosure; Figure 6 is a schematic flowchart of a process for batch finding of archaic and modern synonyms in a corpus according to an embodiment of the present disclosure; Figure 7 is an exemplary schematic diagram of the functions of a corpus system according to an embodiment of the present disclosure; Figure 8 is a schematic diagram of the statistical usage cases of each semantic item of the character "哥" in historical documents according to an embodiment of the present disclosure; Figure 9 is a schematic diagram of the statistical usage cases of each semantic item of the character "虫" in historical documents according to an embodiment of the present disclosure; Figure 10 is a schematic diagram of the statistical usage cases of each semantic item of the character "荷" in historical documents according to an embodiment of the present disclosure; Figure 11 is a schematic diagram of the statistical usage cases of each semantic item of the character "景" in historical documents according to an embodiment of the present disclosure; Figure 12 is a schematic diagram of the statistical usage cases of each member of the concepts related to "China" in historical documents according to an embodiment of the present disclosure; Figure 13 is a schematic diagram of a construction device for an archaic Chinese diachronic semantic corpus according to an embodiment of the present disclosure; and Figure 14 is a schematic diagram of another construction device for an archaic Chinese diachronic semantic corpus according to an embodiment of the present disclosure. Detailed implementation manners

[0011] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0012] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0013] The diachronic complexity of the Chinese semantic system not only makes it difficult for modern people to accurately learn and understand word meanings, but also poses a significant challenge to machine-based semantic analysis. Therefore, constructing a large-scale semantic evolution resource database is of great importance. In recent years, the international academic community has actively constructed such databases. However, in current technology, on the one hand, existing large-scale corpus constructions mostly focus on modern Chinese, with only a small number covering classical Chinese. Annotating large corpora is an extremely challenging task; therefore, most corpora primarily include unannotated raw data, and even when annotation is conducted, it is mainly at the level of word segmentation and part-of-speech classification. On the other hand, semantic annotation research also focuses on modern Chinese, with relatively little research in the field of classical Chinese. Therefore, the construction of a large-scale diachronic semantic corpus resource for classical Chinese remains an urgent issue.

[0014] To address the aforementioned problems in the prior art, this invention provides a method, apparatus, and storage medium for constructing a diachronic semantic corpus of Classical Chinese. First, approximately 160 million characters of Classical Chinese text from the pre-Qin period to the Ming and Qing dynasties are collected and organized, supplemented with nearly 30 million additional characters from different periods, totaling nearly 200 million characters. Then, considering the characteristics of Classical Chinese texts, a high-precision semantic annotation and alignment algorithm is developed using large language models and vector representation technology to automatically annotate all texts in the corpus and align them with the definitions in the *Comprehensive Dictionary of Chinese Characters*. Finally, the definitions are automatically clustered, forming a multi-level language resource corpus comprising "concept-word form-word meaning-corpus".

[0015] It should be noted that in this invention, the semantic terms and the semantic terms of words have the same meaning and can be used interchangeably.

[0016] It should be noted that, in this invention, the duration percentage information includes at least one of the following percentage information: 1. The ratio of the frequency of different meanings of the same word to the total frequency of the word in different historical periods; and 2. The ratio of the frequency of different words with the same meaning to the frequency of all words with the same meaning in different historical periods.

[0017] Example 1 According to this embodiment, an embodiment of a method for constructing a classical Chinese diachronic semantic corpus is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0018] The method embodiments provided in this example can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Figure 1 A hardware block diagram of a computing device for implementing a method for constructing a diachronic semantic corpus of Classical Chinese is shown. Figure 1 As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include a display, keyboard, and cursor control device connected to the input / output interface. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0019] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element in a computing device. As involved in the embodiments of this disclosure, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0020] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for constructing a classical Chinese diachronic semantic corpus in this embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the above-mentioned method for constructing a classical Chinese diachronic semantic corpus. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computing device via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0021] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computing device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0022] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of the computing device.

[0023] It should be noted here that, in some optional embodiments, the above... Figure 1 The computing device shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computing devices.

[0024] Figure 2 This is a schematic diagram of the Classical Chinese diachronic semantic corpus system described in this embodiment. (Refer to...) Figure 2As shown, the system includes: a user terminal 100 for user interaction, where users can input commands to use the corpus system and receive results returned by the system. For example, users can input commands for semantic-level corpus retrieval, lexical semantic evolution analysis, or semantic clustering analysis on the user terminal 100, which are transmitted to the computing center 200 via a network. The computing center 200 includes a computing server running the ancient Chinese diachronic semantic corpus system of this invention, as well as large-scale modeling systems (LLMs) such as the Ancient Chinese Large Language Model and the Ancient Chinese Semantic Vector Model. The computing center 200 is connected to a database 300 via a network, which stores databases related to the ancient Chinese diachronic semantic corpus, such as ancient Chinese corpora containing time information, all entries in the *Comprehensive Dictionary of Chinese Characters*, and intermediate processing results during the operation of the ancient Chinese diachronic semantic corpus system. Figure 2 The ancient Chinese diachronic semantic corpus system shown can realize all the functions and steps of this invention.

[0025] Under the aforementioned operating environment, according to the first aspect of this embodiment, a method for constructing a classical Chinese diachronic semantic corpus is provided. This method comprises... Figure 1 The hardware structure shown is implemented, or is made by Figure 2 The system implementation shown. Figure 3 A flowchart illustrating the method is shown below. (Refer to...) Figure 3 As shown, the method includes: S302. Collect ancient Chinese corpus containing time information and preprocess the ancient Chinese corpus; S304. Use a pre-trained language model to automatically annotate the meanings of words in the pre-processed Classical Chinese corpus. S306. Perform word sense alignment on the automatically labeled words, and then perform word sense clustering on the word sense aligned words; S308. Summarize the results of word meaning alignment and word meaning clustering, and statistically analyze the historical proportion of word meanings to obtain the ancient Chinese historical word meaning corpus.

[0026] Optionally, in S302, the collection of ancient Chinese corpus containing time information involves preprocessing the ancient Chinese corpus. To construct a diachronic semantic corpus of ancient Chinese, a large-scale ancient Chinese corpus covering different historical periods and literary styles was collected. For example, approximately 160 million characters of ancient Chinese texts from the pre-Qin period to the Ming and Qing dynasties were collected and organized, and nearly 30 million characters of period texts were added, totaling nearly 200 million characters.

[0027] Optionally, the preprocessing includes one or a combination of the following: Remove duplicate data; Replace special characters; Standardize the punctuation marks in the text; Mark the corpus with dynasty information.

[0028] For example, a large amount of unannotated text was first collected from "Daizhi Ge" and "Wikisource". Considering that the meanings of many words in poems are relatively vague and the context information is not rich enough, the collected corpus mainly consists of texts such as prose, novels, dramas, and official documents. To improve the quality of the corpus, duplicate corpus was removed, special characters were replaced, and various punctuation marks in the text were unified. Optionally, the corpus can also be marked with dynasty information. It should be noted that the writing time of historical books such as the Twenty-Four Histories usually lags behind by one dynasty, and their language features mainly reflect the characteristics of the writing era. Therefore, the marked time is based on the dynasty in which the text author wrote. Considering the corpus balance in each period comprehensively, the corpus is divided into pre-Qin, Han, Wei-Jin-Northern and Southern Dynasties, Tang, Song, Yuan, Ming, and Qing according to time.

[0029] Furthermore, the large-scale diachronic dictionary "Hanyu Da Cidian" is selected as the reference for sense alignment. "Hanyu Da Cidian" contains more than 300,000 entries and more than 400,000 senses. The process of disyllabification in ancient Chinese makes the distinction between words and phrases at the diachronic level relatively complex, such as "qi / zi" and "qizi", and "Hanyu Da Cidian" not only interprets words from a practical perspective but also interprets the phrases before disyllabification. Therefore, following the practice of "Hanyu Da Cidian", both words and phrases are marked.

[0030] Optionally, in S304, using the pre-trained language model to automatically mark the meanings of words in the preprocessed ancient Chinese corpus includes: Input the ancient Chinese corpus and any candidate target word into the pre-trained language model to obtain the result of interpretation in context, where the pre-trained language model is an ancient Chinese large language model.

[0031] Optionally, the pre-trained language model is a pre-trained ancient Chinese large language model. In the present invention, the pre-trained ancient Chinese large language model is a large language model that is trained using ancient Chinese corpus on the basis of the basic large language model and the parameters of the large language model are adjusted. This ancient Chinese large language model can refer to, for example, the deep learning language model disclosed in the patent with the application number 2024102954319 and the invention name "An Ancient Chinese Text Processing Method, Device and Storage Medium Based on Language Model".

[0032] The construction of the existing sense-annotated corpus mainly relies on manual annotation, which not only has a high cost but also faces problems such as weak consistency. In the embodiments of the present invention, first, the ancient Chinese large language model is used to interpret the target words in the corpus, and then the semantic vector representation technology is used for sense alignment, and the word sense with the highest matching score is taken as the alignment result, thereby realizing the matching with the dictionary sense. As shown in Table 1 for example: Table 1: Examples of word sense annotation and alignment

[0033] Among them, in Example 1 of Table 1 above, for the corpus "Placing the generals side by side, killing the officials, invading the commanderies and counties, and Fang Guozhen had already risen first at sea. Other bandits held troops and occupied territories, and there were many plunders. The world was in great chaos." that contains the target word "兵" (see the first row of the first column in Table 1), the ancient Chinese large language model is used to determine the semantic meaning of the target word in the corpus as "soldiers, army" (see the first row of the second column in Table 1) through context-based interpretation. Then, the results of context-based interpretation are matched with each semantic item of the target word "兵" in the dictionary to determine its matching score (for example, it can be cosine similarity). For example, in the first row of Table 1, "兵器。0.707" in the third column means that the semantic item "兵器" in the Chinese Dictionary matches the "士兵,軍隊。" in the context-based interpretation of the large model with a matching score of 0.707; "兵卒;军队。0.878" means that the semantic item "兵卒;军队" in the Chinese Dictionary matches the "士兵,軍隊。" in the context-based interpretation of the large model with a matching score of 0.878; and so on.

[0034] In Example 2, for the corpus "When Emperor Gaozu heard about it, he became even more afraid, so he indulged in drinking and accepted bribes to hide himself." that contains the target word "縱酒", the ancient Chinese large language model is used to determine the semantic meaning of the target word in the corpus as "indulging in drinking" (see the second row of the second column in Table 1) through context-based interpretation. Then, the results of context-based interpretation are matched with each semantic item of the target word "縱酒" in the dictionary. For example, "酗酒,任意狂饮。0.901" in the third column of the second row in Table 1 means that the semantic item "酗酒,任意狂饮。" in the Chinese Dictionary matches the "酗酒,任性而飲" in the context-based interpretation of the large model with a matching score of 0.901.

[0035] In large-scale word sense annotation and alignment, it is necessary to traverse the words in a vast amount of corpus and analyze them one by one. All the entries in the Chinese Dictionary are candidate target words. At this time, the problem of word segmentation ambiguity will be faced in the definition of the target word. For example: Xiang Yu killed Song Yi, merged his troops and crossed the river, and made himself the supreme general, and all the generals such as Ying Bu belonged to him. (From "The Annals of Emperor Gao of Han, Volume 1"). In the above sentence, although "上将" is included in the dictionary, it is not the target word, but "上将军" is. Another example: In the future, when one grows old, nothing is accomplished. (From "The Tides of Xupu, Chapter 82"). In the above sentence, "后年" is not the target word, and the correct word segmentation result should be "日后 / 年华 / 老大".

[0036] If a classical Chinese word segmentation tool is used to pre-segment the text, the tool's vocabulary needs to match the *Comprehensive Chinese Dictionary* to cover dictionary entries. However, there is currently no dedicated word segmentation tool for this purpose. Therefore, this invention proposes a word meaning alignment algorithm that can resolve this word segmentation ambiguity problem during word meaning annotation and alignment.

[0037] like Figure 4 As shown, semantic alignment of the automatically annotated corpus includes: 402. Inputting the large language model definition and dictionary definition into the Modern Chinese Vector Representation Model yields vectors v1 and v2, respectively.

[0038] In this step, the semantic representation of the large language model in context is performed using a general modern Chinese text vector model to obtain vector v1, and each dictionary entry is semantically represented using a general modern Chinese text vector model to obtain vector v2.

[0039] The general modern Chinese text vector model is the Qwen3-Embedding model, and the dictionary is a large Chinese dictionary.

[0040] S404. Calculate the cosine similarity s1 between v1 and v2.

[0041] To confirm whether the cosine similarity s1 can reflect the matching effect between the model's definition and the dictionary definition, this embodiment of the invention sets two thresholds: a first threshold Th1 and a second threshold Th2, where the first threshold Th1 is greater than the second threshold Th2. If the cosine similarity s1 is greater than the first threshold Th1, it is determined that the contextual definition and the dictionary definition match, and the dictionary definition is output as the matched definition. If the cosine similarity s1 is less than or equal to the second threshold Th2, it is determined that the contextual definition and the dictionary definition do not match. If the cosine similarity s1 is greater than the second threshold Th2 and less than or equal to the first threshold Th1, the dictionary example sentences are further queried to determine whether there is a match. The first threshold Th1 and the second threshold Th2 are preset based on the effect of historical word meaning alignment.

[0042] Optionally, further querying the dictionary example sentences to determine whether a match exists includes: If there are no example sentences in the dictionary, it is determined that the in-text definition matches the dictionary entry, and the dictionary entry is output as the matching entry. If the dictionary contains example sentences, then steps S406 and S408 are executed.

[0043] In other words, if the cosine similarity s1 is greater than the second threshold Th2 and less than or equal to the first threshold Th1, and there are no example sentences in the dictionary, then the in-text interpretation result is determined to match the dictionary definition, and the dictionary definition is output as the matching definition; if the cosine similarity s1 is greater than the second threshold Th2 and less than or equal to the first threshold Th1, but there are example sentences in the dictionary, then steps S406 and S408 are further executed.

[0044] S406. If the dictionary contains example sentences, then the pre-trained language model is used to obtain the vector representation v3 of the target word in the context and the vector representation v4 of the target word in the example sentences. S408. Calculate the cosine similarity s2 between v3 and v4; To determine whether the contextual definition and the dictionary definition match based on the cosine similarity s2, this embodiment of the invention sets a third threshold Th3, which is less than the second threshold Th2. If s2 is less than or equal to the third threshold Th3, the contextual definition and the dictionary definition are determined not to match; if s2 is greater than the third threshold Th3, the contextual definition and the dictionary definition match, and the dictionary definition is output as the matched definition. The third threshold Th3 is preset based on the effect of historical word meaning alignment.

[0045] Optionally, after the above S402, S404, S406 and S408 and related judgment steps, if multiple matching terms are obtained, they are sorted from high to low according to the similarity score s1, and the matching term with the highest similarity is taken as the annotation alignment result; if only one matching term is obtained, the matching term is taken as the annotation alignment result.

[0046] Regarding the method steps included in S304 above, an example is provided in this embodiment of the invention. In this example, the first threshold Th1 = 0.9, the second threshold Th1 = 0.8, and the third threshold Th1 = 0.6, as shown below. Figure 5 As shown, it specifically includes: S402 and S404 are implemented using module 1, and S406 and S408 are implemented using module 2.

[0047] First, the original text and any candidate target word are input into a pre-trained language model to obtain the context-specific paraphrase results. Then, in Module 1, the context-specific paraphrase results (such as "the day after tomorrow") and each dictionary entry (such as "next year", "the year after next") are semantically represented using the general modern Chinese text vector model Qwen3-Embedding to obtain vectors v1 and v2, and the cosine similarity s1 between the two is calculated. To confirm whether the cosine similarity s1 can reflect the matching effect between the model paraphrase and the dictionary entry, random sampling is performed on the matching data in three intervals of similarity: (0.7, 0.8], (0.8, 0.9], and (0.9, 1.0]. There are 200 pieces of data in each group, and the matching results are manually evaluated. It is found that the matching accuracy in the interval (0.7, 0.8] is 79%, the matching accuracy in the interval (0.8, 0.9] is 93%, and the matching accuracy in the interval (0.9, 1.0] is 100%. Therefore, the following operation plan is set in the embodiment of the present invention: If s1 is greater than 0.9, it means that the model paraphrase and the dictionary entry are highly similar, and the dictionary entry can be output as the matching entry, such as "alcoholism, drink extravagantly. 0.901" in Example 2 of Table 1 above; If s1 is less than or equal to 0.8, it means that the model paraphrase and the dictionary entry are inconsistent, which may be caused by a mismatch in word meaning or a segmentation error. For example, when asking the model to explain "the year after next", what it gives is a jumbled statement like "the day after tomorrow of next year", which fails to meet the similarity matching requirements. In this case, "the year after next" will no longer be labeled; If s1 is between 0.8 and 0.9, it is necessary to further query the dictionary examples and enter Module 2 for matching. If there are no examples under this entry in the dictionary, it is defaulted that the context-specific paraphrase result and the dictionary entry are similar, and this entry is listed as the matching entry.

[0048] Module 2 needs to further determine the annotation result based on the above matching. First, use the special vector model for classical Chinese to obtain the vector representation v3 of the target word (for example, "ruo dang xu hou nian") in the context and the vector representation v4 of the target word in the example sentence (for example, "ri hou nian hua lao da"), and then calculate the similarity s2 between the two. If s2 is greater than 0.6, this dictionary sense is regarded as the matching sense. It should be noted that in this invention, before the formal annotation, the data in the interval (0.8, 0.9] of the output result of Module 1 was sampled, and 200 of them were randomly selected and verified by manual evaluation to verify the correctness of the annotation result, that is, whether the word sense of the word with s1 exceeding 0.8 accurately corresponds to the word sense of the word in the context. After evaluation, the annotation accuracy rate of the data in the interval (0.8, 0.9] is about 78%. Next, this batch of data is input into Module 2 for processing to obtain the similarity s2 given by the classical Chinese semantic vector model. By gradually adjusting the third threshold of s2, when the third threshold is 0.6, the annotation accuracy rate of the interval (0.8, 0.9] can be increased to 85%, and only 10% of the data needs to be filtered at this time. Therefore, the third threshold of Module 2 is set to 0.6 to ensure a high accuracy rate and recall rate. Among them, the special vector model for classical Chinese is the classical Chinese BERT model.

[0049] After the processing of Module 1 and Module 2, if multiple candidate matching senses are obtained, they are sorted according to the score of the similarity s1, and the word sense with the highest similarity is taken as the annotation alignment result. As shown in Table 1 above, in Example 1 of the first row, the large model gives the contextual interpretation "soldier, army" and gets 6 matching items in the "Chinese Dictionary", and the sense and the value of the cosine similarity s1 are respectively: <o Weapon. 0.707; Soldier; army. 0.878; Military, war. 0.769; Kill with a weapon. 0.627; Still harm. 0.487; Meaning die in battle. 0.556; Among them, the value of the cosine similarity s1 of the second item "soldier; army." is the largest, so the annotation alignment result is "soldier; army.".

[0050] Optionally, the method of this embodiment further includes: S310. Randomly extract N results from the annotation alignment results, and perform manual verification according to the classical Chinese corpus, candidate target words, all senses of the Chinese Dictionary, and the final annotation alignment results; Among them, the content of the manual verification includes: Whether there is an error in the annotation result, whether there is an error in the sense alignment, and whether there is an error in the word segmentation; It should be noted that in the translation of "ri hou nian hua lao da", it is guessed that the original text may be incorrect. It may be more appropriate to adjust it according to the correct text, but this translation is carried out according to the original text. Also, the "o0000225" in the original text is guessed to be "00000225", which is translated according to the guessed content. If there are other correct information, it needs to be adjusted according to the actual situation.N is an integer greater than or equal to 300.

[0051] In this invention, to verify the actual annotation effect of the above algorithm, 500 annotation samples were randomly selected. Specifically: ① First, all (target word, text) combinations were extracted, that is, each target word and its surrounding text (50 characters before and after it), with one target word corresponding to one text; ② 500 samples were randomly selected from the (target word, text) combinations. Then, manual evaluation was conducted. During the evaluation, given the corpus text, target words, all definitions in the *Hanyu Da Cidian* (Comprehensive Dictionary of Chinese), and the final annotation results of the algorithm, the evaluation required analyzing whether there were any errors in the annotation results, including definition alignment errors and word segmentation errors. The evaluation results showed that, on the sampled data, the algorithm's annotation accuracy reached 90.6%.

[0052] Optionally, in S306 of this embodiment, semantic clustering is performed on the semantically aligned words, including: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Let q represent the set of semantic terms for the q-th similar word.

[0053] In this invention, after the above clustering process, a multi-level language resource database consisting of "concept-word form-word meaning-corpus" is finally formed.

[0054] In this step, the final clustering result shows that each cluster represents a concept, and the elements in the clusters are semantic terms related to the concept. For example, the cluster expressing the meanings of brilliance and brightness includes the following words and their meanings: Brilliant: dazzling; radiant. | Glorious: magnificent. | Bright: luminous. | Bright: luminous. | Brilliant: luminous; magnificent. | Glorious: radiant; magnificent. | Radiant: luminous. | Radiant: luminous. | Radiant: luminous. | Glorious: brilliant. | Radiant: radiant. | Radiant: luminous. | Radiant: luminous. | Radiant: luminous. | Radiant: luminous.

[0055] The ancient Chinese diachronic semantic corpus of this invention comprises over 190 million characters, 65,698 annotated words, and a total of 126,865 semantic entries. This large-scale, high-coverage corpus provides a solid data foundation for macro-level ancient Chinese semantic retrieval and research, and helps promote the in-depth application of the "distant reading" method in the field of digital humanities.

[0056] Based on the ancient Chinese diachronic semantic corpus of this invention, multi-level language information of "concept-word form-word meaning-corpus" is integrated to provide three intelligent lexical semantic retrieval and analysis functions: First, semantic-level corpus retrieval. Traditional corpora or ancient text databases primarily retrieve relevant data by inputting keywords or expressions, often requiring researchers to perform extensive manual screening and annotation of the search results. This invention's Ancient Chinese diachronic semantic corpus, based on a large-scale semantically annotated corpus, enables the function of querying semantic terms based on words. Users can further click on specific semantic terms to obtain data expressing a particular meaning of the target word. Given the prevalence of polysemy in Ancient Chinese, this function greatly facilitates the retrieval of polysemous word corpora. Figure 7 (a) presents corpus examples of the word “article” when expressing the meaning of the ritual and music system.

[0057] Second, semantic evolution analysis of words. Based on the ancient Chinese diachronic semantic corpus of this invention, the frequency of semantic terms (i.e., the ratio of the number of times a term appears in a certain dynasty to the total number of characters in the corpus of that dynasty) and the proportion of terms (i.e., the ratio of the number of times a term appears in a certain dynasty to the total frequency of the word in that dynasty) of each word in the diachronic dimension are statistically analyzed and presented in a visual way to intuitively reflect the dynamic process of semantic evolution. Figure 7 (b) Clearly demonstrates the historical trajectory of the meaning of the word “article” gradually evolving from “patterns” and “ritual and music system” to “complete text” and “literary works”.

[0058] Third, semantic clustering analysis. Based on the ancient Chinese diachronic semantic corpus of this invention, semantic terms were clustered on a large scale using a semantic clustering algorithm. Users can click "View Similar Terms" to obtain information on other words expressing similar semantics, facilitating further analysis of the evolution of concept names. For example, in the cluster of "article" expressing the meaning of a written text, words such as "literary writing," "words," and "letters" are also included. Figure 7 As shown in (c).

[0059] The ancient Chinese diachronic semantic corpus based on this invention not only facilitates corpus retrieval but also lays a data foundation for research and practical applications in multiple fields. For example, in dictionary compilation and revision, this corpus can provide quantitative evidence for semantic ranking and summarization, and also help analyze the differentiation and merging of semantic terms in different contexts, improving the scientific accuracy and rationality of dictionary entries. Furthermore, through systematic and diachronic semantic term statistics and contextual presentation, researchers can gain a deeper understanding of the semantic distribution and evolutionary trajectory of words in different periods, which is of significant value for data-driven language teaching and humanities research. In language teaching, words with different meanings in ancient and modern Chinese are a challenge in learning classical Chinese. This corpus allows for the batch extraction of words with different meanings in ancient and modern Chinese and their evolutionary trajectories based on large-scale semantic annotation data, providing support for the construction of relevant teaching resources. In humanities research, semantic corpus retrieval functions can be used to obtain research materials closely related to specific images or concepts; based on the diachronic statistical information of semantics, the emergence and development of key concepts can be located; with the help of semantic clustering functions, the evolution of multiple words under the same concept can be systematically analyzed, providing objective linguistic evidence for changes in specific socio-cultural concepts.

[0060] The rich linguistic and semantic information gathered in the corpus provides intuitive and systematic teaching resources for explaining difficult points in classical Chinese instruction. In classical Chinese teaching, classroom instruction cannot fully cover the differences between ancient and modern meanings of words that appear in extracurricular reading and various exams.

[0061] Optionally, embodiments of the present invention further include S314, which involves statistically analyzing the percentage of semantic duration, including: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning of item m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is...n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.

[0062] JS divergence is used to measure the difference between two proportion distributions. For proportion distributions P and Q, the JS divergence... The calculation method is as follows: ; in, ; It is the KL divergence of P relative to AV. It is the KL divergence of Q relative to AV. The KL divergence (Kullback-Leibler Divergence) is also called relative entropy. The calculation method of KL divergence in this invention is the same as that in the prior art.

[0063] It should be noted that the statistical word meaning duration percentage information in S314 above can be used to calculate the word meaning duration percentage information based on the results of summarizing word meaning alignment, and determine the ratio of the frequency of different meanings of the same word to the total frequency of the word in different historical periods; or it can be used to calculate the word meaning duration percentage information based on the results of summarizing word meaning clustering, and determine the ratio of the frequency of different words with the same meaning to the frequency of all words with that meaning in different historical periods.

[0064] As a concrete example, specifically, for a word W, a total of n meanings (m1, m2, ..., ...) have appeared. m n These meanings exhibit a percentage distribution at time t. By calculating the initial time of each word. and the last time The Jensen–Shannon divergence can detect the degree of change in the meaning of words. The smaller the value, the more stable the word meaning; the larger the value, the higher the degree of change in the word meaning, that is, there are significant differences in the distribution of word senses in two periods. When calculating the logarithm, the base can be set to 2, and the divergence threshold can be set to 0.31 to ensure that the Jensen–Shannon divergence value ranges from 0 to 1. When this occurs, the word can be considered to have undergone an obvious change in its meaning. Since classical Chinese teaching mainly focuses on relatively high-frequency words, in this invention, polysemous words that appear more than 1000 times are screened in the corpus, a total of 4954 words, among which 2374 words are determined to have undergone a change in meaning. These words can be regarded as ancient and modern variant words that need to be focused on in classical Chinese teaching and ancient book annotation.

[0065] For example, the Jensen–Shannon divergence of "ge" is 0.996, which belongs to one of the words with the largest change in meaning. As Figure 8 shown, it was originally an ancient character of "song" and was quite commonly used in the Han and Tang dynasties. Later, "song" was differentiated from "ge", and the mainstream meaning of "ge" became the appellation for an older man or an elder brother. The Jensen–Shannon divergence of "chong" is 0.687, which belongs to the words with a relatively high degree of change. As shown in Figure 9, "chong" did not originally refer to insects. In the Han and Tang dynasties, it mainly referred to the general term for all animals. The meaning of insects has been increasing continuously after the Tang Dynasty and has gradually become the dominant sense. According to the above calculation results of the Jensen–Shannon divergence, "dui", "pai", "ju", "yan", "dou", "dajia", "zuzhi", etc. are all ancient and modern variant words with obvious changes in meaning, and can be focused on in classical Chinese teaching and testing. At the same time, the semantic evolution trajectory provided by the diachronic semantic corpus of ancient Chinese in this invention can also provide strong support for the teaching and research of relevant vocabulary.

[0066] Quantitative analysis of information such as images, themes, and rhetoric is a commonly used method in digital humanities research. In literary research, image analysis is particularly important, and the words used to express images often have polysemy. For example, in plant images, "he" has both the meanings of load and bear, and can also refer to lotus; "lan" can refer to orchid and often appears in names, "li" is both a plant and a surname, and "zhu" can refer to bamboo or a musical instrument. Figure 10 is the usage frequency of "he" counted according to its senses. In most cases, "he" is mainly used to express the meaning of bear, and the meaning of lotus is relatively less. If researchers manually screen the literature materials, it will take a lot of time, while the diachronic semantic corpus of ancient Chinese in this invention can provide accurate word sense-level corpus retrieval, greatly improving the research efficiency. In addition, it can also quantitatively track the distribution and evolution of images in different contexts, providing an important data basis for the horizontal and vertical comparative research on the use of literary images.

[0067] In the study of literature and history, it is often necessary to examine specific concepts, and semantic analysis can often be introduced as a support. For example, "the blending of scene and emotion" is often mentioned in poetry and literary criticism throughout history. From the perspective of semantic analysis, it can provide more objective evidence for relevant arguments. As Figure 11 shown, the original meaning of "scene" is sunlight. Its meanings of scenery and scene emerged in the Song Dynasty and became the mainstream meanings since the Yuan Dynasty. Therefore, the emergence and development of the concept of "the blending of scene and emotion" conform to the objective facts of language use. In summary, the diachronic semantic corpus of ancient Chinese in the present invention can provide linguistic evidence for the emergence and development of key concepts, thus making the study of literature and history more scientific and rigorous.

[0068] In the diachronic semantic corpus of ancient Chinese in the present invention, the semantic clustering function provides convenient conditions for systematically analyzing the evolution of multiple words under the same concept. For example, when exploring the issue of the formation of the awareness of the Chinese nation community, words related to the meaning of "China" can be retrieved through the corpus, including "Huaxia", "Central Plains", "China", "Zhongxia", "Zhonghua", "Zhuxia", etc. The frequency statistics results are shown in Figure 12. Before the Qing Dynasty, the usage frequencies of these words were relatively stable. Since the Qing Dynasty, the usage frequency of the word "China" has increased significantly, and at the same time, the word "Zhonghua" has also changed significantly.

[0069] In addition, according to the second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program runs, the method described in any one of the above is executed by a processor.

[0070] Thus, according to this embodiment, functions such as semantic item-level corpus retrieval, lexical semantic evolution analysis, and semantic clustering analysis can be achieved, laying a data foundation for research in multiple fields and practical applications, improving the research efficiency, and at the same time reducing the teaching difficulty of classical Chinese teaching, etc. This not only helps to reveal the laws and mechanisms of the semantic changes of Chinese vocabulary, but also provides important data resources for classical Chinese teaching, dictionary compilation, and related humanities research fields, greatly improving the teaching and research efficiency.

[0071] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0072] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0073] Example 2 Figure 13 An apparatus 1500 for constructing a classical Chinese diachronic semantic corpus according to the first aspect of this embodiment is shown, which corresponds to the method described according to the first aspect of Embodiment 1. Reference Figure 13 As shown, the device 1500 includes: Preprocessing module 1510 is used to collect ancient Chinese corpus containing time information and preprocess the ancient Chinese corpus. The annotation module 1520 is used to automatically annotate the meanings of words in the pre-processed Classical Chinese corpus using a pre-trained language model. The alignment and clustering module 1530 is used to align the meanings of automatically labeled words and perform word meaning clustering on the aligned words; and The sorting module 1540 is used to summarize the results of word meaning alignment and word meaning clustering, and to statistically analyze the historical proportion of word meanings to obtain the ancient Chinese historical word meaning corpus.

[0074] Optionally, the preprocessing module 1510 is also used to perform one or a combination of the following preprocessing steps: Remove duplicate data; Replace special characters; Standardize the punctuation marks in the text; The corpus is tagged with dynasty information.

[0075] Optionally, the annotation module 1520 is also used to input the ancient Chinese corpus and any candidate target word into the pre-trained language model to obtain the in-text interpretation result, wherein the pre-trained language model is an ancient Chinese large language model.

[0076] Optionally, the alignment clustering module 1530 also includes module one 1531, for: The in-text definitions and each dictionary entry are semantically represented using a general modern Chinese text vector model to obtain vectors v1 and v2, respectively. The cosine similarity s1 between v1 and v2 is calculated. The general modern Chinese text vector model is the Qwen3-Embedding model, and the dictionary is a Chinese dictionary. If s1 is greater than the first threshold, then it is determined that the in-text interpretation result matches the dictionary definition, and the dictionary definition is output as the matching definition; If s1 is less than or equal to the second threshold, then it is determined that the in-text interpretation result does not match the dictionary definition. If s1 is greater than the second threshold and less than or equal to the first threshold, then the dictionary example sentences are further queried to determine whether a match is found; If multiple matching terms are obtained, they are sorted from high to low according to the similarity score s1, and the matching term with the highest similarity is taken as the annotation alignment result; if only one matching term is obtained, the matching term is taken as the annotation alignment result. The second threshold is less than the first threshold.

[0077] Optionally, the alignment clustering module 1530 further includes module two 1532, used to further query the dictionary example sentences to determine whether a match is found, including: If there are no example sentences in the dictionary, it is determined that the in-text definition matches the dictionary entry, and the dictionary entry is output as the matching entry. If the dictionary contains example sentences, the pre-trained classical Chinese-specific vector model is used to obtain the target word's vector representation v3 in the context and the target word's vector representation v4 in the example sentences, and the cosine similarity s2 between v3 and v4 is calculated, wherein the classical Chinese-specific vector model is the classical Chinese BERT model. If s2 is less than or equal to the third threshold, it is determined that the contextual interpretation result does not match the dictionary definition; if s2 is greater than the third threshold, the contextual interpretation result matches the dictionary definition, and the dictionary definition is output as the matched definition. The third threshold is less than the second threshold.

[0078] Optionally, the ancient Chinese diachronic semantic corpus construction device 1500 in this embodiment further includes a verification module 1550, used for: N results are randomly selected from the annotation alignment results and manually verified based on the ancient Chinese corpus, candidate target words, all definitions in the Chinese dictionary, and the final annotation alignment results. The manual verification process includes: Are there any errors in the annotation results? Are there any errors in the semantic alignment? Are there any errors in the word segmentation? N is an integer greater than or equal to 300.

[0079] Optionally, the ancient Chinese diachronic semantic corpus construction device 1500 in this embodiment further includes a clustering module 1560, used for: Semantic clustering is performed on the semantically aligned words, including: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Let q represent the set of semantic terms for the q-th similar word.

[0080] Optionally, the ancient Chinese diachronic semantic corpus construction device 1500 in this embodiment further includes a semantic change analysis module 1570, used to statistically analyze the diachronic proportion of semantic information, including: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning of item m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is...n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.

[0081] Therefore, according to this embodiment, functions such as semantic corpus retrieval, lexical semantic evolution analysis, and semantic clustering analysis can be realized, laying a data foundation for research and practical applications in multiple fields, improving research efficiency, and reducing the difficulty of teaching classical Chinese.

[0082] It should be noted that the apparatus of Embodiment 2 can implement all the methods and steps of Embodiment 1, solve the same technical problems, and achieve the same technical effects. The similarities will not be repeated here.

[0083] Example 3 Figure 14 An apparatus 1600 is shown for implementing the method for constructing a classical Chinese diachronic semantic corpus according to the first aspect of this embodiment, the apparatus 1600 corresponding to the method described according to the first aspect of Embodiment 1. Reference Figure 14 As shown, the device 1600 includes: Processor 1610; and Memory 1620, connected to the processor, is used to provide the processor 1620 with instructions to perform the following processing steps: Collect ancient Chinese corpus containing time information, and preprocess the ancient Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The words after automatic annotation are aligned with their meanings, and then the aligned words are clustered according to their meanings. The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the ancient Chinese historical semantic corpus.

[0084] Optionally, the preprocessing includes one or a combination of the following: Remove duplicate data; Replace special characters; Standardize the punctuation marks in the text; The corpus is tagged with dynasty information.

[0085] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: The ancient Chinese corpus and any candidate target word are input into the pre-trained language model to obtain the in-text interpretation result, wherein the pre-trained language model is a large ancient Chinese language model.

[0086] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: The in-text definitions and each dictionary entry are semantically represented using a general modern Chinese text vector model to obtain vectors v1 and v2, respectively. The cosine similarity s1 between v1 and v2 is calculated. The general modern Chinese text vector model is the Qwen3-Embedding model, and the dictionary is a Chinese dictionary. If s1 is greater than the first threshold, then it is determined that the in-text interpretation result matches the dictionary definition, and the dictionary definition is output as the matching definition; If s1 is less than or equal to the second threshold, then it is determined that the in-text interpretation result does not match the dictionary definition. If s1 is greater than the second threshold and less than or equal to the first threshold, then the dictionary example sentences are further queried to determine whether a match is found; If multiple matching terms are obtained, they are sorted from high to low according to the similarity score s1, and the matching term with the highest similarity is taken as the annotation alignment result; if only one matching term is obtained, the matching term is taken as the annotation alignment result. The second threshold is less than the first threshold.

[0087] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: Further querying the dictionary example sentences to determine if a match is found includes: If there are no example sentences in the dictionary, it is determined that the in-text definition matches the dictionary entry, and the dictionary entry is output as the matching entry. If the dictionary contains example sentences, the pre-trained classical Chinese-specific vector model is used to obtain the target word's vector representation v3 in the context and the target word's vector representation v4 in the example sentences, and the cosine similarity s2 between v3 and v4 is calculated, wherein the classical Chinese-specific vector model is the classical Chinese BERT model. If s2 is less than or equal to the third threshold, it is determined that the contextual interpretation result does not match the dictionary definition; if s2 is greater than the third threshold, the contextual interpretation result matches the dictionary definition, and the dictionary definition is output as the matched definition. The third threshold is less than the second threshold.

[0088] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: N results are randomly selected from the annotation alignment results and manually verified based on the ancient Chinese corpus, candidate target words, all definitions in the Chinese dictionary, and the final annotation alignment results. The manual verification process includes: Are there any errors in the annotation results? Are there any errors in the semantic alignment? Are there any errors in the word segmentation? N is an integer greater than or equal to 300.

[0089] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: Semantic clustering is performed on the semantically aligned words, including: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Let q represent the set of semantic terms for the q-th similar word.

[0090] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning of item m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is... n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.

[0091] Therefore, according to this embodiment, functions such as semantic corpus retrieval, lexical semantic evolution analysis, and semantic clustering analysis can be realized, laying a data foundation for research and practical applications in multiple fields, improving research efficiency, and reducing the difficulty of teaching classical Chinese.

[0092] It should be noted that the apparatus of Embodiment 3 can implement all the methods and steps of Embodiment 1, solve the same technical problems, and achieve the same technical effects. The similarities will not be repeated here.

[0093] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0094] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0095] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0097] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0098] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0099] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for constructing a diachronic semantic corpus of Classical Chinese, characterized in that, include: Collect ancient Chinese corpus containing time information, and preprocess the ancient Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The words after automatic annotation are aligned with their meanings, and then the aligned words are clustered according to their meanings. The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the ancient Chinese historical semantic corpus.

2. The method according to claim 1, characterized in that, The preprocessing includes one or a combination of the following: Remove duplicate data; Replace special characters; Standardize the punctuation marks in the text; The corpus is tagged with dynasty information.

3. The method according to claim 1, characterized in that, The automatic word meaning annotation of the pre-processed Classical Chinese corpus using a pre-trained language model includes: The ancient Chinese corpus and any candidate target word are input into the pre-trained language model to obtain the in-text interpretation result, wherein the pre-trained language model is an ancient Chinese large language model. The word sense alignment of the automatically annotated corpus includes: The in-text interpretation results and each dictionary entry are semantically represented using a general modern Chinese text vector model to obtain vectors v1 and v2 respectively, and the cosine similarity s1 between v1 and v2 is calculated. If s1 is greater than the first threshold, then it is determined that the in-text interpretation result matches the dictionary definition, and the dictionary definition is output as the matching definition; If s1 is less than or equal to the second threshold, then it is determined that the in-text interpretation result does not match the dictionary definition. If s1 is greater than the second threshold and less than or equal to the first threshold, then the dictionary example sentences are further queried to determine whether a match is found; If multiple matching terms are obtained, they are sorted from high to low according to the similarity score s1, and the matching term with the highest similarity is taken as the annotation alignment result; if only one matching term is obtained, the matching term is taken as the annotation alignment result. The second threshold is less than the first threshold.

4. The method according to claim 3, characterized in that, The further query of the dictionary example sentences to determine whether a match exists includes: If there are no example sentences in the dictionary, it is determined that the in-text definition matches the dictionary entry, and the dictionary entry is output as the matching entry. If the dictionary contains example sentences, the pre-trained classical Chinese vector model is used to obtain the target word's vector representation v3 in the context and the target word's vector representation v4 in the example sentences, and the cosine similarity s2 between v3 and v4 is calculated, wherein the classical Chinese vector model is the classical Chinese BERT model. If s2 is less than or equal to the third threshold, it is determined that the contextual interpretation result does not match the dictionary definition; if s2 is greater than the third threshold, the contextual interpretation result matches the dictionary definition, and the dictionary definition is output as the matched definition. The third threshold is less than the second threshold.

5. The method according to claim 3 or 4, characterized in that, Also includes: N results are randomly selected from the annotation alignment results and manually verified based on the ancient Chinese corpus, candidate target words, all definitions in the Chinese dictionary, and the final annotation alignment results. The manual verification process includes: Are there any errors in the annotation results? Are there any errors in the semantic alignment? Are there any errors in the word segmentation? N is an integer greater than or equal to 300.

6. The method according to any one of claims 1 to 4, characterized in that, Semantic clustering of words after semantic alignment includes: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Let q represent the set of semantic terms for the q-th similar word.

7. The method according to any one of claims 1 to 4, characterized in that, Also includes: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning of item m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is... n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.

8. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, the method described in any one of claims 1 to 7 is performed by a processor.

9. A device for constructing a diachronic semantic corpus of Classical Chinese, characterized in that, include: The preprocessing module is used to collect ancient Chinese corpus containing time information and to preprocess the ancient Chinese corpus. The annotation module is used to automatically annotate the meanings of words in the pre-processed Classical Chinese corpus using a pre-trained language model. The alignment and clustering module is used to align the meanings of automatically labeled words and then perform word meaning clustering on the aligned words. as well as The sorting module is used to summarize the results of word meaning alignment and word meaning clustering, and to statistically analyze the historical proportion of word meanings to obtain the ancient Chinese historical word meaning corpus.

10. A device for constructing a diachronic semantic corpus of Classical Chinese, characterized in that, include: processor; as well as A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: Collect ancient Chinese corpus containing time information, and preprocess the ancient Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The words after automatic annotation are aligned with their meanings, and then the aligned words are clustered according to their meanings. The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the ancient Chinese historical semantic corpus.