A method and device for constructing an ancient Chinese diachronic word meaning corpus and a storage medium
Patent Information
- Application Number
- CN202512041685.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-12-31
AI Technical Summary
[0003]但是现有技术中,现有大型语料库建设多关注现代汉语,仅少量语料库覆盖到古代汉语,词义标注研究同样以现代汉语为核心,并且没有体现古汉语词汇的词义随时代的演变
[0009]在本公开实施例中,首先收集含时间信息的古代汉语语料。然后,针对古汉语文本特点,利用大语言模型和向量表征技术研发了高精度的词义标注、对齐算法,对库中全部文本进行自动标注,并与《汉语大词典》义项进行对齐。最后,对义项进行自动聚类,形成了“概念-词形-词义-语料”多层级的语言资源库,实现了义项级的语料检索,可视化呈现词语义项的历时演变轨迹,词义聚类查询和概念名称演变分析等功能,提高了研究的效率,降低了教学的难度。
Smart Images

Figure CN122045324B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and storage medium for constructing a diachronic semantic corpus of classical Chinese. Background Technology
[0002] Lexical semantic evolution is an important dimension of language evolution, reflecting the development of human society and the changes in civilization. The semantic system of Chinese vocabulary has been passed down for thousands of years, and its processes and mechanisms of change are particularly complex. This is not only a key and difficult issue in the study of Chinese language history, but also brings many challenges to ancient Chinese information processing, language teaching, lexicography, and humanities research. Therefore, the construction of a large-scale semantic evolution resource database is of great significance.
[0003] However, current technologies for constructing large-scale corpora primarily focus on modern Chinese, with only a small number covering classical Chinese. Semantic annotation research also centers on modern Chinese and fails to reflect the evolution of classical Chinese vocabulary meanings over time. Therefore, constructing a large-scale diachronic corpus of classical Chinese semantics is a pressing technical problem that needs to be solved. Summary of the Invention
[0004] The embodiments of this disclosure provide a method, apparatus, and storage medium for constructing a diachronic semantic corpus of Classical Chinese, so as to at least solve the technical problem of how to construct a large-scale diachronic semantic corpus of Classical Chinese in the prior art.
[0005] According to one aspect of the present disclosure, a method for constructing a classical Chinese diachronic semantic corpus is provided, comprising: collecting classical Chinese corpus containing time information, and preprocessing the classical Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The words after automatic annotation are aligned with their meanings, and then the aligned words are clustered according to their meanings. The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the ancient Chinese historical semantic corpus.
[0006] According to another aspect of the present disclosure, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is executed, a processor performs any of the methods described above.
[0007] According to another aspect of the present disclosure, an apparatus for constructing a classical Chinese diachronic semantic corpus is also provided, comprising: a preprocessing module for collecting classical Chinese corpus containing time information and preprocessing the classical Chinese corpus; The annotation module is used to automatically annotate the meanings of words in the pre-processed Classical Chinese corpus using a pre-trained language model. The alignment and clustering module is used to align the meanings of automatically labeled words and then perform word meaning clustering on the aligned words; and The sorting module is used to summarize the results of word meaning alignment and word meaning clustering, and to statistically analyze the historical proportion of word meanings to obtain the ancient Chinese historical word meaning corpus.
[0008] According to another aspect of the present disclosure, an apparatus for constructing a classical Chinese diachronic semantic corpus is also provided, comprising: a processor; and A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: Collect ancient Chinese corpus containing time information, and preprocess the ancient Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The words after automatic annotation are aligned with their meanings, and then the aligned words are clustered according to their meanings. The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the ancient Chinese historical semantic corpus.
[0009] In this embodiment, firstly, ancient Chinese corpus containing time information is collected. Then, considering the characteristics of ancient Chinese texts, a high-precision semantic annotation and alignment algorithm is developed using large language models and vector representation technology. This algorithm automatically annotates all texts in the database and aligns them with the definitions in the *Comprehensive Dictionary of Chinese Characters*. Finally, the definitions are automatically clustered, forming a multi-level language resource database of "concept-word form-word meaning-corpus." This enables definition-level corpus retrieval, visualization of the diachronic evolution of semantic terms, semantic clustering queries, and concept name evolution analysis, improving research efficiency and reducing teaching difficulty. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure. In the drawings: Figure 1 This is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of this disclosure; Figure 2 This is a schematic diagram of a classical Chinese diachronic semantic corpus system according to an embodiment of the present disclosure; Figure 3 This is a flowchart illustrating the method for constructing a classical Chinese diachronic semantic corpus according to an embodiment of this disclosure; Figure 4 is a schematic flowchart of a word meaning alignment method according to an embodiment of the present disclosure; Figure 5 is an exemplary schematic diagram of a word meaning alignment method according to an embodiment of the present disclosure; Figure 6 is a schematic flowchart of finding diachronic heterosemies in batches in a corpus according to an embodiment of the present disclosure; Figure 7 is a schematic diagram of an example of corpus system functions according to an embodiment of the present disclosure; Figure 8 is a schematic diagram of statistics of usage examples of various meaning items of the Chinese character "ge" in historical documents according to an embodiment of the present disclosure; Figure 9 is a schematic diagram of statistics of usage examples of various meaning items of the Chinese character "chong" in historical documents according to an embodiment of the present disclosure; Figure 10 is a schematic diagram of statistics of usage examples of various meaning items of the Chinese character "he" in historical documents according to an embodiment of the present disclosure; Figure 11 is a schematic diagram of statistics of usage examples of various meaning items of the Chinese character "jing" in historical documents according to an embodiment of the present disclosure; Figure 12 is a schematic diagram of statistics of usage examples of related concept members of "Zhongguo" in historical documents according to an embodiment of the present disclosure; Figure 13 is a schematic diagram of an apparatus for constructing a diachronic semantic corpus of ancient Chinese according to an embodiment of the present disclosure; and Figure 14 is a schematic diagram of another apparatus for constructing a diachronic semantic corpus of ancient Chinese according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0011] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some, but not all, embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.
[0012] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0013] The diachronic complexity of the Chinese semantic system not only makes it difficult for modern people to accurately learn and understand word meanings, but also poses a significant challenge to machine-based semantic analysis. Therefore, constructing a large-scale semantic evolution resource database is of great importance. In recent years, the international academic community has actively constructed such databases. However, in current technology, on the one hand, existing large-scale corpus constructions mostly focus on modern Chinese, with only a small number covering classical Chinese. Annotating large corpora is an extremely challenging task; therefore, most corpora primarily include unannotated raw data, and even when annotation is conducted, it is mainly at the level of word segmentation and part-of-speech classification. On the other hand, semantic annotation research also focuses on modern Chinese, with relatively little research in the field of classical Chinese. Therefore, the construction of a large-scale diachronic semantic corpus resource for classical Chinese remains an urgent issue.
[0014] To address the aforementioned problems in the prior art, this invention provides a method, apparatus, and storage medium for constructing a diachronic semantic corpus of Classical Chinese. First, approximately 160 million characters of Classical Chinese text from the pre-Qin period to the Ming and Qing dynasties are collected and organized, supplemented with nearly 30 million additional characters from other periods, totaling nearly 200 million characters. Then, considering the characteristics of Classical Chinese texts, a high-precision semantic annotation and alignment algorithm is developed using large language models and vector representation technology to automatically annotate all texts in the corpus and align them with the definitions in the *Comprehensive Dictionary of Chinese Characters*. Finally, the definitions are automatically clustered, forming a multi-level language resource corpus comprising "concept-word form-word meaning-corpus".
[0015] It should be noted that in this invention, the meaning and semantic term represent the same meaning and can be used interchangeably.
[0016] It should be noted that, in this invention, the duration percentage information includes at least one of the following percentage information: 1. The ratio of the frequency of different meanings of the same word to the total frequency of the word in different historical periods; and 2. The ratio of the frequency of different words with the same meaning to the frequency of all words with the same meaning in different historical periods.
[0017] Example 1 According to this embodiment, an embodiment of a method for constructing a classical Chinese diachronic semantic corpus is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0018] The method embodiments provided in this example can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Figure 1 A hardware block diagram of a computing device for implementing a method for constructing a diachronic semantic corpus of Classical Chinese is shown. Figure 1 As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), memory for storing data, transmission devices for communication functions, and input / output interfaces. The memory, transmission devices, and input / output interfaces are connected to the processor via a bus. In addition, it may also include a display, keyboard, and cursor control device connected to the input / output interfaces. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0019] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element in a computing device. As involved in the embodiments of this disclosure, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).
[0020] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for constructing a classical Chinese diachronic semantic corpus in this embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the above-mentioned method for constructing a classical Chinese diachronic semantic corpus. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computing device via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0021] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computing device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0022] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of the computing device.
[0023] It should be noted here that, in some optional embodiments, the above... Figure 1 The computing device shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computing devices.
[0024] Figure 2 This is a schematic diagram of the Classical Chinese diachronic semantic corpus system described in this embodiment. (Refer to...) Figure 2As shown, the system includes: a user terminal 100 for user interaction, where users can input commands to use the corpus system and receive results returned by the system. For example, users can input commands for semantic-level corpus retrieval, lexical semantic evolution analysis, or semantic clustering analysis on the user terminal 100, which are transmitted to the computing center 200 via a network. The computing center 200 includes a computing server running the ancient Chinese diachronic semantic corpus system of this invention, as well as large-scale modeling systems (LLMs) such as the Ancient Chinese Large Language Model and the Ancient Chinese Semantic Vector Model. The computing center 200 is connected to a database 300 via a network, which stores databases related to the ancient Chinese diachronic semantic corpus, such as ancient Chinese corpora containing time information, all entries in the *Comprehensive Dictionary of Chinese Characters*, and intermediate processing results during the operation of the ancient Chinese diachronic semantic corpus system. Figure 2 The ancient Chinese diachronic semantic corpus system shown can realize all the functions and steps of this invention.
[0025] Under the aforementioned operating environment, according to the first aspect of this embodiment, a method for constructing a classical Chinese diachronic semantic corpus is provided. This method comprises... Figure 1 The hardware structure shown is implemented, or is made by Figure 2 The system implementation shown. Figure 3 A flowchart illustrating the method is shown below. (Refer to...) Figure 3 As shown, the method includes: S302. Collect ancient Chinese corpus containing time information and preprocess the ancient Chinese corpus; S304. Use a pre-trained language model to automatically annotate the meanings of words in the pre-processed Classical Chinese corpus. S306. Perform word sense alignment on the automatically labeled words, and then perform word sense clustering on the word sense aligned words; S308. Summarize the results of word meaning alignment and word meaning clustering, and statistically analyze the historical proportion of word meanings to obtain the ancient Chinese historical word meaning corpus.
[0026] Optionally, in S302, the collection of ancient Chinese corpus containing time information involves preprocessing the ancient Chinese corpus. To construct a diachronic semantic corpus of ancient Chinese, a large-scale ancient Chinese corpus covering different historical periods and literary styles was collected. For example, approximately 160 million characters of ancient Chinese texts from the pre-Qin period to the Ming and Qing dynasties were collected and organized, and nearly 30 million characters of period texts were added, totaling nearly 200 million characters.
[0027] Optionally, the preprocessing includes one or a combination of the following: Remove duplicate data; Replace special characters; Standardize the punctuation marks in the text; Mark the dynasty information of the corpus.
[0028] For example, a large amount of unlabeled text is first collected from "Daizhige" and "Chinese Wikisource". Considering that many words in poetry have relatively ambiguous meanings and insufficient contextual information, the collected corpus is mainly composed of texts such as proses, novels, dramas and official documents. To improve corpus quality, duplicate corpora are removed, special characters are replaced, and various punctuation marks in the text are unified. Optionally, dynasty information of the corpus can also be marked. It should be noted that the writing time of historical books such as the Twenty-Four Histories is usually one dynasty later than the recorded events, and their linguistic characteristics mainly reflect the characteristics of the writing era. Therefore, the marked time is subject to the writing dynasty of the text author. With comprehensive consideration of the balance of corpus from various periods, the corpus is divided into Pre-Qin, Han, Wei-Jin, Northern and Southern Dynasties, Tang, Song, Yuan, Ming and Qing according to time.
[0029] Further, the large-scale diachronic dictionary *Hanyu Da Cidian* is selected as the reference for sense alignment. *Hanyu Da Cidian* includes more than 300,000 entries and more than 400,000 senses. The process of disyllabification in ancient Chinese makes the distinction between words and phrases on the diachronic level relatively complicated, such as "wife / children" and "wife". From a practical perspective, *Hanyu Da Cidian* not only provides definitions for words, but also provides definitions for phrases before disyllabification. Therefore, in accordance with the practice of *Hanyu Da Cidian*, both words and phrases are labeled.
[0030] Optionally, in S304, using a pre-trained language model to automatically annotate word senses for words in the preprocessed ancient Chinese corpus comprises: Inputting the ancient Chinese corpus and any candidate target word into the pre-trained language model to obtain a context-dependent word sense result, wherein the pre-trained language model is an ancient Chinese large language model.
[0031] Optionally, the pre-trained language model is a pre-trained ancient Chinese large language model. In the present invention, the pre-trained ancient Chinese large language model refers to a large language model that is trained on the basis of a basic large language model using ancient Chinese corpus, and the parameters of the large language model have been adjusted. The ancient Chinese large language model can refer to, for example, the deep learning language model disclosed in the patent with application number 2024102954319 and invention title "Ancient Chinese text processing method, device and storage medium based on language model".
[0032] The construction of existing word sense annotated corpora mainly relies on manual annotation, which is not only costly, but also faces problems such as weak consistency. In the embodiment of the present invention, the ancient Chinese large language model is firstly used to explain the target words in the corpus, and then semantic vector representation technology is used to perform word sense alignment, and the word sense with the highest matching score is taken as the alignment result, thereby realizing matching with dictionary senses. As shown in Table 1: Table 1: Example of word sense annotation and alignment
[0033] Wherein, in Example 1 of the above Table 1, for the corpus containing the target word "兵 (bing)" which reads "並置將帥,殺吏,侵略郡縣,而方國珍已先起海上。他盜擁兵據地,寇掠甚衆。天下大亂" (see the first row of the first column in Table 1), context-dependent word sense disambiguation is performed using an ancient Chinese large language model to determine that the semantic meaning of the target word in the corpus is "soldier, army" (see the first row of the second column in Table 1). Then, the context-dependent word sense result is matched with each sense entry of the target word "兵" in the dictionary to determine the matching score (which may be cosine similarity, for example). For example, in the first row of Table 1, "Weapon. 0.707" in the third column indicates that the matching score between the word sense entry "weapon" from *Hanyu Da Cidian (Comprehensive Chinese Dictionary)* and "soldier, army." in the context-dependent word sense obtained by the large language model is 0.707; "Soldier; army. 0.878" indicates that the matching score between the word sense entry "soldier; army" from *Hanyu Da Cidian* and "soldier, army." in the context-dependent word sense obtained by the large language model is 0.878; and so on.
[0034] In Example 2, for the corpus containing the target word "縱酒 (zong jiu)" which reads "高祖聞之益懼,因縱酒納賂以自晦", context-dependent word sense disambiguation is performed using an ancient Chinese large language model to determine that the semantic meaning of the target word in the corpus is "indulging in drinking, drinking unrestrainedly" (see the second row of the second column in Table 1). Then, the context-dependent word sense result is matched with each sense entry of the target word "縱酒" in the dictionary. For example, "Excessive drinking, drinking unrestrainedly at will. 0.901" in the third column of the second row of Table 1 indicates that the matching score between the word sense entry "Excessive drinking, drinking unrestrainedly at will." from *Hanyu Da Cidian* and "Excessive drinking, drinking willfully." in the context-dependent word sense obtained by the large language model is 0.901.
[0035] In large-scale word sense annotation and alignment, it is required to traverse words in massive corpora for analysis one by one, and all entries in *Hanyu Da Cidian* are candidate target words. At this time, word segmentation ambiguity will be encountered in the definition of target words. For example: Xiang Yu killed Song Yi, crossed the Yellow River with his army, made himself a supreme general, and all generals including Ying Bu belonged to his command. (from *Book of Han • Imperial Annals of Emperor Gaozu, Part I*). In the above sentence, although "shang jiang (senior general)" is included in the dictionary, it is not a target word, and "shang jiang jun (supreme general)" is. For another example: After many years, I have grown old and accomplished nothing. (from *Xiepu Chao • Chapter 82*). In the above sentence, "hou nian (the year after next)" is not a target word, and the correct word segmentation result should be "ri hou (later) / nian hua (time) / lao da (old)".
[0036] If a classical Chinese word segmentation tool is used to pre-segment the text, the tool's vocabulary needs to match the *Comprehensive Chinese Dictionary* to cover dictionary entries. However, there is currently no dedicated word segmentation tool for this purpose. Therefore, this invention proposes a word meaning alignment algorithm that can resolve this word segmentation ambiguity problem during word meaning annotation and alignment.
[0037] like Figure 4 As shown, semantic alignment of the automatically annotated corpus includes: 402. Inputting the large language model definition and dictionary definition into the Modern Chinese Vector Representation Model yields vectors v1 and v2, respectively.
[0038] In this step, the semantic representation of the large language model in context is performed using a general modern Chinese text vector model to obtain vector v1, and each dictionary entry is semantically represented using a general modern Chinese text vector model to obtain vector v2.
[0039] The general modern Chinese text vector model is the Qwen3-Embedding model, and the dictionary is a large Chinese dictionary.
[0040] S404. Calculate the cosine similarity s1 between v1 and v2.
[0041] To confirm whether the cosine similarity s1 can reflect the matching effect between the model's definition and the dictionary definition, this embodiment of the invention sets two thresholds: a first threshold Th1 and a second threshold Th2, where the first threshold Th1 is greater than the second threshold Th2. If the cosine similarity s1 is greater than the first threshold Th1, the contextual definition result and the dictionary definition are determined to match, and the dictionary definition is output as the matched definition. If the cosine similarity s1 is less than or equal to the second threshold Th2, the contextual definition result and the dictionary definition are determined to not match. If the cosine similarity s1 is greater than the second threshold Th2 and less than or equal to the first threshold Th1, the dictionary example sentences are further queried to determine whether there is a match. The first threshold Th1 and the second threshold Th2 are preset based on the effect of historical word meaning alignment.
[0042] Optionally, further querying the dictionary example sentences to determine whether a match exists includes: If there are no example sentences in the dictionary, it is determined that the in-text definition matches the dictionary entry, and the dictionary entry is output as the matching entry. If the dictionary contains example sentences, then steps S406 and S408 are executed.
[0043] In other words, if the cosine similarity s1 is greater than the second threshold Th2 and less than or equal to the first threshold Th1, and there are no example sentences in the dictionary, then the in-text interpretation result is determined to match the dictionary definition, and the dictionary definition is output as the matching definition; if the cosine similarity s1 is greater than the second threshold Th2 and less than or equal to the first threshold Th1, but there are example sentences in the dictionary, then steps S406 and S408 are further executed.
[0044] S406. If the dictionary contains example sentences, then the pre-trained language model is used to obtain the vector representation v3 of the target word in the context and the vector representation v4 of the target word in the example sentences. S408. Calculate the cosine similarity s2 between v3 and v4; To determine whether the contextual definition and the dictionary definition match based on the cosine similarity s2, this embodiment of the invention sets a third threshold Th3, which is less than the second threshold Th2. If s2 is less than or equal to the third threshold Th3, the contextual definition and the dictionary definition are determined not to match; if s2 is greater than the third threshold Th3, the contextual definition and the dictionary definition match, and the dictionary definition is output as the matched definition. The third threshold Th3 is preset based on the effect of historical word meaning alignment.
[0045] Optionally, after the above S402, S404, S406 and S408 and related judgment steps, if multiple matching terms are obtained, they are sorted from high to low according to the similarity score s1, and the matching term with the highest similarity is taken as the annotation alignment result; if only one matching term is obtained, the matching term is taken as the annotation alignment result.
[0046] Regarding the method steps included in S304 above, an example is provided in this embodiment of the invention. In this example, the first threshold Th1 = 0.9, the second threshold Th1 = 0.8, and the third threshold Th1 = 0.6, as shown below. Figure 5 As shown, it specifically includes: S402 and S404 are implemented using module 1, and S406 and S408 are implemented using module 2.
[0047] First, input the original text and any candidate target word into a pre-trained language model to obtain the contextual disambiguation result. Then, in module 1, the contextual disambiguation result (e.g., the day after tomorrow) and each dictionary sense (e.g., next year, the year after next) are semantically represented using the general modern Chinese text vector model Qwen3-Embedding to obtain vectors v1 and v2, and the cosine similarity s1 between the two is calculated. To verify whether cosine similarity s1 can reflect the matching effect between the model's definition and the dictionary sense, random sampling is performed on matching data whose similarities fall into three intervals (0.7, 0.8], (0.8, 0.9], and (0.9, 1.0] respectively, with 200 pieces of data in each group, and manual evaluation is conducted on the matching results. It is found that the matching accuracy is 79% for the interval (0.7, 0.8], 93% for the interval (0.8, 0.9], and 100% for the interval (0.9, 1.0]. Therefore, the embodiment of the present invention provides the following operation scheme: If s1 is greater than 0.9, it indicates that the model's definition is highly similar to the dictionary sense, and the dictionary sense can be output as a matched sense, such as "excessive drinking, drink to immoderation. 0.901" in Example 2 of Table 1 above; If s1 is less than or equal to 0.8, it indicates that the model's definition is inconsistent with the dictionary sense, which may be caused by word sense mismatch or word segmentation error. For example, when the model is required to interpret "the year after next", it gives a semantically disordered expression "the tomorrow of next year", which cannot meet the similarity matching requirement. In this case, "the year after next" will no longer be annotated; If s1 is between 0.8 and 0.9, it is necessary to further query dictionary example sentences and enter module 2 for matching. If there is no dictionary example sentence under this sense, it is defaulted that the contextual disambiguation result is similar to the dictionary sense, and this sense is listed as a matched sense.
[0048] Module 2 needs to further determine the annotation result based on the above matching. First, the dedicated ancient Chinese vector model is used to obtain the vector representation v3 of the target word (e.g., "若当须后年") in the context and the vector representation v4 of the target word in the example sentence (e.g., "日后年华老大"), then the similarity s2 between the two vectors is calculated. If s2 is greater than 0.6, the dictionary sense is regarded as a matched sense. It should be noted that in the present invention, before formal annotation, sampling is performed on the data output by Module 1 that falls within the interval (0.8, 0.9], 200 pieces of data are randomly selected from them to verify the correctness of the annotation results by manual evaluation, that is, to check whether the word sense with s1 exceeding 0.8 accurately corresponds to the word sense of the word in the context. After evaluation, the data annotation accuracy of data in the interval (0.8, 0.9] is approximately 78%. Next, this batch of data is input into Module 2 for processing, and the similarity s2 given by the ancient Chinese semantic vector model is obtained. By gradually adjusting the third threshold of s2, when the third threshold is 0.6, the annotation accuracy of data in the interval (0.8, 0.9] can be increased to 85%, and only 10% of the data needs to be filtered at this time. Therefore, the third threshold of Module 2 is set to 0.6 to ensure high precision and recall. Wherein, the dedicated ancient Chinese vector model is an ancient Chinese BERT model.
[0049] After processing by Module 1 and Module 2, if multiple candidate matching senses are obtained, they are sorted according to the similarity score of s1, and the word sense with the highest similarity is taken as the annotation alignment result. For example, as shown in the above Table 1, in Example 1 of the first row, the large model gives the contextual interpretation "soldiers, army". 6 matching items are obtained in *Hanyu Da Cidian*, and the senses and the cosine similarity s1 values are respectively: Weapons. 0.707; Soldiers; army. 0.878; Military affairs, war. 0.769; Killing with weapons. 0.627; Still meaning injury. 0.487; Refers to death in battle. 0.556; Among them, the second item "Soldiers; army." has the largest cosine similarity s1 value, so the annotation alignment result is "Soldiers; army.".
[0050] Optionally, the method of this embodiment further includes: S310, randomly extracting N results from said annotation alignment results, and performing manual verification based on said ancient Chinese corpus, candidate target words, all senses in *Hanyu Da Cidian* and the final annotation alignment results; Wherein, the content of manual verification includes: whether there are errors in the annotation results, whether there are errors in sense alignment and whether there are errors in word segmentation; N is an integer greater than or equal to 300.
[0051] In this invention, to verify the actual annotation effect of the above algorithm, 500 annotation samples were randomly selected. Specifically: ① First, all (target word, text) combinations were extracted, that is, each target word and its surrounding text (50 characters before and after it), with one target word corresponding to one text; ② 500 samples were randomly selected from the (target word, text) combinations. Then, manual evaluation was conducted. During the evaluation, given the corpus text, target words, all definitions in the *Hanyu Da Cidian* (Comprehensive Dictionary of Chinese), and the final annotation results of the algorithm, the evaluation required analyzing whether there were any errors in the annotation results, including definition alignment errors and word segmentation errors. The evaluation results showed that, on the sampled data, the algorithm's annotation accuracy reached 90.6%.
[0052] Optionally, in S306 of this embodiment, semantic clustering is performed on the semantically aligned words, including: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Let q represent the set of semantic terms for the q-th similar word.
[0053] In this invention, after the above clustering process, a multi-level language resource database of "concept-word form-word meaning-corpus" is finally formed.
[0054] In this step, the final clustering result shows that each cluster represents a concept, and the elements in the clusters are semantic terms related to the concept. For example, the cluster expressing the meanings of brilliance and brightness includes the following words and their meanings: Brilliant: dazzling; radiant. | Glorious: magnificent. | Bright: luminous. | Bright: luminous. | Brilliant: luminous; magnificent. | Glorious: radiant; magnificent. | Radiant: luminous. | Radiant: luminous. | Radiant: luminous. | Glorious: brilliant. | Radiant: radiant. | Radiant: luminous. | Radiant: luminous. | Radiant: luminous. | Radiant: luminous.
[0055] The ancient Chinese diachronic semantic corpus of this invention comprises over 190 million characters, 65,698 annotated words, and a total of 126,865 semantic entries. This large-scale, high-coverage corpus provides a solid data foundation for macro-level ancient Chinese semantic retrieval and research, and helps promote the in-depth application of the "distant reading" method in the field of digital humanities.
[0056] Based on the ancient Chinese diachronic semantic corpus of this invention, multi-level language information of "concept-word form-word meaning-corpus" is integrated to provide three intelligent lexical semantic retrieval and analysis functions: First, semantic-level corpus retrieval. Traditional corpora or ancient text databases primarily retrieve relevant data by inputting keywords or expressions, often requiring researchers to perform extensive manual screening and annotation of the search results. This invention's Ancient Chinese diachronic semantic corpus, based on a large-scale semantically annotated corpus, enables the function of querying semantic terms based on words. Users can further click on specific semantic terms to obtain data expressing a particular meaning of the target word. Given the prevalence of polysemy in Ancient Chinese, this function greatly facilitates the retrieval of polysemous word corpora. Figure 7 (a) presents corpus examples of the word “article” when expressing the meaning of the ritual and music system.
[0057] Second, semantic evolution analysis of words. Based on the ancient Chinese diachronic semantic corpus of this invention, the frequency of semantic terms (i.e., the ratio of the number of times a term appears in a certain dynasty to the total number of characters in the corpus of that dynasty) and the proportion of terms (i.e., the ratio of the number of times a term appears in a certain dynasty to the total frequency of the word in that dynasty) of each word in the diachronic dimension are statistically analyzed and presented in a visual way to intuitively reflect the dynamic process of semantic evolution. Figure 7 (b) Clearly demonstrates the historical trajectory of the meaning of the word “article” gradually evolving from “patterns” and “ritual and music system” to “complete text” and “literary work”.
[0058] Third, semantic clustering analysis. Based on the ancient Chinese diachronic semantic corpus of this invention, semantic terms were clustered on a large scale using a semantic clustering algorithm. Users can click "View Similar Terms" to obtain information on other words expressing similar semantics, facilitating further analysis of the evolution of concept names. For example, in the cluster of "article" expressing the meaning of a written text, words such as "literary writing," "words," and "letters" are also included. Figure 7 As shown in (c).
[0059] The ancient Chinese diachronic semantic corpus based on this invention not only facilitates corpus retrieval but also lays a data foundation for research and practical applications in multiple fields. For example, in dictionary compilation and revision, this corpus can provide quantitative evidence for semantic ranking and summarization, and also help analyze the differentiation and merging of semantic terms in different contexts, improving the scientific accuracy and rationality of dictionary entries. Furthermore, through systematic and diachronic semantic term statistics and contextual presentation, researchers can gain a deeper understanding of the semantic distribution and evolutionary trajectory of words in different periods, which is of significant value for data-driven language teaching and humanities research. In language teaching, words with different meanings in ancient and modern Chinese are a challenge in learning classical Chinese. This corpus allows for the batch extraction of words with different meanings in ancient and modern Chinese and their evolutionary trajectories based on large-scale semantic annotation data, providing support for the construction of relevant teaching resources. In humanities research, semantic corpus retrieval functions can be used to obtain research materials closely related to specific images or concepts; based on the diachronic statistical information of semantics, the emergence and development of key concepts can be located; with the help of semantic clustering functions, the evolution of multiple words under the same concept can be systematically analyzed, providing objective linguistic evidence for changes in specific socio-cultural concepts.
[0060] The rich linguistic and semantic information gathered in the corpus provides intuitive and systematic teaching resources for explaining difficult points in classical Chinese instruction. In classical Chinese teaching, classroom instruction cannot fully cover the differences between ancient and modern meanings of words that appear in extracurricular reading and various exams.
[0061] Optionally, embodiments of the present invention further include S314, which involves statistically analyzing the percentage of semantic duration, including: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is...n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.
[0062] JS divergence is used to measure the difference between two proportion distributions. For proportion distributions P and Q, the JS divergence... The calculation method is as follows: ; in, ; It is the KL divergence of P relative to AV. It is the KL divergence of Q relative to AV. The KL divergence (Kullback-Leibler Divergence) is also called relative entropy. The calculation method of KL divergence in this invention is the same as that in the prior art.
[0063] It should be noted that the statistical word meaning duration percentage information in S314 above can be used to calculate the word meaning duration percentage information based on the results of summarizing word meaning alignment, and determine the ratio of the frequency of different meanings of the same word to the total frequency of the word in different historical periods; it can also be used to calculate the word meaning duration percentage information based on the results of summarizing word meaning clustering, and determine the ratio of the frequency of different words with the same meaning to the frequency of all words with that meaning in different historical periods.
[0064] As a concrete example, specifically, for a word W, a total of n meanings (m1, m2, ..., ...) have appeared. m n These meanings exhibit a percentage distribution at time t. By calculating the initial time of each word. and the last time Jensen–Shannon divergence (JS divergence) between can detect the degree of word meaning change. A smaller value indicates a more stable word meaning, while a larger value indicates a higher degree of word meaning change, that is, the word meaning distribution of the two periods shows a significant difference. The base 2 can also be used for logarithmic calculation, and the divergence threshold is set to 0.31, which ensures that the value range of JS divergence is between 0 and 1. When , the word can be identified as having obvious word meaning change. Since classical Chinese teaching mainly targets relatively high-frequency words, in the present invention, polysemous words that appear more than 1000 times are screened from the corpus, with a total of 4954 words, among which 2374 words are determined to have meaning changes. These words can be regarded as words with different meanings in ancient and modern times that need to be focused on in classical Chinese teaching and ancient book annotation.
[0065] For example, the JS divergence of "哥 (ge)" is 0.996, which is one of the words with the largest word meaning change, as Figure 8 shows, it was originally the ancient character of "歌 (song)", which was quite commonly used in the Han and Tang dynasties. Later, "歌" differentiated from "哥", and the mainstream meaning of "哥" became the address for an older man or an elder brother. The JS divergence of "虫 (chong, insect)" is 0.687, which is a word with a high degree of change. As shown in Figure 9, "虫" did not originally refer to insects. In the Han and Tang dynasties, it mainly referred to the general term of all animals. The insect meaning continued to increase after the Tang Dynasty, and gradually became the dominant meaning. According to the above JS divergence calculation results, "队 (dui, team)", "派 (pai, faction)", "局 (ju, bureau)", "衍 (yan, spread)", "豆 (dou, bean)", "大家 (da jia, everyone)", "组织 (zu zhi, organization)" are all words with different meanings in ancient and modern times whose meanings have changed significantly, which can be focused on explanation and assessment in classical Chinese teaching and testing. At the same time, the word meaning evolution track provided by the diachronic word meaning corpus of ancient Chinese based on the present invention can also provide strong support for the teaching and research of related vocabulary.
[0066] Quantitative analysis of information such as imagery, themes and rhetoric is a common method in digital humanities research. In literary research, imagery analysis is particularly important, and words used to express imagery are often polysemous. For example, among plant imageries, "荷 (he)" has the meaning of bearing and enduring, and can also represent lotus; "兰 (lan)" can refer to orchid and is often used in names, "李 (li)" is both a plant and a surname, "竹 (zhu)" can refer to bamboo and also can refer to musical instruments. Figure 10 is the usage frequency of "he" counted by meaning. In most cases, "he" is mainly used to express the meaning of bearing, and the meaning of lotus is relatively rare. If researchers manually sort out literature materials, it will consume a lot of time, while the diachronic word meaning corpus of ancient Chinese of the present invention can provide accurate word meaning-level corpus retrieval, which greatly improves research efficiency. In addition, it can also enable the distribution and evolution of imagery in different contexts to be tracked quantitatively, providing an important data basis for horizontal and vertical comparative research on the use of literary imagery.
[0067] In cultural and historical research, it is often necessary to investigate specific concepts, and word meaning analysis can often be introduced as support. For example, poetry and literary criticism across dynasties frequently mentions "the integration of feeling and scenery". If analyzed from the perspective of word meaning, more objective evidence can be provided for related arguments. As Figure 11 shown, the original meaning of the character "景" is sunlight, and its meanings of scenery and landscape emerged in the Song Dynasty and became the mainstream meaning since the Yuan Dynasty. Therefore, the emergence and development of the concept of "integration of feeling and scenery" conforms to the objective facts of language usage. To sum up, the diachronic word meaning corpus of ancient Chinese of the present invention can provide linguistic evidence for the emergence and development of key concepts, thereby making cultural and historical research more scientific and rigorous.
[0068] In the diachronic word meaning corpus of ancient Chinese of the present invention, the word meaning clustering function provides convenient conditions for systematically analyzing the evolution of multiple words under the same concept. For example, when exploring the formation of the consciousness of the Chinese national community, words related to the meaning of "Zhongguo (China)" can be queried through the corpus, including "Huaxia", "Zhongyuan", "Zhongguo", "Zhongxia", "Zhonghua", "Zhuxia", etc. The frequency statistics result is shown in Figure 12. Before the Qing Dynasty, the use frequency of these words was relatively stable. Starting from the Qing Dynasty, the use frequency of the word "Zhongguo" increased significantly, and the word "Zhonghua" also underwent obvious changes.
[0069] Furthermore, according to the second aspect of the present embodiment, a storage medium is provided. The storage medium includes a stored program, wherein when the program runs, the processor executes the method described in any one of the above items.
[0070] Therefore, according to the present embodiment, functions such as sense-level corpus retrieval, lexical semantic evolution analysis and word meaning clustering analysis can be realized, which lays a data foundation for multi-field research and practical application, improves research efficiency, and reduces the teaching difficulty of classical Chinese and other courses. This not only helps to reveal the laws and mechanisms of semantic changes in Chinese vocabulary, but also provides important data resources for classical Chinese teaching, dictionary compilation and related humanities research fields, greatly improving the efficiency of teaching and research.
[0071] It should be noted that, for the foregoing method embodiments, they are all described as a series of action combinations for the sake of simple description. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the involved actions and modules are not necessarily required by the present invention.
[0072] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0073] Example 2 Figure 13 An apparatus 1500 for constructing a classical Chinese diachronic semantic corpus according to the first aspect of this embodiment is shown, which corresponds to the method described according to the first aspect of Embodiment 1. Reference Figure 13 As shown, the device 1500 includes: Preprocessing module 1510 is used to collect ancient Chinese corpus containing time information and preprocess the ancient Chinese corpus. The annotation module 1520 is used to automatically annotate the meanings of words in the pre-processed Classical Chinese corpus using a pre-trained language model. The alignment and clustering module 1530 is used to align the meanings of automatically labeled words and perform word meaning clustering on the aligned words; and The sorting module 1540 is used to summarize the results of word meaning alignment and word meaning clustering, and to statistically analyze the historical proportion of word meanings to obtain the ancient Chinese historical word meaning corpus.
[0074] Optionally, the preprocessing module 1510 is also used to perform one or a combination of the following preprocessing steps: Remove duplicate data; Replace special characters; Standardize the punctuation marks in the text; The corpus is tagged with dynasty information.
[0075] Optionally, the annotation module 1520 is also used to input the ancient Chinese corpus and any candidate target word into the pre-trained language model to obtain the in-text interpretation result, wherein the pre-trained language model is an ancient Chinese large language model.
[0076] Optionally, the alignment clustering module 1530 also includes module one 1531, for: The in-text definitions and each dictionary entry are semantically represented using a general modern Chinese text vector model to obtain vectors v1 and v2, respectively. The cosine similarity s1 between v1 and v2 is calculated. The general modern Chinese text vector model is the Qwen3-Embedding model, and the dictionary is a Chinese dictionary. If s1 is greater than the first threshold, then it is determined that the in-text interpretation result matches the dictionary definition, and the dictionary definition is output as the matching definition; If s1 is less than or equal to the second threshold, then it is determined that the in-text interpretation result does not match the dictionary definition. If s1 is greater than the second threshold and less than or equal to the first threshold, then the dictionary example sentences are further queried to determine whether a match is found; If multiple matching terms are obtained, they are sorted from high to low according to the similarity score s1, and the matching term with the highest similarity is taken as the annotation alignment result; if only one matching term is obtained, the matching term is taken as the annotation alignment result. The second threshold is less than the first threshold.
[0077] Optionally, the alignment clustering module 1530 further includes module two 1532, used to further query the dictionary example sentences to determine whether a match is found, including: If there are no example sentences in the dictionary, it is determined that the in-text definition matches the dictionary entry, and the dictionary entry is output as the matching entry. If the dictionary contains example sentences, the pre-trained classical Chinese-specific vector model is used to obtain the target word's vector representation v3 in the context and the target word's vector representation v4 in the example sentences, and the cosine similarity s2 between v3 and v4 is calculated, wherein the classical Chinese-specific vector model is the classical Chinese BERT model. If s2 is less than or equal to the third threshold, it is determined that the contextual interpretation result does not match the dictionary definition; if s2 is greater than the third threshold, the contextual interpretation result matches the dictionary definition, and the dictionary definition is output as the matched definition. The third threshold is less than the second threshold.
[0078] Optionally, the ancient Chinese diachronic semantic corpus construction device 1500 in this embodiment further includes a verification module 1550, used for: N results are randomly selected from the annotation alignment results and manually verified based on the ancient Chinese corpus, candidate target words, all definitions in the Chinese dictionary, and the final annotation alignment results. The manual verification process includes: Are there any errors in the annotation results? Are there any errors in the semantic alignment? Are there any errors in the word segmentation? N is an integer greater than or equal to 300.
[0079] Optionally, the ancient Chinese diachronic semantic corpus construction device 1500 in this embodiment further includes a clustering module 1560, used for: Semantic clustering is performed on the semantically aligned words, including: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Let q represent the set of semantic terms for the q-th similar word.
[0080] Optionally, the ancient Chinese diachronic semantic corpus construction device 1500 in this embodiment further includes a semantic change analysis module 1570, used to statistically analyze the diachronic proportion of semantic information, including: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning of item m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is...n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.
[0081] Therefore, according to this embodiment, functions such as semantic corpus retrieval, lexical semantic evolution analysis, and semantic clustering analysis can be realized, laying a data foundation for research and practical applications in multiple fields, improving research efficiency, and reducing the difficulty of teaching classical Chinese.
[0082] It should be noted that the apparatus of Embodiment 2 can implement all the methods and steps of Embodiment 1, solve the same technical problems, and achieve the same technical effects. The similarities will not be repeated here.
[0083] Example 3 Figure 14 An apparatus 1600 is shown for implementing the method for constructing a classical Chinese diachronic semantic corpus according to the first aspect of this embodiment, the apparatus 1600 corresponding to the method described according to the first aspect of Embodiment 1. Reference Figure 14 As shown, the device 1600 includes: Processor 1610; and Memory 1620, connected to the processor, is used to provide the processor 1620 with instructions to perform the following processing steps: Collect ancient Chinese corpus containing time information, and preprocess the ancient Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The words after automatic annotation are aligned with their meanings, and then the aligned words are clustered according to their meanings. The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the ancient Chinese historical semantic corpus.
[0084] Optionally, the preprocessing includes one or a combination of the following: Remove duplicate data; Replace special characters; Standardize the punctuation marks in the text; The corpus is tagged with dynasty information.
[0085] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: The ancient Chinese corpus and any candidate target word are input into the pre-trained language model to obtain the in-text interpretation result, wherein the pre-trained language model is a large ancient Chinese language model.
[0086] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: The in-text definitions and each dictionary entry are semantically represented using a general modern Chinese text vector model to obtain vectors v1 and v2, respectively. The cosine similarity s1 between v1 and v2 is calculated. The general modern Chinese text vector model is the Qwen3-Embedding model, and the dictionary is a Chinese dictionary. If s1 is greater than the first threshold, then it is determined that the in-text interpretation result matches the dictionary definition, and the dictionary definition is output as the matching definition; If s1 is less than or equal to the second threshold, then it is determined that the in-text interpretation result does not match the dictionary definition. If s1 is greater than the second threshold and less than or equal to the first threshold, then the dictionary example sentences are further queried to determine whether a match is found; If multiple matching terms are obtained, they are sorted from high to low according to the similarity score s1, and the matching term with the highest similarity is taken as the annotation alignment result; if only one matching term is obtained, the matching term is taken as the annotation alignment result. The second threshold is less than the first threshold.
[0087] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: Further querying the dictionary example sentences to determine if a match is found includes: If there are no example sentences in the dictionary, it is determined that the in-text definition matches the dictionary entry, and the dictionary entry is output as the matching entry. If the dictionary contains example sentences, the pre-trained classical Chinese-specific vector model is used to obtain the target word's vector representation v3 in the context and the target word's vector representation v4 in the example sentences, and the cosine similarity s2 between v3 and v4 is calculated, wherein the classical Chinese-specific vector model is the classical Chinese BERT model. If s2 is less than or equal to the third threshold, it is determined that the contextual interpretation result does not match the dictionary definition; if s2 is greater than the third threshold, the contextual interpretation result matches the dictionary definition, and the dictionary definition is output as the matched definition. The third threshold is less than the second threshold.
[0088] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: N results are randomly selected from the annotation alignment results and manually verified based on the ancient Chinese corpus, candidate target words, all definitions in the Chinese dictionary, and the final annotation alignment results. The manual verification process includes: Are there any errors in the annotation results? Are there any errors in the semantic alignment? Are there any errors in the word segmentation? N is an integer greater than or equal to 300.
[0089] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: Semantic clustering is performed on the semantically aligned words, including: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Let q represent the set of semantic terms for the q-th similar word.
[0090] Optionally, the memory 1620 is also configured to provide the processor 1620 with instructions for performing the following processing steps: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning of item m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is... n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.
[0091] Therefore, according to this embodiment, functions such as semantic corpus retrieval, lexical semantic evolution analysis, and semantic clustering analysis can be realized, laying a data foundation for research and practical applications in multiple fields, improving research efficiency, and reducing the difficulty of teaching classical Chinese.
[0092] It should be noted that the apparatus of Embodiment 3 can implement all the methods and steps of Embodiment 1, solve the same technical problems, and achieve the same technical effects. The similarities will not be repeated here.
[0093] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0094] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0095] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0098] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0099] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for constructing a diachronic semantic corpus of Classical Chinese, characterized in that, include: Collect ancient Chinese corpus containing time information, and preprocess the ancient Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The automatically labeled words are aligned in terms of word meaning, and then the aligned words are clustered in terms of word meaning. The word meaning clustering of the aligned words includes: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Represents the set of semantic terms for the q-th similar word; The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the aforementioned ancient Chinese historical semantic corpus, which also includes: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is... n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.
2. The method according to claim 1, characterized in that, The preprocessing includes one or a combination of the following: Remove duplicate data; Replace special characters; Standardize the punctuation marks in the text; The corpus is tagged with dynasty information.
3. The method according to claim 1, characterized in that, The automatic word meaning annotation of the pre-processed Classical Chinese corpus using a pre-trained language model includes: The ancient Chinese corpus and any candidate target word are input into the pre-trained language model to obtain the in-text interpretation result, wherein the pre-trained language model is an ancient Chinese large language model. The word sense alignment of the automatically annotated corpus includes: The in-text interpretation results and each dictionary entry are semantically represented using a general modern Chinese text vector model to obtain vectors v1 and v2 respectively, and the cosine similarity s1 between v1 and v2 is calculated. If s1 is greater than the first threshold, then it is determined that the in-text interpretation result matches the dictionary definition, and the dictionary definition is output as the matching definition; If s1 is less than or equal to the second threshold, then it is determined that the in-text interpretation result does not match the dictionary definition. If s1 is greater than the second threshold and less than or equal to the first threshold, then the dictionary example sentences are further queried to determine whether a match is found; If multiple matching terms are obtained, they are sorted from high to low according to the similarity score s1, and the matching term with the highest similarity is taken as the annotation alignment result; if only one matching term is obtained, the matching term is taken as the annotation alignment result. The second threshold is less than the first threshold.
4. The method according to claim 3, characterized in that, The further query of the dictionary example sentences to determine whether a match exists includes: If there are no example sentences in the dictionary, it is determined that the in-text definition matches the dictionary entry, and the dictionary entry is output as the matching entry. If the dictionary contains example sentences, the pre-trained classical Chinese vector model is used to obtain the target word's vector representation v3 in the context and the target word's vector representation v4 in the example sentences, and the cosine similarity s2 between v3 and v4 is calculated, wherein the classical Chinese vector model is the classical Chinese BERT model. If s2 is less than or equal to the third threshold, it is determined that the contextual interpretation result does not match the dictionary definition; if s2 is greater than the third threshold, the contextual interpretation result matches the dictionary definition, and the dictionary definition is output as the matched definition. The third threshold is less than the second threshold.
5. The method according to claim 3 or 4, characterized in that, Also includes: N results are randomly selected from the annotation alignment results and manually verified based on the ancient Chinese corpus, candidate target words, all definitions in the Chinese dictionary, and the final annotation alignment results. The manual verification process includes: Are there any errors in the annotation results? Are there any errors in the semantic alignment? Are there any errors in the word segmentation? N is an integer greater than or equal to 300.
6. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, the method described in any one of claims 1 to 5 is performed by a processor.
7. A device for constructing a diachronic semantic corpus of Classical Chinese, characterized in that, include: The preprocessing module is used to collect ancient Chinese corpus containing time information and to preprocess the ancient Chinese corpus. The annotation module is used to automatically annotate the meanings of words in the pre-processed Classical Chinese corpus using a pre-trained language model. The alignment and clustering module is used to align the meanings of automatically labeled words and to cluster the meanings of the aligned words. The alignment and clustering module is also used to perform the following operations: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Represents the set of semantic terms for the q-th similar word; The processing module is used to summarize the results of word meaning alignment and word meaning clustering, and to statistically analyze the historical proportion of word meanings to obtain the ancient Chinese historical word meaning corpus. The device also includes: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is... n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.
8. A device for constructing a diachronic semantic corpus of Classical Chinese, characterized in that, include: processor; as well as A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: Collect ancient Chinese corpus containing time information, and preprocess the ancient Chinese corpus; Automatic semantic annotation of words in pre-processed Classical Chinese corpus is performed using a pre-trained language model. The automatically labeled words are aligned in terms of word meaning, and then the aligned words are clustered in terms of word meaning. The word meaning clustering of the aligned words includes: For each sense i, the sense description is encoded using a general modern Chinese text vector model to obtain the sense vector h. i ; For each sense vector h i Calculate h i Other semantic terms h j cosine similarity k ij ; If k ij If the value is greater than 0.75, then sense j is considered a similar sense to sense i, and sense j and sense i form a set A of similar senses. q ; The set A was analyzed using the agglomerative hierarchical clustering algorithm AHC. q Clustering is performed, and for any two sets, if the overlap of elements is greater than or equal to 50%, they are merged. Each cluster represents a concept, and the elements in the cluster are semantic terms related to the concept. i and j are the semantic item numbers of the word, which are integers greater than or equal to 1; q is the item set number, which is an integer greater than or equal to 1. h i h represents the semantic term of the i-th word. j Let A represent the semantic term of the j-th word. q Represents the set of semantic terms for the q-th similar word; The results of semantic alignment and semantic clustering are summarized, and the historical proportion of semantic information is statistically analyzed to obtain the aforementioned ancient Chinese historical semantic corpus, which also includes: For a specific word W, count the set M of semantic terms that appear, where M = (m1, m2, ..., m3). m n Let m1, m2, ..., mn be the number of different meanings of the word W. m n n is an integer greater than or equal to 2; Determine the percentage distribution of the set of senses M at time t. ,in, The percentage of the word W that is the meaning m1 at time t. The percentage of the word W that is a meaning m2 at time t. For word W at time t, the meaning m is... n The proportion; Determine the initial time of the first occurrence of the word and the last time it appeared ,and < ; According to the aforementioned proportion distribution Calculate word W at the initial time and the last time JS divergence between ; If the above If the value is greater than or equal to the preset JS divergence threshold, then the word W is determined to have undergone a significant semantic change; otherwise, the word W is determined not to have undergone a significant semantic change.
Citation Information
Patent Citations
Topic analysis method for Chinese ancient books
CN111581964A