Ancient book text data processing method and device, medium and electronic equipment
By using prompt templates in ancient texts to guide the entity relationship recognition model to understand the entity relationship in ancient texts, the problem of insufficient accuracy of entity relationship recognition in ancient texts in the existing technology is solved, and higher recognition accuracy and relationship extraction effects are achieved.
Patent Information
- Application Number
- CN202510154682.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to accurately identify the relationship types between entities in ancient texts. This is mainly due to the complex semantics of ancient texts, the scarce labeling data and the small amount of context information, which makes it difficult for large language models based on modern text corpus training to be directly applicable to the field of ancient texts.
By obtaining the target ancient book text and the specified entity to be processed, the entity is written into the pre-constructed prompt template, and the target ancient book text and prompt template are input into the pre-constructed entity relationship recognition model, guiding the model to understand the relationship type between entities according to the set reasoning logic.
The accuracy of the recognition of relationship types between entities in ancient texts is improved, and the model is guided to gradually think about and understand the context semantics, reducing noise information interference, and improving the accuracy and overall performance of relationship extraction.
Smart Images

Figure CN120068872A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence and text data processing technology, and in particular, relates to an ancient book text data processing method, device, medium and electronic equipment. Background Art
[0002] With the advent of a new round of scientific and technological revolution represented by cutting-edge technologies such as artificial intelligence, digital humanities has gradually become a research hotspot in the humanities and social sciences. Its core goal is to conduct in-depth analysis and mining of ancient books and historical documents through intelligent means. Among them, the extraction of relationships between entities in ancient texts (i.e., ancient classics texts) is an important part of digital humanities research, aiming to identify and extract the relationship types between entities from ancient texts. This work is of great significance to the construction of ancient book standard knowledge bases and the construction of ancient book knowledge graphs. At the same time, it can assist researchers in clarifying historical contexts and cultural contexts, thereby further promoting the digital inheritance and protection of ancient book resources. However, compared with modern Chinese texts, ancient texts often have significant characteristics such as complex and obscure semantics, scarce annotated data, concise and concise expressions, and less contextual information. This makes it difficult for large language models trained based on modern text corpora to be directly applied to the field of ancient texts, making it difficult to accurately identify the relationship types between entities. Therefore, how to improve the accuracy of identifying the relationship types between entities in the field of ancient texts is a technical problem that needs to be solved urgently. Summary of the invention
[0003] The embodiments of the present application provide a method, device, computer program product or computer program, computer-readable medium and electronic device for processing ancient book text data, which can improve the accuracy of identifying the relationship types between entities in the field of ancient book texts to a certain extent.
[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0005] According to one aspect of an embodiment of the present application, a method for processing ancient book text data is provided, the method comprising: obtaining a target ancient book text to be processed, and a specified entity in the target ancient book text; writing the specified entity into a pre-constructed prompt template, and inputting the target ancient book text and the prompt template into a pre-constructed entity relationship recognition model, the prompt template being used to guide the entity relationship recognition model to understand the relationship type between the specified entities in the target ancient book text according to a set reasoning logic; and outputting the relationship type between the specified entities in the target ancient book text identified by the entity relationship recognition model.
[0006] In some embodiments of the present application, based on the foregoing solution, the prompt template includes a task definition sub-prompt template, an entity type determination sub-prompt template, and a relationship type determination sub-prompt template; wherein, the task definition sub-prompt template is used to guide the entity relationship recognition model to focus on the specified entity in the target ancient book text; the entity type determination sub-prompt template is used to guide the entity relationship recognition model to identify the entity type of the specified entity; the relationship type determination sub-prompt template is used to guide the entity relationship recognition model to identify the relationship type between the specified entities.
[0007] In some embodiments of the present application, based on the foregoing solution, the entity relationship recognition model is constructed through the following steps: obtaining a pre-trained language model, where the pre-trained language model has the ability to understand the semantics of modern texts; obtaining a pre-constructed ancient book training sample set, where each ancient book training sample in the ancient book training sample set is labeled with the actual entity type of the specified entity and the actual relationship type between the specified entities; iteratively training the pre-trained language model based on the ancient book training sample set to obtain the entity relationship recognition model.
[0008] In some embodiments of the present application, based on the foregoing solution, before iteratively training the pre-trained language model based on the ancient book training sample set, the method further includes: defining virtual words for each entity type belonging to the ancient book text and each entity relationship type belonging to the ancient book text, where the virtual words are used to represent the semantic connotation of the corresponding entity type or the semantic connotation of the corresponding entity relationship type; determining the word embedding vectors of each virtual word and adding the word embedding vectors of each virtual word to the word embedding matrix of the pre-trained language model.
[0009] In some embodiments of the present application, based on the foregoing solution, the iteratively training the pre-trained language model based on the ancient book training sample set to obtain the entity relationship recognition model includes: determining a training set and a validation set in the ancient book training sample set; training the pre-trained language model based on the training set to train the ability of the pre-trained language model to identify the entity type of the specified entity in the ancient book text and to identify the relationship type between the specified entities; based on the validation set, respectively verifying the model ability of the pre-trained language model before training and the model ability of the pre-trained language model after training; using the pre-trained language model with the optimal model ability as the latest pre-trained language model, and returning to execute the step of training the pre-trained language model based on the training set until a preset condition is met to obtain the entity relationship recognition model.
[0010] In some embodiments of the present application, based on the foregoing solution, training the pre-trained language model based on the training set includes: writing the specified entities labeled in the ancient book training samples in the training set into the prompt template; inputting the ancient book training samples in the training set and the prompt template into the pre-trained language model, so that the pre-trained language model identifies the entity types of the specified entities in the ancient book training samples in the training set and the relationship types between the specified entities; based on the entity types of the specified entities in the ancient book training samples identified by the pre-trained language model, the relationship types between the specified entities, the actual entity types of the specified entities labeled in the ancient book training samples in the training set, and the actual relationship types between the specified entities, the model parameters of the pre-trained language model are updated reversely based on a preset loss function.
[0011] In some embodiments of the present application, based on the foregoing solution, the method further includes: performing a masking process on the specified entities before inputting the ancient book training samples in the training set and the prompt template into the pre-trained language model.
[0012] According to one aspect of the embodiments of the present application, there is provided an ancient book text data processing device, the device includes: an acquisition unit, configured to acquire a target ancient book text to be processed and the specified entities in the target ancient book text; an input unit, configured to write the specified entities into a pre-constructed prompt template and input the target ancient book text and the prompt template into a pre-constructed entity relationship recognition model, where the prompt template is used to guide the entity relationship recognition model to understand the relationship types between the specified entities in the target ancient book text according to a set inference logic; an output unit, configured to output the relationship types between the specified entities in the target ancient book text recognized by the entity relationship recognition model.
[0013] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium and are adapted to be read and executed by a processor, so that a computer device having the processor executes the method as described above.
[0014] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium, in which at least one program code is stored, and the at least one program code is loaded and executed by a processor to implement the operations performed by the method as described above.
[0015] According to one aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes one or more processors and one or more memories. At least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the method as described above.
[0016] Based on the technical solution proposed in the present application, by writing the specified entities in the target ancient book text into a pre-constructed prompt template, and then inputting the target ancient book text and the prompt template into a pre-built entity relationship recognition model, the entity relationship recognition model can be guided to think step by step. On the basis of fully understanding the context semantics of the specified entities in the target ancient book text and the entity type judgment, the relationship type between the specified entities in the target ancient book text can be accurately recognized.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0019] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown;
[0020] Figure 2 A flowchart of the ancient book text data processing method in the embodiments of the present application is shown;
[0021] Figure 3 An application scenario diagram of the ancient book text data processing method in the embodiments of the present application is shown;
[0022] Figure 4 A block diagram of the ancient book text data processing device in the embodiments of the present application is shown;
[0023] Figure 5 A structural diagram of the electronic device in the embodiments of the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0025] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0026] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0027] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may be changed according to the actual situation.
[0028] At present, with the advent of a new round of scientific and technological revolution represented by cutting-edge technologies such as artificial intelligence, digital humanities has gradually become a research hotspot in the fields of humanities and social sciences. Its core goal is to deeply analyze and mine ancient Chinese historical materials and documents through intelligent means. Among them, the extraction of relationships between entities in ancient Chinese texts (i.e., ancient classical texts) is an important part of digital humanities research, aiming to identify and extract the types of relationships between entities from ancient Chinese texts. This work is of great significance for the construction of an ancient Chinese standard knowledge base and the construction of an ancient Chinese knowledge graph, and can also assist researchers in clarifying the historical context and cultural context, thereby further promoting the digital inheritance and protection of ancient Chinese resources. However, compared with modern Chinese texts, ancient Chinese texts often have significant characteristics such as complex and obscure semantics, scarce labeled data, concise and condensed expressions, and less context information. This makes it difficult for large language models trained on modern text corpora to be directly applicable to the field of ancient Chinese texts, and thus difficult to accurately identify the types of relationships between entities therein. In this case, the present application proposes a technical solution for processing ancient Chinese text data to improve the accuracy of identifying the types of relationships between entities in the field of ancient Chinese texts.
[0029] Figure 1 The figure shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied.
[0030] As Figure 1 shown, the system architecture may include a terminal device (such as Figure 1 one or more of the smart phone 101, tablet computer 102, and portable computer 103 shown in the figure. Of course, it may also be a desktop computer, etc., but is not limited thereto, and the present application does not make any restrictions here), a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as wired communication links, wireless communication links, and the like.
[0031] It should be noted that the ancient Chinese text data processing method provided by the embodiments of the present application may be executed by the server 105. Correspondingly, the ancient Chinese text data processing device is generally arranged in the server 105. Of course, in other cases, the ancient Chinese text data processing method may also be executed by the terminal device.
[0032] It should also be noted that Figure 1 the numbers of the terminal device, network, and server in are merely illustrative. According to the implementation requirements, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0033] The implementation details of the ancient book text data processing method according to the embodiments of the present application are elaborated in detail as follows:
[0034] Figure 2 The flowchart of the ancient book text data processing method according to the embodiments of the present application is shown. This ancient book text data processing method can be executed by a device with computing and processing capabilities, such as the Figure 1 server 105 shown in Figure 2 As shown, this ancient book text data processing method at least includes steps 210 to 230, which are introduced in detail as follows:
[0035] In step 210, a target ancient book text to be processed and specified entities in the target ancient book text are obtained.
[0036] In the present application, the target ancient book text may be an ancient book text input by a user for identifying the relationship type between entities therein, and the specified entities may be entities to be identified in the target ancient book text specified by the user. It should be noted that the specified entities are the entity words in the target ancient book text.
[0037] To enable those skilled in the art to better understand the present application, the following Figure 3 is described by taking a specific embodiment as an example.
[0038] See Figure 3 , which shows an application scenario diagram of the ancient book text data processing method according to the embodiments of the present application.
[0039] As Figure 3 shown in sub - figure (a) of
[0040] , the user can input in the terminal interface: "Please identify the relationship type between 'Deng Fei' and 'Fu River' in 'The people of Zitong County, Deng Fei, Wang Linggong, etc. recruited more than ten thousand people in the township, and again set up camps ten miles east of the prefecture, north of the Fu River, to respond to them'".
[0041] Continuing to refer to Figure 2 , in step 220, the specified entities are written into a pre - constructed prompt template, and the target ancient book text and the prompt template are input into a pre - constructed entity relationship recognition model. The prompt template is used to guide the entity relationship recognition model to understand the relationship type between the specified entities in the target ancient book text according to the set reasoning logic.
[0042] In this application, the pre-constructed prompt template can give corresponding prompts to the entity relationship recognition model, enabling the entity relationship recognition model to understand the relationship type between the specified entities in the target ancient book text according to the reasoning logic defined in the prompt template. In other words, the prompt template can organize the original input into a prompt in the form of "fill in the blanks", that is, realize the mapping P: X→T from the input to the prompt token sequence * , and the token sequence after mapping contains at least one or more special [mask] tokens.
[0043] In this application, it should be noted that the "[mask]" here is used to indicate the target content that the entity relationship recognition model needs to understand, that is, the "blank" to be filled in the "fill in the blanks".
[0044] In this application, the prompt template can include three types of sub-prompt templates, which generally guide the model to think step by step, and on the basis of fully understanding the context semantics and entity type judgment, obtain the final judgment result of the relationship type between the specified entities. Specifically, the three types of sub-prompt templates can include a task definition sub-prompt template, an entity type determination sub-prompt template, and a relationship type determination sub-prompt template.
[0045] Among them, the task definition sub-prompt template is used to guide the entity relationship recognition model to pay attention to the specified entities in the target ancient book text (the task definition sub-prompt template is also used to define the task for the entity relationship recognition model); the entity type determination sub-prompt template is used to guide the entity relationship recognition model to identify the entity types of the specified entities; the relationship type determination sub-prompt template is used to guide the entity relationship recognition model to identify the relationship types between the specified entities. These sub-prompt templates can guide the entity relationship recognition model to analyze and reason step by step, and obtain the judgment result of the relationship type between the specified entities on the basis of fully understanding the entity types and context semantic information of the specified entities in the ancient book text.
[0046] Next, this application will introduce the above-mentioned sub-prompt templates in detail.
[0047] (1) Task definition sub-prompt template.
[0048] In some embodiments, the specific content of the task definition sub-prompt template can refer to the following example:
[0049] "In the following text, it is necessary to pay attention to the two entities <e1_start>Entity 1<e1_end> and <e2_start>Entity 2<e2_end> in the text, and judge the relationship type between these two entities".
[0050] Among them, "Entity 1" and "Entity 2" are two entities to be specified. For example, if the target ancient book text is "Deng Fei, Wang Linggong and others in Zitong County recruited more than ten thousand people in the township, and again, ten li east of the prefecture, north of the Fu River, they set up stockades to respond to them", and the specified entities are "Deng Fei" and "Fu River". Then, writing the specified entities into the task definition sub-prompt template will result in: "In the following text, it is necessary to pay attention to the two entities <e1_start>Deng Fei<e1_end> and <e2_start>Fu River<e2_end> in the text, and judge the relationship type between these two entities."
[0051] Furthermore, "<e1_start>", "<e1_end>" and "<e2_start>", "<e2_end>" are special entity marker tokens, which respectively represent the start and end of the subject (i.e., the head entity) and the object (i.e., the tail entity) in a relationship description instance. These special tokens are added to the vocabulary of the pre-trained language model (i.e., the form of the entity relationship recognition model before training) as learnable tokens during the model training stage and are continuously updated as the model is trained. This sub-prompt template can indicate the task type and the entity pairs that need to be concerned to the model, guide the model to narrow the scope of attention, weaken the interference of noise information, so as to guide the pre-trained language model to pay full attention to the key entity information in the context, and at the same time make the model clear about the final task goal, which can significantly stimulate the model's relationship extraction knowledge and ability, and at the same time reduce the interference of irrelevant information.
[0052] (2) Entity type determination sub-prompt template. This sub-prompt template can require the model to judge the entity types of the head and tail entities that constitute the entity relationship.
[0053] In some embodiments, the specific content of the entity type determination sub-prompt template can refer to the following example:
[0054] "Among them, <e1_start>Entity 1<e1_end> is <mask>Type of entity. <e2_start>Entity 2<e2_end> is <mask>Type of entity.”
[0055] In this application, for example, if the specified entities are "Deng Fei" and "Fu River", through the entity type determination sub - prompt template, it can be obtained that: "Among them, <e1_start>Deng Fei<e1_end> is an entity of <person> type, and <e2_start>Fu River<e2_end> is an entity of <location> type.”
[0056] In this application, judging the entity types of the specified entities, that is, requiring the model to first identify the entity types of the subject and object that constitute the entity relationship type before obtaining the determination result of the final relationship type between the specified entities, can enhance the accuracy of the model in subsequent judgment of the entity relationship type.
[0057] (3) Relationship type determination sub - prompt template. This sub - prompt template can explicitly guide the entity relationship recognition model to make a judgment on the entity relationship type based on the relationship description context information and entity types.
[0058] In some embodiments, the specific content of the relationship type determination sub - prompt template can refer to the following example:
[0059] "Considering the context and entity types where the two entities are located, <e1_start>Entity 1<e1_end> <mask><e2_start>Entity 2<e2_end>.”
[0060] In this application, based on the understanding information guided by the task definition sub-prompt template and the entity type determination sub-prompt template, the entity relationship recognition model can fully understand the context semantic connotation where the entity relationship is located and the entity type. On this basis, through the relationship type determination sub-prompt template, the entity relationship recognition model can be guided to infer the prediction result of the relationship type between the specified entities in the target ancient book text.
[0061] Continue to refer to Figure 2 , in step 230, output the relationship type between the specified entities in the target ancient book text recognized by the entity relationship recognition model.
[0062] In this application, as Figure 3 shown in subgraph (a) and subgraph (b) therein, after the user inputs "Please identify the relationship type between 'Deng Fei' and 'Fu River' in 'Zitong County citizen Deng Fei, Wang Linggong, etc. recruited more than ten thousand people in the township, and again set up a stockade ten miles east of the prefecture, north of the Fu River, to respond to it.'" in the terminal interface and requests an answer, an answer that the relationship between 'Deng Fei' and 'Fu River' is a garrison relationship can be output.
[0063] In this application, by writing the specified entities in the target ancient book text into a pre-constructed prompt template, and then inputting the target ancient book text and the prompt template into a pre-constructed entity relationship recognition model, the entity relationship recognition model can be guided to think and reason step by step. In this process, the model can not only fully understand the context semantics of the specified entities in the target ancient book text, but also make judgments on the entity types, thereby providing more comprehensive background information for the recognition of the relationship type.
[0064] Specifically, first, the technical solution proposed in this application can enhance the model's understanding of context. That is, through the guidance of the prompt template, the model can better capture the context information of the specified entity, thereby understanding the mutual relationship between entities and their functions in ancient book texts. This context understanding is crucial for accurately identifying complex relationships. Especially in ancient book texts, the meaning of entities often depends on their context. Second, it can improve the accuracy of entity type recognition. Before identifying entity relationships, the model first determines the entity types. This process can effectively reduce errors in the recognition process. For example, clarifying that "Deng Fei" is of the "person" type and "Fu River" is of the "location" type can help the model more accurately infer the relationship between them and avoid confusing entities of different types. Third, it can reduce the interference of noise information. Through the clear prompt template, the model can focus on the key information related to the specified entity, thereby effectively reducing the interference of irrelevant information. This focusing ability enables the model to more accurately extract valuable relationship information when processing ancient book texts. Fourth, it can promote the effective transfer of knowledge. The pre-constructed entity relationship recognition model can utilize the knowledge and experience obtained in other text fields and make adaptive adjustments in combination with the characteristics of ancient book texts. This knowledge transfer can improve the performance of the model in ancient book texts, especially when the training data is scarce. Fifth, it can enhance the overall relationship extraction ability. By systematically guiding the model to understand and analyze the structure of the text, the ability of the model to extract entity relationships in ancient book texts can be significantly improved in the end. This not only helps in the analysis of a single text but also provides a solid foundation for the construction of knowledge graphs and the improvement of standard knowledge bases.
[0065] Next, this application will make a detailed description of the construction process of the entity relationship recognition model. Specifically, it includes the following steps 201 to step 203:
[0066] Step 201: Obtain a pre-trained language model that has the ability to understand the semantics of modern texts.
[0067] Step 202: Obtain a pre-constructed ancient book training sample set, where each ancient book training sample in the set is labeled with the actual entity type of the specified entity and the actual relationship type between the specified entities.
[0068] Step 203: Iteratively train the pre-trained language model based on the ancient book training sample set to obtain the entity relationship recognition model.
[0069] In this application, the pre-trained language model refers to a model trained on large-scale text data. Specifically, by training on rich modern text corpora, the pre-trained language model can capture the usage habits of contemporary language, popular vocabulary, context changes, etc., so as to better understand the actual meaning of the text. That is, by learning the structural, grammatical, semantic and other features of the language, these models can perform well in a variety of natural language processing tasks, and their applications can significantly improve the effect of entity relationship extraction in ancient book texts. In some embodiments, the pre-trained language model can be a BERT model, a GPT model, or a RoBERTa model. Specifically, this application does not make too many limitations in this regard. In this application, the pre-trained language model can be trained to obtain the entity relationship recognition model.
[0070] In this application, for the ancient book training sample set, it can be constructed in the following way: Given a relation extraction data set D = {X, Y}, X represents the ancient book training sample set, and Y represents the corresponding entity relationship type label set. For each ancient book training sample x in X, it can be represented as a context sequence {w, w s , w o , w n}, where |n| is the length of the sample context sequence, w s is the subject of the relation instance (a specified entity, i.e., the head entity), and w o is the object of the relation instance (another specified entity, i.e., the tail entity). Y can be represented as {y 1 , y 2 , …, y m}, where m is the number of entity relationship types. The goal of model training is to predict the relationship type between the subject w s and the object w o according to the sample context sequence. Where
[0071] In this application, before step 203 above, that is, before iteratively training the pre-trained language model based on the ancient book training sample set, the following steps 2001 to 2002 can also be executed:
[0072] Step 2001, define virtual words for each entity type belonging to the ancient book text and each entity relationship type belonging to the ancient book text. The virtual words are used to represent the semantic connotation of the corresponding entity type or the semantic connotation of the corresponding entity relationship type.
[0073] Step 2002, determine the word embedding vectors of each virtual word, and add the word embedding vectors of each virtual word to the word embedding matrix of the pre-trained language model.
[0074] In some embodiments of the present application, the entity types belonging to ancient book texts may include "person", "location", "official position", "book", etc., and the entity relationship types belonging to ancient book texts may include "hostile attack", "holding office", "superior-subordinate", "political support", "colleague", "arrival", "management", "born in a certain place", "garrison", "alias", "parents", "brothers", etc.
[0075] In the present application, there are significant differences between the entity types and relationship types existing in ancient book texts and modern Chinese texts. For example, the entity type of "official position" commonly found in ancient book texts, as well as relationship types such as "garrison" and "alias", are very rare in modern Chinese. Taking the "alias" relationship as an example, in ancient book texts, it can be described as someone having other names such as a "style name" or "courtesy name". This relationship is very common when introducing the life of ancient figures, but in the context of modern Chinese, this relationship is extremely rare. The existence of this phenomenon causes the pre-trained language model to seriously lack understanding of the relationship types and entity types existing in ancient book texts, resulting in poor performance in relation extraction. To solve this problem, the present application combines the characteristics of prompt learning and designs two ancient book external domain knowledge injection mechanisms, namely ancient book entity type knowledge injection and ancient book relationship type knowledge injection. Next, a detailed introduction to these two knowledge injection mechanisms will be given.
[0076] First, ancient book entity type knowledge injection.
[0077] In the present application, to help the pre-trained language model better understand and distinguish entities in ancient book texts, virtual words are designed for each entity type belonging to ancient book texts to represent their semantic connotations, that is, it is assumed that there is a virtual vocabulary in the vocabulary of the pre-trained language model that can represent all the semantic words corresponding to a certain relationship type. Specifically, for each entity type e, the present application first tokenizes the type description of its dataset to obtain the corresponding token sequence Subsequently, the word embedding vector of the corresponding virtual word a is initialized according to the following formula e :
[0078]
[0079] where E(.) represents the virtual word embedding vector. Then, these new word embedding vectors are added to the word embedding matrix of the original pre-trained language model.
[0080] Second, ancient book entity relationship type knowledge injection.
[0081] In this application, to help the model better understand the domain relationship knowledge existing in ancient Chinese texts, similar to the mechanism of injecting ancient Chinese entity type knowledge, this application can also define virtual words for each entity relationship type belonging to ancient Chinese texts, and initialize the word embedding vectors with the corresponding type descriptions of the entity relationship types in the same way, so as to integrate external relationship type knowledge, and add these new word embedding vectors to the word embedding matrix of the pre-trained language model.
[0082] To enable those skilled in the art to better understand the injection of ancient Chinese entity relationship type knowledge, the following takes the ancient Chinese entity relationship type "alias" as an example for illustration.
[0083] The first step is to obtain the relationship type description: Assume that the type description of "alias" is "an entity that represents the names of ancient Chinese figures other than personal names".
[0084] The second step is to perform word segmentation on the type description of "alias" to obtain the relevant token sequence: "an entity that represents the names of ancient Chinese figures other than personal names". The tokens obtained through word segmentation are: "except", "personal name", "other than", "represent", "China", "ancient", "figure", "name", "of", "entity".
[0085] The third step is to obtain word embeddings, that is, to obtain the word vectors corresponding to these tokens from the word embedding matrix of the pre-trained language model: E("except"), E("personal name"), E("other than"), E("represent"), E("China"), E("ancient"), E("figure"), E("name"), E("of"), E("entity").
[0086] The fourth step is to calculate the average value, that is, to take the average of these word vectors E("except"), E("personal name"), E("other than"), E("represent"), E("China"), E("ancient"), E("figure"), E("name"), E("of"), E("entity") to obtain the word embedding vector of the "alias" virtual word.
[0087] The fifth step is to add to the word embedding matrix, add the word embedding vector of the virtual word to the word embedding matrix of the pre-trained language model to ensure that the model can recognize and use these new words (i.e., the virtual words defined for each entity relationship type).
[0088] In this application, before iteratively training the pre-trained language model based on the ancient book training sample set, by defining specific virtual words for each entity type and each entity relationship type in the ancient book text, the semantic connotations of the corresponding entity types and entity relationship types can be effectively represented, thereby enriching the model's understanding of the ancient book text. Next, the word embedding vector for each virtual word needs to be determined, and these word embedding vectors are integrated into the word embedding matrix of the pre-trained language model. In this way, the prompt information can be enhanced and enriched, thus improving the effect of relationship extraction between entities in the ancient book text.
[0089] In this application, the iterative training of the pre-trained language model based on the ancient book training sample set to obtain the entity relationship recognition model can be performed according to the following steps 2031 to 2034:
[0090] Step 2031, determine the training set and the validation set in the ancient book training sample set.
[0091] Step 2032, train the pre-trained language model based on the training set to train the ability of the pre-trained language model to recognize the entity types of specified entities in the ancient book text and the relationship types between the specified entities.
[0092] Step 2033, based on the validation set, respectively verify the model capabilities of the pre-trained language model before training and the pre-trained language model after training.
[0093] Step 2034, use the pre-trained language model with the optimal model capability as the latest pre-trained language model, and return to execute the step of training the pre-trained language model based on the training set until a preset condition is met to obtain the entity relationship recognition model.
[0094] In this application, by dividing the ancient book training sample set into a training set for model training and a validation set for evaluating the model performance, it can effectively prevent the model from overfitting during the training process and ensure the generalization ability of the model on unseen data. By training the pre-trained language model with the training set, the parameters of the model can be adjusted to enable it to have the ability to identify specific entity types in ancient book texts and the ability to understand and identify the relationship types between different entities. Through in-depth learning of ancient book texts, the model can capture the semantic features in the texts, thereby improving its recognition accuracy. The performance of the pre-trained language model before and after training is evaluated using the validation set. By comparing the changes in the model's capabilities, the impact of the training process on the model performance can be intuitively understood. This verification process can not only identify the advantages and disadvantages of the model but also provide a basis for subsequent model adjustment. Finally, the pre-trained language model with the optimal model capabilities is selected as the latest model and returned to the training step for further iterative training until the model meets the preset conditions. In this way, the continuous optimization of the model can be ensured, making the finally obtained entity relationship recognition model have higher accuracy and robustness in the processing of ancient book texts.
[0095] It should be noted that the preset conditions can be set according to actual needs. For example, it can be that the number of iterative training reaches the set number, or it can be that the recall rate (or F1) of the model reaches the set value.
[0096] In summary, through a systematic iterative training process and combined with the characteristics of ancient book texts, this application effectively improves the performance of the obtained entity relationship recognition model in entity relationship recognition tasks, providing strong technical support for ancient book research and related applications.
[0097] In this application, step 2032 above, that is, training the pre-trained language model based on the training set, can be executed according to the following steps 20321 to 20323:
[0098] Step 20321, write the specified entities labeled in the ancient book training samples in the training set into the prompt template.
[0099] Step 20322, input the ancient book training samples in the training set and the prompt template into the pre-trained language model, so that the pre-trained language model can identify the entity types of the specified entities in the ancient book training samples in the training set and the relationship types between the specified entities.
[0100] Step 20323: Based on the entity types of the specified entities in the ancient book training samples in the training set recognized by the pre-trained language model, the relationship types between the specified entities, the actual entity types of the specified entities annotated in the ancient book training samples in the training set, and the actual relationship types between the specified entities, the model parameters of the pre-trained language model are updated backward based on a preset loss function.
[0101] In the present application, during the model training process, one of the optimization objectives can be to calculate the cross-entropy loss between the predicted probability distribution of the virtual word at the [MASK] position in the entity type determination sub-prompt template and the actual entity type annotated for the specified entity in the ancient book training sample, so as to update the model parameters. In this way, the model's understanding of entity types in ancient book texts can be significantly enhanced, and the accuracy of entity type classification in ancient book texts can be greatly improved. At the same time, as an auxiliary task for entity relationship type judgment, the judgment of entity types also provides valuable hint signals for relationship extraction. In some embodiments, the entity type judgment loss function L entity (i.e., the preset loss function) is as follows:
[0102] L entity = 0.5 * CELoss(entity1_logits, entity1_labels)
[0103] + 0.5 * CELoss(entity2_logits, entity2_labels)(2)
[0104] where CELoss is the cross-entropy loss function, entity1_logits and entity1_labels are respectively the predicted probability distribution of the entity type and the actual entity type annotated for the head entity (i.e., the specified entity) therein; entity2_logits and entity2_labels are respectively the predicted probability distribution of the entity type and the actual entity type annotated for the tail entity (i.e., the specified entity) therein. In the present application, during the model training process, one of the optimization objectives can be to calculate the cross-entropy loss between the predicted probability distribution of the virtual word at the [MASK] position in the relationship type determination sub-prompt template and the actual relationship type between the specified entities in the ancient book training sample. During the inference process, the fine-tuned pre-trained language model predicts the probability distribution of the virtual word at the [MASK] position in the relationship type determination sub-prompt template, and finds the entity relationship type corresponding to the virtual word with the highest probability according to the relationship semantic mapping as the final prediction result. In some embodiments, the entity relationship type judgment loss function L relation is:
[0105] L relation =CELoss(rel_logits,rel_labels)(3)
[0106] Among them, rel_logits and rel_labels are the predicted probability distribution of entity relationship types and the actual relationship types between the labeled specified entities, respectively.
[0107] Furthermore, in this application, the overall training loss function is:
[0108]
[0109] In this application, the following step 20320 may also be performed:
[0110] Step 20320, before inputting the ancient book training samples in the training set and the prompt template into the pre-trained language model, mask processing is performed on the designated entity.
[0111] In this application, considering the lack of samples in the ancient book training sample set, the model is very likely to overfit the training set during the training process, resulting in a decrease in its generalization ability. Based on this, this application proposes an entity mask mechanism to add regularization constraints to model training to enhance the generalization of the model and prevent overfitting. Specifically, during the training process, in addition to the combined prompt template mentioned above, this application can also construct an entity mask prompt template for each sample.
[0112] In some embodiments, during model training, two different prompt templates (ie, a combined prompt template and an entity mask prompt template) may be input to the model at the same time, or only the entity mask prompt template may be input to the model. This application does not make specific limitations on this.
[0113] The entity mask prompt template masks all entities based on the combined prompt template, requiring the pre-trained language model to predict the relationship between two entities based only on contextual information and the location of the entity without considering the entity. This mechanism can effectively prevent the pre-trained language model from predicting the relationship type by "memorizing" the correspondence between a certain entity and a certain relationship (for example, preventing the model from directly judging that the relationship between entity A and entity B is a "hostile attack" relationship based solely on the entity "Cao Cao" and the entity "Liu Bei" and ignoring the contextual semantics of the ancient text where the entity "Cao Cao" and the entity "Liu Bei" are located, because in some ancient texts, there is a "superior-subordinate" relationship between Cao Cao and Liu Bei), forcing the model to mine implicit semantic clues from the context for relationship reasoning, playing a good regularization role, and significantly improving the model's effect on extracting entity relationship types in ancient texts.
[0114] In this application, by introducing the prompt learning method into the task of extracting relationships between entities in ancient Chinese texts and combining the unique semantic features of ancient Chinese texts, a prompt learning technology based on inference guidance and knowledge enhancement is designed and proposed. Aiming at the characteristics of complex and obscure semantics and concise expressions in ancient Chinese texts, this application proposes an inference-guided prompt template synthesis technology to help the entity relationship recognition model think and reason step by step, deeply understand the context semantic connotations of ancient Chinese texts, thereby significantly enhancing the attention interaction between each other, reducing information loss, and then being able to more accurately complete the task of extracting relationships between entities in ancient Chinese texts. In view of the significant semantic differences between ancient Chinese texts and modern Chinese, this application also designs a prompt enhancement strategy that integrates the semantic knowledge of ancient Chinese texts. By introducing external knowledge unique to the field of ancient Chinese, the understanding of ancient Chinese entities and relationship types by the model is improved, that is, external ancient Chinese semantic knowledge is introduced through the injection of ancient Chinese entity type knowledge and the injection of ancient Chinese relationship type knowledge to enhance and enrich the prompt information and improve the effect of relationship extraction. In addition, in order to alleviate the problems of scarce training samples of ancient Chinese texts and easy overfitting of the model, this application also designs a regularization training mechanism based on entity masking, which significantly enhances the robustness and generalization ability of the model.
[0115] To prove the progressiveness and superiority of this application, the inventors of this application also conducted experiments using the Chinese ancient historical literature information extraction dataset. The results show that the method proposed in this application significantly improves the effect of extracting relationships between entities in ancient Chinese texts, and the accuracy rate reaches 90.25%, demonstrating excellent entity relationship extraction performance.
[0116] The following introduces the device embodiments of this application, which can be used to execute the ancient Chinese text data processing method in the above embodiments of this application. For the details not disclosed in the device embodiments of this application, please refer to the embodiments of the ancient Chinese text data processing method above.
[0117] Figure 4 The block diagram of the ancient Chinese text data processing device according to an embodiment of this application is shown.
[0118] Refer to Figure 4 As shown, the ancient Chinese text data processing device 400 according to an embodiment of this application includes: an acquisition unit 401, an input unit 402, and an output unit 403.
[0119] Among them, an acquisition unit 401 is configured to acquire a target ancient book text to be processed and specified entities in the target ancient book text; an input unit 402 is configured to write the specified entities into a pre-constructed prompt template, and input the target ancient book text and the prompt template into a pre-constructed entity relationship recognition model, where the prompt template is used to guide the entity relationship recognition model to understand the relationship type between the specified entities in the target ancient book text according to a set inference logic; an output unit 403 is configured to output the relationship type between the specified entities in the target ancient book text recognized by the entity relationship recognition model.
[0120] In some embodiments of the present application, based on the foregoing solution, the prompt template includes a task definition sub-prompt template, an entity type determination sub-prompt template, and a relationship type determination sub-prompt template; wherein, the task definition sub-prompt template is used to guide the entity relationship recognition model to focus on the specified entities in the target ancient book text; the entity type determination sub-prompt template is used to guide the entity relationship recognition model to recognize the entity types of the specified entities; the relationship type determination sub-prompt template is used to guide the entity relationship recognition model to recognize the relationship type between the specified entities.
[0121] In some embodiments of the present application, based on the foregoing solution, the device further includes: a training unit, configured to construct the entity relationship recognition model through the following steps: acquire a pre-trained language model, where the pre-trained language model has the ability to understand the semantics of modern texts; acquire a pre-constructed ancient book training sample set, where each ancient book training sample in the ancient book training sample set is labeled with the actual entity type of the specified entity and the actual relationship type between the specified entities; perform iterative training on the pre-trained language model based on the ancient book training sample set to obtain the entity relationship recognition model.
[0122] In some embodiments of the present application, based on the foregoing solution, the device further includes: an ancient book domain knowledge enhancement unit, configured to, before performing iterative training on the pre-trained language model based on the ancient book training sample set, define virtual words for each entity type belonging to the ancient book text and each entity relationship type belonging to the ancient book text, where the virtual words are used to represent the semantic connotation of the corresponding entity type or the semantic connotation of the corresponding entity relationship type; determine the word embedding vectors of each virtual word, and add the word embedding vectors of each virtual word to the word embedding matrix of the pre-trained language model.
[0123] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: determine a training set and a validation set in the ancient book training sample set; train the pre-trained language model based on the training set to train the ability of the pre-trained language model to identify the entity types of specified entities in ancient book texts and the relationship types between specified entities; based on the validation set, verify the model ability of the pre-trained language model before training and the model ability of the pre-trained language model after training respectively; use the pre-trained language model with the optimal model ability as the latest pre-trained language model, and return to execute the step of training the pre-trained language model based on the training set until a preset condition is met to obtain the entity relationship recognition model.
[0124] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: write the specified entities annotated in the ancient book training samples in the training set into the prompt template; input the ancient book training samples in the training set and the prompt template into the pre-trained language model, so that the pre-trained language model can identify the entity types of the specified entities in the ancient book training samples in the training set and the relationship types between the specified entities; based on the entity types of the specified entities in the ancient book training samples in the training set identified by the pre-trained language model, the relationship types between the specified entities, the actual entity types of the specified entities annotated in the ancient book training samples in the training set, and the actual relationship types between the specified entities, update the model parameters of the pre-trained language model in reverse based on a preset loss function.
[0125] In some embodiments of the present application, based on the foregoing solution, the device further includes: a masking processing unit, configured to perform masking processing on the specified entities before inputting the ancient book training samples in the training set and the prompt template into the pre-trained language model.
[0126] As another embodiment of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, which are stored in a computer-readable storage medium and are adapted to be read and executed by a processor, so that a computer device having the processor can execute the method as described above.
[0127] As another embodiment of the present application, there is also provided a computer-readable storage medium. The computer-readable storage medium may be included in the electronic device described in the above embodiments; or it may exist alone without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by an electronic device, the electronic device can implement the method described in the above embodiments.
[0128] Based on the same inventive concept, an embodiment of the present application further provides an electronic device. Referring to Figure 5 , a schematic structural diagram of the electronic device in the embodiment of the present application is shown. The electronic device includes one or more memories 503, one or more processors 502, and at least one computer program (program code) stored on the memory 503 and executable on the processor 502. When the processor 502 executes the computer program, the method described above is implemented.
[0129] Among them, in Figure 5 , the bus architecture (represented by bus 500), the bus 500 may include any number of interconnected buses and bridges. The bus 500 links various circuits including one or more processors represented by the processor 502 and the memory represented by the memory 503 together. The bus 500 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art. Therefore, the present application will not further describe them. The bus interface 505 provides an interface between the bus 500 and the receiver 501 and the transmitter 504. The receiver 501 and the transmitter 504 may be the same element, that is, a transceiver, providing a unit for communicating with various other devices on the transmission medium. The processor 502 is responsible for managing the bus 500 and general processing, and the memory 503 may be used to store data used by the processor 502 when executing operations.
[0130] The functions described in the present application may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium. Other examples and implementations are within the scope and spirit of the present application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination of these. In addition, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.
[0131] In several embodiments provided by this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0132] The units described as separate components may or may not be physically separated. The components serving as control devices may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0134] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.< / mask> < / mask> < / mask>
Claims
1. A method for processing ancient book text data, characterized in that: The method comprises: Acquire a target ancient book text to be processed and a designated entity in the target ancient book text; Writing the designated entity into a pre-constructed prompt template, and inputting the target ancient book text and the prompt template into a pre-constructed entity relationship recognition model, wherein the prompt template is used to guide the entity relationship recognition model to understand the relationship type between the designated entities in the target ancient book text according to a set reasoning logic; Output the relationship type between the specified entities in the target ancient book text identified by the entity relationship recognition model.
2. The method according to claim 1, characterized in that The prompt template includes a task definition sub-prompt template, an entity type determination sub-prompt template, and a relationship type determination sub-prompt template; wherein, The task definition sub-prompt template is used to guide the entity relationship recognition model to focus on the specified entity in the target ancient book text; The entity type determination sub-prompt template is used to guide the entity relationship recognition model to recognize the entity type of the specified entity; The relationship type determination sub-prompt template is used to guide the entity relationship recognition model to recognize the relationship type between the specified entities.
3. The method according to claim 2, characterized in that The entity relationship recognition model is constructed by the following steps: Obtain a pre-trained language model, wherein the pre-trained language model has the ability to understand the semantics of modern text; Acquire a pre-constructed ancient book training sample set, wherein each ancient book training sample in the ancient book training sample set is annotated with an actual entity type of a specified entity and an actual relationship type between the specified entities; The pre-trained language model is iteratively trained based on the ancient book training sample set to obtain the entity relationship recognition model.
4. The method according to claim 3, characterized in that Before iteratively training the pre-trained language model based on the ancient book training sample set, the method further includes: For each entity type belonging to the ancient text, and each entity relationship type belonging to the ancient text, a virtual word is defined, wherein the virtual word is used to represent the semantic connotation of the corresponding entity type, or the semantic connotation of the corresponding entity relationship type; Determine a word embedding vector for each virtual word, and add the word embedding vector for each virtual word to the word embedding matrix of the pre-trained language model.
5. The method according to claim 4, characterized in that The iterative training of the pre-trained language model based on the ancient book training sample set to obtain the entity relationship recognition model includes: Determine a training set and a validation set in the ancient book training sample set; Training the pre-trained language model based on the training set to train the pre-trained language model to recognize entity types of specified entities in ancient texts and to recognize relationship types between specified entities; Based on the verification set, verifying the model capability of the pre-trained language model before training and the model capability of the pre-trained language model after training respectively; The pre-trained language model with the best model capability is used as the latest pre-trained language model, and the step of training the pre-trained language model based on the training set is returned to execute until the preset conditions are met to obtain the entity relationship recognition model.
6. The method according to claim 5, characterized in that The step of training the pre-trained language model based on the training set includes: Writing the designated entities annotated by the ancient book training samples in the training set into the prompt template; Inputting the ancient book training samples in the training set and the prompt template into the pre-trained language model, so that the pre-trained language model can identify the entity type of the specified entity in the ancient book training samples in the training set, and the relationship type between the specified entities; According to the entity types of the specified entities in the ancient book training samples in the training set and the relationship types between the specified entities recognized by the pre-trained language model, and the actual entity types of the specified entities marked in the ancient book training samples in the training set and the actual relationship types between the specified entities, the model parameters of the pre-trained language model are reversely updated based on a preset loss function.
7. The method according to claim 6, characterized in that The method further comprises: Before inputting the ancient book training samples in the training set and the prompt template into the pre-trained language model, mask processing is performed on the designated entity.
8. An ancient book text data processing device, characterized in that: The device comprises: An acquisition unit, used for acquiring a target ancient book text to be processed and a designated entity in the target ancient book text; An input unit, used to write the specified entity into a pre-constructed prompt template, and input the target ancient book text and the prompt template into a pre-constructed entity relationship recognition model, wherein the prompt template is used to guide the entity relationship recognition model to understand the relationship type between the specified entities in the target ancient book text according to a set reasoning logic; An output unit is used to output the relationship type between the specified entities in the target ancient book text identified by the entity relationship recognition model.
9. A computer-readable storage medium, characterized in that: At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by the processor to implement the operations performed by the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device includes one or more processors and one or more memories, wherein at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the method according to any one of claims 1 to 7.