Historical figure knowledge proofreading method and system, storage medium and electronic equipment

By building a proofreading method that combines a multivariate database and assertion detection with a large language model, the problems of redundant information and high false positive rate in the proofreading of historical figure knowledge are solved, and efficient and accurate proofreading results are achieved.

CN120611041AActive Publication Date: 2025-09-09北京蜜度信息技术有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511106815.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-09
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

When proofreading knowledge about historical figures, existing technologies contain redundant and irrelevant retrieval information, making it impossible to proofread efficiently and accurately. In addition, the dependency extraction model results in poor robustness and a high false alarm rate.

Method used

Build the first and second knowledge databases of historical figures, obtain identification fields through assertion detection, use a large language model combined with a multi-database for proofreading, eliminate interference from irrelevant content, and improve proofreading accuracy and robustness.

Benefits of technology

It achieves efficient and accurate proofreading of historical figures’ knowledge, reduces the false alarm rate, and improves the accuracy and anti-interference ability of proofreading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611041A_ABST
    Figure CN120611041A_ABST
Patent Text Reader

Abstract

The invention provides a historical character knowledge proofreading method and system, a storage medium and electronic equipment. The historical character knowledge proofreading method comprises the steps of constructing a first historical character knowledge database and a second historical character knowledge database; obtaining a to-be-proofread text, and performing assertion detection on the to-be-proofread text to obtain an assertion detection result; obtaining an identification field in the to-be-proofreading text based on the assertion detection result; according to the identification field, the corresponding historical figure is retrieved in the first historical figure knowledge database online, and related reference information of the historical figure is retrieved in the second historical figure knowledge database online; and obtaining a proofreading result of the historical character knowledge based on the historical character, the related reference information and the to-be-proofreading text. According to the method, the historical figure knowledge database is constructed offline, and in the online proofreading process, the historical figures are determined through the identification fields, and then the corresponding related reference information is retrieved, so that the problem that the retrieved information is redundant and irrelevant is effectively solved, and the proofreading efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of deep learning, and in particular relates to a method, system, storage medium and electronic device for proofreading knowledge of historical figures. Background Art

[0002] Proofreading knowledge about historical figures is a key link in ensuring the accuracy and credibility of historical information. Its importance is mainly reflected in two aspects: first, through multi-source data comparison (such as official history, archaeological discoveries, and academic papers) and logical verification (timeline conflict detection, character relationship contradiction analysis), it effectively corrects the distortion of historical figure information caused by errors in document copying, subjective interpretation bias, or modern misinterpretation, providing a reliable basis for academic research and cultural communication; second, in the digital age, automated proofreading technology can efficiently process massive amounts of historical data, prevent the spread of erroneous information in the Internet environment, and maintain the objectivity of historical narratives. For example, it can automatically identify and correct common sense errors such as "Zhang Fei was a civil servant in the Three Kingdoms period" to avoid misleading public cognition.

[0003] The application areas of historical figure knowledge proofreading span a wide range of academic, educational, and cultural sectors. In academia, it helps researchers quickly verify key information such as birth and death dates and official positions in digitized ancient texts, improving textual research efficiency. In education, it ensures the accuracy of historical figure descriptions in textbooks and museum commentaries. In cultural communication, it provides fact-checking services for film, television, and historical novels, balancing artistic creativity with historical facts. In digital humanities projects, it serves as a fundamental step in building high-precision historical knowledge graphs. For example, by resolving discrepancies in different texts regarding the number of Zheng He's voyages to the West, reliable standardized data can be generated. These applications collectively promote the scientific inheritance and innovative use of historical knowledge. When proofreading historical figure knowledge, the retrieved content is often inconsistent, containing a large amount of redundant and irrelevant information, and lacks reasoning capabilities. Therefore, how to efficiently and accurately proofread historical figure knowledge is a technical challenge that those skilled in the art are urgently seeking to address. Summary of the Invention

[0004] The purpose of this application is to provide a method, system, storage medium and electronic device for proofreading knowledge of historical figures, which can effectively avoid redundant and irrelevant retrieval information and proofread the knowledge of historical figures efficiently and accurately.

[0005] In a first aspect, the present application provides a method for proofreading knowledge about historical figures, the method comprising: Constructing a primary knowledge database of historical figures and a secondary knowledge database of historical figures; Acquire a text to be proofread, and perform assertion detection on the text to be proofread to obtain an assertion detection result; Acquire an identification field in the text to be proofread based on the assertion detection result; Searching online for a corresponding historical figure in the first historical figure knowledge database according to the identification field, and searching online for relevant reference information of the historical figure in the second historical figure knowledge database; A proofreading result of the historical figure knowledge is obtained based on the historical figure, the relevant reference information and the text to be proofread.

[0006] In one implementation of the first aspect, constructing a first knowledge database of historical figures and a second knowledge database of historical figures includes: Get data on articles about historical figures; Extracting entities and relationships from the historical figure article data as data fields, and extracting assertion expressions as data expressions; Storing the historical figure article data and the data fields as the historical figure first knowledge database; The data representation is converted into a data representation vector, and the data representation vector and the historical figure article data to which the data representation belongs are stored as the second knowledge database of the historical figure.

[0007] In an implementation of the first aspect, obtaining an identification field in the text to be proofread based on the assertion detection result includes: When the text to be proofread includes an assertion statement and the assertion detection result is passed, the identification field in the text to be proofread is extracted to proofread the text for historical figure knowledge; When the text to be proofread does not include an assertion statement, the assertion detection result is failure, and the identification field in the text to be proofread is not extracted.

[0008] In an implementation of the first aspect, obtaining the identification field in the text to be proofread includes: Extracting entities from the text to be proofread as identifying entity fields; Extracting the relationship in the text to be proofread as an identification relationship field; Extracting an assertion statement in the text to be proofread as an identification assertion field; The identification entity field, the identification relationship field and the identification assertion field are used as the identification field.

[0009] In an implementation of the first aspect, searching online for a corresponding historical figure in the first knowledge database of historical figures according to the identification field includes: Acquire a person entity based on the identification entity field and the identification relationship field; A knowledge search is performed on the first knowledge database of historical figures based on the figure entity to obtain the corresponding historical figure.

[0010] In an implementation of the first aspect, searching online for relevant reference information of the historical figure in the second knowledge database of historical figures according to the identification field includes: Performing semantic similarity detection on the identification assertion field and the data expression vector to retrieve assertion expressions related to the historical figure in the second knowledge database of historical figures; A semantic consistency check is performed on the retrieved assertion expressions and the historical figure article data, so that the assertion expressions with high confidence scores are used as the relevant reference information.

[0011] In an implementation of the first aspect, obtaining a proofreading result of the historical figure knowledge based on the historical figure, the relevant reference information, and the text to be proofread includes: The identification field, the relevant reference information and the text to be proofread are input into a preset proofreading prompt, so that the large language model proofreads the text to be proofread based on the proofreading prompt to obtain a proofreading result of historical figure knowledge.

[0012] In a second aspect, the present application provides a proofreading system for historical figure knowledge, the system comprising: A construction module, used to construct a first knowledge database of historical figures and a second knowledge database of historical figures; An acquisition module is used to acquire a text to be proofread and perform assertion detection on the text to be proofread to obtain an assertion detection result; A second acquisition module, configured to acquire an identification field in the text to be proofread based on the assertion detection result; An online search module, configured to search online for a corresponding historical figure in the first historical figure knowledge database according to the identification field, and to search online for relevant reference information of the historical figure in the second historical figure knowledge database; The proofreading module is used to obtain a proofreading result of the historical figure knowledge based on the historical figure, the relevant reference information and the text to be proofread.

[0013] In a third aspect, the present application provides an electronic device, comprising: a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned method for proofreading knowledge of historical figures.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by an electronic device, implements the above-mentioned method for proofreading knowledge of historical figures.

[0015] As described above, the method, system, storage medium, and electronic device for proofreading historical figure knowledge described in this application have the following beneficial effects: (1) This application can utilize the contextual semantic information of historical figures' articles, has semantic understanding and reasoning capabilities, and effectively improves the accuracy of proofreading.

[0016] (2) This application does not rely entirely on the relational extraction model, but instead achieves accurate proofreading based on a multivariate database, which is highly robust.

[0017] (3) This application can identify relevant historical figures and reference information based on identification fields, and perform proofreading based on key information, effectively eliminating the interference of irrelevant content on proofreading judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Shown is a flowchart of a method for proofreading knowledge of historical figures of the present application in one embodiment.

[0019] Figure 2 Shown is a flow chart of constructing a historical figure knowledge database in one embodiment.

[0020] Figure 3 Shown is a flow chart of searching for corresponding historical figures in a historical figure knowledge database according to identification fields in one embodiment.

[0021] Figure 4 Shown is a flow chart of searching for relevant reference information of historical figures in a historical figure knowledge database according to identification fields in one embodiment.

[0022] Figure 5 Shown is a schematic structural diagram of a proofreading system for historical figure knowledge in one embodiment of the present application.

[0023] Figure 6 Shown is a schematic structural diagram of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION

[0024] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0025] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.

[0026] Verifying knowledge about historical figures involves comparing authoritative data from multiple sources and conducting logical verification to ensure the accuracy and consistency of information about the person's life and deeds. Traditional verification methods for historical figure information rely on triple extraction and rule matching. Specifically, a relation extraction model is used to extract relationships within the text, such as "Li Bai - Zi - Taibai," and then the knowledge base is queried to determine if the relationship holds. This approach faces two problems: First, the rigid rule-based verification logic is complex, resulting in high development and maintenance costs. It also fails to leverage contextual semantic information, lacks semantic understanding and reasoning capabilities, and is prone to false positives. Second, it relies heavily on the effectiveness of the extraction model, completely relying on the results of the extraction model to determine relationships within the text. Any extraction errors inevitably lead to erroneous results, resulting in extremely poor robustness. Furthermore, recently emerging verification methods rely on large models and knowledge bases. These retrieval methods primarily rely on embedding models, retrieving the most relevant textual knowledge based on text semantic similarity. This retrieval approach suffers from inconsistent content, with irrelevant content significantly interfering with the large model's judgment.

[0027] In order to at least solve the above-mentioned technical problems, the present application provides a proofreading method, system, storage medium and electronic device for historical figure knowledge, which can build a historical figure knowledge database offline, extract and integrate historical figure knowledge, and in the process of online proofreading, first determine the historical figure through the identification field, and then retrieve the relevant reference information corresponding to the historical figure, effectively solving the problem of redundant and irrelevant retrieved information, and improving proofreading efficiency and accuracy.

[0028] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.

[0029] Figure 1The flowchart of the proofreading method of the historical figure knowledge provided by the embodiment of the present application is shown. Figure 1 As shown, the method for proofreading knowledge of historical figures includes steps S1 to S5.

[0030] Step S1: Construct a first knowledge database of historical figures and a second knowledge database of historical figures.

[0031] This application organizes the knowledge related to historical figures, performs knowledge cleaning, knowledge extraction and knowledge integration to build a historical figure knowledge database offline. Figure 2 The flowchart of the proofreading method of the historical figure knowledge provided by the embodiment of the present application is shown. Figure 2 As shown, building a historical figure knowledge database includes steps S11 to S14.

[0032] Step S11: Obtain historical figure article data.

[0033] Historical figure article data refers to structured or unstructured textual information centered around historical figures, typically encompassing key elements such as the figure's life story, deeds, historical context, and social relationships. This data can be sourced from a variety of sources, including ancient texts, academic papers, biographies, digital archives, and online encyclopedias. In some embodiments, historical figure article data can be obtained from highly reliable data sources, such as the Cihai (Encyclopedia of Chinese), Dabaike (Encyclopedia of Encyclopedias), or other open-source knowledge bases.

[0034] Step S12: extracting entities and relationships from the historical figure article data as data fields, and extracting assertion expressions as data expressions.

[0035] An entity refers to an objective object or concept in a text that has independent meaning and can be clearly referred to. It usually includes concrete or abstract elements such as people, places, organizations, time, and numbers. In the knowledge system of historical figures, entities are the basic nodes that constitute the information network, such as "Qin Shi Huang", "221 BC", "Qin Dynasty", etc. Its core characteristics include uniqueness (distinguishing entities with the same name through unique identifiers), typing (classification by predefined categories), and attribute association (such as the person entity is accompanied by attributes such as birth and death years, and place of origin). In some embodiments, entities can be extracted from historical figure article data through an entity extraction model. For example, the entity extraction model can be used to extract the person entities "Li Bai" and "Du Fu" and the place entity "Chang'an" from "Li Bai and Du Fu met in Chang'an".

[0036] In some embodiments, a relationship extraction model can be used to extract relationships from historical figure articles. Relationships describe semantic associations or interactions between entities and are represented by a "subject-predicate-object" triple structure. In historical figure analysis, relationships can be categorized into three types: static attributes (e.g., "Ying Zheng - was - Qin Shi Huang"), dynamic events (e.g., "Yue Fei - led - the Yue Family Army"), and abstract connections (e.g., "Confucius - influenced - Confucianism").

[0037] In some embodiments, assertion extraction models can be used to extract assertions from historical figures' articles. Assertion statements are statements in text that explicitly express a definitive judgment or claim. They typically contain strong affirmative or negative language (such as "necessarily" or "absolutely impossible") and are used to convey unquestionable opinions or conclusions. These statements possess three core attributes: high certainty (such as "the imperial examination system began in the Sui Dynasty"), verifiability (falsifiable through historical sources), and structural characteristics (often supported by data, citations, and other supporting evidence).

[0038] Step S13: storing the historical figure article data and the data fields as the historical figure first knowledge database.

[0039] In some embodiments, the historical figure article data and data fields are converted into JSON format and stored in an Elasticsearch database to serve as the historical figure first knowledge database. Specifically, the historical figure first knowledge database includes the historical figure article data in JSON format, which should include entity and relationship triples. Furthermore, assertion expressions in JSON format are also stored in the historical figure first knowledge database.

[0040] Step S14: convert the data representation into a data representation vector, and store the data representation vector and the historical figure article data to which the data representation belongs as the second knowledge database of the historical figure.

[0041] In some embodiments, the data representation is embedded, thereby converting the data representation into a continuous data representation vector, which enables mathematical calculations and similarity comparisons in a high-dimensional space by capturing the semantic or feature associations between the data. After obtaining the data representation vector, the data representation vector is stored in the Milvs database, which serves as the second knowledge database of the historical figures. At the same time, the json to which the data representation vector belongs, i.e., the historical figure article data, is recorded. Among them, the Milvus library can achieve semantic generalization, and can cope with changes and expansions at the semantic level to a certain extent, ensuring effective knowledge proofreading under different semantic expressions.

[0042] It should be noted that when forming the first knowledge database of historical figures and the second knowledge database of historical figures, the knowledge data finally stored in the database needs to be manually reviewed by experts to ensure the accuracy of the knowledge, and then form the historical figure knowledge database.

[0043] Step S2: Obtain the text to be proofread, and perform assertion detection on the text to be proofread to obtain an assertion detection result.

[0044] Step S3: Obtain an identification field in the text to be proofread based on the assertion detection result.

[0045] In some embodiments, the input sources of the text to be proofread can include various types such as user-generated historical texts, web scraped content, academic document fragments, etc. Such texts often contain errors in historical knowledge, such as misplaced relationships between characters, incorrect birth and death dates, etc.

[0046] In some embodiments, after obtaining the text to be proofread, assertion detection is performed on the text to determine whether to proofread knowledge about historical figures. For example, assertion detection can be performed using the BERT model. The BERT (Bidirectional Encoder Representations from Transformers) model, through a deep bidirectional attention mechanism, effectively captures the contextual dependencies and deterministic characteristics of assertive statements in text, enabling high-precision assertion detection. Specifically, the BERT model converts the input text to be proofread (e.g., "The imperial examination system must have begun in the Sui Dynasty") into context-sensitive word vectors. Using a fine-tuned classification head, it identifies assertive words (e.g., "must have") and their semantic strength. It also combines the semantic logic of the entire sentence to determine the degree of certainty (e.g., classifying it as "absolute assertion" or "speculative statement").

[0047] Furthermore, obtaining the identification field in the text to be proofread based on the assertion detection result includes: when the text to be proofread includes an assertion statement, the assertion detection result is passed, and the identification field in the text to be proofread is extracted to proofread the text to be proofread for knowledge of historical figures; when the text to be proofread does not include an assertion statement, the assertion detection result is failed, and the identification field in the text to be proofread is not extracted.

[0048] Furthermore, when the assertion detection result is passed, the identification fields in the text to be proofread are extracted, including extracting the entities in the text to be proofread as identification entity fields; extracting the relationships in the text to be proofread as identification relationship fields; extracting the assertion expressions in the text to be proofread as identification assertion fields; and using the identification entity field, the identification relationship field and the identification assertion field as the identification field.

[0049] Step S4: searching online for the corresponding historical figure in the first historical figure knowledge database according to the identification field, and searching online for relevant reference information of the historical figure in the second historical figure knowledge database.

[0050] In some embodiments, the present application first searches for the corresponding historical figure in the first knowledge database of historical figures through the identification field, and then performs proofreading of the relevant knowledge. Figure 3 The flowchart of the proofreading method of the historical figure knowledge provided by the embodiment of the present application is shown. Figure 3 As shown, searching for the corresponding historical figure in the historical figure knowledge database according to the identification field includes step S41 and step S42.

[0051] Step S41: Acquire a person entity based on the entity identification field and the relationship identification field.

[0052] Step S42: performing a knowledge search in the first knowledge database of historical figures based on the figure entity to obtain the corresponding historical figure.

[0053] In some embodiments, the person entities involved in the identification entity field and the identification relationship field can be extracted, and a knowledge search can be performed in the first knowledge database of historical figures based on the obtained person entities, thereby excluding irrelevant people and obtaining relevant historical figures. For example, person entities such as "Reporter Li", "Li Bai", "Professor Xu", and "Du Fu" can be extracted from the identification entity field and the identification relationship field. Based on the above person entities, a knowledge search can be performed in the first knowledge database of historical figures to obtain "Li Bai" and "Du Fu" as historical figures.

[0054] In some embodiments, after obtaining the historical figure, the present application searches online for relevant reference information of the historical figure in the second knowledge database of the historical figure in order to proofread the text to be proofread. Figure 4 The flowchart of the proofreading method of the historical figure knowledge provided by the embodiment of the present application is shown. Figure 4 As shown, online searching for relevant reference information of the historical figure in the second knowledge database of historical figures according to the identification field includes step S43 and step S44.

[0055] Step S43: Perform semantic similarity detection on the identification assertion field and the data expression vector to retrieve assertion expressions related to the historical figure in the second knowledge database of historical figures.

[0056] Semantic similarity testing quantifies the semantic proximity between two text segments by calculating the spatial distance between text vectors (such as cosine similarity or Euclidean distance). Its core goal is to transcend superficial lexical differences and capture deeper connections in meaning. In some embodiments, the BERT model can be used to convert the identification assertion field into a query vector. The cosine similarity between the query vector and the data representation vector stored in the second knowledge database of historical figures is then calculated. Based on the cosine similarity, performance representations related to the historical figure are retrieved.

[0057] In some embodiments, assertions related to historical figures can be retrieved by threshold determination. That is, when the cosine similarity between a data representation vector and a vector representation of the converted identification assertion field exceeds a set threshold, the data representation vector is considered to be an assertion related to a historical figure.

[0058] Step S44: performing semantic consistency detection on the retrieved assertion expression and the historical figure article data, so as to use the assertion expression with a high confidence score as the relevant reference information.

[0059] In some embodiments, the assertion statements retrieved from the second knowledge database of historical figures and the historical figure article data can be subjected to time conflict detection, logical contradiction analysis, and other tests. For example, it can be determined whether the assertion statements retrieved are supported by multiple historical figure article data, that is, the degree of matching with the historical figure article data, thereby calculating the confidence score, and using the assertion statements with high confidence scores as the relevant reference information.

[0060] Step S5: obtaining a proofreading result of the historical figure knowledge based on the historical figure, the relevant reference information and the text to be proofread.

[0061] After obtaining the historical figures and related reference information in the text to be proofread, the historical figures, the related reference information and the text to be proofread can be input into a preset proofreading prompt, so that the large language model proofreads the text to be proofread based on the proofreading prompt to obtain the proofreading results of the historical figure knowledge.

[0062] A Large Language Model (LLM) is a Transformer-based model containing billions (or more) of parameters. Its massive parameter count enables it to learn complex patterns and semantic relationships within massive amounts of text data, resulting in powerful semantic understanding and reasoning capabilities. For example, when processing natural language text, it can understand the logical relationships between words and sentences, accurately capture the contextual meaning, and conduct in-depth analysis and processing of the text. A Large Language Model is essentially a generative model. When given a paragraph as input, it generates a corresponding paragraph based on its learned knowledge and patterns. However, without reference input, the content generated directly by a Large Language Model based on a question may not be very accurate. This is because its generation relies heavily on its own prior knowledge, which may be affected by limitations or ambiguities in the training data. For example, for certain domain-specific or time-sensitive questions, the model's own memory alone may not be able to provide an accurate answer.

[0063] Therefore, in order to reduce the occurrence of knowledge hallucinations in large language models and to be able to efficiently and accurately proofread the knowledge of historical figures, this application uses retrieval enhancement generation technology to first search for relevant reference information based on the input text to be proofread, and then input the retrieved historical figures, relevant reference information and the text to be proofread into the preset proofreading prompt words of the large language model. In this way, the large language model can infer a proofreading result that is more in line with the context based on the determined historical figures and relevant reference information, effectively improving the accuracy of the answer and preventing the occurrence of knowledge hallucinations. In addition, because the relevant reference information is closely related to the historical figures, even if the text to be proofread has some ambiguity or irregularity in the language expression or is subject to a certain degree of noise interference, as long as the knowledge of the historical figures mentioned is consistent with the content of the relevant reference information, proofreading can be performed, thereby improving the robustness and anti-interference ability of the model.

[0064] The protection scope of the proofreading method of historical figure knowledge described in the embodiment of this application is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, subtracting, or replacing steps in the existing technology based on the principles of this application are included in the protection scope of this application.

[0065] An embodiment of the present application also provides a proofreading system for historical figure knowledge, which can implement the proofreading method for historical figure knowledge described in this application. However, the implementation device of the proofreading system for historical figure knowledge described in this application includes but is not limited to the structure of the proofreading system for historical figure knowledge listed in this embodiment. All structural deformations and replacements of the existing technology made according to the principles of this application are included in the protection scope of this application.

[0066] like Figure 5As shown, in one embodiment, the proofreading system of historical figure knowledge of the present application includes a construction module 41 , an acquisition module 42 , a second acquisition module 43 , an online search module 44 and a proofreading module 45 .

[0067] The construction module 41 is used to construct a first knowledge database of historical figures and a second knowledge database of historical figures.

[0068] The acquisition module 42 is configured to acquire the text to be proofread and perform assertion detection on the text to be proofread to obtain an assertion detection result.

[0069] The second acquisition module 43 is configured to acquire an identification field in the text to be proofread based on the assertion detection result.

[0070] The online search module 44 is used to search online for the corresponding historical figure in the first historical figure knowledge database according to the identification field, and to search online for the relevant reference information of the historical figure in the second historical figure knowledge database.

[0071] The proofreading module 45 is configured to obtain a proofreading result of the historical figure knowledge based on the historical figure, the relevant reference information and the text to be proofread.

[0072] Among them, the structures and principles of the construction module 41, the acquisition module 42, the second acquisition module 43, the online search module 44 and the proofreading module 45 correspond one to one with the steps in the proofreading method of the historical figure knowledge mentioned above, so they are not repeated here.

[0073] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices or methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules / units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules or units, which can be electrical, mechanical or other forms.

[0074] The modules / units described as separate components may or may not be physically separate, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / units in the various embodiments of the present application may be integrated into a processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into a single module / unit.

[0075] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0076] The present application also provides a computer-readable storage medium. Persons skilled in the art will appreciate that all or part of the steps in the methods of the above embodiments can be performed by instructing a processor through a program. The program can be stored in a computer-readable storage medium, which is a non-transitory medium, such as random access memory, read-only memory, flash memory, a hard disk, a solid-state drive, magnetic tape, a floppy disk, an optical disc, or any combination thereof. The storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0077] An embodiment of the present application further provides an electronic device comprising a processor and a memory.

[0078] The memory is used to store computer programs.

[0079] The memory includes various media that can store program codes, such as ROM, RAM, magnetic disk, USB flash drive, memory card or optical disk.

[0080] The processor is connected to the memory and is used to execute the computer program stored in the memory so that the electronic device executes the above-mentioned method for proofreading knowledge of historical figures.

[0081] Preferably, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0082] like Figure 6 As shown, the electronic device of the present application is implemented as a general-purpose computing device. Components of the electronic device may include, but are not limited to, one or more processors or processing units 51, a memory 52, and a bus 53 connecting different system components (including the memory 52 and the processing unit 51).

[0083] Bus 53 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0084] Electronic devices typically include a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, removable and non-removable media.

[0085] The memory 52 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 521 and / or cache memory 522. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 523 may be used to read and write non-removable, non-volatile magnetic media ( Figure 6 Not shown, usually called a "hard drive"). Although Figure 6Although not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), as well as an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to bus 53 via one or more data media interfaces. Memory 52 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present application.

[0086] A program / utility 524 having a set (at least one) of program modules 5241 may be stored, for example, in memory 52. ​​Such program modules 5241 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 5241 generally implement the functions and / or methods of the embodiments described herein.

[0087] The electronic device may also communicate with one or more external devices (e.g., a keyboard, a pointing device, a display, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via input / output (I / O) interface 54. Furthermore, the electronic device may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via network adapter 55. Figure 6 As shown, the network adapter 55 communicates with other modules of the electronic device via the bus 53. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0088] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.

Claims

1. A method for proofreading knowledge about historical figures, characterized in that: The method comprises: Constructing a primary knowledge database of historical figures and a secondary knowledge database of historical figures; Acquire a text to be proofread, and perform assertion detection on the text to be proofread to obtain an assertion detection result; Acquire an identification field in the text to be proofread based on the assertion detection result; Searching online for a corresponding historical figure in the first historical figure knowledge database according to the identification field, and searching online for relevant reference information of the historical figure in the second historical figure knowledge database; A proofreading result of the historical figure knowledge is obtained based on the historical figure, the relevant reference information and the text to be proofread.

2. The method for proofreading knowledge of historical figures according to claim 1, characterized in that: Constructing the first knowledge database of historical figures and the second knowledge database of historical figures includes: Get data on articles about historical figures; Extracting entities and relationships from the historical figure article data as data fields, and extracting assertion expressions as data expressions; Storing the historical figure article data and the data fields as the historical figure first knowledge database; The data representation is converted into a data representation vector, and the data representation vector and the historical figure article data to which the data representation belongs are stored as the second knowledge database of the historical figure.

3. The method for proofreading knowledge of historical figures according to claim 1, characterized in that: Acquiring the identification field in the text to be proofread based on the assertion detection result includes: When the text to be proofread includes an assertion statement and the assertion detection result is passed, the identification field in the text to be proofread is extracted to proofread the text for historical figure knowledge; When the text to be proofread does not include an assertion statement, the assertion detection result is failure, and the identification field in the text to be proofread is not extracted.

4. The method for proofreading knowledge of historical figures according to claim 2, characterized in that: Obtaining the identification field in the text to be proofread includes: Extracting entities from the text to be proofread as identifying entity fields; Extracting the relationship in the text to be proofread as an identification relationship field; Extracting an assertion statement in the text to be proofread as an identification assertion field; The identification entity field, the identification relationship field and the identification assertion field are used as the identification field.

5. The method for proofreading knowledge of historical figures according to claim 4, characterized in that: Searching online for a corresponding historical figure in the first historical figure knowledge database according to the identification field includes: Acquire a person entity based on the identification entity field and the identification relationship field; A knowledge search is performed on the first knowledge database of historical figures based on the figure entity to obtain the corresponding historical figure.

6. The method for proofreading knowledge of historical figures according to claim 4, characterized in that: Online searching for relevant reference information of the historical figure in the second knowledge database of historical figures according to the identification field includes: Performing semantic similarity detection on the identification assertion field and the data expression vector to retrieve assertion expressions related to the historical figure in the second knowledge database of historical figures; A semantic consistency check is performed on the retrieved assertion expressions and the historical figure article data, so that the assertion expressions with high confidence scores are used as the relevant reference information.

7. The method for proofreading knowledge of historical figures according to claim 1, characterized in that: The proofreading results of the historical figure knowledge obtained based on the historical figure, the relevant reference information and the text to be proofread include: The identification field, the relevant reference information and the text to be proofread are input into a preset proofreading prompt, so that the large language model proofreads the text to be proofread based on the proofreading prompt to obtain a proofreading result of historical figure knowledge.

8. A proofreading system for historical figure knowledge, characterized in that: The system comprises: A construction module, used to construct a first knowledge database of historical figures and a second knowledge database of historical figures; An acquisition module is used to acquire a text to be proofread and perform assertion detection on the text to be proofread to obtain an assertion detection result; A second acquisition module, configured to acquire an identification field in the text to be proofread based on the assertion detection result; An online search module, configured to search online for a corresponding historical figure in the first historical figure knowledge database according to the identification field, and to search online for relevant reference information of the historical figure in the second historical figure knowledge database; The proofreading module is used to obtain a proofreading result of the historical figure knowledge based on the historical figure, the relevant reference information and the text to be proofread.

9. An electronic device, characterized in that: The electronic device includes: a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the computer program stored in the memory, so as to enable the electronic device to execute the method for proofreading knowledge of historical figures according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by an electronic device, the method for collating knowledge of historical figures according to any one of claims 1 to 7 is realized.

Citation Information

Patent Citations

  • Historical figure information knowledge base updating method and system, medium and electronic equipment

    CN117743357A

  • Text factual proofreading method and system based on large language model

    CN119149720A

  • Event content proofreading method and device based on large model and storage medium

    CN120145075A

  • Text factuality proofreading method and system, electronic equipment and medium

    CN120278145A

  • Personal knowledge graph population from declarative user utterances

    US20170024375A1