Text factuality proofreading method and system, electronic equipment and medium

Through the combination of a multi-knowledge base and a large language model, the problems of strong dependence on hard rules and insufficient semantic understanding in the existing technology are solved, efficient and accurate text factual proofreading is achieved, false positive rates are reduced and proofreading quality is improved.

CN120278145AActive Publication Date: 2025-07-08北京蜜度信息技术有限公司

Patent Information

Application Number
CN202510772321.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The prior art relies on hard rules in text factual proofreading, resulting in high development and maintenance costs and lack of semantic understanding, high false positive rates, making it difficult to adapt to diverse scenarios.

Method used

Using a method of combining multiple knowledge bases and large language models, we recall relevant factual description information from the multiple knowledge base through search and enhancement generation technology, and use large language models for semantic proofing to break the dependence on rules and achieve intelligent and accurate text proofing.

Benefits of technology

It significantly improves the accuracy and efficiency of text proofreading, reduces the false positive rate, improves the quality and generalization of proofreading, and reduces the dependence on rules and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278145A_ABST
    Figure CN120278145A_ABST
Patent Text Reader

Abstract

The invention provides a text factuality proofreading method and system, electronic equipment and a medium. The method comprises the steps that a to-be-corrected text input by a user is acquired; based on a retrieval enhancement generation technology, recalling reference information related to the text to be corrected from a pre-constructed multivariate knowledge base; filling the text to be corrected and the reference information into a target template; and correcting factual errors existing in the text to be corrected by using the trained large language model according to the reference information filled in the target template. According to the method, reference information supplemented by the multivariate knowledge base and excellent semantic comprehension and analysis capabilities of the large language model are utilized, so that the text proofreading accuracy is remarkably improved; by cooperating with knowledge resources of the multivariate knowledge base, the accuracy and comprehensiveness of fact retrieval are improved; according to the method, the dependence on rules is avoided, text factual proofreading can be intelligently and accurately carried out, the false alarm rate is effectively reduced, and the text proofreading efficiency and quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of text processing, and relates to a method, system, electronic device and medium for text factual proofreading. Background Art

[0002] With the development of Internet technology and the improvement of social informatization, the volume of text data has shown an explosive growth. Text data, with its rich information content and wide application fields, has become an important form for people to acquire knowledge, express opinions and transmit information. However, incorrect factual statements may lead to information misleading.

[0003] Traditional text factual proofreading methods are mainly implemented based on information extraction and rule matching. The specific operation process is to first extract triples related to factuality from the text, and then proofread the extracted triples according to pre-set hard rules to determine whether the text content conforms to the facts.

[0004] However, although this method is simple and direct, it has obvious limitations in practical applications. First, the traditional rule-based method relies on hard rules, with complex proofreading logic, high development and maintenance costs, and lack of generalization, making it difficult to adapt to diverse scenarios. Second, the traditional method cannot utilize context semantic information and does not have semantic understanding and reasoning capabilities, resulting in a high false alarm rate. Especially when dealing with texts involving entity ambiguity, complex semantics or requiring deep reasoning, the error correction effect is not good. Summary of the Invention

[0005] The purpose of this application is to provide a method, system, electronic device and medium for text factual proofreading, which is used to solve the technical problems of strong rule dependence, high maintenance cost and insufficient semantic understanding in the prior art.

[0006] In a first aspect, this application provides a method for text factual proofreading, including: obtaining the text to be corrected input by the user; based on retrieval-augmented generation technology, recalling factual description information related to the text to be corrected from a pre-constructed multi-source knowledge base as reference information; the multi-source knowledge base includes a sentence-entity joint knowledge base, a triple knowledge base, and a sentence vector knowledge base; according to the semantic type of the text to be corrected, selecting a target template from a preset plurality of text proofreading prompt word templates, and filling the text to be corrected and the reference information into the target template; using a trained large language model to proofread the factual errors in the text to be corrected according to the reference information filled into the target template.

[0007] In an implementation of the first aspect, based on the retrieval-augmented generation technology, recalling the factual description information related to the text to be corrected from the pre-constructed multi-source knowledge base includes: splitting the text to be corrected into several sentences to be corrected; based on an entity extraction model, extracting entities in each of the sentences to be corrected to obtain a first entity extraction result; according to the first entity extraction result, screening out the sentences to be corrected that include at least two entities as candidate sentences to be corrected; based on a cross-knowledge-base collaborative retrieval strategy, retrieving knowledge related to the candidate sentences to be corrected from the multi-source knowledge base; and taking the retrieved knowledge as the factual description information related to the text to be corrected.

[0008] In an implementation of the first aspect, the construction process of the multi-source knowledge base includes: splitting the factual description articles stored in the document resource base into several clauses; based on an entity extraction model, extracting entities in each of the clauses to obtain a second entity extraction result; based on the second entity extraction result, screening out the clauses that include at least two entities as candidate clauses; storing the candidate clauses and the corresponding entities in a database to form a sentence-entity joint knowledge base; based on a relation extraction model, extracting relation triples of each of the candidate clauses; storing the extracted relation triples in a database to form a triple knowledge base; based on a vectorization model, performing vectorization processing on each of the candidate clauses to obtain sentence vectors of each of the candidate clauses; and storing the sentence vectors of each of the candidate clauses in a database to form a sentence vector knowledge base.

[0009] In an implementation of the first aspect, based on a cross-knowledge-base collaborative retrieval strategy, retrieving knowledge related to the candidate sentences to be corrected from the multi-source knowledge base includes: calculating the semantic similarity between the candidate sentences to be corrected and the sentence vectors of each of the candidate clauses in the sentence vector knowledge base; recalling candidate clauses with high semantic similarity to the candidate sentences to be corrected from the sentence-entity joint knowledge base; calculating the structural similarity between the candidate sentences to be corrected and the relation triples of each of the candidate clauses in the triple knowledge base; and recalling relation triples with high structural feature similarity to the candidate sentences to be corrected from the triple knowledge base.

[0010] In an implementation of the first aspect, using the trained large language model, proofreading the factual errors in the text to be corrected according to the reference information filled into the target template includes: detecting semantic conflicts between the text to be corrected and the reference information, and quantifying the degree of semantic conflict between the text to be corrected and the reference information through a confidence score; if the confidence score is lower than a preset threshold, it is determined that there are factual errors in the text to be corrected; marking the source of the conflict in the target template; correcting the factual errors in the text to be corrected, and outputting the text after factual proofreading; otherwise, it is determined that there are no factual errors in the text to be corrected.

[0011] In an implementation of the first aspect, after proofreading the factual errors in the text to be corrected, it further includes: verifying whether the entity types in the text after factual proofreading are within the specified range; if so, it is determined that the verification passes, and the text after factual proofreading is used as new reference information and refilled into the target template for secondary proofreading; otherwise, it is determined that the verification fails, and an interception process is performed in a timely manner.

[0012] In an implementation of the first aspect, the training process of the large language model includes: pre-training the large language model based on a general domain dataset; fine-tuning the pre-trained large language model based on a vertical domain dataset.

[0013] In a second aspect, the present application provides a text factual proofreading system, including: a text acquisition module for acquiring the text to be corrected input by the user; an information acquisition module for recalling factual description information related to the text to be corrected from a pre-constructed multi-source knowledge base as reference information based on retrieval-augmented generation technology; the multi-source knowledge base includes a sentence and entity joint knowledge base, a triple knowledge base, and a sentence vector knowledge base; a template filling module for selecting a target template from a preset plurality of text proofreading prompt templates according to the semantic type of the text to be corrected, and filling the text to be corrected and the reference information into the target template; a factual comparison module for using the trained large language model to proofread the factual errors in the text to be corrected according to the reference information filled into the target template.

[0014] In a third aspect, the present application provides an electronic device, including: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in any one of the above.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any one of the above is implemented.

[0016] As described above, the text factual proofreading method, system, electronic device and medium described in the present application have the following beneficial effects: (1) By utilizing the reference information supplemented by the multi-source knowledge base and the excellent semantic understanding and analysis capabilities of the large language model, a significant improvement in the accuracy of text proofreading has been achieved; (2) By collaborating with the knowledge resources of the multi-source knowledge base, the accuracy and comprehensiveness of fact retrieval have been improved; (3) It gets rid of the dependence on rules, can perform text factual proofreading more intelligently and accurately, effectively reduces the false alarm rate, and improves the efficiency and quality of proofreading. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It shows a schematic structural diagram of the mobile terminal described in the present application in an embodiment.

[0018] Figure 2 It shows a flowchart of the text factual proofreading method described in the embodiments of the present application in an embodiment.

[0019] Figure 3 It shows a schematic diagram of the principle of the retrieval enhanced generation technology described in the embodiments of the present application in an embodiment.

[0020] Figure 4 It shows a flowchart of the construction of the multi-source knowledge base described in the embodiments of the present application in an embodiment.

[0021] Figure 5 It shows a flowchart of the text factual proofreading method described in the embodiments of the present application in an embodiment.

[0022] Figure 6 It shows a schematic structural diagram of the text factual proofreading system described in the embodiments of the present application in an embodiment.

[0023] Figure 7 It shows a schematic structural diagram of the electronic device described in the embodiments of the present application in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following specific examples illustrate the implementation manners of the present application. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0025] It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0026] In addition, in the present application, descriptions such as "first" and "second" are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions conflicts with each other or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0027] As a key link in information quality control, text factual proofreading aims to ensure that the text content is highly consistent with objective facts through systematic verification and precise correction. Existing text factual proofreading methods mainly rely on structured hard rule bases when performing rule matching. Although this verification mechanism based on the fact database has clear judgment criteria, it is prone to false alarm problems caused by triple extraction deviation in actual application scenarios. For example, when proofreading the text "The Yuan Dynasty poet Sa Du La commented on Li Bai's 'Thoughts in the Silent Night'", ideally, correct triples such as "Yuan Dynasty - poet - Sa Du La", "Tang Dynasty - poet - Li Bai", and "'Thoughts in the Silent Night' - author - Li Bai" should be accurately parsed. However, due to the limitations of natural language processing technology, incorrect associations may occur. For example, the dynasty attribute may be incorrectly bound to "Yuan Dynasty - poet - Li Bai", or abnormal triples that violate temporal and spatial logic such as "Yuan Dynasty - critic - 'Thoughts in the Silent Night'" may be generated during the parsing of the comment relationship.

[0028] The following embodiments of the present application provide a text factual proofreading method, system, electronic device, and medium. The present application utilizes the reference information supplemented by the multi-source knowledge base and the excellent semantic understanding and analysis capabilities of the large language model to significantly improve the accuracy of text proofreading; by collaborating with the knowledge resources of the multi-source knowledge base, the accuracy and comprehensiveness of fact retrieval are improved; it gets rid of the dependence on rules and can perform text factual proofreading more intelligently and accurately, effectively reducing the false alarm rate and improving the efficiency and quality of proofreading.

[0029] The text factual proofreading method provided by the embodiments of the present application can run on similar devices such as mobile terminals and computer terminals. Taking running on the mobile terminal as an example, Figure 1 is the hardware structure block diagram of the mobile terminal for the text factual proofreading method, asFigure 1 As shown, the mobile terminal may include: a processor and a memory. The processor may be a central processing unit, and the memory is used to store data. Figure 1 The mobile terminal in [reference] is only for illustration and does not limit the specific structure of the mobile terminal.

[0030] Optionally, the mobile terminal may further include: a communication transmission device and an input / output device.

[0031] Optionally, the memory may be used to store computer programs, such as software programs and modules of application software. The memory may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0032] Optionally, the communication transmission device may be used to receive or send data via a network, and the network may include a wireless network provided by a communication provider of the mobile terminal. The communication transmission device may include a NIC (Network Interface Controller), which may be connected to other network devices through a base station and thus communicate with the Internet.

[0033] Next, the technical solutions in the embodiments of the present application will be described in detail with reference to the accompanying drawings in the embodiments of the present application.

[0034] Please refer to Figure 2 , which shows a flowchart of the text factual proofreading method in an embodiment of the present application. As Figure 2 shown, the embodiment of the present application provides a text factual proofreading method including the following steps S100 to S400.

[0035] In step S100, obtain the text to be corrected input by the user.

[0036] In this embodiment, the input source of the text to be corrected may cover various types such as user-created text, web-scraped content, academic literature fragments, etc. Such texts usually contain explicit factual errors, implicit logical contradictions, and ambiguities caused by vague expressions.

[0037] Specifically, the text to be corrected may contain the following six types of typical factual deviations: Spatio-temporal coordinate error. For example, "Zheng He's voyages to the Western Seas reached the Cape of Good Hope in Africa".

[0038] Mismatch in the relationship between characters. For example, "Su Shi and Lu You jointly advocated literary reform."

[0039] Inaccurate numerical precision. For example, "The total length of the Yangtze River is 5,400 kilometers."

[0040] Reversal of the order of events. For example, "The An Lushan Rebellion broke out before the establishment of the Tang Dynasty."

[0041] Confusion in the concept category. For example, classifying photosynthesis as a chemical reaction rather than a biological process.

[0042] Fallacy in the citation source. For example, "The Analects of Confucius records Mencius' thoughts."

[0043] In step S200, based on the retrieval-augmented generation technology, recall the factual description information related to the text to be corrected from the pre-constructed multi-source knowledge base as reference information; the multi-source knowledge base includes a sentence-entity joint knowledge base, a triple knowledge base, and a sentence vector knowledge base.

[0044] Retrieval-Augmented Generation (RAG) is a model that combines retrieval and generation technologies. In this embodiment, RAG provides rich background knowledge and context information for the text proofreading process, improves the quality and accuracy of the text proofreading results, and effectively prevents the generation of knowledge hallucinations.

[0045] Please refer to Figure 3 , which shows the schematic diagram of the retrieval-augmented generation technology in an embodiment described in the embodiments of the present application. As Figure 3 shown, based on the retrieval-augmented generation technology, the steps of recalling the factual description information related to the text to be corrected from the pre-constructed multi-source knowledge base include the following steps S201 to S205.

[0046] In step S201, split the text to be corrected into several sentences to be corrected.

[0047] Specifically, the splitting operation of the text to be corrected can be completed by means of a sentence segmentation model. For example, the sentence segmentation model adopted can be constructed based on punctuation rules or implemented using the natural language processing (NLP) tool Spacy. This model can reasonably decompose complex long sentences into key semantic units while retaining complete subject-predicate-object structure sentences.

[0048] In step S202, based on the entity extraction model, extract the entities in each of the sentences to be corrected to obtain the first entity extraction result.

[0049] In this embodiment, the entity extraction model has the ability to extract multiple types of entities. The entities it can extract cover a wide range, including not only common types such as person entities, work entities, and time entities, but also domain-specific entities. The domain-specific entities here specifically involve entity contents with specific domain attributes such as historical events and technical terms.

[0050] In step S203, according to the first entity extraction result, the to-be-corrected sentences including at least two entities are screened out as candidate to-be-corrected sentences.

[0051] In this embodiment, through the screening operation, semantic units without entities or with only a single entity can be filtered out. For example, for non-fact-related expressions such as "Today is really a nice day", since such sentences lack verifiable fact relevance, by filtering out such sentences, the retrieval efficiency can be significantly improved and invalid knowledge base query operations can be reduced.

[0052] In step S204, based on the cross-knowledge base collaborative retrieval strategy, knowledge related to the candidate to-be-corrected sentence is retrieved from the multiple knowledge bases.

[0053] In an embodiment of the present application, retrieving knowledge related to the candidate to-be-corrected sentence from the multiple knowledge bases based on the cross-knowledge base collaborative retrieval strategy includes: calculating the semantic similarity between the candidate to-be-corrected sentence and the sentence vectors of each candidate clause in the sentence vector knowledge base; recalling candidate clauses with high semantic similarity to the candidate to-be-corrected sentence from the sentence and entity joint knowledge base; calculating the structural similarity between the candidate to-be-corrected sentence and the relational triples of each candidate clause in the triple knowledge base; recalling relational triples with high structural feature similarity to the candidate to-be-corrected sentence from the triple knowledge base.

[0054] In step S205, the retrieved knowledge is used as factual description information related to the to-be-corrected text.

[0055] In this implementation manner, the limitation of a single knowledge base is broken, and by collaborating the knowledge resources of multiple knowledge bases, the accuracy and comprehensiveness of fact retrieval are improved.

[0056] In an embodiment of the present application, recalling factual description information related to the to-be-corrected text from a pre-constructed multiple knowledge bases based on the retrieval enhanced generation technology further includes step S206.

[0057] In step S206, the retrieved factual description information related to the to-be-corrected text is subjected to cleaning, screening, and integration processing.

[0058] Please refer to Figure 4, which shows the construction flow chart of the multi-source knowledge base in an embodiment of the present application. As Figure 4 shown, the construction process of the multi-source knowledge base includes: splitting the factual description articles stored in the document resource library into several clauses; extracting entities in each of the clauses based on an entity extraction model to obtain a second entity extraction result; screening out the clauses including at least two entities based on the second entity extraction result as candidate clauses; storing the candidate clauses and the corresponding entities in a database to form a sentence-entity joint knowledge base; extracting relation triples of each of the candidate clauses based on a relation extraction model; storing the extracted relation triples in a database to form a triple knowledge base; performing vectorization processing on each of the candidate clauses based on a vectorization model to obtain sentence vectors of each of the candidate clauses; and storing the sentence vectors of each of the candidate clauses in a database to form a sentence vector knowledge base.

[0059] In practical applications, the sentence-entity joint knowledge base can adopt an LMDB library, the triple knowledge base can adopt a Neo4j library, and the sentence vector knowledge base can adopt a Milvus library. Among them, the LMDB library can provide accurate entity-sentence pairs and provide accurate association information of entities and corresponding sentences for fact verification; the Neo4j library has the function of supporting relation reasoning, which helps to mine potential logical relations between entities, so as to more comprehensively verify facts; the Milvus library can realize semantic generalization and can, to a certain extent, cope with semantic changes and expansions to ensure effective fact verification under different semantic expressions. The cooperation of these three knowledge bases can cover various complex scenarios in the fact verification process.

[0060] It should be noted that the multi-source knowledge base involved in this embodiment can complete the construction work in an offline state, so that the overhead brought by real-time processing can be effectively avoided, and the efficiency and stability of the entire system can be improved.

[0061] In this implementation manner, by introducing retrieval-enhanced generation technology and a multi-source knowledge base, the complex maintenance work of the traditional rule base is avoided. The multi-source knowledge base covers multi-dimensional fact information from fine-grained entities to coarse-grained semantics, can quickly recall reference information related to the text to be corrected through semantic retrieval, reduces the dependence on customized rules, improves the generalization of the method, and reduces the development and maintenance costs.

[0062] In step S300, according to the semantic type of the text to be corrected, select a target template from a plurality of preset text proofreading prompt word templates, and fill the text to be corrected and the reference information into the target template.

[0063] Specifically, the semantic categories of the text to be corrected include time, location, and character relationship categories. Each semantic category corresponds to a predefined text proofreading prompt template.

[0064] This application can intelligently match or dynamically select the most suitable prompt template according to the semantic type of the text to be corrected, breaking the limitations of traditional fixed templates, achieving a problem-oriented template adaptation effect, greatly improving the accuracy of text proofreading, and optimizing the user experience.

[0065] In step S400, using the trained large language model, according to the reference information filled into the target template, proofread the factual errors in the text to be corrected.

[0066] A large language model (LLM) refers to a Transformer architecture model containing billions (or more) of parameters. Its huge parameter scale enables it to learn complex patterns and semantic relationships in a vast amount of text data, thus possessing powerful semantic understanding and reasoning capabilities. For example, when processing natural language text, it can understand the logical relationships between words and sentences, accurately capture the meaning of the context, and then conduct in-depth analysis and processing of the text content.

[0067] By combining the semantic understanding and reasoning capabilities of the large language model, this application can make full use of context information for error correction. For example, when filling the target template, the model can infer more context-compliant error correction results based on the reference information and the semantic type of the text to be corrected. Compared with traditional rule-based methods, this semantic reasoning-based error correction method can effectively solve problems of entity ambiguity (such as "Li Bai" referring to different people) and complex semantic errors (such as factual contradictions in negative sentences), significantly reducing the false alarm rate.

[0068] Essentially, a large language model is a generative model. When given a passage as input, it will generate a passage based on the knowledge and patterns it has learned. However, in the absence of reference information input, the content directly generated by the large language model according to the question may not be very accurate. This is because its generation results mainly rely on its own prior knowledge and may be affected by the limitations or ambiguities of the training data. For example, in some issues involving specific fields or strong timeliness, relying solely on the model's own memory may not be able to give accurate answers.

[0069] To make the content generated by large language models more accurate, this application fills reference information in the target template. The filling of reference information provides a more specific and targeted context for large language models, guiding them to generate answers that are more in line with the actual situation. For example, in the field of literary knowledge, if you directly ask a large language model "Li Bai is a poet of the Song Dynasty", due to possible knowledge confusion or inaccurate memory of some details in the model, it may answer correctly. When supplementary reference information "It is known that Li Bai's dynasty is the Tang Dynasty" is added and the same question is then asked, the model will make a judgment based on the supplementary reference information and thus give the correct answer.

[0070] In addition, due to the stability and authority of reference information, even if there are some ambiguities, non-standardities in language expression or a certain degree of noise interference in the text to be corrected, as long as its core factual content is consistent with the reference information, the model can accurately judge, thereby improving the robustness and anti-interference ability of the model. For example, there may be some colloquial expressions or a small number of typos in an article, but as long as the facts conveyed are consistent with the reference information, the model will not wrongly determine that the text has factual errors because of these minor issues, ensuring that the model can make accurate judgments in different situations.

[0071] In an embodiment of this application, the training process of the large language model includes: pre-training the large language model based on a general domain dataset; and fine-tuning the pre-trained large language model based on a vertical domain dataset.

[0072] In this embodiment, the general domain dataset usually covers a wide range of topics and rich language expression forms, such as large-scale network texts, books, news reports, etc. The large language model automatically mines statistical laws and semantic information in language through an unsupervised learning method.

[0073] The vertical domain dataset refers to being focused on a specific discipline or application scenario. For example, in the medical field, the vertical domain dataset may include a large number of medical literature, case reports, professional term explanations, etc.; in the legal field, there will be relevant data such as laws and regulations articles, judicial case analyses, etc. The vertical domain data can provide more specific and accurate knowledge information for the model, making the model more familiar with the specific concepts, terms and facts in this field.

[0074] In this implementation method, a two-stage training paradigm is adopted. By performing model pre-training on a huge amount of data, world knowledge can be compressed into the parameters of the large language model, enabling the large language model to have good general basic capabilities and downstream task migration capabilities; by performing model fine-tuning on vertical domain data, the parameters of the large language model can be made most suitable for solving the factual proofreading task based on reference information in the current application scenario.

[0075] See also Figure 5 , which is a flow chart of another embodiment of the text factual proofreading method described in the embodiment of the present application.

[0076] like Figure 5 As shown, using the trained large language model, based on the reference information filled in the target template, proofreading the factual errors existing in the text to be corrected includes: performing semantic conflict detection on the text to be corrected and the reference information, and quantifying the degree of semantic conflict between the text to be corrected and the reference information through a confidence score; if the confidence score is lower than a preset threshold, it is determined that there are factual errors in the text to be corrected; marking the source of the conflict in the target template; correcting the factual errors in the text to be corrected, and outputting the fact-proofread text; otherwise, it is determined that there are no factual errors in the text to be corrected.

[0077] In the use of reference information, the large language model has demonstrated its wide applicability and strong generalization ability. It can not only refer to structured triple information, but also flexibly respond to the input of related sentences. In contrast, the triple extraction model in the prior art is relatively rigid, and a strict framework needs to be defined in advance to determine which types of rules can be extracted. If certain key information is not defined in the framework, such as the dynasty in which the poet lives or representative works, the model will not be able to extract relevant triples to support factual proofreading. The large language model is different. As long as the reference information contains relevant content, regardless of whether the content is pre-defined in the framework, the model can make accurate judgments with its powerful semantic understanding and reasoning capabilities, thereby providing users with more comprehensive and accurate factual proofreading services.

[0078] In one embodiment of the present application, after proofreading the factual errors existing in the text to be corrected, it also includes: verifying whether the entity type in the text after the factual proofreading is within the specified range; if so, it is determined that the verification is passed, and the text after the factual proofreading is used as new reference information and re-filled into the target template for secondary proofreading; otherwise, it is determined that the verification has failed, and interception processing is performed in a timely manner.

[0079] In this implementation, through an effective verification and interception mechanism, erroneous judgments caused by entity type mismatch are avoided, the accuracy and reliability of the proofreading results are guaranteed, the stability and credibility of the entire proofreading system are improved, and users are provided with more accurate and high-quality text proofreading services.

[0080] It should be noted that the protection scope of the text factual proofreading method described in the embodiments of the present application is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or reducing steps of the prior art and replacing steps according to the principle of the present application is included in the protection scope of the present application.

[0081] Please refer to Figure 6 , which shows the structural schematic diagram of the text factual proofreading system described in the embodiments of the present application in one embodiment. As Figure 6 shown, the embodiments of the present application provide a text factual proofreading system, including a text acquisition module, an information acquisition module, a template filling module, and a factual comparison module.

[0082] The text acquisition module is used to acquire the text to be corrected input by the user.

[0083] The information acquisition module is used to recall the factual description information related to the text to be corrected from a pre-constructed multi-source knowledge base based on retrieval augmentation generation technology as reference information; the multi-source knowledge base includes a sentence-entity joint knowledge base, a triple knowledge base, and a sentence vector knowledge base.

[0084] The template filling module is used to select a target template from a preset plurality of text proofreading prompt word templates according to the semantic type of the text to be corrected, and fill the text to be corrected and the reference information into the target template.

[0085] The factual comparison module is used to proofread the factual errors in the text to be corrected by using a trained large language model according to the reference information filled into the target template.

[0086] It should be noted that the structures and principles of the text acquisition module, information acquisition module, template filling module, and factual comparison module in the embodiments of the present application correspond one by one to the steps in the above text factual proofreading method, so they will not be elaborated here.

[0087] The text factual proofreading system provided by the embodiments of the present application can implement the text factual proofreading method described in the present application. However, the implementation devices of the text factual proofreading method described in the present application include but are not limited to the structures of the text factual proofreading system listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principle of the present application are included in the protection scope of the present application.

[0088] Please refer to Figure 7 , which shows the structural schematic diagram of the electronic device described in the embodiments of the present application in one embodiment.

[0089] As Figure 7 shown, the embodiments of the present application provide an electronic device, including: a processor and a memory.

[0090] The memory is used to store a computer program; The processor is configured to execute the computer program stored in the memory, so that the electronic device performs the method described in any one of the above.

[0091] Preferably, the processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disc.

[0092] This embodiment further includes one or more of a multimedia component, an input / output (I / O) interface, and a communication component.

[0093] The multimedia component may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals. The received audio signals may be further stored in the memory or sent through the communication component. The audio component further includes at least one speaker for outputting audio signals. The I / O interface provides an interface between the processor and other interface modules, and the other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component is used for the timer to communicate with other devices in a wired or wireless manner. The wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them. Accordingly, the communication component may include: a Wi-Fi module, a Bluetooth module, and an NFC module.

[0094] In several embodiments provided in this application, it should be understood that the disclosed system, apparatus or method may be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces, and the indirect coupling or communication connection of devices or modules or units may be in an electrical, mechanical or other form.

[0095] The modules / units described as separate components may or may not be physically separated, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of this application. For example, in each embodiment of this application, the functional modules / units may be integrated in a processing module, or each module / unit may exist physically alone, or two or more modules / units may be integrated in one module / unit.

[0096] Those of ordinary skill in the art should also be able to further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0097] The embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in any one of the above is implemented. Those of ordinary skill in the art can understand that all or part of the steps in the method of implementing the above embodiments can be completed by instructing the processor through a program. The described program can be stored in a computer-readable storage medium. The storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid-state disk (SSD)).

[0098] The embodiments of this application can also provide a computer program product, which includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the processes or functions described in the embodiments of this application are fully or partially generated. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, or data center to another website, computer, or data center in a wired manner (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.).

[0099] When the computer program product is executed by a computer, the computer executes the method described in the foregoing method embodiments. The computer program product may be a software installation package. In the case where the foregoing method needs to be used, the computer program product can be downloaded and executed on the computer.

[0100] The descriptions of the processes or structures corresponding to the foregoing respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.

[0101] The foregoing embodiments merely illustrate the principles and effects of the present application, rather than limiting the present application. Any person familiar with this technology can modify or change the foregoing embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present application should still be covered by the claims of the present application.

Claims

1. A text factual proofreading method, characterized in that, Including: Obtain the text to be corrected entered by the user; Based on the retrieval-augmented generation technology, recall the factual description information related to the text to be corrected from a pre-constructed multi-source knowledge base as reference information; the multi-source knowledge base includes a sentence-entity joint knowledge base, a triple knowledge base, and a sentence vector knowledge base; According to the semantic type of the text to be corrected, select a target template from a preset plurality of text proofreading prompt word templates, and fill the text to be corrected and the reference information into the target template; Use the trained large language model to proofread the factual errors in the text to be corrected according to the reference information filled in the target template.

2. The method according to claim 1, characterized in that Based on the retrieval-augmented generation technology, recalling the factual description information related to the text to be corrected from a pre-constructed multi-source knowledge base includes: Split the text to be corrected into several sentences to be corrected; Based on an entity extraction model, extract the entities in each sentence to be corrected to obtain a first entity extraction result; According to the first entity extraction result, screen out the sentences to be corrected that include at least two entities as candidate sentences to be corrected; Based on a cross-knowledge base collaborative retrieval strategy, retrieve the knowledge related to the candidate sentences to be corrected from the multi-source knowledge base; Use the retrieved knowledge as the factual description information related to the text to be corrected.

3. The method according to claim 1, wherein The construction process of the multi-source knowledge base includes: Split the factual description articles stored in the document resource library into several clauses; Based on an entity extraction model, extract the entities in each clause to obtain a second entity extraction result; Based on the second entity extraction result, screen out the clauses that include at least two entities as candidate clauses; Store the candidate clauses and the corresponding entities in a database to form the sentence-entity joint knowledge base; Based on a relation extraction model, extract the relation triples of each candidate clause; Store the extracted relation triples in a database to form the triple knowledge base; Based on a vectorization model, perform vectorization processing on each candidate clause to obtain the sentence vectors of each candidate clause; Store the sentence vectors of each candidate clause in a database to form the sentence vector knowledge base.

4. The method according to claim 2, wherein Based on a cross-knowledge base collaborative retrieval strategy, retrieving the knowledge related to the candidate sentences to be corrected from the multi-source knowledge base includes: Calculate the semantic similarity between the candidate sentences to be corrected and the sentence vectors of each candidate clause in the sentence vector knowledge base; Recall the candidate clauses with high semantic similarity to the candidate sentences to be corrected from the sentence-entity joint knowledge base; Calculate the structural similarity between the candidate sentences to be corrected and the relation triples of each candidate clause in the triple knowledge base; Recall the relation triples with high structural feature similarity to the candidate sentences to be corrected from the triple knowledge base.

5. The method according to claim 1, wherein Using the trained large language model to proofread the factual errors in the text to be corrected according to the reference information filled in the target template includes: Perform semantic conflict detection on the text to be corrected and the reference information, and quantify the degree of semantic conflict between the text to be corrected and the reference information through a confidence score; If the confidence score is lower than the preset threshold, it is determined that there are factual errors in the text to be corrected; mark the conflict source in the target template; correct the factual errors in the text to be corrected, and output the text after factual proofreading; Otherwise, it is determined that there are no factual errors in the text to be corrected.

6. The method according to claim 1, wherein After proofreading the factual errors in the text to be corrected, it further includes: Verify whether the entity types in the text after factual proofreading are within the specified range; If so, it is determined that the verification is passed, and the text after factual proofreading is used as new reference information and refilled into the target template for secondary proofreading; Otherwise, it is determined that the verification fails, and interception processing is performed in a timely manner.

7. The method according to claim 1, characterized in that, The training process of the large language model includes: Pre-train the large language model based on a general domain dataset; Fine-tune the pre-trained large language model based on a vertical domain dataset.

8. A text fact-checking system, characterized in that, It includes: A text acquisition module for acquiring the text to be corrected input by the user; An information acquisition module for recalling factual description information related to the text to be corrected from a pre-constructed multi-source knowledge base based on retrieval-augmented generation technology as reference information; the multi-source knowledge base includes a sentence and entity joint knowledge base, a triple knowledge base, and a sentence vector knowledge base; A template filling module for selecting a target template from a preset plurality of text proofreading prompt templates according to the semantic type of the text to be corrected, and filling the text to be corrected and the reference information into the target template; A factual comparison module for using the trained large language model to proofread the factual errors in the text to be corrected according to the reference information filled into the target template.

9. An electronic device, characterized in that, It includes: A processor and a memory; The memory is used to store a computer program; The processor is used to execute the computer program stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Legal cognition method and device based on multi-level multi-dimension semantic comprehension and medium

    CN108073569A

  • Intelligent question answering method and device based on medical knowledge graph

    CN113505243A

  • Knowledge proposition error correction method and system based on knowledge graph

    CN118709695A

  • Text factual proofreading method and system based on large language model

    CN119149720A

  • Large model text generation method and system based on adaptive cue words

    CN119623475A

Cited By

  • Historical figure knowledge proofreading method and system, storage medium and electronic equipment

    CN120611041A

  • File digital governance method and system based on large model

    CN121092759A

  • Search system updating method and device based on large language model, equipment and storage medium

    CN121116353A

  • Ffact index database construction method, fact error correction method and system, terminal and medium

    CN122470682A