Event content proofreading method and device based on large model and storage medium
Through the event content proofreading method based on the big model, the problem of event argument extraction deviation in the existing technology is solved, and the accuracy of event content verification is improved. By building an event knowledge base and retrieval mechanism, the accuracy of event content verification is improved.
Patent Information
- Application Number
- CN202510633845.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
When facing word ambiguity and complex grammatical scenarios, it is difficult for the prior art to accurately distinguish the semantics of words in different events, resulting in event argument extraction bias and low accuracy of event content verification.
The event content proofreading method based on the big model is used to identify events in the target text through the big model, rewritten according to the target rewriting rules, forming reference events with unified structure and clear semantics, and storing them in the event knowledge base. Based on the event to be proofreaded, the associated target reference events are retrieved from the event knowledge base, and the big model is compared to improve the verification accuracy.
Through the semantic understanding ability of the big model, events can be accurately identified and rewrited, a high-quality event knowledge base is built, event argument extraction bias is reduced, and event content verification is improved.
Smart Images

Figure CN120145075A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular, relates to a method, device and storage medium for event content proofreading based on a large model. Background Art
[0002] Event content verification is a technology that extracts key event information from text and conducts comparative analysis of events in different texts. Its core lies in accurately identifying event types, extracting event elements, and establishing associations. In digital publishing and intelligent auditing scenarios, event content verification can be used to verify event consistency across documents and identify illegal information or false content.
[0003] In related technologies, event content verification usually uses structured event extraction technology. Structured event extraction mainly relies on traditional natural language processing technologies such as trigger word detection, entity recognition, and dependency syntax analysis to identify event types from unstructured texts and extract relevant event arguments for comparison. However, when faced with the ambiguity of words and complex grammatical scenarios, the above methods are difficult to accurately distinguish the semantics of these words in different events, resulting in deviations in event argument extraction and low accuracy in event content verification. Summary of the invention
[0004] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a method, device and storage medium for event content verification based on a large model to improve the accuracy of event content verification.
[0005] In a first aspect, the present application provides a method for event content proofreading based on a large model, comprising: Identify events in a target text through a large model, and rewrite the events in the target text according to target rewriting rules to obtain reference events; the target text is standard data that has been reviewed; Storing the reference event in an event knowledge base; Retrieving a target reference event associated with the event to be proofread from the event knowledge base based on the event to be proofread; The event to be checked is compared with a target reference event through the large model to obtain a comparison result indicating whether the event to be checked is correct.
[0006] The event content proofreading method based on a large model provided by an embodiment of the present application identifies events in a target text through the large model, and rewrites the events in the target text according to target rewriting rules to obtain reference events; the target text is standard data that has been proofread; stores the reference events in an event knowledge base; retrieves a target reference event associated with the event to be proofread from the event knowledge base based on the event to be proofread; compares the event to be proofread and the target reference event through the large model to obtain a comparison result indicating whether the event to be proofread is correct. The embodiment of the present application utilizes the powerful semantic understanding ability of the large model, can accurately identify events in the standard data, and performs rewriting processing on the events to form reference events with unified structure and clear semantics, thereby being able to build a high-quality event knowledge base, reducing the problem of deviation in event argument extraction in related technologies. When proofreading the event to be proofread, the target reference event can be retrieved from the event knowledge block according to relevance, and then compared with the help of the large model, improving the accuracy of event content verification.
[0007] According to an embodiment of the present application, the method further includes: Obtain data to be proofread; Identify events in the data to be proofread through the large model, and rewrite the events in the data to be proofread according to the target rewriting rules to obtain events to be proofread.
[0008] In this embodiment, by utilizing the powerful semantic understanding ability of the large model, events in the data to be proofread can be accurately identified, so as to accurately extract key event information, rewrite the identified events according to the target rewriting rules, and convert the original events into events to be proofread with unified structure and clear semantics, realizing the standardization of event expressions, facilitating subsequent comparison, reducing misjudgment caused by expression differences, and further improving the accuracy of the verification result.
[0009] According to an embodiment of the present application, the method further includes: Vectorize the reference events to obtain a vector representation of the reference events; Store the vector representation in the event knowledge base.
[0010] In this embodiment, through the vectorization processing of the reference events, complex event information can be converted into a numerical vector representation that is easy for a computer to process, reducing the dimension of the data and optimizing the data storage structure. When facing an event to be proofread, the similarity between the event to be proofread and the reference event vector can be quickly calculated based on the vector-based retrieval method. Compared with the traditional text matching method, the relevant target reference event can be more accurately located from a large amount of event data.
[0011] According to an embodiment of the present application, before identifying an event in the target text through a large model, it includes: Preprocess the target file; the preprocessing includes: splitting the target file by chapter, and supplementing the markup symbols and / or annotation content in the target file into the original text to obtain formatted data corresponding to the target file; the formatted data includes the article name, chapter title, and chapter content.
[0012] In this embodiment, by splitting the target file by chapter, a long text can be disassembled into smaller, more manageable sub-units, which not only reduces the difficulty for the large model to process complex text, but also helps to accurately locate the specific context where the event occurs, reducing the problem of missing key event information due to overly long text. Moreover, supplementing the markup symbols or annotation content into the original text can fill the information gaps that may affect semantic understanding, making the text semantics more complete and coherent, forming formatted data including the article name, chapter title, and chapter content, providing structured input for the large model, helping the large model to more efficiently understand the text logic and semantic relationships, and thus more accurately identify the event type and extract event elements.
[0013] According to an embodiment of the present application, before storing the reference event in the event knowledge base, it includes: Verify whether the reference event is consistent with the event in the target text based on the large model; In the case where the reference event is inconsistent with the event in the target text, modify the reference event so that the modified reference event is consistent with the event in the target text.
[0014] In this embodiment, before storing the reference event in the knowledge base, the large model is used to verify the consistency between the reference event and the original event in the target text, capturing differences in the expression, elements, etc. of the event, reducing the problem of incorrect or deviated information being stored in the event knowledge base. When it is found that the reference event is inconsistent with the original event, the reference event is modified by the large model, so that the reference events stored in the event knowledge base are all highly accurate.
[0015] According to an embodiment of the present application, identifying an event in the target text through a large model and rewriting the event in the target text according to the target rules includes: Construct a prompt instruction for the large model; the prompt instruction includes the target rewriting rule; the target rewriting rule includes: identifying event content from the article body of the target file, complementing the event elements of the event content based on the article name, chapter title, and chapter content of the target file and performing coreference resolution, and simplifying the event content. Input the prompt instruction into the large model, so that the large model can identify the events in the target text according to the prompt instruction, and rewrite the events in the target text according to the target rewriting rules to obtain the reference events.
[0016] In this embodiment, by constructing a prompt instruction containing the target rewriting rules, the direction and standard of event processing are specified for the large model. The event content is identified from the article text, and the event elements are complemented and disambiguated based on the article name, chapter title, and chapter content. The context information of the text is utilized to reduce problems such as word ambiguity and unclear reference, making the event elements more complete and accurate. Simplifying the event content removes redundant information and highlights the core elements, making the event expression clearer and more concise.
[0017] According to an embodiment of the present application, retrieving the target reference event associated with the event to be proofread from the event knowledge base includes: Performing keyword matching from the event knowledge base according to the event to be proofread to obtain a first retrieval result; the first retrieval result includes multiple first candidate reference events sorted based on the matching degree; Vectorize the event to be proofread, and perform similarity matching from the event knowledge base based on the vectorized event to be proofread to obtain a second retrieval result; the second retrieval result includes multiple second candidate reference events sorted based on similarity; Assign weights to the first candidate reference events and the second candidate reference events based on the event element type; Screen the target reference event from the first candidate reference events and the second candidate reference events according to the weights, matching degrees, and similarities.
[0018] In this embodiment, by using keyword matching to obtain the first retrieval result, it is possible to quickly retrieve the preliminary candidate reference events related to the event to be proofread from the event knowledge base. After vectorizing the event to be proofread, similarity matching is performed to obtain the second retrieval result. The vector calculation is used to deeply mine similar events from the semantic level, making up for the possible lack of semantic understanding in keyword matching. Further, weights are assigned to the candidate reference events based on the event element type, and the key components of the event are considered differently according to their importance. The target reference event is screened from the candidate reference events, which not only reduces the limitation problem of a single retrieval method but also highlights the key elements through weight assignment, improving the accuracy of retrieving the target reference event.
[0019] In a second aspect, the present application provides an event content proofreading device based on a large model. The device includes: An identification and rewriting module, configured to identify events in the target text through a large model, and rewrite the events in the target text according to target rewriting rules to obtain reference events; the target text is the reviewed standard data; A storage module, configured to store the reference events in an event knowledge base; A retrieval module, configured to retrieve target reference events associated with the event to be verified from the event knowledge base based on the event to be verified; A verification module, configured to compare the event to be verified and the target reference events through the large model to obtain a comparison result indicating whether the event to be verified is correct.
[0020] The event content verification method based on a large model provided by the embodiments of the present application identifies events in the target text through the large model, and rewrites the events in the target text according to target rewriting rules to obtain reference events; the target text is the reviewed standard data; stores the reference events in an event knowledge base; retrieves target reference events associated with the event to be verified from the event knowledge base based on the event to be verified; compares the event to be verified and the target reference events through the large model to obtain a comparison result indicating whether the event to be verified is correct. The embodiments of the present application utilize the powerful semantic understanding ability of the large model to accurately identify events in the standard data, and perform rewriting processing on the events to form reference events with unified structure and clear semantics, thereby being able to construct a high-quality event knowledge base, reducing the problem of event argument extraction deviation in related technologies. When verifying the event to be verified, the target reference event can be retrieved from the event knowledge block according to the relevance, and then compared with the help of the large model, improving the accuracy of event content verification.
[0021] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the event content verification method based on a large model as described in the first aspect above is implemented.
[0022] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the event content verification method based on a large model as described in the first aspect above is implemented.
[0023] In a fifth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the event content verification method based on a large model as described in the first aspect above is implemented.
[0024] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects: The event content proofreading method based on a large model provided by an embodiment of the present application identifies events in a target text through the large model, and rewrites the events in the target text according to target rewriting rules to obtain reference events; the target text is standard data that has been proofread; stores the reference events in an event knowledge base; retrieves target reference events associated with the event to be proofread from the event knowledge base based on the event to be proofread; and compares the event to be proofread and the target reference event through the large model to obtain a comparison result indicating whether the event to be proofread is correct. The embodiment of the present application utilizes the powerful semantic understanding ability of the large model to accurately identify events in the standard data, and rewrite the events to form reference events with a unified structure and clear semantics, so as to be able to build a high-quality event knowledge base, reduce the problem of event argument extraction deviation in related technologies, and when proofreading the event to be proofread, the target reference event can be retrieved from the event knowledge block according to the relevance, and then compared with the help of the large model, improving the accuracy of event content verification.
[0025] Furthermore, by utilizing the powerful semantic understanding ability of the large model, events in the data to be proofread can be accurately identified, so as to accurately extract the key information of the events, rewrite the identified events according to the target rewriting rules, and convert the original events into events to be proofread with a unified structure and clear semantics, realizing the standardization of event expressions, facilitating subsequent comparison, reducing misjudgment caused by expression differences, and further improving the accuracy of the verification result.
[0026] Furthermore, by vectorizing the reference events, complex event information can be converted into a numerical vector representation that is easy for a computer to process, reducing the dimension of the data and optimizing the data storage structure, so that when facing the event to be proofread, the vector-based retrieval method can quickly calculate the similarity between the vector of the event to be proofread and the reference event vector. Compared with the traditional text matching method, it can more accurately locate the relevant target reference event from a large amount of event data.
[0027] Even further, by splitting the target file into chapters, a long text can be disassembled into sub-units that are easier to process, which not only reduces the difficulty of the large model in processing complex texts, but also helps to accurately locate the specific context where the event occurs, reducing the problem of missing key event information due to the overly long text. Moreover, by supplementing the marked symbols or annotation content into the original text, the information gap that may affect semantic understanding can be filled, making the text semantics more complete and coherent, forming formatted data including the article name, chapter title, and chapter content, providing structured input for the large model, helping the large model to more efficiently understand the text logic and semantic relationship, and thus more accurately identify the event type and extract event elements.
[0028] Further, before storing the reference event in the knowledge base, the large model is used to perform consistency verification on the reference event and the original event in the target text, capture the differences in the expression, elements, etc. of the event, reduce the problem of incorrect or deviated information being stored in the event knowledge base. When it is found that the reference event is inconsistent with the original event, the large model is used to modify the reference event so that the reference events stored in the event knowledge base all have high accuracy.
[0029] Furthermore, by constructing a prompt instruction containing the target rewriting rules, the large model is pointed to the direction and standard of event processing. The event content is identified from the article text, and the event elements are complemented and disambiguated based on the article name, chapter title, and chapter content. By using the text context information, problems such as word ambiguity and unclear reference are reduced, making the event elements more complete and accurate; simplifying the event content removes redundant information and highlights the core elements, making the event expression clearer and more concise.
[0030] Furthermore, by using keyword matching to obtain the first retrieval result, the preliminary candidate reference events related to the event to be proofread can be quickly retrieved from the event knowledge base. After vectorizing the event to be proofread, similarity matching is performed to obtain the second retrieval result. By using vector calculation, similar events are deeply mined from the semantic level, making up for the problem of insufficient semantic understanding that may exist in keyword matching. Further, weights are assigned to the candidate reference events based on the event element type, and differential consideration is carried out according to the importance of the key components of the event. The target reference events are screened from the candidate reference events, which not only reduces the limitation problem of a single retrieval method but also highlights the key elements through weight assignment, improving the accuracy of target reference event retrieval.
[0031] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 is a schematic flowchart of the method for proofreading event content based on a large model provided by an embodiment of the present application; Figure 2 is an example of the process of event identification and rewriting provided by an embodiment of the present application; Figure 3It is a schematic diagram of the event comparison process provided by the embodiments of the present application; Figure 4 It is an example of the event proofreading process provided by the embodiments of the present application; Figure 5 It is a schematic structural diagram of an event content proofreading device based on a large model provided by the embodiments of the present application; Figure 6 It is a schematic structural diagram of an electronic device provided by the embodiments of the present application. Specific embodiments
[0034] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.
[0035] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the associated objects before and after.
[0036] Next, in conjunction with the accompanying drawings, the method, device, and storage medium for event content proofreading based on a large model provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.
[0037] Among them, the method for event content proofreading based on a large model can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.
[0038] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablet computers with a touch-sensitive surface (for example, a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer with a touch-sensitive surface (for example, a touch screen display and / or a touchpad).
[0039] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.
[0040] The method for event content proofreading based on a large model provided by the embodiments of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the functions of the method. The electronic devices mentioned in the embodiments of the present application include, but are not limited to, mobile phones, tablet computers, computers, cameras, wearable devices, etc. Hereinafter, taking the electronic device as the execution subject as an example, the method for event content proofreading based on a large model provided by the embodiments of the present application will be described.
[0041] As Figure 1 shown, the method for event content proofreading based on a large model includes: step 110, step 120, step 130, and step 140.
[0042] Step 110: Identify the events in the target text through a large model, and rewrite the events in the target text according to the target rewriting rules to obtain reference events; the target text is the standard data that has been proofread.
[0043] A large language model (LLM) is a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. The large model can handle various natural language tasks, such as text generation, sentiment analysis, question answering systems, etc. In some embodiments, the large model adopted may be Tongyi Qianwen, ChatGLM, LLaMA, etc., or other large models may also be adopted, and the embodiments of this specification do not limit this.
[0044] In the scope of natural language processing, an event refers to a certain behavior performed or a certain situation that occurs by specific participants at a specific time and place. An event usually includes multiple event elements, such as the subject of the event (who participated in the event), the action of the event (what behavior occurred), the time and place of the event, etc. For example, "Xiaoming read a novel in the library yesterday afternoon", where "Xiaoming" is the subject of the event, "read" is the action of the event, "yesterday afternoon" is the time, "library" is the place, and "a novel" is the object of the event. These event elements together constitute a complete event.
[0045] In the embodiments of the present application, the target text is a text that has been strictly proofread and can be considered as standard data or correct data. The target text can be a published book, which has gone through multiple links such as editing and proofreading, and has high content quality. The target text can also be a published article, which has gone through processes such as review before publication to ensure the accuracy and reliability of the content. The target text can be various formats of data, such as e-books, word documents, etc.
[0046] In some embodiments, the target file can be preprocessed; the preprocessing includes: splitting the target file by chapter, and supplementing the markup symbols and / or annotation content in the target file into the original text to obtain the formatted data corresponding to the target file; the formatted data includes the article name, chapter titles, and chapter content.
[0047] In the target text, whether it is a published book or a published article, it often has a relatively large length and a complex structure. Since chapters themselves are important markers for the logical division of text content, and each chapter focuses on a specific theme for discussion or narration, using chapters as the basis for splitting can break down the overall text into relatively independent parts, enabling the large model to focus on the content under a specific theme when identifying events later and reducing the problem of information mixing caused by overly long text. For example, in a historical-themed book, different chapters may respectively describe historical events in different periods. After splitting by chapter, the large model can identify events for the content of each period, thereby improving the recognition efficiency and accuracy.
[0048] The target file usually can also include content such as markup symbols and annotations. Among them, markup symbols are usually used in the text to emphasize, separate, or explain certain content, while annotations are supplementary explanations or interpretations of the main text content. Therefore, markup symbols and annotations are very important for understanding the complete semantics of the text. In the unprocessed text, markup symbols and annotations may be processed or ignored separately. Identifying the markup symbols, annotations, and other content in the target file and supplementing the markup symbols, annotations, and other content into the original text can restore the complete intention during text creation and enable the large model to obtain more complete information. For example, in an academic paper, the annotations contain important reference documents, detailed descriptions of research methods, and other content. Supplementing these annotations into the main text, the large model can better understand the research background and relevant details when identifying events, reducing the problem of misjudgment of events caused by information loss. Another example is that some special markup symbols may represent specific concepts or event boundaries. Supplementing the annotation content into the original text helps the large model more accurately define the event scope and extract complete event elements.
[0049] After the above preprocessing, the formatted data corresponding to the target file can be obtained. The formatted data includes the article name, chapter titles, and chapter content. The article name can clarify the theme and source of the text, facilitating the distinction and management of different texts; the chapter titles further refine the content theme of the text, enabling the large model to quickly understand the core content of each chapter; the chapter content is the main part for event identification and processing.
[0050] In this embodiment, by splitting the target file by chapter, the long text can be disassembled into smaller, more manageable sub-units. This not only reduces the difficulty for the large model to process complex text, but also helps to accurately locate the specific context in which an event occurs, reducing the problem of missing key event information due to overly long text. Moreover, by supplementing the original text with markup symbols or annotation content, the information gap that may affect semantic understanding can be filled, making the text semantics more complete and coherent, forming formatted data that includes the article name, chapter titles, and chapter content, providing structured input for the large model, helping the large model to more efficiently understand the text logic and semantic relationships, and thus more accurately identify event types and extract event elements.
[0051] In some embodiments, the target text can be input into a large model. The large model identifies the events in the target text and rewrites the events in the target text according to the target rewriting rules to obtain reference events. When the large model identifies the events in the target text, it performs word segmentation on the input target text, splitting the text into individual words or phrases. Then, using its powerful semantic understanding ability, it analyzes the semantic roles and context relationships of each word in the sentence. By identifying the grammatical structure, semantic information, and event trigger words in the text, it determines whether there are events in the text and the types of events. For example, if trigger words such as "purchase" or "sale" appear in the text, the large model can preliminarily judge that there may be events related to transactions. Moreover, the large model also combines the context information to determine the various elements of the event, such as the subjects participating in the transaction and the items traded.
[0052] The target rewriting rules are pre-set rules used to convert the identified events into a unified and standardized form. After obtaining the events in the target text, the large model rewrites the events according to the target rewriting rules. Specifically, the large model can adjust the expression of the events to make them more concise and clear. For example, it simplifies complex sentence structures and replaces vague expressions with accurate words. Of course, the large model can also organize and standardize the event elements to make the formats and expressions of the event elements consistent. Describe the events with simple sentences like "In xxx year, xxx happened". After rewriting, reference events with a unified structure and clear semantics can be obtained.
[0053] In some embodiments, identifying the events in the target text by the large model and rewriting the events in the target text according to the target rules includes: Constructing a prompt instruction for the large model; the prompt instruction includes the target rewriting rules; the target rewriting rules include: identifying the event content from the main body of the article of the target file, complementing the event elements of the event content based on the article name, chapter titles, and chapter content of the target file and performing coreference resolution, and simplifying the event content; Input the prompting instruction into the large model so that the large model can identify the events in the target text according to the prompting instruction and rewrite the events in the target text according to the target rewriting rules to obtain the reference events.
[0054] Although large models have powerful language understanding and processing capabilities, they need clear instructions to know what specific tasks to complete. Through the prompting instruction, the working direction and requirements can be defined for the large model, and integrating the target rewriting rules into the prompting instruction provides specific rewriting requirements for the work of the large model.
[0055] The target rewriting rules can include multiple key points. Identify the event content from the main text of the target document, and complement the event elements of the event content and disambiguate the references based on the article title, chapter title, and chapter content of the target document. Among them, the article title and chapter title contain the core theme and background information of the text, and the chapter content includes a detailed elaboration of the theme. The large model can use this information to fill in the key elements that may be missing in the event, such as clarifying the background of the event, supplementing vague time or location, etc. And by comprehensively analyzing this content, it is also possible to reduce the problem of unclear references that may exist in the text, such as clarifying the specific referents of pronouns such as "he", "it", "the same year", "this", etc.
[0056] Simplifying the event content can make the event expression more concise and clear, remove redundant information, highlight the key content, and facilitate subsequent storage and comparison. For example, complex sentence patterns can be simplified, and the expression content can be streamlined, using a simple sentence to describe the event content, thereby reducing the difficulty of understanding and ensuring the unity and comparability of the event expression.
[0057] After constructing the prompting instruction, the prompting instruction can be input into the large model so that the large model can identify the events in the target text according to the prompting instruction and rewrite the events in the target text according to the target rewriting rules to obtain the reference events.
[0058] In one example, the prompting instruction is as follows: #Goal You are an event extraction expert, and your task is to extract the events in the article body according to the content of [Reference Information] and the article body.
[0059] ##Rules First, you need to carefully read [Reference Information] and the article body. [Reference Information] includes <Book Name>, <Chapter Name>, <Previous Content>, etc., and the article body is the input text that needs to extract events.
[0060] Then, you need to extract important events from the main text of the article. If key elements of the event, such as time, location, characters, etc., are missing in the main text of the article, you can obtain them according to [Reference Information].
[0061] Finally, output the extracted events. If there is more than one event in the main text of the article, output multiple events. If there are no important events or the event information is incomplete in the main text of the article, output "No important events".
[0062] Note that the "^" sign appearing in the text indicates that the time is the same as that in the previous paragraph. Note that the extracted events need to strictly conform to [Reference Information] and the main text of the article. If key elements such as time in the input information are unclear or incomplete, do not make random speculations or supplement information, and output "Unable to obtain accurate information". Note that the output event information needs to be complete and clear, without containing referential relationships and unclear information. Describe the event information in a simple sentence like "In xxx year, xxx happened".
[0063] ##input [Reference Information]: <Book Name>: (book_name) <Chapter Name>: (book_title) <Previous Content>: (pre_contents) Please output the extraction result.
[0064] In the above prompt instructions, the "#Goal" part is about the role setting and task description, and the "##Rules" part is the description of various rules, including the goal rewriting rules. These contents together constitute the prompt instructions, which are used to inform the model of the task goal, execution rules, etc., and guide the model on how to extract events from the text. The "##input" part defines the reference information, that is, the target file referred to by the large model.
[0065] In this embodiment, by constructing a prompt instruction containing the goal rewriting rule, the direction and standard for event processing are specified for the large model. The event content is identified from the main text of the article, and the event elements are complemented and disambiguated based on the article name, chapter title, and chapter content. The context information of the text is utilized to reduce problems such as word ambiguity and unclear references, making the event elements more complete and accurate; simplifying the event content removes redundant information and highlights the core elements, making the event expression clearer and more concise.
[0066] In some embodiments, before storing the reference event in the event knowledge base, it includes: Based on the large model, verify whether the reference event is consistent with the event in the target text; In the case where the reference event is inconsistent with the event in the target text, modify the reference event so that the modified reference event is consistent with the event in the target text.
[0067] In this embodiment, since the reference event is used as the benchmark data for subsequent proofreading of the event to be proofread, it is very important whether the benchmark data is accurate. It can be judged whether there may be errors in time reasoning, context reference, and output format during the event rewriting process. If there may be errors, point out the possible error factors and ask the large model to verify and modify. After confirmation, obtain the final rewritten result, thereby improving the accuracy rate of event rewriting.
[0068] Specifically, the large model can deeply analyze the events in the reference event and the target text, starting from each event element, such as the subject of the event, the action of the event, the time and place where the event occurs, etc., and compare the event elements in the reference event with the corresponding event elements in the target text. For example, in an event describing "Xiaoming borrows books in the library", the large model can check whether "Xiaoming" in the reference event is consistent with the subject in the target text, whether the action of "borrowing" is accurate, and whether the information such as the location of the library and the specific time can all correspond.
[0069] If, after verification by the large model, it is found that the reference event is inconsistent with the event in the target text, for example, an event element is missing, perhaps the specific time when the event occurs is omitted in the reference event; or the element is inaccurate, for example, the subject of the event is misattributed, writing what was originally done by "Xiaoming" as done by "Xiaohong". When such problems occur, the reference event can be modified.
[0070] Specifically, the large model can make targeted corrections to the reference event based on the accurate event information in the target text. If an event element is missing, the large model will extract the corresponding information from the target text for supplementation. For example, if it is found that the location where the event occurs is not mentioned in the reference event, the large model will search for relevant content in the target text and add the accurate location information to the reference event. Through the modification of the large model, the reference event can be made consistent with the event in the target text, improving the accuracy and integrity of the reference event.
[0071] In an example, as Figure 2 shown, the article name of the target file is "The Rise and Fall of Ancient Rome" and the chapter name is "On the Eve of Caesar's Assassination". Through the large model's event recognition and rewriting of the article content, the obtained reference event is "In 45 BC, there was a dispute at the Senate meeting, and the next year someone planned to assassinate Caesar." After verification, it is found that the rewritten event is too brief, losing key figures (Cassius, Brutus, Antony), key locations (the Senate), accurate time (March 15, 44 BC), etc., and the verification fails.
[0072] The reference event is further modified by the large model. The modified reference event is "In 45 BC, during a Senate meeting, Cassius accused Caesar, Brutus echoed, and Antony refuted. After the meeting, Cassius and Brutus et al. planned to assassinate Caesar in the Senate on March 15, 44 BC." Further verification of the modified reference event shows that the rewritten event contains the core elements such as key figures, time, and location in the original content, which is consistent with the original text, and the verification passes.
[0073] In this embodiment, before storing the reference event in the knowledge base, the large model is used to perform consistency verification on the reference event and the original event in the target text, capture the differences in the expression, elements, etc. of the event, and reduce the problem of incorrect or deviated information being stored in the event knowledge base. When it is found that the reference event is inconsistent with the original event, the large model is used to modify the reference event so that the reference events stored in the event knowledge base all have high accuracy.
[0074] Step 120: Store the reference event in the event knowledge base.
[0075] In some embodiments, the reference event can be directly stored in the event knowledge base without extracting event elements, avoiding information loss caused by incomplete element extraction and retaining the natural expression of the language.
[0076] In some embodiments, the reference event can also be vectorized to obtain a vector representation of the reference event; the vector representation is stored in the event knowledge base.
[0077] In the event knowledge base, a large number of reference events need to be stored. If stored in the form of the original text, when retrieving reference events related to the event to be proofread, the computer needs to compare a large amount of text one by one, with low efficiency. In the numerical space of the vector representation, the reference events semantically similar to the event to be proofread can be quickly screened out by calculating the distance between vectors (such as cosine similarity and other measurement methods).
[0078] In this embodiment, the reference event can be vectorized through the bag-of-words model, word embedding (such as Word2Vec, GloVe, etc.) or a vector representation method based on a deep learning model to obtain a vector representation of the reference event, and the vector representation of the reference event is stored in the event knowledge base.
[0079] In this embodiment, by vectorizing reference events, complex event information can be converted into a numerical vector representation that is easy for a computer to process, reducing the dimensionality of the data and optimizing the data storage structure. When faced with an event to be verified, the vector-based retrieval method can quickly calculate the similarity between the event to be verified and the reference event vector. Compared with the traditional text matching method, it can more accurately locate relevant target reference events from a large amount of event data.
[0080] Step 130: Retrieve target reference events associated with the event to be verified from the event knowledge base based on the event to be verified.
[0081] In some embodiments, the event to be verified can be an event obtained by the large model identifying and rewriting the data to be verified. For example, the data to be verified can be obtained; the large model can be used to identify the events in the data to be verified and rewrite the events in the data to be verified according to the target rewriting rules to obtain the event to be verified. Of course, the event to be verified can also be an event input by the user or an event obtained from other clients, and this application does not make any limitations in this regard.
[0082] In some embodiments, the data to be verified can be data from different channels, such as news articles, online forum posts, enterprise internal documents, etc. The format of the data to be verified can include but is not limited to plain text format, Word documents, PDF files, web page content, etc. For example, in the intelligent review scenario, the data to be verified can be the dynamic information posted by the user on the social platform; in the field of digital publishing, it can be the manuscript document submitted by the author without strict review.
[0083] After obtaining the data to be verified, the data to be verified can be preprocessed. After preprocessing, the large model can be used to identify the events in the data to be verified and rewrite the events in the data to be verified according to the target rewriting rules to obtain the event to be verified. Among them, the preprocessing method, event identification, and rewriting method of the data to be verified can refer to the preprocessing method, event identification, and rewriting method of the target document, and this application will not elaborate on them one by one.
[0084] In this embodiment, by utilizing the powerful semantic understanding ability of the large model, the events in the data to be verified can be accurately identified, so that the key information of the events can be accurately extracted. The identified events are rewritten according to the target rewriting rules, and the original events are converted into events to be verified with a unified structure and clear semantics, realizing the standardization of event expression, facilitating subsequent comparison, reducing misjudgment caused by expression differences, and further improving the accuracy of the verification result.
[0085] In an embodiment of the present application, after obtaining the event to be proofread, the target reference event associated with the event to be proofread can be retrieved from the event knowledge base according to the storage structure and indexing mechanism of the event knowledge base. The target reference event can be one reference event or multiple reference events. For example, the event knowledge base can be pre-tokenized and an inverted index can be constructed. When querying an event, the event to be proofread parses keywords, matches index entries, and sorts and returns relevant target reference events based on a relevance algorithm.
[0086] In some embodiments, retrieving the target reference event associated with the event to be proofread from the event knowledge base based on the event to be proofread includes: Performing keyword matching from the event knowledge base according to the event to be proofread to obtain a first retrieval result; the first retrieval result includes multiple first candidate reference events sorted based on the matching degree; Vectorizing the event to be proofread, and performing similarity matching from the event knowledge base based on the vectorized event to be proofread to obtain a second retrieval result; the second retrieval result includes multiple second candidate reference events sorted based on the similarity; Assigning weights to the first candidate reference events and the second candidate reference events based on the event element type; Screening the target reference event from the first candidate reference events and the second candidate reference events according to the weights, the matching degree, and the similarity.
[0087] In this embodiment, a full-text retrieval method can be adopted, and through a keyword matching strategy, a preliminary set of candidate reference events can be screened out from the event knowledge base. Specifically, the keywords of the event to be proofread can be parsed. The reference events in the event knowledge base have been preprocessed during storage and also contain keyword fields available for matching. During the matching process, the matching degree score of each reference event can be calculated based on the frequency, position of the keyword, and the degree of association with the event theme. For example, if a reference event completely contains all the keywords of the event to be proofread and the distribution of the keywords in the text is reasonable, its matching degree score will be relatively high; if only some keywords match or the position of the keyword deviates from the core semantics, the score will be low. The reference events are sorted from high to low according to the matching degree score to generate a first retrieval result including multiple first candidate reference events. For example, the top 5 reference events with the highest matching degree scores can be used as the first candidate reference events. The number of first candidate reference events can be set in advance. For example, it can be 1, 2, 5, 10, or other numbers, and the embodiments of the present application do not limit this.
[0088] To more deeply explore the semantic associations between events, a preliminary candidate reference event set can also be screened from the event knowledge base through vector retrieval. Specifically, the vector representations of reference events are stored in the event knowledge base. The event to be proofread can be vectorized using the same vectorization method as the reference events, and the similarity (such as cosine similarity, Euclidean distance, etc.) between the vector of the event to be proofread and the vectors of each reference event can be calculated to evaluate their proximity at the semantic level. The reference events are sorted from high to low according to the similarity scores to generate a second retrieval result containing multiple second candidate reference events. For example, the top 5 reference events with the highest similarity can be used as the second candidate reference events. The number of second candidate reference events can be preset. For example, it can be 1, 2, 5, 10 or other numbers, and the embodiments of the present application do not limit this.
[0089] To more accurately determine the target reference event with the highest degree of relevance to the event to be proofread, weights can be assigned to each candidate reference event based on the event element type. Since different event elements have different importance in event description and judgment, in order to reduce the influence of possible error content (such as time, location, etc.) in the event to be proofread on the degree of relevance and cause the retrieval result to deviate from the actual event relevance, when sorting, the weights of event elements such as time and location can be scored lower, and the weight of the event subject can be scored higher. Weights are assigned to each candidate reference event according to this weight assignment method, and the final score of each candidate reference event is calculated by combining the matching degree and similarity. For example, the matching degree and similarity of each candidate reference event can be added and then multiplied by the weight to obtain the final score. Of course, the final score can also be calculated by other methods. According to the high and low of the final scores, the candidate reference events are sorted again, and the n candidate reference events with the highest final scores are determined as the target reference events. Among them, n can be 1, 2, 5, 10 or other numbers, and the embodiments of the present application do not limit this.
[0090] In this embodiment, by using keyword matching to obtain the first retrieval result, it is possible to quickly retrieve the preliminary candidate reference events related to the event to be proofread from the event knowledge base. After vectorizing the event to be proofread and performing similarity matching, the second retrieval result is obtained. The vector calculation is used to deeply explore similar events at the semantic level, making up for the problem of insufficient semantic understanding that may exist in keyword matching. Further, weights are assigned to the candidate reference events based on the event element type, and differential considerations are made according to the importance of the key components of the event. The target reference events are screened from the candidate reference events, which not only reduces the limitation problem of a single retrieval method but also highlights the key elements through weight assignment, improving the accuracy of retrieving the target reference events.
[0091] Step 140: Compare the event to be verified and the target reference event through a large model to obtain a comparison result indicating whether the event to be verified is correct.
[0092] In the embodiments of the present application, the event comparison process is as Figure 3 shown, Figure 3 The retrieval result in is the target reference event obtained by retrieval. The correlation between the two can be compared through a large model. If it is recognized that the two belong to different events, the comparison result output by the large model is "unable to determine whether the event to be verified is correct". If it is recognized that the two belong to the same event, it is necessary to further compare whether the event elements of the two are the same. If they are the same, the output comparison result is "the event to be verified is correct". If they are different, the output comparison result is "the event to be verified is incorrect".
[0093] In some embodiments, when comparing the event to be verified and the target reference event through a large model, it is also necessary to construct a prompt instruction, and the content included in the promotion instruction can include task prompts, comparison rules, output rules, etc. In one example, the prompt instruction is as follows: #Goal You are now a text proofreader. Your task is to proofread the "text to be verified" based on the "reference materials" I provided, judge whether there are errors in the elements such as the time of the text to be verified, and output the proofreading result.
[0094] ##Rules Explanation of key terms: Text to be verified: The text that needs to be proofread, and it is not certain whether there are errors.
[0095] Reference materials: The completely correct content, but it is not certain whether it is the same event as the text to be verified.
[0096] Same event: The reference material selected from the reference materials, which is the same event as the text to be verified and can fully prove that the time in the text to be verified is correct.
[0097] Related event: The reference material selected from the reference materials, which is not the same event as the text to be verified, but is closely related to the text to be verified Event correlation judgment: Ignoring secondary factors such as time and location, the time and location of the text to be verified may be incorrect. Therefore, these may not necessarily be the elements for judging whether it is the same event.
[0098] The event subject must be the same.
[0099] The event action must be the same.
[0100] The event object must be the same.
[0101] The event results must be consistent.
[0102] Related events: Different stages of an event cannot be judged as the same event. For example, "preparing a meeting" and "officially convening" are different, but they can be regarded as related events.
[0103] Time judgment: If it is possible to determine whether there are any error contents in the text to be proofread based on the reference materials, please output the proofreading result and explain which reference material or materials are used for the judgment.
[0104] If it is impossible to determine whether there are any error contents in the text to be proofread based on the reference materials. For example, the reference materials are not relevant to the text to be proofread, or the reference materials are relevant to the text to be proofread, but due to incomplete expression, ambiguity, etc., it is impossible to clearly determine that they are exactly the same event, and further verification is required. Please output "Unable to judge based on the reference materials".
[0105] If the reference material is an exact time xxxx year xx month xx day, and the text to be proofread is a rough time, xxxx year, or xxxx year xx month, if the times are the same and there is no more precise reference material, it can be output as correct. Conversely, if the text to be proofread is an exact time, xxxx year xx month xx day, and the reference material is a rough time, xxxx year, or xxxx year xx month, even if the times are the same, it can only be output as unable to judge.
[0106] Example of output format: ("Same events": ["xxx"], "Related events": ["xxx"], "Proofreading result": "xxx" ) ##Input Text to be proofread: {input_event} Reference materials: {reference_event} Please output the proofreading result.
[0107] In the above prompt instructions, the text to be proofread can be the event to be proofread, and the reference materials are the retrieval results, that is, the target reference events.
[0108] The following uses an example to illustrate the comparison process. For example Figure 4As shown, the event to be verified is "On May 10, 43 BC, the Roman senators deliberated and decided to prepare to elect a new consul in July of that year." The search results include: Reference 1 "On July 15, 43 BC, the Roman Senate officially elected Marcus as the new consul.", Reference 2 "On May 10, 43 BC, many Roman senators gathered and discussed the election of the consul, and preferred to choose a suitable time in summer for the election.", and Reference 3 "On May 10, 43 BC, the Roman Senate assembled, and many senators participated in the deliberation of the subsequent political personnel arrangements, mentioning the plan to promote the election process of the new consul in July." The specific judgment process of the large model in the event comparison includes: Step 1: Reference 1 mentions that Marcus was officially elected as the new consul on July 15, 43 BC, while the text to be verified states that preparations were made to elect a new consul in July of that year. One is the actual election time and result, and the other is the preparation for the election. They are not exactly the same event.
[0109] Reference 2 mentions that in 43 BC, on May 10, the senators discussed the election of the consul and preferred to choose a time in summer for the election. The main body of the event is the Roman senators, the behavior focuses on the discussion of the election of the consul, and the object is the election of the consul, which is the same as the text to be verified. Although the specific month of July is not clearly mentioned, according to the rules, such time differences can be ignored, and it should belong to the same event.
[0110] Reference 3 mentions that in 43 BC, on May 10, the Senate assembled to discuss personnel arrangements and planned to promote the election process of the new consul in July. The main body, behavior, and object of the event are also the same as those in the text to be verified, belonging to the same event.
[0111] Step 2: The time in the text to be verified is May 10, 43 BC, and both Reference 2 and Reference 3 clearly mention this time, so the time is correct.
[0112] Therefore, the comparison result is: According to Reference 2 and Reference 3, the event to be verified is correct.
[0113] The event content proofreading method based on a large model provided by an embodiment of the present application identifies events in a target text through the large model, and rewrites the events in the target text according to target rewriting rules to obtain reference events; the target text is standard data that has been proofread; stores the reference events in an event knowledge base; retrieves target reference events associated with the event to be proofread from the event knowledge base based on the event to be proofread; compares the event to be proofread and the target reference event through the large model to obtain a comparison result indicating whether the event to be proofread is correct. The embodiment of the present application utilizes the powerful semantic understanding ability of the large model, can accurately identify events in the standard data, and performs rewriting processing on the events to form reference events with unified structure and clear semantics, so as to be able to build a high-quality event knowledge base, reduce the problem of event argument extraction deviation in related technologies, and when proofreading the event to be proofread, can retrieve the target reference event from the event knowledge block according to relevance, and then use the large model for comparison, improving the accuracy of event content verification.
[0114] In the event content proofreading method based on a large model provided by an embodiment of the present application, the execution subject may be an event content proofreading device based on a large model. In the embodiment of the present application, taking the event content proofreading device based on a large model executing the event content proofreading method based on a large model as an example, the event content proofreading device provided by the embodiment of the present application is described.
[0115] An embodiment of the present application also provides an event content proofreading device based on a large model.
[0116] As Figure 5 shown, the event content proofreading device based on a large model includes: An identification and rewriting module 510, configured to identify events in a target text through a large model, and rewrite the events in the target text according to target rewriting rules to obtain reference events; the target text is standard data that has been proofread; A storage module 520, configured to store the reference events in an event knowledge base; A retrieval module 530, configured to retrieve target reference events associated with the event to be proofread from the event knowledge base based on the event to be proofread; A proofreading module 540, configured to compare the event to be proofread and the target reference event through a large model to obtain a comparison result indicating whether the event to be proofread is correct.
[0117] The event content proofreading method based on a large model provided by an embodiment of the present application identifies events in a target text through the large model, and rewrites the events in the target text according to target rewriting rules to obtain reference events; the target text is standard data that has been proofread; stores the reference events in an event knowledge base; retrieves target reference events associated with the event to be proofread from the event knowledge base based on the event to be proofread; compares the event to be proofread and the target reference event through the large model to obtain a comparison result indicating whether the event to be proofread is correct. The embodiment of the present application utilizes the powerful semantic understanding ability of the large model to accurately identify events in the standard data, and performs rewriting processing on the events to form reference events with a unified structure and clear semantics, thereby being able to build a high-quality event knowledge base, reducing the problem of event argument extraction deviation in related technologies. When proofreading the event to be proofread, the target reference event can be retrieved from the event knowledge block according to the relevance, and then compared with the help of the large model, improving the accuracy of event content verification.
[0118] The event content proofreading device based on a large model in an embodiment of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application does not make specific limitations.
[0119] The event content proofreading device based on a large model in an embodiment of the present application may be a device with an operating system. The operating system may be a Microsoft (Windows) operating system, an Android operating system, an IOS operating system, or other possible operating systems. The embodiment of the present application does not make specific limitations.
[0120] In some embodiments, such as Figure 6As shown in the figure, the embodiment of the present application further provides an electronic device 600, including a processor 601, a memory 602, and a computer program stored on the memory 602 and executable on the processor 601. When the program is executed by the processor 601, it implements each process of the above-mentioned embodiment of the event content proofreading method based on the large model and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0121] It should be noted that the electronic device in the embodiment of the present application includes the above-mentioned mobile electronic device and non-mobile electronic device.
[0122] The embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the event content proofreading method based on the large model and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0123] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disk or optical disc, etc.
[0124] The embodiment of the present application further provides a computer program product, including a computer program, which implements the above-mentioned event content proofreading method based on the large model when executed by a processor.
[0125] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disk or optical disc, etc.
[0126] The embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the event content proofreading method based on the large model and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0127] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0128] It should be noted that in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0130] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
[0131] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0132] Although embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.
Claims
1. A method for event content proofreading based on a large model, characterized in that: include: Recognize events in the target text through the large model, and rewrite the events in the target text according to the target rewriting rules to obtain reference events; The target text is standard data that has been reviewed; Storing the reference event in an event knowledge base; Retrieving a target reference event associated with the event to be proofread from the event knowledge base based on the event to be proofread; The event to be checked is compared with a target reference event through the large model to obtain a comparison result indicating whether the event to be checked is correct.
2. The method according to claim 1, characterized in that The method further comprises: Obtain the data to be proofread; Events in the data to be proofread are identified through the large model, and the events in the data to be proofread are rewritten according to the target modification rules to obtain the events to be proofread.
3. The method according to claim 1, characterized in that: The method further comprises: Vectorizing the reference event to obtain a vector representation of the reference event; The vector representation is stored in the event knowledge base.
4. The method according to claim 1, characterized in that: Before identifying events in the target text through the large model, including: The target file is preprocessed; the preprocessing includes: dividing the target file into chapters, and adding the marking symbols and / or annotation content in the target file to the original text to obtain formatted data corresponding to the target file; the formatted data includes the article name, chapter title and chapter content.
5. The method according to claim 1, characterized in that Before storing the reference event in the event knowledge base, the following steps are included: Verify whether the reference event is consistent with the event in the target text based on the large model; In the case that the reference event is inconsistent with the event in the target text, the reference event is modified so that the modified reference event is consistent with the event in the target text.
6. The method according to claim 1, characterized in that The identifying of events in the target text by using the large model and rewriting the events in the target text according to the target rules includes: Constructing a prompt instruction for the large model; the prompt instruction includes the target rewriting rule; the target rewriting rule includes: identifying event content from the article body of the target file, completing and disambiguating event elements of the event content based on the article name, chapter title and chapter content of the target file, and simplifying the event content; The prompt instruction is input into the large model, so that the large model recognizes the event in the target text according to the prompt instruction, and rewrites the event in the target text according to the target rewriting rule to obtain the reference event.
7. The method according to claim 1, characterized in that The step of retrieving a target reference event associated with the event to be proofread from the event knowledge base based on the event to be proofread includes: Perform keyword matching from the event knowledge base according to the event to be proofread to obtain a first search result; the first search result includes a plurality of first candidate reference events sorted based on matching degree; Vectorizing the event to be proofread, and performing similarity matching from the event knowledge base based on the vectorized event to be proofread to obtain a second search result; the second search result includes a plurality of second candidate reference events sorted based on similarity; assigning weights to the first candidate reference event and the second candidate reference event based on event element types; The target reference event is obtained by screening the first candidate reference event and the second candidate reference event according to the weight, matching degree and similarity.
8. A method for proofreading event content based on a large model, characterized in that: include: An identification and rewriting module is used to identify events in a target text through a large model, and rewrite the events in the target text according to a target rewriting rule to obtain a reference event; The target text is standard data that has been reviewed; A storage module, used for storing the reference event in an event knowledge base; A retrieval module, used for retrieving a target reference event associated with the event to be proofread from the event knowledge base based on the event to be proofread; The proofreading module is used to compare the event to be proofread with the target reference event through the large model to obtain a comparison result indicating whether the event to be proofread is correct.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the event content proofreading method based on a large model as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the event content proofreading method based on a large model as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Question and answer method and device based on large model
CN118885579A
Text factual proofreading method and system based on large language model
CN119149720A
Method, system and equipment for large language model fact verification based on retrieval enhancement and medium
CN119204025A
Event extraction method, device and storage medium
US20220300546A1
Generative event extraction method based on ontology guidance
US20240143633A1
Cited By
Method and system for multi-dimensionally summarizing classroom live-recorded texts
CN120319248A
Historical figure knowledge proofreading method and system, storage medium and electronic equipment
CN120611041A
Historical figure knowledge proofreading method and system, storage medium and electronic device
CN120611041B