NLP semantic extraction of unstructured text to generate logical axes and optimize computing power methods

The trial documents are cut and formatted through a large model, and hypothetical time content is filtered, structured data is generated and logical axis is displayed. This solves the problem of insufficient logical relationship identification, improves the accuracy and efficiency of logical axis, and supports dynamic incremental processing.

CN120218084BActive Publication Date: 2025-08-15JIANGSU ZHONGWEI TECH SOFTWARE SYST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510697068.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-15
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

During the daily security inspection process, it is difficult for the existing technology to accurately identify the complex semantic and implicit logical relationships in court trial documents and transcript documents, resulting in insufficient accuracy and completeness of the generated logical axis, which affects the quality of subsequent tasks.

Method used

The text content is cut and formatted by large models, extracting time-related statements, filtering hypothetical time content, generating structured data through semantic understanding and inference, and displaying the logical axis in real time if computing power allows. Otherwise, the logical relationship is pre-processed and displayed, and dynamic incremental processing is supported.

Benefits of technology

It improves the accuracy and completeness of the generation of logical axes, reduces the consumption of computing resources, realizes efficient services for large models in practical applications, and supports users to dynamically add elements according to their needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218084B_ABST
    Figure CN120218084B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for NLP semantic extraction of unstructured text to generate logical axes and optimize computing power, extracting data from text, promoting the functionalization process of large-scale models in practical applications, and enabling these models to serve various business scenarios more accurately. By presetting elements, semantic information is extracted from text and logical axes are constructed, which reduces computing power requirements, reduces the consumption of computing resources, and improves overall processing efficiency. During use, users can add elements at any time according to file content and needs, and they are added to the generated logical axis according to preset logic, realizing dynamic incremental processing of content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of NLP, and in particular to a method for generating logical axes and optimizing computing power by extracting unstructured text from NLP semantics. Background Art

[0002] Text contains a wealth of information, making information mining crucial. Named Entity Recognition (NER) is a key task in natural language processing (NLP), involving the identification and classification of certain types of information elements, known as named entities (NEs). NER's primary goal is to identify and classify entity information from text, such as names of people, places, organizations, and amounts. It is widely used in fields such as healthcare, railway construction, and geology. In the era of information explosion, NER is not only crucial for information extraction and data analysis, but also plays a fundamental role in multiple application scenarios, including sentiment analysis, question-answering systems, and machine translation.

[0003] When relevant departments handle a large number of court documents and transcripts during daily security checks, their ability to capture the complex semantics and implicit logical relationships in the documents is limited. It is difficult to accurately identify complex logic such as cause and effect and transitions in the text. The accuracy and completeness of the generated logical axis need to be improved, and logical relationships are often missing or erroneous, affecting the quality of subsequent tasks. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for generating logical axes and optimizing computing power by extracting unstructured text using NLP semantics, so as to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for extracting unstructured text from NLP semantics to generate logical axes and optimize computing power, characterized by comprising the following steps:

[0006] Step S1: Use the large model to cut the text content, cut out the content of different roles to form new text files for different roles;

[0007] Step S2: Using the large model, extract the sentences related to time from each paragraph in the new text file and form an array;

[0008] Step S3: Filter and remove the assumed time content in the time-related sentences in the array of step S2;

[0009] Step S4: traverse the array in step S2, unify the formatting, extract the time information, and derive the accurate date based on semantic reasoning;

[0010] Step S5: Use the large model to convert the inferred time, characters, content elements, relationships, and mind maps into structured data, and concatenate them with the array in step S2 to form a structured array corresponding to each character. Reverse-verify each inferred element to see if it is correct. If it is incorrect, remove the element.

[0011] Step S6: Display the data on a logical axis based on computing power allocation. Clicking on the data on the logical axis allows you to jump back to the original text.

[0012] Preferably, when filtering out the hypothetical time content in the time-related statements in step S3, when the text file contains statements with hypothetical time descriptions, the content related to time predictions and assumptions is intelligently identified and eliminated, and only the information that has occurred and is factual is retained after processing by the large model.

[0013] Preferably, when extracting time information in step S4, the specific holidays or anniversaries contained in the text file are converted into corresponding dates through semantic understanding. For holidays or anniversaries with unfixed dates, they are converted into corresponding months or quarters. For holidays or anniversaries that cannot be determined to a specific year, month or quarter, they are marked as "unknown".

[0014] Preferably, when time information is extracted in step S4, when the text file contains statements such as "next year", "Nth year after", "last year", "the year before", and "Nth year before", the large model first identifies and extracts the base year, and then, according to predefined rules, the model understands "next year" as the year after the base year, "Nth year after" as the base year plus N years, "last year" as the year before the base year, "the year before" as the two years before the base year, and "Nth year before" as the base year minus N years. If the large model cannot complete the reasoning, these time expressions are displayed as they are.

[0015] Preferably, when the data is displayed as a logical axis according to the computing power in step S6, the results are generated in real time to form the logical axis if the computing power is sufficient.

[0016] In the case of insufficient computing power, the logical relationship between the data can be run in advance, and the logical axis display can be generated through preset elements during use.

[0017] Preferably, when the data is logically displayed in step S6, when a character repeatedly states the data at the same time point, the latest data describing the time point for each character is displayed first to ensure that the logical axis is clear and the information is accurate. At the same time, in order to retain complete data records, all historical statements are displayed in a superimposed manner. In addition, the system supports data comparison of different characters at the same time point, and presents the differences in views of each character through parallel display or difference annotation.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] (1) This invention extracts data from text, promotes the functionalization of large-scale models in practical applications, and enables these models to serve various business scenarios more accurately.

[0020] (2) The present invention extracts semantic information from text by presetting elements and constructs logical axes, which reduces computing power requirements, reduces computing resource consumption, and improves overall processing efficiency.

[0021] (3) During the use of the present invention, users can add elements at any time according to the file content and needs, and they are added to the generated logical axis according to the preset logic, thereby realizing dynamic incremental processing of content. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0023] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention. Example

[0024] During routine security inspections, relevant departments need to process a large number of court documents and transcripts. The proposed NLP semantic extraction method for unstructured text, generating logical axes, and optimizing computing power is implemented as follows:

[0025] Step S1: Use the large model to cut the text content, cut out the content of different roles to form new text files for different roles;

[0026] Step S2: extract the sentences related to time from each paragraph in the new text file using the big model, form an array, and store it in the memory space in the format of [{''time:'','relship','event':'','person','address',''}....];

[0027] Step S3: Filter and remove the hypothetical time content in the time-related statements in the array of step S2. When filtering and removing the hypothetical time content in the time-related statements, intelligently identify and remove the content related to time predictions and assumptions when the text file contains hypothetical time descriptions. After processing through the big model, only the information that has already occurred and is factual is retained. For example, in a trial regarding a commercial cooperation breach dispute, the plaintiff stated: "If the defendant can complete the delivery of the first batch of goods before March 1st as agreed in the contract, then our company's subsequent production plan can proceed according to the March 5th time node, and there will be no production line stoppage due to raw material shortages, resulting in losses of up to 500,000 yuan." The defendant responded: "Even if we delivered the goods on March 1st, a sudden failure of your production equipment would also cause the production line to stop, and the responsibility cannot be attributed entirely to us." Both the defendant and the plaintiff used hypothetical time statements to express their views. After the big model screened all time statements, such hypothetical time was removed.

[0028] Step S4: Traverse the array in step S2, unify the formatting, extract the time information, and convert the specific holidays or anniversaries contained in the text file into corresponding dates through semantic understanding. For holidays or anniversaries with unfixed dates, convert them to the corresponding months or quarters. For holidays or anniversaries that cannot be determined to a specific year, month or quarter, they will be marked as "unknown". The original description statement will be retained for the time marked as "unknown". When the text file contains the statements "next year", "the Nth year after", "last year", "the year before", and "the Nth year before", the large model will first identify and extract the base year. Then, according to the predefined rules, the model will understand "next year" as the year after the base year, "the Nth year after" as the base year plus N years, "last year" as the year before the base year, "the year before" as the two years before the base year, and "the Nth year before" as the base year minus N years. If the large model cannot complete the reasoning, these time expressions will be displayed as they are;

[0029] Step S5: Use the big model to convert the inferred time, characters, content elements, relationships, and mind maps into structured data, and splice them with the array in step S2 to form a structured array corresponding to each character. Reverse-verify whether each inferred element is correct. If it is incorrect, the element is eliminated. The big model performs reasoning and question-answering on the cut text file in step S1 to obtain the answer. During reverse verification, each element is compared with the obtained answer. If the answer is inconsistent, the element is eliminated.

[0030] Step S6: Display the data on a logical axis according to the computing power allocation. When the computing power is sufficient, the results are generated in real time to form the logical axis. When the computing power is insufficient, the logical relationship between the data is run in advance. For example, when the user needs to process a large number of files, there will be insufficient computing power, and the time to generate the logical axis will be too long. The split files will be multi-threaded and parallelized, and the calculation will be completed in one go. When displaying, it is only necessary to filter and display according to the elements to reduce the time complexity. During use, the logical axis is generated by the preset elements. Clicking the data on the logical axis supports jumping back to the original text. The generated logical axis can be time logic, task relationship logic, event logic, and user-defined logic;

[0031] For example, in a court hearing regarding a dispute over payment for a sales contract, the plaintiff stated: "On August 10, 2022, our company signed a sales contract with the defendant, agreeing to supply electronic products worth 150,000 yuan, with payment due within 30 days of arrival. On August 20, 2022, our company delivered the goods as agreed, along with a certificate of conformity. When the payment came due on September 20, 2022, the defendant delayed payment, citing financial constraints. After repeated demands for payment, the defendant paid 30,000 yuan in November 2022, leaving the remaining 120,000 yuan unpaid. The plaintiff then sued for the remaining payment plus interest." The defendant responded: "After receiving the goods, we discovered that some of the products had quality issues that affected their use, so we delayed payment. The 30,000 yuan payment in November 2022 was a partial payment, and payment will be discussed later after the quality issues have been resolved through negotiation."

[0032] When generating the logical axis, the time information described by the plaintiff and defendant in the above time is extracted, and a timeline is generated in chronological order. The description of the event that occurred at the time node corresponding to the time node on the timeline is displayed. When clicking on the time node on the timeline to read the event description, you can go back to the original text of the trial document to view the event description of the node.

[0033] Step S7: When the data is displayed logically, if a character repeatedly states the data at the same time point, the latest data describing the time point for each character will be displayed first to ensure that the logical axis is clear and the information is accurate. At the same time, in order to retain the complete data record, all historical statements are displayed in a superimposed manner. In addition, the system supports data comparison of different characters at the same time point, and presents the differences in the views of each character through parallel display or difference annotation. If the plaintiff's description of the event occurring at the same time point is different in the first and second instance, the plaintiff's description of the event in the second instance will be displayed first on the generated timeline, while retaining the different versions of the plaintiff's description of the event in the first instance. When the defendant describes the event at that time point, the description will be added to that time point. When the user reads the event at that time point, the defendant and the plaintiff's descriptions will be displayed side by side, and the conflicting content will be marked in red for the user to read comparatively.

[0034] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. NLP semantic extraction of unstructured text to generate logical axes and optimize computing power method, characterized by: The following steps are involved: Step S1: Use the large model to cut the text content, cut out the content of different roles to form new text files for different roles; Step S2: Using the large model, extract the sentences related to time from each paragraph in the new text file and form an array; Step S3: Filtering and removing the hypothetical time content in the time-related sentences in the array of step S2. When filtering and removing the hypothetical time content in the time-related sentences, sentences containing hypothetical time descriptions in the text file are intelligently identified and removed from the content related to time predictions and assumptions. After processing with the large model, only information that has already occurred and is factual is retained. Step S4: traverse the array in step S2, unify the formatting, extract the time information, and derive the accurate date based on semantic reasoning; Step S5: Use the large model to convert the inferred time, characters, content elements, relationships, and mind maps into structured data, and concatenate them with the array in step S2 to form a structured array corresponding to each character. Reverse-verify each inferred element to see if it is correct. If it is incorrect, remove the element. Step S6: Display the data on a logical axis based on computing power allocation. Clicking on the data on the logical axis allows you to jump back to the original text.

2. The method for generating logical axes and optimizing computing power from NLP semantic extraction of unstructured text according to claim 1 is characterized by: When extracting time information in step S4, the specific holidays or anniversaries contained in the text file are converted into corresponding dates through semantic understanding. For holidays or anniversaries with unfixed dates, they are converted into corresponding months or quarters. For holidays or anniversaries that cannot be determined to a specific year, month or quarter, they are marked as "unknown".

3. The method for generating logical axes and optimizing computing power from NLP semantic extraction of unstructured text according to claim 1 is characterized by: When extracting time information in step S4, if the text file contains statements such as "next year", "Nth year after", "last year", "the year before", and "Nth year before", the large model first identifies and extracts the base year. Then, based on predefined rules, the model interprets "next year" as the year after the base year, "Nth year after" as the base year plus N years, "last year" as the year before the base year, "the year before" as the two years before the base year, and "Nth year before" as the base year minus N years. If the large model cannot complete the reasoning, these time expressions are displayed as they are.

4. The method for generating logical axes and optimizing computing power from NLP semantic extraction of unstructured text according to claim 1 is characterized by: In step S6, when the data is displayed in a logical axis according to the computing power, the results are generated in real time to form a logical axis if the computing power is sufficient. In the case of insufficient computing power, the logical relationship between the data can be run in advance, and the logical axis display can be generated through preset elements during use.

5. The method for generating logical axes and optimizing computing power from NLP semantic extraction of unstructured text according to claim 1 is characterized by: When the data is logically displayed in step S6, if a character repeatedly states the data at the same time point, the latest data describing the time point for each character will be displayed first to ensure that the logical axis is clear and the information is accurate. At the same time, in order to retain complete data records, all historical statements are displayed in a superimposed manner. In addition, the system supports data comparison of different characters at the same time point, and presents the differences in the views of each character through parallel display or difference annotation.

Citation Information

Patent Citations

  • Inference task processing method, question and answer method and server

    CN117194639A

  • Log analysis system and method based on event link multi-dimensional visualization

    CN118708560A