NLP semantic extraction unstructured text generation logic axis and computing power optimization method
Through the large model cutting and extracting text content, combining semantic understanding and reasoning, generating logical axis and optimizing computing power, the problem of identifying complex logical relationships in the existing technology is solved, and the accuracy and efficiency of logical axis are improved.
Patent Information
- Application Number
- CN202510697068.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
When handling trial documents and transcripts in the prior art, it is difficult to accurately identify complex logical relationships in the text, resulting in the accuracy and completeness of the logical axis that needs to be improved, and there are often missing or incorrect logical relationships.
The text content is cut and extracted by a large model, forming structured data of different roles, and extracting time information and logical relationships through semantic understanding and reasoning, generating logical axes, and optimizing computing power allocation to improve processing efficiency.
It realizes accurate identification and extraction of complex logical relationships, improves the accuracy and completeness of the logic axis, reduces the consumption of computing resources, and improves the overall processing efficiency.
Smart Images

Figure CN120218084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of NLP, and particularly to a method for generating a logical axis from unstructured text in NLP semantic extraction and optimizing computing power. Background Art
[0002] The text contains rich information, and it is of great significance to mine the information therein. Named Entity Recognition (NER) is a key task in Natural Language Processing (NLP), which includes identifying and classifying certain types of information elements, namely Named Entities (NE). The main goal of NER is to identify and classify entity information from text, such as personal names, place names, organization names, and amounts, etc., and it is widely used in many fields such as medical treatment, railway construction, and geological science. In the context of the information explosion era, NER is not only crucial for information extraction and data analysis, but also plays a fundamental role in multiple application scenarios such as sentiment analysis, question answering systems, and machine translation.
[0003] During the daily security inspection process of relevant departments when processing a large number of court trial documents and transcript documents, the ability to capture complex semantics and implicit logical relationships in court trial documents and transcript documents is limited. It is difficult to accurately identify complex logics such as causality and turning points in the text, and the accuracy and integrity of generating the logical axis need to be improved. There are often situations where logical relationships are missing or incorrect, which affects the quality of subsequent tasks. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for generating a logical axis from unstructured text in NLP semantic extraction and optimizing computing power to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for generating a logical axis from unstructured text in NLP semantic extraction and optimizing computing power, which is characterized by including the following steps: Step S1: Use a large model to cut the text content, and cut out the content of different roles to form new text files of different roles; Step S2: Use a large model to extract the sentences related to time from each paragraph in the new text file to form an array; Step S3: After screening, remove the assumed time content in the sentences related to time in the array of Step S2; Step S4: Traverse the array in Step S2, format it uniformly, extract the time information, and infer the accurate date based on semantics; Step S5: Use the big model to convert the inferred time, characters, content elements, relationships, and mind maps into structured data, and concatenate them with the array in step S2 to form a structured array corresponding to each character. Reversely verify whether each inferred element is correct. If it is not correct, remove the element. Step S6: Display the data on a logical axis according to the computing power allocation. Clicking on the data on the logical axis supports jumping back to the original text.
[0006] Preferably, when filtering out hypothetical time content in time-related statements in step S3, when the text file contains statements with hypothetical time descriptions, the content related to time predictions and assumptions is intelligently identified and eliminated, and only the information that has occurred and is factual is retained after processing by the large model.
[0007] Preferably, when the time information is extracted in step S4, the specific holidays or anniversaries contained in the text file are converted into corresponding dates through semantic understanding. For holidays or anniversaries with unfixed dates, they are converted into corresponding months or quarters. For holidays or anniversaries that cannot be determined to a specific year, month or quarter, they are marked as "unknown".
[0008] Preferably, when time information is extracted in step S4, if the text file contains statements such as "next year", "Nth year after", "last year", "the year before", and "Nth year before", the large model first identifies and extracts the base year, and then according to predefined rules, the model understands "next year" as the year after the base year, "Nth year after" as the base year plus N years, "last year" as the year before the base year, "the year before" as the two years before the base year, and "Nth year before" as the base year minus N years. If the large model cannot complete the reasoning, these time expressions are displayed as is.
[0009] Preferably, when the data is displayed in a logical axis according to the computing power in step S6, the results are run in real time to form the logical axis if the computing power is sufficient. In the case of insufficient computing power, the logical relationship between the data can be run in advance, and the logical axis display can be generated through preset elements during use.
[0010] Preferably, when the data is logically displayed in step S6, when a character repeatedly states the data at the same time point, priority is given to displaying the latest data describing the time point for each character, to ensure that the logical axis is clear and the information is accurate. At the same time, in order to retain complete data records, all historical statements are displayed in a superimposed manner. In addition, the system supports comparison of data of different characters at the same time point, and presents the differences in viewpoints of each character by means of parallel display or difference annotation.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention extracts data from the text, promotes the functionalization process of large models in practical applications, and enables these models to serve various business scenarios more accurately.
[0012] (2) The present invention extracts semantic information from the text by presetting elements and constructs a logical axis, reducing the computing power requirements, reducing the consumption of computing resources, and improving the overall processing efficiency.
[0013] (3) During the use of the present invention, the user can add elements at any time according to the file content and requirements, and they will be added to the generated logical axis according to the preset logic, realizing dynamic incremental processing of the content. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Embodiment
[0016] In the daily security inspection process of relevant departments, a large number of court trial documents and transcript documents need to be processed. The implementation steps of the NLP semantic extraction method for generating a logical axis from unstructured text and optimizing computing power are as follows: Step S1: Use a large model to cut the text content, and cut out the content of different roles to form new text files of different roles; Step S2: Use a large model to extract the sentences related to time from each paragraph in the new text file to form an array, and store it in the memory space in the format of [{ 'time': '', 'relship', 'event': '', 'person', 'address': ''}....]; Step S3: Remove the assumed time content in the time-related statements in the array obtained in Step S2. When removing the assumed time content in the time-related statements, if the text file contains statements with hypothetical time descriptions, the content related to time prediction and assumption is identified and removed through intelligent recognition. After being processed by the large model, only the information that has occurred and is a fact is retained. For example, in a court trial regarding a commercial cooperation breach dispute, the plaintiff stated, "If the defendant can complete the delivery of the first batch of goods before March 1st as stipulated in the contract, then the subsequent production plan of our company can proceed according to the time node of March 5th, and there will be no stagnation of the production line due to raw material shortages, resulting in a loss of up to 500,000 yuan." The defendant replied, "Even if we deliver the goods on March 1st, but your company's production equipment suddenly breaks down, it will also lead to the stagnation of the production line. You can't blame us entirely." Both the plaintiff and the defendant used hypothetical time statements to elaborate their views. After the large model screened all the time statements that appeared, such assumed times were removed; Step S4: Traverse the array in Step S2, format it uniformly, extract the time information, and convert the specific festivals or anniversaries contained in the text file into corresponding dates through semantic understanding. For festivals or anniversaries with uncertain dates, they are converted to the corresponding months or quarters. For festivals or anniversaries that cannot be determined to a specific year, month, or quarter, they are marked as "unknown". For the time marked as "unknown", the original description statement is retained. When the text file contains statements such as "the following year", "the Nth year after", "last year", "the year before last", "the Nth year before", the large model first identifies and extracts the reference year, and then according to the predefined rules, the model understands "the following year" as the next year of the reference year, "the Nth year after" as the reference year plus N years, "last year" as the previous year of the reference year, "the year before last" as the two previous years of the reference year, and "the Nth year before" as the reference year minus N years. If the large model cannot complete the reasoning, these time expressions are displayed as they are; Step S5: Use the large model to form structured data from the inferred time, characters, content elements, relationships, and mind maps, and splice them with the array in Step S2 to form a structured array corresponding to each role. Reverse verify whether each inferred element is correct. If it is incorrect, remove the element. The large model conducts reasoning and answering on the segmented text file in Step S1 to obtain the answer. During reverse verification, each element is compared with the obtained answer. If the answer is inconsistent, the element is removed; Step S6: Display the data on the logical axis according to the computing power allocation. When the computing power is sufficient, the results are generated in real time to form a logical axis. When the computing power is insufficient, the logical relationships between the data are pre-run in advance. For example, when the user needs to process a large number of case files, the computing power may be insufficient, and the time taken to generate the logical axis is too long. The split files are then calculated in parallel using multiple threads and completed in one go. When displaying, only filter and display according to the elements to reduce the time complexity. During use, generate a logical axis through preset elements. Clicking on the data on the logical axis supports jumping back to the original text for traceability. The generated logical axis can be a time logic, a task relationship logic, an event logic, or a user-defined logic; For example: In a court trial regarding a dispute over the payment for a sales contract, the plaintiff stated, "On August 10, 2022, our company signed a sales contract with the defendant, agreeing to supply electronic products worth 150,000 yuan, and the payment was to be made within 30 days after the goods arrived; on August 20, 2022, our company delivered the goods as agreed and attached a certificate of conformity. On September 20, 2022, the payment deadline expired, but the defendant delayed on the grounds of tight funds. After repeated reminders, the defendant paid 30,000 yuan in November 2022, and the remaining 120,000 yuan was unpaid. Therefore, we filed a lawsuit to claim the remaining payment and interest." The defendant responded, "After receiving the goods, we found that some of the products had quality problems and affected the use, so we did not pay in a timely manner. The 30,000 yuan paid in November 2022 was a partial payment, and we will negotiate to resolve the quality problem before discussing further payment." When generating the logical axis, after extracting the time information described by the plaintiff and the defendant during the above time period, generate a timeline in chronological order, and display the event description that occurred at that time node on the time node corresponding to the timeline. When reading the event description by clicking on the time node on the timeline, you can trace back to the original text of the court trial document to view the event description at that node.
[0017] Step S7: When the data is logically displayed and a role repeatedly states the data at the same time point, the latest data described by each role at that time point is preferentially displayed to ensure that the logical axis is clear and the information is accurate. At the same time, to retain the complete data record, all historical statements are superimposed and displayed. In addition, the system supports comparing the data of different roles at the same time point, presenting the differences in the viewpoints of each role by means of side-by-side display or difference marking; if the plaintiff describes the event that occurred at the same time point differently in the first and second instances, the description of the plaintiff in the second instance of the event is preferentially displayed on the generated timeline, while retaining the different versions of the plaintiff's description of the event in the first instance. When the defendant describes the event at that time point, add the description to that time point. When the user reads the event at that time point, the descriptions of the defendant and the plaintiff are displayed side by side, and the conflicting content between the two descriptions is highlighted in red for the user to compare and read.
[0018] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A method for NLP semantic extraction to generate a logical axis from unstructured text and optimize computing power, characterized in that The following steps are involved: Step S1: Use a large model to cut the text content, cut out the content of different roles to form new text files of different roles; Step S2: extract the sentences related to time from each paragraph in the new text file using the big model to form an array; Step S3: Screen and remove the assumed time content in the time-related sentences in the array of step S2; Step S4: traverse the array in step S2, format it uniformly, extract the time information, and derive the accurate date based on semantic reasoning; Step S5: Use the big model to convert the inferred time, characters, content elements, relationships, and mind maps into structured data, and concatenate them with the array in step S2 to form a structured array corresponding to each character. Reversely verify whether each inferred element is correct. If it is not correct, remove the element. Step S6: Display the data on a logical axis according to the computing power allocation. Clicking on the data on the logical axis supports jumping back to the original text.
2. The method for generating a logical axis from unstructured text for NLP semantic extraction and optimizing computing power according to claim 1, characterized in that: When filtering out hypothetical time content in time-related statements in step S3, when a text file contains statements describing hypothetical time, the content related to time prediction and assumptions is intelligently identified and eliminated, and only information that has occurred and is factual is retained after processing by a large model.
3. The method for generating a logical axis and optimizing computing power by extracting unstructured text in NLP semantics according to claim 1, wherein: When extracting time information in step S4, the specific holidays or anniversaries contained in the text file are converted into corresponding dates through semantic understanding. For holidays or anniversaries with unfixed dates, they are converted into corresponding months or quarters. For holidays or anniversaries that cannot be determined to a specific year, month or quarter, they are marked as "unknown".
4. The method for generating a logical axis and optimizing computing power by extracting unstructured text in NLP semantics according to claim 1, wherein: When extracting time information in step S4, if the text file contains statements such as "next year", "Nth year after", "last year", "the year before", and "Nth year before", the big model first identifies and extracts the base year, and then according to predefined rules, the model interprets "next year" as the year after the base year, "Nth year after" as the base year plus N years, "last year" as the year before the base year, "the year before" as the two years before the base year, and "Nth year before" as the base year minus N years. If the big model cannot complete the reasoning, it displays these time expressions as is.
5. The method for generating a logical axis from unstructured text for NLP semantic extraction and optimizing computing power according to claim 1, characterized in that: When the data is displayed in a logical axis according to the computing power in step S6, if the computing power is sufficient, the results are run in real time to form the logical axis. If the computing power is insufficient, the logical relationship between the data is run in advance, and the logical axis display is generated by preset elements during use.
6. The method for generating a logical axis from unstructured text for NLP semantic extraction and optimizing computing power according to claim 1, characterized in that: When the data is logically displayed in step S6, if a character repeatedly states the data at the same time point, the latest data describing the time point for each character is displayed first to ensure that the logical axis is clear and the information is accurate. At the same time, in order to retain the complete data record, all historical statements are displayed in a superimposed manner. In addition, the system supports data comparison of different characters at the same time point, and presents the differences in viewpoints of each character through parallel display or difference annotation.
Citation Information
Patent Citations
Inference task processing method, question and answer method and server
CN117194639A
Log analysis system and method based on event link multi-dimensional visualization
CN118708560A
Lightweight large model-based electric power knowledge system construction and intelligent question and answer method
CN119597864A
Text processing method and device, electronic equipment, computer readable storage medium and computer program product
CN119692462A
Network threat intelligence analysis attack scene graph (ASG) generation method based on deep learning and natural language processing
CN119966695A