Document data analysis method and device of large language model, equipment and medium
By generating structured input data and using a large language model for hierarchical structure analysis, combined with long-term thinking chain technology, the problem of high manual maintenance costs in existing technologies is solved, intelligent classification and dynamic grading of data are achieved, and management efficiency and the interpretability of results are improved.
Patent Information
- Application Number
- CN202510821632.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
The existing data classification and grading system based on large language models relies on manually written prompt words, resulting in high maintenance costs and delayed responses, making it difficult to achieve intelligent classification and dynamic grading of data.
By determining the preset data management rules file, generating structured input data based on the hierarchical configuration file, and using the large language model to perform multi-dimensional text understanding and analysis, identifying the hierarchical structure, generating the target hierarchical configuration file, automatically parsing the document to be analyzed, and combining the long thinking chain technology to provide a logical reasoning process.
It realizes the automated parsing of data management details, improves management efficiency, ensures the accuracy and interpretability of classification results, supports real-time multi-source data verification, and reduces manual intervention.
Smart Images

Figure CN120671638A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of document content detection, and in particular to a document data analysis method, device, equipment and medium for a large language model. Background Art
[0002] With the rapid development of the information age, all sectors of society are increasingly reliant on smart devices and the internet, and the scale and speed of data flows within the internet are growing exponentially. Data classification and grading, as a core component of a data security management system, systematically organizes and standardizes the management of data assets at the source, scientifically defines data value attributes and risk levels, and provides a theoretical basis and practical guidance for the subsequent construction of a multi-layered security protection system.
[0003] How to combine business scenarios, compliance requirements and risk characteristics to achieve intelligent classification and dynamic grading of data, break through the limitations of manual reliance, reduce management costs, and at the same time support refined protection of sensitive data and provide technical support for security management of the entire data life cycle is an important innovation direction to adapt to the data governance needs of the digital age.
[0004] Current data classification and grading technologies primarily rely on methods such as keyword matching, semantic similarity calculation, and deep neural network classification. While new solutions based on large-scale pre-trained language models have emerged in recent years, their practical application still faces significant technical challenges. For example, in the case of corporate trade secret data management, detailed data management regulations typically cover a large number of clauses covering technical secrets, business information, customer data, personal privacy, and other multi-dimensional content, and require frequent updates in response to regulatory policies and corporate strategy adjustments. Existing classification and grading systems based on large language models still rely heavily on manually written prompts, resulting in significant issues such as high maintenance costs and delayed system response.
[0005] To sum up, how to realize intelligent classification and dynamic grading of data is an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the present invention aims to provide a document data analysis method, apparatus, device, and medium for a large language model, which can realize intelligent classification and dynamic grading of data. The specific scheme is as follows:
[0007] In a first aspect, the present application provides a document data analysis method based on a large language model, comprising:
[0008] Determining a preset data management rule file, performing a preset hierarchical configuration operation on the preset data management rule file based on a preset hierarchical configuration file establishment method to obtain a corresponding configuration result, and then adding a corresponding preset prompt word to the configuration result to generate corresponding structured input data;
[0009] Inputting the structured input data into a preset large language model, performing a preset multi-dimensional text understanding and analysis operation on the structured input data using the preset large language model to identify a hierarchical structure in the structured input data, and generating a corresponding target hierarchical configuration file based on the hierarchical structure;
[0010] Obtain the document to be analyzed, input the document to be analyzed and the target hierarchical configuration file into a preset large language model, and use the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain the target analysis result and the corresponding model reasoning process information.
[0011] Optionally, the preset hierarchical configuration file is established by manually establishing the hierarchical configuration file or automatically establishing the hierarchical configuration file;
[0012] Accordingly, the preset hierarchical configuration operation is performed on the preset data management rules file based on the preset hierarchical configuration file establishment method to obtain a corresponding configuration result, including:
[0013] If the preset hierarchical configuration file is established manually, the user terminal directly performs the preset hierarchical configuration operation on the preset interactive page according to the data management rules document to obtain the corresponding configuration result;
[0014] If the preset hierarchical configuration file establishment method is to automatically establish a hierarchical configuration file, the data management detail document is uploaded, and a preset hierarchical configuration operation is performed on the data management detail document to obtain a corresponding configuration result.
[0015] Optionally, after generating a corresponding target hierarchical configuration file based on the hierarchical structure, the method further includes:
[0016] Determining whether the target level configuration file meets a preset configuration condition;
[0017] If the target level configuration file does not meet the preset configuration conditions, modifying the target level configuration file through the user terminal;
[0018] If the target-level configuration file meets a preset configuration condition, the step of inputting the document to be analyzed and the target-level configuration file into a preset large language model is triggered.
[0019] Optionally, the using the preset large language model to perform a preset multi-dimensional text understanding and analysis operation on the structured input data to identify a hierarchical structure in the structured input data, and generating a corresponding target hierarchical configuration file based on the hierarchical structure includes:
[0020] Using the preset large language model to perform a preset multi-dimensional text understanding and analysis operation on the structured input data to obtain data categories in the structured input data;
[0021] The structured input data is classified based on the data category to obtain a hierarchical structure in the structured input data, and a corresponding target hierarchical configuration file is generated based on the hierarchical structure.
[0022] Optionally, the using the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain a target analysis result and corresponding model reasoning process information includes:
[0023] Determine the current layer node corresponding to each layer according to the hierarchical structure of the target layer configuration file;
[0024] Parsing the document to be analyzed using the current layer nodes corresponding to each layer to obtain the parsed data of each layer and the corresponding model reasoning process information;
[0025] The target analysis result corresponding to the document to be analyzed is determined based on the parsed data at each level.
[0026] Optionally, the document to be analyzed is parsed using the preset large language model based on the hierarchical structure in the target hierarchical configuration file to obtain corresponding model reasoning process information, including:
[0027] Using the long thought chain technology, the step of parsing the document to be analyzed using the preset large language model based on the hierarchical structure in the target hierarchical configuration file is decomposed into various intermediate links;
[0028] Perform preset verification and integration operations on each intermediate link to obtain the corresponding model reasoning path.
[0029] Optionally, the using the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain a target analysis result and corresponding model reasoning process information includes:
[0030] Using a preset tool interface to perform a preset verification operation on the data in the document to be analyzed to obtain a target response result;
[0031] According to the target response result, the preset large language model is used to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain a target analysis result and corresponding model reasoning process information.
[0032] In a second aspect, the present application provides a document data analysis device based on a large language model, comprising:
[0033] a data generation module, configured to determine a preset data management rule file, perform a preset hierarchical configuration operation on the preset data management rule file based on a preset hierarchical configuration file establishment method to obtain a corresponding configuration result, and then add a corresponding preset prompt word to the configuration result to generate corresponding structured input data;
[0034] a file generation module, configured to input the structured input data into a preset large language model, perform a preset multi-dimensional text understanding and analysis operation on the structured input data using the preset large language model, so as to identify a hierarchical structure in the structured input data, and generate a corresponding target hierarchical configuration file based on the hierarchical structure;
[0035] The result and process acquisition module is used to obtain the document to be analyzed, input the document to be analyzed and the target hierarchical configuration file into a preset large language model, and use the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain the target analysis result and the corresponding model reasoning process information.
[0036] In a third aspect, the present application provides an electronic device, comprising:
[0037] Memory, used to store computer programs;
[0038] A processor is used to execute the computer program to implement the document data analysis method based on the large language model as described above.
[0039] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned document data analysis method based on a large language model.
[0040] In summary, the present application first determines the preset data management rules file, performs a preset hierarchical configuration operation on the preset data management rules file based on the preset hierarchical configuration file establishment method to obtain the corresponding configuration result, and then adds the corresponding preset prompt word to the configuration result to generate the corresponding structured input data; the structured input data is input into the preset large language model, and the preset large language model is used to perform a preset multi-dimensional text understanding analysis operation on the structured input data to identify the hierarchical structure in the structured input data, and generate the corresponding target hierarchical configuration file based on the hierarchical structure; obtains the document to be analyzed, inputs the document to be analyzed and the target hierarchical configuration file into the preset large language model, and uses the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain the target analysis result and the corresponding model reasoning process information. As can be seen from the above, the present application first determines the preset data management rules file, performs a preset hierarchical configuration operation on it according to the establishment method based on the preset hierarchical configuration file, so as to obtain the corresponding configuration result, and then adds the corresponding preset prompt word to the configuration result to generate the structured input data. Then, the structured input data is input into the preset large language model, and the model is used to perform a preset multi-dimensional text understanding and analysis operation to identify the hierarchical structure therein, and generate a target hierarchical configuration file based on this hierarchical structure. After that, the document to be analyzed is obtained, and the document to be analyzed and the target hierarchical configuration file are input into the preset large language model together. The model is used to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file, and finally the target analysis result and the corresponding model reasoning process information are obtained. In this way, the present application realizes the automated parsing and processing of data management details without the need for manual intervention analysis, which significantly improves management efficiency. By establishing a topological relationship model at the document category level through the hierarchical reasoning mechanism, the hierarchical attribution of the category and level of the document can be systematically determined, and the visual presentation and logical traceability of the classification results can be realized. In response to the interpretability defects of the output results of the large language model, on the basis of ensuring the accuracy of data category and level identification, relying on the long thinking chain technology of the large model, this system provides a reason explanation and a complete thinking process for users to check and approve step by step, thereby enhancing the authority and accuracy of the judgment results. It supports real-time connection with third-party databases. When encountering specific terms and conditions, the intelligent agent can initiate multi-source data verification requests. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0042] Figure 1 This is a system architecture diagram of a document data analysis method for a large language model disclosed in this application;
[0043] Figure 2 This is a flow chart of a document data analysis method for a large language model disclosed in this application;
[0044] Figure 3 A schematic diagram of a specific data management rules parsing service method disclosed in this application;
[0045] Figure 4 A schematic diagram of a specific data-level reasoning service method disclosed in this application;
[0046] Figure 5 A business flow chart of a specific large language model document data analysis method disclosed in this application;
[0047] Figure 6 This is a schematic diagram of the structure of a document data analysis device for a large language model disclosed in this application;
[0048] Figure 7 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0050] At present, current data classification and grading technologies mainly rely on methods such as keyword matching, semantic similarity calculation and deep neural network classification. Although new solutions based on large-scale pre-trained language models have emerged in recent years, their practical application still faces significant technical challenges. Taking the management of corporate trade secret data as an example, the company's data management regulations usually cover a large number of clauses in multiple dimensions such as technical secrets, business information, customer data, personal privacy, etc., and need to be frequently updated with laws, regulations and corporate strategy adjustments. The existing classification and grading system based on large language models is still highly dependent on manually written prompt words, and there are prominent problems such as high manual maintenance costs and delayed system response. In order to solve the above technical problems, the present application discloses a document data analysis method, device, equipment and medium for a large language model, which can realize intelligent classification and dynamic grading of data.
[0051] The system framework used in the document data analysis method based on the large language model of this application can be found in detail. Figure 1As shown in the figure, the system architecture adopts a layered design pattern, which consists of the engine layer, capability layer, service layer and product layer from bottom to top.
[0052] The engine layer, as the technical foundation of the system, primarily manages basic resources and supports core algorithms. The core of the engine layer is the large language model, which provides the capability layer with technical capabilities such as text classification, text generation, logical reasoning, and natural language understanding.
[0053] The capability layer is built on top of the engine layer. The capability layer's document reading service is responsible for reading user-uploaded documents, the data management rules parsing service is responsible for converting user-uploaded company data management rules documents into hierarchical configuration files, and the data hierarchical reasoning service implements the core functionality of document hierarchical reasoning.
[0054] The top layer is the product layer, the final user interface presented by the system. It utilizes a web-based interactive architecture to achieve visual operation throughout the entire process. This layer integrates front-end display technology with back-end service call interfaces, providing core functional modules such as document upload, processing progress monitoring, grading result visualization, and configuration file modification, forming a complete data classification and grading intelligent system solution.
[0055] See also Figure 2 As shown, the embodiment of the present invention discloses a document data analysis method based on a large language model, comprising:
[0056] Step S11: determine a preset data management rule file, perform a preset hierarchical configuration operation on the preset data management rule file based on a preset hierarchical configuration file establishment method to obtain a corresponding configuration result, and then add a corresponding preset prompt word to the configuration result to generate corresponding structured input data.
[0057] In this embodiment, first obtain the preset data management rules file that needs to be analyzed, and then perform the preset hierarchical configuration operation on the preset data management rules file according to the preset hierarchical configuration file establishment method, wherein the preset hierarchical configuration file establishment method is to manually establish a hierarchical configuration file or automatically establish a hierarchical configuration file. If the preset hierarchical configuration file establishment method is to manually establish a hierarchical configuration file, the user terminal directly performs the preset hierarchical configuration operation on the preset interactive page according to the data management rules document to obtain the corresponding configuration result; if the preset hierarchical configuration file establishment method is to automatically establish a hierarchical configuration file, the data management rules document is uploaded, and the preset hierarchical configuration operation is performed on the data management rules document to obtain the corresponding configuration result. Specifically, Figure 3As shown, there are two ways to configure the preset hierarchical configuration for the pre-set data management rules file: First, you can manually configure the hierarchical configuration file based on the company's data management rules document on the interactive page. Second, you can upload the company's data management rules document and call the document reading service and the data management rules parsing service to perform structured analysis on the rules content to automatically generate the corresponding configuration results. Then, after reading the document and adding preset prompts for the task description, structured input data that meets the processing specifications of the large language model is generated.
[0058] Step S12: input the structured input data into a preset large language model, and use the preset large language model to perform a preset multi-dimensional text understanding and analysis operation on the structured input data to identify the hierarchical structure in the structured input data, and generate a corresponding target hierarchical configuration file based on the hierarchical structure.
[0059] In this embodiment, after obtaining the structured input data, the structured input data is input into a preset large language model, and the preset large language model can be used to perform a preset multi-dimensional text understanding and analysis operation on the structured input data to obtain the data category in the structured input data; the structured input data is classified based on the data category to obtain the hierarchical structure in the structured input data, and a corresponding target hierarchical configuration file is generated based on the hierarchical structure. Specifically, the pre-processed structured input data is submitted to the large language model for multi-dimensional text understanding and analysis to ensure that the hierarchical structure in the data classification system can be accurately identified. After the large language model processing is completed, a target hierarchical configuration file that complies with predefined specifications will be automatically generated. For example:
[0060] {
[0061] "Technical Information": {
[0062] "Product R&D":None,
[0063] "Operation and Maintenance Management": None
[0064] },
[0065] "Business Information": {
[0066] "Personnel Management":None,
[0067] "Strategic Planning": None
[0068] }
[0069] }
[0070] The use of the None value follows strict semantic conventions, clearly identifying the node as a terminal classification unit that cannot be further subdivided, namely a leaf node, to ensure the rigor and enforceability of the hierarchical structure of the configuration file.
[0071] It is understood that after the large language model outputs the target-level configuration file, it is necessary to determine whether the target-level configuration file meets the preset configuration conditions; if the target-level configuration file does not meet the preset configuration conditions, the target-level configuration file is modified by the user terminal; if the target-level configuration file meets the preset configuration conditions, the step of inputting the document to be analyzed and the target-level configuration file into the preset large language model is triggered. Specifically, for the target-level configuration file generated by the detailed analysis, the user terminal can modify and confirm the operation through the interactive interface provided by the system. When the target-level configuration file does not meet the user terminal's requirements, the target-level configuration file can be modified.
[0072] Step S13: obtain the document to be analyzed, input the document to be analyzed and the target hierarchical configuration file into a preset large language model, and use the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain the target analysis result and the corresponding model reasoning process information.
[0073] In this embodiment, when the user needs to perform data identification, he / she uploads the document to be analyzed. The text to be analyzed uploaded by the user and the target level configuration file obtained by parsing are passed to the data level reasoning service, that is, a large language model is preset, and then, the current level nodes corresponding to each level are determined according to the hierarchical structure of the target level configuration file; the document to be analyzed is parsed using the current level nodes corresponding to each level to obtain the corresponding parsed data of each level and the corresponding model reasoning process information; the target analysis result corresponding to the document to be analyzed is determined based on the parsed data of each level. Specifically, Figure 4 As shown, the document to be identified and the target level configuration file are uploaded as input to the data level reasoning service. At the same time, the user provides a tool interface that can be called as needed. The control module serves as the scheduling center. It determines whether the current layer needs to expand nodes based on the level configuration file. When there are too many nodes, it triggers block processing. It also builds and distributes adaptation prompts for the reasoning nodes at each level based on the configuration. Based on the prompt of the control module, the current layer node calls the large language model to parse the document in parallel, and outputs the "classification result of the current layer", that is, the parsed data and model reasoning process information of each level. Finally, the target analysis result of the document to be analyzed is determined based on the parsed data of each level.
[0074] Furthermore, the long chain of thought technology can be used to decompose the steps of parsing the document to be analyzed using the preset large language model based on the hierarchical structure in the target hierarchical configuration file into various intermediate links; and preset verification and integration operations are performed on each intermediate link to obtain the corresponding model reasoning path. Specifically, combined with LoCoT (Long Chain of Thought), by presenting a systematic and step-by-step logical reasoning process, complex reasoning tasks are gradually deconstructed and a structured reasoning path is generated. The long chain of thought technology decomposes the hierarchical structure in the target hierarchical configuration file into multiple intermediate links by simulating the progressive reasoning mode of human thinking, and finally forms a rigorous and traceable model reasoning path through layer-by-layer verification and integration.
[0075] In addition, a preset tool interface can be used to perform a preset verification operation on the data in the document to be analyzed to obtain a target response result; based on the target response result, the preset large language model is used to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain a target analysis result and corresponding model reasoning process information. Specifically, the required third-party tool interface is dynamically called to perform a preset verification operation on the document to be analyzed. After obtaining the response result, the hierarchical configuration file is deeply analyzed to output the target analysis result and model reasoning process information. For example, the data management details provided by the user end may clearly require that "personal identity information" be limited to "personal identity information of corporate customers" and provide an identity verification interface and method of using the interface to query whether an individual is a company customer based on their name, i.e., a preset tool interface provided by a third party, to complete identity verification. At this time, if the large language model determines that the document to be analyzed involves personal identity information elements, it will automatically call the customer identity verification interface provided by the user end to initiate a query request through the name field. The interface response result will serve as the key basis for classification determination: when the returned result confirms that the individual corresponding to the name is a company customer, it will be ultimately determined that the information belongs to the business secret category of personal identity information; otherwise, the classification will be excluded.
[0076] As can be seen from the above, the embodiment of the present application first determines the preset data management rules file, and performs the preset hierarchical configuration operation on it according to the establishment method based on the preset hierarchical configuration file, thereby obtaining the corresponding configuration result, and then adds the corresponding preset prompt words in the configuration result to generate structured input data. The structured input data is then input into the preset large language model, and the preset multi-dimensional text understanding and analysis operation is performed on it with the help of the model to identify the hierarchical structure therein, and the target hierarchical configuration file is generated based on this hierarchical structure. After that, the document to be analyzed is obtained, and the document to be analyzed and the target hierarchical configuration file are input into the preset large language model together, and the model is used to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file, and finally the target analysis result and the corresponding model reasoning process information are obtained. In this way, the embodiment of the present application realizes the automated parsing and processing of data management rules, without the need for manual intervention analysis, and significantly improves management efficiency. By establishing a topological relationship model at the document category level through the hierarchical reasoning mechanism, the hierarchical attribution of the category and level to which the document belongs can be systematically determined, and the visual presentation and logical traceability of the classification results can be realized. To address the interpretability limitations of large language model outputs, while ensuring accurate identification of data categories and levels, this system leverages the model's long-term thought chain technology to provide explanations and a complete thought process for users to step-by-step review and approval, enhancing the authority and accuracy of judgment results. It also supports real-time integration with third-party databases. When encountering specific clause judgment requirements, the intelligent agent can initiate multi-source data verification requests.
[0077] Based on the previous embodiment, this application discloses a document data analysis method based on a large language model, which can realize intelligent classification and dynamic classification of data. At the same time, this application can be used in the field of medical health, such as systematic management of sensitive information such as patient medical records and clinical trial data. Figure 5 The document data analysis method based on the large language model shown is described in detail.
[0078] First, the hospital's medical records are used as the preset data management rules file. By calling the document reading service and the data management rules parsing service, a structured analysis of the rules content is performed to automatically generate the corresponding configuration results. Then, after document reading and adding preset prompts for task descriptions, structured input data that meets the processing specifications of the large language model is formed. The pre-processed structured input data is submitted to the large language model for multi-dimensional text understanding analysis to ensure that the hierarchical structure in the data classification system can be accurately identified and the target hierarchical configuration file is generated. The hierarchical node configuration file obtained by parsing is shown below:
[0079] {
[0080] "Personal basic information": None,
[0081] "Personally identifiable information": None,
[0082] "Personal biometric information":None,
[0083] "Network identity information":None,
[0084] "Personal health and physiological information": {
[0085] "Health Information":None,
[0086] "Personal medical information":None,
[0087] },
[0088] "Personal education and work information": {
[0089] "Personal education information":None,
[0090] "Personal work information":None,
[0091] },
[0092] }
[0093] Next, the user-uploaded text to be analyzed and the target level configuration file obtained by parsing are passed to the data level reasoning service, that is, the large language model is preset, and the user provides a tool interface that can be called as needed. The current layer node of the control module uses the long thinking chain technology to call the large language model to parse the text to be analyzed in parallel, output the parsed data and model reasoning process information of each level, and finally determine the patient's medical record based on the parsed data of each level. The configuration file of the hierarchical category description is as follows:
[0094] {“Basic Personal Information”: “Basic information about a natural person, such as name, birthday, age, gender, ethnicity, nationality, place of origin, marital status, family relationships, address, personal telephone number, email address, interests and hobbies, etc.”,
[0095] “Personal identity information”: “Information that can directly identify a natural person, such as ID card, military ID card, passport, driver’s license, work permit, access pass, social security card, residence permit, Hong Kong, Macau and Taiwan pass, etc., document numbers, document validity periods, and ID photos or photocopies.”
[0096] “Personal biometric information”: “Biometric raw information and comparison information, such as face, fingerprint, gait, voiceprint, gene, iris, handwriting, palm print, ear, eye print, etc.”
[0097] "Online identity information" means "information that can directly identify the identity of a network or communication user and account-related information, such as user account number, user ID, instant messaging account number, online social user account number, user profile picture, nickname, personal signature, IP address, account opening date, etc."
[0098] “Health information”: “General information related to personal health status, such as weight, height, body temperature, lung capacity, blood pressure, blood type, etc.”
[0099] “Personal medical information”: “Records related to personal illness and treatment, such as symptoms, hospitalization records, doctor’s orders, test reports, physical examination reports, surgery and anesthesia records, nursing records, medication records, drug and food allergy information, fertility information, past medical history, diagnosis and treatment, family medical history, current medical history, history of infectious diseases, smoking history, etc.”
[0100] “Personal education information”: “Information related to personal education and training, such as academic qualifications, degrees, educational experience, transcripts, qualification certificates, training records, etc.”
[0101] “Personal work information”: “Information related to personal job search and work status, such as personal occupation, position, title, work unit, work location, work experience, salary, work performance, resume, etc.”
[0102] }
[0103] Correspondingly, the generated medical record is as follows:
[0104] **MEDICAL RECORDS**
[0105] **Hospital Name**: **First Affiliated Hospital
[0106] **Department**: Respiratory Medicine
[0107] **Patient Name**:**
[0108] **Gender**: Male
[0109] Age: 35
[0110] **ID number**: 320586xxxxxxxxxxxx
[0111] **Contact number**: 138xxxxxxxx
[0112] **Home address**: No.******, ** Road, ** District, ** City
[0113] **Date of consultation**: **** year ** month ** day
[0114] **Outpatient number**: 202*******0123
[0115] ###**Chief Complaint**
[0116] Repeated cough and sputum accompanied by low fever for 3 days, aggravated by chest pain for 1 day.
[0117] History of Present Illness
[0118] Three days ago, the patient developed an irritating, dry cough without apparent cause, which was more pronounced at night and occasionally produced small amounts of white, sticky sputum. He also had a low-grade fever (37.5°C) in the afternoon. Treatment with cefixime and Tylenol did not alleviate his symptoms. One day ago, his cough worsened, producing yellow sputum (approximately 50 ml / day), and his temperature rose to 38.3°C. He also had tenderness on pressure in the right lower chest, but no hemoptysis, dyspnea, or impaired consciousness. He presented to our hospital today. A routine blood test revealed elevated white blood cell counts, and a chest CT scan revealed consolidation in the right lower lung.
[0119] Correspondingly, the model reasoning process information for generating medical records is as follows:
[0120] OK, now we need to analyze the medical record text provided by the user and classify it into six categories. First, I need to read the text carefully and then check whether each category is included one by one.
[0121] First, the document is titled "Medical Record," which includes the hospital name, department, patient name, gender, age, ID number, contact number, home address, consultation date, and clinic number. The Chief Complaint and Present Illness History section details the patient's symptoms, medical history, and examination results.
[0122] Next, I need to check each category individually:
[0123] 1. **Personal basic information**: This includes name, gender, and age. These are all basic information, so it should be "yes".
[0124] 2. **Personal identity information**: The ID number clearly appears in the text and is considered identity information, so the answer is “yes”.
[0125] 3. **Personal biometric information**: Biometrics usually refers to fingerprints, irises, etc. This information is not included in the text, so the answer is “No”.
[0126] 4. **Network identity information**: such as account number, password, etc. There is no relevant content here, so "No".
[0127] 5. **Personal health and physiological information**: The chief complaint and history of present illness describe in detail the patient's symptoms, body temperature, test results, and other health information. They fall into this category, so the answer is "yes."
[0128] 6. **Personal Education and Work Information**: The text does not mention educational background or work unit, so “No”.
[0129] Each judgment needs to be confirmed to ensure it meets the requirements. For example, an ID number is considered identity information, while symptoms and test results in health information fall under the health and physiology category. Other categories are not considered. Therefore, the final classification result should be "yes" for these three categories and "no" for the rest.
[0130] {
[0131] "category": {
[0132] "Personal basic information": "Yes",
[0133] "Personally Identifiable Information": "Yes",
[0134] "Personal biometric information": "No",
[0135] "Network Identity Information": "No",
[0136] "Personal health information": "Yes",
[0137] "Personal education and work information": "No"
[0138] },
[0139] "Inference Process": "1. The text is a complete medical record, including the patient's personal information, medical records, and health status.\n\n2. Basic Personal Information: The text clearly lists the patient's name, gender, and age, which constitutes basic identity information.\n\n3. Personal Identity Information: The ID number, as a legal identification number, is directly disclosed, meeting this category definition.\n\n4. Personal Biometric Information: Biometric data such as fingerprints and facial recognition features is not included.\n\n5. Online Identity Information: Online identity credentials such as email addresses, account numbers, and passwords are not included.\n\n6. Personal Health and Physiological Information: The chief complaint and current medical history include detailed records of temperature fluctuations (37.5°C / 38.3°C), cough symptoms, chest CT scan results, and other physiological health indicators.\n\n7. Personal Education and Work Information: The text does not mention relevant information such as education, occupation, or work unit."
[0140] }
[0141] In this way, structured and clear instructions in each module help the large language model better understand the requirements and goals of the task; parameter fine-tuning allows the large language model to better handle rule-based data classification and grading tasks; and through long-term thinking chain technology, the problem of unexplainable data category-level judgment results is effectively solved, and the complete thinking process of the large model is fully presented.
[0142] See also Figure 6 As shown, an embodiment of the present invention discloses a document data analysis device based on a large language model, comprising:
[0143] The data generation module 11 is configured to determine a preset data management rule file, perform a preset hierarchical configuration operation on the preset data management rule file based on a preset hierarchical configuration file establishment method to obtain a corresponding configuration result, and then add a corresponding preset prompt word to the configuration result to generate corresponding structured input data;
[0144] A file generation module 12 is configured to input the structured input data into a preset large language model, perform a preset multi-dimensional text understanding and analysis operation on the structured input data using the preset large language model, so as to identify a hierarchical structure in the structured input data, and generate a corresponding target hierarchical configuration file based on the hierarchical structure;
[0145] The result and process acquisition module 13 is used to obtain the document to be analyzed, input the document to be analyzed and the target hierarchical configuration file into a preset large language model, and use the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain the target analysis result and the corresponding model reasoning process information.
[0146] As can be seen from the above, this application first determines the preset data management rules file, and performs the preset hierarchical configuration operation on it according to the establishment method based on the preset hierarchical configuration file, so as to obtain the corresponding configuration result, and then adds the corresponding preset prompt words in the configuration result to generate structured input data. The structured input data is then input into the preset large language model, and the preset multi-dimensional text understanding and analysis operation is performed on it with the help of the model to identify the hierarchical structure therein, and the target hierarchical configuration file is generated based on this hierarchical structure. After that, the document to be analyzed is obtained, and the document to be analyzed and the target hierarchical configuration file are input into the preset large language model together, and the model is used to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file, and finally the target analysis result and the corresponding model reasoning process information are obtained. In this way, the present application realizes the automated parsing and processing of data management rules without the need for manual intervention analysis, which significantly improves management efficiency. By establishing a topological relationship model at the document category level through the hierarchical reasoning mechanism, the hierarchical attribution of the category and level to which the document belongs can be systematically determined, and the visual presentation and logical traceability of the classification results can be realized. To address the interpretability limitations of large language model outputs, while ensuring accurate identification of data categories and levels, this system leverages the model's long-term thought chain technology to provide explanations and a complete thought process for users to step-by-step review and approval, enhancing the authority and accuracy of judgment results. It also supports real-time integration with third-party databases. When encountering specific clause judgment requirements, the intelligent agent can initiate multi-source data verification requests.
[0147] In some specific implementations, the preset hierarchical configuration file is established by manually establishing the hierarchical configuration file or automatically establishing the hierarchical configuration file;
[0148] Accordingly, the data generation module 11 may specifically include:
[0149] A first configuration result obtaining unit is configured to, if the preset hierarchical configuration file establishment method is manual establishment of the hierarchical configuration file, directly perform the preset hierarchical configuration operation on the preset interactive page according to the data management rule document through the user terminal to obtain a corresponding configuration result;
[0150] The second configuration result obtaining unit is configured to upload the data management rule document if the preset hierarchical configuration file establishment mode is to automatically establish the hierarchical configuration file, and perform a preset hierarchical configuration operation on the data management rule document to obtain a corresponding configuration result.
[0151] In some specific implementations, the document data analysis device based on a large language model may further include:
[0152] A target level configuration file determination module is used to determine whether the target level configuration file meets a preset configuration condition;
[0153] A file modification module, configured to modify the target-level configuration file through a user terminal if the target-level configuration file does not meet a preset configuration condition;
[0154] The file input step triggering module is used to trigger the step of inputting the document to be analyzed and the target level configuration file into a preset large language model if the target level configuration file meets the preset configuration conditions.
[0155] In some specific implementations, the file generation module 12 may specifically include:
[0156] a data category acquisition unit, configured to perform a preset multi-dimensional text understanding and analysis operation on the structured input data using the preset large language model to obtain data categories in the structured input data;
[0157] The target hierarchical configuration file generating unit is configured to classify the structured input data based on the data category to obtain a hierarchical structure in the structured input data, and generate a corresponding target hierarchical configuration file based on the hierarchical structure.
[0158] In some specific implementations, the result and process acquisition module 13 may specifically include:
[0159] a current layer node determining unit, configured to determine the current layer node corresponding to each layer according to the hierarchical structure of the target layer configuration file;
[0160] A parsed data and reasoning process information acquisition unit, configured to parse the document to be analyzed using the current layer nodes corresponding to the respective layers, to obtain the parsed data of the respective layers and the corresponding model reasoning process information;
[0161] The target analysis result determining unit is configured to determine the target analysis result corresponding to the document to be analyzed based on the parsed data at each level.
[0162] In some specific implementations, the unit for acquiring parsed data and reasoning process information may specifically include:
[0163] a step decomposition subunit, configured to decompose the step of parsing the document to be analyzed using the preset large language model based on the hierarchical structure in the target hierarchical configuration file into various intermediate links using a long thought chain technique;
[0164] The model reasoning path acquisition subunit is used to perform preset verification and integration operations on each intermediate link to obtain the corresponding model reasoning path.
[0165] In some specific implementations, the result and process acquisition module 13 may specifically include:
[0166] A target response result obtaining unit, configured to perform a preset verification operation on the data in the document to be analyzed using a preset tool interface to obtain a target response result;
[0167] The result and process acquisition unit is used to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file using the preset large language model according to the target response result to obtain the target analysis result and the corresponding model reasoning process information.
[0168] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0169] Figure 7 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the document data analysis method for a large language model disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0170] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0171] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0172] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, NetWare, Unix, Linux, etc. In addition to including a computer program capable of implementing the document data analysis method of a large language model performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0173] Furthermore, this application discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned method for analyzing document data using a large language model. The specific steps of this method can be found in the corresponding content disclosed in the aforementioned embodiments and will not be further described here.
[0174] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0175] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0176] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0177] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0178] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A document data analysis method based on a large language model, characterized in that: include: Determining a preset data management rule file, performing a preset hierarchical configuration operation on the preset data management rule file based on a preset hierarchical configuration file establishment method to obtain a corresponding configuration result, and then adding a corresponding preset prompt word to the configuration result to generate corresponding structured input data; Inputting the structured input data into a preset large language model, performing a preset multi-dimensional text understanding and analysis operation on the structured input data using the preset large language model to identify a hierarchical structure in the structured input data, and generating a corresponding target hierarchical configuration file based on the hierarchical structure; Obtain the document to be analyzed, input the document to be analyzed and the target hierarchical configuration file into a preset large language model, and use the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain the target analysis result and the corresponding model reasoning process information.
2. The document data analysis method based on a large language model according to claim 1, characterized in that: The preset hierarchical configuration file is established in a manner of manually establishing the hierarchical configuration file or automatically establishing the hierarchical configuration file; Accordingly, the preset hierarchical configuration operation is performed on the preset data management rules file based on the preset hierarchical configuration file establishment method to obtain a corresponding configuration result, including: If the preset hierarchical configuration file is established manually, the user terminal directly performs the preset hierarchical configuration operation on the preset interactive page according to the data management rules document to obtain the corresponding configuration result; If the preset hierarchical configuration file establishment method is to automatically establish a hierarchical configuration file, the data management detail document is uploaded, and a preset hierarchical configuration operation is performed on the data management detail document to obtain a corresponding configuration result.
3. The document data analysis method based on a large language model according to claim 1, characterized in that: After generating the corresponding target level configuration file based on the hierarchical structure, the method further includes: Determining whether the target level configuration file meets a preset configuration condition; If the target level configuration file does not meet the preset configuration conditions, modifying the target level configuration file through the user terminal; If the target-level configuration file meets a preset configuration condition, the step of inputting the document to be analyzed and the target-level configuration file into a preset large language model is triggered.
4. The document data analysis method based on a large language model according to claim 1, characterized in that: The method of performing a preset multi-dimensional text understanding and analysis operation on the structured input data using the preset large language model to identify a hierarchical structure in the structured input data and generating a corresponding target hierarchical configuration file based on the hierarchical structure includes: Using the preset large language model to perform a preset multi-dimensional text understanding and analysis operation on the structured input data to obtain data categories in the structured input data; The structured input data is classified based on the data category to obtain a hierarchical structure in the structured input data, and a corresponding target hierarchical configuration file is generated based on the hierarchical structure.
5. The document data analysis method based on a large language model according to any one of claims 1 to 4, characterized in that: The using of the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain a target analysis result and corresponding model reasoning process information includes: Determine the current layer node corresponding to each layer according to the hierarchical structure of the target layer configuration file; Parsing the document to be analyzed using the current layer nodes corresponding to each layer to obtain the parsed data of each layer and the corresponding model reasoning process information; The target analysis result corresponding to the document to be analyzed is determined based on the parsed data at each level.
6. The document data analysis method based on a large language model according to claim 5, characterized in that: The document to be analyzed is parsed using the preset large language model based on the hierarchical structure in the target hierarchical configuration file to obtain corresponding model reasoning process information, including: Using the long thought chain technology, the step of parsing the document to be analyzed using the preset large language model based on the hierarchical structure in the target hierarchical configuration file is decomposed into various intermediate links; Perform preset verification and integration operations on each intermediate link to obtain the corresponding model reasoning path.
7. The document data analysis method based on a large language model according to any one of claims 1 to 4, characterized in that: The using of the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain a target analysis result and corresponding model reasoning process information includes: Using a preset tool interface to perform a preset verification operation on the data in the document to be analyzed to obtain a target response result; According to the target response result, the preset large language model is used to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain a target analysis result and corresponding model reasoning process information.
8. A document data analysis device based on a large language model, characterized in that: include: a data generation module, configured to determine a preset data management rule file, perform a preset hierarchical configuration operation on the preset data management rule file based on a preset hierarchical configuration file establishment method to obtain a corresponding configuration result, and then add a corresponding preset prompt word to the configuration result to generate corresponding structured input data; a file generation module, configured to input the structured input data into a preset large language model, perform a preset multi-dimensional text understanding and analysis operation on the structured input data using the preset large language model, so as to identify a hierarchical structure in the structured input data, and generate a corresponding target hierarchical configuration file based on the hierarchical structure; The result and process acquisition module is used to obtain the document to be analyzed, input the document to be analyzed and the target hierarchical configuration file into a preset large language model, and use the preset large language model to parse the document to be analyzed based on the hierarchical structure in the target hierarchical configuration file to obtain the target analysis result and the corresponding model reasoning process information.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the document data analysis method based on a large language model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program; wherein, when the computer program is executed by a processor, the document data analysis method based on a large language model as described in any one of claims 1 to 7 is implemented.