Structured text generation and quality control method based on knowledge-driven large model
Through a multi-round iterative optimization method combining knowledge graphs and ontology structures, the problem of quality control in electronic medical record systems is solved, and structured text generation with high reliability and interpretability is achieved, which is suitable for medical, legal and scientific research fields.
Patent Information
- Application Number
- CN202510510182.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
The existing electronic medical record system has problems such as misdiagnosis, missing key information, and logical contradictions in quality control. The quality of medical records in primary medical institutions is particularly outstanding, and the existing technology is difficult to effectively cover the update of complex semantics and adaptation guidelines in clinical medicine, resulting in low recognition accuracy and frequent generation errors.
A large language model based on knowledge is adopted, combining knowledge graphs and ontology structures, structured text generation and quality control are carried out through multiple iterative optimization processes, including entity recognition, difference retrieval, evidence retrieval and multiple iterative corrections to ensure the reliability and interpretability of the generated content.
It realizes a semantic-driven text structure that is controllable, has credible content and is standardized in structure, significantly reduces generation errors, improves the reliability and reviewability of quality control of electronic medical records, and is suitable for highly credible scenarios such as medical care, law, and scientific research.
Smart Images

Figure CN120387452A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text data processing, and particularly relates to a method for generating and quality controlling structured text based on a knowledge-driven large model. Background Art
[0002] With the continuous development of medical informatization in China, electronic medical records play an increasingly important role in clinical diagnosis and treatment. The electronic medical records contain non-explicit disease implicit information such as patient examination abnormalities, index change trends, and disease risk factors, which have important clinical value for clinical diagnosis and treatment, risk warning, and early screening. As a standardized information carrier throughout the entire process of diagnosis and treatment, electronic medical records have significantly improved the efficiency and quality of medical services through structured data entry, intelligent information retrieval, and cross-platform sharing mechanisms. The real-time data sharing function has broken the information silos and achieved multi-disciplinary collaborative diagnosis and treatment. The medical record mining based on big data analysis has provided a solid foundation for clinical decision support, medical quality control, and research data utilization.
[0003] However, the current electronic medical record system still faces many challenges in quality control in actual applications. First, there are quality problems in medical record records such as misdiagnosis, missing key information, and logical contradictions. Second, the system has insufficient intelligence and lacks a real-time quality monitoring mechanism. More seriously, the uneven distribution of medical resources in China has led to particularly prominent medical record quality problems in primary medical institutions: the professional levels of primary doctors vary widely, and there is a lack of effective quality control tools. These problems not only directly affect the accuracy of clinical decisions but may also lead to serious medical disputes.
[0004] The automatic quality inspection of electronic medical records, as a core part of medical informatization construction, plays a crucial role in ensuring the standardization, integrity, and clinical rationality of medical record data, which is directly related to medical quality control and the improvement of diagnosis and treatment efficiency. However, the current mainstream technical solutions still face significant technical bottlenecks in practical applications: Although the traditional method based on a rule engine has clear logical judgment ability, it highly depends on manually writing rules, making it difficult to effectively cover the complex semantic expressions and diverse diagnosis and treatment scenarios in clinical medicine, and even less able to adapt to the frequent updates of clinical guidelines and diagnosis and treatment specifications; The solution using traditional NLP models (such as BERT, etc.) has certain semantic understanding ability, but requires a large amount of professionally annotated medical data for supervised training. Medical data annotation is not only costly but also requires the in-depth participation of clinical experts, and the trained models often have problems with insufficient generalization ability, with a significant decline in recognition accuracy when facing rare disease terms or special clinical manifestations; Although the large language model technology that has emerged in recent years shows powerful natural language understanding and generation capabilities, it still has obvious limitations in the medical quality inspection scenario, including being prone to factual errors (hallucinations), and due to the timeliness limitations of the model training data, it is difficult to obtain and apply the latest medical knowledge and clinical guidelines.
[0005] In response to this, the inventor proposes a structured text generation and quality control method based on a knowledge-driven large model to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a structured text generation and quality control method based on a knowledge-driven large model to solve the problems raised in the above background technology.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] A structured text generation and quality control method based on a knowledge-driven large model includes the following steps:
[0009] S1. Construct a domain entity, attribute, and relationship network to obtain a knowledge graph; perform sentence splitting, denoising, and normalization processing on the original dialogue or document to obtain preprocessed text;
[0010] S2. Perform named entity recognition on the preprocessed text and align it with the entities in the knowledge graph, conduct entity recognition performance evaluation to obtain an entity mapping set; perform hypernym and hyponym and association relationship reasoning on the entity mapping set to obtain entity relationship pairs; construct a chained prompt template based on the entity relationship pairs to obtain an initial Prompt;
[0011] S3. Initial structured generation, which is used to input the Prompt into the large model to generate a structured text draft; calculate the entity coverage rate and template matching degree metrics for the structured text draft to obtain an evaluation result;
[0012] S4. Compare the evaluation result with a preset threshold to identify information missing and logic inconsistent items to obtain difference items; use the difference items to retrieve relevant evidence fragments in the knowledge graph or external library, and obtain an evidence set through Softmax normalization;
[0013] S5. Jointly construct a refined and corrected Prompt with the evidence set and the draft, obtain a corrected prompt by calculating the RAG joint probability; input the corrected prompt into the large model to rewrite the draft locally or globally to generate a corrected text;
[0014] S6. Repeat steps S3 to S5 for the corrected text, and use the termination judgment formula for multi-round closed-loop optimization until the corrected text meets the preset thresholds of semantic integrity and logical consistency, and output a high-quality text; apply multi-dimensional evaluation metrics to the high-quality text to generate a quality report.
[0015] Preferably, the formula for evaluating the entity recognition performance is:
[0016]
[0017] Among them, Epred: the set of entities automatically extracted by the model, diseases, examinations, medications;
[0018] Egold: the set of reference standard entities manually annotated;
[0019] ∣Epred∩Egold∣: the number of entities correctly recognized.
[0020] Preferably, the formula for Softmax normalization is:
[0021]
[0022] Among them, vq: the query vector constructed by the difference item H, composed of text or entity combinations;
[0023] vdi: the vector representation of the i-th piece of evidence in the knowledge base;
[0024] T: the temperature coefficient, which controls the distribution smoothness, and the smaller the value, the more focused;
[0025] k: the number of candidate evidence documents.
[0026] Preferably, the calculation expression for the RAG joint probability is:
[0027]
[0028] Among them, x: the original or pre-revision structured text draft;
[0029] di: the i-th relevant evidence retrieved from the knowledge base;
[0030] P(di|x): the matching degree between the text and the evidence;
[0031] P(y|x, di): the probability that the large model generates the verification suggestion or rewritten fragment y under the condition of introducing the evidence di.
[0032] Preferably, the termination judgment formula for the multi-round closed-loop optimization is:
[0033] t* = min{t|Recall(K (t) ) ≥ τr ∧ F1(K (t) ) ≥ τf}
[0034] Among them, K(t): the structured text generated in the t-th iteration;
[0035] Recall(K (t) ), F1(K (t) ): the quality indicators calculated according to the aforementioned formula one;
[0036] τr, τf: the preset entity recall rate and F1 threshold (such as 0.93, 0.92);
[0037] t*: the minimum number of iterations required to meet the quality standard.
[0038] Preferably, in the step S1, the domain ontology engineering and knowledge graph construction technology is adopted to collect multi-source terms and relationships and construct a knowledge graph covering entities-attributes-relationships;
[0039] The text preprocessing method combining rules and statistics is used to segment the original text, remove redundant symbols and unify professional terms to obtain the preprocessed text.
[0040] Preferably, in the step S2, the named entity recognition model based on ontology constraints is used to accurately extract industry entities in the preprocessed text and map the extraction results to the entities in the knowledge graph to obtain the entity mapping set.
[0041] Preferably, three indicators of entity coverage rate, template matching degree and semantic consistency are used to automatically evaluate the draft to obtain the evaluation result.
[0042] Preferably, based on the threshold comparison algorithm, the evaluation result is compared with the preset threshold in turn to identify missing entities, format deviations and logical conflict items.
[0043] Preferably, in step S6, the preset thresholds of entity coverage and semantic consistency are used as termination conditions, and the evaluation, difference extraction, retrieval enhancement and difference correction steps are cyclically executed until the generated text meets the preset quality requirements and high-quality text is output; five evaluation indicators, namely Precision, Recall, F1, semantic consistency and format compliance, are used to conduct a multi-dimensional comprehensive evaluation of high-quality text and generate a quality report.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] (1) The present invention deeply integrates knowledge graphs, ontology structures and large language models, and introduces structured knowledge constraints in the entire process of text generation. This makes the model output no longer rely solely on language pattern matching, but has precise control capabilities at the entity, relationship and context semantic levels, thereby realizing a semantically driven text construction with "controllable generation, credible content and standardized structure", including but not limited to automatic quality inspection of electronic medical records.
[0046] (2) This invention leverages the Retrieval-Augmented Generation (RAG) mechanism to provide a traceable knowledge basis for large models. Each piece of generated or revised content can be mapped to a clear knowledge source. This knowledge-based approach significantly reduces the model's "hallucination" phenomenon, enabling structured text generation with a clear and interpretable path, effectively enhancing the reliability and auditability of the generated results. This approach is particularly suitable for high-trust scenarios such as healthcare, law, and scientific research.
[0047] (3) This invention introduces a multi-round iterative optimization chain of "generate-verify-repair-regenerate," dynamically comparing entities, identifying logical conflicts, and retrieving supporting evidence in each round, guiding the large model to self-correct and semantically improve. This mechanism breaks through the limitations of the traditional "single-round generation as the end point" approach and implements a dynamic, adaptive optimization process, which can continuously improve output quality and consistency. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a flow chart of the structured text generation and quality control method based on the knowledge-driven large model of the present invention. DETAILED DESCRIPTION
[0049] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0050] Example 1:
[0051] Please refer to Figure 1 As shown, the structured text generation and quality control method based on the knowledge-driven large model includes the following steps:
[0052] S1. Construct a domain entity, attribute, and relationship network to obtain a knowledge graph; perform sentence splitting, denoising, and normalization on the original dialogue or document to obtain preprocessed text;
[0053] S2. Perform named entity recognition on the preprocessed text and align it with the entities in the knowledge graph, conduct entity recognition performance evaluation to obtain an entity mapping set; perform hypernym and hyponym and association relationship reasoning on the entity mapping set to obtain entity relationship pairs; construct a chained prompt template based on the entity relationship pairs to obtain an initial Prompt;
[0054] The formula for the entity recognition performance evaluation is:
[0055]
[0056] Among them, Epred: the set of entities automatically extracted by the model, diseases, examinations, medications;
[0057] Egold: the set of reference standard entities manually annotated;
[0058] ∣Epred∩Egold∣: the number of correctly recognized entities;
[0059] It is used to measure the recognition accuracy of the model for industry-specific terms, ensure that the medical entities extracted by the model cover key clinical information; accurately evaluate the phenomena of missed extraction, over-extraction, or mis-extraction by the model, and provide accurate input for subsequent relationship reasoning and generation;
[0060] S3. Initial structured generation, which is used to input the Prompt into the large model to generate a structured text draft; calculate the entity coverage rate and template matching degree indicators for the structured text draft to obtain an evaluation result;
[0061] S4. Compare the evaluation result with a preset threshold to identify missing information and logically inconsistent items to obtain difference items; use the difference items to retrieve relevant evidence fragments in the knowledge graph or external library, and obtain an evidence set through Softmax normalization;
[0062] The formula for the Softmax normalization is:
[0063]
[0064] Among them, vq: the query vector constructed by the difference item H, which is composed of text or entity combinations;
[0065] vdi: vector representation of the i-th piece of evidence in the knowledge base;
[0066] T: Temperature coefficient, controls the smoothness of the distribution, the smaller the value, the more focused;
[0067] k: the number of candidate evidence documents;
[0068] Implement knowledge retrieval driven by structured differential terms; improve the authority and interpretability of generated content, and significantly reduce the "model hallucination" phenomenon;
[0069] S5. Combine the evidence set and the draft to construct a refined revision prompt, and obtain a revision prompt by calculating the RAG joint probability; input the revision prompt into the large model to perform local or global rewriting on the draft to generate a revised text;
[0070] The calculation expression of the RAG joint probability is:
[0071]
[0072] Where, x: original or revised draft of structured text;
[0073] di: the i-th relevant evidence retrieved from the knowledge base;
[0074] P(di|x): the matching degree between text and evidence;
[0075] P(y|x,di): The probability that the large model generates a verification suggestion or rewrites the fragment y under the condition of introducing evidence di;
[0076] It is used to integrate multi-source evidence to drive high-quality content generation. It ensures that the generated content meets semantic logic while strictly adhering to the knowledge source; it provides a clear basis for each generated fragment, facilitating quality control and result traceability;
[0077] S6. Repeat steps S3 to S5 on the revised text, using multiple rounds of closed-loop optimization using a termination judgment formula, until the revised text meets preset thresholds for semantic integrity and logical consistency, and outputs a high-quality text; and generate a quality report for the high-quality text using multi-dimensional evaluation indicators.
[0078] The termination judgment formula of the multi-round closed-loop optimization is:
[0079] t*=min{t|Recall(K (t) )≥τr∧F1(K (t) )≥τf}
[0080] Where K(t): structured text generated in the tth round of iteration;
[0081] Recall(K (t) ), F1(K (t) ): The quality index calculated according to the aforementioned Formula 1;
[0082] τr, τf: Preset entity recall rate and F1 threshold (such as 0.93, 0.92);
[0083] t*: The minimum number of iterations required to meet the quality standard;
[0084] It is used to determine whether the termination condition is reached, thereby controlling the number of iterations of the generation process and improving efficiency; avoiding lack of quality or over-generation, and ensuring that the final output meets the structured specifications and knowledge integrity.
[0085] Specifically, in the step S1, domain ontology engineering and knowledge graph construction technologies are adopted to collect multi-source terms and relationships and construct a knowledge graph covering entity-attribute-relationship;
[0086] The text preprocessing method combining rules and statistics is adopted to perform sentence segmentation, remove redundant symbols and unify professional terms on the original text to obtain preprocessed text.
[0087] Specifically, in the step S2, an ontology-constrained named entity recognition model is used to accurately extract industry entities in the preprocessed text and map the extraction results to entities in the knowledge graph to obtain an entity mapping set.
[0088] Specifically, three indicators of entity coverage rate, template matching degree and semantic consistency are adopted to automatically evaluate the draft to obtain an evaluation result.
[0089] Specifically, a threshold comparison algorithm is used to sequentially compare the evaluation result with a preset threshold to identify missing entities, format deviations and logical conflict items.
[0090] Specifically, in step S6, the preset thresholds of entity coverage rate and semantic consistency are used as termination conditions to loop through the steps of evaluation, difference extraction, retrieval enhancement and difference correction until the generated text meets the preset quality requirements and output high-quality text; five evaluation indicators of Precision, Recall, F1, semantic consistency and format compliance are used to comprehensively evaluate the high-quality text in multiple dimensions and generate a quality report.
[0091] As can be seen from the above, the present invention deeply integrates the knowledge graph, ontology structure and large language model, introduces structured knowledge constraints in the whole process of text generation, so that the model output no longer only depends on language pattern matching, but has precise control ability at the entity, relationship and context semantic levels, thus realizing semantic-driven text construction of "controllable generation, credible content and standardized structure";
[0092] With the retrieval-augmented generation (RAG) mechanism, this method provides a traceable knowledge basis for large models, and each piece of generated or corrected content can be mapped to a clear knowledge source. By means of knowledge support, the phenomenon of model "hallucination" is significantly reduced, enabling the generation of structured text to have a clear interpretable path, effectively enhancing the reliability and reviewability of the generation results, and being particularly suitable for high-trust scenarios such as medical, legal, and scientific research;
[0093] Introduce a multi-round iterative optimization chain of "generate - verify - repair - regenerate". In each round, entities are dynamically compared, logical conflicts are identified, and evidence support is retrieved to guide the large model to self-correct and semantic improvement. This mechanism breaks through the limitation of the traditional "single-round generation as the end point", realizes a dynamic adaptive optimization process, and can continuously improve the output quality and consistency;
[0094] Through the collaborative design of prompt engineering and ontology constraints, this method realizes the intelligent structured transformation of unstructured text, including natural language conversations and medical records, automatically completes information extraction, format alignment, and hierarchical organization, not only improving the generation efficiency of structured documents but also providing a standardized data basis for downstream decision support, logical reasoning, and automatic analysis;
[0095] Construct a multi-dimensional quality evaluation index system, including entity coverage, format consistency, logical integrity, etc., to realize automatic, quantitative, and traceable evaluation of the quality of structured text. Compared with the method relying on manual experience judgment, this solution significantly improves the objectivity, accuracy, and executability of quality control;
[0096] By modularly combining mechanisms such as ontology-driven, RAG retrieval, and prompt construction, the method has good domain transferability and template flexibility. Just replace the knowledge graph and prompt structure, and it can quickly adapt to the structured text generation and verification requirements in different fields such as medical, financial, government affairs, and education.
[0097] Example Two:
[0098] Quality control of medical records for "chronic cough" in the Department of Respiratory Medicine
[0099] 1. Data acquisition
[0100] Medical record samples: Randomly select 50 outpatient electronic medical records of "chronic cough" in the Department of Respiratory Medicine of our hospital;
[0101] Manual annotation: Two deputy chief physicians perform entity annotation according to the "Diagnosis and Treatment Guidelines for Respiratory Diseases in China (2024)";
[0102] Knowledge source:
[0103] "Diagnosis and Treatment Guidelines for Respiratory Diseases in China (2024)"
[0104] In-hospital Vectorized Index Database of 500 Historical "Chronic Cough" Cases
[0105] 2. Parameter Settings
[0106] Threshold of Entity Recognition Model: Confidence ≥ 0.75
[0107] Number of Candidates for Vector Retrieval k = 3
[0108] Softmax Temperature T = 0.1
[0109] Quality Threshold: Recall ≥ 0.93, F1 ≥ 0.92
[0110] 3. Entity Recognition Metrics are Shown in Table 1 Below:
[0111] Table 1
[0112]
[0113]
[0114] Evidence Retrieval Probability
[0115] Calculate Cosine Similarity:
[0116] sim(q,d) = {0.85, 0.78, 0.73}
[0117] Normalized Probability:
[0118]
[0119] Joint Generation Probability (let p(y∣x,d) = {0.90, 0.85, 0.80}:
[0120] Multi-round Iterative Convergence is Shown in Table 2 Below
[0121] Table 2
[0122]
[0123]
[0124] As can be seen from the above, the high-quality standard is achieved with only 3 rounds of iteration, reducing the average manual proofreading workload by 60% compared to single-round generation; the RAG joint probability of 0.8694 significantly suppresses the "hallucination" output.
[0125] Example 3:
[0126] Quality Control of "Chest Pain" Medical Records in the Department of Cardiology
[0127] 1. Data Acquisition
[0128] Medical record samples: 60 "chest pain" hospitalization records in the Department of Cardiology of our hospital were selected;
[0129] Manual annotation: Three experts in the Department of Cardiology performed entity and logic annotation according to the "Cardiology Guidelines (2023)";
[0130] Knowledge sources:
[0131] "Cardiology Guidelines (2023)"
[0132] National Pharmacopoeia of Medical Standards
[0133] 2. Parameter Setting
[0134] Confidence threshold of entity recognition model: ≥0.80
[0135] Number of candidates for vector retrieval k = 4
[0136] Softmax temperature T = 0.2
[0137] Quality threshold: Recall ≥ 0.93, F1 ≥ 0.92
[0138] 3. Entity recognition metrics are shown in Table 3 below;
[0139] Table 3
[0140]
[0141]
[0142] Evidence retrieval probability
[0143] Similarity {0.92, 0.88, 0.75, 0.70}, obtained with temperature T = 0.2
[0144] p(di∣q) ≈ {0.385, 0.303, 0.187, 0.125}
[0145] Conditional generation probability p(y∣x,d) = {0.93, 0.89, 0.85, 0.80}
[0146] Joint generation probability
[0147]
[0148] Multi-round iteration convergence is shown in Table 4 below:
[0149] Table 4
[0150]
[0151] Among them, only 1 round of iteration is required to meet the high-quality standard; the joint probability of RAG is 0.9039, effectively improving the credibility of the verification suggestions, and the overall quality inspection efficiency is improved by about 75%.
[0152] As can be seen from the above, both the second and third embodiments quickly converge to high-quality structured text output in different clinical scenarios. The key indicators Recall, F1, and the joint probability of RAG are all far ahead of the traditional single-round generation method, fully verifying the significant advantages of this solution in terms of semantic integrity, logical consistency, and generation credibility.
[0153] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0154] In the drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments are involved, and other structures can refer to the general design. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other.
[0155] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for generating and quality controlling structured text based on a knowledge-driven large model, characterized in that It includes the following steps: S1. Construct a domain entity, attribute, and relationship network to obtain a knowledge graph; perform sentence splitting, denoising, and normalization on the original dialogue or document to obtain preprocessed text; S2. Perform named entity recognition on the preprocessed text and align it with the entities in the knowledge graph, conduct entity recognition performance evaluation to obtain an entity mapping set; perform hypernym and hyponym and association relationship reasoning on the entity mapping set to obtain entity relationship pairs; construct a chained prompt template based on the entity relationship pairs to obtain an initial Prompt; S3. Initial structured generation, which is used to input the Prompt into a large model to generate a structured text draft; Calculate entity coverage and template matching degree metrics for the structured text draft to obtain an evaluation result; S4. Compare the evaluation result with a preset threshold to identify information missing and logically inconsistent items to obtain difference items; It is used to retrieve relevant evidence fragments in the knowledge graph or external library according to the difference items, and obtain an evidence set through Softmax normalization; S5. Jointly construct a refined and corrected Prompt with the evidence set and the draft, and obtain a corrected prompt by calculating the RAG joint probability; Input the corrected prompt into the large model to locally or globally rewrite the draft to generate a corrected text; S6. Repeat steps S3 to S5 for the corrected text, and use the termination judgment formula for multi-round closed-loop optimization until the corrected text meets the preset thresholds of semantic integrity and logical consistency, and output a high-quality text; Apply multi-dimensional evaluation metrics to the high-quality text to generate a quality report.
2. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, wherein The formula for the entity recognition performance evaluation is: where Epred: the set of entities automatically extracted by the model, diseases, examinations, medications; Egold: the set of reference standard entities manually annotated; ∣Epred∩Egold∣: the number of correctly recognized entities; It is used to measure the recognition accuracy of the model for industry-specific terms, ensure that the medical entities extracted by the model cover key clinical information; accurately evaluate the phenomena of missed extraction, over-extraction, or mis-extraction by the model, and provide accurate input for subsequent relationship reasoning and generation.
3. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, characterized in that The formula for the Softmax normalization is: where vq: the query vector constructed by the difference item H, composed of text or entity combinations; vdi: the vector representation of the i-th evidence in the knowledge base; T: the temperature coefficient, which controls the distribution smoothness, and the smaller the value, the more focused; k: the number of candidate evidence documents.
4. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, characterized in that, The calculation expression for the RAG joint probability is: where x: the original or pre-revision structured text draft; di: the i-th relevant evidence retrieved from the knowledge base; P(di∣x): the matching degree between the text and the evidence; P(y∣x,di): the probability that the large model generates a verification suggestion or a rewritten fragment y under the condition of introducing the evidence di.
5. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, characterized in that The termination judgment formula for the multi-round closed-loop optimization is: t* = min{t|Recall(K (t) ) ≥ τr ∧ F1(K (t) ) ≥ τf} where K(t): the structured text generated in the t-th iteration; Recal l(K (t) ),F1(K (t) ): The quality index calculated according to the foregoing Formula 1; τr,τf: the preset entity recall rate and F1 threshold; t*: the minimum number of iterations required to meet the quality standard.
6. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, characterized in that In the step S1, the domain ontology engineering and knowledge graph construction technology are adopted to collect multi-source terms and relationships and construct a knowledge graph covering entities-attributes-relationships; The rule-and-statistics hybrid text preprocessing method is adopted to perform sentence segmentation, remove redundant symbols and unify professional terms on the original text to obtain preprocessed text.
7. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, characterized in that In the step S2, the named entity recognition model based on ontology constraints is used to accurately extract industry entities in the preprocessed text and map the extraction results to the entities in the knowledge graph to obtain an entity mapping set.
8. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, characterized in that, Three indicators of entity coverage rate, template matching degree and semantic consistency are adopted to automatically evaluate the draft to obtain an evaluation result.
9. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, characterized in that Based on the threshold comparison algorithm, the evaluation result is compared with the preset threshold in turn to identify missing entities, format deviations and logical conflict items.
10. The structured text generation and quality control method based on a knowledge-driven large model according to claim 1, characterized in that, In step S6, the preset thresholds of entity coverage rate and semantic consistency are used as termination conditions to loop through the evaluation, difference extraction, retrieval enhancement and difference correction steps until the generated text K meets the preset quality requirements and output high-quality text; Five evaluation indicators of Precision, Recall, F1, semantic consistency and format compliance are used to comprehensively evaluate the high-quality text in multiple dimensions and generate a quality report.
Citation Information
Cited By
Generation method for improving quality of content generated by RAG technology based on semantic matching
CN121599129A
Progressive text quality improvement method and device based on large language model and medium
CN122133660A
A large model-based text generation optimization system
CN122693645A