Weakly supervised knowledge graph automatic construction method for professional text

By converting professional text into a structured format and extracting entities, relationships, and attributes using a large language model, identifying pattern conflicts, and dynamically updating the domain graph pattern, the problem of high cost and rigid patterns in traditional knowledge graph construction is solved, achieving low-cost and efficient construction of professional text knowledge graphs.

CN121390253BActive Publication Date: 2026-04-14STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional professional text knowledge graph construction relies on strongly supervised learning, which is costly and difficult to adapt to the dynamic update requirements of professional knowledge. Existing weakly supervised methods suffer from rigid preset patterns and lack of effective quality assurance, resulting in problems such as redundant entities, incorrect relationships, and logical conflicts in the knowledge graph.

Method used

Professional texts are converted into structured text formats. Entities, relationships, and attributes are extracted in a weakly supervised manner using a large language model. Pattern conflicts are identified, and the domain graph pattern is dynamically updated through a multi-dimensional confidence evaluation mechanism to optimize and reconstruct the knowledge graph until a preset quality threshold is reached.

Benefits of technology

It achieves low-cost and high-efficiency knowledge graph construction, adapts to emerging knowledge types in professional texts, ensures the dynamic updating capability and adaptability of the knowledge graph to professional scenarios, and outputs accurate and comprehensive core knowledge of professional texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121390253B_ABST
    Figure CN121390253B_ABST
Patent Text Reader

Abstract

The application provides a weakly supervised knowledge graph automatic construction method for professional text, relates to the technical field of knowledge graph construction, and comprises the following steps: converting a professional text into a structured text format to generate a preprocessed text set; based on a preset domain graph pattern, constructing a structured knowledge unit by using a large language model, identifying a pattern conflict, and generating a candidate new pattern description document; evaluating the comprehensive confidence of the candidate new pattern description document through a multi-dimensional confidence evaluation mechanism, and dynamically updating the domain graph pattern; optimizing and reconstructing the structured knowledge unit, calculating the quality index of the reconstructed weakly supervised knowledge graph, repeatedly constructing the weakly supervised knowledge graph until the quality index of the weakly supervised knowledge graph reaches a preset quality stability threshold, and obtaining a weakly supervised knowledge graph of a professional text. The technical problem that the preset domain graph pattern in the prior art is rigid and difficult to adapt to emerging knowledge types in the construction of a weakly supervised knowledge graph is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge graph construction, and more particularly to an automated method for constructing weakly supervised knowledge graphs for specialized texts. Background Technology

[0002] In professional fields such as law, medicine, and finance, professional texts carry a massive amount of core knowledge with rigorous structure and complex semantics. This knowledge is an important foundation for decision support and intelligent analysis in these fields, and there is an urgent need to achieve structured presentation and efficient utilization through knowledge graphs.

[0003] However, traditional methods for constructing specialized textual knowledge graphs often rely on strongly supervised learning, requiring significant manpower for data annotation and rule formulation. This approach is not only costly and time-consuming but also struggles to keep pace with the dynamic updates needed for professional knowledge as industries evolve. Meanwhile, while existing weakly supervised methods reduce the dependence on annotation, they generally suffer from rigid pre-defined domain graph models and a lack of effective quality assurance and iteration mechanisms. This results in knowledge graphs prone to entity redundancy, relational errors, and logical conflicts, failing to meet the stringent requirements of professional domains for rigorous, accurate, and timely knowledge representation.

[0004] Therefore, there is an urgent need for an automated method for constructing weakly supervised knowledge graphs for professional texts, in order to solve the technical pain points in the construction of knowledge graphs for professional texts. Summary of the Invention

[0005] This invention addresses the technical problems in existing weakly supervised knowledge graph construction, such as rigid preset domain graph patterns that are difficult to adapt to emerging knowledge types and the lack of effective quality assessment and iteration mechanisms. It provides an automated method for constructing weakly supervised knowledge graphs for professional texts.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0007] This invention provides an automated method for constructing weakly supervised knowledge graphs for specialized texts, including:

[0008] Convert professional text into structured text format and generate a preprocessed text collection;

[0009] Based on a preset domain graph pattern, entities, relations and attributes are extracted from the preprocessed text in a weakly supervised manner using a large language model to construct structured knowledge units, and pattern conflicts between the structured knowledge units and the domain graph pattern are identified to generate candidate new pattern description documents.

[0010] The comprehensive confidence level of the candidate new pattern description documents is evaluated through a multi-dimensional confidence assessment mechanism, and the domain graph pattern is dynamically updated.

[0011] The structured knowledge units are optimized and reconstructed based on the updated domain graph pattern to construct a weakly supervised knowledge graph, and the quality index of the reconstructed weakly supervised knowledge graph is calculated.

[0012] Repeatedly construct the weakly supervised knowledge graph until the quality index of the weakly supervised knowledge graph reaches the preset quality stability threshold, and output the final professional text weakly supervised knowledge graph.

[0013] The beneficial effects of this invention are:

[0014] Compared to existing technologies, this application first converts professional text into a structured text format, generating a preprocessed text set. This preprocessed text set accurately adapts to the text input requirements of weakly supervised knowledge graph construction, improving the accuracy and efficiency of subsequent entity, relation, and attribute extraction stages. Secondly, based on a preset domain graph pattern, a large language model is used to extract entities, relations, and attributes from the preprocessed text in a weakly supervised manner, constructing structured knowledge units. Pattern conflicts between these structured knowledge units and the domain graph pattern are identified, generating candidate new pattern description documents. This effectively solves the problem of existing technologies struggling to automatically identify, evaluate, and incorporate new knowledge beyond the domain graph pattern in professional text, ensuring the dynamic updating capability and adaptability of the knowledge graph to professional scenarios. Thirdly, a multi-dimensional confidence evaluation mechanism assesses the comprehensive confidence of the candidate new pattern description documents, dynamically updating the domain graph pattern. This effectively balances the rigor of the domain graph pattern with its adaptability to emerging professional text scenarios, solving the problem of inaccurate pattern updates caused by single-dimensional or manual evaluation in existing technologies. Furthermore, the structured knowledge units are optimized and reconstructed based on the updated domain graph model. The quality index of the reconstructed weakly supervised knowledge graph is calculated to ensure that it can meet the practical application needs of professional texts for knowledge representation. Finally, the weakly supervised knowledge graph is repeatedly constructed until its quality index reaches a preset quality stability threshold. The final professional text weakly supervised knowledge graph is then output, resulting in a professional text weakly supervised knowledge graph that can accurately and comprehensively carry the core knowledge of the professional text, thus meeting the practical application needs of knowledge graphs in the professional domain.

[0015] Through the above technical solution, this application eliminates the reliance on large-scale manual annotation in strongly supervised methods, effectively solving the problems of high cost and low efficiency in traditional construction models. Simultaneously, by identifying conflicts between structured knowledge units and preset domain graph patterns and generating candidate new pattern description documents, combined with a multi-dimensional confidence evaluation mechanism, it achieves dynamic updates of the domain graph patterns, overcoming the pain points of rigid preset patterns and inability to adapt to emerging knowledge types in professional texts in existing weakly supervised methods. Furthermore, the updated domain graph patterns are used to optimize and reconstruct structured knowledge units and calculate quality indicators. Through iterative construction until the quality indicators reach a preset stable quality threshold, problems such as insufficient knowledge extraction accuracy and entity redundancy or relational errors in the graph are effectively avoided. The final output of a weakly supervised knowledge graph for professional texts combines the advantages of low cost and high efficiency with good adaptability to professional scenarios, and possesses stable accuracy, completeness, and low conflict rate, meeting the rigor and practicality requirements of knowledge graphs in professional fields such as law and medicine. Attached Figure Description

[0016] Figure 1 A flowchart illustrating the automated construction method for weakly supervised knowledge graphs for professional texts provided by this invention;

[0017] Figure 2 This is a schematic diagram illustrating the process of generating a preprocessed text set in the automated construction method for weakly supervised knowledge graphs for professional texts provided by this invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0020] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0021] Examples, such as Figure 1 As shown, this embodiment of the invention provides an automated method for constructing weakly supervised knowledge graphs for specialized texts, including:

[0022] S10: Convert professional text into structured text format and generate a preprocessed text collection.

[0023] Currently, professional texts, such as legal judgments and regulations, are mostly in unstructured or semi-structured formats such as PDF and DOCX. These formats are highly heterogeneous and contain redundant information, making it difficult to directly and accurately identify core content and semantic relationships during the knowledge graph construction process.

[0024] Meanwhile, traditional manual processing methods are inefficient and costly, while general text processing methods lack adaptability to the terminology and structural features of professional fields, and cannot meet the requirements of weakly supervised knowledge graph construction for the structuring and standardization of input text.

[0025] To address the aforementioned issues, this application converts professional text into a structured text format and generates a preprocessed text set.

[0026] Specifically, such as Figure 2 As shown, step S10 in the method includes:

[0027] Professional texts are uniformly converted into standardized Markdown or XML formats to form intermediate structured text;

[0028] The intermediate structured text is segmented using a semantic integrity segmentation algorithm. During the segmentation process, hierarchical structure information is extracted, including chapter path information, semantic type information, and reference relationship information.

[0029] Based on the chapter path information, semantic type information, and reference relationship information, add chapter path markers, semantic type markers, and reference relationship markers to each text block, and assign a unique text block identifier to each text block;

[0030] The text blocks, which have been marked with chapter path tags, semantic type tags, reference relationship tags, and text block identifiers, are organized into a preprocessed text set.

[0031] In this embodiment, professional text is first uniformly converted into standardized Markdown or XML format to form intermediate structured text. Specifically, the format type of the professional text is first identified, such as PDF, DOCX, TXT, OCR text converted from scanned documents, etc. Then, format conversion tools, such as PDF to XML tools and DOCX to Markdown tools, are used to uniformly convert professional text of different formats into Markdown or XML format. During the conversion process, the structural features of the professional text, such as chapter titles, paragraph separators, and lists, must be preserved to avoid losing key format information, ultimately forming intermediate structured text with a unified format and parsable structure.

[0032] Among them, the Markdown format is adapted to lightweight text hierarchical expression, which can quickly represent the chapter level of professional text through concise markup symbols, such as the cause of action, plaintiff's claims, and the court's opinion in legal texts, balancing readability and structure; the XML format is adapted to fine-grained structured tag definitions, which can accurately encapsulate the core semantic units of professional text through custom semantic tags, meeting the fine-grained requirements of subsequent weakly supervised extraction for text structure parsing; converting professional text into standardized Markdown or XML formats can eliminate the heterogeneity problem caused by different source formats, providing a uniform and parsable intermediate structured text carrier for subsequent steps such as block processing.

[0033] For example, if the professional text is a civil tort judgment document in PDF format, the scanned PDF text can first be processed by OCR character correction to correct key content such as incorrectly identified legal terms, clause numbers, or party information; then, the corrected text can be uniformly converted into XML format using a format conversion tool, and the text content can be structurally encapsulated using custom legal text semantic tags to finally form an intermediate structured text with uniform format and accurate content.

[0034] Secondly, the intermediate structured text is segmented using a semantic integrity segmentation algorithm. During segmentation, hierarchical structure information is extracted, including chapter path information, semantic type information, and citation relationship information. Specifically, a semantic integrity segmentation algorithm model optimized for professional texts is first loaded. When the algorithm is executed, it first uses periods, semicolons, and newlines as basic segmentation points to split the intermediate structured text into initial text fragments. Then, the similarity of the keywords in each initial text fragment is calculated. For example, the semantic vector similarity of words within a fragment is calculated based on the LawBERT model. Adjacent fragments with keyword similarity greater than or equal to a preset similarity threshold are merged to ensure that each final text block contains only one core semantic topic, such as evidence acceptance explanations or legal clause application analysis, avoiding semantic fragmentation. During segmentation, hierarchical structure information is simultaneously extracted from the context and structured tags of each text block, specifically including chapter path information, semantic type information, and citation relationship information, ensuring that the location, semantic attributes, and associated information of each text block are traceable. The preset similarity threshold can be dynamically set according to specific application scenarios and actual needs. Preferably, this application recommends setting the preset similarity threshold to 0.85, which takes into account both screening accuracy and screening efficiency.

[0035] The semantic integrity segmentation algorithm refers to a segmentation algorithm optimized for the semantic features of professional texts. Its core logic is based on basic segmentation and semantic merging, integrating domain-specific knowledge such as professional terminology dictionaries and domain text structure templates to avoid splitting core semantic units and ensure the semantic integrity of each text block. The chapter path information refers to the hierarchical position path of the text block in the original professional text, such as a civil judgment in a legal context—"This Court Holds"—"Applicable Law Section"—"Citation Analysis of Article 1194 of the Civil Code." The semantic type information refers to the core semantic category of the text block, such as the description of the cause of action, basic information of the parties, explanation of evidence acceptance, legal clause citation analysis, and judgment result determination in a legal context. The citation relationship information refers to the external resource information cited in the text block or the association identifier of other internal text blocks, such as citing Article 1194 of the Civil Code or referring to the evidence description text block in paragraph 5 of this judgment in a legal context.

[0036] For example, for the intermediate structured text converted to XML format, the semantic integrity segmentation algorithm first locates the "Court's Opinion" section and splits it into three initial text segments: fact finding, application of law, and liability allocation. Then, it calculates the semantic similarity of keywords between each segment, finding that the similarity between fact finding and application of law is 0.7, and the similarity between application of law and liability allocation is 0.68, both below the preset similarity threshold of 0.85. Therefore, it determines that the core semantics of the three segments are independent and retains them as three independent text blocks. Simultaneously, the hierarchical structure information of each text block is extracted during the segmentation process. Taking the application of law text block as an example, the extracted chapter path information is Civil Judgment Document - Court's Opinion - Application of Law; the semantic type information is legal clause citation; and the citation relationship information is citation of Article 1194 of the Civil Code.

[0037] For example, a semantic integrity segmentation algorithm model adapted to professional scenarios can be constructed based on the semantic features of professional texts. Taking the legal scenario as an example: First, a large number of diverse legal professional texts are collected as a sample training dataset. Simultaneously, the sample training data is finely annotated. For instance, a combination of manual annotation and machine pre-annotation can be used to accurately mark semantic separation features, such as landmark semantic demarcation sentences like "This court believes," "In summary," and "In accordance with the preceding paragraph," as well as implicit separation features such as paragraph breaks after legal clause citations, end markers of evidence listing, and starting sentences of liability division, forming a sample supervision label set. Second, a semantic integrity segmentation algorithm model is constructed based on the LawBERT pre-trained model in the legal domain. By embedding a legal terminology dictionary, the semantic encoding layer of the model is optimized, enabling it to accurately identify professional semantic relationships in legal texts. Next, the sample training dataset and the corresponding sample supervision label set are divided into a training set and a validation set in an 8:2 ratio. The training set is input into the semantic integrity segmentation algorithm model for iterative training. During training, the semantic similarity threshold of the model is dynamically adjusted using the semantic integrity score of text blocks as an indicator, and the segmentation effect is verified in real time through the validation set to correct the model parameters. The final trained model is a semantic integrity segmentation algorithm model for legal texts. This model can accurately identify semantic boundaries such as evidence description, legal application, fact-finding, and liability division, avoiding the splitting of complete legal semantic units.

[0038] Next, based on the chapter path information, semantic type information, and citation relationship information, chapter path tags, semantic type tags, and citation relationship tags are added to each text block, and a unique text block identifier is assigned to each text block. Specifically, the chapter path tag refers to converting the chapter path information into standardized encoding; the semantic type tag refers to converting the semantic type information into standardized encoding; and the citation relationship tag refers to converting the citation relationship information into standardized encoding.

[0039] For example, a tagging and encoding rule for professional texts can be established first. Taking a legal scenario as an example: chapter path tags can adopt an encoding format of domain identifier, document type, chapter level, and sequence number; semantic type tags can adopt a legal semantic category abbreviation encoding; and citation relationship tags can adopt a format of citation type and citation object encoding. Based on the tagging and encoding rule, the extracted chapter path information, semantic type information, and citation relationship information are converted into corresponding chapter path tags, semantic type tags, and citation relationship tags, respectively, and embedded into the non-text area of ​​each text block. At the same time, a unique text block identifier is assigned to each text block. The text block identifier can adopt a globally unique format of text block type, timestamp, and random number to ensure that different professional texts and text blocks processed in different batches will not have duplicate identifiers, and that the identifier can be associated with all tagging information of the text block.

[0040] For example, taking a legal scenario as an example, based on the tagging and encoding rules, the chapter path information: Civil Judgment - This Court's Opinion - Application of Law can be transformed into the corresponding chapter path tag: FL-MSPJ-BYSR-FLYS-001, where FL = legal field, MSPJ = civil judgment, BYSR = this court's opinion, FLYS = application of law, and 001 = serial number; the semantic type information: legal clause citation can be transformed into the corresponding semantic type tag: FLTKYY; the citation relationship information: citation of Article 1194 of the Civil Code can be transformed into the corresponding REF-FLTF-MF-1194; at the same time, a unique text block identifier is assigned to the text block, such as TB-20250520-897654; finally, the chapter path tag, semantic type tag, and citation relationship tag are embedded into the non-text area of ​​each text block.

[0041] Finally, the text blocks, with added chapter path markers, semantic type markers, reference relationship markers, and text block identifiers, are organized into a preprocessed text collection. Specifically, each text block can be stored in the order of text block identifier, chapter path marker, semantic type marker, reference relationship marker, and text content, ensuring that the data and text content are decoupled but associative. Multiple text blocks are then integrated into a unified preprocessed text collection, which can be stored in a serializable format such as JSONL or CSV.

[0042] In summary, compared to existing technologies, this application converts professional text into a structured text format and generates a preprocessed text set. This effectively eliminates the format heterogeneity of professional text, transforming unstructured / semi-structured raw text into a standardized, semantically complete, and parsable structured text form. The generated preprocessed text set can accurately adapt to the text input requirements of weakly supervised knowledge graph construction, improving the accuracy and processing efficiency of subsequent entity, relation, and attribute extraction stages.

[0043] S20: Based on a preset domain graph pattern, entities, relationships, and attributes are extracted from the preprocessed text in a weakly supervised manner using a large language model to construct structured knowledge units, and pattern conflicts between the structured knowledge units and the domain graph pattern are identified to generate candidate new pattern description documents.

[0044] In professional text fields such as law and medicine, the knowledge system is both rigorous and constantly updated with industry development. New knowledge generated by a large number of emerging scenarios often exceeds the coverage of the pre-set domain map model.

[0045] However, when dealing with such problems, existing technologies either rely on manual identification of out-of-pattern knowledge one by one, which is inefficient and prone to bias due to subjective judgment; or they lack standardized automated mechanisms, which cannot accurately complete the process of identifying new out-of-pattern knowledge and updating the knowledge graph. As a result, valuable new knowledge in professional texts is difficult to be efficiently and systematically absorbed into the existing knowledge graph, which restricts the dynamic updating capability of the graph and fails to meet the professional field's requirements for the timeliness and completeness of knowledge.

[0046] To address the aforementioned issues, this application, based on a preset domain graph pattern, utilizes a large language model to extract entities, relationships, and attributes from the preprocessed text in a weakly supervised manner, constructs structured knowledge units, identifies pattern conflicts between the structured knowledge units and the domain graph pattern, and generates candidate new pattern description documents.

[0047] Specifically, step S20 in the method includes:

[0048] A weak supervision prompt template is constructed based on a preset domain graph pattern, wherein the weak supervision prompt template includes a set of entity type definitions, a set of relation type definitions, and a set of attribute definitions;

[0049] The preprocessed text set and the weakly supervised prompt template are input into the large language model to obtain the entity extraction result set, the relation extraction result set and the attribute extraction result set, and the extraction confidence score of each extraction result in the entity extraction result set and the relation extraction result set is recorded.

[0050] The attribute value format is validated on the attribute extraction result set to generate an attribute validity score;

[0051] Filter out extraction results with a confidence score lower than the preset extraction confidence threshold and an attribute validity score lower than the preset attribute validity threshold;

[0052] Based on the filtered entity extraction results, relation extraction results, and attribute extraction results, a structured knowledge unit containing a qualified entity set, a qualified relation set, and a qualified attribute set is constructed.

[0053] Identify pattern conflicts between the structured knowledge units and the domain graph patterns, and generate candidate new pattern description documents.

[0054] In this embodiment, a weakly supervised prompt template is first constructed based on a preset domain graph pattern. This template includes a set of entity type definitions, a set of relationship type definitions, and a set of attribute definitions. The domain graph pattern refers to a structured and standardized framework for a knowledge graph defined for a specific professional domain. For example, in a legal scenario, the domain graph pattern may include a systematic pattern composed of predefined legal domain entity types, inter-entity relationship types, and entity / relationship attribute rules.

[0055] The weakly supervised prompt template refers to a set of prompt words that guides a large language model to complete knowledge extraction tasks without relying on large-scale manually labeled samples, using only domain-specific type definitions and a small number of typical extracted examples. It includes a set of entity type definitions, a set of relation type definitions, and a set of attribute definitions. The entity type definition set defines the categories and characteristics of core entities within the domain, such as subdividing them into causes of action, parties, and legal clauses in a legal context. The relation type definition set clearly defines the core semantic relationships between entities within the domain, such as party-claim-cause of action, cause of action-applicable-legal clause, and party-evidence-evidence relationships in a legal context. The attribute definition set defines the core attributes and format requirements of entities or relationships, such as the format requirements for cause of action-court of trial, legal clause-clause number, and party-identity information in a legal context.

[0056] For example, taking a legal scenario as an example, the weak supervision prompt template constructed based on the preset domain graph pattern is as follows: the entity type definition includes cause of action, parties, and legal clauses; the relationship type definition includes the association between cause of action and legal clauses, such as cause of action - applicable - legal clauses, the association between parties and cause of action, such as parties - claims - cause of action, and the association between parties and evidence, such as parties - submission - evidence; the attribute definition includes cause of action, parties, and legal clauses, etc.

[0057] Secondly, the preprocessed text set and the weakly supervised prompt template are input into the large language model to obtain entity extraction result sets, relation extraction result sets, and attribute extraction result sets. The extraction confidence score for each extraction result in the entity extraction result set and relation extraction result set is recorded. Specifically, the preprocessed text set is first split into batches of 50 text blocks each to avoid overloading the large language model. Then, the weakly supervised prompt template is appended to each batch of text blocks and input into the large language model adapted to the professional scenario. The large language model uses the weakly supervised prompt template as a reference to perform semantic analysis on each text block, identify matching knowledge elements, and categorize them. Finally, it outputs the entity extraction result set, relation extraction result set, and attribute extraction result set. Simultaneously, the large language model generates extraction confidence scores for the entity and relation extraction results respectively.

[0058] Among them, the large language model can dynamically select a general large language model according to the needs of professional scenarios, such as GPT-4, Wenxin Yiyan, and Tongyi Qianwen. It can also be fine-tuned based on the domain data of the general large language model to obtain a more suitable domain-specific large language model for the semantic characteristics of specific domains. Taking the legal scenario as an example, in addition to using the general large language model after legal fine-tuning, the pre-trained model LawBERT in the legal domain can also be used directly. It has natural advantages in legal terminology recognition and semantic association analysis, which can improve the accuracy of knowledge extraction.

[0059] The extraction confidence score refers to the quantified probability value representing the reliability of the entity and relation extraction results output by the large language model. The score ranges from 0 to 1; the closer the extraction confidence score is to 1, the higher the large language model's confidence in the result. Its calculation logic involves the large language model calculating multiple dimensions of indicators, such as the matching degree between the extracted results and the weakly supervised prompt template, and the semantic support strength of the results in the text, and then fusing these results through an internal algorithm.

[0060] For example, taking a preprocessed text set of legal documents as an example, the content of a certain preprocessed text block is: Plaintiff Zhang claims that defendant a certain e-commerce company infringed his right to disseminate information online. This case should be governed by Article 1194 of the Civil Code. Through the weak supervision prompt template, it can be clearly stated that: the entity type includes the cause of action being a dispute over the right to disseminate information online, the parties being the plaintiff and defendant, and the legal clause being a clause of the Civil Code; the relationship type includes parties-claims-cause of action, cause of action-application-legal clause; after inputting both into the large language model, the output entity extraction result set includes: dissemination of information online. The extracted results set includes: Rights dispute, confidence score 0.98; Zhang, confidence score 0.99; an e-commerce company, confidence score 0.97; Article 1194 of the Civil Code, confidence score 0.99; the relationship extraction result set includes: Zhang - claim - dispute over the right to disseminate information online, confidence score 0.96; dispute over the right to disseminate information online - applicable - Article 1194 of the Civil Code, confidence score 0.95; the attribute extraction result set includes: Article 1194 of the Civil Code - clause number: 1194; Zhang - identity type: plaintiff.

[0061] Next, the attribute value format of the attribute extraction result set is validated to generate an attribute validity score. Specifically, based on a preset domain graph model and weakly supervised prompt template, an attribute value format validation rule base is constructed. The rule base needs to clearly define the format standards, value ranges, and compliance judgment conditions for different attribute types. Taking the legal scenario as an example, specific rules can be formulated for core attributes such as legal clause-clause number, cause of action-filing time, and party-ID number. Subsequently, each attribute entry in the attribute extraction result set is traversed, and the attribute value is compared with the corresponding validation rule to complete the validation from three dimensions: format compliance, value range rationality, and semantic adaptability. Finally, an attribute validity score is generated based on the validation results using a quantification algorithm, and the attribute validity score is bound to each attribute entry.

[0062] The attribute validity score is an evaluation value that quantifies the compliance of the format and the reasonableness of the value range of the attribute extraction results. The score ranges from 0 to 1, with a higher score indicating that the attribute value better meets the requirements of the professional scenario. For example, the calculation logic for the attribute validity score is as follows: attributes with fully compliant format and reasonable value range receive 1 point; attributes with minor format flaws but complete core information receive 0.6-0.9 points; attributes with serious format errors receive 0.1-0.5 points; and attributes whose value and type are completely mismatched receive 0 points.

[0063] For example, taking the attribute extraction result set of a legal scenario as an example, the weak supervision prompt template has clearly defined the attribute definition set as follows: the format of legal clause - clause number is number + clause, such as clause 1194; the format of cause of action - filing time is YYYY-MM-DD; the format of party - ID number is 18 digits. The verification process of the attribute extraction result set and the generation of attribute validity scores are as follows: Article 1194 of the Civil Code - Clause Number: 1194, the attribute value 1194 completely matches the format of number + clause, and is strongly related to the legal clause entity and has high semantic adaptability, with an attribute validity score of 0.99; Zhang - Identity Type: Plaintiff, the attribute value Plaintiff belongs to the legal subject type preset by the template, has no format deviation, and is consistent with the text semantics of Zhang being the claimant in the case, with an attribute validity score of 0.98.

[0064] Furthermore, the system filters out extraction results whose confidence scores are lower than a preset extraction confidence threshold and whose attribute validity scores are lower than a preset attribute validity threshold. Specifically, for entity extraction result sets and relation extraction result sets, the system compares whether the extraction confidence score is lower than the preset extraction confidence threshold; for attribute extraction result sets, the system compares whether the attribute validity score is lower than the preset attribute validity threshold. The preset extraction confidence threshold is a pre-set critical value used to determine the credibility of entity / relation extraction results, ranging from 0 to 1. It can be dynamically set based on professional domain characteristics, semantic complexity of domain text, and business needs. For example, in a legal scenario, the preset extraction confidence threshold can be set to 0.9. The preset attribute validity threshold is a pre-set critical value used to determine the compliance of attribute value format, ranging from 0 to 1. It can be dynamically set based on professional domain characteristics and business needs. For example, in a legal scenario, the preset attribute validity threshold can be set to 0.8.

[0065] For example, in a legal scenario, if the preset extraction confidence threshold is 0.9 and the preset attribute validity threshold is 0.8, the entity extraction confidence score in the entity extraction result set is compared to see if it is lower than the preset extraction confidence threshold. For example, the extraction confidence score for the information network dissemination right dispute is 0.98, which is not lower than the preset extraction confidence threshold, and is therefore retained. At the same time, the relationship extraction confidence score in the relationship extraction result set is compared to see if it is lower than the preset extraction confidence threshold. For example, the extraction confidence score for Zhang's claim on the information network dissemination right dispute is 0.96, which is not lower than the preset extraction confidence threshold, and is therefore retained. Then, the attribute validity score in the attribute extraction result set is compared to see if it is lower than the preset attribute validity threshold. For example, the attribute validity score for Article 1194 of the Civil Code is 0.99, which is not lower than the preset attribute validity threshold, and is therefore retained.

[0066] In this way, by filtering out extraction results with confidence scores below the preset extraction confidence threshold and attribute validity scores below the preset attribute validity threshold, low-quality knowledge elements generated during weakly supervised extraction can be eliminated. These include low-confidence entities / relationships with insufficient model trust, as well as invalid attributes with format violations and semantic mismatches, thereby improving the data quality for subsequent construction of structured knowledge units.

[0067] Furthermore, based on the filtered entity extraction, relation extraction, and attribute extraction results, structured knowledge units containing sets of qualified entities, qualified relations, and qualified attributes are constructed. In this way, the filtered, highly reliable, and compliant entity, relation, and attribute extraction results can be integrated into clearly defined and traceable structured knowledge units, providing standardized foundational data for the accurate construction of the subsequent knowledge graph.

[0068] Finally, pattern conflicts between structured knowledge units and domain graph patterns are identified, generating candidate new pattern description documents. This allows for the accurate identification of new entity / relationship types within structured knowledge units that exceed domain graph patterns, providing standardized candidate solutions for dynamic updates of graph patterns, effectively adapting to emerging professional text scenarios, and ensuring the scalability and adaptability of the knowledge graph.

[0069] Specifically, the phrase "constructing a structured knowledge unit containing a set of qualified entities, a set of qualified relationships, and a set of qualified attributes based on the filtered entity extraction results, relation extraction results, and attribute extraction results" includes:

[0070] Each entity extraction result after filtering is assigned a unique entity identifier, and the entity type and entity name are recorded to form a qualified entity set;

[0071] Each filtered relation extraction result is assigned a unique relation identifier, and the relation type and relation name are recorded to form a qualified relation set;

[0072] The filtered attribute extraction results are associated with the corresponding entity identifiers or relation identifiers to form a qualified attribute set.

[0073] Based on entity identifiers, relation identifiers, and corresponding attribute extraction results, construct an entity-relation topology;

[0074] Based on the set of qualified attributes, the text block identifier and extraction confidence score corresponding to each qualified entity and qualified relationship are recorded and integrated to form a structured knowledge unit.

[0075] In this embodiment, a unique entity identifier is first assigned to each filtered entity extraction result, and the entity type and entity name are recorded to form a qualified entity set. Specifically, entity identifier encoding rules can be formulated based on the characteristics of a professional domain. Then, the filtered entity extraction result set is traversed, and an identifier is assigned to each entity one by one, while the entity type and entity name are recorded simultaneously. Finally, the triples of entity identifier, entity type, and entity name are sorted and organized according to entity type to form a non-repeating and searchable qualified entity set.

[0076] Entity identifiers are globally unique string identifiers assigned to entities. An example of the format is: XX-Entity type abbreviation-Serial number, where XX is an abbreviation of the professional field name, such as FL representing law. Entity identifiers can directly associate all attributes and relationships of an entity, enabling full-chain traceability of the entity.

[0077] For example, for disputes concerning the right to disseminate information online in the filtered entity extraction results, a unique entity identifier can be assigned: FL-AY-0001, where FL represents law, AY represents cause of action, and 0001 represents serial number. This records the entity type as cause of action and the entity name as dispute concerning the right to disseminate information online. Following the same logic and method, the filtered entity extraction result set is traversed to form a qualified entity set.

[0078] Secondly, each extracted relation is assigned a unique relation identifier, and the relation type and name are recorded to form a qualified relation set. For example, the process of constructing a qualified entity set can be referenced to assign a unique relation identifier to each relation. An example format is: XX-Relationship Type Abbreviation-Sequence Number, where XX is an abbreviation of the professional field name, such as FL representing law. The relation identifier can directly associate all attributes and relationships of the relation, enabling full-link traceability of the relation.

[0079] For example, for the filtered relation extraction results of Zhang Mou-claim-information network dissemination right dispute, a unique relation identifier can be assigned: FL-ZZ-0001, where FL represents law, ZZ represents claim, 0001 represents sequence number, the record relation type is claim, the relation name is claim, the associated head entity is Zhang Mou, and the associated tail entity is information network dissemination right dispute. Following the same logic and method, the filtered relation extraction result set is traversed to form a qualified relation set.

[0080] Next, the filtered attribute extraction results are associated with the corresponding entity identifiers or relation identifiers to form a qualified attribute set.

[0081] Specifically, the attribution rules for attributes are first determined based on the characteristics of the specific professional field. Taking the legal scenario as an example, attributes such as ID number and identity type in the attribute extraction results belong to the party entity, clause number and validity status belong to the legal clause entity, and the court of trial and filing time belong to the cause of action entity. Related evidence numbers can be attributed to the party-claim-cause of action relationship. Then, the filtered attribute extraction results are traversed, and the corresponding entity or relationship is matched based on the attribute name and attribution rules. Finally, the entity / relationship name is associated with the corresponding entity identifier or relationship identifier to form a set of qualified attributes with clear attribution.

[0082] For example, in the filtered attribute extraction results, Article 1194 of the Civil Code - Article No.: 1194, Zhang Mou - Identity Type: Plaintiff, according to the attribute attribution rules, Article No.: 1194 matches the legal article entity and is associated with the entity identifier FL-TK-0001 representing Article 1194 of the Civil Code.

[0083] Furthermore, based on entity identifiers, relation identifiers, and the corresponding attribute extraction results, an entity-relationship topology is constructed. Specifically, a directed graph structure of nodes and edges can be used to construct the entity-relationship topology: the entities corresponding to the entity identifiers are used as nodes, each node is labeled with the entity name, entity type, and attributes, and the relations corresponding to the relation identifiers are used as directed edges, with the direction of the edges pointing from the head entity to the tail entity. The edge labels include the relation name, relation type, and associated attributes. Here, the head entity refers to the entity that is the starting point of the relation in the semantic relationship between entities; it is the starting point of the head entity-relationship-tail entity association link and is the initiator or subject of the relation. The tail entity refers to the entity that is the target of the relation and is the receiver or object of the relation. For example, in the relation: Zhang - Claim - Information Network Dispute over the Right to Disseminate Information, Zhang is the head entity, representing the initiator of the claim relation, and the Information Network Dispute over the Right to Disseminate Information is the tail entity, representing the object of the claim relation.

[0084] For example, when constructing an entity-relationship topology, the basic links of head entity nodes, relation edges, and tail entity nodes are first established, and then the corresponding attributes are attached to the nodes or edges, ultimately forming an entity-relationship topology that can intuitively reflect the relationships between knowledge elements.

[0085] Finally, based on the set of qualified attributes, the text block identifier and extraction confidence score corresponding to each qualified entity and qualified relation are recorded and integrated to form a structured knowledge unit. Specifically, according to the association effect of the set of qualified attributes, a corresponding text block identifier and extraction confidence score are added to each qualified entity and qualified relation. Then, the qualified entities and qualified relations with the supplemented information are integrated to form a structured knowledge unit.

[0086] For example, if the set of qualified attributes includes identity type - plaintiff - 0.98 - FL - DSR - 001, the qualified entity is FL - DSR - 001 (representative: Zhang, party), the qualified relationship is FL - ZZ - 001 (representative: claim, Zhang - dispute over right of dissemination of information on the Internet), and it is known that Zhang and the claim relationship both come from the text block identifier FL - 20250601-001, the extraction confidence score of Zhang is 0.99, and the extraction confidence score of the claim relationship is 0.96, then the integrated structured knowledge unit includes: FL - DSR - 001 [Zhang, party, FL - 20250601-001, 0.99], FL - ZZ - 001 [claim, FL - 20250601-001, 0.96], and the attribute: identity type - plaintiff - 0.98 - FL - DSR - 001.

[0087] Specifically, the step of "identifying pattern conflicts between the structured knowledge units and the domain graph patterns, and generating candidate new pattern description documents" includes:

[0088] Based on the qualified entity set, calculate the semantic similarity between the entity types in the qualified entity set and the predefined entity types in the domain graph pattern, and take the maximum semantic similarity as the entity type semantic similarity.

[0089] When the semantic similarity of the entity type is lower than the preset entity type conflict threshold, it is identified as an entity-level pattern conflict and an entity-level pattern conflict instance is generated.

[0090] Based on the qualified relation set, calculate the semantic similarity between the relation types in the qualified relation set and the predefined relation types in the domain graph pattern, and take the maximum semantic similarity as the relation type semantic similarity;

[0091] When the semantic similarity of the relation type is lower than the preset relation type conflict threshold, it is identified as a relation-level pattern conflict and a relation-level pattern conflict instance is generated.

[0092] Based on the entity-level and relation-level pattern conflict instances, the semantic context of pattern conflict is analyzed using a large language model.

[0093] Based on the semantic analysis results of the pattern conflict context, a candidate new pattern description document containing candidate new entity type definitions and candidate new relation type definitions is generated by using a large language model adapted to professional scenarios.

[0094] In this embodiment, firstly, based on the qualified entity set, the semantic similarity between the entity type in the qualified entity set and the predefined entity type in the domain graph pattern is calculated, and the maximum semantic similarity is taken as the entity type semantic similarity. When the entity type semantic similarity is lower than the preset entity type conflict threshold, it is identified as an entity-level pattern conflict and an entity-level pattern conflict instance is generated. Specifically, firstly, all entity types are extracted from the qualified entity set, and predefined entity types are extracted from the domain graph pattern. Then, a semantic model adapted to the professional domain is used to convert the extracted entity types and each predefined entity type into computer-recognizable semantic vectors. The semantic similarity between the vectors is calculated by calculating the cosine similarity between them, and the maximum semantic similarity is taken as the entity type semantic similarity. Then, the entity type semantic similarity is compared with a preset entity type conflict threshold. If it is lower than the preset entity type conflict threshold, it means that the core semantic association between the extracted entity type and the predefined entity type in the domain graph pattern is weak, and its scene features and semantic connotations have exceeded the coverage of the preset pattern and cannot be effectively adapted to the existing pattern. In this case, it is determined that the entity type conflicts with the domain graph pattern, and a corresponding entity-level pattern conflict instance is generated.

[0095] Entity type semantic similarity is an indicator that quantifies the degree of semantic association between an entity type and a predefined entity type. Its value ranges from 0 to 1. The closer the semantic similarity is to 1, the higher the semantic fit between the two entities in terms of core connotation and scenario attributes. To ensure computational accuracy, the semantic model needs to be adapted to the characteristics of the specific domain. Different general semantic models can be selected based on the domain characteristics. For example, in legal scenarios, the LawBERT semantic model, pre-trained for legal texts, can be chosen.

[0096] The preset entity type conflict threshold can be dynamically set based on the semantic characteristics of the professional field. The setting process needs to balance the standardization of the pattern with the inclusiveness of new types. It should not be set too low, which would lead to the generalization of the pattern and loss of its standardization value, nor should it be set too high, which would miss valuable new types. Taking the legal scenario as an example, because the rigor of legal terminology requires precise definition of entity types, semantic deviations may lead to legal logic errors. Therefore, the preset entity type conflict threshold can be set to 0.7.

[0097] For example, taking a legal scenario, if the preset entity type conflict threshold is 0.7, the entity type extracted from the qualified entity set, "Live-streaming e-commerce false advertising dispute," has a semantic similarity of 0.65 with the predefined entity type "false advertising dispute" in the domain graph pattern calculated by the LawBERT model. This is less than the preset entity type conflict threshold of 0.7, indicating that the semantic association may not meet the preset entity type conflict threshold because the live-streaming e-commerce scenario attribute of the extracted entity type is not covered by the predefined type. Therefore, it is determined to be an entity-level pattern conflict, generating an entity-level pattern conflict instance: Entity-level pattern conflict instance-002: Extracted entity type: Live-streaming e-commerce false advertising dispute, predefined comparison type: false advertising dispute, entity type semantic similarity 0.65, conflict reason: the extracted type contains the live-streaming e-commerce scenario limitation, which exceeds the semantic range of the predefined type.

[0098] Secondly, based on the qualified relation set, the semantic similarity between the relation types in the qualified relation set and the predefined relation types in the domain graph pattern is calculated, and the maximum semantic similarity is taken as the relation type semantic similarity. When the semantic similarity of a relation type is lower than a preset relation type conflict threshold, it is identified as a relation-level pattern conflict, and a relation-level pattern conflict instance is generated. Specifically, all relation types are extracted from the qualified relation set, and predefined relation types are extracted from the domain graph pattern. Using the same semantic model of the professional domain, the extracted relation types and each predefined relation type are transformed into semantic vectors. The semantic similarity is calculated using cosine similarity, and the maximum value is selected as the relation type semantic similarity. Then, it is compared with the preset relation type conflict threshold. If the semantic similarity of a relation type is lower than the preset relation type conflict threshold, it means that the core semantic association between the extracted relation type and the predefined relation type in the domain graph pattern is loose, and the scenario orientation and association logic it carries have exceeded the coverage of the preset pattern and cannot be effectively adapted to the existing relation type specifications. In this case, it is identified as a relation-level pattern conflict, and a relation-level pattern conflict instance is generated.

[0099] Among them, relation type semantic similarity is an indicator that quantifies the degree of semantic association between a relation type and a predefined relation type. Its value ranges from 0 to 1; the closer the relation type semantic similarity is to 1, the higher the semantic fit between the two types of relations. To ensure computational accuracy, the semantic model needs to be adapted to the characteristics of the specific domain. Different general semantic models can be selected based on the domain characteristics. For example, in legal scenarios, the LawBERT semantic model, pre-trained for legal texts, can be chosen.

[0100] The preset relationship type conflict threshold can be dynamically set based on the semantic characteristics of a professional field. The setting process needs to balance the normativity of the pattern with the inclusiveness of new types. It should not be set too low, which would lead to the generalization of the pattern and loss of normative value, nor should it be set too high, which would miss valuable new types. Taking the legal scenario as an example, the preset relationship type conflict threshold can be set to 0.7.

[0101] For example, if the preset relationship type conflict threshold is 0.7, the relationship type extracted from the qualified relationship set is: live-streaming e-commerce association. The predefined relationship types in the domain graph pattern are association, subject-behavior, and product-sales. The semantic similarity of the relationship type between live-streaming e-commerce association and association is calculated to be 0.62, the semantic similarity of the relationship type with subject-behavior is 0.58, and the semantic similarity of the relationship type with product-sales is 0.61. Finally, the semantic similarity of the relationship type is determined to be 0.62, which is less than the preset relationship type conflict threshold of 0.7. Therefore, it is judged as a relationship-level pattern conflict, and a relationship-level pattern conflict instance is generated as follows: Relationship-level pattern conflict instance-001: Extracted relationship type: live-streaming e-commerce association, most relevant predefined relationship type: association, relationship type semantic similarity 0.62, conflict reason: the extracted relationship type contains live-streaming e-commerce scenario attributes.

[0102] Secondly, based on entity-level and relation-level pattern conflict instances, the semantic context of pattern conflicts is analyzed using a large language model. Specifically, all entity-level and relation-level pattern conflict instances can be organized into a conflict set, and corresponding contextual information can be added to each conflict instance, such as the original text block content where the conflicting entity / relationship is located, and information about other related entities / attributes. Then, the conflict set and contextual information are input into a large language model adapted to the professional scenario. The large language model outputs the semantic context of the conflict analysis pattern conflict through deep semantic analysis, clarifying the cause of the conflict, the core features of the conflicting entities / relationships, and the scope of applicable professional scenarios. Among them, the large language model can be dynamically selected according to the characteristics of the domain. For general scenarios, general large models such as GPT-4, Wenxin Yiyan, Tongyi Qianwen, and Xunfei Xinghuo can be used, while for legal scenarios, pre-trained models in the legal domain, such as LawGPT or Wenxin Yiyan, can be selected first.

[0103] For example, after inputting the aforementioned entity-level and relation-level pattern conflict instances and their corresponding contextual information into the large language model, the semantic context of the pattern conflict is analyzed and output: The essence of the conflict is that live-streaming e-commerce, as an emerging business model, has related dispute types and relationships that exceed the coverage of traditional legal graph models; the conflict entity characteristics are that it uses "live-streaming sales" as a scenario prefix, which is a sub-type of false advertising disputes, requiring association with new entities such as live-streaming platforms and anchors; the conflict relationship characteristics are that live-streaming sales are associated with a clear binding relationship between the entity and the live-streaming scenario, which is different from general association relationships; the applicable scenario is the judgment documents of false advertising and product quality disputes in the field of live-streaming e-commerce.

[0104] Finally, based on the semantic analysis results of the pattern conflict context, a candidate new pattern description document is generated by using a large language model adapted to the professional scenario, which includes definitions of candidate new entity types and candidate new relation types. Specifically, the semantic analysis results of the pattern conflict context have clarified the nature of the conflict and the core features of the conflicting entities / relationships. A large language model adapted to the professional scenario, such as LawGPT and GPT-4, is invoked. Based on the semantic logic of the professional domain, entity types and relation types that exceed the preset scope of the domain graph pattern are standardized and defined. Candidate new entity type definitions containing elements such as type name, core features, and parent type are generated, as well as candidate new relation type definitions containing elements such as type name, head / tail entity type restrictions, and core semantics. The two types of definitions are then integrated into a structured candidate new pattern description document.

[0105] Among them, the candidate new entity type definition refers to the standardized definition formulated for entity-level pattern conflict instances, which is the same as the predefined relation type format in the domain graph pattern, ensuring that the new entity type is both adaptable to emerging scenarios and logically compatible with the original pattern; the candidate new relation type definition is the standardized definition formulated for relation-level pattern conflict instances, which is the same as the predefined relation type format in the domain graph pattern, ensuring that the new relation can accurately carry the scenario-based association logic between entities; the candidate new pattern description document is a carrier that records the candidate new entity type definition and the candidate new relation type definition in a structured format, such as XML / JSON.

[0106] In summary, compared to existing technologies, this application, based on a preset domain graph pattern, utilizes a large language model to extract entities, relationships, and attributes from the preprocessed text in a weakly supervised manner, constructing structured knowledge units. It then identifies pattern conflicts between these structured knowledge units and the domain graph pattern, generating candidate new pattern description documents. This allows for efficient extraction of high-value knowledge from professional texts and construction of structured knowledge units without large-scale manual annotation. Simultaneously, it automatically identifies pattern conflicts beyond the preset domain graph pattern and generates candidate new pattern description documents. This effectively solves the problem of existing technologies struggling to automatically identify, evaluate, and incorporate new knowledge beyond the domain graph pattern in professional texts, ensuring the dynamic updating capability and adaptability of the knowledge graph to professional scenarios.

[0107] S30: Evaluate the overall confidence level of the candidate new pattern description documents through a multi-dimensional confidence assessment mechanism, and dynamically update the domain graph pattern.

[0108] In the construction of knowledge graphs in professional text fields such as law and medicine, the value of candidate new pattern description documents varies significantly. Some are high-value types adapted to emerging scenarios, while others are isolated cases or semantically ambiguous invalid types. Existing technologies lack standardized multi-dimensional quantitative evaluation mechanisms and can only judge the value of candidate new patterns through a single dimension or human subjective judgment. This can easily lead to the inclusion of low-value types in the graph pattern, causing redundancy, and the omission of high-value emerging types. Furthermore, there is a lack of hierarchical dynamic update rules, which cannot balance the rigor of the domain graph pattern with its adaptability to emerging professional scenarios.

[0109] To address the aforementioned issues, this application employs a multi-dimensional confidence assessment mechanism to evaluate the overall confidence level of the candidate new pattern description documents and dynamically updates the domain graph patterns.

[0110] Specifically, step S30 in the method includes:

[0111] Based on the candidate new pattern description document, the candidate new entity type definition and the candidate new relation type definition are parsed out;

[0112] The log-normalized value of the frequency of occurrence of the candidate new entity type definition in all text blocks of the preprocessed text set is calculated as the frequency confidence of the candidate new entity type definition;

[0113] Calculate the maximum cosine similarity between the semantic vectors of the candidate new entity type definition and all predefined entity type definitions in the domain graph pattern, and use it as the semantic consistency confidence of the candidate new entity type definition;

[0114] Calculate the confidence score of the relation structure rationality in the definition of candidate new relation types;

[0115] The frequency confidence, semantic consistency confidence, and relational structure rationality confidence are weighted and fused according to preset weights to obtain the comprehensive confidence of the candidate new pattern description document.

[0116] In this embodiment, the candidate new entity type definition and candidate new relation type definition are first parsed from the candidate new pattern description document. Specifically, based on the candidate new pattern description document, information such as type name and core features are extracted from the candidate new entity type definition, and information such as type name and head / tail entity type restrictions are extracted from the candidate new relation type definition.

[0117] For example, the following can be parsed from the candidate new pattern description document: Candidate new entity type definition: Type name: Live Streaming E-commerce False Advertising Dispute, Parent type: False Advertising Dispute; and candidate new relationship type definition: Type name: Live Streaming Scenario Association, Head entity type restriction: Live Streaming Subject, Tail entity type restriction: Live Streaming Platform.

[0118] Secondly, the log-normalized value of the frequency of candidate new entity type definitions appearing in all text blocks of the preprocessed text set is calculated as the frequency confidence score of the candidate new entity type definitions. Specifically, all text blocks in the preprocessed text set are traversed to obtain the total number of text blocks, and the frequency of the type name or core feature description of the candidate new entity type definition appearing in the text is counted. For example, the dispute over false advertising in live-streaming e-commerce appears 45 times in 1000 legal text blocks. Then, the frequency is log-normalized using the formula: Frequency Confidence Score = ln(1 + Frequency of Occurrence) / ln(1 + Total Number of Text Blocks), converting the frequency of occurrence into a standardized value of 0-1 to eliminate the interference of differences in the total number of text blocks. Among them, the frequency confidence score is an indicator that quantifies the prevalence of new entity types in professional texts. The closer the value is to 1, the more common it is in the scenario.

[0119] For example, if a dispute over false advertising in live-streaming e-commerce occurs 45 times in 1000 legal text blocks, the frequency confidence level is calculated as ln(1+45) / ln(1+1000)≈0.55.

[0120] Next, the maximum cosine similarity between the semantic vectors of the candidate new entity type definition and all predefined entity type definitions in the domain graph pattern is calculated as the semantic consistency confidence score of the candidate new entity type definition. Specifically, a domain-pretrained semantic model is used to convert the core feature text of the candidate new entity type definition and the text of all predefined entity type definitions in the domain graph pattern into semantic vectors respectively. Then, the cosine similarity of each pair of vectors is calculated, and the maximum cosine similarity is taken as the semantic consistency confidence score. The higher the semantic consistency confidence score, the closer the semantic association between the new entity type and the original pattern, and the better the compatibility. The semantic model can be dynamically selected according to the domain characteristics. For general scenarios, general large models such as GPT-4, Wenxin Yiyan, Tongyi Qianwen, and Xunfei Xinghuo can be used.

[0121] For example, the core features of the live-streaming e-commerce false advertising dispute in the candidate new entity type definition are compared with all predefined entity type definitions in the domain graph pattern: false advertising dispute, online shopping dispute, and tort liability dispute. The cosine similarity calculated by LawBERT is 0.72, 0.58, and 0.45, respectively. The maximum value of 0.72 is taken as the semantic consistency confidence score.

[0122] Furthermore, the confidence level of the relation structure rationality defined by the candidate new relation type is calculated. Combining the extraction credibility of the entities connected by the relation with the type normativity, the rationality of the relation structure is quantified. The higher the entity credibility and the more the type conforms to the preset pattern, the greater the confidence level of the relation structure rationality.

[0123] Finally, the frequency confidence score, semantic consistency confidence score, and relational structure rationality confidence score are weighted and fused according to preset weights to obtain the comprehensive confidence score of the candidate new pattern description document. The preset weights need to be dynamically determined based on domain requirements, and the sum of the preset weights for frequency confidence score, semantic consistency confidence score, and relational structure rationality confidence score is 1. For example, for rigorous domains such as law and medicine, the weight of semantic consistency confidence score can be set to the highest, such as 0.4, to ensure that the new type does not violate domain logic; the weight of relational structure rationality confidence score is next, such as 0.35, to ensure the rigor of relational logic; and the weight of frequency confidence score is the lowest, such as 0.25, to avoid overlooking niche but important new types.

[0124] For example, if the frequency confidence score is 0.55, the semantic consistency confidence score is 0.72, and the relational structure rationality confidence score is 0.7, and the preset weights of the frequency confidence score, semantic consistency confidence score, and relational structure rationality confidence score are 0.25, 0.4, and 0.35 respectively, then the overall confidence score of the candidate new pattern description document is calculated as 0.55 × 0.25 + 0.72 × 0.4 + 0.7 × 0.35 = 0.6705. The comprehensive confidence score reflects the overall value and compliance of the candidate new entity type definitions and candidate new relation type definitions in the candidate new pattern description document. It covers three dimensions: first, frequency confidence score reflects the universality of candidate new entity types in professional text scenarios; second, semantic consistency confidence score reflects the semantic compatibility of new entity types with predefined entity types in the domain graph pattern; and third, relation structure rationality confidence score reflects the reliability of new relation types connecting entities and the standardization of type combinations. Essentially, it is a quantitative indicator for measuring whether candidate new patterns are suitable for inclusion in the domain graph pattern and for adapting to the dynamic update needs of professional text knowledge graphs.

[0125] Specifically, the "calculation of the relation structure rationality confidence of the candidate new relation type definition" includes:

[0126] Extract the head entity identifier and tail entity identifier connected to the candidate new relation type definition from the relation-level schema conflict instance on which the candidate new schema description document is based;

[0127] In the qualified entity set of the structured knowledge unit, query the head entity type and head entity extraction confidence score corresponding to the head entity identifier, and the tail entity type and tail entity extraction confidence score corresponding to the tail entity identifier.

[0128] Calculate the geometric mean of the confidence scores of the head entity extraction and the tail entity extraction, and use it as the confidence score of the basic relation.

[0129] Calculate the semantic similarity between the head entity type and all entity types in the predefined entity type set of the domain graph pattern, and take the maximum value as the head entity type support.

[0130] Calculate the semantic similarity between the tail entity type and all entity types in the predefined entity type set of the domain graph pattern, and take the maximum value as the tail entity type support;

[0131] Calculate the geometric mean of the head entity type support and the tail entity type support, and use it as the type support coefficient;

[0132] Multiplying the basic relation confidence score by the type support coefficient yields the relation structure rationality confidence score.

[0133] In this embodiment, the head entity identifier and tail entity identifier connected to the candidate new relation type definition are first extracted from the relation-level schema conflict instances upon which the candidate new schema description document is based. Specifically, the relation-level schema conflict instances upon which the candidate new schema description document is based have recorded the association information of the candidate new relation type, head entity, and tail entity. For example, in the case of Zhang Mou - Claim - Information Network Dissemination Right Dispute, Zhang Mou is the head entity, and the Information Network Dissemination Right Dispute is the tail entity. By matching the head / tail entity names with the qualified entity set, the corresponding head entity identifier and tail entity identifier are extracted, thus achieving precise binding between the relation and the entity.

[0134] Secondly, within the qualified entity set of the structured knowledge unit, the system queries the header entity type and header entity extraction confidence score corresponding to the header entity identifier, and the tail entity type and tail entity extraction confidence score corresponding to the tail entity identifier. Specifically, based on the header and tail entity identifiers, the system locates the corresponding entity entries in the qualified entity set and extracts the entity type and extraction confidence score.

[0135] Next, calculate the geometric mean of the confidence scores for the head entity extraction and the tail entity extraction, as the foundational relation confidence score. Specifically, the geometric mean, which is good at handling product-type indicators, is used to calculate the average of the confidence scores for the head and tail entities. The formula is: Foundational Relation Confidence Score = This ensures that the results accurately reflect the reliability of the entity combination.

[0136] For example, if the confidence score for the head entity extraction is 0.98 and the confidence score for the tail entity extraction is 0.95, then the confidence score of the basic relation = The confidence level of the basic relationship is approximately 0.965. It reflects the overall reliability of the extraction results of the head and tail entities connected by the candidate new relationship type definition. The higher the confidence level of the basic relationship, the more reliable the entity base on which the candidate new relationship is attached. Conversely, it indicates that the entity extraction results of the relationship connection itself have credibility flaws, which will directly weaken the rationality of the relationship structure.

[0137] Furthermore, the semantic similarity between the head entity type and all entity types in the predefined entity type set of the domain graph pattern is calculated, and the maximum value is taken as the head entity type support. For example, the semantic similarity between the head entity type and all predefined entity types in the domain graph pattern can be calculated using a semantic model, and the maximum value is taken as the head entity type support. A higher head entity type support indicates better compatibility between the head entity type and the existing pattern. The semantic model can be dynamically selected based on domain characteristics; for general scenarios, general large models such as GPT-4, Wenxin Yiyan, Tongyi Qianwen, and Xunfei Xinghuo can be used.

[0138] For example, if the semantic similarity between the head entity type and all entity types in the predefined set of entity types in the domain graph pattern is calculated to be 0.78, 0.65, ..., the maximum value of 0.78 is selected as the head entity type support.

[0139] Furthermore, the semantic similarity between the tail entity type and all entity types in the predefined set of entity types in the domain graph pattern is calculated, and the maximum value is taken as the tail entity type support. Specifically, the calculation logic and method of the tail entity type support are the same as those of the head entity type support. For example, the tail entity type support is calculated and selected to be 0.82. The higher the tail entity type support, the better the compatibility between the tail entity type and the existing pattern.

[0140] Furthermore, the geometric mean of the head entity type support and the tail entity type support is calculated as the type support coefficient. Specifically, following the same calculation logic for the confidence of the basic relations, the geometric mean is used to calculate the average of the head entity type support and the tail entity type support. The calculation formula is: Type Support Coefficient = This ensures that the results reflect the overall standardization of the entity type combination.

[0141] For example, if the type support of the head entity is 0.78 and the type support of the tail entity is 0.82, the calculated type support coefficient is = The type support coefficient is approximately 0.8. It can reflect the overall semantic adaptability and standardization of the head entity type, tail entity type and predefined entity type set in the domain graph pattern connected by the candidate new relation type definition. It is a key dimension for evaluating the rationality of the new relation type structure.

[0142] Finally, the confidence score of the basic relationship is multiplied by the type support coefficient to obtain the confidence score of the relationship structure rationality. Specifically, the confidence score of the relationship structure rationality = confidence score of the basic relationship × type support coefficient. This confidence score reflects both the reliability of the entities themselves and the standardization of the type combination, thus comprehensively reflecting the rationality of the relationship structure.

[0143] For example, if the confidence level of the basic relation is 0.965 and the type support coefficient is 0.8, then the confidence level of the relation structure rationality is 0.965 × 0.8 = 0.772.

[0144] Furthermore, the "dynamically updated domain graph mode" includes:

[0145] Set a pattern adoption threshold and a pattern storage threshold, wherein the pattern adoption threshold is greater than the pattern storage threshold;

[0146] The overall confidence level of the candidate new pattern description documents is compared with the pattern adoption threshold and the pattern storage threshold;

[0147] When the overall confidence level is greater than the pattern adoption threshold, the candidate new entity type definition in the candidate new pattern description document is added to the entity type definition set of the domain graph pattern, and the candidate new relation type definition is added to the relation type definition set of the domain graph pattern, thus forming an updated domain graph pattern.

[0148] When the overall confidence level is less than or equal to the pattern adoption threshold and greater than the pattern temporary storage threshold, the candidate new entity type definition and candidate new relation type definition of the candidate new pattern description document are added as supplementary information to the weak supervision prompt template.

[0149] When the overall confidence level is less than or equal to the pattern temporary storage threshold, the candidate new pattern description document is discarded.

[0150] In this embodiment, a pattern adoption threshold and a pattern storage threshold are first set, wherein the pattern adoption threshold is greater than the pattern storage threshold. Specifically, the pattern adoption threshold is the minimum confidence requirement for directly incorporating a pattern into the domain graph, and should be set relatively high, such as 0.8, to ensure that the incorporated new types fully conform to the domain logic. The pattern storage threshold is the minimum requirement for retaining observations, and can be set relatively low, such as 0.5, to leave sufficient space for new types that have potential but are not yet widespread, and must satisfy the condition that the pattern adoption threshold is greater than the pattern storage threshold. The pattern adoption threshold and the pattern storage threshold can be dynamically adjusted according to the characteristics of the professional domain. For example, in a legal scenario, the pattern adoption threshold can be set to 0.8 and the pattern storage threshold to 0.5.

[0151] Secondly, the overall confidence level of the candidate new schema description document is compared with the schema adoption threshold and the schema temporary storage threshold. When the overall confidence level is greater than the schema adoption threshold, the candidate new entity type definitions in the candidate new schema description document are added to the entity type definition set of the domain graph schema, and the candidate new relation type definitions are added to the relation type definition set of the domain graph schema, forming an updated domain graph schema. Specifically, candidate new entity type definitions are extracted from the candidate new schema description document, added to the entity type definition set of the domain graph schema, and their parent type associations are updated, such as treating live-streaming e-commerce false advertising disputes as a subtype of false advertising disputes; and candidate new relation type definitions are extracted and added to the relation type definition set, clarifying their head / tail entity type restrictions.

[0152] For example, if the overall confidence level of the candidate new pattern description document is 0.85, which is greater than the pattern adoption threshold of 0.8, then the candidate new entity type definition in the candidate new pattern description document is added to the entity type definition set of the domain graph pattern, and the candidate new relation type definition is added to the relation type definition set of the domain graph pattern, thus forming the updated domain graph pattern.

[0153] Conversely, when the overall confidence level is less than or equal to the pattern adoption threshold but greater than the pattern temporary storage threshold, the candidate new entity type definitions and candidate new relation type definitions from the candidate new pattern description document are added as supplementary information to the weak supervision prompt template. Specifically, the candidate new entity type definitions and candidate new relation type definitions from the candidate new pattern description document are converted into natural language descriptions and added to the weak supervision prompt template. The large language model will then refer to this template when extracting knowledge. If the frequency of new types continues to increase, the confidence level will improve in the next evaluation, potentially reaching the pattern adoption threshold.

[0154] For example, if the overall confidence level of the candidate new schema description document is 0.772, which is between the schema adoption threshold and the schema temporary storage threshold, then the candidate new entity type definition and candidate new relation type definition of the candidate new schema description document will be added as supplementary information to the weak supervision prompt template.

[0155] Conversely, when the overall confidence level is less than or equal to the pattern temporary storage threshold, the candidate new pattern description document is discarded. Specifically, if the overall confidence level is less than or equal to the pattern temporary storage threshold, it indicates that the new type may be an isolated case or semantically ambiguous. The candidate new pattern description document is directly discarded, and the reason for discarding is recorded, such as semantic ambiguity and low frequency of occurrence, to ensure that the processing is traceable.

[0156] In summary, compared to existing technologies, this application assesses the overall confidence level of the candidate new pattern description documents through a multi-dimensional confidence evaluation mechanism and dynamically updates the domain graph patterns. Thus, by using multi-dimensional quantitative evaluation to accurately determine the overall value and compliance of candidate new pattern description documents, and combining this with stratification thresholds to achieve dynamic stratified updates of the domain graph patterns, this approach filters out low-value and invalid new patterns to avoid graph redundancy, while also incorporating high-value emerging types and temporarily stored potential types. This effectively balances the rigor of the domain graph patterns with their adaptability to emerging professional text scenarios, solving the problem of inaccurate pattern updates caused by single-dimensional or manual evaluation in existing technologies.

[0157] S40: Optimize and reconstruct the structured knowledge units based on the updated domain graph pattern to construct a weakly supervised knowledge graph, and calculate the quality index of the reconstructed weakly supervised knowledge graph.

[0158] In the aforementioned steps, entities, relationships, and attributes were extracted from professional texts using a weakly supervised approach to construct structured knowledge units. Conflicts between these units and domain graph patterns were identified, and candidate new patterns were generated. Dynamic updates of the domain graph patterns were achieved through multi-dimensional confidence evaluation.

[0159] Therefore, based on the updated domain graph pattern, structured knowledge units can be optimized and reconstructed through type identification updates, entity fusion, and other means to build a weakly supervised knowledge graph. Quality indicators can then be calculated from the dimensions of accuracy, conflict rate, and completeness to ensure that the weakly supervised knowledge graph of professional texts is suitable for practical applications.

[0160] To address the aforementioned issues, this application optimizes and reconstructs the structured knowledge units based on the updated domain graph model, constructs a weakly supervised knowledge graph, and calculates the quality index of the reconstructed weakly supervised knowledge graph.

[0161] Specifically, step S40 in the method includes:

[0162] Based on the updated domain graph pattern, the type identifiers of the qualified entity set and qualified relation set in the structured knowledge unit are updated, and the entity type of qualified entity and the relation type of qualified relation are mapped to the standard type definition in the updated domain graph pattern.

[0163] The qualified entity set after the type identifier is updated is processed by entity fusion, merging multiple qualified entities that point to the same real-world object, and synchronously updating the associated qualified relationship set and qualified attribute set;

[0164] Based on the qualified entity set, qualified relation set, and qualified attribute set after entity fusion processing, a weakly supervised knowledge graph is constructed.

[0165] Calculate the quality index of the reconstructed weakly supervised knowledge graph.

[0166] In this embodiment, firstly, based on the updated domain graph pattern, the type identifiers of the qualified entity set and qualified relation set in the structured knowledge unit are updated, mapping the entity type of qualified entities and the relation type of qualified relations to the standard type definitions in the updated domain graph pattern. Specifically, since the updated domain graph pattern has included high-confidence candidate new types, the qualified entity set of the structured knowledge unit must be traversed first, and the original type identifier of each entity is replaced with the standard type name in the updated pattern: if the entity type is a newly added candidate type, the new type definition is directly adopted; if it is still an existing predefined type, the identifier is confirmed to be consistent with the pattern. For the qualified relation set, the relation type is also mapped to the standard relation definition in the updated pattern to ensure that the type identifiers of entities and relations conform to the latest specifications.

[0167] For example, in the updated domain graph pattern in the legal scenario, a new entity type, "Live Streaming E-commerce False Advertising Dispute," is added. In the structured knowledge unit, an entity whose original type was identified as "Live Streaming E-commerce False Advertising Dispute (Temporary)," is updated to the standard type "Live Streaming E-commerce False Advertising Dispute" through type mapping, and mapped to the standard type: "Live Streaming Scenario Association."

[0168] Secondly, entity fusion processing is performed on the qualified entity set after the type identifier is updated, merging multiple qualified entities pointing to the same real-world object, and simultaneously updating the associated qualified relationship set and qualified attribute set. Specifically, entity fusion needs to be based on the dual logic of attribute matching and semantic association: first, the core attributes of the entities after the type identifier is updated are extracted, such as the subject name of legal entities, etc., and then the attribute similarity and semantic similarity between entities are calculated through a semantic model. If both are higher than the preset fusion threshold (e.g., 0.9), they are determined to be the same entity and merged. During merging, the attribute information with the highest confidence level extracted from each entity must be retained, and all associated qualified relationship sets and qualified attribute sets are uniformly associated with the unique entity after fusion to avoid the loss of relationships and attributes. Among them, the semantic model can be dynamically selected according to the domain characteristics. For general scenarios, general large models such as GPT-4, Wenxin Yiyan, Tongyi Qianwen, and Xunfei Xinghuo can be used; the preset fusion threshold can be dynamically determined according to the actual application scenario and needs.

[0169] Secondly, a weakly supervised knowledge graph is constructed based on the qualified entity set, qualified relation set, and qualified attribute set after entity fusion processing. Specifically, the fused qualified entity set serves as the node, and the association links between entities are established through the qualified relation set. Then, the attribute information in the qualified attribute set is used as the attribute label for the entity. During the construction process, a graph database, such as the standard format of Neo4j, can be used to ensure that each entity and relation has a unique identifier, and that attribute information is precisely bound to entities and relations, forming a weakly supervised knowledge graph that is machine-parsable and human-understandable.

[0170] Finally, the quality metrics of the reconstructed weakly supervised knowledge graph are calculated. By calculating the quality metrics from multiple dimensions, the accuracy, standardization, and completeness of the reconstructed weakly supervised knowledge graph can be evaluated.

[0171] Furthermore, the "quality indicators for the reconstructed weakly supervised knowledge graph" include:

[0172] A predetermined number of relation samples are randomly selected from the qualified relation set of the reconstructed weakly supervised knowledge graph. The number of correct relation samples is obtained through sampling evaluation, and the ratio of the number of correct relation samples to the total number of relation samples is calculated as the graph accuracy.

[0173] The total number of all pattern conflict instances generated during the process of generating the reconstructed weakly supervised knowledge graph is counted. The ratio of the total number of pattern conflict instances to the sum of the total number of qualified entities and qualified relations in the weakly supervised knowledge graph is calculated as the graph conflict rate.

[0174] The number of entities whose entity types belong to the predefined entity type set of the domain graph pattern in the qualified entity set of the reconstructed weakly supervised knowledge graph is counted, and the ratio of this number to the total number of entities in the qualified entity set is calculated as the graph completeness.

[0175] The graph accuracy, graph conflict rate, and graph completeness are weighted and fused according to preset weights to obtain the graph quality index of the reconstructed weakly supervised knowledge graph.

[0176] In this embodiment, a predetermined number of relation samples are first randomly selected from the qualified relation set of the reconstructed weakly supervised knowledge graph. The number of correct relation samples is obtained through sampling evaluation, and the ratio of the number of correct relation samples to the total number of extracted relation samples is calculated as the graph accuracy. Specifically, a predetermined number of relation samples, such as 5% of the total number of relations, are extracted from the qualified relation set of the reconstructed graph using stratified random sampling, covering all relation types. During sampling, it is necessary to ensure a balanced sample ratio between newly added relation types and existing relation types. Subsequently, the correctness of the samples is evaluated through manual review and professional text tracing: if the entities connected by the relation have clear association evidence in the original professional text and conform to domain logic, they are determined to be correct relation samples; otherwise, they are determined to be incorrect relation samples. Finally, the graph accuracy is calculated using the formula: Graph Accuracy = Number of Correct Relation Samples / Total Number of Extracted Relation Samples. The closer the graph accuracy is to 1, the higher the accuracy of the graph relations.

[0177] For example, if the reconstructed weakly supervised knowledge graph has a set of 2000 qualified relations, and 100 samples are extracted, and after manual review and text tracing, 92 sample relations are confirmed to be true and valid, then the graph accuracy = 92 / 100 = 0.92.

[0178] Secondly, the total number of all pattern conflict instances generated during the generation of the reconstructed weakly supervised knowledge graph is counted. The ratio of this total number of pattern conflict instances to the sum of the total number of qualified entities and qualified relations in the weakly supervised knowledge graph is calculated as the graph conflict rate. Specifically, the graph conflict rate is calculated as follows: Graph Conflict Rate = Total Number of Pattern Conflict Instances / (Total Number of Qualified Entities + Total Number of Qualified Relations). The closer the graph conflict rate is to 0, the fewer logical conflicts there are in the reconstructed weakly supervised knowledge graph, and the stronger its standardization.

[0179] The total number of all schema conflict instances generated during the process of generating the reconstructed weakly supervised knowledge graph refers to the sum of all entity-level and relation-level conflict instances that were not effectively resolved throughout the entire schema update process.

[0180] For example, if the total number of qualified entities in the weakly supervised knowledge graph is 800, the total number of qualified relations is 2000, and the total number of all pattern conflict instances generated during the process of generating the reconstructed weakly supervised knowledge graph is 14, then the graph conflict rate = 14 / (800+2000) = 0.005.

[0181] Next, the number of entities in the qualified entity set of the reconstructed weakly supervised knowledge graph whose entity type belongs to the predefined entity type set of the domain graph pattern is counted, and the ratio of this number to the total number of entities in the qualified entity set is calculated as the graph completeness. A higher graph completeness indicates that the reconstructed weakly supervised knowledge graph covers more core domain entities, and thus has better knowledge integrity.

[0182] For example, if the total number of qualified entities in the entity set is 800, and the number of entities in the qualified entity set of the reconstructed weakly supervised knowledge graph whose entity type belongs to the predefined entity type set of the domain graph pattern is 720, then the graph completeness = 720 / 800 = 0.9.

[0183] Finally, the graph accuracy, graph conflict rate, and graph completeness are weighted and fused according to preset weights to obtain the graph quality index of the reconstructed weakly supervised knowledge graph. The preset weights can be dynamically determined based on the specific professional field. For example, for rigorous fields such as law and medicine, the graph accuracy weight can be set to the highest, such as 0.5, to prioritize correct relationships; the graph conflict rate weight is next, set to 0.3, to strictly control logical conflicts; and the graph completeness weight is the lowest, set to 0.2, to balance coverage and accuracy. The graph quality index is calculated as follows: Graph Quality Index = w1 × Graph Accuracy + w2 × (1 - Graph Conflict Rate) + w3 × Graph Completeness, where w1, w2, and w3 are the preset weights for graph accuracy, graph conflict rate, and graph completeness, respectively.

[0184] For example, if the graph accuracy, graph conflict rate, and graph completeness are 0.92, 0.005, and 0.9 respectively, and the preset weights for the graph accuracy, graph conflict rate, and graph completeness are 0.5, 0.3, and 0.2 respectively, then the graph quality index = 0.92×0.5 + (1-0.005)×0.3 + 0.9×0.2 = 0.9385, which means that the quality of the reconstructed weakly supervised knowledge graph is relatively excellent.

[0185] In summary, compared to existing technologies, this application optimizes and reconstructs the structured knowledge units based on the updated domain graph model to build a weakly supervised knowledge graph, and calculates the quality indicators of the reconstructed weakly supervised knowledge graph. Thus, by optimizing and reconstructing the structured knowledge units based on the updated domain graph model, achieving unified type identification and entity fusion, the constructed weakly supervised knowledge graph effectively solves problems such as inconsistent types, redundant entities, and chaotic relational logic. Simultaneously, by quantifying the graph quality through quality indicators, its standardization, accuracy, and completeness are accurately assessed, ensuring that the weakly supervised knowledge graph can meet the practical application needs of professional texts for knowledge expression.

[0186] S50: Repeatedly construct the weakly supervised knowledge graph until the quality index of the weakly supervised knowledge graph reaches the preset quality stability threshold, and output the final professional text weakly supervised knowledge graph.

[0187] In this embodiment, an iterative reconstruction process is initiated. After each round of weakly supervised knowledge graph construction, the obtained graph quality index is compared with a preset quality stability threshold. If the preset quality stability threshold is not reached, the process is backtracked to the type identifier update stage to re-encode the structured knowledge unit optimization and reconstruction, weakly supervised knowledge graph construction, and quality index calculation. The above iterative process is repeated until the quality index of the weakly supervised knowledge graph reaches or exceeds the preset quality stability threshold after a certain reconstruction. Then, the iteration is stopped and the final professional text weakly supervised knowledge graph is output.

[0188] The preset quality stability threshold is a quantitative standard dynamically set based on the knowledge application needs of specific professional fields. It is the minimum value for judging that the graph quality meets the requirements of actual use. For example, in legal scenarios, the preset quality stability threshold can be set to 0.9. The preset quality stability threshold defines a clear benchmark for the graph quality, avoiding low-quality graph output. The iterative reconstruction method effectively solves the problems that may exist in a single reconstruction, such as incomplete type alignment, inaccurate entity fusion, and substandard quality. The final weakly supervised knowledge graph output is highly adapted to the updated domain graph pattern and has stable accuracy, low conflict rate, and high completeness.

[0189] In this way, we can obtain a weakly supervised knowledge graph of professional texts that can accurately and comprehensively carry the core knowledge of professional texts, thus meeting the practical application needs of knowledge graphs in professional fields.

[0190] In summary, the embodiments of this application have at least the following technical effects:

[0191] Compared to existing technologies, this application first converts professional text into a structured text format and generates a preprocessed text set. This effectively eliminates the format heterogeneity of professional text, transforming unstructured / semi-structured raw text into a standardized, semantically complete, and parsable structured text form. The generated preprocessed text set can accurately adapt to the text input requirements of weakly supervised knowledge graph construction, improving the accuracy and processing efficiency of subsequent entity, relation, and attribute extraction stages.

[0192] Secondly, based on a preset domain graph pattern, this application utilizes a large language model to extract entities, relationships, and attributes from the preprocessed text in a weakly supervised manner, constructing structured knowledge units. It then identifies pattern conflicts between these structured knowledge units and the domain graph pattern, generating candidate new pattern description documents. This allows for efficient extraction of high-value knowledge from professional texts and construction of structured knowledge units without large-scale manual annotation. Simultaneously, it automatically identifies pattern conflicts beyond the preset domain graph pattern and generates candidate new pattern description documents. This effectively solves the problem of existing technologies struggling to automatically identify, evaluate, and incorporate new knowledge beyond the domain graph pattern in professional texts, ensuring the dynamic updating capability and adaptability of the knowledge graph to professional scenarios.

[0193] Furthermore, this application employs a multi-dimensional confidence assessment mechanism to evaluate the overall confidence level of the candidate new pattern description documents and dynamically updates the domain graph patterns. In this way, multi-dimensional quantitative evaluation accurately determines the overall value and compliance of candidate new pattern description documents. Combined with stratification thresholds, it achieves dynamic stratified updates of the domain graph patterns. This not only filters out low-value and invalid new patterns to avoid graph redundancy but also incorporates high-value emerging types and temporarily stored potential types, effectively balancing the rigor of the domain graph patterns with their adaptability to emerging professional text scenarios. This solves the problem of inaccurate pattern updates caused by single-dimensional or manual evaluation in existing technologies.

[0194] Furthermore, this application optimizes and reconstructs the structured knowledge units based on the updated domain graph model to construct a weakly supervised knowledge graph, and calculates the quality indicators of the reconstructed weakly supervised knowledge graph. Thus, by optimizing and reconstructing the structured knowledge units based on the updated domain graph model, achieving unified type identification and entity fusion, the constructed weakly supervised knowledge graph effectively solves problems such as inconsistent types, redundant entities, and chaotic relational logic. Simultaneously, by quantifying the graph quality through quality indicators, its standardization, accuracy, and completeness are accurately assessed, ensuring that the weakly supervised knowledge graph can meet the practical application needs of professional texts for knowledge expression.

[0195] Finally, this application repeatedly constructs the weakly supervised knowledge graph until its quality index reaches a preset quality stability threshold, outputting the final weakly supervised knowledge graph for the professional text. In this way, a weakly supervised knowledge graph for the professional text can be obtained that accurately and comprehensively carries the core knowledge of the professional text, meeting the practical application needs of knowledge graphs in professional fields.

[0196] Through the above technical solution, this application eliminates the reliance on large-scale manual annotation in strongly supervised methods, effectively solving the problems of high cost and low efficiency in traditional construction models. Simultaneously, by identifying conflicts between structured knowledge units and preset domain graph patterns and generating candidate new pattern description documents, combined with a multi-dimensional confidence evaluation mechanism, it achieves dynamic updates of the domain graph patterns, overcoming the pain points of rigid preset patterns and inability to adapt to emerging knowledge types in professional texts in existing weakly supervised methods. Furthermore, the updated domain graph patterns are used to optimize and reconstruct structured knowledge units and calculate quality indicators. Through iterative construction until the quality indicators reach a preset stable quality threshold, problems such as insufficient knowledge extraction accuracy and entity redundancy or relational errors in the graph are effectively avoided. The final output of a weakly supervised knowledge graph for professional texts combines the advantages of low cost and high efficiency with good adaptability to professional scenarios, and possesses stable accuracy, completeness, and low conflict rate, meeting the rigor and practicality requirements of knowledge graphs in professional fields such as law and medicine.

[0197] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0198] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0199] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0200] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0201] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0202] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0203] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. An automated method for constructing weakly supervised knowledge graphs for specialized texts, characterized in that: The method includes: Convert professional text into structured text format and generate a preprocessed text collection; Based on a preset domain graph pattern, entities, relations and attributes are extracted from the preprocessed text in a weakly supervised manner using a large language model to construct structured knowledge units, and pattern conflicts between the structured knowledge units and the domain graph pattern are identified to generate candidate new pattern description documents. The comprehensive confidence level of the candidate new pattern description documents is evaluated through a multi-dimensional confidence assessment mechanism, and the domain graph pattern is dynamically updated. The structured knowledge units are optimized and reconstructed based on the updated domain graph pattern to construct a weakly supervised knowledge graph, and the quality index of the reconstructed weakly supervised knowledge graph is calculated. Repeatedly construct the weakly supervised knowledge graph until the quality index of the weakly supervised knowledge graph reaches the preset quality stability threshold, and output the final professional text weakly supervised knowledge graph. Specifically, based on a preset domain graph pattern, a large language model is used to extract entities, relationships, and attributes from the preprocessed text in a weakly supervised manner to construct structured knowledge units, including: A weak supervision prompt template is constructed based on a preset domain graph pattern, wherein the weak supervision prompt template includes a set of entity type definitions, a set of relation type definitions, and a set of attribute definitions; The preprocessed text set and the weakly supervised prompt template are input into the large language model to obtain the entity extraction result set, the relation extraction result set and the attribute extraction result set, and the extraction confidence score of each extraction result in the entity extraction result set and the relation extraction result set is recorded. The attribute value format is validated on the attribute extraction result set to generate an attribute validity score; Filter out extraction results with a confidence score lower than the preset extraction confidence threshold and an attribute validity score lower than the preset attribute validity threshold; Based on the filtered entity extraction results, relation extraction results, and attribute extraction results, a structured knowledge unit containing a qualified entity set, a qualified relation set, and a qualified attribute set is constructed. The process includes identifying pattern conflicts between the structured knowledge units and the domain graph patterns, and generating candidate new pattern description documents, including: Based on the qualified entity set, calculate the semantic similarity between the entity types in the qualified entity set and the predefined entity types in the domain graph pattern, and take the maximum semantic similarity as the entity type semantic similarity; When the semantic similarity of the entity type is lower than the preset entity type conflict threshold, it is identified as an entity-level pattern conflict and an entity-level pattern conflict instance is generated. Based on the qualified relation set, calculate the semantic similarity between the relation types in the qualified relation set and the predefined relation types in the domain graph pattern, and take the maximum semantic similarity as the relation type semantic similarity; When the semantic similarity of the relation type is lower than the preset relation type conflict threshold, it is identified as a relation-level pattern conflict and a relation-level pattern conflict instance is generated. Based on the entity-level and relation-level pattern conflict instances, the semantic context of pattern conflict is analyzed using a large language model. Based on the semantic analysis results of the pattern conflict context, a candidate new pattern description document containing candidate new entity type definitions and candidate new relation type definitions is generated by using a large language model adapted to professional scenarios.

2. The method for automated construction of weakly supervised knowledge graphs for specialized texts according to claim 1, characterized in that, Convert professional text into structured text format and generate a preprocessed text collection, including: Professional texts are uniformly converted into standardized Markdown or XML formats to form intermediate structured text; The intermediate structured text is segmented using a semantic integrity segmentation algorithm. During the segmentation process, hierarchical structure information is extracted, including chapter path information, semantic type information, and reference relationship information. Based on the chapter path information, semantic type information, and reference relationship information, add chapter path markers, semantic type markers, and reference relationship markers to each text block, and assign a unique text block identifier to each text block; The text blocks, which have been marked with chapter path tags, semantic type tags, reference relationship tags, and text block identifiers, are organized into a preprocessed text set.

3. The method for automated construction of weakly supervised knowledge graphs for specialized texts according to claim 1, characterized in that, Based on the filtered entity extraction results, relation extraction results, and attribute extraction results, a structured knowledge unit is constructed, comprising a qualified entity set, a qualified relation set, and a qualified attribute set, including: Each entity extraction result after filtering is assigned a unique entity identifier, and the entity type and entity name are recorded to form a qualified entity set; Each filtered relation extraction result is assigned a unique relation identifier, and the relation type and relation name are recorded to form a qualified relation set; The filtered attribute extraction results are associated with the corresponding entity identifiers or relation identifiers to form a qualified attribute set. Based on entity identifiers, relation identifiers, and corresponding attribute extraction results, construct an entity-relation topology; Based on the set of qualified attributes, the text block identifier and extraction confidence score corresponding to each qualified entity and qualified relationship are recorded and integrated to form a structured knowledge unit.

4. The method for automated construction of weakly supervised knowledge graphs for specialized texts according to claim 1, characterized in that, The overall confidence level of the candidate novel pattern description documents is evaluated using a multi-dimensional confidence assessment mechanism, including: Based on the candidate new pattern description document, the candidate new entity type definition and the candidate new relation type definition are parsed out; The log-normalized value of the frequency of occurrence of the candidate new entity type definition in all text blocks of the preprocessed text set is calculated as the frequency confidence of the candidate new entity type definition; Calculate the maximum cosine similarity between the semantic vectors of the candidate new entity type definition and all predefined entity type definitions in the domain graph pattern, and use it as the semantic consistency confidence of the candidate new entity type definition; Calculate the confidence score of the relation structure rationality in the definition of candidate new relation types; The frequency confidence, semantic consistency confidence, and relational structure rationality confidence are weighted and fused according to preset weights to obtain the comprehensive confidence of the candidate new pattern description document.

5. The method for automated construction of weakly supervised knowledge graphs for specialized texts according to claim 4, characterized in that, Calculate the confidence score of the relation structure rationality defined in the candidate new relation type definition, including: Extract the head entity identifier and tail entity identifier connected to the candidate new relation type definition from the relation-level schema conflict instance on which the candidate new schema description document is based; In the qualified entity set of the structured knowledge unit, query the head entity type and head entity extraction confidence score corresponding to the head entity identifier, and the tail entity type and tail entity extraction confidence score corresponding to the tail entity identifier. Calculate the geometric mean of the confidence scores of the head entity extraction and the tail entity extraction, and use it as the confidence score of the basic relation. Calculate the semantic similarity between the head entity type and all entity types in the predefined entity type set of the domain graph pattern, and take the maximum value as the head entity type support. Calculate the semantic similarity between the tail entity type and all entity types in the predefined set of entity types in the domain graph pattern, and take the maximum value as the tail entity type support. Calculate the geometric mean of the head entity type support and the tail entity type support, and use it as the type support coefficient; Multiplying the basic relation confidence score by the type support coefficient yields the relation structure rationality confidence score.

6. The method for automated construction of weakly supervised knowledge graphs for specialized texts according to claim 4, characterized in that, Dynamically updated domain graph patterns include: Set a pattern adoption threshold and a pattern storage threshold, wherein the pattern adoption threshold is greater than the pattern storage threshold; The overall confidence level of the candidate new pattern description documents is compared with the pattern adoption threshold and the pattern storage threshold; When the overall confidence level is greater than the pattern adoption threshold, the candidate new entity type definition in the candidate new pattern description document is added to the entity type definition set of the domain graph pattern, and the candidate new relation type definition is added to the relation type definition set of the domain graph pattern, thus forming an updated domain graph pattern. When the overall confidence level is less than or equal to the pattern adoption threshold and greater than the pattern temporary storage threshold, the candidate new entity type definition and candidate new relation type definition of the candidate new pattern description document are added as supplementary information to the weak supervision prompt template. When the overall confidence level is less than or equal to the pattern temporary storage threshold, the candidate new pattern description document is discarded.

7. The method for automated construction of weakly supervised knowledge graphs for specialized texts according to claim 1, characterized in that, The structured knowledge units are optimized and reconstructed based on the updated domain graph pattern to construct a weakly supervised knowledge graph. The quality metrics of the reconstructed weakly supervised knowledge graph are then calculated, including: Based on the updated domain graph pattern, the type identifiers of the qualified entity set and qualified relation set in the structured knowledge unit are updated, and the entity type of qualified entity and the relation type of qualified relation are mapped to the standard type definition in the updated domain graph pattern. The qualified entity set after the type identifier is updated is processed by entity fusion, merging multiple qualified entities that point to the same real-world object, and synchronously updating the associated qualified relationship set and qualified attribute set; Based on the qualified entity set, qualified relation set, and qualified attribute set after entity fusion processing, a weakly supervised knowledge graph is constructed. Calculate the quality index of the reconstructed weakly supervised knowledge graph.

8. The method for automated construction of weakly supervised knowledge graphs for specialized texts according to claim 7, characterized in that, Calculate the quality metrics of the reconstructed weakly supervised knowledge graph, including: A predetermined number of relation samples are randomly selected from the qualified relation set of the reconstructed weakly supervised knowledge graph. The number of correct relation samples is obtained through sampling evaluation, and the ratio of the number of correct relation samples to the total number of relation samples is calculated as the graph accuracy. The total number of all pattern conflict instances generated during the process of generating the reconstructed weakly supervised knowledge graph is counted. The ratio of the total number of pattern conflict instances to the sum of the total number of qualified entities and qualified relations in the weakly supervised knowledge graph is calculated as the graph conflict rate. The number of entities whose entity types belong to the predefined entity type set of the domain graph pattern in the qualified entity set of the reconstructed weakly supervised knowledge graph is counted, and the ratio of this number to the total number of entities in the qualified entity set is calculated as the graph completeness. The graph accuracy, graph conflict rate, and graph completeness are weighted and fused according to preset weights to obtain the graph quality index of the reconstructed weakly supervised knowledge graph.

Citation Information

Patent Citations

  • Power industry knowledge graph construction method fused with large-scale language model

    CN118627604A

  • Power equipment fault diagnosis method and system based on dynamic knowledge graph and large model collaborative reasoning

    CN120996202A