Intelligent knowledge base system based on multi-domain model collaboration and construction method
The multi-domain model collaborative architecture with a three-tier storage system and dynamic weight algorithm addresses adaptability and compliance issues in intelligent knowledge management, enhancing efficiency and reducing risks in specialized fields.
Patent Information
- Application Number
- CN202510456304.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-15
AI Technical Summary
The existing intelligent knowledge base system has problems such as insufficient field adaptability, low knowledge generation efficiency, weak dynamic optimization capabilities and high compliance risks in the fields of medical care, law, etc.
The multi-domain model collaborative architecture is adopted, combined with a three-dimensional matching matrix and a three-level storage system, and precise generation and management of cross-domain knowledge is achieved through domain intention analysis, dynamic model routing, knowledge cleaning pipelines and dynamic weighting algorithms.
The accuracy and efficiency of knowledge generation have been improved, compliance has been guaranteed, and the risk of data leakage has been greatly reduced.
Smart Images

Figure CN120317345A_ABST
Abstract
Description
1. Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to an intelligent knowledge base system and a construction method based on the collaboration of multi-domain models, which are applicable to knowledge management and service scenarios in vertical fields such as medical, legal, and financial. The utilization rate in small company scenarios is the highest, and the maximum utility of this project can be exerted. 2. Background Art
[0002] The existing intelligent knowledge base systems generally have the following technical defects: Insufficient domain adaptability: Traditional systems rely on a single general model and cannot meet the high-precision requirements for professional knowledge in fields such as medical and legal. For example, the medical field needs to follow the FDA certification standard, and the legal field needs to distinguish the levels of legal effect, while the general model is difficult to accurately handle these professional requirements.
[0003] Low knowledge generation efficiency: The knowledge production process lacks an automated cleaning mechanism, and the cost of manual review is high and the efficiency is low. For example, a certain legal knowledge base processes 100,000 pieces of knowledge per day on average, and the time-consuming proportion of manual conflict verification reaches 40%, resulting in a knowledge update cycle as long as 72 hours.
[0004] Weak dynamic optimization ability: The knowledge storage system is static and cannot adjust the knowledge quality in real time according to user feedback. For example, 80% of the low-frequency and low-confidence knowledge in a certain financial knowledge base occupies 60% of the storage resources, resulting in a 30% decrease in the retrieval efficiency of core knowledge.
[0005] Compliance risk: The knowledge transmission in sensitive fields (such as medical and legal) lacks an encryption mechanism, and there is a risk of data leakage. For example, a certain medical system did not encrypt patient data, resulting in the leakage of 2,000 medical records and triggering a HIPAA compliance lawsuit. 3. Summary of the Invention
[0006] 1. Object of the Invention The present invention aims to solve the problems of insufficient domain adaptability, low knowledge generation efficiency, weak dynamic optimization ability, and high compliance risk in the prior art, and provides an intelligent knowledge base system and a construction method based on the collaboration of multi-domain models to achieve accurate generation, efficient cleaning, dynamic storage, and compliance management of cross-domain knowledge.
[0007] 2. Technical Solution The core technical solution of the present invention includes: Collaborative architecture of multi-domain models: The system is divided into three major modules: an intelligent scheduling layer, a knowledge production layer, and an intelligent knowledge base. Through technologies such as domain intention parsing, model dynamic routing, and knowledge cleaning pipelines, accurate generation and management of cross-domain knowledge are achieved.
[0008] Three-dimensional matching matrix: The model dynamic routing module establishes a three-dimensional matching matrix that includes the domain, model, and credibility, and automatically selects the optimal model according to the query domain and demand type. For example, for clinical guideline queries in the medical field, the Med-PaLM 2 model certified by the FDA is preferentially called, and for contract review in the legal field, the power-law intelligent contract review model is preferentially called.
[0009] Three-level storage system: The intelligent knowledge base adopts a three-level storage structure of a core library, an extended library, and a temporary library, and dynamically adjusts the storage location according to the knowledge score. The core library stores high-frequency and high-confidence knowledge (user satisfaction > 95%), the extended library stores medium-confidence knowledge (80% ≤ satisfaction < 95%), and the temporary library stores knowledge to be verified (satisfaction < 80%).
[0010] Dynamic weight algorithm: The feedback optimization module calculates the knowledge score in real time through a comprehensive calculation of explicit scores (60%), implicit behavior scores (30%), and domain expert calibration (10%), triggering the promotion, demotion, or elimination of knowledge.
[0011] 3. Beneficial effects Compared with the prior art, the present invention has the following remarkable advantages: Precise domain adaptation: The three-dimensional matching matrix realizes the precise mapping between the model and the domain, and the knowledge generation accuracy rate is increased to 98% (medical field) and 95% (legal field), which is 20% and 15% higher than that of traditional general models respectively.
[0012] Improved knowledge generation efficiency: The knowledge cleaning pipeline automatically processes processes such as format standardization and conflict verification, with a daily processing volume of up to 500,000 items, the proportion of time-consuming for manual review is reduced to less than 5%, and the knowledge update cycle is shortened to within 2 hours.
[0013] Optimized resource allocation: The three-level storage system improves the retrieval efficiency of core knowledge by 40%, increases the storage resource utilization rate by 30%, and the elimination rate of low-frequency knowledge reaches 60%.
[0014] Compliance guarantee: In the medical field, knowledge transmission uses AES-256 encryption, and in the legal field, the legal effect level is marked, which fully complies with compliance standards such as HIPAA and GDPR, and the data leakage risk is reduced to less than 0.01%. IV. Description of the Drawings
[0015] Figure 1 is the abstract drawing, which is the system architecture diagram of the intelligent knowledge base. In the form of a black-and-white line drawing, it clearly shows the three-level drive architecture and interaction relationship of the intelligent scheduling layer, the knowledge production layer, and the intelligent knowledge base.
[0016] Intelligent Scheduling Layer (Tag 100): It includes a domain intent parsing engine (101) and a model dynamic routing module (102). The domain intent parsing engine (101) completes query domain recognition (such as medical, legal, financial) and requirement grading (distinguishing factual queries from decision-making requirements) within 500ms through a pre-built 20+ vertical domain term library (keyword graph) and context semantic analysis; the model dynamic routing module (102) establishes a three-dimensional matching matrix of "domain - model - credibility", and dynamically selects the target model according to the parsing result - for example, in the medical clinical guideline scenario, the Med-PaLM 2 model certified by the FDA is preferentially called, and in the legal contract review scenario, the PowerLaw intelligent contract review model is preferentially called.
[0017] Knowledge Production Layer (Tag 200): It includes a professional model cluster (201) and a knowledge cleaning pipeline (202). The professional model cluster (201) accesses third-party authoritative models in fields such as medical, legal, and financial (an API compliance agreement needs to be signed); the knowledge cleaning pipeline (202) performs three-stage processing on the model output: ① format standardization (uniformly converted to the JSON-LD format), ② conflict verification (manual review is triggered when the differences in the conclusions of multiple models exceed the preset threshold), ③ timeliness marking (marking time attributes such as the guideline version number and drug approval date corresponding to the knowledge).
[0018] Intelligent Knowledge Base (Tag 300): It includes a three-level storage system (core library 301, extension library 302, temporary library 303) and a feedback optimization module (304). The core library (301) stores high-frequency and high-confidence knowledge that has passed double-blind cross-verification and has a user satisfaction rate higher than 95%, using distributed key-value storage (response time less than 50ms); the extension library (302) stores medium-confidence knowledge that has passed single-model verification and has a satisfaction rate between 80% - 95%, using a graph database for storage; the temporary library (303) stores unverified knowledge with a satisfaction rate lower than 80% for a single model output. The feedback optimization module (304) calculates the knowledge score in real time and triggers promotion or demotion through a dynamic weight algorithm of explicit scoring (user 1-5 star rating), implicit behavior score (calculated based on buried point data such as query conversion rate, page stay duration, and secondary retrieval rate), and domain expert calibration weight (regularly reviewed and assigned by the domain expert committee) - for example, knowledge with a score higher than 90 points for 30 consecutive days is promoted from the extension library to the core library, and knowledge with a score lower than 60 points and no feedback for 45 days is marked for elimination.
[0019] Figure 2 is a schematic diagram of the three-dimensional matching matrix, intuitively presenting the dynamic mapping logic of domain, model, and credibility.
[0020] The matrix has "domain", "model", and "credibility" as three dimensions, with specific classifications marked for each dimension: The domain dimension includes more than 20 vertical domains such as healthcare, law, and finance; the model dimension marks third-party authoritative models such as Med-PaLM 2 and the power-law intelligent contract review model; the credibility dimension includes parameters such as the historical accuracy rate of the model, response time, and authoritative certification status (such as FDA certification).
[0021] The routing rules distinguish priorities through different shades of gray lines: Models with authoritative certifications (such as FDA-certified models in the healthcare domain) use dark gray nodes, and their weight coefficients are 1.5 times that of ordinary models; ordinary models use light gray nodes. When a user queries for specific domain requirements, the matrix automatically matches the corresponding domain and requirement type, and preferentially selects target models with high historical accuracy rates and fast response speeds (such as preferentially calling FDA-certified models for healthcare clinical guideline scenarios and professional contract review models for legal contract review scenarios), and generates the final knowledge through cross-model confidence weighted fusion - that is, weighted calculation is performed based on the historical accuracy rate and output confidence of each model to ensure the accuracy and authority of the results.
[0022] Figure 3 is a flowchart of the dynamic weight algorithm, showing the multi-dimensional calculation logic of knowledge scoring.
[0023] The explicit scoring module (accounting for 60%): Users rate from 1 to 5 stars through the interaction interface, directly reflecting their satisfaction with the knowledge. For example, a certain medical guideline receives a 4.8-star rating, corresponding to a relatively high explicit score, indicating a high level of user recognition of this knowledge.
[0024] The implicit behavior module (accounting for 30%): Calculated through buried-point data, it includes three core indicators: ① Query-click conversion rate (the ratio of the number of clicks to the number of queries, the higher the conversion rate, the higher the score), ② Page stay duration (the time users view the knowledge, the longer the stay time, the higher the score), ③ Secondary search rate (the frequency at which users re-search related questions, the higher the frequency, the lower the score). The implicit behavior score is generated by integrating these three indicators, comprehensively reflecting the potential needs of users and the actual value of the knowledge.
[0025] The expert calibration module (accounting for 10%): The domain expert committee conducts manual reviews of knowledge in high-risk domains (such as medical treatment suggestions and legal article interpretations) every quarter, and adjusts and calibrates the weights according to the authority and accuracy of the knowledge. For example, the calibration weight of medical knowledge sourced from FDA guidelines increases, while the calibration weight of knowledge from non-authoritative sources decreases, ensuring the reliability of knowledge in professional domains.
[0026] The three modules generate the final score through weighted calculation: the explicit score accounts for 60%, the implicit behavior score accounts for 30%, and the expert calibration weight accounts for 10%. The final score drives the dynamic flow of knowledge among the three-level libraries - high-score knowledge (such as scoring above 90 for 30 consecutive days) is promoted to the core library, and low-score knowledge (such as scoring below 60 and having no feedback for a long time) is marked for elimination, forming a self-evolution mechanism for knowledge quality. V. Specific Implementation Modes
[0027] Domain intention parsing: The domain intention parsing engine identifies the domain of the user's query "Treatment Guidelines for Acute Myocardial Infarction" as "medical" and the demand type as "clinical guidelines" through a keyword graph (including more than 3,000 medical terms) and context semantic analysis.
[0028] Model dynamic routing: The three-dimensional matching matrix selects the Med-PaLM 2 model certified by the FDA to generate knowledge based on the domain (medical), demand type (clinical guidelines), and credibility (the historical accuracy rate of the model > 95%).
[0029] Knowledge cleaning: The knowledge cleaning pipeline standardizes the format of the treatment plan output by the model (such as unifying it to the JSON format) and performs conflict verification with authoritative guidelines (such as the ACC / AHA guidelines). It automatically passes when the difference rate < 1%.
[0030] Knowledge storage: The cleaned knowledge is stored in the core library, using a distributed key-value storage (such as Redis), and the retrieval response time < 50ms.
[0031] Feedback tuning: The explicit score (4.8 / 5 stars) and implicit behavior score (query conversion rate 92%) of the user for the treatment plan trigger the update of the knowledge score and maintain the storage status of the core library.
[0032] Domain intention parsing: Parse the domain of the user's query "Review of the Legality of Labor Contract Clauses" as "law" and the demand type as "contract review".
[0033] Model dynamic routing: Select the power-law intelligent contract review model (historical accuracy rate 98%) to generate review opinions.
[0034] Knowledge cleaning: Perform conflict verification on the review opinions output by multiple models. When the difference rate > 5%, trigger manual review to ensure the accurate marking of the legal effect level.
[0035] Knowledge storage: The review opinions are stored in the extended library (user satisfaction 90%), using a graph database (such as Neo4j) for storage, supporting association relationship queries.
[0036] Feedback Optimization: If the review opinion score has been >90 for 30 consecutive days, it will be promoted to the core library; if the score <60 and there is no feedback for 45 days, it will be marked as to be phased out and trigger manual review. VI. Advantages and Innovation Points of the Invention
[0037] Technological Innovation: Multi-domain Model Collaboration: Through a three-dimensional matching matrix, precise matching of domain, model, and credibility is achieved, breaking through the domain limitations of traditional systems.
[0038] Dynamic Knowledge Management: The combination of a three-level storage system and a dynamic weight algorithm realizes real-time optimization of knowledge and efficient allocation of resources.
[0039] Compliance Design: Encrypted transmission of knowledge in sensitive fields and legal effect level marking meet the industry's compliance requirements.
[0040] Application Value: Efficiency Improvement: The knowledge generation efficiency is more than 10 times higher than that of traditional systems, and the manual review cost is reduced by 80%.
[0041] Quality Assurance: The knowledge accuracy rate is increased to over 95%, and the retrieval efficiency of core knowledge is increased by 40%.
[0042] Risk Control: The data leakage risk is reduced to 0.01%, and no compliance risks occur.
Claims
1. An intelligent knowledge base system based on the collaboration of multi-domain models, characterized in that, Including: An intelligent scheduling layer, configured to parse the domain intent of a user query and dynamically schedule an adapted professional domain model. The intelligent scheduling layer includes: (1) A domain intent parsing engine that identifies the query domain through a keyword graph and context semantic analysis, distinguishing factual queries from decision-making requirements; (2) A model dynamic routing module that establishes a three-dimensional matching matrix including domain, model, and credibility, and selects a target model according to the domain intent and requirement level. A knowledge production layer that generates professional domain knowledge through the target model and performs cleaning processing. The knowledge production layer includes: (1) A professional model cluster that contains at least three third-party authoritative models in vertical domains such as medical, legal, and finance; (2) A knowledge cleaning pipeline that standardizes the format, verifies conflicts, and marks timeliness of the knowledge output by the model. An intelligent knowledge base that stores hierarchical knowledge and dynamically optimizes it according to user feedback. The intelligent knowledge base includes: (1) A three-level storage system, where the core library stores high-frequency and high-credibility knowledge, the expansion library stores medium-confidence knowledge, and the temporary library stores knowledge to be verified; (2) A feedback tuning module that collects explicit user feedback and implicit behavior data, calculates knowledge scores through a dynamic weight algorithm, and triggers the promotion, demotion, or elimination of knowledge.
2. The system according to claim 1, wherein, The domain identification process of the domain intent parsing engine includes: Construct a keyword graph containing more than 20 vertical domains, associating domain-specific terms with general semantics; Based on the keyword graph and context semantic analysis, accurately locate the query domain within 500ms.
3. The system according to claim 1, characterized in that, The three-dimensional matching matrix of the model dynamic routing module satisfies: When the query domain is medical and the demand type is clinical guidelines, preferentially call the Med-PaLM 2 model certified by the FDA; When the query domain is legal and the demand type is contract review, preferentially call the power-law intelligent contract review model; The weight coefficient of the authoritative certification model is 1.5 times that of the ordinary model, and the weight is dynamically adjusted according to the historical accuracy of the model.
4. The system according to claim 1, wherein The conflict verification mechanism of the knowledge cleaning pipeline includes: Detect claim conflicts for the same knowledge output by multiple models, and trigger manual review when the conclusion difference exceeds a preset threshold; Perform confidence-weighted fusion on conflict-free knowledge, and the fusion method is the sum of the products of the historical accuracy weights of each model and the output confidence.
5. The system according to claim 1, characterized in that, The access criteria of the three-level storage system are: Core library: Knowledge needs to pass double-blind cross-verification and the user satisfaction is greater than 95%, and uses distributed key-value storage; Expansion library: Knowledge needs to pass single-model verification and the user satisfaction is between 80% and 95%, and uses graph database storage; Temporary library: Store knowledge output by a single model and with user satisfaction lower than 80%, and uses distributed file system storage.
6. The system according to claim 1, characterized in that, The dynamic weight algorithm of the feedback tuning module is: The knowledge score is composed of an explicit score (60%), an implicit behavior score (30%), and a domain expert calibration weight (10%). Among them, the explicit score is the user's 1-5 star rating, and the implicit behavior score is calculated based on the query conversion rate and the stay duration.
7. The system according to claim 1, characterized in that, It also includes a compliance guarantee module, configured to: Perform AES-256 encryption on the knowledge transmission in the medical field, which complies with HIPAA compliance standards; Mark the legal effect levels of the knowledge in the legal field to distinguish sources such as administrative regulations and judicial interpretations.
8. An intelligent knowledge base construction method based on multi-domain model collaboration, characterized in that, It includes the following steps: Identify the field and type of requirements of the user's query through the domain intention parsing engine; Select the target model from the professional model cluster through the model dynamic routing module; Standardize the format, check for conflicts, and mark the timeliness of the knowledge output by the target model through the knowledge cleaning pipeline; Store the cleaned knowledge in the corresponding level library of the intelligent knowledge base, and dynamically adjust the knowledge score according to user feedback, triggering the promotion, demotion, or elimination of the knowledge.
9. The method according to claim 8, characterized in that, The triggering condition for knowledge promotion in step 4 is: When the score of a certain piece of knowledge is greater than 90 points for 30 consecutive days, it is promoted from the extended library to the core library; When the score of a certain piece of knowledge is lower than 60 points and there is no new feedback within 45 days, it is marked as to be eliminated and triggers manual review.
10. The method according to claim 8, characterized in that The selection process of the target model in step 2 includes: Construct a three-dimensional matching matrix including the field, model, and credibility; Select the model with the highest historical accuracy and a response time less than 500ms from the three-dimensional matching matrix according to the query field and type of requirements.
Citation Information
Cited By
Aerospace big data intelligent data governance system and method based on four-database collaboration and knowledge generation
CN121834012A