A knowledge base management method and device, computer equipment and a storage medium

By performing quality checks and discrepancy analysis on the knowledge base, expired or conflicting knowledge data is automatically identified and updated, solving the problems of untimely updates and low automation in knowledge base management, and achieving efficient and reliable knowledge iteration.

CN122114103APending Publication Date: 2026-05-29PING AN TECH (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-14
Publication Date
2026-05-29

Smart Images

  • Figure CN122114103A_ABST
    Figure CN122114103A_ABST
Patent Text Reader

Abstract

The application discloses a kind of knowledge base management method, device, computer equipment and storage medium, the knowledge base management method includes steps: quality detection is carried out to knowledge data in knowledge database, obtains quality detection result;According to quality detection result, determine the target knowledge data needing to update;Get reference knowledge data, and calculate the knowledge difference degree between reference knowledge data and target knowledge data;According to the corresponding relationship of knowledge difference degree and data update strategy, determine target data update strategy, update target knowledge data.This method realizes the precision and automation of knowledge update by the dual mechanism of quality detection and difference degree analysis, effectively improves the reliability and efficiency of knowledge base management, is suitable for dynamic maintenance and intelligent optimization of knowledge base system in financial technology, medical health and other business fields, can provide effective knowledge support for decision support and intelligent service in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge management, and more particularly to a knowledge base management method, apparatus, computer equipment, and storage medium. Background Technology

[0002] Currently, financial insurance, legal compliance, and healthcare / elderly care systems all include modules for managing knowledge bases, which are essential for their business operations. The effectiveness of knowledge data management directly impacts the efficiency of business processes. However, existing knowledge base management methods have several shortcomings: the validity of some knowledge content is difficult to identify efficiently, leading to outdated content not being updated in a timely manner, potentially affecting the accuracy of knowledge application; the basis for knowledge updates is not clear enough, making it difficult to guarantee the reliability of updates; and the level of automation in knowledge base management is insufficient, making it difficult to adapt to the high-frequency, large-scale needs of knowledge iteration. These problems restrict the efficiency and effectiveness of knowledge base management, necessitating a better solution. Summary of the Invention

[0003] This invention provides a knowledge base management method, apparatus, computer equipment, and storage medium to solve the problems of untimely knowledge updates, poor update effects, and low automation in existing knowledge base management methods.

[0004] A knowledge base management method includes the following steps: performing quality inspection on knowledge data in a knowledge database to obtain quality inspection results; determining, based on the quality inspection results, outdated or conflicting knowledge data from the knowledge database as target knowledge data that needs to be updated; obtaining benchmark knowledge data corresponding to the target knowledge data from an external data source and calculating the knowledge difference between the benchmark knowledge data and the target knowledge data; determining a target data update strategy corresponding to the target knowledge data based on the correspondence between the knowledge difference and data update strategies, and updating the target knowledge data according to the target data update strategy.

[0005] A knowledge base management device includes: a quality inspection module for inspecting the quality of knowledge data in a knowledge database and obtaining quality inspection results; a target determination module for determining, based on the quality inspection results, the knowledge data in the knowledge database that needs to be removed due to knowledge expiration or knowledge conflict as target knowledge data for updating; an auditing module for obtaining benchmark knowledge data corresponding to the target knowledge data from an external data source and calculating the knowledge difference between the benchmark knowledge data and the target knowledge data; and a knowledge update module for determining a target data update strategy corresponding to the target knowledge data based on the correspondence between the knowledge difference and data update strategies, and updating the target knowledge data according to the target data update strategy.

[0006] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned knowledge base management method.

[0007] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned knowledge base management method.

[0008] The aforementioned technical solutions for knowledge base management methods, devices, computer equipment, and storage media include the following steps: quality inspection of knowledge data in a knowledge database to obtain quality inspection results; identification of outdated or conflicting knowledge data from the knowledge database as target knowledge data requiring updating based on the quality inspection results; acquisition of benchmark knowledge data corresponding to the target knowledge data from an external data source, and calculation of the knowledge difference degree between the benchmark knowledge data and the target knowledge data; determination of a target data update strategy corresponding to the target knowledge data based on the correspondence between the knowledge difference degree and the data update strategy, and updating the target knowledge data according to the target data update strategy. This method, through a dual mechanism of quality inspection and difference degree analysis, achieves precise and automated knowledge updates, effectively improving the reliability and efficiency of knowledge base management, and providing a replicable technical path for high-frequency, large-scale knowledge iteration. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of an application environment for a knowledge base management method according to an embodiment of the present invention; Figure 2 This is a flowchart of a knowledge base management method according to an embodiment of the present invention; Figure 3 This is a flowchart of step S1 in a knowledge base management method according to an embodiment of the present invention; Figure 4 This is another flowchart of step S1 in the knowledge base management method of one embodiment of the present invention; Figure 5 This is a schematic diagram of a knowledge base management device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] The knowledge base management method provided in this embodiment of the invention can be applied to, for example... Figure 1 The application environment shown. Specifically, this knowledge base management method is applied in a knowledge base management system, which includes, for example, […]. Figure 1 The diagram illustrates a client and server that communicate over a network to effectively manage knowledge data in the knowledge database, ensuring timely and accurate updates and guaranteeing the data's accuracy and timeliness. The client, also known as the user terminal, is the program that provides local services to the client, corresponding to the server. Clients can be installed on, but are not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be a standalone server or a server cluster consisting of multiple servers.

[0013] In one embodiment, such as Figure 2 As shown, a knowledge base management method is provided, which is applied to... Figure 1 Taking the server in the example, the following steps are included: Step S1: Perform quality checks on the knowledge data in the knowledge database and obtain the quality check results.

[0014] It should be noted that a knowledge database refers to a collection of information with practical application value in a specific field, including but not limited to text, images, structured data, and metadata. It is classified, labeled, and associated using preset rules, supporting efficient knowledge extraction, organization, updating, and retrieval. Therefore, before step S1, the knowledge database needs to be initialized and configured, including data source access, standardization processing, and index construction, to ensure the accuracy and efficiency of subsequent knowledge updates and retrieval.

[0015] In this embodiment, internal and external knowledge sources (such as regulatory policies, product terms, customer service documents, etc.) are integrated into a knowledge database. Raw knowledge data is collected from multiple sources, standardized, and then transformed into a unified knowledge structure, which is subsequently stored in the knowledge database. The knowledge structure includes knowledge content, publication time, data source, and metadata tags. i A knowledge structure can be represented as: ;in, Indicates the first i The knowledge content of each knowledge data item Indicates the first i The publication time of the knowledge data. Indicates the first i The data sources for each knowledge item include official sources, internal sources, and manual review. Indicates the first i Metadata tags for each knowledge data item (including domain classification, scope of application, etc.).

[0016] Furthermore, a vector database is established using vectorized retrieval and semantic indexing technology. First, knowledge data is vectorized to generate corresponding semantic vectors. These semantic vectors are then divided into knowledge topic clusters. Based on these clusters, knowledge data is categorized and stored, and a cross-source association index is established. This allows multiple semantically similar knowledge data points in the database to be associated with the same knowledge topic (reflected by metadata tags), thus achieving efficient organization and real-time updates of multi-source heterogeneous knowledge data and supporting semantic similarity queries and version association. For example, the Sentence-BERT (SBERT) embedding model is used to encode knowledge content into high-dimensional vectors and store them in the vector database. Then, cosine similarity calculation is used to achieve efficient semantic-level retrieval, ensuring that users can accurately match relevant knowledge fragments during queries.

[0017] For example, the knowledge database serves as a financial knowledge base, aggregating multi-source data including insurance industry regulations, product terms, and claims cases. The vector database constructed using the aforementioned methods can accurately respond to frequently requested topics such as "critical illness insurance payout standards." When a user enters "Is thyroid cancer considered a critical illness?", the system can retrieve the latest regulatory documents and product terms from different sources based on semantic matching, and identify the scope of application using a tagging system, returning a consistent conclusion: since 2023, in the vast majority of critical illness insurance products, thyroid cancer is determined according to TNM staging, with stage III and above included in the coverage. This mechanism significantly improves the consistency and timeliness of knowledge services.

[0018] After the knowledge database is established, step S1 is performed to conduct quality checks on the knowledge data in the database. This involves identifying and filtering out knowledge data that is outdated or conflicting, and generating corresponding quality check results. The quality check includes two dimensions: first, timeliness checks based on timestamps to determine if the knowledge data is outdated; knowledge data exceeding a preset update cycle or whose publication date is too far in the past is marked as "expired" and awaiting verification; second, logical consistency checks based on content conflicts, by comparing multiple knowledge data from different sources under the same knowledge topic to determine if there are knowledge conflicts among them, identifying knowledge data with contradictory expressions, overlapping scopes of application, or conflicting policy versions, and marking them as "knowledge conflicting."

[0019] like Figure 3 As shown, timeliness detection specifically includes the following sub-steps: Step S111: Evaluate the credibility of the knowledge data to obtain the knowledge credibility of the knowledge data.

[0020] In this embodiment, the knowledge credibility of knowledge data is calculated based on the data source, semantic consistency of knowledge content, and external verification results. Specifically, the higher the authority of the data source, the closer the semantic consistency is to domain consensus, and the higher the knowledge credibility of knowledge data verified by authoritative external channels, the better.

[0021] Specifically, the knowledge data's source score is determined based on its origin; a consistency score is obtained based on the semantic consistency between the knowledge data and a pre-defined knowledge set; and a verification score is obtained by verifying the knowledge data through user feedback or automatic system verification. Finally, the knowledge credibility of the knowledge data is calculated using a weighted fusion method based on the source score, consistency score, and verification score. (Corresponding to the...) i The knowledge credibility of a knowledge data point can be represented as: ; in, Indicates the first i The source score of each knowledge data item Indicates the first i Consistency score of knowledge data Indicates the first i The validation score of each knowledge data point , and These are the weight coefficients for the corresponding scores, and they satisfy... .

[0022] It should be noted that the source score is obtained through a tiered evaluation based on the data source, with official sources receiving the highest score, followed by internal sources, and then those subject to manual review. The consistency score is obtained by calculating the semantic similarity between the knowledge data and existing authoritative knowledge sets. The verification score combines the frequency of user feedback with the results of automatic system verification, reflecting the reliability of knowledge in practical applications. The weighting coefficients are determined by the characteristics of the knowledge domain and the application scenario to ensure the rationality and adaptability of the credibility assessment.

[0023] In this embodiment, the source score, consistency score, and verification score are weighted and summed according to preset weights to obtain the knowledge credibility. This can be used to quantitatively evaluate the reliability of each knowledge data. The higher the knowledge credibility, the stronger the reliability of the knowledge data, and the more likely it is to be used for subsequent knowledge services.

[0024] Step S112: Use knowledge credibility to evaluate the timeliness of knowledge data and obtain a timeliness score for the knowledge data.

[0025] It should be noted that knowledge credibility is used as the initial input for timeliness assessment. By introducing a time decay coefficient, the decay effect of knowledge credibility over time is modeled, and the current timeliness score after the knowledge credibility decays over time is calculated. This allows for dynamic assessment of the timeliness of knowledge data, quantifying the credibility loss caused by knowledge over time, and dynamically adjusting the priority of knowledge data in knowledge services. The lower the timeliness score, the more significant the decline in reliability of the knowledge data due to time. The time decay coefficient is dynamically set according to the update frequency of the knowledge domain, with higher decay rates assigned to domains with high-frequency updates to more sensitively reflect knowledge aging.

[0026] In this embodiment, it is necessary to first obtain the time decay coefficient of the knowledge data and the historical update time of the last update of the knowledge data. The historical update time is obtained through the timestamp field in the knowledge data. The time decay coefficient is determined based on the difference between the current time and the historical update time, combined with the time sensitivity of the knowledge domain. Then, a timeliness score is calculated based on the knowledge credibility, the time decay coefficient, and the historical update time. Specifically, the current knowledge credibility is multiplied by an exponential decay function of the time decay coefficient to obtain the timeliness score at the current moment after decay. This score comprehensively reflects the reliability changes of the knowledge data over time, ensuring that highly timely knowledge is prioritized for serving the application system.

[0027] Specifically, the timeliness score can be expressed as: ; in, Indicates time Time i Timeliness score of knowledge data Indicates the first i Knowledge credibility of individual knowledge data Indicates the first i The time decay coefficient of knowledge data Indicates the first i The current time of each piece of knowledge data Indicates the first i The historical update time of each piece of knowledge data.

[0028] In some embodiments, to prevent the overall decay of knowledge data from affecting data stability, a rolling reinforcement mechanism can be adopted. If knowledge data is frequently accessed or verified as effective, its decay coefficient can be dynamically reduced, thereby slowing down its timeliness decay and creating a "hot knowledge" reinforcement effect to enhance the lasting influence of key knowledge. By introducing user feedback and usage frequency as reinforcement signals, high-value knowledge can be dynamically identified and its lifespan appropriately extended. This ensures timeliness while preventing high-quality data from being prematurely marginalized due to time decay, achieving adaptive optimization of knowledge base management. For example, in a financial risk control knowledge base, a rule for judging abnormal market fluctuations, after being repeatedly accessed and verified as effective by a real-time transaction monitoring system, has a significantly higher usage frequency than similar rules. Based on this, its decay coefficient is automatically reduced, maintaining its timeliness score at a high level and extending its lifespan. This rule accumulates positive feedback through continuous verification, further strengthening its weight in risk identification and forming a virtuous cycle of "the more effective, the more frequently used, the more lasting."

[0029] Step S113: Obtain the quality inspection results based on the timeliness score.

[0030] In this embodiment, the quality inspection result is determined by a set timeliness threshold. If the timeliness score of the knowledge data is lower than the timeliness threshold, the corresponding knowledge data is marked as "knowledge expired" and pending verification, and the subsequent data review process is triggered; otherwise, it is determined as "available".

[0031] It should be noted that when knowledge data is marked as "expired," a review task is automatically pushed to the relevant maintenance entity, and the duration of the pending verification status is recorded. For knowledge data in the pending verification state, its access permissions in critical decision-making scenarios are restricted, allowing only degraded use in low-risk application scenarios, and users are simultaneously alerted to the risk of information timeliness issues. After review and confirmation of updates, its credibility and timeliness scores are recalculated, restoring it to an "available" state; if verification cannot be completed within the specified period, it is temporarily removed to prevent outdated knowledge from causing misleading information. This mechanism achieves closed-loop management of the knowledge lifecycle, ensuring the dynamic reliability of the knowledge service system.

[0032] For example, an entry about drug dosage in a medical knowledge base, due to prolonged inactivity and updated medical guidelines, gradually saw its timeliness score drop below the timeliness threshold, automatically marking it as "expired knowledge" and triggering a review reminder. Upon receiving the reminder, relevant experts reviewed the entry, confirmed its deviation from the latest clinical guidelines, and immediately initiated an update process, revising the dosage parameters and supplementing the cited evidence. The updated knowledge data, after review, was re-entered into the database, its timeliness score rebounded, restoring its "usable" status, and simultaneously notifying the business systems that had previously accessed the data of a risk warning. This process ensures the accuracy and authority of medical decision support information, effectively preventing diagnostic and treatment risks caused by outdated knowledge, and simultaneously enhancing the knowledge base's self-evolution capability through a closed-loop feedback mechanism.

[0033] like Figure 4 As shown, logical consistency checking specifically includes the following sub-steps: Step S121: Perform similarity assessment on multiple knowledge data for each knowledge topic to obtain the knowledge similarity of multiple knowledge data for each knowledge topic.

[0034] It should be noted that a knowledge database comprises multiple knowledge datasets, each corresponding to a specific knowledge topic. Knowledge similarity refers to the degree of semantic similarity between different knowledge datasets under the same knowledge topic. By calculating the semantic similarity of multiple knowledge datasets under the same topic, potential risks of knowledge conflicts can be identified. Higher knowledge similarity indicates greater knowledge overlap and a lower likelihood of potential conflicts.

[0035] In this embodiment, a knowledge consistency detection mechanism is designed to detect knowledge conflicts, calculating the implicit logical consistency between knowledge data based on a semantic vector model. For a set of knowledge data under the same knowledge topic, pairwise comparisons are performed to generate a knowledge similarity matrix. Each element in the knowledge similarity matrix represents the knowledge similarity between two corresponding knowledge data. The knowledge similarity between two knowledge data under the same knowledge topic can be expressed as: ; in, Representing the first under the same knowledge topic i The knowledge data and the first j Knowledge similarity between knowledge data Indicates the first i The first knowledge data i Semantic vectors Indicates the first j The first knowledge data j Semantic vectors Indicates the first i The semantic vector of the first i Vector magnitude, Indicates the first j The semantic vector of the first j Vector magnitude.

[0036] In this embodiment, for each knowledge topic, the cosine of the angle between the semantic vectors of all pairwise knowledge data under it is calculated as the knowledge similarity, forming a symmetrical knowledge similarity matrix. The diagonal elements of the knowledge similarity matrix are all 1s, representing the semantic consistency of the same knowledge data. Furthermore, based on the knowledge similarity matrix, the overall similarity distribution of knowledge data under that knowledge topic can be obtained, identifying semantic outliers (knowledge data) that deviate from the main group, serving as a warning of potential conflicts or abnormal data.

[0037] Step S122: Based on knowledge similarity, obtain multiple knowledge data on the same knowledge topic but with knowledge conflicts.

[0038] In this embodiment, by analyzing the knowledge similarity matrix corresponding to each knowledge topic, multiple knowledge data points with knowledge conflicts are obtained for each knowledge topic. Specifically, if two pieces of knowledge data within the same knowledge topic have a knowledge similarity lower than a preset similarity threshold, and the time interval between their publications is less than a preset time threshold, then the two pieces of knowledge data are determined to have potential expression discrepancies or logical conflicts. That is, if there are... Then determine the first i The first knowledge data and the first j There are several pieces of knowledge data that share the same knowledge topic but conflict with each other. Indicates the similarity threshold. Indicates the first i The release time of each piece of knowledge data. Indicates the first j The release time of each piece of knowledge data. This represents the time threshold. Simultaneously, the first... i The first knowledge data and the first j Some knowledge data are marked as potential conflict items. These knowledge data with low similarity and close time may reflect the replacement of old and new knowledge or information contradictions, and the two can be further semantically verified.

[0039] Step S123: Perform multiple credibility assessments on multiple knowledge data to obtain the corresponding multiple knowledge credibility levels.

[0040] In this embodiment, multiple independent credibility assessments are performed on the obtained knowledge data with knowledge conflicts, thereby obtaining multiple knowledge credibility levels for each corresponding knowledge data. Each assessment involves different large models or agents scoring the authenticity, accuracy, and source reliability of the knowledge data based on their built-in knowledge base and reasoning capabilities. The scoring results are normalized to the [0,1] interval, and each assessment result serves as a knowledge credibility level for that knowledge data.

[0041] Step S124: Obtain the quality inspection results based on multiple knowledge credibility levels.

[0042] It should be noted that multiple credibility levels for each knowledge data point can constitute a credibility sequence. This credibility sequence reflects the level of consistency in perception among different evaluation sources regarding the same knowledge data. The smaller the variance, the more similar the models' judgments of its authenticity; the higher the mean credibility, the greater the likelihood that the knowledge data is widely accepted. Therefore, the stability and consistency of knowledge data can be measured by statistically analyzing the mean and variance of the credibility sequence.

[0043] In this embodiment, a multi-model voting method is used to fuse multiple knowledge credibility scores to generate a comprehensive knowledge credibility value. Among them, there are a total of Several large models participate in the evaluation, and each large model outputs a normalized knowledge credibility score. The corresponding comprehensive knowledge credibility score is the arithmetic mean of the credibility scores of each knowledge model, i.e.: ; in, Indicates the first The large model for the first i The credibility of each knowledge data point is evaluated to obtain the corresponding knowledge credibility.

[0044] It should be noted that a higher overall knowledge credibility score indicates greater reliability and consistency of the knowledge data under multi-source verification, making it suitable for use as high-quality knowledge in subsequent application processes. Conversely, a lower score suggests significant controversy or uncertainty, requiring the intervention of domain experts for judgment. This mechanism automates and intelligently filters knowledge, improving the accuracy and robustness of the knowledge base update process.

[0045] In this embodiment, it is also necessary to calculate the distribution variance of multiple knowledge credibility to obtain the consistency score variance.

[0046] Among them, using the first i The credibility sequence of knowledge data is calculated, and its variance is used to obtain the consistency score variance, which measures the impact of different evaluation models on the credibility of the first knowledge data. iThe smaller the variance in the credibility judgment of knowledge data, the more consistent the credibility assessment of the knowledge data by each model is, and the more stable the result is; conversely, it reflects that there is a large discrepancy in the judgment between models, and there may be implicit semantic ambiguity or insufficient evidence.

[0047] Furthermore, quality inspection results are obtained based on the comprehensive knowledge credibility value and the consistency score variance. The quality inspection results are jointly determined by the comprehensive knowledge credibility value and the consistency score variance, which constitute a two-dimensional criterion: knowledge data with a high comprehensive knowledge credibility value and a low consistency score variance is judged to be of high quality and can be directly adopted; knowledge data with a high comprehensive knowledge credibility value but a high consistency score variance, although generally having high acceptance, has discrepancies between models and needs to be marked with a warning and supplemented by manual review; knowledge data with a low comprehensive knowledge credibility value is considered unreliable regardless of the consistency score variance and is removed or temporarily suspended from entering the database.

[0048] Specifically, for knowledge data with a high overall knowledge credibility value, when the variance of its corresponding consistency score exceeds the variance threshold, the knowledge data is automatically marked as a "knowledge conflict" sample.

[0049] For example, in a medical knowledge base, a newly entered statement that "Drug A can treat disease X" has a comprehensive knowledge credibility score of 0.88 after evaluation by five major models. However, its consistency score variance reaches 0.15, exceeding the set variance threshold of 0.12. It is immediately marked as a "knowledge conflict" sample and a review process is triggered. Expert review found that this conclusion is based solely on a single clinical trial, posing a risk of insufficient evidence. Ultimately, it was decided to postpone its entry into the database until supplementary multi-center research data is available for verification. This mechanism significantly reduces the risk of misjudgment due to model bias or data deviation, enhancing the prudence and scientific rigor of knowledge entry. In highly sensitive fields such as finance and law, similar strategies can effectively identify potentially controversial clauses or unconventional judgments, avoiding cascading errors caused by overconfidence in a single model. By continuously monitoring the dynamics of comprehensive credibility and variance, it is also possible to capture abnormal fluctuations in the knowledge evolution process, supporting periodic re-evaluation and version tracking of existing knowledge, thereby constructing a dynamic knowledge system with self-correcting capabilities.

[0050] In this embodiment, the comprehensive knowledge credibility value is combined with the consistency score variance. This quality detection mechanism effectively balances the efficiency and security of automated evaluation, ensuring that the knowledge database maintains high accuracy and strong consistency during dynamic updates. At the same time, it provides a reliable guarantee for subsequent knowledge reasoning and application, and the quality detection results further support the hierarchical management and differentiated application of knowledge data.

[0051] In some other embodiments, the overall knowledge credibility value can also be calculated by a weighted average, with the weights determined by the historical accuracy of each evaluation model, to ensure that the high-credibility model has a greater impact on the results.

[0052] Step S2: Based on the quality inspection results, identify outdated or conflicting knowledge data from the knowledge database as the target knowledge data that needs to be updated.

[0053] It should be noted that the target knowledge data includes knowledge data marked as "knowledge outdated" and / or "knowledge conflicting". The former needs to be replaced or deleted due to decreased timeliness or being refuted by new evidence, while the latter needs further verification due to significant discrepancies in model judgments.

[0054] In this embodiment, based on the quality inspection results, it is determined whether the knowledge data is outdated or conflicting. If the knowledge data is outdated or conflicting, it is identified as target knowledge data for subsequent update processes. If the knowledge data is not outdated or conflicting, its state in the knowledge database is maintained, and no update process needs to be triggered. Specifically, by periodically scanning the quality inspection markers in the knowledge database, the aforementioned two types of target knowledge data are automatically selected, triggering subsequent audit and update processes. This ensures that the knowledge base content remains accurate, reliable, and consistent, while reducing manual screening costs and improving the intelligence level of knowledge maintenance.

[0055] In this application, for knowledge data identified as target knowledge data, an update task queue can be automatically generated and sorted according to knowledge category, scope of influence, and priority to ensure that key knowledge is processed first. The above mechanism supports activation at a set period or in an event-driven manner, ensuring timely response to new data injection, external knowledge source updates, or system feedback anomalies.

[0056] Step S3: Obtain the benchmark knowledge data corresponding to the target knowledge data from the external data source, and calculate the knowledge difference between the benchmark knowledge data and the target knowledge data.

[0057] It should be noted that benchmark knowledge data refers to baseline information obtained through authoritative external knowledge sources, real-time data acquisition channels, or user feedback mechanisms, which may be used to replace or supplement target knowledge data in the knowledge database. Knowledge difference is an indicator used to quantify the degree of semantic content deviation between benchmark knowledge data and target knowledge data. It can be calculated through semantic alignment and vector similarity to quantify the degree of difference between the target knowledge data and benchmark knowledge data.

[0058] In this embodiment, after determining that target knowledge data that needs to be updated exists in the knowledge database, an update query is initiated to an external data source (such as the latest regulatory website) to obtain the latest published benchmark knowledge data. If benchmark knowledge data is not obtained, the original record is maintained and marked as "pending review," pending re-inspection or manual intervention in subsequent cycles. If benchmark knowledge data is obtained, it is semantically aligned with the target knowledge data. Key semantic vectors are extracted using a pre-trained language model, and the cosine similarity between the two in the semantic space is calculated. The degree of difference is comprehensively evaluated by combining indicators such as keyword changes and factual attribute shifts to obtain the knowledge difference degree.

[0059] In this application, the calculation result of knowledge difference will serve as the key basis for whether to perform subsequent update operations. When the knowledge difference exceeds the preset difference threshold, it is determined that the target knowledge data needs to be replaced or corrected, and the update process is automatically triggered. If the knowledge difference is lower than the difference threshold, it is regarded as information redundancy or minor fluctuation, and the original knowledge data can be retained and the comparison log is recorded.

[0060] For example, in a medical knowledge base, when the target knowledge data is "the recommended dosage of a certain drug is 5 mg twice daily," while the baseline knowledge data shows "the recommended dosage has been adjusted to 10 mg once daily," semantic alignment identifies the dual change in medication frequency and dosage. Combined with keyword differences and factual attribute offsets, a weighted calculation is performed to determine a high knowledge difference score, classifying the update as a substantial change and requiring inclusion in the priority update queue. If the knowledge difference score is below a preset difference threshold, it is considered as having no significant change, the original data is retained, and a comparison log is recorded. The entire process requires no manual intervention; the server automatically completes the difference identification and decision-making, significantly improving the timeliness and accuracy of knowledge base updates.

[0061] Step S4: Based on the correspondence between knowledge difference and data update strategy, determine the target data update strategy corresponding to the target knowledge data, and update the target knowledge data according to the target data update strategy.

[0062] It should be noted that the data update strategy includes a first data update strategy, a second data update strategy, and a third data update strategy. The first data update strategy includes: correcting outdated values ​​or erroneous fields in the target knowledge data based on the benchmark knowledge data, while preserving the original semantic structure. The second data update strategy includes: generating revised knowledge data through a large model based on the benchmark knowledge data, and replacing the corresponding parts of the target knowledge data with the revised parts of the revised knowledge data. The third data update strategy includes: completely replacing the target knowledge data with the acquired benchmark knowledge data. The difference threshold is dynamically set according to the sensitivity of knowledge updates, typically ranging from 0.7 to 0.9, with lower values ​​indicating lower tolerance for knowledge changes.

[0063] In this embodiment, the difference threshold is set to 0.7. If the knowledge difference exceeds the difference threshold of 0.7, it is determined that an update operation needs to be performed; otherwise, the current knowledge data is kept unchanged, and the comparison result is recorded for subsequent traceability.

[0064] Furthermore, when the knowledge difference exceeds the difference threshold of 0.7, the corresponding update method is selected based on the degree of knowledge difference and the preset data update strategy. Specifically, if the knowledge difference is small and concentrated in local fields, the first data update strategy is executed; if the semantic structure of the knowledge data changes significantly but does not overturn the original logic, the second data update strategy is initiated, and the large model generates a semantically coherent revised version of the knowledge data, using the revised part of the revised knowledge data to replace the core paragraph of the original target knowledge data; if the knowledge difference is extremely high or involves changes in key facts, the third data update strategy is triggered, completely replacing the target knowledge data with the benchmark knowledge data, and simultaneously updating the associated indexes and reference relationships to ensure the consistency and timeliness of the knowledge database.

[0065] Furthermore, during update operations, a version chain and digital signature are generated to ensure that each change is traceable and tamper-proof. The digital signature uses asymmetric encryption technology, binding the operation timestamp and the executing entity information to enhance the credibility and auditability of knowledge updates. The version chain records the input source, difference analysis results, strategies used, and snapshots of previous and subsequent versions for each update, forming a complete evolution path. This supports on-demand rollback or comparative analysis, further improving the standardization and security of knowledge base management. In addition, the execution process of data update strategies supports a rollback mechanism, which can automatically restore to the most recent stable version when anomaly detection or verification fails, ensuring the continuity and reliability of knowledge services. Rollback operations are also recorded in the version chain and generate an independent digital signature to prevent malicious tampering or the spread of misoperations. The generation of a version chain and digital signature for each knowledge revision ensures that the update process is fully traceable, auditable, and verifiable, effectively preventing data pollution and illegal tampering risks. This meets the stringent requirements of highly sensitive industries such as finance and insurance regulation for the authenticity and integrity of knowledge, ensuring the maintenance of the authority and compliance of the knowledge system in a dynamically evolving environment.

[0066] In other embodiments, multi-level difference thresholds can be set to dynamically match corresponding data update strategies based on the specific numerical range of knowledge difference. For example, 0.75 can be set as the boundary between light and moderate updates, and 0.85 as the boundary between moderate and heavy updates, thereby achieving more refined update decision control. When the knowledge difference is between 0.75 and 0.85, a second data update strategy is initiated to perform structural optimization and partial rewriting of the knowledge data. When the difference exceeds 0.85, a heavy update process is executed to ensure that changes to key information are synchronized in a timely and complete manner. The entire process follows the principle of "minimum necessary updates," effectively improving the accuracy and responsiveness of knowledge database maintenance while reducing the risk of system disturbances caused by frequent fine-tuning.

[0067] For example, in a financial institution's knowledge base, an entry regarding foreign exchange management regulations changed due to policy adjustments. When comparing the old and new knowledge, the knowledge difference score was measured at 0.85, exceeding the set threshold of 0.7, thus triggering a second data update strategy. The corresponding large model generates a semantically coherent revised text based on the latest regulatory documents, replacing the core content of the original entry, and automatically updating related compliance guidelines and risk warnings. This change is fully recorded in the version chain, accompanied by a digital signature, ensuring traceability and tamper-proofing, and guaranteeing the continued compliance and trustworthiness of knowledge services in a heavily regulated environment. Simultaneously, the updated knowledge entry undergoes end-to-end verification to ensure its compatibility and consistency with the existing knowledge network.

[0068] By combining natural language reasoning and graph association analysis, this mechanism automatically identifies potential conflicts or redundant information and pushes it to auditors for review. In practical use, this mechanism effectively reduces decision-making risks caused by knowledge conflicts, significantly improves the efficiency and accuracy of compliance knowledge iteration, reduces the average response time by 60%, and lowers the error rollback rate to below 0.3%, providing a replicable technical path for building a highly reliable and self-evolving knowledge infrastructure. With the continuous acceleration of knowledge update frequency and the constant upgrading of regulatory requirements, the above mechanism can further integrate real-time public opinion monitoring and multi-source heterogeneous data fusion capabilities to achieve more granular knowledge evolution perception and adaptive update decisions. By introducing a dynamic confidence assessment model, it can automatically identify the time-degradation trend of knowledge data, provide early warnings of potential obsolescence risks, and initiate pre-update processes, thereby transforming passive response into proactive intervention. In financial scenarios, this proactive maintenance significantly reduces the operational risks caused by compliance lag and enhances institutions' ability to quickly adapt to external policy changes. Simultaneously, combined with a federated learning framework, cross-organizational knowledge updates can be collaboratively promoted while ensuring data privacy, driving the formation of an industry-level knowledge consensus network. This evolutionary direction not only improves the intelligence level of individual systems, but also lays the foundation for building a trustworthy, controllable, and explainable knowledge ecosystem.

[0069] In some other embodiments, the difference threshold can be dynamically adjusted according to the knowledge type to ensure that updates in sensitive domains are more sensitive, while updates in general domains remain stable.

[0070] Furthermore, in the knowledge database proposed in this application, after each retrieval call, the access count of the referenced knowledge data is updated, and its priority in subsequent management is adjusted to achieve feedback optimization of "high-frequency and useful knowledge being presented first". At the same time, the decay coefficient of the corresponding knowledge data is dynamically changed, and the above-mentioned timeliness detection, auditing and updating steps are automatically executed.

[0071] Furthermore, reinforcement learning algorithms (RLHF / RLAIF) can be used to dynamically adjust the weights of knowledge data in the knowledge database based on the correctness of the retrieval results. The core logic is to use user feedback and system verification results as reward signals to continuously optimize the confidence score and retrieval ranking strategy of knowledge data. For high-value knowledge data that is repeatedly cited and correctly verified, its weight is automatically increased and it is included in the core knowledge set. For knowledge data that has not been cited for a long time or frequently leads to erroneous reasoning, a review mechanism is triggered to assess its necessity for retention. This mechanism enables the knowledge database to dynamically self-evolve, ensuring that the content is not only accurate and reliable but also highly adaptable to the changing needs of real-world application scenarios.

[0072] In summary, the above method, through time decay and consensus voting mechanisms, enables the knowledge database to automatically eliminate outdated knowledge and retain high-confidence content. Combined with dynamic weight adjustment and multi-level difference threshold control, it achieves refined and intelligent knowledge updates. While ensuring stability, it also possesses strong adaptability, continuously optimizing its internal knowledge structure based on actual usage feedback, thereby constructing a closed-loop knowledge base management system with self-learning, self-verification, and self-optimization capabilities. Faced with a rapidly changing information environment, this system can flexibly adjust the update granularity and response rhythm according to actual business scenarios, balancing efficiency and reliability to ensure that knowledge services are always in optimal condition, achieving continuous knowledge evolution and traceable compliance auditing. Therefore, this method can be widely applied to high-compliance scenarios such as insurance clause update monitoring, regulatory compliance checks, and internal knowledge document auditing, significantly improving the level of intelligent knowledge base management in enterprises.

[0073] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0074] In one embodiment, a knowledge base management device is provided, which corresponds one-to-one with the knowledge base management method described in the above embodiments. For example... Figure 5As shown, the knowledge base management device includes a quality inspection module 101, a target determination module 102, a knowledge audit module 103, and a knowledge update module 104. Detailed descriptions of each functional module are as follows: The quality inspection module 101 is used to perform quality inspection on the knowledge data in the knowledge database and obtain the quality inspection results.

[0075] The target determination module 102 is used to determine, based on the quality inspection results, the knowledge data that is outdated or conflicting from the knowledge database as the target knowledge data that needs to be updated.

[0076] The knowledge audit module 103 is used to obtain benchmark knowledge data corresponding to the target knowledge data from an external data source and to calculate the knowledge difference between the benchmark knowledge data and the target knowledge data.

[0077] The knowledge update module 104 is used to determine the target data update strategy corresponding to the target knowledge data based on the correspondence between knowledge difference degree and data update strategy, and update the target knowledge data according to the target data update strategy.

[0078] Specific limitations regarding the knowledge base management device can be found in the limitations of the knowledge base management method described above, and will not be repeated here. Each module in the aforementioned knowledge base management device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0079] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores the aforementioned knowledge data and version update records, supporting the storage and retrieval of both structured and unstructured data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a knowledge base management method.

[0080] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the knowledge base management method described in the above embodiment, for example... Figure 2 S1-S4, as shown, will not be described again here to avoid repetition. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the knowledge base management device, for example... Figure 5 The functions of the quality inspection module 101, target determination module 102, knowledge audit module 103, and knowledge update module 104 shown are not described again here to avoid duplication.

[0081] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the knowledge base management method described in the above embodiment, for example... Figure 2 S1-S4, as shown, will not be described again here to avoid repetition. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the knowledge base management device, for example... Figure 5 The functions of the quality inspection module 101, target determination module 102, knowledge audit module 103, and knowledge update module 104 shown are not described again here to avoid repetition. The computer-readable storage medium can be non-volatile or volatile.

[0082] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0083] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0084] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A knowledge base management method, characterized in that, Including the following steps: Perform quality checks on the knowledge data in the knowledge database and obtain the quality check results; Based on the quality inspection results, the knowledge data that is outdated or conflicting is identified from the knowledge database as the target knowledge data that needs to be updated. Obtain benchmark knowledge data corresponding to the target knowledge data from an external data source, and calculate the knowledge difference between the benchmark knowledge data and the target knowledge data; Based on the correspondence between the knowledge difference degree and the data update strategy, a target data update strategy corresponding to the target knowledge data is determined, and the target knowledge data is updated according to the target data update strategy.

2. The knowledge base management method according to claim 1, characterized in that, The step of updating the target knowledge data according to the target data update strategy includes: If the target data update strategy is determined to be the first data update strategy, then the outdated values ​​or erroneous fields in the target knowledge data are corrected based on the benchmark knowledge data. If the target data update strategy is determined to be the second data update strategy, then a revised knowledge data is generated based on the baseline knowledge data, and the revised part is used to replace the part of the target knowledge data. If the target data update strategy is determined to be the third data update strategy, then the target knowledge data is directly replaced with the baseline knowledge data.

3. The knowledge base management method according to claim 1, characterized in that, The process of performing quality checks on the knowledge data in the knowledge database and obtaining the quality check results includes: The credibility of the knowledge data is evaluated to obtain the knowledge credibility of the knowledge data; The timeliness of the knowledge data is evaluated using the knowledge credibility, and a timeliness score of the knowledge data is obtained. The quality inspection result is obtained based on the timeliness score.

4. The knowledge base management method according to claim 3, characterized in that, The credibility of the knowledge data is evaluated to obtain the knowledge credibility of the knowledge data, including: Based on the data source of the knowledge data, determine the source score of the knowledge data; Based on the semantic consistency between the knowledge data and the preset knowledge set, obtain the consistency score of the knowledge data; The knowledge data is verified, and a verification score is obtained for the knowledge data. The knowledge credibility is calculated based on the source score, the consistency score, and the verification score.

5. The knowledge base management method according to claim 3, characterized in that, The timeliness of the knowledge data is evaluated using the knowledge credibility to obtain a timeliness score for the knowledge data, including: Obtain the time decay coefficient of the knowledge data, and the historical update time of the last update of the knowledge data; The timeliness score is calculated based on the knowledge credibility, the time decay coefficient, and the historical update time.

6. The knowledge base management method according to claim 1, characterized in that, The knowledge database includes multiple pieces of knowledge data, each piece of knowledge data corresponding to a knowledge topic. The step of performing quality checks on the knowledge database and obtaining quality check results includes: The similarity of multiple knowledge data for each knowledge topic is evaluated to obtain the knowledge similarity of multiple knowledge data for each knowledge topic. Based on the knowledge similarity, obtain multiple knowledge data that share the same knowledge topic but conflict with each other; Multiple credibility assessments are performed on the various knowledge data to obtain corresponding multiple knowledge credibility levels; The quality inspection results are obtained based on the credibility of the multiple knowledge bases.

7. The knowledge base management method according to claim 6, characterized in that, The process of obtaining the quality inspection result based on the multiple knowledge credibility levels includes: By integrating the multiple knowledge credibility scores, a comprehensive knowledge credibility value is generated. Calculate the variance of the distribution of the credibility of the multiple knowledge points to obtain the variance of the consistency score; The quality inspection result is obtained based on the comprehensive knowledge credibility value and the consistency score variance.

8. A knowledge base management device, characterized in that, include: The quality inspection module is used to inspect the knowledge data in the knowledge database and obtain the quality inspection results. The target determination module is used to determine, based on the quality inspection results, the knowledge data that is outdated or conflicting from the knowledge database as the target knowledge data that needs to be updated. The knowledge audit module is used to obtain benchmark knowledge data corresponding to the target knowledge data from external data sources and calculate the knowledge difference between the benchmark knowledge data and the target knowledge data. The knowledge update module is used to determine the target data update strategy corresponding to the target knowledge data based on the correspondence between the knowledge difference degree and the data update strategy, and to update the target knowledge data according to the target data update strategy.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the knowledge base management method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the knowledge base management method as described in any one of claims 1 to 7.