Multi-source data-oriented quality management knowledge base incremental updating method and system
By employing a three-level filtering system and explicit metadata processing, the problem of untimely incremental updates to the knowledge base during the screening and grading of multi-source data was solved, enabling efficient and accurate updates to the quality management knowledge base and ensuring the timeliness and accuracy of data processing.
Patent Information
- Application Number
- CN202511063114.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-07-31
AI Technical Summary
In existing technologies, the screening and classification of multi-source data requires multiple steps such as format unification, cleaning, standardization, and benchmark dimension evaluation, resulting in a lengthy processing chain, complex confidence calculation, and untimely incremental updates of the knowledge base, which affects the timeliness of quality management.
A three-level filtering system is adopted, which realizes progressive quality control of data through explicit metadata processing, scanning verification and incremental update analysis. It quantifies the accuracy and timeliness of explicit metadata processing, dynamically compares knowledge base content, automatically identifies and separates conflicting information, and ensures that data that meets the requirements is directly used for knowledge base updates.
It enables efficient value transformation of multi-source data, ensures that the knowledge base presents the latest quality status in real time, improves the timeliness of quality decisions and the accuracy of data processing, reduces interference from low-quality data, and optimizes the knowledge base update process.
Smart Images

Figure CN120874997A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a method and system for incremental updating of a quality management knowledge base for multi-source data. Background Technology
[0002] In the digital transformation of enterprise quality management, with the popularization of industrial internet, Internet of Things (IoT) and intelligent inspection technologies, the amount of multi-source quality data accumulated by enterprises (such as production equipment parameters, quality inspection records, customer complaints, process logs, etc.) is growing exponentially. In existing technologies, multiple screening steps are used to classify the acquired quality management data. First, multi-source data is collected and integrated through a unified data format. Then, multi-source data is preprocessed and feature-engineered, such as data cleaning, data standardization, and data feature extraction. Next, the acquired enterprise quality data is evaluated in three dimensions, including real-time evaluation, accuracy evaluation, and confidence evaluation. Knowledge base construction, knowledge fusion, and conflict resolution are carried out based on existing rules and associations based on full data. Among these, the confidence evaluation of the data is an important prerequisite for conflict resolution. Finally, knowledge verification and storage applications of the knowledge base are carried out in combination with the confidence evaluation of the data.
[0003] For example, Chinese invention patent CN119884102B discloses a quality management method based on standard data, including: constructing a dynamic standard database, collecting production data in real time, handling data missingness and noise by adopting a hybrid anomaly detection and dynamic interpolation strategy, generating a standardized matching code based on hash and polynomial ring operation to ensure data consistency, using a dynamic threshold model and fuzzy logic to determine quality deviations, and combining weighted Euclidean distance to quantify deviations for graded early warning and response measures, and a closed-loop feedback mechanism to continuously update model parameters and knowledge base.
[0004] For example, Chinese invention patent with announcement number CN108717426B discloses a method, apparatus, computer equipment, and storage medium for updating enterprise data, including: obtaining incremental data identifiers of enterprise data, including the enterprise identifier and data dimension of the incremental enterprise data, sending them to a data update server according to the enterprise identifier and data dimension, and updating the enterprise data in a document-type database, wherein the relational database server sends them to the data update server according to the enterprise identifier and data dimension.
[0005] The above-mentioned technology has at least the following technical problems: In existing technologies, multi-source data is acquired from multiple detection devices, resulting in large volumes of data in various formats. The data screening and grading process requires multiple steps, including format unification, cleaning, standardization, and benchmark dimension evaluation (accuracy evaluation, real-time evaluation, and integrity evaluation), leading to a lengthy processing chain. This makes the confidence calculation in the benchmark dimension evaluation process complex, and the lack of parallel processing mechanisms in each evaluation stage further prolongs the processing cycle due to the serial execution logic. In addition, when the acquired multi-source data conflicts with the existing knowledge base, the resolution stage relies on the delayed confidence evaluation results, causing data that meets the grading criteria to not be prioritized in a timely manner. These delays make it difficult for the effective data after screening and grading to trigger incremental updates to the knowledge base in a timely manner, causing newly added quality rules and relationships to remain in the processing queue. Ultimately, this results in the knowledge base not reflecting the latest quality status in real time, affecting the timeliness of knowledge base-based quality decisions. There is a problem of untimely incremental updates to the knowledge base corresponding to quality management data during the screening and grading process. Summary of the Invention
[0006] To address the technical problem of untimely incremental updates to the knowledge base corresponding to quality management data during the screening and grading process in existing technologies, this invention provides a method and system for incremental updates of a quality management knowledge base for multi-source data. The technical solution is as follows: On the one hand, a method for incremental updating of a quality management knowledge base oriented towards multi-source data is provided, including the following steps: A1, the quality management platform receives acquired multi-source quality management data from enterprises and performs an accuracy assessment of the explicit metadata processing process. Simultaneously, it performs a first data screening and grading based on the proportion of processing time for the acquired explicit metadata. The accuracy assessment quantifies the precision of explicit metadata processing for enterprise multi-source quality management data, and the first data screening and grading is used to initially identify data from multi-source quality management data that meets the basic metadata quality requirements; A2, the qualified quality management data after the first data screening and grading undergoes a scan verification integrity assessment, and the data is updated based on the scan verification integrity assessment... The results are used for a second data screening and grading. The scan verification integrity assessment is used to quantify the metadata integrity of the qualified data after the first screening. The second data screening and grading is used to further screen out data with complete metadata that meets the basic format requirements for knowledge base updates, reducing knowledge base update deviations caused by missing data. A3, the qualified quality management data after the second data screening and grading is uploaded to the knowledge base for a third data screening and grading. At the same time, the results of the third data screening and grading are input into the knowledge base for incremental updates, and incremental update effectiveness analysis is performed. Incremental update effectiveness analysis is used to quantify the similarity matching degree between the qualified data after the second screening and the existing data in the knowledge base.
[0007] On the other hand, a quality management knowledge base incremental update system for multi-source data is provided, including: a quality management data identification module, a quality management data scanning and verification module, and an incremental update effectiveness analysis module. The quality management data identification module receives acquired multi-source quality management data from enterprises and assesses the accuracy of the explicit metadata processing process. It also performs a first-stage data screening and grading based on the proportion of processing time for the acquired explicit metadata. The quality management data scanning and verification module scans and verifies the completeness of the qualified quality management data after the first screening and grading, and performs a second-stage data screening and grading based on the results. The incremental update effectiveness analysis module uploads the qualified quality management data after the second screening and grading to the knowledge base for a third-stage data screening and grading. It also inputs the results of the third-stage screening and grading into the knowledge base for incremental updates and performs incremental update effectiveness analysis.
[0008] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. By constructing a three-tiered filtering system of data identification, scanning verification, and update analysis, progressive quality control is achieved in the stages of metadata quality screening, integrity verification, and knowledge base matching, thereby realizing the efficient value transformation of multi-source data. In the data identification stage, the accuracy and timeliness of explicit metadata processing are quantified to filter raw data that meets the basic quality threshold in real time. In the scanning verification stage, a dynamic repair mechanism for metadata defects is established based on the integrity assessment results. When analyzing new data, the system automatically compares the duplication of existing content in the knowledge base, intelligently identifies and separates conflicting information, and ensures that new data that meets the requirements is directly used to improve the content of the knowledge base.
[0009] 2. By accurately quantifying the proportion of explicit metadata processing time and integrating a time compression factor, the interference of abnormally short processing times is suppressed, improving the reliability of accuracy assessment. The actual processing time is dynamically compared with preset standards, automatically filtering distorted data caused by unidentified explicit labels, eliminating assessment bias. The time compression factor intelligently fits the normal fluctuation range, strengthening the ability to resist interference from outliers. This mechanism provides a high-precision judgment basis for the initial data screening, ensuring accurate differentiation of data that complies with metadata processing regulations and meets the basic requirements for incremental updates to the knowledge base. It intercepts low-quality data at the source, improving the purity and usability of the initial data pool, ensuring that effective data triggers updates in a timely manner, allowing the knowledge base to present the latest quality status in real time, and guaranteeing the timeliness of quality decisions.
[0010] 3. By integrating cosine similarity and the reliability of data sources as supplementary factors, a high-precision confidence score is generated to quantify the comprehensive similarity between the data to be updated and the knowledge base content, avoiding misleading update decisions by a single indicator; by automatically prioritizing high-confidence data for updates, key information is ensured to be entered into the database first, while suppressing interference from low-quality data, strengthening the integrated judgment of the reliability of data sources, improving the ability to identify conflicting data, and thus effectively preventing useless or potentially problematic data from entering the knowledge base, prioritizing the most critical information during the update process. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating an incremental update method for a quality management knowledge base oriented towards multi-source data, provided in an embodiment of the present invention; Figure 2 A flowchart for identifying quality data provided in this embodiment of the invention; Figure 3 The flowchart corresponding to the quality management data scanning verification provided in the embodiments of the present invention; Figure 4 This is a flowchart corresponding to the incremental update validity analysis and determination provided in the embodiments of the present invention; Figure 5 This is a schematic diagram of the structure of an incremental update system for a quality management knowledge base oriented towards multi-source data, provided in an embodiment of this application. Detailed Implementation
[0013] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0014] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0015] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0016] This invention provides a method for incremental updating of a quality management knowledge base oriented towards multi-source data, such as... Figure 1 The flowchart shown illustrates a method for incremental updates of a quality management knowledge base oriented towards multi-source data. The process can include the following steps: A1, the quality management platform receives acquired multi-source quality management data from the enterprise and performs an accuracy assessment of the explicit metadata processing. Simultaneously, based on the proportion of explicit metadata processing time obtained during the accuracy assessment, a first data screening and grading is conducted. The accuracy assessment quantifies the precision of explicit metadata processing for the enterprise's multi-source quality management data, providing a basis for determining whether the data qualifies for subsequent incremental updates to the knowledge base. The first data screening and grading initially identifies data from the multi-source quality management data that meets the basic metadata quality requirements, thus creating an initial data pool with basic reliability for the knowledge base incremental update; A2, the qualified quality management data after the first data screening and grading is then processed... A. Perform a scan verification integrity assessment, and based on the results of the scan verification integrity assessment, conduct a second data screening and grading. The scan verification integrity assessment is used to quantify the metadata integrity of qualified data after the first screening. The second data screening and grading is used to further screen out data with complete metadata that meets the basic format requirements for knowledge base updates, eliminate or supplement incomplete data, provide higher quality candidate data for incremental updates of the knowledge base, and reduce knowledge base update deviations caused by missing data. A3. Upload the qualified quality management data after the second data screening and grading to the knowledge base for a third data screening and grading. At the same time, input the results of the third data screening and grading into the knowledge base for incremental updates, and conduct incremental update effectiveness analysis. The incremental update effectiveness analysis is used to quantify the degree of matching between the qualified data after the second screening and the existing data in the knowledge base.
[0017] In this embodiment, a hierarchical and progressive dynamic filtering and multi-dimensional verification system is constructed. In the initial stage, the data filtering foundation is established based on the accuracy assessment of explicit metadata to ensure the basic reliability of the original data pool. In the secondary filtering, field completeness diagnosis and structured repair technology are adopted to form a standardized dataset with complete contextual association, eliminating knowledge architecture bias caused by data incompleteness. Traditional existing methods cannot establish accurate quantitative indicators and use relatively simple missing value removal strategies, which lack repair capabilities and can eliminate knowledge bias caused by data incompleteness.
[0018] In the quality management scenarios of automobile manufacturing enterprises, traditional methods, which simply remove missing values, lead to data chain breaks (such as missing key timestamps or process association information) when updating the knowledge base. This results in errors in process parameter correlation analysis, low accuracy in fault prediction, and a lack of data repair capabilities, highlighting significant knowledge architecture bias issues. By adopting this method, the A1 stage automatically identifies and downgrades abnormal data sources through explicit metadata evaluation, constructing an initial data pool with higher basic reliability. The A2 stage utilizes field completeness diagnosis and structured repair technology to automatically supplement missing related fields (such as process IDs and equipment numbers), generating a standardized dataset with complete contextual logic, completely eliminating knowledge architecture bias caused by data incompleteness. During incremental updates in the A3 stage, the system improves the matching degree between the knowledge base and real-time process parameters based on the repaired complete data chain, making fault warnings more accurate and significantly reducing redundant data storage. Simultaneously, dynamic filtering and multi-dimensional verification ensure the correctness and security of the updated knowledge base content, solving the core problems of low data reliability and difficulty in eliminating knowledge bias in traditional methods.
[0019] Furthermore, the first data screening and classification are performed based on the proportion of explicit metadata processing time obtained. Specifically, the explicit metadata processing time is obtained from the accuracy assessment process, and the preset explicit metadata processing time is obtained from the database. The obtained explicit metadata processing time is used as the numerator, and the preset explicit metadata processing time in the database is used as the denominator. The proportion of explicit metadata processing time is obtained by proportional processing. At the same time, the time compression factor in the database is combined for product processing to suppress the deviation caused by abnormal processing time (such as excessively short time due to failure to identify explicit identifiers).
[0020] Calculating the percentage by using the actual explicit metadata processing time as the numerator and the preset time as the denominator provides a clear picture of processing efficiency. Combining this with a time compression factor optimizes the evaluation. It effectively suppresses biases caused by excessively short processing times due to anomalies such as the failure to identify explicit identifiers, preventing these abnormal data from unduly impacting the evaluation results. This results in a more objective and accurate assessment, more realistically reflecting the actual situation of explicit metadata processing in accuracy evaluation. Furthermore, it provides a more precise quantitative basis for the initial data screening and grading, ensuring the effective differentiation of data from multi-source quality management data that has accurately processed metadata and meets the basic requirements for incremental knowledge base updates.
[0021] In this embodiment, the explicit metadata processing time is obtained through a built-in timer. The preset explicit metadata processing time is represented by the summation and averaging of the historical explicit metadata processing times in the historical identification accuracy assessment process. The time compression factor is preset in the database and is used to reflect the ratio between the obtained explicit metadata processing time and the actual required explicit metadata processing time. It corrects the coefficient of the relationship between the actual time and the obtained time, suppresses the influence of abnormal special cases on the deviation of statistical results, and makes the obtained time ratio more in line with the actual logic. In practical applications, the corresponding explicit metadata processing time can be input to obtain the corresponding time compression factor, providing a quantitative basis for the assessment accuracy and helping to make accurate calculations.
[0022] like Figure 2 The flowchart shown is a process for identifying quality data according to an embodiment of the invention. The first data screening and classification is performed based on the proportion of explicit metadata processing time after fitting. The data corresponding to the proportion of explicit metadata processing time is classified into first-level quality management data, second-level quality management data and third-level quality management data. Corresponding first-level measures, second-level measures and third-level measures are taken. Only the first-level measures are qualified quality management data in the first data screening and classification, and then the second data screening and classification is performed.
[0023] It is further important to understand that the first data screening and grading is based on data shards. The specific process includes: if the percentage of explicit metadata processing time after fitting is less than the minimum percentage of historical explicit metadata processing time in the database, meaning the first data screening and grading is qualified, then the corresponding enterprise multi-source quality management data is recorded as Level 1 quality management data, and Level 1 measures are taken; Level 1 measures mean that the qualified Level 1 quality management data after the first data screening and grading is subjected to a second data screening and grading; if the percentage of explicit metadata processing time after fitting is greater than the maximum percentage of historical explicit metadata processing time in the database, meaning the first data screening and grading is unqualified, then the corresponding enterprise multi-source quality management data is recorded as Level 2 quality management data, and Level 2 measures are taken; Level 2 measures mean that the unqualified Level 2 quality management data after the first data screening and grading is cached and processed later.
[0024] If the percentage of explicit metadata processing time after fitting is within the range of the percentage of explicit metadata processing time in the database, the corresponding enterprise multi-source quality management data will be recorded as Level 3 quality management data, and Level 3 measures will be taken. Specifically, the data compression ratio will be mapped in the database based on the obtained explicit metadata processing time percentage deviation, which will be used to quantify the degree of compression of enterprise multi-source quality management data to reduce data processing time. The explicit metadata processing time percentage deviation represents the difference between the obtained explicit metadata processing time percentage and the preset explicit metadata processing time percentage in the database. The explicit metadata processing time percentage range represents the closed interval corresponding to the maximum and minimum values of the historical explicit metadata processing time percentage in the database. The maximum and minimum values of the historical explicit metadata processing time percentage are represented by the average of the sum of the maximum and minimum values of the historical explicit metadata processing time percentage in the database.
[0025] If the corresponding enterprise multi-source quality management data is not fully classified within the specified explicit metadata identification time, data sharding optimization is performed. Specifically, the data sharding adjustment value is obtained by mapping the acquired explicit metadata identification time deviation in the database. This value is used to quantify the specific ratio that needs to be adjusted for explicit metadata identification processing sharding. It is the specific value for correcting the current data sharding and directly reflects the adjustment range of the current data sharding. The explicit metadata identification time deviation represents the absolute value of the difference between the acquired explicit metadata identification time and the preset explicit metadata identification time in the database. The explicit metadata identification time represents the actual time consumed from the time the quality management platform receives the enterprise multi-source quality management data to the completion of the identification and recognition of explicit metadata (metadata that directly describes the structured information of the data) in the data (i.e., clarifying the type, fields, format, and other structured information of the metadata). The specified explicit metadata identification time is represented by the sum and average of the time consumed in the historical explicit metadata identification process.
[0026] In this embodiment, the consistent hashing algorithm is used to increase the data compression rate and data sharding ratio in a timely manner based on the current data compression rate and data sharding adjustment ratio until the classification of first-level quality management data and all enterprise multi-source quality management data is completed. At the same time, the historical data compression rate and data sharding adjustment ratio are used as sample data and input into the random forest model to train the compression rate-sharding random forest model based on the consistent hashing algorithm. The obtained data compression rate and data sharding ratio are input into the compression rate-sharding random forest model to output the corresponding data compression rate adjustment value and data sharding ratio adjustment value.
[0027] By establishing a three-level classification based on dynamic data characteristics, the priority of data streams is autonomously labeled and allocated in the initial processing stage. First-level data enters the deep verification process through a fast channel to ensure high-timeliness processing. Second-level redundant data is automatically transferred to the buffer to avoid the risk of resource crowding. Third-level regular data achieves a dynamic balance between time consumption and accuracy through an adaptive compression strategy, forming a processing energy efficiency optimization architecture.
[0028] Simultaneously, a real-time sharding feedback adjustment mechanism is introduced to automatically reconstruct the data sharding granularity based on the identified progress deviation, eliminating the processing channel congestion caused by data scale fluctuations, ensuring dynamic adaptation between the explicit metadata identification process and throughput capacity, establishing the relationship between quality data fluctuations and management resource scheduling from the source, improving the value of multi-source data and knowledge transformation efficiency through multi-dimensional quality screening, and providing steady-state evolution support for the quality knowledge base.
[0029] Furthermore, a second data screening and grading process is conducted based on the results of the scan verification integrity assessment. Specifically, the verification metadata processing rate is obtained from the scan verification integrity assessment process, and a preset verification metadata processing rate is simultaneously retrieved from the database and compared. The comparison result is then fitted with a duration integration factor to obtain a verification metadata processing rate score. This is to suppress deviations caused by abnormal verification rates (such as excessively short duration due to the identification of data required for verification), thereby providing high-quality preparatory data for the third data screening and grading of the knowledge base, ensuring that the data meets the complete data required for incremental updates of the knowledge base.
[0030] In this embodiment, the verification metadata processing rate is obtained by timing the time required for the specified amount of data. The time is obtained through a timer. The verification metadata processing rate preset in the database is represented by the sum and average of the historical verification metadata processing rates during the historical scan verification integrity assessment process. The duration integration factor is preset in the database and is used to reflect the ratio between the obtained verification metadata processing rate and the actual verification metadata processing rate. It corrects the coefficient of the relationship between the actual rate and the obtained rate, suppresses the influence of abnormal special cases on the result deviation, and makes the reflected rate score more in line with the actual logic. In practical applications, inputting the verification metadata processing rate can obtain the corresponding duration integration factor, providing a quantitative basis for assessing the scan verification integrity.
[0031] like Figure 3The diagram shows a flowchart of the quality management data scanning verification process according to an embodiment of the present invention. Specifically, a second data screening and classification is performed based on the obtained verification metadata processing rate score. The data is divided into three categories: a first category, a second category, and a third category. The data corresponding to the first category can be further classified as secondary screening data samples, the data corresponding to the second category is data requiring assistance, and the data corresponding to the third category is invalid data. Only the secondary screening data samples can enter the third data screening and classification process.
[0032] It is further important to understand that the second data screening and grading is based on a multi-threaded parallel processing window. The specific process includes: First, when the obtained verification metadata processing rate score is greater than the maximum historical verification metadata processing rate score in the database, the corresponding Level 1 quality management data is recorded as a second-stage screening data sample and directly transmitted to the knowledge base for a third data screening and grading; Second, when the obtained verification metadata processing rate score is within the allowable range of the verification metadata processing rate score in the database, the corresponding Level 1 quality management data is recorded as data requiring assistance and sent to the quality management platform for background supplementation and validity verification; Third, when the verification metadata processing rate score is less than the minimum historical verification metadata processing rate score in the database, the corresponding Level 1 quality management data is recorded as invalid data, and the first data screening and grading is performed again.
[0033] The specific backend supplementation process is as follows: Based on the fields corresponding to the data to be assisted, a prompt is made to call the quality management platform API (Application Programming Interface) to supplement the missing fields; if the supplemented verification metadata processing rate score meets the first category, it is classified as a secondary screening data sample that can be transmitted to the knowledge base, and a third data screening and grading process is performed; if the supplemented and re-acquired verification metadata processing rate score does not meet the first category, the deviation of the re-acquired verification metadata processing rate score is mapped in the database to obtain the corresponding multi-parallel processing window adjustment value, which is used to quantify the current degree of increasing the multi-parallel processing window to reduce the processing time of verification metadata; if, after adjusting the multi-parallel processing window for a period of time, the re-acquired verification metadata processing rate score does not meet the first category, a missing field supplementation warning is issued; otherwise, it is determined that a third data screening and grading process can be performed.
[0034] In this embodiment, the verification metadata processing rate score deviation represents the difference between the obtained verification metadata processing rate and the preset verification metadata processing rate in the database. The MulWinTCP algorithm is used to adjust and increase the processing window in a timely manner based on the obtained current multi-parallel processing window until the obtained verification metadata processing rate score meets the first type of case. At the same time, the historical multi-parallel processing windows are used as sample data and input into the convolutional neural network model. The model is trained based on the MulWinTCP algorithm to obtain the multi-parallel processing-neural network model. The obtained multi-parallel processing window is input into the multi-parallel processing-neural network model and the corresponding multi-parallel processing window adjustment value is output.
[0035] By quickly identifying and diverting high-processing-rate primary data samples directly into the knowledge base, the flow time of high-quality data is shortened. For "data awaiting assistance," the platform API is called to supplement missing fields, and combined with an intelligent verification mechanism, potentially usable data is transformed, reducing the risk of potential errors. The dynamic multi-parallel processing window adjustment mechanism accurately quantifies resource investment based on real-time performance deviations, intelligently improving processing throughput and avoiding blind consumption of resources. This maximizes data processing speed, reduces invalid operations, and reliably improves the quality of data entering the knowledge base, ultimately ensuring the efficient and robust operation of the data processing workflow.
[0036] Further, the incremental update effectiveness analysis specifically involves: calculating the cosine similarity between the normalized secondary screening data samples and the corresponding reference screening data samples stored in the knowledge base. The closer the cosine similarity is to 1, the more similar the two sets of data are; the closer it is to 0, the greater the difference. This is typically used to quantify the similarity between secondary reference data and reference data stored in the knowledge base. If the obtained cosine similarity is greater than the preset cosine similarity in the database, the similarity assessment is deemed unqualified. In this case, the secondary screening data samples are not stored in the knowledge base for incremental updates, and the obtained secondary screening data samples are uploaded to the cache for manual review. If the obtained cosine similarity is not greater than the preset cosine similarity in the database, the similarity assessment is deemed qualified, and conflict judgment is performed.
[0037] Specifically, the conflict judgment process involves the following steps: During the calculation of cosine similarity, the result is used as the base variable and fitted with supplementary confidence scores to obtain a supplementary confidence score. These supplementary confidence scores are then sorted in descending order from 1 to 0. The supplementary confidence score quantifies the similarity between the secondary screening data samples that passed the similarity assessment and the corresponding reference screening data in the knowledge base. The supplementary confidence data represents the data source reliability supplementary factor obtained by inputting the secondary screening data samples that passed the similarity assessment into the database. The descending order is used to visualize the update priority of the secondary screening data samples that passed the confidence assessment in the knowledge base, allowing decision-makers to intuitively compare the update value of different data samples. During the incremental update of the quality management knowledge base, prioritizing the updating of data samples with higher rankings significantly improves the accuracy and timeliness of the knowledge base content while reducing the risk of data conflicts. If the obtained supplementary confidence score is greater than the supplementary confidence score corresponding to the reference screening data stored in the knowledge base, a confidence judgment is performed; otherwise, caching is implemented.
[0038] In this embodiment, the confidence supplementary data is preset in the database and is used to reflect the reliability of the acquired data and assess the confidence of the actual acquired secondary screening data. In practical applications, the corresponding confidence supplementary data can be obtained by inputting the cosine similarity of the corresponding data. Specifically, a mapping relationship between data cosine similarity and confidence supplementary data is pre-established in the database. When the actual secondary screening data is acquired, the cosine similarity of the corresponding data is calculated and used as a query condition. The corresponding confidence supplementary data can be found by searching and matching in the preset mapping relationship in the database, thereby knowing the reliability of the acquired data and providing a basis for subsequent operations such as data reliability-based analysis and evaluation.
[0039] This example uses cosine similarity to quantify the differences between old and new data, filtering out data samples highly consistent with the knowledge base content to avoid duplicate or low-value information from being added to the database. It intercepts data exceeding the similarity threshold, transferring it to caching and manual review to prevent highly similar data samples from appearing in the knowledge base. The introduction of confidence-based supplementary data deeply integrates the data source credibility dimension, enabling conflict judgment to break through the limitations of a single similarity indicator. The descending sorting of confidence-based supplementary scores intelligently prioritizes updates, ensuring that high-value data takes effect first. This reduces the risk of knowledge base content conflicts, enhances the accuracy and consistency of incremental updates, and provides a reliable safety net for key decisions through manual review, ultimately achieving intelligent and risk-controllable dynamic maintenance of the knowledge base.
[0040] It's important to understand that caching involves storing data in a cache to facilitate subsequent operations, such as reverting to the previous process. This caching process has an automatic deletion mechanism to prevent excessive caching from increasing the load and slowing down the system. The specific process of caching includes writing data to the cache management system, performing periodic cleanup tasks, and handling cache access requests. Access records are updated and processing requests are performed simultaneously with the requests. In the periodic cleanup tasks, it is necessary to determine whether the deletion conditions are met based on the set time interval. If they are met, the cached items are deleted.
[0041] like Figure 4 The flowchart shown is the incremental update effectiveness analysis flowchart provided in the embodiment of the present invention. The cosine similarity obtained is used to compare and determine whether to perform conflict judgment or manual review. Conflict judgment is carried out by obtaining confidence score to perform a third data screening and classification. The operation is carried out according to whether the data is in the confidence interval specified in the database. For example, if it is in the first confidence interval, the knowledge base is incrementally updated and the existing knowledge base is covered.
[0042] Further understanding is needed regarding the confidence level determination, which involves the following three steps: S1, when the obtained confidence score falls within the preset first confidence interval in the database, the existing knowledge base is overwritten and incrementally updated, while cross-validation is performed on the incrementally updated knowledge base; S2, when the obtained confidence score falls within the preset second confidence interval in the database, the existing knowledge base is retained and the corresponding secondary filtered data is marked as pending verification, prompting the pre-set personnel to review it; S3, when the obtained confidence score falls within the preset third confidence interval in the database, the existing knowledge base is retained and the corresponding secondary filtered data is marked as pending caching. Here, the first confidence interval is typically set to [0.8, 1], the second confidence interval is typically set to [0.6, 0.8], and the third confidence interval is typically set to [0, 0.6]. In industrial production line scenarios, the pre-set personnel can fine-tune the range of the above-mentioned confidence interval settings based on the actual application situation.
[0043] After an incremental update, cross-validation of the knowledge base is required. The specific steps are as follows: Before the next incremental update of the quality management knowledge base, the baseline database and the incremental database are compared, and data filtering and classification are performed simultaneously based on the baseline database and the incremental database. Data filtering and classification includes first data filtering and classification, second data filtering and classification, and third data filtering and classification. The baseline database represents the knowledge base before the incremental update, and the incremental database represents the knowledge base after the incremental update. If the amount of data corresponding to the secondary filtering data within the first confidence interval in the incremental database is greater than the initial amount of data in the baseline database, the incremental update is deemed valid, and the knowledge base after the incremental update is marked as qualified. Otherwise, an early warning is triggered and automatic updates are suspended to prompt the designated personnel to conduct a check.
[0044] In this embodiment, confidence interval division enables refined update control: high-confidence data (second-stage filtered data samples of the first confidence interval) directly triggers coverage updates and initiates cross-validation; medium-confidence data (second-stage filtered data samples of the second confidence interval) is transferred to the manual review channel; and low-confidence data (second-stage filtered data samples of the third confidence interval) is intelligently cached to avoid invalid operations. Subsequent cross-validation steps rigorously compare the data volume changes between the incremental and baseline databases to accurately determine the validity of the update. If the update is valid, it is marked for inclusion in the database; otherwise, an alert and pause process are triggered. Through visualized priority scheduling, confidence-level hierarchical processing, and closed-loop verification mechanisms, the efficiency of knowledge base updates is maximized while ensuring the reliability and data quality stability of each incremental update, reducing the workload of manual review, and preventing systemic update risks.
[0045] like Figure 5 The diagram shown is a structural schematic of an incremental update system for a quality management knowledge base oriented towards multi-source data, provided in an embodiment of this application. The system may include the following modules: a quality management data identification module, a quality management data scanning and verification module, and an incremental update effectiveness analysis module. The quality management data identification module receives acquired multi-source quality management data from enterprises and assesses the accuracy of the explicit metadata processing process. It also performs a first-stage data screening and grading based on the proportion of processing time for the acquired explicit metadata. The quality management data scanning and verification module performs a scan verification integrity assessment on the qualified quality management data after the first data screening and grading, and performs a second-stage data screening and grading based on the results of the scan verification integrity assessment. The incremental update effectiveness analysis module uploads the qualified quality management data after the second-stage data screening and grading to the knowledge base for a third-stage data screening and grading. It also inputs the results of the third-stage data screening and grading into the knowledge base for incremental updates and performs incremental update effectiveness analysis.
[0046] In this embodiment, a hierarchical, progressive processing mechanism significantly improves the reliability of incremental updates to the quality management knowledge base. The quality management data identification module performs initial screening based on explicit metadata accuracy assessment, ensuring the fundamental quality of the data from the source. The scanning verification module, through integrity assessment and structured repair, forms a standardized dataset with complete contextual relationships, effectively eliminating knowledge architecture biases caused by incomplete data in traditional methods. The incremental update analysis module ensures the accuracy and effectiveness of the updated content by quantifying the matching degree between the data and the knowledge base. This system solves the problems of low reliability and difficulty in eliminating knowledge biases in multi-source data fusion, achieving secure, efficient, and accurate dynamic updates to the knowledge base.
[0047] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device; the computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0048] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0049] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0050] It should be understood that, in various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0051] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0052] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0053] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the explicit or discussed couplings or direct couplings or communication connections between devices or units can be through some interfaces, and indirect couplings or communication connections between devices or units can be electrical, mechanical, or other forms.
[0054] The units described as separate components may or may not be physically separate. Explicit components as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0055] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0056] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0057] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for incremental updating of a quality management knowledge base oriented towards multi-source data, characterized in that, Includes the following steps: A1. The quality management platform receives the acquired multi-source quality management data from the enterprise and performs an accuracy assessment of the explicit metadata processing process. At the same time, it performs the first data screening and classification based on the proportion of the acquired explicit metadata processing time. The accuracy assessment is used to quantify the precision of the explicit metadata processing of the enterprise's multi-source quality management data. The first data screening and classification is used to initially identify data that meets the basic metadata quality requirements from the multi-source quality management data. A2. After the first data screening and grading, qualified quality management data are scanned and verified for completeness assessment. Based on the results of the scanned and verified completeness assessment, a second data screening and grading is performed. The scanned and verified completeness assessment is used to quantify the metadata completeness of qualified data after the first screening. The second data screening and grading is used to further screen out data with complete metadata that meets the basic format requirements for knowledge base updates, thereby reducing knowledge base update deviations caused by missing data. A3. Upload the qualified quality management data after the second data screening and grading to the knowledge base, and perform a third data screening and grading. At the same time, input the results of the third data screening and grading into the knowledge base for incremental updates, and perform incremental update effectiveness analysis. The incremental update effectiveness analysis is used to quantify the similarity matching degree between the qualified data after the second screening and the existing data in the knowledge base.
2. The incremental update method for a quality management knowledge base oriented towards multi-source data as described in claim 1, characterized in that, The first data filtering and classification based on the proportion of processing time for the acquired explicit metadata is as follows: The explicit metadata processing time is obtained from the accuracy assessment process. Simultaneously, the preset explicit metadata processing time is obtained from the database and compared to obtain the proportion of explicit metadata processing time. At the same time, the time compression factor in the database is combined for fitting processing to suppress the deviation caused by abnormal processing time and improve the reliability of explicit metadata processing accuracy assessment. The first data filtering and classification is based on data shards, and the specific process includes: If the percentage of processing time for the explicit metadata after fitting is less than the minimum percentage of processing time for the historical explicit metadata in the database, then the corresponding enterprise multi-source quality management data will be recorded as Level 1 quality management data, and Level 1 measures will be taken. The first-level measure refers to a second-level data screening and grading of qualified first-level quality management data after the first data screening and grading. If the percentage of explicit metadata processing time after fitting is greater than the maximum percentage of historical explicit metadata processing time in the database, then the corresponding enterprise multi-source quality management data will be recorded as secondary quality management data, and secondary measures will be taken. The secondary measures refer to caching and delaying the processing of secondary quality management data that fail the initial data screening and grading.
3. The incremental update method for a quality management knowledge base oriented towards multi-source data as described in claim 2, characterized in that, The first data screening and classification also includes: If the percentage of explicit metadata processing time after fitting is within the percentage of explicit metadata processing time in the database, then the corresponding enterprise multi-source quality management data will be recorded as Level 3 quality management data, and Level 3 measures will be taken. The three-level measures are as follows: based on the deviation of the processing time of the acquired explicit metadata, the corresponding data compression ratio is mapped in the database to quantify the degree of compression of the enterprise's multi-source quality management data, so as to reduce the data processing time. If the corresponding enterprise multi-source quality management data is not fully classified within the specified explicit metadata identification time, data fragmentation optimization will be performed. The data sharding optimization specifically involves mapping the acquired explicit metadata identifier identification time deviation in the database to obtain the corresponding data sharding adjustment value. This value is used to quantify the specific ratio that needs to be adjusted for sharding that performs explicit metadata identifier processing, directly reflecting the adjustment range of the current data sharding.
4. The incremental update method for a quality management knowledge base oriented towards multi-source data as described in claim 1, characterized in that, The second data screening and grading based on the results of the scan verification integrity assessment is as follows: The verification metadata processing rate is obtained during the scanning verification integrity assessment process. Simultaneously, the preset verification metadata processing rate is obtained from the database and compared. The comparison result is fitted with the duration integration factor to obtain the verification metadata processing rate score, so as to suppress the deviation caused by abnormal verification rate and improve the accuracy assessment of verification metadata processing integrity. The second data filtering and grading is performed based on a multi-threaded parallel processing window, and the specific process includes: In the first category, when the obtained verification metadata processing rate score is greater than the maximum historical verification metadata processing rate score in the database, the corresponding first-level quality management data is recorded as a second-level screening data sample and directly transmitted to the knowledge base for a third-level data screening and classification. The second category is when the obtained verification metadata processing rate score is within the allowed range of the verification metadata processing rate score in the database, the corresponding first-level quality management data is recorded as data to be assisted and sent to the quality management platform for background supplementation process and validity verification. The third category is when the verification metadata processing rate score is less than the minimum historical verification metadata processing rate score in the database, the corresponding first-level quality management data is recorded as invalid data, and the first data screening and classification are carried out again.
5. The incremental update method for a quality management knowledge base oriented towards multi-source data as described in claim 4, characterized in that, The background supplementation process is as follows: Based on the fields corresponding to the data to be assisted, the system prompts to call the quality management platform API to supplement the missing fields. If the supplemented verification metadata processing rate score meets the first category, it is classified as a secondary screening data sample that can be transmitted to the knowledge base, and a third data screening and classification is carried out. If the processing rate score of the verification metadata after supplementation does not conform to the first case, the corresponding multi-parallel processing window adjustment value is mapped in the database based on the deviation of the processing rate score of the verification metadata after supplementation. This value is used to quantify the current degree of increasing the multi-parallel processing window in order to improve the processing time of the verification metadata. If the processing rate score of the newly acquired verification metadata does not meet the first category after the multi-parallel processing window has been adjusted for a period of time, a warning will be issued for missing fields; otherwise, it will be determined that a third data filtering and classification can be performed.
6. The incremental update method for a quality management knowledge base oriented towards multi-source data as described in claim 1, characterized in that, The incremental update validity analysis specifically includes: Calculate the cosine similarity between the normalized secondary screening data samples and the corresponding reference screening data samples stored in the knowledge base. If the obtained cosine similarity is greater than the preset cosine similarity in the database, it is determined that the similarity evaluation is unqualified, and the obtained secondary screening data samples are uploaded to the cache for manual review. If the obtained cosine similarity is not greater than the preset cosine similarity in the database, it is determined that the similarity assessment is qualified and a conflict judgment is made.
7. The incremental update method for a quality management knowledge base oriented towards multi-source data as described in claim 6, characterized in that, The conflict determination specifically includes: In the calculation of cosine similarity, the calculation result of cosine similarity is used as the basic variable. By fitting it with the confidence supplementary data, the confidence supplementary score is obtained, and the obtained confidence supplementary scores are sorted in descending order from 1 to 0. The confidence supplement score is used to quantify the degree of similarity between the secondary screening data corresponding to the secondary screening data samples that have passed the similarity assessment and the corresponding reference screening data in the knowledge base. The confidence supplement data refers to inputting the secondary screening data samples that have passed the similarity assessment into the database to obtain the corresponding data source reliability supplement factor; The descending sort is used to visualize the update priority of secondary screening data samples that have passed the confidence assessment in the knowledge base.
8. The incremental update method for a quality management knowledge base oriented towards multi-source data as described in claim 7, characterized in that, The conflict determination also includes: If the obtained confidence score is greater than the confidence score corresponding to the reference filter data stored in the knowledge base, then a confidence judgment is performed; otherwise, caching is performed. The confidence level determination consists of the following three steps: S1, when the obtained confidence score is within the first confidence interval preset in the database, the existing knowledge base is overwritten and incrementally updated, and cross-validation is performed on the incrementally updated knowledge base. S2, when the obtained confidence score is within the second confidence interval preset in the database, the existing knowledge base is retained and the corresponding secondary screening data is marked as pending verification, which is used to prompt the preset personnel to review it; S3. When the obtained confidence score is within the third confidence interval preset in the database, the existing knowledge base is retained and the corresponding secondary filtered data is marked as to be cached.
9. The incremental update method for a quality management knowledge base oriented towards multi-source data as described in claim 8, characterized in that, The incrementally updated knowledge base undergoes cross-validation, specifically through the following steps: Before the next incremental update of the quality management knowledge base, the baseline library and the incremental library are compared, and data filtering and classification are performed simultaneously based on the baseline library and the incremental library. The data filtering and classification includes first data filtering and classification, second data filtering and classification, and third data filtering and classification. The baseline library represents the knowledge base before the incremental update, and the incremental library represents the knowledge base after the incremental update. If the amount of data corresponding to the secondary screening data within the first confidence interval in the incremental database is greater than the initial amount of data in the baseline database, the incremental update is deemed valid, and the updated knowledge base is marked as qualified. Otherwise, an alert is triggered and automatic updates are suspended to prompt designated personnel to conduct an investigation.
10. A system applying the incremental update method for a quality management knowledge base oriented towards multi-source data as described in any one of claims 1-9, comprising: The module includes a quality management data identification module, a quality management data scanning and verification module, and an incremental update validity analysis module. The quality management data identification module is used to receive the acquired multi-source quality management data of the enterprise, and to evaluate the accuracy of the identification of the explicit metadata processing process. At the same time, it performs the first data screening and classification based on the proportion of the acquired explicit metadata processing time. The quality management data scanning and verification module is used to scan and verify the integrity of qualified quality management data after the first data screening and grading, and to perform a second data screening and grading based on the results of the scanning and verification integrity assessment. The incremental update effectiveness analysis module is used to upload qualified quality management data after the second data screening and grading to the knowledge base, perform a third data screening and grading, and input the results of the third data screening and grading into the knowledge base for incremental updates, and perform incremental update effectiveness analysis.
Citation Information
Patent Citations
Enterprise data updating methods, devices, computer equipment and storage media
CN108717426B
A quality management method based on standard data
CN119884102B
Knowledge base data updating method and knowledge base data updating device
CN106776635A
Health medical data management method and system, electronic equipment and storage medium
CN115274122A
Software operation and maintenance quality evaluation method and system based on operation and maintenance quality management database
CN116245406A