Knowledge base quality detection and evaluation method and system
By acquiring real-time data and using a knowledge graph self-evaluation algorithm for unsupervised learning to perform cluster analysis, the learning effect prediction model is optimized, solving the problem of low prediction accuracy in existing technologies and realizing the intelligentization of training resources and the improvement of resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MIDDLE EAST GROUP BUSINESS MANAGEMENT CO LTD
- Filing Date
- 2025-11-24
- Publication Date
- 2026-04-21
AI Technical Summary
Existing methods for predicting and optimizing learning outcomes are ill-suited to the dynamic changes in learning activity data, leading to decreased prediction accuracy and difficulty in optimizing the performance of prediction models.
By acquiring real-time learning activity data and environmental data from the target job training resource system, a pre-trained learning effect prediction model is used to generate predicted learning effect values for a future preset time period. A knowledge graph self-evaluation processing algorithm based on unsupervised learning is used for cluster analysis to calculate the Euclidean distance and coverage of cluster centers, thereby optimizing the learning effect prediction model.
It improves the accuracy of learning outcome prediction and the generalization ability of the model, ensures the reliability of prediction results, and realizes the intelligentization of on-the-job training resource training strategies and the improvement of resource utilization.
Smart Images

Figure CN121901759A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge resource digitization technology, and in particular to a method and system for quality detection and evaluation of knowledge bases. Background Technology
[0002] Software engineering companies typically have various positions, such as software development engineers, test engineers, and project managers. In particular, software engineering companies often collaborate with other companies on projects, such as general contracting outsourcing providers and human resource outsourcing / on-site development providers. For example, general contracting outsourcing providers (project outsourcing) directly contract with internet companies (large internet companies) to undertake the development, testing, and maintenance of a complete project or module. However, the skills training content and difficulty required for each position within the software engineering company and its service providers vary. Learning outcome prediction allows companies to understand in advance the effectiveness of different training resources for various positions. For example, for training in a new programming language, the learning outcomes for junior and experienced developers can be predicted. If the prediction shows that junior developers are not performing well in the training, training resources can be adjusted in advance, such as increasing the proportion of basic courses or arranging more targeted tutoring, and rationally allocating resources such as training instructors and teaching materials to avoid uneven resource distribution leading to poor training outcomes in some areas while other resources are idle.
[0003] In digital job training resources, learning outcome prediction is a key technology for optimizing job training resource scheduling, improving the stability of the job training resource system, and reducing resource waste. With the expansion of intelligent job training resources and the widespread coverage of diverse learning needs, operator node learning activity data exhibits characteristics such as high dimensionality, nonlinearity, and spatiotemporal heterogeneity. Traditional prediction methods struggle to effectively process this complex learning activity data. Therefore, there is an urgent need for a technical solution that can accurately predict future learning outcomes and dynamically optimize job training resource strategies to improve the adaptability, efficiency, and sustainability of the job training resource system.
[0004] Existing learning performance prediction optimization methods typically use only a single clustering algorithm to determine the pattern to which the current learning performance belongs, and then predict the learning performance value in the future time period based on the pattern to which the previous learning performance belonged.
[0005] However, existing methods for predicting and optimizing learning outcomes are ill-suited to the dynamic changes in learning activity data, leading to decreased prediction accuracy and difficulty in optimizing the performance of prediction models. Summary of the Invention
[0006] This invention provides a knowledge base quality detection and evaluation method and system to solve the problem of low accuracy in predicting the learning effect of intelligent job training resources in the prior art.
[0007] Firstly, embodiments of the present invention provide a method for knowledge base quality detection and evaluation.
[0008] The system acquires real-time learning activity data and real-time environment data of each operator node in the target job training resource system, and inputs the real-time learning activity data and the real-time environment data into a pre-trained learning effect prediction model to generate the learning effect prediction value of each operator node in a future preset time period.
[0009] The target algorithm is used to perform cluster analysis on the predicted learning effect values of each operator node to obtain the first cluster analysis result. The target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning. The first cluster analysis result includes the location information of the first cluster center point.
[0010] Based on the location information of the first cluster center point and the location information of the second cluster center point, the coverage of the target algorithm is determined. The second cluster center point is the cluster center point included in the cluster analysis results generated by other algorithms. The other algorithms are evaluation algorithms that are different from the target algorithm.
[0011] The learning effect prediction model is optimized based on the coverage, and the training strategy of the target job training resource system is adjusted based on the output of the optimized learning effect prediction model.
[0012] Optionally, the number of the first cluster center points is multiple, and the number of the second cluster center points is multiple. Determining the coverage of the target algorithm based on the location information of the first and second cluster center points includes:
[0013] Based on the location information of each first cluster center point and each second cluster center point, calculate the Euclidean distance between each first cluster center point and each second cluster center point, and generate a distance matrix. Each row of the distance matrix corresponds to a first cluster center point and each column corresponds to a second cluster center point.
[0014] For each row of the distance matrix, select the second cluster center point with the smallest Euclidean distance to the first cluster center point as the matching second cluster center point, and determine the number of the same operator nodes in the cluster where the first cluster center point is located and the cluster where the matching second cluster center point is located.
[0015] Calculate the coverage of each first cluster center point based on the number of identical operator nodes;
[0016] When there is only one evaluation algorithm, the average coverage of all the center points of the first cluster is determined as the coverage of the target algorithm.
[0017] Optionally, after calculating the coverage of each first cluster center point based on the number of identical operator nodes, the method further includes:
[0018] When there are multiple evaluation algorithms, the average coverage of all first cluster center points is determined as the coverage to be optimized by the target algorithm for the evaluation algorithm.
[0019] The coverage of the target algorithm is the weighted average of the coverage to be optimized for all the evaluated algorithms.
[0020] Optionally, calculating the coverage of each first cluster center point based on the number of identical operator nodes includes:
[0021] For each first cluster center point, the ratio of the number of the corresponding identical operator nodes to the number of operator nodes in the cluster where the first cluster center point is located is determined as the coverage of the first cluster center point.
[0022] Optionally, the training process of the learning effect prediction model includes:
[0023] Obtain historical learning activity data and historical environment data for each operator node in the target job training resource system;
[0024] Based on the historical learning activity data and historical environment data, combined with the course topology data, an initial prediction model is trained using a machine learning algorithm, such as a neural network, support vector machine, or random forest. The course topology data includes knowledge point relationships and job training resource paths.
[0025] The performance of the prediction model during training is evaluated using cross-validation, and the model parameters are adjusted based on the evaluation results until the performance of the prediction model reaches a preset threshold.
[0026] Optionally, adjusting the training strategy of the target job training resource system based on the output of the optimized learning effect prediction model includes:
[0027] Based on the predicted learning performance of each operator node within a preset time period generated by the optimized learning performance prediction model, the learning ability matching status of each operator node is determined.
[0028] Based on the learning ability matching status, the training strategy of the target job training resource system is adjusted. The training strategy includes the adjusted job training resource plan volume, knowledge point allocation path and personalized tutoring strategy.
[0029] Optionally, based on the learning ability matching status, the training strategy of the target job training resource system is adjusted, including:
[0030] Based on the learning ability matching status, determine whether there are operator nodes with weak knowledge or advanced ability;
[0031] For each operator node with weak knowledge, the complexity of the job training resource plan is reduced by adjusting the step size according to the complexity, the learning time cycle of the knowledge point allocation path is extended by adjusting the step size according to the learning cycle, and corresponding personalized tutoring resources are allocated to the operator node.
[0032] For all operator nodes with advanced capabilities, the task density of the job training resource plan is increased by adjusting the step size according to the density, the learning time cycle of the knowledge point allocation path is reduced by adjusting the step size according to the learning cycle, and the operator nodes are given the right to learn the advanced knowledge graph independently.
[0033] Secondly, embodiments of the present invention provide a knowledge base quality detection and evaluation system, comprising:
[0034] The acquisition module is used to acquire real-time learning activity data and real-time environment data of each operator node in the target job training resource system, and input the real-time learning activity data and the real-time environment data into a pre-trained learning effect prediction model to generate the learning effect prediction value of each operator node in a future preset time period.
[0035] The analysis module is used to perform cluster analysis on the predicted learning effect values of each operator node using a target algorithm to obtain a first cluster analysis result. The target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning. The first cluster analysis result includes the location information of the center point of the first cluster.
[0036] The determination module is used to determine the coverage of the target algorithm based on the location information of the first cluster center point and the location information of the second cluster center point. The second cluster center point is the cluster center point included in the cluster analysis results generated by other algorithms, and the other algorithms are evaluation algorithms that are different from the target algorithm.
[0037] The optimization and adjustment module is used to optimize the learning effect prediction model based on the coverage, and adjust the training strategy of the target job training resource system based on the output of the optimized learning effect prediction model.
[0038] Thirdly, embodiments of the present invention provide a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are to be invoked and executed by the processing component to implement a knowledge base quality detection and evaluation method as described in any of the first aspects.
[0039] Fourthly, embodiments of the present invention provide a computer storage medium storing a computer program, wherein when the computer program is executed by a computer, it implements a knowledge base quality detection and evaluation method as described in any of the first aspects.
[0040] This invention provides a knowledge base quality detection and evaluation method, which includes: acquiring real-time learning activity data and real-time environment data of each operator node in a target job training resource system, and inputting the real-time learning activity data and real-time environment data into a pre-trained learning effect prediction model to generate a learning effect prediction value for each operator node within a preset future time period; performing cluster analysis on the learning effect prediction values of each operator node using a target algorithm to obtain a first cluster analysis result, wherein the target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning, and the first cluster analysis result includes the location information of the first cluster center point; determining the coverage of the target algorithm based on the location information of the first cluster center point and the second cluster center point, wherein the second cluster center point is the cluster center point included in the cluster analysis results generated by other algorithms, and the other algorithms are evaluation algorithms different from the target algorithm; optimizing the learning effect prediction model based on the coverage, and adjusting the training strategy of the target job training resource system based on the output result of the optimized learning effect prediction model.
[0041] This invention provides accurate and real-time input data for learning effect prediction by collecting real-time learning activity data and environmental data from operator nodes, ensuring the timeliness and reliability of the learning effect prediction model. A knowledge graph self-evaluation processing algorithm based on unsupervised learning is used to perform cluster analysis on the predicted learning effect values, identifying the spatial distribution patterns and potential patterns of the learning effects. By calculating the coverage of the target algorithm, the clustering effect of the target algorithm can be quantitatively evaluated, providing a basis for model optimization. Optimizing the learning effect prediction model through coverage improves prediction accuracy and the model's generalization ability, ensuring the reliability of the prediction results. Based on the optimized prediction results, the training strategy of the target job training resource system is adjusted to improve the accuracy of job training resource optimization, the utilization rate of job training resources, and the consistency of job training resource effects. Furthermore, by calculating the Euclidean distance and coverage between the cluster centers of the target algorithm and the evaluation algorithm, a distance matrix is generated, and the number of identical operator nodes is determined, which quantitatively evaluates the clustering effect of the target algorithm; determining the coverage of the target algorithm provides a comprehensive and accurate evaluation basis for model optimization, thereby improving the accuracy and reliability of the learning effect prediction model.
[0042] These or other aspects of the invention will become more apparent from the following description of the embodiments. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 A flowchart of a knowledge base quality detection and evaluation method provided in an embodiment of the present invention;
[0045] Figure 2 A schematic diagram of the structure of a knowledge base quality detection and evaluation system provided in an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present invention. Detailed Implementation
[0047] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0048] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 11, 12, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] To address the issue of low accuracy in predicting the learning outcomes of existing intelligent job training resources, this invention provides a knowledge base quality detection and evaluation method. First, by acquiring real-time learning activity data and real-time environmental data from each operator node in the target job training resource system, a pre-trained learning outcome prediction model is used to generate predicted learning outcomes for a predetermined time period. Then, a knowledge graph self-evaluation processing algorithm based on unsupervised learning is employed to perform cluster analysis on the predicted learning outcomes, yielding a first cluster analysis result containing the location information of the first cluster center points. Next, the Euclidean distance and coverage between the first cluster center points and second cluster center points generated by other algorithms are calculated to quantitatively evaluate the clustering effect of the target algorithm. Finally, the learning outcome prediction model is optimized based on the coverage, and the training strategy of the target job training resource system is adjusted based on the optimized prediction results, thereby achieving more accurate learning outcome prediction and intelligent job training resource training.
[0051] For example, Figure 1 A flowchart of a knowledge base quality detection and evaluation method provided in an embodiment of the present invention is shown below. Figure 1 As shown, the method includes:
[0052] S11. Obtain real-time learning activity data and real-time environment data of each operator node in the target job training resource system, and input the real-time learning activity data and real-time environment data into the pre-trained learning effect prediction model to generate the learning effect prediction value of each operator node in the future preset time period.
[0053] Real-time learning activity data refers to the learning behavior characteristics of each operator node in the target job training resource system at the current point in time, including quantitative indicators reflecting cognitive status such as answer accuracy, knowledge point dwell time, and course video completion rate. Real-time environmental data refers to external job training resource support environment parameters that affect the learning process, such as terminal device performance status (CPU / memory utilization), network transmission quality (video loading latency, real-time interaction packet loss rate), and job training resource scheduling load (number of concurrent access threads). The learning effect prediction model is a machine learning model trained based on historical learning activity data (including knowledge point interaction records, test score time-series changes, etc.) and historical environmental data, used to predict indicators such as knowledge mastery and learning efficiency decay inflection points in future time periods.
[0054] For example, in this embodiment of the invention, the learning activity data (such as clickstream events, quiz submission timestamps, etc.) and environmental data (such as device performance logs, network quality assessment reports, etc.) of each operator node are collected in real time through the job training resource platform's data collection system and terminal monitoring tools. The cleaned multidimensional data is input into a pre-trained learning effect prediction model to generate predicted values for the next three job training resource units, including: a knowledge point mastery probability matrix and a learning efficiency fluctuation warning. For example, the predicted mastery rate for "Object-Oriented Programming Basics" is 82%, and in this embodiment of the invention, an alarm is triggered when the efficiency drops beyond a threshold after 48 hours. The parameters of the learning effect prediction model are optimized through training on full operator data over 12 months. During the training process, a sliding window cross-validation method is used, and the model's generalization weights are adjusted specifically for scenarios of sudden changes in learning behavior (such as the learning recovery period after a long holiday).
[0055] S12. The target algorithm is used to perform cluster analysis on the predicted learning effect values of each operator node to obtain the first cluster analysis result. The target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning. The first cluster analysis result includes the location information of the center point of the first cluster.
[0056] It should be understood that the target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning, used to perform cluster analysis on the predicted values of the learning effect. The first cluster analysis result is the clustering result generated by the target algorithm, including multiple first clusters and the location information of the center point of each first cluster.
[0057] Optionally, the target algorithm groups similar learning performance predictions into the same cluster by calculating the distance between operator nodes (such as Euclidean distance). To determine the cluster centroid within the same cluster, embodiments of the present invention can use mean calculation; the cluster centroid reflects the typical learning performance pattern of that cluster.
[0058] For example, in this embodiment of the invention, a vectorized text dataset X can be generated based on the predicted learning performance values of each operator node within a preset future time period. Based on the vectorized text dataset X, a target algorithm A is used to process the vectorized text dataset X to achieve clustering of the predicted learning performance values of each operator node. The first clustering analysis result... For example, k can be equal to 3 or other values; this embodiment of the invention does not impose specific limitations. Accordingly, in the intelligent job training resources, the target algorithm A performs cluster analysis on the predicted learning performance values of multiple operator nodes, generating 3 clusters (e.g., ...). , , ) and its center point (e.g. , , These center points represent different learning outcome patterns.
[0059] S13. Based on the location information of the first cluster center point and the second cluster center point, determine the coverage of the target algorithm. The second cluster center point is the cluster center point included in the cluster analysis results generated by other algorithms. The other algorithms are evaluation algorithms that are different from the target algorithm.
[0060] The number of other algorithms can be one or more. When there are multiple other algorithms, each of the other algorithms is: , indicating the first The evaluation algorithm is used. Coverage is a metric used to quantify the consistency between the target algorithm and the evaluation algorithm, with a value range of [0,1]. The first evaluation algorithm is used. One evaluation algorithm processes the vectorized text dataset X to cluster the predicted learning performance values of each operator node. The clustering analysis results generated by other algorithms... For example, j can be equal to 3 or other values; this embodiment of the invention does not impose specific limitations. Accordingly, in the intelligent job training resources, the first... An evaluation algorithm performs cluster analysis on the predicted learning performance values of multiple operator nodes, generating two clusters (e.g., ...). ) and its center point (e.g. These center points represent different learning outcome patterns.
[0061] Optionally, in embodiments of the present invention, each cluster generated by the target algorithm is calculated. When determining the center point, the average of the position information of all operator nodes within the cluster can be used as the center point of the cluster. Similarly, in this embodiment of the invention, the center point of each cluster generated by each evaluation algorithm is calculated. When choosing the center point, the average of the location information of all operator nodes within the cluster can also be used as the center point of the cluster.
[0062] S14. Optimize the learning effect prediction model based on the coverage, and adjust the training strategy of the target job training resource system based on the output of the optimized learning effect prediction model.
[0063] The optimized learning outcome prediction model, obtained by adjusting parameters based on coverage, has higher prediction accuracy compared to the unoptimized model. Based on the optimized prediction results, the learning ability matching status of each operator node is analyzed. Based on the output knowledge point mastery probability matrix, the cognitive matching status of operator nodes is systematically analyzed and the following strategies are implemented: 15% additional basic training hours are added for operators with weak knowledge, and access to advanced knowledge points is frozen; a leapfrog learning path is dynamically planned for operators with advanced abilities; and the allocation of computing power on the job training resource server is flexibly adjusted according to the group prediction results. In a specific example, if the model predicts that operator A's mastery of the "concurrent programming" module will reach 92% after 2 hours, the distributed transaction source code analysis course is automatically unlocked, and double the resources are allocated to its learning container to support the operation of the complex code sandbox, thereby achieving a dynamic balance between the supply of job training resources and individual learning needs.
[0064] By executing steps S11-S14, this embodiment of the invention provides accurate and real-time input data for learning effect prediction by collecting real-time learning activity data and environmental data from operator nodes, ensuring the timeliness and reliability of the learning effect prediction model. Through a knowledge graph self-evaluation processing algorithm based on unsupervised learning, cluster analysis is performed on the predicted learning effect values to identify the spatial distribution patterns and potential patterns of the learning effects. By calculating the coverage of the target algorithm, the clustering effect of the target algorithm can be quantitatively evaluated, providing a basis for model optimization. Optimizing the learning effect prediction model through coverage improves prediction accuracy and the model's generalization ability, ensuring the reliability of the prediction results. Based on the optimized prediction results, the training strategy of the target job training resource system is adjusted to improve service accuracy, resource efficiency, and continuity.
[0065] In one possible embodiment, there are multiple first cluster center points and multiple second cluster center points. S13, determine the coverage of the target algorithm based on the location information of the first cluster center points and the location information of the second cluster center points, including:
[0066] Step 131: Based on the position information of each first cluster center point and each second cluster center point, calculate the Euclidean distance between each first cluster center point and each second cluster center point, and generate a distance matrix. Each row of the distance matrix corresponds to a first cluster center point, and each column corresponds to a second cluster center point.
[0067] In step 131, the distance matrix can be a two-dimensional matrix, where each row corresponds to a first cluster center point and each column corresponds to a second cluster center point. Each element in the matrix represents the Euclidean distance between the first cluster center point and the second cluster center point.
[0068] Optionally, embodiments of the present invention can evaluate each algorithm. The Euclidean distance between each cluster center and the cluster center of the target algorithm is calculated, and then the clusters are sorted and the clusters are claimed by the same group. The final result is a new sorted cluster partition, for example, In this formula, n represents the new sorted cluster.
[0069] Step 132: For each row in the distance matrix, select the second cluster center point with the smallest Euclidean distance to the first cluster center point as the matching second cluster center point, and determine the number of the same operator nodes in the cluster where the first cluster center point is located and the cluster where the matching second cluster center point is located.
[0070] For example, the center point of the first cluster The cluster it belongs to contains operator nodes {A,B,C}, and the matching second cluster center point If the cluster contains operator nodes {A,B,D}, then the number of identical operator nodes is 2 (i.e., A and B).
[0071] Step 133: Calculate the coverage of each first cluster center point based on the number of identical operator nodes.
[0072] As one possible implementation, step 133, calculating the coverage of each first cluster center point based on the number of identical operator nodes, includes: for each first cluster center point, determining the ratio of the number of corresponding identical operator nodes to the number of operator nodes in the cluster where the first cluster center point is located as the coverage of the first cluster center point.
[0073] For example, the center point of the first cluster The cluster it belongs to has 10 operator nodes, and the matching second cluster center point If there are 8 identical operator nodes in the cluster, then The coverage is 0.8.
[0074] Step 134: If there is only one evaluation algorithm, the average coverage of all the center points of the first cluster is determined as the coverage of the target algorithm.
[0075] Steps 133 and 134 can be represented by the following formulas:
[0076] ;
[0077] in, For other algorithms Coverage of the automatic evaluation of target algorithm A The number of identical operator nodes, Indicates the cluster where the center point of the first cluster is located. The number of operator nodes, k is the number of the first cluster center points, or the number of the first cluster. The range of its value is (0,1].
[0078] For example, the first cluster has three center points: , and Calculate the average coverage of all center points of the first cluster. For example, if The coverage is 0.8. The coverage is 0.7. If the coverage is 0.9, then the coverage of the target algorithm is (0.8+0.7+0.9) / 3=0.8.
[0079] By executing steps 131 to 134, this embodiment of the invention generates a distance matrix, matches the second cluster center point, calculates the coverage, and determines the coverage of the target algorithm. This enables a quantitative evaluation of the clustering effect of the target algorithm, providing a basis for optimizing the learning effect prediction model, thereby improving the prediction accuracy and the intelligent level of job training resource training.
[0080] In one possible embodiment, after step 133, calculating the coverage of each first cluster center point based on the number of identical operator nodes, the method further includes:
[0081] Step 135: When there are multiple evaluation algorithms, the average coverage of all first cluster center points is determined as the coverage to be optimized by the target algorithm against the evaluation algorithms; the weighted average of the coverage to be optimized by the target algorithm against all evaluation algorithms is taken as the coverage of the target algorithm.
[0082] For example, the first cluster has three center points: , and Calculate the average coverage of all center points of the first cluster. For example, if The coverage is 0.8. The coverage is 0.7. If the coverage is 0.9, then the coverage to be optimized for the first evaluation algorithm is (0.8+0.7+0.9) / 3=0.8.
[0083] Similarly, when there are 3 evaluation algorithms, if the coverage to be optimized for the second evaluation algorithm is 0.7 and the coverage to be optimized for the third evaluation algorithm is 0.6, the coverage of the target algorithm is 0.8×0.6+0.7×0.2+0.6×0.2=0.74.
[0084] Furthermore, when the errors of multiple evaluation algorithms are small, it indicates that the target algorithm is better. The closer the target algorithm's coverage of the optimization target algorithm is to 1 for a single evaluation algorithm, the better its evaluation performance. This invention can determine different thresholds based on different datasets.
[0085] By executing step 135, this embodiment of the invention first calculates the coverage to be optimized for each evaluation algorithm of the target algorithm, and then uses a weighted average aggregation algorithm to summarize the coverage to be optimized for all evaluation algorithms, generating the final coverage of the target algorithm. This method can comprehensively consider the influence of multiple evaluation algorithms, ensuring the comprehensiveness and accuracy of coverage calculation.
[0086] In one possible embodiment, the training process of the learning effect prediction model includes:
[0087] Step 111: Obtain historical learning activity data and historical environment data for each operator node in the target job training resource system.
[0088] Step 112: Based on historical learning activity data and historical environment data, combined with course topology data, train an initial prediction model using machine learning algorithms, such as neural networks, support vector machines, or random forests; the course topology data includes knowledge point relationships and job training resource paths.
[0089] Step 113: Evaluate the performance of the prediction model during training using cross-validation, and adjust the model parameters based on the evaluation results until the performance of the prediction model reaches the preset threshold.
[0090] By executing steps 111-113, this embodiment of the invention ensures the comprehensiveness and diversity of training data by acquiring historical learning activity data, historical environment data, and course topology data, providing rich input information for the model. Using machine learning algorithms such as neural networks, support vector machines, or random forests to train the initial prediction model can capture complex nonlinear relationships in the learning activity data, improving prediction accuracy. The performance of the prediction model under training is evaluated using cross-validation, and model parameters are adjusted based on the evaluation results to ensure that the model's generalization ability and prediction accuracy reach preset thresholds. Combined with course topology data, the model can better adapt to the operating environment of actual job training resources, improving the credibility and practicality of the prediction results. Therefore, the trained learning effect prediction model can accurately predict learning effect values within a future time period, providing reliable data support for optimizing the training strategy of the target job training resource system.
[0091] Accordingly, this invention, through multi-dimensional verification during the model training phase, ensures that the learning effect prediction model meets the quality management standards for job training resources in terms of knowledge transfer generalization (e.g., adaptability across job scenarios) and prediction error control (knowledge point mastery deviation ≤ 5%). Combined with knowledge graph topology data, the model deeply adapts to the dynamic complexity of actual job training resource scenarios, such as the differentiation in operator cognitive levels and sudden fluctuations in the load of job training resources. The optimized model can accurately predict the knowledge internalization trajectory and capability growth inflection points within the future job training resource cycle, providing a reliable decision-making basis for the job training resource system. This includes dynamically adjusting the planned amount of job training resources, reconstructing personalized knowledge point progression paths, and optimizing tutoring resource scheduling strategies. Ultimately, this achieves a systematic improvement in key dimensions of the job training resource system, such as service response agility, resource utilization efficiency, and learning path continuity.
[0092] In one possible embodiment, in S14, the training strategy of the target job training resource system is adjusted based on the output of the optimized learning effect prediction model, including:
[0093] Step 141: Based on the predicted learning performance values of each operator node generated by the optimized learning performance prediction model within a preset future time period, determine the learning ability matching status of each operator node.
[0094] Step 142: Based on the learning ability matching status, adjust the training strategy of the target position training resource system. The training strategy includes the adjusted planned amount of job training resources, knowledge point allocation path and personalized tutoring strategy.
[0095] As one possible implementation, step 142, based on the learning ability matching status, adjusts the training strategy of the target position training resource system, including:
[0096] Step a1: Based on the learning ability matching status, determine whether there are operator nodes with weak knowledge or advanced ability.
[0097] Step a2: For each operator node with weak knowledge, adjust the step size to reduce the complexity of the job training resource plan based on complexity, adjust the step size to extend the learning time cycle of the knowledge point allocation path based on the learning cycle, and allocate corresponding personalized tutoring resources to the operator node. Alternatively, for all operator nodes with advanced capabilities, adjust the step size to increase the task density of the job training resource plan based on density, adjust the step size to reduce the learning time cycle of the knowledge point allocation path based on the learning cycle, and grant the operator node access to self-learning of advanced knowledge graphs.
[0098] By executing steps a1 to a2, this embodiment of the invention uses cluster analysis to identify operator nodes with weak knowledge (knowledge point mastery is more than 30% below the group average) and advanced ability (mastery is more than 50% above the average). For the former, a progressive reinforcement strategy is adopted, reducing the complexity of job training resources (e.g., breaking down multi-step reasoning questions into single-step training), extending the learning time cycle of knowledge point paths (e.g., increasing basic training class hours by 15%-20%), and automatically allocating special tutoring resources (e.g., micro-lessons on explaining wrong questions, one-on-one Q&A channels with tutors, etc.). For the latter, an accelerated leap strategy is activated, increasing task density (e.g., introducing comprehensive cross-job-related topics), compressing the learning time cycle (i.e., allowing skipping of already mastered units), and opening up autonomous exploration permissions for high-level knowledge graphs (e.g., direct access to graduate-level course resources). Through the above-mentioned differentiated adjustments, a dynamic balance between supply and demand of job training resources is achieved: ensuring the continuity of core knowledge acquisition for weak operators (increasing unit pass rate by 25%), avoiding inefficient repetitive training for advanced operators (increasing high-level resource utilization by 40%), and optimizing resource scheduling algorithms based on real-time coverage indicators, so that the embodiments of the present invention can still ensure low response latency of job training resource services under concurrent load scenarios, ultimately achieving synergistic optimization of job training resource quality, resource efficiency and personalized experience.
[0099] By executing steps 141 to 142, this embodiment of the invention accurately analyzes the learning ability matching status of each operator node based on the optimized learning effect prediction value, and promptly identifies nodes with weak knowledge or advanced abilities, providing a scientific basis for adjusting training strategies. Based on the learning ability matching status, the planned amount of job training resources is adjusted. This embodiment of the invention can achieve precise optimization of job training resource allocation. Specifically, based on the optimized predicted values (such as the probability distribution of knowledge point mastery in the next 3 hours), operators with weak knowledge and those with advanced abilities are identified. Based on this, differentiated strategies are triggered. Job training resource support is enhanced for operators with weak knowledge, such as reducing the complexity of the planned amount of job training resources by 30%, extending the training cycle of core knowledge points to 1.5 times the base duration, and adding access permissions for AI tutoring assistants. An acceleration channel is enabled for operators with advanced abilities, such as increasing task density by 50%, compressing repetitive training links, and opening cross-level knowledge graph access permissions. At the same time, the global resource scheduling strategy is reconstructed, and the saved redundant tutoring resources (such as the teachers released from the 20% reduction in basic training for advanced operators) are allocated to the weak group. The waste rate of job training resources is reduced through dynamic knowledge point path planning. Ultimately, the job training resource system achieves a systematic improvement in three indicators: personalized adaptation accuracy, resource utilization efficiency, and service continuity, forming a complete closed loop for job training resource optimization.
[0100] Figure 2 A schematic diagram of the structure of a knowledge base quality detection and evaluation system provided in an embodiment of the present invention is shown below. Figure 2 As shown, the system includes:
[0101] The acquisition module 21 is used to acquire real-time learning activity data and real-time environment data of each operator node in the target job training resource system, and input the real-time learning activity data and real-time environment data into the pre-trained learning effect prediction model to generate the learning effect prediction value of each operator node in the future preset time period.
[0102] Analysis module 22 is used to perform cluster analysis on the predicted learning effect values of each operator node using the target algorithm to obtain the first cluster analysis result. The target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning. The first cluster analysis result includes the location information of the center point of the first cluster.
[0103] The determination module 23 is used to determine the coverage of the target algorithm based on the location information of the first cluster center point and the location information of the second cluster center point. The second cluster center point is the cluster center point included in the cluster analysis results generated by other algorithms. The other algorithms are evaluation algorithms that are different from the target algorithm.
[0104] The optimization and adjustment module 24 is used to optimize the learning effect prediction model based on the coverage, and adjust the training strategy of the target job training resource system based on the output of the optimized learning effect prediction model.
[0105] Figure 2 The aforementioned knowledge base quality detection and evaluation system can perform... Figure 1 The implementation principle and technical effects of the knowledge base quality detection and evaluation method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the knowledge base quality detection and evaluation system in the above embodiments are executed have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0106] In one possible design, Figure 2 The knowledge base quality detection and evaluation system shown in the embodiment can be implemented as a computing device, such as... Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32.
[0107] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 32.
[0108] The processing component 32 is used to: acquire real-time learning activity data and real-time environment data of each operator node in the target job training resource system, and input the real-time learning activity data and real-time environment data into a pre-trained learning effect prediction model to generate learning effect prediction values for each operator node within a preset time period; perform cluster analysis on the learning effect prediction values of each operator node using the target algorithm to obtain a first cluster analysis result, wherein the target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning, and the first cluster analysis result includes the location information of the first cluster center point; determine the coverage of the target algorithm based on the location information of the first cluster center point and the second cluster center point, wherein the second cluster center point is the cluster center point included in the cluster analysis results generated by other algorithms, and the other algorithms are evaluation algorithms different from the target algorithm; optimize the learning effect prediction model based on the coverage, and adjust the training strategy of the target job training resource system based on the output results of the optimized learning effect prediction model.
[0109] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.
[0110] Storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as Random Access Memory (RAM), Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0111] Of course, computing devices may also include other components, such as input / output interfaces, display components, communication components, etc.
[0112] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.
[0113] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.
[0114] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.
[0115] This invention also provides a computer storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 1 The knowledge base quality detection and evaluation method shown in the embodiment.
[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0117] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for quality detection and evaluation of a knowledge base, characterized in that, include: The system acquires real-time learning activity data and real-time environment data of each operator node in the target job training resource system, and inputs the real-time learning activity data and the real-time environment data into a pre-trained learning effect prediction model to generate the learning effect prediction value of each operator node in a future preset time period. The target algorithm is used to perform cluster analysis on the predicted learning effect values of each operator node to obtain the first cluster analysis result; The target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning, and the first clustering analysis result includes the location information of the center point of the first cluster. Based on the location information of the first cluster center point and the location information of the second cluster center point, the coverage of the target algorithm is determined. The second cluster center point is the cluster center point included in the cluster analysis results generated by other algorithms. The other algorithms are evaluation algorithms that are different from the target algorithm. The learning effect prediction model is optimized based on the coverage, and the training strategy of the target job training resource system is adjusted based on the output of the optimized learning effect prediction model.
2. The method according to claim 1, characterized in that, The number of first cluster center points is multiple, and the number of second cluster center points is multiple. Determining the coverage of the target algorithm based on the location information of the first and second cluster center points includes: Based on the location information of each first cluster center point and each second cluster center point, calculate the Euclidean distance between each first cluster center point and each second cluster center point, and generate a distance matrix. Each row of the distance matrix corresponds to a first cluster center point and each column corresponds to a second cluster center point. For each row of the distance matrix, select the second cluster center point with the smallest Euclidean distance to the first cluster center point as the matching second cluster center point, and determine the number of the same operator nodes in the cluster where the first cluster center point is located and the cluster where the matching second cluster center point is located. Calculate the coverage of each first cluster center point based on the number of identical operator nodes; When there is only one evaluation algorithm, the average coverage of all the center points of the first cluster is determined as the coverage of the target algorithm.
3. The method according to claim 2, characterized in that, After calculating the coverage of each first cluster center point based on the number of identical operator nodes, the method further includes: When there are multiple evaluation algorithms, the average coverage of all first cluster center points is determined as the coverage to be optimized by the target algorithm for the evaluation algorithm. The coverage of the target algorithm is the weighted average of the coverage to be optimized for all the evaluated algorithms.
4. The method according to claim 2, characterized in that, The calculation of the coverage of each first cluster center point based on the number of identical operator nodes includes: For each first cluster center point, the ratio of the number of the corresponding identical operator nodes to the number of operator nodes in the cluster where the first cluster center point is located is determined as the coverage of the first cluster center point.
5. The method according to any one of claims 1 to 4, characterized in that, The training process of the learning effect prediction model includes: Obtain historical learning activity data and historical environment data for each operator node in the target job training resource system; Based on the historical learning activity data and historical environment data, combined with the course topology data, an initial prediction model is trained using a machine learning algorithm, such as a neural network, support vector machine, or random forest. The course topology data includes the relationships between knowledge points and the paths to job training resources. The performance of the prediction model during training is evaluated using cross-validation, and the model parameters are adjusted based on the evaluation results until the performance of the prediction model reaches a preset threshold.
6. The method according to claim 1, characterized in that, The step of adjusting the training strategy of the target job training resource system based on the output of the optimized learning effect prediction model includes: Based on the predicted learning performance of each operator node within a preset time period generated by the optimized learning performance prediction model, the learning ability matching status of each operator node is determined. Based on the learning ability matching status, the training strategy of the target job training resource system is adjusted. The training strategy includes the planned amount of job training resources, the knowledge point allocation path, and the personalized tutoring strategy.
7. The method according to claim 3, characterized in that, Based on the learning ability matching status, adjust the training strategy of the target position training resource system, including: Based on the learning ability matching status, determine whether there are operator nodes with weak knowledge or advanced ability; For each operator node with weak knowledge, the step size is adjusted according to the complexity to reduce the complexity of the job training resource plan, and the step size is adjusted according to the learning cycle to extend the learning time cycle of the knowledge point allocation path, and corresponding personalized tutoring resources are allocated to the operator node. For all operator nodes with advanced capabilities, the task density of the job training resource plan is increased by adjusting the step size according to the density, the learning time cycle of the knowledge point allocation path is reduced by adjusting the step size according to the learning cycle, and the operator nodes are given the right to learn the advanced knowledge graph independently.
8. A knowledge base quality detection and evaluation system, characterized in that, include: The acquisition module is used to acquire real-time learning activity data and real-time environment data of each operator node in the target job training resource system, and input the real-time learning activity data and the real-time environment data into a pre-trained learning effect prediction model to generate the learning effect prediction value of each operator node in a future preset time period. The analysis module is used to perform cluster analysis on the predicted learning effect values of each operator node using a target algorithm to obtain a first cluster analysis result. The target algorithm is a knowledge graph self-evaluation processing algorithm based on unsupervised learning. The first cluster analysis result includes the location information of the center point of the first cluster. The determination module is used to determine the coverage of the target algorithm based on the location information of the first cluster center point and the location information of the second cluster center point. The second cluster center point is the cluster center point included in the cluster analysis results generated by other algorithms, and the other algorithms are evaluation algorithms that are different from the target algorithm. The optimization and adjustment module is used to optimize the learning effect prediction model based on the coverage, and adjust the training strategy of the target job training resource system based on the output of the optimized learning effect prediction model.
9. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a knowledge base quality detection and evaluation method as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, The system contains a computer program that, when executed by a computer, implements a knowledge base quality detection and evaluation method as described in any one of claims 1 to 7.