A grid worker skills evaluation method and system

Through deep reinforcement learning and knowledge graph technology, a multi-dimensional evaluation model and personalized learning recommendation system are built, which solves the problem of low personalization of existing grid worker skills evaluation, and achieves efficient and continuous improvement of grid worker skills evaluation.

CN118966872BActive Publication Date: 2025-05-13ZHEJIANG WUXINSHUKE INFORMATION IND CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410990537.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-05-13
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

The existing grid worker skill evaluation mechanism lacks a data-driven closed loop of evaluation learning, and fails to make full use of massive grid work data, resulting in low evaluation personalization and inability to effectively improve grid worker skills.

Method used

The deep reinforcement learning algorithm is used to dynamically optimize the skill evaluation strategy of grid workers, and introduce knowledge graph inference and link prediction. By collecting and preprocessing the basic, work results and learning data of grid workers, a multi-dimensional evaluation model is built, personalized skill evaluation results are generated, and personalized learning content is recommended based on the knowledge graph.

Benefits of technology

It has improved the pertinence and effectiveness of grid workers' evaluation, formed a virtuous cycle of "evaluation-learning-re-evaluation-re-learning", continuously optimized the evaluation and learning mechanism, and promoted the continuous improvement of grid workers' skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118966872B_ABST
    Figure CN118966872B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for evaluating the skills of grid workers, which relates to the field of grid worker evaluation, including: collecting grid worker skill data; using a machine learning algorithm based on reinforcement learning to construct a multidimensional evaluation model for grid workers; using the collected skill data as input, and obtaining personalized skill evaluation results through the constructed multidimensional evaluation model calculation; inputting the obtained personalized skill evaluation results into a recommendation algorithm based on a knowledge graph to generate personalized recommended learning content; obtaining the updated skill data of the grid worker after completing the recommended learning content, and preprocessing the updated skill data as a second training sample set; using a distributed incremental learning method based on a parameter server, and using the second training sample set to adjust the constructed multidimensional evaluation model to obtain an updated multidimensional evaluation model. In view of the low degree of personalization of grid worker skill evaluation in the prior art, the present application improves the pertinence of grid worker evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of grid worker evaluation, and in particular to a grid worker skill evaluation method and system. Background Art

[0002] With the continuous improvement of the modern urban governance system, grid management has become an important measure to improve the level of urban refined governance. As an important executor of grid management, the skill level of grid workers is directly related to the quality and efficiency of grid work. In order to ensure the orderly development of grid worker team building and grid management work, a scientific and effective grid worker skill evaluation and improvement mechanism is urgently needed.

[0003] At present, the skill evaluation of grid workers mainly adopts the traditional assessment and evaluation method, which conducts qualitative and quantitative evaluation by setting key performance indicators (KPIs) and task completion. The training and learning for grid workers also adopts a "one-size-fits-all" unified course. The learning content is out of touch with the actual work needs of grid workers, and the learning effect is difficult to guarantee. The root cause of these problems is that the existing grid worker management mechanism lacks a data-driven evaluation and learning closed loop, fails to make full use of massive grid work data, and has not established an effective connection between evaluation results and recommended learning. There is a gap between evaluation and learning, and lacks continuous optimization capabilities.

[0004] In the related art, for example, Chinese patent document CN117540903A provides a method, system and application of automatic allocation of community grid workers, including dividing the city into community grids on a map; establishing evaluation indicators for each grid; determining the complexity coefficient of each evaluation indicator according to the importance of the evaluation indicator; establishing an automatic allocation model based on the complexity coefficient and the evaluation indicator; entering the information of the grid workers according to the initial grid worker allocation value and realizing association through the common attribute number of the database field, and obtaining a new grid worker allocation value according to the data change of each evaluation indicator; normalizing the initial grid worker allocation value and the new grid worker allocation value; performing complexity correction calculation on the normalized data according to the number of evaluation indicators and the update of the influencing factors to obtain the latest grid worker allocation value; adjusting the number of grid workers according to the latest grid worker allocation value; and displaying on the map according to the adjusted results. However, the construction of the automatic allocation model of this scheme depends on the data change of the evaluation indicator. If these data cannot fully reflect the actual skill level or experience accumulation of the grid workers, the grid worker allocation value output by the model may be inaccurate or insufficient to meet personalized needs. Summary of the invention

[0005] In response to the problem of low personalization in grid worker skill evaluation in the prior art, the present application provides a grid worker skill evaluation method and system, which dynamically optimizes the grid worker skill evaluation strategy through a deep reinforcement learning algorithm, and introduces knowledge graph reasoning and link prediction, thereby improving the pertinence of grid worker evaluation.

[0006] One aspect of this specification provides a method for evaluating grid worker skills, including:

[0007] Collecting grid worker skill data ; Skill data X includes basic data B, work results data W and learning data L of grid workers; among which, basic data B includes name b1, gender b2, age b3, education b4 and major b5; work results data W includes event reporting data U and inspection data C, among which, event reporting data U includes M1 major event type u1 and M2 minor event type u2, and inspection data C includes inspection indicator list ; Learning data L includes daily learning status .

[0008] Preprocess the collected skill data X as the first training sample set , n=1, 2, ..., N; the first training sample set is trained using a machine learning algorithm based on reinforcement learning Perform feature extraction and model training to build a multi-dimensional evaluation model for grid workers ; Among them, reinforcement learning adopts setting reward function , establish a mapping relationship between the grid worker’s evaluation result y and the reward value r, and take the cumulative sum of the reward values ​​R (τ) as the optimization target;

[0009] Skill data to be collected As input, the multidimensional evaluation model constructed Calculate and obtain personalized skill evaluation results .

[0010] Input the obtained personalized skill evaluation result Y into the recommendation algorithm based on knowledge graph Generate personalized recommended learning content ; Among them, the recommendation algorithm Leverage pre-built grid work knowledge graphs According to the personalized skill evaluation results Y, through knowledge reasoning and link prediction, personalized recommended learning content Z is generated according to different grid workers; among them, the grid work knowledge graph G contains a concept node set V and an association relationship edge set E.

[0011] Obtain the updated skill data of the grid worker after completing the recommended learning content Z , the updated skill data Preprocess as the second training sample set , n=1, 2, ..., N'; using a distributed incremental learning method based on a parameter server , using the second training sample set The multidimensional evaluation model constructed Make adjustments to obtain an updated multidimensional evaluation model .

[0012] Further, the steps are to build a multi-dimensional evaluation model for grid workers. , including: defining the state space S, which contains the skill data xi of the grid worker; defining the action space A, which contains the skill data xi of the grid worker Evaluation action , i=1,2,...,N, j=1,2,...,Mi; define the reward function , used to establish the mapping relationship between the grid worker's evaluation result y and the reward value r; construct a deep neural network as the evaluation strategy function , randomly initialize the parameters θ of the evaluation strategy function π, and let the initial state ;

[0013] According to the current status And the parameter θ of the evaluation strategy function π, each evaluation action is calculated through a deep neural network The probability distribution of , and according to the probability distribution Select an evaluation action ; Execute evaluation action , get instant rewards , and obtain the new state ;Will It is stored in the experience replay pool D as an experience sample.

[0014] Randomly sample a batch of experience samples from the experience replay pool D , take advantage of instant rewards and the next state The estimated Q value of Compute the current state-action pair The target Q value ; and use the target Q value and estimated Q value The mean square error is used as the loss function , update the parameter θ of the evaluation strategy function π through the gradient descent algorithm; let t=t+1, repeat the above iterations until the preset number of training rounds T is reached or the evaluation strategy function π converges; output the final evaluation strategy function As a multidimensional evaluation model for grid workers .

[0015] Further, select an evaluation action , including: the current state Input into the deep neural network corresponding to the evaluation strategy function π, and after multiple layers of nonlinear transformation, the state is obtained The feature representation vector ;

[0016] Represent the feature vector Input to the output layer of the evaluation strategy function π, and converted into each action in the evaluation action space A through the softmax function The probability distribution of ; According to the probability distribution , select an evaluation action through the ε-greedy strategy .

[0017] Further, select an evaluation action , also includes: setting the entropy regularization term , the evaluation action The probability distribution of The entropy of is added to the loss function as a regularization constraint We get a new loss function , , where β is the entropy regularization coefficient;

[0018] Calculate the new loss function The gradient of the parameter θ about the evaluation policy function π , update the parameter θ using the gradient descent algorithm, , where α is the learning rate; repeat the above iterations until the evaluation strategy function π converges or reaches the preset number of updates.

[0019] Furthermore, the parameter θ of the evaluation strategy function π is updated by the gradient descent algorithm, including:

[0020] Randomly sample a batch of experience samples from the experience replay pool D ;

[0021] For each experience sample , using the target evaluation strategy function Calculate the next state The estimated Q value of , where a is the next state All possible actions; use Bellman optimal equation to calculate the current state-action pair The target Q value , , where γ is the discount factor; Indicates in status Take action The optimal Q value estimate of . Indicates in status Take action γ is the discount factor used to decay the value of future rewards, and its value range is usually between [0, 1]. Indicates that in the next state Estimate the maximum Q value among all possible actions a.

[0022] Using the evaluation strategy function Compute the current state-action pair The estimated Q value of ;

[0023] Using the target Q value and estimated Q value The mean square error constructs the loss function , ; Calculate the loss function The gradient of the parameter θ about the evaluation policy function π , update the parameter θ using the gradient descent algorithm, , where α is the learning rate;

[0024] Repeat the above iterations until the loss function Convergence or reaching the preset number of updates.

[0025] Furthermore, generate personalized recommended learning content Z, including: pre-building grid work knowledge graph , the knowledge graph G contains a concept node set V and an association edge set E; among them, the concept node Represents the abstract concept of grid work, with associated edges Represents a concept node and The semantic association between

[0026] The obtained personalized skill evaluation result Y is used as the input query CX, and the target node set T related to the query CX semantics is searched in the knowledge graph G. The target node set , set T is a subset of set V; taking the nodes in the target node set T as the starting nodes, using knowledge reasoning and link prediction algorithms, inferring the extended node set semantically associated with the target node set T in the knowledge graph G , expand the node set ; According to the target node set T and the expansion node set The concept nodes contained in are combined with the pre-configured recommendation rule base R to generate personalized recommended learning content for different grid workers through pattern matching and filtering sorting. .

[0027] Preferably, according to the target node set T and the extended node set Combined with the pre-configured recommendation rule base R, through pattern matching and filtering sorting, personalized recommended learning content L for different grid workers is generated, including: according to the target node set T and the extended node set The types, attributes and associations of the concept nodes and entity nodes contained in are used to perform pattern matching in the recommendation rule library R to retrieve the recommendation rules that match the personalized skill evaluation results of the current grid worker; the placeholders in the recommendation rules are replaced with specific concept nodes or entity nodes to generate preliminary personalized recommendation learning content. ; According to the pre-configured filtering conditions in the recommendation rule library R, ​​the preliminary personalized recommendation learning content Filter to obtain the filtered personalized recommended learning content ; According to the pre-configured sorting strategy in the recommendation rule library R, ​​the filtered personalized recommendation learning content is Sort and generate the final personalized recommended learning content L.

[0028] Furthermore, using knowledge reasoning and link prediction algorithms, we can infer the extended node set semantically associated with the target node set T in the knowledge graph G. , including: taking each node in the target node set T As the starting node, use the path-based reasoning algorithm to search for the starting node in the knowledge graph G. As the head node, other nodes in the knowledge graph G The inference path of the tail node ; Among them, the reasoning path By the association edge set Composition, reasoning path Length Less than or equal to the preset maximum length threshold M.

[0029] According to the reasoning path Length , association edge The semantic relevance weight , and the tail node Importance , calculate the inference path The semantic relevance score of ;

[0030] Selecting a semantic relevance score Reasoning paths greater than the preset threshold δ The tail node in , as the starting node Semantically associated extension nodes are added to the extension node set In; to expand the node set Each node in is a new starting node, and the process is repeated until the node set is expanded. No new nodes are added or the preset number of iterations is reached.

[0031] Further, according to the reasoning path Length , association edge The semantic relevance weight , and the tail node Importance , calculate the inference path The semantic relevance score of , including: Calculating the inference path Length score L(P(i, j)), L(P(i, j)) = 1 / log(1+len(P(i, j))); calculate the inference path Each associated relationship edge The semantic relevance weight The sum of the inference path The semantic relevance score of , ; represents an inference path from node i to node j. Indicates the reasoning path An edge on , connecting node p and node q. Represents edge The semantic relevance weight of Each edge on , and its corresponding semantic relevance weight Add together to get the semantic relevance score of the entire reasoning path The object of summation is the inference path All edges on The semantic relevance weight By summing up, we can calculate the semantic relevance score of the entire reasoning path. The first edge on , take its semantic relevance weight For the inference path The second edge on , take its semantic relevance weight , and add it to the result of the first step. And so on, for the reasoning path The semantic relevance weights of each edge on the path are accumulated. The sum of the semantic relevance weights of all edges is used as the inference path The semantic relevance score of . Get the semantic relevance score of the entire reasoning path. This score can be used to measure the strength of the semantic relevance of the reasoning path.

[0032] Compute inference paths The tail node Importance score , ,in The tail node The degree in the knowledge graph G, that is, the degree of the node The number of directly connected relationship edges;

[0033] Score the length , semantic relevance score and importance score Linear weighted summation to obtain the inference path The semantic relevance score of .

[0034] Furthermore, the relationship edge The semantic relevance weight Calculated by the following steps:

[0035] In the concept node set V of the knowledge graph G, count any two concept nodes and Frequency of co-occurrence , as a node and The co-occurrence frequency between them; for each associated relationship edge in the knowledge graph G , get the two concept nodes it connects and , using point wise mutual information (PMI) to calculate and The semantic relevance weight , ;in, and Respectively represent nodes and The frequency of occurrence in the knowledge graph G, Representation Node and co-occurrence frequency; weight the semantic relevance Normalized to the interval [0, 1] as the association edge The final semantic relevance weight of .

[0036] Another aspect of the present specification also provides a grid worker skill evaluation system for executing a grid worker skill evaluation method.

[0037] Using deep reinforcement learning algorithms, the evaluation strategy function is trained through the state-behavior-reward data generated by grid workers in their work practice, so that the evaluation results can objectively reflect the skill level of grid workers in the actual working environment. Traditional subjective evaluation methods often rely on expert experience and subjective judgment, and the evaluation results are easily affected by human factors. This application adopts a data-driven reinforcement learning method, which adaptively optimizes the evaluation strategy from environmental feedback through learning and refining massive work data, thus overcoming the limitations of subjective evaluation.

[0038] The deep reinforcement learning framework is introduced, and the evaluation strategy function can be adaptively optimized and updated through continuous strategy learning iteration. Traditional evaluation methods usually use fixed evaluation index systems and weights, which are difficult to adapt to the complex and changeable grid working environment. This application models the grid worker skill evaluation as a Markov decision process, and through instant reward feedback on the grid worker's behavior, the evaluation strategy function can be dynamically adjusted to continuously adapt to changes in the grid worker's skill level and changes in the working environment, and realize adaptive optimization of the evaluation strategy.

[0039] On the basis of personalized skill evaluation, knowledge graph technology is used to realize personalized generation of recommended learning content. Through knowledge reasoning and link prediction algorithms, learning content that highly matches the skill shortcomings of grid workers is discovered from the massive grid work knowledge graph, which improves the personalization of learning recommendations. Traditional learning recommendation methods are often based on learning paths or popular learning resources summarized by domain experts, and the recommended content lacks pertinence and effectiveness. This application integrates personalized skill evaluation results and uses semantic association analysis to recommend and filter out tailored learning content from the knowledge graph, making grid workers' learning more efficient and targeted, and effectively improving the effectiveness of learning recommendations.

[0040] The skill evaluation of grid workers is closely combined with targeted learning, forming a virtuous cycle of "evaluation-learning-re-evaluation-re-learning". Through personalized skill shortcoming analysis and recommended learning, an effective path for grid workers to continuously improve their skills is provided. At the same time, work practice data and learning feedback data are used to continuously optimize the evaluation and learning mechanism to promote the continuous improvement of grid workers' skill levels. This application provides effective theoretical and technical support for the continuous improvement of grid workers' skills and has significant application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1is an exemplary flow chart of a method for evaluating grid worker skills according to some embodiments of this specification;

[0042] Figure 2 is an exemplary flow chart for constructing a multidimensional evaluation model according to some embodiments of this specification;

[0043] Figure 3 is an exemplary flow chart of obtaining personalized recommended learning content according to some embodiments of this specification;

[0044] Figure 4 is an exemplary flow chart of obtaining an extended node set according to some embodiments of this specification;

[0045] Figure 5 It is an exemplary module diagram of a grid worker skill evaluation system shown in some embodiments of this specification. DETAILED DESCRIPTION

[0046] The method and system provided in the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0047] Figure 1This is an exemplary flow chart of a grid worker skill evaluation method according to some embodiments of the present specification, a grid worker skill evaluation method, comprising: collecting grid worker skill data; the skill data includes basic data, work performance data and learning data of the grid worker; wherein the basic data includes name, gender, age, education and major; the work performance data includes event reporting data and inspection data, wherein the event reporting data includes M1 major category and M2 minor category event types, and the inspection data includes an inspection indicator list; the learning data includes daily learning status; the collected skill data is pre-processed as a first training sample set; a machine learning algorithm based on reinforcement learning is used to extract features and train a model for the first training sample set, and a multi-dimensional evaluation model for the grid worker is constructed; wherein the reinforcement learning adopts the setting of a reward function to establish a mapping relationship between the evaluation results of the grid worker and the reward value, and uses The cumulative sum of the reward values ​​is the optimization target; the collected skill data is used as input, and the personalized skill evaluation results are calculated through the constructed multidimensional evaluation model; the obtained personalized skill evaluation results are input into the recommendation algorithm based on the knowledge graph to generate personalized recommended learning content; wherein, the recommendation algorithm uses the pre-constructed grid work knowledge graph, and according to the personalized skill evaluation results, through knowledge reasoning and link prediction, generates personalized recommended learning content according to different grid workers; wherein, the grid work knowledge graph contains the concepts, entities and associations of grid work; the skill data updated after the grid worker completes the recommended learning content is obtained, and the updated skill data is pre-processed as the second training sample set; a distributed incremental learning method based on a parameter server is adopted, and the constructed multidimensional evaluation model is adjusted using the second training sample set to obtain an updated multidimensional evaluation model.

[0048] Through a variety of data collection channels, we obtain data related to grid worker skills and construct a grid worker skill data set X = {x1, x2, ..., xN}, where N is the total number of grid workers. The skill data xi of each grid worker consists of basic data B, work results data W and learning data L. The basic attribute information of grid workers, including name b1, gender b2, age b3, education b4 and major b5, is extracted from the personnel information database of the grid worker management system to form the basic data B of grid workers. These data are usually entered into the system when grid workers join the company and can be directly obtained. The work results data W consists of two parts: event reporting data U and inspection data C. For the event reporting data U, the event data reported by each grid worker within a certain period of time is obtained from the event reporting module of the grid management work platform. According to the predefined event classification system, the event type is divided into M1 major category and M2 minor category to form the event reporting data U of each grid worker, where u1 represents the M1 major category event type reported by the grid worker, and u2 represents the M2 minor category event type reported. In this embodiment, M1 takes 21; M2 takes 512. For the inspection data C, the data records of each grid member in the daily inspection work are obtained from the daily inspection module of the grid management work platform. According to the grid inspection indicator system, a list C={c1, c2,..., cK} containing K inspection indicators is set, and the data of each grid member on each inspection indicator is extracted to form the inspection data C of the grid member. The learning behavior data of each grid member in a certain period of time is obtained from the learning record database of the grid member online learning platform. The daily learning situation is recorded as a T-dimensional vector L={l1, l2,..., lT}, where lt represents the learning behavior data of the grid member on the tth day, which may include information such as learning time, learning courses, learning progress, test scores, etc.

[0049] Figure 2It is an exemplary flowchart for constructing a multidimensional evaluation model according to some embodiments of this specification, defining the basic elements of reinforcement learning. Definition of state space S: The skill data xi of the grid worker is taken as an element of state space S, and each xi represents the skill state of a grid worker. The skill data xi includes basic data B, work result data W and learning data L, that is, xi=(Bi, Wi, Li). The basic data Bi contains attribute information such as the name, gender, age, education and major of the grid worker. The work result data Wi contains event reporting data Ui and inspection data Ci, reflecting the work performance of the grid worker. The learning data Li contains the daily learning situation of the grid worker, reflecting its learning behavior and effect. The dimension of state space S is spliced ​​by the skill data of all grid workers, that is, S=(x1, x2,..., xN). Definition of action space A: The evaluation actions ai,j of each skill xi of the grid worker are taken as elements of action space A. Each evaluation action ai,j represents the evaluation of the jth skill of the i-th grid worker. i=1, 2,..., N represents the number of the grid worker, and N is the total number of grid workers. j=1, 2, ..., Mi represents the number of evaluation actions of the i-th grid worker, and Mi represents the number of skill items of the i-th grid worker. Evaluation actions ai, j can be discrete grade evaluations (such as excellent, good, qualified, unqualified) or continuous score evaluations (such as 0-100 points). The dimension of the action space A is the sum of the number of evaluation actions of all grid workers, that is, A= (a1, 1, ..., a1, M1, a2, 1, ..., a2, M2, ..., aN, 1, ..., aN, MN). Definition of reward function r (s, a): The reward function r (s, a) is used to establish a mapping relationship between the evaluation result y of the grid worker and the reward value r. The evaluation result y represents the skill evaluation result of the grid worker obtained by taking action a in state s, which can be a comprehensive evaluation or a sub-item evaluation. The reward value r represents the degree of goodness of the evaluation result y, and is usually designed as a numerical function, such as r=f (y). The design of the reward function needs to consider factors such as the weight of the evaluation indicators and the distribution of the evaluation results, so that the reward value can reasonably reflect the quality of the evaluation results. Common reward function design methods include linear functions, exponential functions, logarithmic functions, etc., or rule functions designed based on expert knowledge.

[0050] Construct a deep neural network as an approximate representation of the evaluation policy function π(a|s, θ). Randomly initialize the parameter θ of the evaluation policy function π. The evaluation policy function π(a|s, θ) takes the state s as input and the probability of taking action a as output, and the parameter θ is the weight parameter to be learned. Use the Xavier initialization method or the He initialization method to randomly generate the initial value of the parameter θ so that the output variance of each layer of the neural network is the same. The value of the initial parameter θ usually obeys a uniform distribution or a normal distribution, and the initialization range and variance can be set by hyperparameters. Set the initial state s0, s0∈S. Randomly select an initial state s0 from the state space S as the starting point of reinforcement learning. The initial state s0 can be the skill data of a grid worker or a combination of skill data of multiple grid workers.

[0051] According to the current state st and the evaluation strategy function π, the evaluation action probability distribution is calculated through the deep neural network, and the evaluation action at is selected. The current state st is input into the deep neural network, and the feature representation vector ht is obtained after multiple layers of nonlinear transformation. The current state st represents the skill data of the grid worker, including basic data, work results data and learning data, that is, st= (Bt, Wt, Lt). The state st is input into the input layer of the deep neural network, and the activation value of the hidden layer is obtained through forward propagation calculation. The calculation formula of the hidden layer is: ht=f(Wh*st+bh), where Wh is the weight matrix of the hidden layer, bh is the bias vector of the hidden layer, and f is the activation function (such as ReLU, tanh, etc.). After multiple hidden layers of nonlinear transformation, the original state st is mapped to a high-dimensional feature representation vector ht, and the key information of the state is extracted. The feature representation vector ht is input into the output layer, and the evaluation action probability distribution π(a|st, θ) is obtained through the softmax function. The feature representation vector ht of the hidden layer is input into the output layer of the neural network to calculate the score of each evaluation action. The calculation formula of the output layer is: ot=Wo*ht+bo, where Wo is the weight matrix of the output layer and bo is the bias vector of the output layer. The score vector ot of the output layer is converted into the probability distribution of the evaluation action π(a|st,θ) through the softmax function. The calculation formula of the softmax function is: π(ai|st,θ)=exp(ot[i]) / ∑jexp(ot[j]), where ot[i] represents the score of the i-th evaluation action. According to the ε-greedy strategy, an action is randomly selected with a probability of ε, and the action with the highest probability is selected with a probability of 1-ε to obtain the evaluation action at. The exploration rate ε (0<ε<1) is introduced to balance exploration and utilization. Generate a random number r between 0 and 1. If r<ε, randomly select an evaluation action at; otherwise, select the evaluation action at with the highest probability. The probability of randomly selecting an action is ε / |A|, where |A| is the size of the evaluation action space. The probability of selecting the action with the highest probability is 1-ε+ε / |A|, that is, adding a small exploration probability to the greedy action. The entropy regularization term H(π(*|st, θ)) is introduced, and the entropy of the action probability distribution is added to the loss function as a regularization constraint to obtain a new loss function L'(θ). The entropy regularization term H(π(*|st, θ)) represents the entropy of the action probability distribution π(a|st, θ), and the calculation formula is: H(π(*|st, θ)) = -∑aπ(a|st, θ)logπ(a|st, θ). The entropy regularization term is added as an additional constraint to the original loss function L(θ), and a new loss function L'(θ) = L(θ) - β*H(π(*|st, θ)) is obtained. The entropy regularization coefficient β controls the regularization strength. The larger β is, the more it encourages the exploration and randomness of the policy function.

[0052] Calculate the gradient ∇θL'(θ) of L'(θ) with respect to the parameter θ, and update the parameter θ by the gradient descent algorithm. Use the back propagation algorithm to calculate the gradient ∇θL'(θ) of the new loss function L'(θ) with respect to the neural network parameter θ. Update the parameter θ according to the update formula of the gradient descent algorithm: θ=θ-α*∇θL'(θ), where α is the learning rate. The learning rate α controls the step size of the parameter update and needs to be set reasonably to balance the convergence speed and stability. Repeat the iteration until the policy function converges or reaches the preset number of updates. Set a threshold for the number of updates or a convergence condition, and repeat the iteration to continuously update the parameters of the policy function. When the parameters of the policy function change very little or reach the preset number of updates, the policy function is considered to have converged and the update process is stopped. The converged policy function can give a reasonable probability distribution of evaluation actions based on the state data and complete the learning optimization of the evaluation strategy. Execute the evaluation action at to obtain the immediate reward rt and the new state st+1. Evaluate the skill data of the grid worker based on the selected evaluation action at. The evaluation action at can be a score, rating or ranking of a skill dimension. After executing the evaluation action at, the immediate reward rt is calculated according to the preset reward function r (s, a). The immediate reward rt reflects the quality of the action at under the state st, which can be a scalar value or a vector value. At the same time, after executing the action at, the skill state of the grid worker will also change, and a new state st+1 will be obtained. The new state st+1 contains the impact of the action at on the grid worker's skills and the updated skill data.

[0053] The four-tuple (st, at, rt, st+1) is stored in the experience replay pool D as an experience sample. The state st, action at, reward rt and next state st+1 form a four-tuple (st, at, rt, st+1), which represents a complete evaluation interaction process. The four-tuple (st, at, rt, st+1) is used as an experience sample and stored in the experience replay pool D. The experience replay pool D is a fixed-size buffer for storing historical interaction data. When the experience replay pool D is full, the sample can be updated by first-in-first-out (FIFO) or random replacement. The experience replay mechanism can break the temporal correlation between data and improve sample utilization efficiency and training stability.

[0054] Sample a batch of experience samples from the experience replay pool D, calculate the target Q value and estimated Q value, construct the loss function and update the policy function parameters. Randomly sample a batch of experience samples {(sk, ak, rk, sk+1)} from the experience replay pool D. Randomly sample a batch of experience samples of size K {(sk, ak, rk, sk+1)} from the experience replay pool D, k=1, 2,..., K. The sampling process usually uses uniform random sampling or priority-based sampling methods. The sampled experience samples are used to calculate the loss function and update the parameters of the policy function. Use the target policy function π' to calculate the estimated Q value Q'(sk+1, a) of the next state sk+1. Using the target policy function π'(a|s, θ'), calculate the estimated Q value Q'(sk+1, a) of each action a according to the next state sk+1. The parameter θ' of the target policy function π' is a delayed update version of the parameter θ of the current policy function π, which is used to improve the stability of the estimate. Q'(sk+1, a) represents the expected long-term reward of taking action a in state sk+1, which is estimated based on the target policy function π'. The target Q value Q*(sk, ak) is calculated using the Bellman optimal equation. Using the Bellman optimal equation, the target Q value Q*(sk, ak) of the state-action pair (sk, ak) is calculated. Bellman optimal equation: Q*(sk, ak)=rk+γ*max_a[Q'(sk+1, a)], where γ is a discount factor with a value range of [0, 1], indicating the degree of decay of future rewards. max_a[Q'(sk+1, a)] represents the estimated Q value corresponding to the action a with the largest Q value selected in the next state sk+1. The estimated Q value Q'(sk, ak, θ) is calculated using the current policy function π. Using the current policy function π(a|s, θ), the estimated Q value Q'(sk, ak, θ) is calculated based on the state-action pair (sk, ak). Q'(sk, ak, θ) represents the expected long-term reward of taking action ak in state sk, which is estimated based on the current policy function π. The loss function L(θ) is constructed using the mean square error between the target Q value and the estimated Q value. The mean square error loss function L(θ) is constructed using the target Q value Q*(sk, ak) and the estimated Q value Q'(sk, ak, θ). The loss function L(θ) = E[(Q*(sk, ak)-Q'(sk, ak, θ))^2], where E[] represents the average of the sampled empirical samples. The loss function L(θ) measures the difference between the estimated Q value and the target Q value, reflecting the evaluation error of the current policy function π.

[0055] Calculate the gradient ∇θL(θ) of L(θ) with respect to the parameter θ, and update the parameter θ by the gradient descent algorithm. Use the backpropagation algorithm to calculate the gradient ∇θL(θ) of the loss function L(θ) with respect to the policy function parameter θ. According to the gradient descent algorithm, update the parameter θ: θ=θ-α*∇θL(θ), where α is the learning rate, which controls the step size of each parameter update. By continuously iteratively updating the parameter θ, the estimated Q value Q'(sk, ak, θ) gradually approaches the target Q value Q*(sk, ak), improving the evaluation quality of the policy function. Repeat the iteration until the loss function converges or reaches the preset number of updates. Set a threshold for the number of updates or a convergence condition, and repeat the execution of continuously updating the parameters of the policy function. When the value of the loss function changes very little or reaches the preset number of updates, it is considered that the policy function has converged and the update process stops. The converged policy function can give a more accurate estimate of the state-action value, providing a reliable basis for action selection. Let t=t+1, and repeat until the preset number of training rounds T is reached or the evaluation policy function converges. Add 1 to the time step t and enter the next evaluation decision stage. Repeat the execution to continuously interact with the environment and update the policy function until the preset number of training rounds T is reached or the evaluation policy function converges. The number of training rounds T controls the total duration of policy learning and needs to be reasonably set according to the complexity of the task and the convergence speed. Take the collected skill data X'={x'1, x'2, ..., x'N'} as input, and calculate the personalized skill evaluation result Y={y1, y2, ..., yN'} through the constructed multidimensional evaluation model f(x, θ); obtain a batch of skill data X'={x'1, x'2, ..., x'N'} of the grid workers to be evaluated from the grid worker management system or database. The skill data x'i of each grid worker includes information in multiple dimensions such as basic data, work results data and learning data. The basic data includes the personal attributes of the grid worker, such as name, gender, age, education, years of work, etc. The work results data includes the work performance indicators of the grid worker, such as the number of incident reports, problem solving rate, customer satisfaction, etc. The learning data includes the grid workers' training participation, test scores, skill certification and other ability improvement indicators. The preprocessed skill data X' is used as input and sent to the trained multidimensional evaluation model f (x, θ). The multidimensional evaluation model f (x, θ) is trained by deep reinforcement learning method. The model parameters θ include the weights and biases of the deep neural network and the hyperparameters of the reinforcement learning algorithm. The model f (x, θ) can evaluate the ability level and development potential of grid workers in different dimensions based on their skill data. Through the forward propagation algorithm, the skill data X' is input into the deep neural network of the multidimensional evaluation model f (x, θ). The neural network maps the input data X' to the output result Y through layer-by-layer calculation. The output result Y = {y1, y2, ..., yN'} represents the personalized skill evaluation result of each grid worker.The evaluation result yi can be a scalar value, indicating the comprehensive ability score of the grid worker; or it can be a vector, indicating the ability level or ranking of the grid worker in different dimensions.

[0056] Figure 3 This is an exemplary flow chart for obtaining personalized recommended learning content according to some embodiments of this specification, and recommended learning: pre-constructing a grid work knowledge graph G = (V, E), the knowledge graph G includes a concept node set V and an association relationship edge set E; wherein, the concept node vi∈V represents the abstract concept of grid work, and the association relationship edge ei,j∈E represents the semantic association between the concept nodes vi and vj. Pre-construct a knowledge graph G covering the grid work field through knowledge engineering or automatic construction methods. The knowledge graph G is represented in the form of a directed graph or an undirected graph, and includes a concept node set V and an association relationship edge set E. The concept node vi represents an abstract concept in grid work, such as "power equipment", "fault handling", "safe operation", etc. The association relationship edge ei,j represents the semantic association between two concept nodes vi and vj, such as "belong to", "include", "cause and effect", etc. The knowledge graph G forms a structured domain knowledge representation by modeling the core concepts of grid work and their association relationships. The obtained personalized skill evaluation result Y is used as the input query CX, and the target node set T semantically related to the query CX is searched in the knowledge graph G, and the target node set T⊆V. The obtained personalized skill evaluation result Y is used as the input query CX. The query CX can be a vector representation, expressing the evaluation score or level of the grid worker in different skill dimensions. In the knowledge graph G, semantic search algorithms such as semantic similarity calculation and keyword matching are used to retrieve concept nodes semantically related to the query CX. The retrieved concept nodes are combined into a target node set T, which represents the knowledge points closely related to the current skill level of the grid worker.

[0057] Figure 4This is an exemplary flow chart for obtaining an extended node set according to some embodiments of the present specification, taking a node in the target node set T as a starting node, using knowledge reasoning and link prediction algorithms, inferring an extended node set E' semantically associated with the target node set T in the knowledge graph G, and the extended node set E'⊆V. Using knowledge reasoning and link prediction algorithms, inferring an extended node set E' semantically associated with the target node set T in the knowledge graph G, including: taking each node vi in ​​the target node set T as a starting node, using a path-based reasoning algorithm, searching in the knowledge graph G for a reasoning path P (i, j) with the starting node vi as the head node and other nodes vj in the knowledge graph G as the tail node; wherein the reasoning path P (i, j) is composed of an association relationship edge set {ep, q|(vp, vq)∈P (i, j)}, and the length len (P (i, j)) of the reasoning path P (i, j) is less than or equal to a preset maximum length threshold M. For each node vi in ​​the target node set T, the path-based reasoning algorithm is executed. Taking node vi as the starting node, search the knowledge graph G for the reasoning path P(i, j) with vi as the head node and other nodes vj as the tail nodes. The reasoning path P(i, j) consists of a series of association edges ep, q, representing a connection path from node vi to node vj. The length of the reasoning path len(P(i, j)) is limited to not exceed the preset maximum length threshold M to avoid the reasoning process being too lengthy. The searched reasoning path P(i, j) reflects the multi-hop semantic association relationship between nodes vi and vj.

[0058] According to the length len(P(i, j) of the reasoning path P(i, j), the semantic relevance weight wp,q of the association edge ep,q, and the importance deg(vj) of the tail node vj, the semantic relevance score S(P(i, j)) of the reasoning path P(i, j) is calculated. According to the length len(P(i, j) of the reasoning path P(i, j), the semantic relevance weight wp,q of the association edge ep,q, and the importance deg(vj) of the tail node vj, the semantic relevance score S(P(i, j)) of the reasoning path P(i, j) is calculated, including: calculating the length score L(P(i, j) of the reasoning path P(i, j), L(P(i, j))=1 / log(1+len(P(i, j))). The length score L(P(i, j)) of the reasoning path is inversely proportional to the path length len(P(i, j)). The logarithmic function log(1+len(P(i, j))) is used to smooth the path length to avoid a low score when the length is too long. The length score L(P(i, j)) reflects the simplicity of the reasoning path. The shorter the path, the more direct the semantic association, and the higher the score.

[0059] Calculate the sum of the semantic relevance weights wp,q of each association edge ep,q in the reasoning path P(i,j) to obtain the semantic relevance score R(P(i,j)) of the reasoning path P(i,j), R(P(i,j))=∑(ep,q)∈P(i,j)(wp,q). The semantic relevance weights wp,q of the association edge ep,q are calculated by the following steps: In the concept node set V of the knowledge graph G, count the frequency f(vi,vj) of any two concept nodes vi and vj appearing together as the co-occurrence frequency between nodes vi and vj. Traverse all concept node pairs (vi,vj) in the knowledge graph G and count their co-occurrence frequency f(vi,vj) in the knowledge graph G. The co-occurrence frequency f(vi,vj) reflects the correlation between two concept nodes in the knowledge graph G. The higher the frequency, the stronger the correlation.

[0060] For each association edge ep,q in the knowledge graph G, obtain the two concept nodes vp and vq connected by it, and use the point wise mutual information (PMI) to calculate the semantic relevance weight wp,q of vp and vq, wp,q = log[f(vp,vq) / (f(vp)*f(vq))]; where f(vp) and f(vq) represent the frequencies of nodes vp and vq appearing in the knowledge graph G, respectively, and f(vp,vq) represents the co-occurrence frequency of nodes vp and vq. For each association edge ep,q in the knowledge graph G, calculate the semantic relevance weight wp,q of the two concept nodes vp and vq connected by it. Use point mutual information (PMI) to measure the relevance of vp and vq. The larger the PMI value, the stronger the relevance of the two nodes. Normalize the semantic relevance weight wp,q to the interval [0, 1] as the final semantic relevance weight of the association edge ep,q. Normalize all the calculated semantic relevance weights wp,q and map them to the interval [0, 1]. Normalization can be achieved using methods such as minimum-maximum normalization or sigmoid function. The normalized semantic relevance weights wp,q represent the relevance strength of the association edge ep,q. The larger the weight, the stronger the relevance. Calculate the importance score I(vj) of the tail node vj of the reasoning path P(i, j), I(vj) = log(1+deg(vj)), where deg(vj) is the degree of the tail node vj in the knowledge graph G, that is, the number of association edges directly connected to the node vj. The importance score I(vj) of the tail node vj is proportional to its degree deg(vj) in the knowledge graph G. Use the logarithmic function log(1+deg(vj)) to smooth the node degree to avoid too high a score when the degree is too large. The importance score I(vj) reflects the centrality and influence of the tail node vj in the knowledge graph G. The larger the degree, the higher the importance.

[0061] The length score L(P(i, j)), the semantic relevance score R(P(i, j)) and the importance score I(vj) are linearly weighted and summed to obtain the semantic relevance score S(P(i, j)) of the reasoning path P(i, j), S(P(i, j)) = α*L(P(i, j)) + β*R(P(i, j)) + γ*I(vj), where α, β, and γ are weight coefficients, and α+β+γ=1. The semantic relevance score S(P(i, j)) of the reasoning path P(i, j) consists of three parts: the length score L(P(i, j)), the semantic relevance score R(P(i, j)) and the importance score I(vj). The weight coefficients α, β, and γ are used to linearly weight the three scores to obtain the final semantic relevance score S(P(i, j)). The weight coefficients α, β, and γ can be adjusted according to actual needs to balance the impact of different factors on the semantic relevance. Select the tail node vj in the reasoning path P(i, j) whose semantic association score S(P(i, j)) is greater than the preset threshold δ as the extended node semantically associated with the starting node vi, and add it to the extended node set E'. For each starting node vi, select the reasoning path P(i, j) whose semantic association score S(P(i, j)) is greater than the preset threshold δ. The tail node vj of the reasoning path P(i, j) that meets the conditions is used as the extended node semantically associated with the starting node vi. Add the extended node vj to the extended node set E' to represent the extended knowledge point semantically associated with the target node set T.

[0062] Take each node in the extended node set E' as a new starting node, and repeat until the extended node set E' no longer adds new nodes or reaches the preset number of iterations. Take the nodes in the extended node set E' as new starting nodes, and recursively perform the knowledge reasoning and link prediction process. Repeat the search in the knowledge graph G for new nodes that are semantically associated with the extended nodes. The iterative process continues until the extended node set E' no longer adds new nodes or reaches the preset number of iterations.

[0063] According to the concept nodes contained in the target node set T and the extended node set E', combined with the pre-configured recommendation rule base R, through pattern matching and filtering sorting, personalized recommended learning content Z={z1, z2,..., zM} for different grid workers is generated. Preferably, according to the concept nodes and entity nodes contained in the target node set T and the extended node set E', combined with the pre-configured recommendation rule base R, through pattern matching and filtering sorting, personalized recommended learning content L for different grid workers is generated, including: according to the types, attributes and associations of the concept nodes and entity nodes contained in the target node set T and the extended node set E', pattern matching is performed in the recommendation rule base R, and the recommendation rules that match the personalized skill evaluation results of the current grid worker are retrieved. A series of recommendation rules are pre-configured in the recommendation rule base R, which are used to generate personalized recommended learning content according to the skill evaluation results of the grid worker. Each recommendation rule defines a specific combination pattern of concept nodes and entity nodes, and corresponding recommended learning content. According to the types, attributes and associations of nodes in the target node set T and the extended node set E', pattern matching is performed in the recommendation rule base R. Retrieve recommendation rules that match the personalized skill evaluation results of the current grid worker as the basis for generating recommended learning content.

[0064] Replace the placeholders in the recommendation rules with specific concept nodes or entity nodes to generate preliminary personalized recommended learning content L'. The recommendation rules may contain placeholders, which means that they can be replaced with specific concept nodes or entity nodes. Replace the nodes in the target node set T and the extended node set E' with the placeholder positions of the recommendation rules. Through placeholder replacement, generate preliminary personalized recommended learning content L', which contains specific knowledge points and learning resources. According to the pre-configured filtering conditions in the recommendation rule base R, filter the preliminary personalized recommended learning content L' to obtain the filtered personalized recommended learning content L'. Some filtering conditions can be pre-configured in the recommendation rule base R to filter the generated recommended learning content. The filtering conditions can be set based on factors such as personal attributes, learning preferences, and knowledge mastered by grid workers. According to the pre-configured sorting strategy in the recommendation rule base R, sort the filtered personalized recommended learning content L' to generate the final personalized recommended learning content L. The recommendation rule base R can be pre-configured with a sorting strategy to prioritize the filtered recommended learning content. The sorting strategy can consider factors such as the importance, difficulty, and matching degree of learning content with the skill level of grid workers. The filtered personalized recommended learning content L' is sorted according to the sorting strategy to generate the final personalized recommended learning content L. The sorted recommended learning content L is presented to grid workers in order of priority from high to low, which is convenient for grid workers to choose and learn. Based on the personalized skill evaluation results of grid workers, personalized recommended learning content is generated in combination with the pre-built grid work knowledge graph and recommendation rule library. The recommended learning content covers knowledge points that are closely related to the current skill level of grid workers, as well as extended knowledge points discovered through knowledge reasoning and link prediction, providing learning paths and resources that are both targeted and comprehensive.

[0065] Iterative update: Obtain the updated skill data X'={X'1, X'2, ..., X'N'} after the grid worker recommends the learning content Z: After the grid worker completes the recommended personalized learning content Z, collect the latest skill data X' of the grid worker. The updated skill data X' can be obtained through various channels, such as learning records, work performance evaluation, skill assessment, etc. of the grid worker learning platform. The updated skill data X' reflects the skill improvement and change of the grid worker after completing the recommended learning content. The collected updated skill data are formed into a set X'={X'1, X'2, ..., X'N'}, where N' represents the number of samples of the updated data. The updated skill data X' is preprocessed as the second training sample set D2={(X'n, Y'n)}, n=1, 2, ..., N': The updated skill data X' is preprocessed and converted into a format that meets the model training requirements. The preprocessing step can include data cleaning, feature extraction, data standardization, etc., which is similar to the preprocessing of step S1. The preprocessed updated skill data X'n is paired with the corresponding skill evaluation label Y'n to form a second training sample set D2. The second training sample set D2 = {(X'n, Y'n)}, n = 1, 2, ..., N', where (X'n, Y'n) represents the feature vector and evaluation label of the nth updated sample.

[0066] The distributed incremental learning method h(D2, f) based on parameter server is adopted to adjust the constructed multidimensional evaluation model f(x, θ) using the second training sample set D2 to obtain the updated multidimensional evaluation model f'(x, θ'): The distributed incremental learning method h(D2, f) based on parameter server is adopted to update the multidimensional evaluation model. The incremental learning method h(D2, f) takes the second training sample set D2 as input and the original multidimensional evaluation model f(x, θ) as the basis to adjust and optimize the model parameters. The parameter server is a distributed machine learning architecture that stores model parameters in a central server and multiple working nodes perform model training and updating in parallel. During the incremental learning process, the working node obtains the latest model parameters from the parameter server, calculates the gradient and updates the parameters using the second training sample set D2, and then sends the updated parameters back to the parameter server. Through the collaboration and parallel computing of multiple working nodes, the incremental learning process of the model can be accelerated and the training efficiency can be improved. The incremental learning method h(D2, f) adjusts the original model f(x, θ) by using the new training sample set D2 to adapt the model to the changes and improvements in the skills of grid workers. After incremental learning, the updated multidimensional evaluation model f'(x, θ') is obtained, where θ' represents the updated model parameters. The updated multidimensional evaluation model f'(x, θ') reflects the changes and improvements in the skills of grid workers after completing the recommended learning content. Through incremental learning, the model can continuously learn and adapt to the dynamic changes in grid workers' skills, providing more accurate and timely personalized evaluation and recommendations. The iterative update process can set the number of iterations or iteration conditions according to actual needs. Each iteration will use the latest skill data to adjust and optimize the model, so that the model can continuously adapt to the changes in grid workers' skills and provide continuous and effective personalized learning recommendation services.

[0067] Figure 5This is an exemplary module diagram of a grid worker skill evaluation system shown in some embodiments of the present specification, a grid worker skill evaluation system, including: a state perception module, obtaining the grid worker's work environment state information s, s includes the grid worker's attribute characteristics x, the grid worker's behavior characteristics a, and the comprehensive performance index o of the grid area; an evaluation strategy learning module, using a deep reinforcement learning algorithm, based on the work environment state information s and immediate reward r obtained by the state perception module, through a policy gradient method or a value function approximation method, to train and obtain a grid worker skill evaluation strategy function π; a skill evaluation module, inputting the real-time work environment state information s obtained by the state perception module into the evaluation strategy function π trained by the evaluation strategy learning module, to obtain the grid worker's personalized skill evaluation result y; a recommendation learning module, using the personalized skill evaluation result y output by the skill evaluation module as a query condition, performing semantic association reasoning and link prediction in a pre-constructed grid work knowledge graph G, recommending and filtering out concept nodes semantically related to the query condition y from the knowledge graph G, and generating personalized recommended learning content z.

Claims

1. A method for evaluating grid worker skills, comprising: Collect grid worker skill data; skill data includes grid worker basic data, work results data and learning data; basic data includes name, gender, age, education and major; work results data includes event reporting data and inspection data, event reporting data includes M1 major category and M2 minor category event types, inspection data includes inspection indicator list; learning data includes daily learning status; The collected skill data is preprocessed as the first training sample set; a machine learning algorithm based on reinforcement learning is used to extract features and train models on the first training sample set, and a multi-dimensional evaluation model for grid workers is constructed; wherein, reinforcement learning uses a reward function to establish a mapping relationship between the evaluation results of grid workers and the reward value, and the cumulative sum of the reward values ​​is used as the optimization target; The collected skill data is used as input, and the personalized skill evaluation results are calculated through the constructed multi-dimensional evaluation model; The obtained personalized skill evaluation results are input into the recommendation algorithm based on the knowledge graph to generate personalized recommended learning content; wherein the recommendation algorithm uses the pre-built grid work knowledge graph, and according to the personalized skill evaluation results, generates personalized recommended learning content according to different grid workers through knowledge reasoning and link prediction; wherein the grid work knowledge graph includes the concepts, entities and associations of grid work; Obtain the updated skill data of the grid workers after they complete the recommended learning content, and pre-process the updated skill data as the second training sample set; adopt a distributed incremental learning method based on a parameter server, use the second training sample set to adjust the constructed multidimensional evaluation model, and obtain an updated multidimensional evaluation model; Construct a multi-dimensional evaluation model for grid workers, including: Define the state space S, which contains the skill data of the grid workers; define the action space A, which contains the evaluation actions of the grid workers' skills; define the reward function R, which is used to establish the mapping relationship between the grid workers' evaluation results and the reward value; Constructing a deep neural network as an evaluation strategy function , randomly initialize the evaluation strategy function Parameters , let the initial state ; According to the current status and the evaluation strategy function parameter , the probability distribution of each evaluation action is calculated through a deep neural network , and according to the probability distribution Select an evaluation action ; Execute evaluation action , get instant rewards , and obtain the new state ; Will Stored as an experience sample in the experience replay pool D; Randomly sample a batch of experience samples from the experience replay pool D and use the immediate reward and the next state The estimated Q value of Compute the current state-action pair The target Q value ; and use the target Q value and estimated Q value The mean square error is used as the loss function , update the evaluation strategy function through the gradient descent algorithm Parameters ; make , repeat the above steps until the preset number of training rounds T or the evaluation strategy function is reached convergence; Output the final evaluation strategy function As a multidimensional evaluation model for grid workers.

2. The grid worker skill evaluation method according to claim 1 is characterized in that: Select an evaluation action ,include: The current state Input to the evaluation strategy function In the corresponding deep neural network, after multiple layers of nonlinear transformation, the state is obtained The feature representation vector ; Represent the feature vector Input to the evaluation strategy function The output layer is converted into the probability distribution of each action in the action space A through the softmax function. ; According to the probability distribution , select an evaluation action through a greedy strategy .

3. The grid worker skill evaluation method according to claim 2 is characterized in that: Select an evaluation action , also includes: Set the entropy regularization term , the probability distribution of the evaluation action The entropy of is added to the loss function as a regularization constraint In the above example, we get a new loss function , ,in is the entropy regularization coefficient; Calculate the new loss function About the evaluation strategy function Parameters Gradient , using the gradient descent algorithm to update the parameters , ,in is the learning rate.

4. The grid worker skill evaluation method according to claim 1, characterized in that: Update the parameters of the evaluation strategy function through the gradient descent algorithm ,include: Randomly sample a batch of experience samples from the experience replay pool D ; For each experience sample , using the target evaluation strategy function Calculate the next state The estimated Q value of ,in For the next state All possible actions of Use Bellman's optimal equation to calculate the current state-action pair The target Q value , , where γ is the discount factor; Using the evaluation strategy function Compute the current state-action pair The estimated Q value of ; Using the target Q value and estimated Q value The mean square error constructs the loss function , ; Calculating the loss function About the evaluation strategy function Parameters Gradient , using the gradient descent algorithm to update the parameters , ,in is the learning rate.

5. The grid worker skill evaluation method according to claim 1, characterized in that: Generate personalized recommended learning content, including: Pre-construct a grid work knowledge graph G, which includes concept nodes, entity nodes, and association edges of grid work; wherein the concept nodes represent the abstract concept of grid work, the entity nodes represent the specific instances of grid work, and the association edges represent the semantic associations between concept nodes and entity nodes and between entity nodes; The obtained personalized skill evaluation results are used as the input query CX, and the target node set T related to the query CX semantics is searched in the knowledge graph G. The target node set T contains concept nodes and entity nodes related to the query CX semantics; Taking the nodes in the target node set T as the starting node, using knowledge reasoning and link prediction algorithms, infer the extended node set semantically associated with the target node set T in the knowledge graph G. , expand the node set Contains concept nodes and entity nodes that are directly or indirectly semantically associated with the target node set T; According to the target node set T and the expanded node set The concept nodes and entity nodes contained in are combined with the pre-configured recommendation rule base R to generate personalized recommended learning content L for different grid workers through pattern matching and filtering sorting.

6. The method for evaluating grid worker skills according to claim 5, characterized in that: Using knowledge reasoning and link prediction algorithms, we can infer the extended node set that is semantically associated with the target node set T in the knowledge graph G. ,include: Taking each node in the target node set T as the starting node, using the path-based reasoning algorithm, search the knowledge graph G for a reasoning path P with the starting node as the head node and other nodes in the knowledge graph G as the tail nodes; wherein the reasoning path P is formed by connecting multiple association edges end to end, and the length of the reasoning path P is less than or equal to the preset maximum length threshold M; Calculate the semantic relevance score S of the reasoning path P according to the length of the reasoning path P, the semantic relevance of the associated relationship edge, and the importance of the tail node; Select the tail node in the reasoning path P whose semantic association score S is greater than the preset threshold as the extension node semantically associated with the target node set T, and add the extension node to the extension node set middle.

7. The grid worker skill evaluation method according to claim 6 is characterized in that: According to the reasoning path Length , association edge The semantic relevance weight , and the tail node Importance , calculate the inference path The semantic relevance score of ,include: Compute inference paths Length score , ; Compute inference paths Each relationship edge The semantic relevance weight The sum of the inference path The semantic relevance score of , ; Compute inference paths The tail node Importance score , ,in, The tail node The degree in the knowledge graph G, that is, the degree of the node The number of directly connected relationship edges; Score the length , semantic relevance score and importance score Linear weighted summation to obtain the inference path The semantic relevance score of , ,in is the weight coefficient.

8. The grid worker skill evaluation method according to claim 7 is characterized in that: Semantic relevance weight Obtained by counting co-occurrence frequencies.

Citation Information

Patent Citations

  • Community grid member automatic distribution method and system and application thereof

    CN117540903A

  • Detection method, device and apparatus for decision engine

    CN113779263A

  • Teaching activity digitization system

    CN114819523A