Virtual workplace multi-dimensional assessment method and system based on deep reinforcement learning

By collecting behavioral data in the virtual workplace and using deep reinforcement learning to build an evaluation system, the subjectivity and one-sided problems of traditional evaluation systems are solved, accurate assessment and dynamic reflection of employee abilities are achieved, and scientific career development guidance is provided.

CN120373970AActive Publication Date: 2025-07-25FLASH TURING (HANGZHOU) TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510857522.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Traditional virtual workplace evaluation systems cannot accurately quantify employees' performance in complex task scenarios, and are subjective and one-sided, and lack in-depth analysis and modeling of employee decision-making behaviors, making it difficult to evaluate employees' adaptability and potential in different task situations.

Method used

By collecting users' behavior data in the virtual workplace environment, building a state transition probability matrix, generating user behavior feature vectors, combining deep reinforcement learning to generate interactive decision sequences, building a task execution planning diagram, and performing feature fusion during task execution to calculate professional competency scores.

Benefits of technology

It has achieved an objective and comprehensive assessment of employees' professional abilities, improved the scientificity and fairness of the assessment, dynamically reflected the user's execution ability and adaptability, and provided targeted guidance for career development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373970A_ABST
    Figure CN120373970A_ABST
Patent Text Reader

Abstract

The invention provides a virtual workplace multi-dimensional assessment method and system based on deep reinforcement learning, and relates to the technical field of human resource assessment, and the method comprises the steps: collecting the behavior data of a user in a virtual workplace environment, constructing a state transition probability matrix, generating an interaction decision sequence based on deep reinforcement learning, and obtaining a state transition probability matrix; and finally, according to the check point data in the actual execution process of the user, feature fusion is carried out to calculate the occupational competency score. According to the invention, objective quantitative evaluation of the workplace capability is realized, and the accuracy and fairness of assessment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to human resource assessment technologies, and in particular to a multi-dimensional assessment method and system for a virtual workplace based on deep reinforcement learning. Background Art

[0002] With the in-depth development of the digital transformation of enterprises, the virtual workplace environment has gradually become an important platform for employee training and assessment. Traditional vocational ability assessment methods mainly rely on manual observation and empirical judgment, making it difficult to accurately quantify the performance of employees in complex task scenarios, and the assessment results often have subjectivity and one-sidedness.

[0003] Existing virtual workplace assessment systems usually adopt a fixed assessment index system, which cannot dynamically adjust the assessment strategy according to the behavioral characteristics and ability development of employees. At the same time, due to the lack of in-depth analysis and modeling of employees' decision-making behaviors, it is difficult to accurately assess the adaptability and potential of employees in different task scenarios.

[0004] Therefore, there is an urgent need for a multi-dimensional assessment method for a virtual workplace based on deep reinforcement learning. By collecting and analyzing user behavior data to construct a state transition model, combined with task scenarios to dynamically generate assessment strategies, it realizes the accurate assessment and development guidance of employees' vocational abilities. Summary of the Invention

[0005] Embodiments of the present invention provide a multi-dimensional assessment method and system for a virtual workplace based on deep reinforcement learning, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiments of the present invention, a multi-dimensional assessment method for a virtual workplace based on deep reinforcement learning is provided, including: Collecting the behavior data of users in the virtual workplace environment, performing time series feature extraction, calculating the behavior state sequence of users in different task stages, constructing a state transition probability matrix based on the behavior state sequence, analyzing the stability characteristics of the user behavior pattern according to the state transition probability matrix, and generating a user behavior feature vector; Constructing a state space representation based on the user behavior feature vector, calculating a state transition function in combination with the task scenario state information, setting a reward function according to the task completion degree and execution efficiency, and generating an interaction decision sequence by using the policy iteration method of deep reinforcement learning; Constructing a task execution plan graph according to the interaction decision sequence, and generating a task execution plan with checkpoints based on the task execution plan graph; During the process of a user executing a task execution plan, collect the execution status data of each checkpoint, calculate the task completion trajectory curve based on the execution status data, extract the decision-making choices of the user at each checkpoint in combination with the task execution planning diagram, perform feature fusion on the decision-making choices and the user behavior feature vector to generate an ability dimension vector, calculate the professional competency score based on the ability dimension vector, and generate a professional ability assessment report.

[0007] In an alternative embodiment, Collect the behavior data of the user in the virtual workplace environment, perform time-series feature extraction, calculate the behavior state sequence of the user in different task stages, construct a state transition probability matrix based on the behavior state sequence, analyze the stability characteristics of the user behavior pattern according to the state transition probability matrix, and generate a user behavior feature vector including: Collect the behavior data of the user in the virtual workplace environment, where the behavior data includes operation sequence data, task switching data, problem-solving data, and collaboration mode data; Calculate the behavior complexity index for the behavior data, dynamically determine the time window size according to the behavior complexity index, extract the statistical features and trend features of the behavior data within the time window, and combine them to generate a behavior state vector; Calculate the time-series correlation degree for the behavior state vector, determine the behavior jump threshold according to the time-series correlation degree, and perform density clustering on the behavior state vector based on the behavior jump threshold to generate a behavior state sequence; Calculate the transition frequency between states according to the behavior state sequence, construct a state transition probability matrix, perform matrix decomposition on the state transition probability matrix to obtain a first transition probability sub-matrix and a second transition probability sub-matrix; perform eigenvalue decomposition on the first transition probability sub-matrix and the second transition probability sub-matrix respectively to obtain an eigenvalue set, calculate the behavior entropy value according to the largest eigenvalue in the eigenvalue set, calculate the behavior stability characteristics according to the remaining eigenvalues, and perform feature fusion on the behavior entropy value and the behavior stability characteristics to generate a user behavior feature vector.

[0008] In an alternative embodiment, Calculate the time-series correlation degree for the behavior state vector, determine the behavior jump threshold according to the time-series correlation degree, and perform density clustering on the behavior state vector based on the behavior jump threshold to generate a behavior state sequence including: Obtain the behavior state vector, calculate the vector space distance and time interval distance between adjacent behavior state vectors, combine the vector space distance and the time interval distance to construct a time-series correlation function, and calculate the time-series correlation degree of the behavior state vector through the time-series correlation function; Statistically analyze the temporal correlation degree, obtain the distribution characteristics of the temporal correlation degree, generate a behavior jump threshold according to the distribution characteristics, determine the behavior jump point according to the temporal correlation degree and the behavior jump threshold, and calculate the local density of the behavior state vector at the behavior jump point; Perform density clustering on the behavior state vectors based on the local density to obtain a behavior state clustering result, and arrange and combine the behavior state clustering results in chronological order to generate a behavior state sequence.

[0009] In an alternative embodiment, Construct a state space representation based on the user behavior feature vector, calculate the state transition function in combination with the task scenario state information, set the reward function according to the task completion degree and execution efficiency, and generate an interaction decision sequence by using the policy iteration method of deep reinforcement learning, including: Obtain the task scenario state information, where the task scenario state information includes task difficulty information, resource status information, and environmental constraint information; Construct a state space representation based on the user behavior feature vector and the task scenario state information, calculate the transition probability between states based on the state space representation, and construct a state transition function in combination with the importance weights of the historical state sequence. The state transition function includes the state evolution law and state transition constraints; Monitor the task execution process in real time, calculate the task completion degree index based on the task objective completion rate, task quality score, and task timeliness, calculate the execution efficiency index based on the computing resource utilization rate, storage resource occupancy rate, and operation time utilization rate, and construct a reward function by weighted combination of the task completion degree index and the execution efficiency index; Input the state space representation into a pre-trained deep reinforcement learning model, calculate the state value based on the state transition function, and perform policy iteration optimization according to the reward function and the state value to generate an interaction decision sequence.

[0010] In an alternative embodiment, Input the state space representation into a pre-trained deep reinforcement learning model, calculate the state value based on the state transition function, and perform policy iteration optimization according to the reward function and the state value to generate an interaction decision sequence, including: Obtain historical interaction data, extract the state space representation, action sequence, and reward value from the historical interaction data, construct a training sample set, and pre-train the deep reinforcement learning model based on the training sample set to obtain the policy network parameters and value network parameters; Input the current state space representation into the policy network, generate an action selection probability based on the policy network parameters, sample the current state based on the action selection probability, and obtain a candidate action sequence; Input the state space representation and the candidate action sequence into the value network, calculate the state-action value based on the value network parameters, calculate the state transition probability according to the state transition function, and combine the state-action value and the state transition probability to obtain the expected cumulative value; Calculate the immediate reward of the candidate action sequence according to the reward function, and combine the immediate reward and the expected cumulative value to generate an action evaluation value; Use a recurrent unit to encode the historical state sequence to generate temporal correlation features, combine the temporal correlation features with the action evaluation value to calculate the policy update gradient, iteratively optimize the policy network parameters based on the policy update gradient, and generate an interaction decision sequence according to the optimized policy network parameters and the action evaluation value.

[0011] In an alternative embodiment, Construct a task execution planning graph according to the interaction decision sequence, and generate a task execution plan with checkpoints based on the task execution planning graph, including: Extract the state information and decision actions from the interaction decision sequence, calculate the state transition frequency of the state information under the action of the decision action, and calculate the state transition probability according to the state transition frequency; Construct the nodes of the task execution planning graph based on the state information, construct the edges of the task execution planning graph based on the state transition probability, calculate the set of subsequent nodes that each node can reach and the transition cost, count the number of times each state node in the interaction decision sequence is visited and the number of times each decision action is selected, and calculate the access weight of the node and the selection weight of the decision action in combination with the state transition probability; Construct a path evaluation function based on the access weight, selection weight, and transition cost, use the path evaluation function to solve the optimal execution path in the task execution planning graph, calculate the change rate of the access weight and the change rate of the state transition probability of each node on the optimal execution path, and use the weighted sum of the access weight change rate and the state transition probability change rate as the node importance; Select the node with the locally maximum importance on the optimal execution path as the checkpoint, extract the state information, optional decision actions, and access weight of the node corresponding to the checkpoint to generate checkpoint configuration information, and combine the optimal execution path, checkpoint configuration information, and the state transition probability between each node to generate a task execution plan.

[0012] In an alternative embodiment, Collect the execution status data of each checkpoint, calculate the task completion trajectory curve according to the execution status data, extract the decision choices of the user at each checkpoint in combination with the task execution planning graph, fuse the decision choices with the user behavior feature vector to generate an ability dimension vector, calculate the professional competence score according to the ability dimension vector, and generate a professional ability evaluation report, including: During the process of a user executing a task execution plan, collect the execution status data of checkpoints, where the execution status data includes current state features, selected decision actions, execution results, and timestamp information; Calculate the state deviation by calculating the difference between the state features in the execution status data and the preset states of the task execution planning graph, and calculate the action deviation by calculating the difference between the decision actions in the execution status data and the recommended actions of the task execution planning graph; Calculate the checkpoint completion score based on the weighted sum of the state deviation and the action deviation, calculate the execution efficiency score based on the change in execution results between adjacent timestamps, and combine the completion score and the execution efficiency score with weights to obtain the checkpoint score; Interpolate and fit the scores of each checkpoint and the corresponding timestamps to generate a task completion trajectory curve, and combine the preset states and recommended actions recorded in the task execution planning graph to extract the decision-making choices of the user at each checkpoint; Fuse the decision-making choices and the user behavior feature vector in the corresponding dimensions to generate an ability dimension vector, calculate the normalized scores of the feature values in each dimension of the ability dimension vector, and use the normalized scores as the professional competence scores to generate a professional ability assessment report.

[0013] In the second aspect of the embodiments of the present invention, a virtual workplace multi-dimensional assessment system based on deep reinforcement learning is provided, including: A first unit for collecting the behavior data of a user in a virtual workplace environment, performing temporal feature extraction, calculating the behavior state sequence of the user in different task stages, constructing a state transition probability matrix based on the behavior state sequence, analyzing the stability characteristics of the user behavior pattern according to the state transition probability matrix, and generating a user behavior feature vector; A second unit for constructing a state space representation based on the user behavior feature vector, calculating a state transition function in combination with the task scenario state information, setting a reward function according to the task completion degree and execution efficiency, and generating an interaction decision sequence by using the policy iteration method of deep reinforcement learning; A third unit for constructing a task execution planning graph according to the interaction decision sequence and generating a task execution plan with checkpoints based on the task execution planning graph; A third unit for, during the process of a user executing a task execution plan, collecting the execution status data of each checkpoint, calculating a task completion trajectory curve according to the execution status data, and combining the task execution planning graph to extract the decision-making choices of the user at each checkpoint, fusing the decision-making choices with the user behavior feature vector to generate an ability dimension vector, and calculating the professional competence score according to the ability dimension vector to generate a professional ability assessment report.

[0014] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] In this embodiment, by collecting user behavior data in a virtual workplace environment and extracting time-series features, and combining deep reinforcement learning technology to construct a decision sequence, it is possible to objectively and comprehensively evaluate the performance of users in the workplace environment, avoid subjective biases in traditional assessment methods, and improve the scientificity and fairness of the assessment. Through a task execution scheme with checkpoints, refined monitoring of each key node in the process of task execution by users is achieved. By calculating the task completion trajectory curve, it is possible to dynamically reflect the execution ability and adaptability of users, providing a richer and more reliable data basis for vocational ability assessment. By fusing the user's decision-making choices with the behavioral feature vectors, a multi-dimensional ability dimension vector is generated, making the assessment results more three-dimensional and comprehensive, capable of accurately identifying the advantages and disadvantages of users in different vocational ability dimensions, providing targeted guidance and suggestions for the career development of users, and at the same time providing a scientific basis for the enterprise's talent selection and cultivation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic flowchart of a virtual workplace multi-dimensional assessment method based on deep reinforcement learning according to an embodiment of the present invention; Figure 2 A schematic diagram of the convergence curve of the reward value of the deep reinforcement learning strategy iteration; Figure 3 It is a schematic diagram of task execution planning. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and for the same or similar concepts or processes, they may not be repeated in some embodiments.

[0020] Figure 1 This is a schematic flowchart of the multi-dimensional assessment method for a virtual workplace based on deep reinforcement learning according to an embodiment of the present invention. As Figure 1 shown, the method includes: Collect the behavioral data of the user in the virtual workplace environment, perform temporal feature extraction, calculate the behavioral state sequence of the user in different task stages, construct a state transition probability matrix based on the behavioral state sequence, analyze the stability characteristics of the user's behavioral pattern according to the state transition probability matrix, and generate a user behavioral feature vector; Construct a state space representation based on the user behavioral feature vector, calculate a state transition function in combination with the task scenario state information, set a reward function according to the task completion degree and execution efficiency, and generate an interaction decision sequence by using the policy iteration method of deep reinforcement learning; Construct a task execution plan graph according to the interaction decision sequence, and generate a task execution plan with checkpoints based on the task execution plan graph; During the process of the user executing the task execution plan, collect the execution status data of each checkpoint, calculate the task completion trajectory curve according to the execution status data, extract the decision-making choices of the user at each checkpoint in combination with the task execution plan graph, perform feature fusion on the decision-making choices and the user behavioral feature vector to generate an ability dimension vector, calculate the professional competence score according to the ability dimension vector, and generate a professional ability assessment report.

[0021] In an alternative embodiment, collecting the behavioral data of the user in the virtual workplace environment, performing temporal feature extraction, calculating the behavioral state sequence of the user in different task stages, constructing a state transition probability matrix based on the behavioral state sequence, and analyzing the stability characteristics of the user's behavioral pattern according to the state transition probability matrix to generate a user behavioral feature vector includes: Collect the behavioral data of the user in the virtual workplace environment, where the behavioral data includes operation sequence data, task switching data, problem-solving data, and collaboration mode data; Calculate a behavioral complexity index for the behavioral data, dynamically determine the time window size according to the behavioral complexity index, extract the statistical features and trend features of the behavioral data within the time window, and combine them to generate a behavioral state vector; Calculate the temporal correlation degree of the behavioral state vector, determine a behavioral jump threshold according to the temporal correlation degree, and perform density clustering on the behavioral state vector based on the behavioral jump threshold to generate a behavioral state sequence; Calculate the transition frequency between states according to the behavior state sequence, construct a state transition probability matrix, perform matrix decomposition on the state transition probability matrix to obtain a first transition probability sub-matrix and a second transition probability sub-matrix; respectively perform eigenvalue decomposition on the first transition probability sub-matrix and the second transition probability sub-matrix to obtain an eigenvalue set, calculate the behavior entropy value according to the largest eigenvalue in the eigenvalue set, calculate the behavior stability feature according to the remaining eigenvalues, and fuse the behavior entropy value and the behavior stability feature to generate a user behavior feature vector.

[0022] Exemplarily, when collecting user behavior data in a virtual workplace environment, the operation sequence data, task switching data, problem-solving data, and collaboration mode data of the user will be recorded. The operation sequence data includes basic operations such as clicks, drags, inputs of the user and their timestamps; the task switching data records the switching frequency and switching mode of the user between different tasks; the problem-solving data includes the path selection and time consumption of the user to solve problems; the collaboration mode data records the interaction behaviors of the user with other virtual characters or real users. For example, when a user processes a document task in a virtual office environment, an operation sequence such as "open document A - read for 10 minutes - modify content - send a message to colleague B - switch to task C" will be recorded.

[0023] Calculate the behavior complexity index for the collected behavior data to dynamically determine the time window size. The behavior complexity index is comprehensively calculated through the operation type diversity, operation frequency change rate, task switching frequency, and collaboration interaction complexity. In a specific implementation, if the operation types of the user change frequently during a certain period, the complexity will be rated as high, and the time window will be correspondingly reduced to 30 seconds; if the operation types are single and stable, the complexity will be rated as low, and the time window will be expanded to 120 seconds. Within the determined time window, extract the statistical features and trend features of the behavior data. The statistical features include operation frequency, task duration, switching frequency, etc.; the trend features include operation acceleration, efficiency change trend, etc. Taking the virtual meeting scenario as an example, extract that the number of speeches of the user within a 60-second window is 5 times, the average duration of each speech is 20 seconds, and the problem-posing frequency is 3 times per 10 minutes. These statistical values and their change trends together form a behavior state vector.

[0024] Calculate the temporal correlation of the generated behavior state vectors, that is, analyze the similarity of behavior states within adjacent time windows. Determine the temporal correlation by calculating the distance values between adjacent state vectors, and then determine the behavior jump threshold. For example, for a certain virtual collaboration task, it is found through analysis that the distance values between state vectors are concentrated in the range of 0.1 - 0.3, but occasionally jump above 0.7. Then the behavior jump threshold can be set to 0.5. Based on this threshold, perform density clustering on the behavior state vectors, and group the state vectors with close distances into the same class, thereby generating a behavior state sequence. Taking a user completing a virtual project management task as an example, its behavior state sequence may be expressed as a temporal arrangement of discrete states such as "planning state - execution state - coordination state - inspection state - execution state - planning state".

[0025] According to the generated behavior state sequence, calculate the transition frequencies between states and construct a state transition probability matrix. Each element in this matrix represents the probability of transitioning from one state to another. For example, in a scenario with four states, in the constructed 4×4 matrix, the element M(1, 2) = 0.35 represents the probability of transitioning from state 1 to state 2 is 35%. Decompose the state transition probability matrix to obtain the first transition probability sub - matrix and the second transition probability sub - matrix. The first sub - matrix reflects the state transition characteristics of the user in the stable working mode, and the second sub - matrix reflects the state transition characteristics in abnormal or stressful situations.

[0026] Perform eigenvalue decomposition on the two transition probability sub - matrices respectively to obtain eigenvalue sets. For the first sub - matrix, the eigenvalue set {0.92, 0.45, 0.28, 0.15} may be obtained; for the second sub - matrix, the eigenvalue set {0.75, 0.62, 0.41, 0.22} may be obtained. Calculate the behavior entropy value according to the maximum eigenvalue in the eigenvalue set (such as 0.92 and 0.75). This entropy value represents the predictability of the user's behavior. The lower the entropy value, the more stable the user's behavior pattern. For example, an entropy value of 0.25 represents a highly predictable behavior pattern; the higher the entropy value, such as 0.85, the stronger the randomness of the user's behavior. Calculate the behavior stability characteristics according to the remaining eigenvalues (such as {0.45, 0.28, 0.15} and {0.62, 0.41, 0.22}), including state transition stability, state persistence, and adaptability indicators.

[0027] Fuse the behavior entropy value and the behavior stability characteristics to generate a user behavior feature vector. This feature vector can be expressed as [behavior entropy value, state transition stability, state persistence, adaptability indicator], for example, [0.35, 0.78, 0.62, 0.45]. This feature vector can be used to evaluate the stability of the user's behavior pattern in the virtual workplace environment, and is further applied to career ability assessment, personalized training program design, or team collaboration optimization.

[0028] Based on the above technical solution, it is possible to deeply model and stably analyze the behaviors of users in a virtual workplace environment, thereby accurately characterizing the behavioral characteristics of users in the processes of multitasking, problem-solving, and collaboration. By dynamically adjusting the time window to extract key behavioral characteristics, and combining the calculation of the state transition probability matrix and the behavioral entropy value, a quantitative evaluation of the degree of change and stability of the behavioral pattern is achieved, which helps to distinguish the differences among different users in terms of cognitive load, task concentration, and behavioral consistency, and further provides support for personalized training, behavioral risk warning, and intelligent assisted decision-making.

[0029] In an optional implementation manner, calculate the temporal correlation degree of the behavior state vector, determine the behavior jump threshold according to the temporal correlation degree, and perform density clustering on the behavior state vector based on the behavior jump threshold to generate a behavior state sequence, including: Obtain the behavior state vector, calculate the vector space distance and the time interval distance between adjacent behavior state vectors, combine the vector space distance and the time interval distance to construct a temporal correlation function, and calculate the temporal correlation degree of the behavior state vector through the temporal correlation function; Conduct statistical analysis on the temporal correlation degree, obtain the distribution characteristics of the temporal correlation degree, generate the behavior jump threshold according to the distribution characteristics, determine the behavior jump point according to the temporal correlation degree and the behavior jump threshold, and calculate the local density of the behavior state vector at the behavior jump point; Perform density clustering on the behavior state vector based on the local density to obtain the behavior state clustering result, and arrange and combine the behavior state clustering result in time sequence to generate a behavior state sequence.

[0030] Exemplarily, calculate the vector space distance and the time interval distance between adjacent behavior state vectors. The vector space distance represents the degree of difference between two behavior state vectors in the feature space, and the Euclidean distance can be used for calculation. For example, if two behavior state vectors are [0.8, 0.6, 0.9, 0.7] and [0.7, 0.6, 0.8, 0.5] respectively, then their Euclidean distance is 0.22. The time interval distance represents the difference between the acquisition time points of two behavior state vectors, and minutes or hours can be used as the unit. For example, if the acquisition times of two behavior state vectors are 10:00 and 10:15 respectively, then the time interval is 15 minutes.

[0031] Construct a time-series correlation function by combining the vector space distance and the time interval distance. The time-series correlation function is used to measure the time-series correlation between behavior state vectors, taking into account two factors: the vector space distance and the time interval. A function that comprehensively considers the two factors can be designed. For example, when the vector space distance is small and the time interval is short, the time-series correlation degree is high; when the vector space distance is large or the time interval is long, the time-series correlation degree is low. In specific implementation, the vector space distance can be divided by the time interval to obtain a ratio. The smaller the ratio, the higher the time-series correlation degree. For example, if the vector space distance is 0.22 and the time interval is 15 minutes, the ratio is 0.0147, indicating a high time-series correlation degree between these two behavior states.

[0032] Calculate the time-series correlation degree of behavior state vectors through the time-series correlation function. For all pairs of adjacent behavior state vectors, calculate their time-series correlation degrees to form a time-series correlation degree sequence. For example, within a working day, the system collects behavior state vectors every 15 minutes, and a total of 32 vectors are collected, then 31 time-series correlation degree values will be obtained.

[0033] Conduct statistical analysis on the time-series correlation degree to obtain the distribution characteristics of the time-series correlation degree. The distribution characteristics can include statistical quantities such as mean, variance, and quantiles. For example, for the above 31 time-series correlation degree values, the calculated mean is 0.02, the standard deviation is 0.015, the 25% quantile is 0.01, and the 75% quantile is 0.03. These distribution characteristics reflect the regularity of behavior state changes.

[0034] Generate a behavior jump threshold based on the distribution characteristics. The behavior jump threshold is used to determine whether a significant change has occurred in the behavior state. An appropriate threshold can be selected based on the distribution characteristics of the time-series correlation degree. For example, the 75% quantile 0.03 of the time-series correlation degree can be selected as the behavior jump threshold, that is, when the time-series correlation degree between adjacent behavior state vectors is greater than 0.03, it is considered that a behavior jump has occurred.

[0035] Determine the behavior jump points according to the time-series correlation degree and the behavior jump threshold. The behavior jump points refer to the time points when significant changes occur in the behavior state. For example, if the time-series correlation degree between the 5th and 6th behavior state vectors is 0.035, exceeding the behavior jump threshold of 0.03, then the time point corresponding to the 6th behavior state vector is marked as a behavior jump point. Calculate the local density of the behavior state vector at the behavior jump point. The local density represents the aggregation degree of behavior state vectors in a certain region of the behavior state vector space. The local density can be estimated by calculating the distance between a certain behavior state vector and its surrounding behavior state vectors. For example, for the 6th behavior state vector, calculate its average distance from the 3 behavior state vectors before and after it. The smaller the distance, the higher the local density. Suppose the calculated local density of the 6th behavior state vector is 0.8.

[0036] Perform density clustering on the behavior state vectors based on local density. Density clustering is a density-based clustering algorithm that divides regions with similar and continuous density into one cluster. The system can use density clustering algorithms such as DBSCAN or OPTICS to perform clustering based on the local density of the behavior state vectors. For example, for 32 behavior state vectors, 5 clustering results may be obtained through density clustering, corresponding to different behavior state types respectively. Arrange and combine the behavior state clustering results in chronological order to generate a behavior state sequence. The behavior state sequence refers to the change process of an employee's behavior state over a period of time. For example, if 32 behavior state vectors are clustered into 5 categories, the possible behavior state sequence may be "Category 1 (Vectors 1 - 5) -> Category 2 (Vectors 6 - 12) -> Category 3 (Vectors 13 - 18) -> Category 4 (Vectors 19 - 25) -> Category 5 (Vectors 26 - 32)". This sequence reflects the change of the employee's behavior state at different time periods.

[0037] In practical applications, this method can be used for the analysis and assessment of employees' behaviors in a virtual workplace environment. For example, for a salesperson, the system can collect the behavior state vectors of the salesperson in a day, including multiple dimensions such as the number of customer communications, the number of product demonstrations, and the transaction amount. After generating the behavior state sequence through the above method, different working states of the salesperson, such as the preparation stage, the communication stage, the transaction stage, etc., can be identified, so as to conduct targeted evaluation and assessment.

[0038] Based on the above technical solution, it is possible to achieve fine-grained recognition and dynamic segmentation of the user's behavior state in the time dimension, thereby effectively capturing the mutation points and stable patterns of the behavior state. By introducing a time series correlation function to comprehensively consider the content similarity and time continuity of the behavior state, the accuracy of jump detection is improved. Clustering based on the local density of the jump points can avoid misjudgment of abnormal behaviors or stage transitions by traditional time series partitioning methods, and generate a state sequence that more conforms to the actual behavior evolution law. This provides a data basis with clear structure and distinct levels for subsequent behavior modeling, stage recognition, and behavior prediction, and improves the system's perception ability of complex behavior patterns.

[0039] In an alternative embodiment, a state space representation is constructed based on the user behavior feature vector, a state transition function is calculated in combination with the task scenario state information, a reward function is set according to the task completion degree and execution efficiency, and the strategy iteration method of deep reinforcement learning is used to generate an interaction decision sequence, including: Obtain the task scenario state information, where the task scenario state information includes task difficulty information, resource state information, and environmental constraint information; Construct a state space representation based on the user behavior feature vector and the task scenario status information, calculate the transition probability between states based on the state space representation, and construct a state transition function by combining the importance weights of the historical state sequence. The state transition function includes the state evolution law and the state transition constraint; Monitor the task execution process in real time, calculate the task completion index based on the task objective completion rate, task quality score, and task timeliness, calculate the execution efficiency index based on the computing resource utilization rate, storage resource occupancy rate, and operation time utilization rate, and construct a reward function by weighted combination of the task completion index and the execution efficiency index; Input the state space representation into a pre-trained deep reinforcement learning model, calculate the state value based on the state transition function, and perform policy iteration optimization according to the reward function and the state value to generate an interaction decision sequence.

[0040] In this embodiment, in the stage of obtaining the task scenario status information, task difficulty information, resource status information, and environmental constraint information are collected. The task difficulty information includes the task complexity score (1 - 10 points), the required skill level (beginner / intermediate / advanced), and the estimated completion time (hours); the resource status information includes the available computing resource percentage (0% - 100%), the remaining storage space (GB), and the bandwidth usage rate (%); the environmental constraint information includes the maximum allowed running time (hours), the power consumption limit (watts), and the noise control requirement (decibels). For example, the task scenario status information of a data processing task may be: the task complexity score is 8 points, the intermediate skill level is required, and the estimated completion time is 4 hours; the available computing resource is 75%, the remaining storage space is 500GB, and the bandwidth usage rate is 30%; the maximum allowed running time is 6 hours, the power consumption limit is 100W, and the noise control requirement is not more than 40 decibels.

[0041] In the stage of constructing the state space representation, the user behavior feature vector is fused with the task scenario state information. The user behavior feature vector includes dimensions such as operation frequency, interaction mode, and preference settings. Specifically, the operation frequency records the number of operations per minute of the user; the interaction mode includes types such as mouse clicks, keyboard inputs, touch screen operations, etc. and their proportions; the preference settings include interface layout, response speed, and the usage of assistive functions. The system converts these heterogeneous data into a unified vector representation through the feature embedding method, with a dimension of 128. The transition probability between states is calculated by analyzing historical interaction data. The transition probability from state i to state j is equal to the number of times directly transferred from state i to state j in the historical records divided by the total number of transfers starting from state i. The importance weights of the historical state sequence are calculated in a time-decaying manner. The more recent states are given higher weights. For example, the weight of the state within the last 1 hour is 1.0, the weight of the state within 1 - 3 hours is 0.8, the weight of the state within 3 - 12 hours is 0.5, the weight of the state within 12 - 24 hours is 0.3, and the weight of the state over 24 hours is 0.1. The construction of the state transition function combines the state transition probability and the importance weights, and at the same time introduces state transition constraint conditions, such as the resource utilization rate cannot exceed 95%, and the response time cannot exceed 200 milliseconds, etc.

[0042] In the stage of task execution monitoring, the task completion index and the execution efficiency index are calculated in real time. The task completion index is calculated based on the following parameters: task objective completion rate (number of completed subtasks / total number of subtasks × 100%), task quality score (1 - 10 points, evaluated through preset quality checkpoints), and task timeliness (actual time consumption / estimated time consumption × 100%). The execution efficiency index is calculated based on the following parameters: computing resource utilization rate (percentage of actual used CPU / GPU), storage resource occupancy rate (used storage space / allocated storage space × 100%), and operation time utilization rate (effective operation time / total operation time × 100%). The reward function is constructed through a weighted combination method, with the weight of the task completion index being 0.6 and the weight of the execution efficiency index being 0.4. For example, when the task objective completion rate is 85%, the task quality score is 7 points, and the task timeliness is 90%, the task completion index is 0.85×0.4 + 7 / 10×0.3 + (1 - 0.9)×0.3 = 0.64; when the computing resource utilization rate is 70%, the storage resource occupancy rate is 60%, and the operation time utilization rate is 80%, the execution efficiency index is 0.7×0.4 + 0.6×0.3 + 0.8×0.3 = 0.7; the final reward value is 0.64×0.6 + 0.7×0.4 = 0.664.

[0043] In the stage of generating the interactive decision-making sequence, a deep reinforcement learning model is adopted for policy iteration. The deep reinforcement learning model consists of a value network and a policy network. The value network is used to evaluate the state value, and the policy network is used to generate the action probability distribution. The value network contains 3 fully connected layers, with the number of nodes in each layer being 256, 128, and 64 respectively, and the ReLU activation function is used; the policy network contains 4 fully connected layers, with the number of nodes in each layer being 256, 128, 128, and the dimension of the action space respectively. The ReLU activation function is used in the first three layers, and the Softmax function is used in the last layer to output the action probability. The policy iteration optimization adopts the following steps: input the current state into the value network to obtain the state value estimation; input the current state into the policy network to obtain the probability distribution of each possible action; sample and select an action according to the probability distribution; execute the action and observe the environmental feedback to obtain the reward and the next state; update the parameters of the value network and the policy network according to the reward and the next state. The network parameter update uses the Adam optimizer, and the learning rate is set to 0.0003. After 10,000 rounds of iterative training, the model can generate an interactive decision-making sequence that adapts to the user behavior characteristics and task scenario requirements. For example, for a document editing task, the generated interactive decision-making sequence includes: first open the relevant documents edited recently (to improve relevance), recommend applicable document templates (to reduce operation time), provide intelligent completion during the user input process (to improve efficiency), prompt to save at the appropriate time (to avoid data loss), and recommend relevant reference materials according to the user's edited content (to improve quality).

[0044] In this embodiment, by fusing user behavior characteristics and task scenario state information, a dynamic state space is constructed, and combined with the deep reinforcement learning policy iteration method, the decision-making efficiency and personalized adaptation ability in complex interactive scenarios are effectively improved. In the prior art, the interactive policy is mostly formulated based on preset rules or static behavior models, and it cannot be flexibly adjusted according to the individual differences of users and the dynamic changes of the task environment, which easily leads to unreasonable resource allocation, deviation of task execution from the goal, or degradation of the user experience. This solution takes the user behavior feature vector as the core, constructs the state space by combining the task difficulty, resource status, and environmental constraints, enhances the comprehensiveness and pertinence of the state representation, and at the same time constructs the state transition function by introducing the weights of the historical state sequence to realize the reasonable modeling of the state evolution trend. In terms of the reward function design, by integrating the task completion degree and execution efficiency, it breaks the evaluation mode based only on the result feedback in traditional reinforcement learning, and can better reflect the multi-dimensional performance during the task execution process. Finally, through the policy iteration optimization method of deep reinforcement learning, the system can dynamically adjust the interactive policy to achieve adaptive optimization in different user and task scenarios, effectively improving the intelligence level and decision-making quality of the interactive system.

[0045] Figure 2Schematic diagram of the convergence curve of the reward value in deep reinforcement learning policy iteration. The main curve shows the changing trend of the reward function value with the number of iterations. It can be seen that as the training progresses, the reward value gradually increases and finally converges to a stable value. In the experiment, a deep reinforcement learning architecture combining a value network and a policy network was adopted. The value network (a 3-layer fully connected layer with the number of nodes being 256, 128, and 64 respectively) was used to evaluate the state value, and the policy network (a 4-layer fully connected layer with the number of nodes being 256, 128, 128, and the dimension of the action space respectively) was used to generate the action probability distribution. The Adam optimizer was used during the training process, and the learning rate was 0.0003.

[0046] In this experiment, the state space representation was constructed by fusing the user behavior feature vector (with a dimension of 128) and the task scenario state information (including task difficulty, resource status, and environmental constraints), and the state evolution law was modeled through the state transition function. The reward function design comprehensively considered the task completion degree index (weight 0.6) and the execution efficiency index (weight 0.4). Three key stages can be seen in the figure: the initial learning stage (the first 400 iterations), the starting point of policy convergence (about 600 iterations), and the stable convergence stage (after 850 iterations). The final model achieved a reward value of 0.664 in the test task, where the task completion degree index was 0.64 and the execution efficiency index was 0.70, showing good task execution ability and resource utilization efficiency.

[0047] In an optional implementation manner, the state space representation is input into a pre-trained deep reinforcement learning model, the state value is calculated based on the state transition function, and policy iteration optimization is performed according to the reward function and the state value. The generated interactive decision sequence includes: Obtain historical interaction data, extract the state space representation, action sequence, and reward value from the historical interaction data, construct a training sample set, and pre-train the deep reinforcement learning model based on the training sample set to obtain the policy network parameters and the value network parameters; Input the current state space representation into the policy network, generate the action selection probability based on the policy network parameters, sample the current state based on the action selection probability, and obtain the candidate action sequence; Input the state space representation and the candidate action sequence into the value network, calculate the state-action value based on the value network parameters, calculate the state transition probability according to the state transition function, and combine the state-action value and the state transition probability to obtain the expected cumulative value; Calculate the immediate reward of the candidate action sequence according to the reward function, and combine the immediate reward and the expected cumulative value to generate the action evaluation value; The historical state sequence is encoded by a recurrent unit to generate temporal correlation features. The temporal correlation features are combined with the action evaluation value to calculate the policy update gradient. Based on the policy update gradient, the parameters of the policy network are iteratively optimized. An interaction decision sequence is generated according to the optimized policy network parameters and the action evaluation value.

[0048] Exemplarily, historical interaction data is first obtained, which comes from the historical interaction process between the user and the system. For example, in an intelligent recommendation system, the historical interaction data may include user behavior records such as browsing products, clicking, and purchasing. The state space representation, action sequence, and reward value are extracted from these historical interaction data. The state space representation can be a user feature vector, such as [0.8, 0.2, 0.5, 0.7], representing the user's interest preferences; the action sequence can be a list of recommended product IDs, such as [1001, 1024, 1056]; the reward value can be the feedback of the user's click or purchase behavior, such as assigning 1 for a click and 0 for no click. These extracted data are constructed into a training sample set, and each sample contains the state representation, the executed action, and the obtained reward. Based on the constructed training sample set, the system pre-trains the deep reinforcement learning model to obtain the initial policy network parameters and value network parameters. The policy network parameters can be represented as a set of weight values, such as [-0.2, 0.5, 0.3, -0.1]; the value network parameters can also be represented as another set of weight values, such as [0.4, -0.3, 0.6, 0.2].

[0049] When an interaction decision needs to be generated for the current user, the current state space representation is first input into the policy network. For example, the state representation of the current user is [0.7, 0.3, 0.6, 0.4], representing the user's current interest preferences and behavior characteristics. The policy network generates the selection probabilities of each possible action based on the input state representation and the pre-trained policy network parameters. Suppose the system has three possible actions, and the selection probabilities output by the policy network are [0.2, 0.5, 0.3], indicating that the probability of selecting action 1 is 0.2, the probability of selecting action 2 is 0.5, and the probability of selecting action 3 is 0.3. Sampling is performed based on these action selection probabilities, and the possible sampled candidate action sequence may be [2, 3, 2, 1], indicating that actions 2, 3, 2, and 1 are selected in sequence as the candidate recommendation sequence.

[0050] Input the state space representation and candidate action sequence into the value network, and calculate the state-action value based on the value network parameters obtained through pre-training. Suppose for the action sequence [2, 3, 2, 1], the state-action values calculated by the value network are [3.2, 2.8, 3.0, 2.5] respectively, representing the long-term value expected to be obtained after executing these actions. Calculate the state transition probability according to the state transition function, which describes the probability distribution of the system transferring to a new state after executing a certain action in the current state. For example, after executing action 2, the system has a 0.7 probability of transferring to state A and a 0.3 probability of transferring to state B. Combine the state-action value and the state transition probability to obtain the expected cumulative value. Suppose the calculated expected cumulative value is [2.8, 2.5, 2.7, 2.2], representing the cumulative value expected to be obtained after considering state transitions when executing each action.

[0051] Calculate the immediate reward of the candidate action sequence according to the reward function. The reward function defines the immediate feedback obtained by executing a certain action in a specific state. For example, if the user clicks on the recommended product, the system gets a reward of 1; if the user purchases the recommended product, the reward is 5; if the user ignores the recommendation, the reward is 0. Suppose for the candidate action sequence [2, 3, 2, 1], the calculated immediate rewards are [1, 0, 1, 5] respectively. Combine the immediate reward and the expected cumulative value to generate the action evaluation value. For example, the immediate reward and the expected cumulative value can be weighted and summed, and the weights can be set to 0.3 and 0.7, and the obtained action evaluation value is [2.26, 1.75, 2.19, 3.04], representing the comprehensive value evaluation of executing each action.

[0052] Use a recurrent unit to encode the historical state sequence to generate temporal correlation features. The historical state sequence records the state changes of the user during past interactions. Through a recurrent unit (such as a long short-term memory network LSTM), the temporal dependencies in the state sequence can be captured. Suppose the temporal correlation features encoded by the system are [0.6, 0.4, 0.7, 0.3], representing the temporal pattern of the user's historical behavior. Combine the temporal correlation features with the action evaluation value to calculate the policy update gradient. For example, the temporal correlation features can be used as weights and weighted with the action evaluation value to obtain the adjusted action evaluation value [1.356, 0.7, 1.533, 0.912]. Based on the adjusted action evaluation value, calculate the update gradient of the policy network parameters and iteratively optimize the policy network parameters. The optimized policy network parameters may become [-0.15, 0.55, 0.28, -0.08], indicating that a better decision-making strategy has been learned. Finally, according to the optimized policy network parameters and the action evaluation value, generate the final interaction decision sequence, such as [2, 1, 3, 2], as the recommendation decision for the current user.

[0053] In this embodiment, by fusing user behavior characteristics and task scenario status information, a dynamic state space is constructed, and combined with the deep reinforcement learning policy iteration method, the decision-making efficiency and personalized adaptation ability in complex interaction scenarios are effectively improved. In the prior art, interaction strategies are mostly formulated based on preset rules or static behavior models, which cannot be flexibly adjusted according to individual differences of users and dynamic changes of task environments, easily leading to unreasonable resource allocation, deviation of task execution from the target, or degradation of user experience. This solution takes the user behavior feature vector as the core, constructs a state space by combining task difficulty, resource status, and environmental constraints, enhances the comprehensiveness and pertinence of state representation, and at the same time constructs a state transition function by introducing the weights of historical state sequences to achieve a reasonable modeling of the state evolution trend. In terms of the reward function design, by integrating task completion degree and execution efficiency, it breaks the evaluation mode based only on result feedback in traditional reinforcement learning and can better reflect the multi-dimensional performance during task execution. Finally, through the policy iteration optimization method of deep reinforcement learning, the system can dynamically adjust interaction strategies to achieve adaptive optimization in different user and task scenarios, effectively improving the intelligence level and decision-making quality of the interaction system.

[0054] In an alternative embodiment, constructing a task execution planning graph according to the interaction decision sequence, and generating a task execution plan with checkpoints based on the task execution planning graph includes: Extracting state information and decision actions from the interaction decision sequence, calculating the state transition frequency of the state information under the action of the decision action, and calculating the state transition probability according to the state transition frequency; Constructing nodes of the task execution planning graph based on the state information, constructing edges of the task execution planning graph based on the state transition probability, calculating the set of subsequent nodes reachable by each node and the transition cost, counting the number of times each state node in the interaction decision sequence is visited and the number of times each decision action is selected, and calculating the access weight of the node and the selection weight of the decision action in combination with the state transition probability; Constructing a path evaluation function based on the access weight, selection weight, and transition cost, using the path evaluation function to solve the optimal execution path in the task execution planning graph, calculating the change rate of the access weight and the change rate of the state transition probability of each node on the optimal execution path, and taking the weighted sum of the change rate of the access weight and the change rate of the state transition probability as the node importance; Selecting the node with the locally maximum importance on the optimal execution path as the checkpoint, extracting the state information, optional decision actions, and access weight of the node corresponding to the checkpoint to generate checkpoint configuration information, and combining the optimal execution path, checkpoint configuration information, and the state transition probability between each node to generate a task execution plan.

[0055] In this embodiment, state information and decision actions are extracted from the interaction decision sequence. State information refers to the working state of an employee at a specific time point in a virtual workplace environment, which can include characteristic values in multiple dimensions such as the current task progress, resource occupancy, and work efficiency indicators. For example, for a project manager, his state information can include the project completion rate, team collaboration index, and rationality of resource allocation. Decision actions refer to the actions taken by an employee in a specific state, such as task assignment, resource adjustment, and meeting convening. The interaction decision sequence is a sequence composed of a series of state information and decision actions in chronological order, recording the state changes and decision-making processes of an employee during task execution.

[0056] Calculate the state transition frequency of the state information under the action of the decision action. The state transition frequency represents the number of times of transitioning from one state to another state through a certain decision action in the interaction decision sequence. For example, when a project manager takes the "increase resources" decision in the "project delay" state, the frequency of the state transition to "project progress recovery" may be 5 times, and the frequency of the transition to "project still delayed" may be 2 times.

[0057] Calculate the state transition probability based on the state transition frequency. The state transition probability represents the probability of transitioning to each possible state after taking a certain decision action in a certain state. For example, after a project manager takes the "increase resources" decision in the "project delay" state, the probability of transitioning to "project progress recovery" is 5 / 7, and the probability of transitioning to "project still delayed" is 2 / 7. The state transition probability reflects the effectiveness of the decision action and the uncertainty of state transition.

[0058] Construct nodes of the task execution planning graph based on the state information. The task execution planning graph is a directed graph structure used to represent the state transition relationship during task execution. Each node in the graph corresponds to a state, such as "project start", "requirement analysis", "design and implementation", etc. All different states that appear in the interaction decision sequence are used as the nodes of the planning graph.

[0059] Construct the edges of the task execution planning graph based on the state transition probabilities. The edges in the planning graph represent the transition relationships from one state node to another, and the weights on the edges are the corresponding state transition probabilities. For example, the weight of the edge from the "Project Delayed" node to the "Project Schedule Recovery" node is 5 / 7. In this way, a complete task execution planning graph is constructed, reflecting the overall structure and probability distribution of the state transitions. Calculate the set of subsequent nodes reachable from each node and the transition cost. The set of subsequent nodes refers to all the nodes that can be reached from the current node through one-step transitions. The transition cost refers to the "cost" required to transfer from the current node to the subsequent node, which can be calculated based on the reciprocal of the state transition probability, that is, the higher the transition probability, the lower the transition cost. For example, the transition cost from the "Project Delayed" node to the "Project Schedule Recovery" node is 7 / 5, and the transition cost to the "Project Still Delayed" node is 7 / 2.

[0060] Count the number of times each state node is visited and the number of times each decision action is selected in the interactive decision sequence. The number of node visits indicates the total number of times in a certain state in the interactive decision sequence. The number of decision action selections indicates the total number of times a certain decision action is selected in the interactive decision sequence. For example, the "Project Delayed" state may be visited 10 times, and the "Increase Resources" decision action may be selected 7 times.

[0061] Calculate the access weight of the nodes and the selection weight of the decision actions in combination with the state transition probabilities. The access weight reflects the importance of the state node in the task execution process and can be calculated by dividing the number of node visits by the total number of state transitions. The selection weight of the decision action reflects the importance of the decision action in the task execution process and can be calculated by dividing the number of decision action selections by the total number of decisions. For example, if there are a total of 100 state transitions in the interactive decision sequence and the "Project Delayed" state is visited 10 times, then its access weight is 0.1.

[0062] Construct a path evaluation function based on the access weight, selection weight, and transition cost. The path evaluation function is used to evaluate the "goodness or badness" of different paths in the task execution planning graph, comprehensively considering the access weights of the nodes on the path, the selection weights of the decision actions, and the transition costs of the transitions. For example, a function can be designed such that the higher the sum of the access weights of the nodes on the path, the higher the sum of the selection weights of the decision actions, and the lower the sum of the transition costs, the higher the path evaluation value.

[0063] Solve the optimal execution path in the task execution planning graph using a path evaluation function. The optimal execution path refers to the path with the highest path evaluation function value among all possible paths from the start node to the target node in the task execution planning graph. Dynamic programming or graph search algorithms such as Dijkstra's algorithm or A* algorithm can be used to solve the optimal execution path in the task execution planning graph. For example, for a project management task, the optimal execution path may be "Project startup → Requirements analysis → Design and implementation → Testing and verification → Project delivery".

[0064] Calculate the change rate of access weight and the change rate of state transition probability for each node on the optimal execution path. The change rate of access weight represents the degree of change in access weight between adjacent nodes and can be obtained by dividing the difference in access weights of adjacent nodes by the access weight of the previous node. The change rate of state transition probability represents the degree of change in state transition probability between adjacent transitions and can be obtained by dividing the difference in state transition probabilities of adjacent transitions by the state transition probability of the previous transition.

[0065] Take the weighted sum of the change rate of access weight and the change rate of state transition probability as the node importance. Node importance reflects the key degree of the node in the optimal execution path. Nodes with high importance are often key points or turning points in the task execution process. For example, if a node has a large change rate of access weight and a large change rate of state transition probability, then the importance of this node is high and it may represent a key stage in the task execution process.

[0066] Select the node with the locally maximum importance on the optimal execution path as the checkpoint. A checkpoint refers to a key node that needs special attention and evaluation during the task execution process. By comparing the importance of adjacent nodes, select the node with the locally maximum importance as the checkpoint. For example, in the optimal execution path of "Project startup → Requirements analysis → Design and implementation → Testing and verification → Project delivery", if the importance of the "Design and implementation" node is locally maximum, then set it as the checkpoint.

[0067] Extract the state information, optional decision actions, and access weights of the node corresponding to the checkpoint to generate checkpoint configuration information. The checkpoint configuration information includes the state description of the checkpoint, the available decision actions, and the corresponding weight information, which is used to guide the decision-making and evaluation during the task execution process. For example, for the "Design and implementation" checkpoint, its configuration information may include state information such as "Degree of completion of the design document" and "Solution to technical difficulties", as well as optional decision actions such as "Adjust the design plan" and "Increase technical resources".

[0068] Generate a task execution plan by combining the optimal execution path, checkpoint configuration information, and state transition probabilities between nodes. The task execution plan is a complete task execution guidance document that includes the optimal path from task start to task completion, key checkpoints, and probability information for each state transition. For example, the execution plan for a project management task may include the optimal execution path "Project startup → Requirements analysis → Design and implementation → Testing and verification → Project delivery", the configuration information for the checkpoint "Design and implementation", and the probabilities of each state transition, such as the transition probability from "Requirements analysis" to "Design and implementation" being 0.9.

[0069] During the multi-dimensional assessment process in the virtual workplace, this task execution plan can be used to guide employees' task execution and serve as a benchmark for assessment. The differences between the actual execution path of an employee and the optimal execution path can be compared, the performance of the employee at checkpoints can be evaluated, and the consistency between the employee's decisions and the recommended decisions can be analyzed, thereby enabling a comprehensive and objective multi-dimensional assessment. At the same time, based on the deep reinforcement learning algorithm, the task execution planning diagram and the optimal execution path can be continuously optimized to improve the accuracy and guidance of the assessment.

[0070] This technical solution can achieve the structured expression of the interactive decision-making process and the dynamic calibration of key nodes, thereby enhancing the controllability and robustness of the task execution plan. In the existing technology, task planning mostly relies on static flowcharts or predefined paths, lacking the reflection of users' actual interactive behaviors and being difficult to identify key control points in complex tasks, resulting in weak fault tolerance and lagging feedback responses during the task execution process. This solution constructs a task execution planning diagram by extracting state transition probabilities and behavior weights, comprehensively integrating the path preferences and decision-making habits of users in actual operations, and then introducing a path evaluation function to comprehensively evaluate multiple possible paths to ensure that the selected path achieves a balance between efficiency and stability. By calculating the change rates of access weights and transition probabilities, the fluctuation regions in the task flow are identified, and representative important nodes are selected as checkpoints, effectively enhancing the response ability of the task plan to unexpected situations and the process monitoring ability. The finally generated task execution plan not only reflects the optimal decision-making path of users' behaviors but also clarifies the key control nodes and state transition relationships, providing solid data support for the dynamic management and process optimization of tasks.

[0071] Figure 3 This is a schematic diagram of the task execution plan, presenting an optimal execution path (connected by solid arrows): starting from "Project startup" (weight 0.15), passing through "Requirements analysis" (weight 0.18), reaching the checkpoint "Design and implementation" (weight 0.25), then entering "Testing and verification" (weight 0.22), and finally completing "Project delivery" (weight 0.20). Among them, the "Design and implementation" node is identified as a key checkpoint, which is determined based on the characteristic of the local maximum importance of this node.

[0072] Meanwhile, two alternative paths (connected by dashed arrows) are also shown: one is the path from "Project Initiation" to "Solution Review" (weight 0.12) and then to "Design Implementation"; the other is the path from "Design Implementation" to "Problem Fixing" (weight 0.15) and then to "Project Delivery". The state transition probabilities show that in the "Project Initiation" state, there is a 0.75 probability of directly entering "Requirement Analysis" and a 0.25 probability of transitioning to "Solution Review"; while in the "Design Implementation" state, there is a 0.68 probability of directly entering "Test Verification" and a 0.32 probability of going through "Problem Fixing".

[0073] This task execution planning diagram quantifies the feasibility of different execution paths and identifies key control nodes by analyzing the state transition frequencies and probabilities in the interactive decision sequence, providing data support and guidance for task execution.

[0074] In an alternative implementation, the execution status data of each checkpoint is collected, the task completion trajectory curve is calculated based on the execution status data, and the decision choices of the user at each checkpoint are extracted in combination with the task execution planning diagram. The decision choices are feature fused with the user behavior feature vector to generate an ability dimension vector, and the professional competency score is calculated based on the ability dimension vector, and a professional ability assessment report is generated, including: During the process of the user executing the task execution plan, the execution status data of the checkpoint is collected, and the execution status data includes current state characteristics, selected decision actions, execution results, and timestamp information; Calculate the state deviation by calculating the difference between the state characteristics in the execution status data and the preset state of the task execution planning diagram, and calculate the action deviation by calculating the difference between the decision actions in the execution status data and the recommended actions of the task execution planning diagram; Calculate the checkpoint completion score based on the weighted sum of the state deviation and the action deviation, calculate the execution efficiency score based on the change in execution results between adjacent timestamps, and weight and combine the completion score and the execution efficiency score to obtain the checkpoint score; Interpolate and fit the scores of each checkpoint and the corresponding timestamps to generate a task completion trajectory curve, and extract the decision choices of the user at each checkpoint in combination with the preset state and recommended actions recorded in the task execution planning diagram; Feature fuse the decision choices with the user behavior feature vector in the corresponding dimension to generate an ability dimension vector, calculate the normalized scores of the feature values in each dimension of the ability dimension vector, use the normalized scores as the professional competency scores, and generate a professional ability assessment report.

[0075] Exemplarily, during the process of a user executing a task execution plan, the execution status data of checkpoints is collected. The execution status data refers to the comprehensive record of information such as the user's working status, decision-making behaviors, and execution results at each checkpoint during the task execution. The execution status data includes four key components: the current status feature, the selected decision action, the execution result, and the timestamp information. The current status feature refers to the multi-dimensional description of the user's working status at the checkpoint. For example, for a product manager, the status features may include the completion degree of the product requirement document, the depth of user research, the integrity of prototype design, etc. The selected decision action refers to the specific action selected by the user in the current state, such as "optimize product requirements", "increase the sample size of user research", etc. The execution result refers to the effect evaluation after the decision action is executed, which can be a quantitative or qualitative description, such as "the completion degree of the product requirement document has increased by 20%". The timestamp information records the exact time of status collection for subsequent time series analysis.

[0076] Calculate the difference between the status feature in the execution status data and the preset status of the task execution plan graph to obtain the status deviation. The ideal status feature of each checkpoint is preset in the task execution plan graph, which represents the working status that the user should reach at this checkpoint. The status deviation represents the gap between the user's actual status and the ideal status, and can be obtained by calculating the difference between the status feature vector and the preset status vector. For example, if the actual value of the completion degree of the product requirement document of the product manager at the requirements analysis checkpoint is 75%, while the preset value is 90%, then the status deviation is -15%. The smaller the status deviation, the closer the user's working status is to the ideal status.

[0077] Calculate the difference between the decision action in the execution status data and the recommended action of the task execution plan graph to obtain the action deviation. The optimal decision action is recommended for each checkpoint in the task execution plan graph, and these recommended actions are learned from a large amount of historical data based on a deep reinforcement learning model. The action deviation represents the degree of difference between the decision action actually selected by the user and the recommended action. In actual implementation, the decision action can be represented as a multi-dimensional vector, and each dimension corresponds to the execution intensity of a basic action type. The action deviation can be obtained by calculating the difference between the decision action vector and the recommended action vector. For example, if the recommended action is "increase the sample size of user research and optimize the requirement document", while the user only executes "optimize the requirement document", then there is an action deviation. The smaller the action deviation, the closer the user's decision-making choice is to the optimal strategy.

[0078] Calculate the checkpoint completion score based on the weighted sum of the state deviation and the action deviation. The completion score reflects the degree to which the user's working state and decision-making behavior at the checkpoint conform to the ideal situation. During the calculation process, different weights can be assigned to the state deviation and the action deviation according to different task types and checkpoint characteristics. Usually, the weight of the state deviation is higher because the working state is a direct manifestation of task completion. For example, the weight of the state deviation can be set to 0.7, and the weight of the action deviation can be set to 0.3, and the completion score is obtained through weighted combination. The higher the completion score, the better the user's performance at this checkpoint.

[0079] Calculate the execution efficiency score based on the change in execution results between adjacent timestamps. The execution efficiency score reflects the speed and efficiency at which the user advances the task between adjacent checkpoints. During the calculation process, compare the change in execution results between adjacent timestamps and evaluate it in combination with the time interval. For example, if the user improves the completion rate of the product requirements document from 50% to 80% within two hours, the execution efficiency score will be calculated based on this change and the time interval. The higher the execution efficiency score, the higher the user's work efficiency.

[0080] Weight and combine the completion score and the execution efficiency score to obtain the checkpoint score. The checkpoint score is a quantitative evaluation of the user's comprehensive performance at the checkpoint, comprehensively considering multiple aspects such as the working state, decision-making behavior, and execution efficiency. In practical applications, the weights of the completion score and the execution efficiency score can be adjusted according to different positions and task characteristics. For example, for a R & D engineer, more attention may be paid to execution efficiency, so the weight of the execution efficiency score can be set higher; while for a product manager, more attention may be paid to work quality, so the weight of the completion score can be set higher.

[0081] Interpolate and fit the scores of each checkpoint and the corresponding timestamps to generate a task completion trajectory curve. The task completion trajectory curve is a curve with time as the horizontal axis and the checkpoint score as the vertical axis, reflecting the changing trend of the user's performance during the entire task execution process. Using an interpolation and fitting algorithm, connect the discrete checkpoint scores into a smooth curve. Commonly used interpolation methods include linear interpolation, spline interpolation, etc., and an appropriate interpolation method can be selected according to the data characteristics. The task completion trajectory curve can intuitively display the user's work rhythm and performance fluctuations, helping to identify the user's advantageous stages and areas for improvement.

[0082] Extract the user's decision-making choices at each checkpoint by combining the preset states and recommended actions recorded in the task execution planning diagram. Decision-making choices refer to the specific actions taken by the user at each checkpoint and are important bases for evaluating the user's decision-making ability and problem-solving ability. By comparing the user's decision-making choices with the recommended actions in the task execution planning diagram, the user's decision-making patterns and preferences can be analyzed. For example, it may be found that a certain user tends to choose conservative decision-making actions and rarely takes innovative decision-making actions. This decision-making choice information will be used for subsequent analysis of ability dimensions.

[0083] Fuse the decision-making choices and the user behavior feature vector in the corresponding dimensions to generate an ability dimension vector. The user behavior feature vector is a multi-dimensional feature description formed by long-term observation and recording of the user's behavior patterns in the virtual workplace, including aspects such as communication style, collaboration mode, learning ability, stress resistance ability, etc. Feature fusion refers to integrating the decision-making choice information and the user behavior feature vector to form a more comprehensive ability dimension vector. During the fusion process, the decision-making choices are mapped to the corresponding ability dimensions and weighted combination is performed with the corresponding dimensions in the user behavior feature vector. For example, if the user selects the decision-making action of "increasing resource investment" when facing the risk of project delay, this decision will be mapped to dimensions such as "resource management ability" and "risk response ability" and fused with the corresponding dimensions in the user behavior feature vector.

[0084] Calculate the normalized scores of the feature values of each dimension in the ability dimension vector and use the normalized scores as the professional competency scores. Normalization is the process of converting feature values with different dimensions and ranges into a unified standard to ensure the comparability of the scores of each dimension. Common normalization methods include min-max normalization, Z-score normalization, etc. Through normalization, the feature values of each dimension in the ability dimension vector are converted into professional competency scores between 0 and 100. For example, the user's "communication and coordination ability" score may be 85 points, the "professional skill level" score may be 92 points, and the "innovative thinking ability" score may be 78 points. These scores reflect the performance levels of the user in each ability dimension.

[0085] Generate a professional ability assessment report based on the professional competency scores. The professional ability assessment report is a comprehensive assessment document of the user's professional ability, including multiple parts such as ability dimension analysis, strength identification, improvement suggestions, etc. In the report, the task completion trajectory curve of the user will be shown, the performance characteristics of the user at different checkpoints will be analyzed, and the distribution of the user's professional competency scores in each dimension will be emphasized. Based on the deep reinforcement learning model, the matching degree between the user's ability characteristics and different positions will also be compared to provide career development suggestions.

[0086] In practical applications, this method can provide a scientific basis for the talent assessment and development of enterprises. Through task execution and multi-dimensional assessment in a virtual workplace environment, enterprises can comprehensively understand the ability characteristics and development potential of employees, providing precise guidance for talent cultivation, job matching, and career planning. At the same time, employees can also clearly recognize their strengths and weaknesses through the occupational ability assessment report, and targeted improve key abilities to achieve career development goals.

[0087] In this embodiment, by collecting status data during task execution and comparing and analyzing it with the planning diagram, the task completion situation of the user can be evaluated in real time. By calculating the status deviation and action deviation, the gap between the user's execution effect and the expected goal can be accurately measured. Combining the weighted calculation of the completion score and the execution efficiency score comprehensively reflects the performance level of the user during task execution. By interpolating and fitting the checkpoint scores, a continuous task completion trajectory curve is generated, intuitively showing the user's execution process and the trend of ability improvement. By fusing decision-making choices with behavioral feature vectors, ability dimension vectors are generated in the corresponding dimensions, realizing multi-dimensional quantitative assessment of the user's occupational ability. Finally, through normalization processing, a standardized competency score is obtained, providing an objective basis for talent assessment and job matching. This solution can not only dynamically track and evaluate the user's task execution situation, but also deeply analyze the user's decision-making behavior and ability characteristics, providing data support for the talent development of enterprises.

[0088] In the second aspect of the embodiments of the present invention, a virtual workplace multi-dimensional assessment system based on deep reinforcement learning is provided, and the system includes: A first unit, configured to collect the behavioral data of a user in a virtual workplace environment, perform temporal feature extraction, calculate the behavioral state sequence of the user in different task stages, construct a state transition probability matrix based on the behavioral state sequence, analyze the stability characteristics of the user's behavioral pattern according to the state transition probability matrix, and generate a user behavioral feature vector; A second unit, configured to construct a state space representation based on the user behavioral feature vector, calculate a state transition function in combination with the task scenario state information, set a reward function according to the task completion degree and execution efficiency, and generate an interaction decision sequence by using the policy iteration method of deep reinforcement learning; A third unit, configured to construct a task execution planning diagram according to the interaction decision sequence, and generate a task execution plan with checkpoints based on the task execution planning diagram; The fourth unit is used to collect the execution status data of each checkpoint during the process of the user executing the task execution plan, calculate the task completion trajectory curve based on the execution status data, extract the decision-making choices of the user at each checkpoint in combination with the task execution planning diagram, perform feature fusion on the decision-making choices and the user behavior feature vector to generate an ability dimension vector, calculate the professional competence score based on the ability dimension vector, and generate a professional ability assessment report.

[0089] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0090] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0091] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-dimensional assessment method for virtual workplaces based on deep reinforcement learning, characterized in that Including: Collect the behavioral data of the user in the virtual workplace environment, perform time series feature extraction, calculate the behavioral state sequence of the user in different task stages, construct a state transition probability matrix based on the behavioral state sequence, analyze the stability characteristics of the user's behavioral pattern according to the state transition probability matrix, and generate a user behavioral feature vector; Construct a state space representation based on the user behavioral feature vector, calculate a state transition function in combination with the task scenario state information, set a reward function according to the task completion degree and execution efficiency, and generate an interaction decision sequence by using the policy iteration method of deep reinforcement learning; Construct a task execution plan graph according to the interaction decision sequence, and generate a task execution plan with checkpoints based on the task execution plan graph; During the process of the user executing the task execution plan, collect the execution state data of each checkpoint, calculate the task completion trajectory curve according to the execution state data, extract the decision-making choices of the user at each checkpoint in combination with the task execution plan graph, perform feature fusion on the decision-making choices and the user behavioral feature vector to generate an ability dimension vector, calculate the professional competency score according to the ability dimension vector, and generate a professional ability evaluation report.

2. The method according to claim 1, wherein Collect the behavioral data of the user in the virtual workplace environment, perform time series feature extraction, calculate the behavioral state sequence of the user in different task stages, construct a state transition probability matrix based on the behavioral state sequence, analyze the stability characteristics of the user's behavioral pattern according to the state transition probability matrix, and generate a user behavioral feature vector including: Collect the behavioral data of the user in the virtual workplace environment, where the behavioral data includes operation sequence data, task switching data, problem-solving data, and collaboration mode data; Calculate a behavioral complexity index for the behavioral data, dynamically determine the time window size according to the behavioral complexity index, extract the statistical features and trend features of the behavioral data within the time window, and combine them to generate a behavioral state vector; Calculate the time series correlation degree of the behavioral state vector, determine the behavior jump threshold according to the time series correlation degree, and perform density clustering on the behavioral state vector based on the behavior jump threshold to generate a behavioral state sequence; Calculate the transition frequency between states according to the behavioral state sequence, construct a state transition probability matrix, perform matrix decomposition on the state transition probability matrix to obtain a first transition probability sub-matrix and a second transition probability sub-matrix; perform eigenvalue decomposition on the first transition probability sub-matrix and the second transition probability sub-matrix respectively to obtain an eigenvalue set, calculate the behavior entropy value according to the largest eigenvalue in the eigenvalue set, calculate the behavior stability characteristics according to the remaining eigenvalues, and perform feature fusion on the behavior entropy value and the behavior stability characteristics to generate a user behavioral feature vector.

3. The method according to claim 2, wherein Calculate the time series correlation degree of the behavioral state vector, determine the behavior jump threshold according to the time series correlation degree, and perform density clustering on the behavioral state vector based on the behavior jump threshold to generate a behavioral state sequence including: Obtain the behavioral state vector, calculate the vector space distance and time interval distance between adjacent behavioral state vectors, combine the vector space distance and time interval distance to construct a time series correlation function, and calculate the time series correlation degree of the behavioral state vector through the time series correlation function; Statistically analyze the temporal correlation degree, obtain the distribution characteristics of the temporal correlation degree, generate a behavior jump threshold according to the distribution characteristics, determine the behavior jump points according to the temporal correlation degree and the behavior jump threshold, and calculate the local density of the behavior state vector at the behavior jump points; Based on the local density, perform density clustering on the behavior state vectors to obtain a behavior state clustering result, and arrange and combine the behavior state clustering results in chronological order to generate a behavior state sequence.

4. The method according to claim 1, characterized in that Construct a state space representation based on the user behavior feature vector, calculate the state transition function in combination with the task scenario state information, set the reward function according to the task completion degree and execution efficiency, and generate an interaction decision sequence using the policy iteration method of deep reinforcement learning, including: Obtain the task scenario state information, which includes task difficulty information, resource state information, and environmental constraint information; Construct a state space representation based on the user behavior feature vector and the task scenario state information, calculate the transition probability between states based on the state space representation, and construct a state transition function in combination with the importance weights of the historical state sequence. The state transition function includes the state evolution law and state transition constraints; Monitor the task execution process in real time, calculate the task completion degree index based on the task objective completion rate, task quality score, and task timeliness, calculate the execution efficiency index based on the computing resource utilization rate, storage resource occupancy rate, and operation time utilization rate, and construct a reward function by weighted combination of the task completion degree index and the execution efficiency index; Input the state space representation into a pre-trained deep reinforcement learning model, calculate the state value based on the state transition function, and perform policy iteration optimization according to the reward function and the state value to generate an interaction decision sequence.

5. The method according to claim 4, wherein Input the state space representation into a pre-trained deep reinforcement learning model, calculate the state value based on the state transition function, and perform policy iteration optimization according to the reward function and the state value to generate an interaction decision sequence, including: Obtain historical interaction data, extract the state space representation, action sequence, and reward value from the historical interaction data, construct a training sample set, and pre-train the deep reinforcement learning model based on the training sample set to obtain the policy network parameters and value network parameters; Input the current state space representation into the policy network, generate an action selection probability based on the policy network parameters, sample the current state based on the action selection probability, and obtain a candidate action sequence; Input the state space representation and the candidate action sequence into the value network, calculate the state-action value based on the value network parameters, calculate the state transition probability according to the state transition function, and combine the state-action value and the state transition probability to obtain the expected cumulative value; Calculate the immediate reward of the candidate action sequence according to the reward function, and combine the immediate reward and the expected cumulative value to generate an action evaluation value; A recurrent unit is used to encode the historical state sequence to generate temporal correlation features. The temporal correlation features are combined with the action evaluation value to calculate the policy update gradient. Based on the policy update gradient, the parameters of the policy network are iteratively optimized, and an interaction decision sequence is generated according to the optimized policy network parameters and the action evaluation value.

6. The method according to claim 1, wherein A task execution plan graph is constructed according to the interaction decision sequence. Generating a task execution plan with checkpoints based on the task execution plan graph includes: Extract the state information and decision actions from the interaction decision sequence, calculate the state transition frequency of the state information under the action of the decision action, and calculate the state transition probability according to the state transition frequency; Based on the state information, construct the nodes of the task execution plan graph. Based on the state transition probability, construct the edges of the task execution plan graph. Calculate the set of subsequent nodes that each node can reach and the transition cost. Count the number of times each state node in the interaction decision sequence is visited and the number of times each decision action is selected. Combine the state transition probability to calculate the access weight of the node and the selection weight of the decision action; Construct a path evaluation function based on the access weight, selection weight, and transition cost. Use the path evaluation function to solve the optimal execution path in the task execution plan graph. Calculate the change rate of the access weight and the change rate of the state transition probability of each node on the optimal execution path. Use the weighted sum of the access weight change rate and the state transition probability change rate as the node importance; Select the node with the locally maximum importance on the optimal execution path as the checkpoint. Extract the state information, optional decision actions, and access weight of the node corresponding to the checkpoint to generate checkpoint configuration information. Combine the optimal execution path, checkpoint configuration information, and the state transition probability between each node to generate a task execution plan.

7. The method according to claim 1, characterized in that, Collect the execution status data of each checkpoint, calculate the task completion trajectory curve according to the execution status data, and extract the decision choices of the user at each checkpoint in combination with the task execution plan graph. Perform feature fusion on the decision choices and the user behavior feature vector to generate an ability dimension vector. Calculate the professional competency score according to the ability dimension vector. Generating a professional ability assessment report includes: Collect the execution status data of the checkpoint during the user's execution of the task execution plan. The execution status data includes the current state characteristics, selected decision actions, execution results, and timestamp information; Calculate the difference between the state characteristics in the execution status data and the preset state of the task execution plan graph to obtain the state deviation. Calculate the difference between the decision action in the execution status data and the recommended action of the task execution plan graph to obtain the action deviation; Calculate the checkpoint completion score according to the weighted sum of the state deviation and the action deviation. Calculate the execution efficiency score according to the change in the execution results between adjacent timestamps. Combine the completion score and the execution efficiency score by weighting to obtain the checkpoint score; Interpolate and fit the scores of each checkpoint and the corresponding timestamps to generate a task completion trajectory curve. Combine the preset state and recommended actions recorded in the task execution plan graph to extract the decision choices of the user at each checkpoint; Fuse the decision selection and the user behavior feature vector in the corresponding dimension to generate a capability dimension vector, calculate the normalized scores of the feature values of each dimension in the capability dimension vector, use the normalized scores as the professional competency scores, and generate a professional ability assessment report.

8. A multi-dimensional assessment system for virtual workplaces based on deep reinforcement learning, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that, It includes: The first unit is used to collect the behavior data of the user in the virtual workplace environment, perform time series feature extraction, calculate the behavior state sequence of the user in different task stages, construct a state transition probability matrix based on the behavior state sequence, analyze the stability characteristics of the user behavior pattern according to the state transition probability matrix, and generate a user behavior feature vector; The second unit is used to construct a state space representation based on the user behavior feature vector, calculate a state transition function in combination with the task scenario state information, set a reward function according to the task completion degree and execution efficiency, and generate an interaction decision sequence by using the policy iteration method of deep reinforcement learning; The third unit is used to construct a task execution planning diagram according to the interaction decision sequence and generate a task execution plan with checkpoints based on the task execution planning diagram; The third unit is used to collect the execution status data of each checkpoint during the process of the user executing the task execution plan, calculate the task completion trajectory curve according to the execution status data, extract the decision selection of the user at each checkpoint in combination with the task execution planning diagram, fuse the decision selection and the user behavior feature vector to generate a capability dimension vector, calculate the professional competency score according to the capability dimension vector, and generate a professional ability assessment report.

9. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Non-uniform density data clustering method and device, electronic equipment and storage medium

    CN115130602A

  • Intelligent post matching model establishing method and matching method based on deep reinforcement learning

    CN118861999A

  • Post competency evaluation system based on virtual reality technology

    CN119026971A

  • Method and system for managing visualization tables of databases of different types

    CN119537424A

  • System and method for behavioral model clustering in television usage, targeted advertising via model clustering, and preference programming based on behavioral model clusters

    US20030101449A1