A Method for Characterizing and Dynamically Updating Expert Portraits Based on Reinforcement Learning

Through a method based on reinforcement learning, combined with reinforcement learning and label structure change method, a variable structure network is constructed, which solves the problem that existing expert portrait methods are difficult to reflect changes in expert dynamic behavior patterns and cognitive states, and realizes dynamic portrayal and evolution updates of expert portraits, and improves the system's adaptability and interpretability.

CN119918642BActive Publication Date: 2025-06-20HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422163.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-20
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing expert portrait methods are difficult to truly reflect the changes in the expert's dynamic behavior patterns and cognitive state, cannot meet the needs of continuous modeling and predictive recommendation, and lack the understanding of the deep semantic relationships between multimodal features and the adaptive deformation ability of the portrait structure.

Method used

Using reinforcement learning-based methods, integrating reinforcement learning and label structure change-dimensional methods, a variable structure network is built, and a two-way optimization mechanism driven by behavior feedback is realized to achieve dynamic portrayal and evolutionary updates of expert portraits.

Benefits of technology

It realizes accurate expression, structural adaptability, optimization closed-loop and controllability of expert portraits, and improves the adaptability, real-timeness and interpretability of expert portrait systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918642B_ABST
    Figure CN119918642B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for depicting and dynamically updating an expert portrait based on reinforcement learning, comprising the following steps: S1, collecting expert data and performing preprocessing to generate a knowledge vector, a behavior vector, and a semantic vector; S2, constructing an expert cognitive representation tensor with potential cognitive labels as the core; S3, inputting the expert cognitive representation tensor into a variable-dimensional mimicry generation network to generate an expert portrait candidate; S4, using a multi-strategy enhanced whale optimization algorithm to optimize the expert portrait candidate to generate an optimized expert portrait; S5, generating an expert portrait evolution path map according to the optimized expert portrait; S6, collecting real behavior feedback data of the expert portrait, and combining the expert portrait evolution path map, generating a joint reward signal through a reward function for two-way update to achieve dynamic update and optimized evolution of potential cognitive labels. The present invention combines reinforcement learning with a label structure variable-dimensional method to achieve dynamic depiction and update of the expert portrait.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning, and in particular to a method for characterizing and dynamically updating expert portraits based on reinforcement learning. Background Art

[0002] With the development of artificial intelligence, knowledge graphs and big data analysis technologies, expert portraits, as a means of modeling the multi-dimensional characteristics of individual experts in specific fields, such as knowledge, ability, behavioral habits and research interests, have been widely used in tasks such as scientific research talent recommendation, team intelligent formation, expert scheduling and academic decision support. The existing expert portrait system is mainly based on the idea of ​​static feature modeling, that is, by analyzing structural data (such as papers, projects, titles, research fields) and semantic data (such as academic profiles, research interests, etc.), a relatively fixed set of labels and weight representations are constructed to describe the knowledge and ability of experts. However, in real scenarios, the behavior patterns, research directions and cross-domain migration behaviors of experts have obvious dynamic evolution characteristics. The traditional static portrait method is difficult to truly reflect the changes in the cognitive state of experts and cannot meet the actual needs of continuous modeling and predictive recommendation of experts.

[0003] From a methodological perspective, traditional expert portrait methods mostly use vector space models or expert representations based on topic models. Their fusion methods often rely on simple splicing or linear weighting, and lack an understanding of the deep semantic relationship between multimodal features. Especially when processing behavioral data (such as task responses, click paths, and cross-domain migration), their time series and behavior-driven characteristics are often ignored, and they are only used as auxiliary information for scoring calculations. This results in the inability of the portrait model to extract dynamic label signals with cognitive value from the expert's behavioral trajectory, further limiting the depth of modeling of the label evolution mechanism. In addition, most existing methods do not consider the variability and adaptive evolution of the expert portrait structure itself. Once the portrait label set is determined, it is difficult to expand or shrink the structure as the behavior evolves, which is not conducive to the continuous characterization and discovery of the expert's potential capabilities.

[0004] In terms of optimization mechanism, the parameter learning of current portrait models mostly relies on supervised learning or heuristic rule matching based on feature similarity, lacking a clear feedback loop and self-evolution capabilities. Although some work has attempted to introduce reinforcement learning strategies to optimize the expert recommendation process, it often only reinforces the scores at the output level, without linking with the intrinsic evolution mechanism of the portrait structure, resulting in low feedback utilization efficiency, slow learning convergence, and unclear evolution path. At the same time, the semantic conflicts and redundant label problems in label combinations have not been effectively solved. In the absence of fine structural control, semantically repeated or logically inconsistent label structures are often generated, affecting the interpretability and credibility of the final portrait.

[0005] In addition, the lack of a path-based structure recording mechanism in expert portrait modeling is also a major flaw in current technologies. Since expert portraits may undergo multiple updates and evolutions at multiple stages, if the structural change process between portraits is not explicitly modeled, the system will be unable to provide the basis for portrait evolution and track the logical process of the generation, transfer, or disappearance of a certain label. This seriously hinders the transparent interpretation and behavior traceability of expert portraits in application systems, and also limits the practical application capabilities of expert portrait models in highly reliable recommendation, expert behavior prediction, and long-term learning tasks.

[0006] Based on this, existing expert portrait methods still have obvious deficiencies in the following aspects: First, they lack the modeling of the deep fusion mechanism between structured, behavioral, and semantic multi-source data and cannot construct a high-dimensional cognitive representation with unified semantic expression ability; second, they lack the adaptive deformation and generation ability at the portrait structure level and cannot support the evolution of label structures as expert behaviors change; third, they lack mechanisms for portrait confidence guidance, label conflict suppression, and behavior compatibility control during the optimization process, affecting the optimization effect and rationality of portrait structures; fourth, they fail to establish a complete portrait evolution trajectory map and lack the modeling and recording of the structural evolution relationship within the life cycle of expert portraits; fifth, they lack a joint reward-driven mechanism based on behavior feedback, cannot achieve the bidirectional iterative update of portrait structures and optimization strategies, and lack effective dynamic learning ability.

[0007] Therefore, how to provide a method for depicting and dynamically updating expert portraits based on reinforcement learning is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0008] An object of the present invention is to propose a method for depicting and dynamically updating expert portraits based on reinforcement learning. The present invention integrates reinforcement learning with a variable-dimensional label structure method, constructs a variable structure network including label growth, structure compression, and dimensional deformation, and introduces a behavior feedback-driven two-way optimization mechanism to achieve the dynamic depiction and evolutionary update of expert portraits, with the advantages of accurate expression, structure adaptability, optimization closed-loop, and evolution controllability, and can effectively improve the adaptability, real-time performance, and interpretability of expert portrait systems in complex tasks.

[0009] A method for depicting and dynamically updating expert portraits based on reinforcement learning according to an embodiment of the present invention includes the following steps:

[0010] S1. Collect the structured data, behavioral data, and semantic data of experts, and perform preprocessing to generate knowledge vectors, behavioral vectors, and semantic vectors;

[0011] S2. Integrate the knowledge vectors, behavioral vectors, and semantic vectors to construct an expert cognitive representation tensor centered on potential cognitive labels;

[0012] S3. Input the expert cognitive representation tensor into the variable-dimensional mimicry generation network, which includes a label growth unit, a structure compression unit, and a dimension deformation unit, to generate candidate expert portraits;

[0013] S4. Optimize the candidate expert portraits using the multi-strategy enhanced whale optimization algorithm, with portrait confidence, potential cognitive label consistency, and expert historical behavior compatibility as constraint conditions, to generate optimized expert portraits;

[0014] S5. Generate an expert portrait evolution path map based on the optimized expert portraits, where the nodes represent the states of the optimized expert portraits, and the edges represent growth, contraction, and replacement relationships;

[0015] S6. Collect real behavior feedback data of the expert portraits, and combine with the expert portrait evolution path map to generate a joint reward signal through a reward function, and use the joint reward signal to drive the bidirectional update of the structure parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, to achieve the dynamic update of the expert portraits and the optimized evolution of the potential cognitive labels.

[0016] Optionally, the S2 specifically includes:

[0017] S21. Convert the knowledge vector, behavior vector, and semantic vector to a unified feature dimension through linear mapping;

[0018] S22. Taking the behavior vector as the leading factor, calculate the interaction weights between the behavior vector and the knowledge vector and the semantic vector, perform weighted processing, and splice the weighted knowledge vector, semantic vector, and behavior vector to form a fused feature vector;

[0019] S23. Construct a potential cognitive label embedding matrix, which is obtained by embedding and encoding a set of set labels, and the set of labels is derived from long-term accumulated expert behavior logs, portrait label history, and semantic topic analysis;

[0020] S24. Perform a tensor product calculation on the fused feature vector and the potential cognitive label embedding matrix to generate an expert cognitive representation tensor, which is a third-order tensor, including a label dimension, a feature dimension, and a time dimension, representing the feature combination under the potential cognitive label dimension, the label response degree, and the time change characteristics in the expert behavior evolution process, respectively.

[0021] Optionally, the S3 specifically includes:

[0022] S31. Slice the expert cognitive representation tensor in the label dimension direction to extract a set of potential cognitive label vectors, denoted as , where represents the A potential cognitive label vector, Indicates the total number of potential cognitive label vectors;

[0023] S32. Input each potential cognitive label vector into the label growth unit and calculate the growth probability:

[0024] ;

[0025] Among them, Indicates the th growth probability of the potential cognitive label vector, Indicates the activation function, And Indicates the weight matrix, Indicates the Gaussian error linear unit activation function, And Indicates the bias vector, Indicates the label amplitude adjustment factor, Indicates Of Norm;

[0026] If Exceeds the set threshold, trigger the label expansion operation and add a new label to the label dimension direction of the expert cognitive representation tensor;

[0027] S33. Input the updated expert cognitive representation tensor into the structure compression unit and calculate the retention importance score of each potential cognitive label vector:

[0028] ;

[0029] Among them, Indicates the retention importance score of the th potential cognitive label vector, And Indicates the weight coefficient, Indicates the th average response value of the potential cognitive label vector in the attention mechanism, Indicates the th potential cognitive label vector in the previous round structure activation variance;

[0030] If Is lower than the set threshold, remove it from the expert cognitive representation tensor structure;

[0031] S34. Input the compressed tensor into the dimension deformation unit and perform affine perturbation and nonlinear transformation on each label vector:

[0032] ;

[0033] Among them, represents the th candidate label vector, and represent the structural perturbation matrix, represents the perturbation amplitude adjustment coefficient, represents the offset vector, represents the Gaussian noise perturbation tensor;

[0034] S35. Calculate the structural difference index of the candidate label vector and perform screening to obtain the candidate body of the expert portrait:

[0035] ;

[0036] Among them, represents the comprehensive structural difference degree of the candidate label vector, represents the historical reference label vector, represents the adjustment coefficient of the divergence term, represents the distribution of the candidate label vector and the historical reference label vector between divergence;

[0037] If falls into the preset structural deviation interval, then retain the corresponding expert cognitive representation tensor as the candidate body of the expert portrait.

[0038] Optionally, the S4 specifically includes:

[0039] S41. Allocate a confidence evaluation module to the candidate body of the expert portrait. The confidence evaluation module makes a fusion judgment based on the coverage of the label dimension of the expert cognitive representation tensor in the expert's historical behavior trajectory, the participation frequency of the completed tasks, and the matching degree with the expert semantic description information to form a confidence input source;

[0040] S42. Calculate the structural confidence score of the candidate body of the expert portrait based on the weighted summation method:

[0041] ;

[0042] Among them, represents the th structural confidence score of the candidate body of the expert portrait, represents the total number of potential cognitive label vectors, , and represent the weighted coefficients, represents the th coverage of the potential cognitive label vector in the expert's historical behavior trajectory, represents the The participation frequency of a potential cognitive label vector in the completed tasks, indicating the matching degree between the

[0043] S43. Use the structural confidence score to regulate the search direction control parameter of individuals in the multi-strategy enhanced whale optimization algorithm. Among them, the candidate expert portraits with a structural confidence score greater than the set score threshold will preferentially approach the current optimal structure during the search process; the candidate expert portraits with a structural confidence score less than or equal to the set score threshold will be delayed in participating in the aggregation process during the search process;

[0044] S44. Conduct a pairwise combination consistency review on the potential cognitive label vectors in the candidate expert portraits. The conflict discrimination conditions include:

[0045] Whether there are semantic repetitions, inclusions, or hyponymy relationships in the first potential cognitive label vector;

[0046] Whether there are functional overlaps or logical inconsistencies in the task capabilities mapped by the second potential cognitive label vector;

[0047] S45. When there are two or more potential cognitive label vectors in the candidate expert portraits that meet any of the conflict determination conditions, they are marked as label conflict bodies and will not be used as a reference for updating the positions of the next-generation whale individuals to avoid the spread of incorrect structures.

[0048] S46. In each iteration of the multi-strategy enhanced whale optimization algorithm, only retain the candidate expert portraits that meet the following two conditions to obtain the optimized expert portraits:

[0049] The first structural confidence score is not lower than the median of the structural confidence scores of all candidate expert portraits;

[0050] The second is not a label conflict body.

[0051] Optionally, the S5 specifically includes:

[0052] S51. Sort the optimized expert portraits in the iteration order, and use each optimized expert portrait as a node in the graph. The node data includes the potential cognitive label vector corresponding to the optimized expert portrait and the arrangement order;

[0053] S52. Compare the potential cognitive label vectors of any two adjacent nodes in order, identify the evolutionary behavior according to the differences. When a certain potential cognitive label vector is newly added in the latter node, it is marked as growth, and when it is missing, it is marked as contraction. If the original potential cognitive label vector no longer appears and is replaced by a functionally equivalent or semantically identical potential cognitive label vector, it is marked as replacement;

[0054] S53. Identify the evolutionary behavior results based on differences, establish directed edges between adjacent nodes, and label each edge with the corresponding potential cognitive label vector evolution type, including growth, contraction, and replacement relationships, to generate an expert portrait evolutionary path map.

[0055] Optionally, the S6 specifically includes:

[0056] S61. Collect the real behavior feedback data of the expert portrait, where the real behavior feedback data includes the click situation of the user on the recommended expert portrait, the completion rate of the matching task, and the acceptance or correction operation of the user on the portrait label;

[0057] S62. Use the expert portrait evolutionary path map to determine the structural change method between each expert portrait and the previous state, and for different evolutionary types, count the corresponding behavior feedback data;

[0058] S63. Construct a reward function to calculate the joint reward signal of the expert portrait:

[0059] ;

[0060] Among them, represents the joint reward signal, , , and represent the weight coefficients, represents the click-through rate of the expert portrait recommendation, represents the task completion rate, represents the label acceptance rate, represents the correction frequency of the user on the portrait label;

[0061] S64. Use the generated joint reward signal to update the structural parameters of the variable-dimensional mimicry generation network:

[0062] ;

[0063] Among them, represents the updated structural parameters, represents the structural parameters before update, represents the learning rate of the structural parameters, represents the joint reward signal for the structural parameters before update gradient;

[0064] S65. At the same time, use the joint reward signal to update the search parameters of the multi-strategy enhanced whale optimization algorithm:

[0065] ;

[0066] Among them, represents the updated search parameters, represents the search parameters before the update, represents the update step size;

[0067] S66. Repeat steps S61 to S65. After the dynamic update of the expert portrait is completed in each round of iteration, the behavior feedback data for the next round is collected again, continuously driving the bidirectional update of the structure parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, so as to realize the dynamic update of the expert portrait and the optimized evolution of potential cognitive labels.

[0068] The beneficial effects of the present invention are as follows:

[0069] First of all, the present invention realizes the deep fusion and unified expression of multi-modal heterogeneous features by uniformly collecting and preprocessing structured data, behavioral data, and semantic data, respectively generating knowledge vectors, behavioral vectors, and semantic vectors, and constructing an expert cognitive representation tensor with potential cognitive labels as the core through a unified fusion mechanism, effectively improving the semantic integrity and cognitive representation ability of the expert portrait. In particular, the behavioral vector participates in the fusion as the dominant feature, enabling the expert portrait to have behavioral perception ability and providing a plastic basis for subsequent structure evolution.

[0070] Secondly, the variable-dimensional mimicry generation network proposed by the present invention has structural variability and active evolution ability. The network integrates a label growth unit, a structure compression unit, and a dimension deformation unit inside, and can dynamically expand, contract, and adjust the structure of the labels according to the internal information of the expert cognitive representation tensor, generating multiple portrait candidates with different structures, fully demonstrating the label combination forms of the expert in different cognitive states, and significantly improving the portrayal accuracy of the portrait model for potential cognitive ability and the structural expression flexibility.

[0071] In addition, the present invention introduces a multi-strategy enhanced whale optimization algorithm to optimize the structure of the portrait candidates, and introduces a structure confidence evaluation mechanism and a label conflict discrimination mechanism in the optimization process. By calculating the historical behavior coverage, task participation frequency, and semantic matching degree of the labels in each candidate portrait, the credibility of the portrait structure is dynamically controlled; at the same time, semantic redundancy, logical conflicts, and functional repetitions between candidate labels are identified, and unreasonable candidate bodies are filtered, effectively avoiding the problems of label misuse and semantic overlap, and improving the label rationality and cognitive expression consistency of the optimized portrait.

[0072] Furthermore, the present invention designs a mechanism for generating an evolution path map of expert portraits. The system can automatically identify the relationships of label growth, contraction, and replacement among portraits at different stages, establish a label evolution path in the form of a directed edge, and form a structural map with temporal attributes. This map not only records the change trajectory of the portrait structure during the evolution process but also provides a traceable structural basis for subsequent portrait causal analysis, trend judgment, and evolution prediction, enhancing the interpretability and trust of the portrait system in practical applications.

[0073] Finally, by introducing the idea of reinforcement learning, the present invention designs a joint reward mechanism based on behavioral feedback, which converts the behavioral performances of expert portraits in the real system, such as click-through rate, task completion rate, label acceptance rate, and correction frequency, into quantitative reward signals, and then synchronously drives the joint update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm. This mechanism realizes the full-cycle closed-loop evolution of the expert portrait model from generation to deployment, feedback, and optimization, significantly improving the dynamic update ability and online adaptive level of the expert portrait system. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0075] Figure 1 is the overall flowchart of a method for depicting and dynamically updating expert portraits based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0077] Refer to Figure 1 , a method for depicting and dynamically updating expert portraits based on reinforcement learning, includes the following steps:

[0078] S1. Collect the structural data, behavioral data, and semantic data of experts, and perform preprocessing to generate knowledge vectors, behavioral vectors, and semantic vectors;

[0079] S2. Fuse the knowledge vectors, behavioral vectors, and semantic vectors to construct an expert cognitive representation tensor with potential cognitive labels as the core;

[0080] S3. Input the expert cognitive representation tensor into a variable-dimensional mimicry generation network, which includes a label growth unit, a structure compression unit, and a dimension deformation unit, to generate an expert portrait candidate;

[0081] S4. Use a multi-strategy enhanced whale optimization algorithm to optimize the candidate expert portraits, and generate optimized expert portraits with portrait confidence, potential cognitive label consistency, and expert historical behavior compatibility as constraints;

[0082] S5. Generate an expert portrait evolution path map based on the optimized expert portraits, where the nodes represent the states of the optimized expert portraits, and the edges represent growth, contraction, and replacement relationships;

[0083] S6. Collect real behavior feedback data of the expert portraits, and combine with the expert portrait evolution path map to generate a joint reward signal through a reward function, and use the joint reward signal to drive the bidirectional update of the structure parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm to achieve dynamic update of the expert portraits and optimized evolution of potential cognitive labels.

[0084] The present invention constructs a complete method for depicting and dynamically updating expert portraits with self-adaptive evolution ability, and for the first time realizes the full-process closed-loop portrait modeling from multi-source data collection, deep fusion, structure generation, optimization regulation to behavior feedback drive. This method not only breaks the limitations of traditional static modeling and fixed structure output of expert portraits, but also integrates reinforcement learning strategies to achieve the continuous evolution of expert portraits and cognitive label structures under the influence of actual behaviors. Through the collaborative design of the variable-dimensional mimicry generation network and the multi-strategy enhanced whale optimization algorithm, the flexible adjustment of the portrait structure in the two stages of generation and optimization is realized, and combined with the click behavior, task completion status and user feedback collected in the actual system, the portrait is dynamically corrected. This method has strong plasticity, high expressiveness and behavior response ability, effectively improves the accuracy, timeliness and system adaptability of expert portraits, and has good practical application and promotion value.

[0085] In this embodiment, the S2 specifically includes:

[0086] S21. Convert the knowledge vector, behavior vector, and semantic vector to a unified feature dimension through linear mapping;

[0087] S22. Taking the behavior vector as the leading factor, calculate the interaction weights between the behavior vector and the knowledge vector and the semantic vector and perform weighted processing, and splice the weighted knowledge vector, semantic vector and behavior vector to form a fused feature vector;

[0088] S23. Construct a potential cognitive label embedding matrix, which is obtained by embedding and encoding a set of preset labels, and the set of labels is derived from long-term accumulated expert behavior logs, portrait label history, and semantic topic analysis;

[0089] S24. Calculate the tensor product of the fused feature vector and the latent cognitive label embedding matrix to generate an expert cognitive representation tensor. The expert cognitive representation tensor is a third-order tensor, including a label dimension, a feature dimension, and a time dimension, representing the feature combination under the latent cognitive label dimension, the label response degree, and the time variation characteristics in the expert behavior evolution process respectively.

[0090] Through the construction process of the expert cognitive representation tensor in step S2, the problems of inconsistent multi-modal feature fusion dimensions and unclear weak semantic expression of behavior features in the existing expert portrait methods are solved. In the present invention, the knowledge vector, the behavior vector, and the semantic vector are mapped to a unified feature space through linear transformation, and the key interaction relationships are extracted by using the behavior vector-dominated attention mechanism, thereby significantly improving the combination degree between the label semantics and the behavior logic. At the same time, the present invention introduces a latent cognitive label embedding matrix, making the portrait construction not only depend on the explicit label set, but also able to mine latent cognitive intentions from historical behaviors and semantic themes, constructing a third-order expert cognitive tensor with clear structure and distinct expression levels, providing a highly plastic input basis for subsequent structure generation and dynamic optimization, and enhancing the adaptability and expression accuracy of the model to the evolution of the cognitive structure.

[0091] In this embodiment, S3 specifically includes:

[0092] S31. Slice the expert cognitive representation tensor in the label dimension direction to extract a set of latent cognitive label vectors, denoted as , where represents the th latent cognitive label vector, and represents the total number of latent cognitive label vectors;

[0093] S32. Input each latent cognitive label vector into the label growth unit to calculate the growth probability:

[0094] ;

[0095] where, represents the growth probability of the th latent cognitive label vector, represents the activation function, and represent weight matrices, represents the Gaussian error linear unit activation function, and represent bias vectors, represents the label amplitude adjustment factor, represents 's norm;

[0096] If If it exceeds the set threshold, the label expansion operation is triggered to add new labels to the label dimension of the expert cognitive representation tensor;

[0097] S33. Input the updated expert cognitive representation tensor into the structure compression unit and calculate the retention importance score of each potential cognitive label vector:

[0098] ;

[0099] in, Indicates The preservation importance score of the latent cognitive label vector, and represents the weight coefficient, Indicates The average response value of the latent cognitive label vector in the attention mechanism, Indicates Potential cognitive label vectors are in front Activation variance in wheel structure;

[0100] like If it is lower than the set threshold, it is removed from the expert cognitive representation tensor structure;

[0101] S34. Input the compressed tensor into the dimension deformation unit, and perform affine perturbation and nonlinear transformation on each label vector:

[0102] ;

[0103] in, Indicates candidate label vectors, and represents the structural perturbation matrix, represents the disturbance amplitude adjustment coefficient, represents the offset vector, represents the Gaussian noise perturbation tensor;

[0104] S35. Calculate the structural difference index of the candidate label vectors and screen them to obtain the expert portrait candidates:

[0105] ;

[0106] in, Represents the comprehensive structural difference of the candidate label vector, represents the historical reference label vector, express The adjustment coefficient of the divergence term, Represents the candidate label vector distribution Vector with historical reference labels Between Divergence;

[0107] If it falls within the preset structural deviation range, the corresponding expert cognitive representation tensor is retained as a candidate for the expert portrait.

[0108] By disassembling the structure of the variable-dimensional mimicry generation network in step S3, a label growth unit, a structure compression unit, and a dimension deformation unit are respectively constructed, realizing the controllable transformation and multi-form reconstruction of the structure of the expert cognitive representation tensor, effectively solving the problems of rigidity of the label structure and lack of adaptive generation ability in the expert portrait. The label growth unit can trigger the growth of new potential labels according to the current cognitive state of the expert, the structure compression unit effectively eliminates low-frequency weakly associated labels, and the dimension deformation unit provides the deformation ability of the label representation in multiple spaces, enabling the portrait to have the ability of non-linear structure migration and tensor space reconstruction. This network can generate multiple candidates with different structures, thereby improving the coverage rate of the portrait for the diverse cognitive expressions of experts and providing a structural selection space for subsequent optimization. It is the basic module for achieving the core goal of cognitive label evolution in the present invention.

[0109] In this embodiment, S4 specifically includes:

[0110] S41. Configure a confidence evaluation module for the expert portrait candidate. The confidence evaluation module makes a fusion judgment based on the coverage of the label dimension of the expert cognitive representation tensor in the expert's historical behavior trajectory, the participation frequency in the completed tasks, and the matching degree with the expert's semantic description information, forming a confidence input source;

[0111] S42. Calculate the structural confidence score of the expert portrait candidate based on the weighted summation method:

[0112] ;

[0113] Wherein, represents the structural confidence score of the th expert portrait candidate, represents the total number of potential cognitive label vectors, , and represent the weighting coefficients, represents the coverage of the th potential cognitive label vector in the expert's historical behavior trajectory, represents the participation frequency of the th potential cognitive label vector in the completed tasks, represents the matching degree of the th potential cognitive label vector with the expert's semantic description information;

[0114] S43. Use the structural confidence score to regulate the search direction control parameters of individuals in the multi-strategy enhanced whale optimization algorithm, where the candidate expert portraits with a structural confidence score greater than the set score threshold are prioritized to approach the current optimal structure during the search process; the candidate expert portraits with a structural confidence score less than or equal to the set score threshold are delayed in participating in the aggregation process during the search process;

[0115] S44. Pairwise combine the potential cognitive label vectors in the candidate expert portraits for consistency review, and the conflict discrimination conditions include:

[0116] Whether there are semantic repetitions, inclusions, or hyponymy relationships in the first potential cognitive label vector;

[0117] Whether there are functional overlaps or logical inconsistencies in the task capabilities mapped by the second potential cognitive label vector;

[0118] S45. When there are two or more potential cognitive label vectors in the candidate expert portraits that meet any of the conflict determination conditions, they are marked as label conflict bodies and will not be used as a reference for updating the positions of the next-generation whale individuals, avoiding the spread of incorrect structures;

[0119] S46. In each iteration of the multi-strategy enhanced whale optimization algorithm, only retain the candidate expert portraits that meet the following two conditions to obtain the optimized expert portraits:

[0120] The first structural confidence score is not lower than the median of the structural confidence scores of all candidate expert portraits;

[0121] The second is not a label conflict body.

[0122] In step S4, a portrait confidence guidance and label conflict discrimination mechanism is introduced in the optimization link, significantly improving the accuracy and label combination rationality of candidate expert portraits during the structural optimization process. Through the confidence module, a weighted evaluation is carried out on the label historical behavior coverage, task participation frequency, and semantic matching degree of each portrait, enabling the optimization algorithm to preferentially retain candidate bodies with credible structures and avoiding the misleading convergence of high-deviation structures. The label conflict discrimination module can identify semantic redundancy, hyponymy conflicts, and functional overlap problems between labels, filtering out semantically inconsistent label combinations from the source, effectively improving the logical consistency and semantic clarity of the optimized portraits. The two work together to enhance the convergence robustness of the whale optimization algorithm in the complex label structure space, reducing structural redundancy, deviation, and ambiguity phenomena during the optimization process, and ensuring that the finally output expert portraits are more accurate, stable, and interpretable.

[0123] In this embodiment, S5 specifically includes:

[0124] S51. Sort the optimized expert portraits in the iteration order, and use each optimized expert portrait as a node in the graph. The node data includes the potential cognitive label vector corresponding to the optimized expert portrait and the arrangement order.

[0125] S52. Compare the potential cognitive label vectors of any two adjacent nodes in order, identify the evolution behavior based on the differences. When a certain potential cognitive label vector is newly added in the latter node, it is marked as growth; when it is missing, it is marked as shrinkage. If the original potential cognitive label vector no longer appears and is replaced by a functionally equivalent or semantically identical potential cognitive label vector, it is marked as replacement.

[0126] S53. Based on the results of identifying the evolution behavior from the differences, establish directed edges between adjacent nodes, and label each edge with the corresponding evolution type of the potential cognitive label vector, including growth, shrinkage, and replacement relationships, to generate an expert portrait evolution path graph.

[0127] In step S5, by constructing an expert portrait evolution path graph, the complete modeling of the structural evolution trajectory of the portrait during the dynamic generation process is realized, filling the gap in the traditional expert portrait method that lacks a mechanism for recording and tracking structural evolution. The present invention takes the portrait structure state as nodes and the growth, shrinkage, and replacement of potential cognitive labels as edges to establish a directed path graph, which not only enables the system to track the process of label introduction, elimination, and replacement, but also helps with portrait causal interpretation, label evolution prediction, and structural traceability analysis. In addition, the introduction of the graph-based structure also provides the system with the ability to manage label structure versions, supports the visual display and time series modeling of expert portraits, enhances the controllability and interpretability of the model, and provides a strong evolutionary evidence basis for the intelligent decision-making recommendation system.

[0128] In this embodiment, S6 specifically includes:

[0129] S61. Collect the real behavior feedback data of the expert portrait, where the real behavior feedback data includes the click situation of the user on the recommended expert portrait, the completion rate of the matching task, and the acceptance or correction operation of the user on the portrait label.

[0130] S62. Use the expert portrait evolution path graph to determine the structural change method between each expert portrait and the previous state, and count the corresponding behavior feedback data for different evolution types.

[0131] S63. Construct a reward function to calculate the joint reward signal of the expert portrait:

[0132] ;

[0133] where represents the joint reward signal, , , and represent weight coefficients, represents the click-through rate of the expert portrait recommendation, represents the task completion rate, represents the label acceptance rate, represents the correction frequency of the user for the portrait label;

[0134] S64. Update the structural parameters of the variable-dimensional mimicry generation network using the generated joint reward signal:

[0135] ;

[0136] wherein, represents the updated structural parameters, represents the structural parameters before update, represents the learning rate of the structural parameters, represents the joint reward signal for the structural parameters before update gradient;

[0137] S65. At the same time, update the search parameters of the multi-strategy enhanced whale optimization algorithm using the joint reward signal:

[0138] ;

[0139] wherein, represents the updated search parameters, represents the search parameters before update, represents the update step size;

[0140] S66. Repeat steps S61 to S65. After each round of iteration completes the dynamic update of the expert portrait, re-collect the next round of behavior feedback data, and continuously drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, so as to realize the dynamic update of the expert portrait and the optimization evolution of the potential cognitive labels.

[0141] Step S6 realizes the two-way dynamic adjustment of the portrait model structure and optimization parameters by collecting the real behavior feedback of the expert portrait in actual applications and constructing a joint reward function based on the expert portrait evolution path map. The reward function comprehensively considers multiple key behavior indicators such as user click-through rate, task completion rate, label acceptance rate, and correction frequency, constituting a quantitative signal that accurately reflects the quality of the expert portrait. By this reward signal, the structure parameters of the variable-dimensional mimicry generation network and the search parameters of the whale optimization algorithm are updated simultaneously, enabling the model to have the capabilities of behavior response and self-adjustment, and breaking through the problems of broken feedback chains, separated parameter updates and structure evolution in traditional model optimization. This mechanism realizes the real-time closed-loop linkage between the portrait generation process and behavior performance, enabling the expert portrait to continuously evolve adaptively and ultimately achieving the unified goal of optimal structure, accurate semantics, and behavioral consistency.

[0142] Example 1:

[0143] To verify the feasibility of the present invention in implementation, the present invention is applied to the scenario of an expert intelligent recommendation system in a scientific research institution to solve the long-existing problems of strong static nature of expert portraits, inability to dynamically evolve and update portrait labels, and low portrait accuracy. This scientific research institution has a traditional expert portrait system that has long relied on simple structured data (papers, projects, and titles published by experts) and semantic data (text descriptions of experts' research directions) to model expert portraits, and is unable to capture the behavior changes and cross-domain migrations of experts in real time. Therefore, problems such as insufficient recommendation accuracy and low expert response rate often occur in scenarios such as recommending experts to participate in scientific research projects, review tasks, and conference reports, and the overall user satisfaction is not high. In addition, due to the inability to update expert portraits in real time and dynamically, the manual cycle for each update of the expert portrait in this system is at least one month, seriously lagging behind the changes in the actual behavior status of experts and reducing the practical value of the recommendation service.

[0144] To address the above problems, the present invention has been specifically applied and practiced in the expert intelligent recommendation system of this scientific research institution. Specifically, in this embodiment, the structured data (papers, projects, titles, research fields), behavioral data (response records of clicking on expert recommendation tasks, accepting tasks, and cross-domain task migrations), and semantic data (research direction descriptions, abstracts, speech manuscripts) of a total of 300 experts in the whole hospital are first collected to generate three types of feature vectors: knowledge, behavior, and semantics. Subsequently, using the fusion method of the present invention, an expert cognitive representation tensor with potential cognitive labels as the core is constructed to deeply fuse complex heterogeneous features into a unified portrait structure.

[0145] Furthermore, the variable-dimensional mimicry generation network of the present invention is applied to the expert system, and the label structure is updated once a week based on newly collected expert data. During the application process, the label growth unit in the network automatically generates new potential cognitive labels according to the activity level of experts in new fields or hot issues; the structure compression unit deletes corresponding labels when the activity level of experts in a certain field decreases significantly; the dimension deformation unit flexibly adjusts the portrait structure dimension according to the actual situation, so that the portrait of each expert always accurately reflects their current state.

[0146] Meanwhile, the present invention introduces a multi-strategy enhanced whale optimization algorithm. Based on the historical behavior coverage, task participation frequency, and semantic matching degree of the expert portrait, it calculates the credibility of the portrait and the label conflict situation, dynamically optimizes and filters the expert portrait every week, eliminates redundant labels and unreasonable combinations, and ensures that the final expert portrait structure has a rigorous logic and clear semantics. Then, through the expert portrait evolution path graph, the historical changes of each expert's portrait are visually recorded, ensuring that the change process of the expert portrait is traceable and providing a more powerful basis for recommendation decisions.

[0147] The actual application statistical data over a one-year period shows that the present invention has significantly improved the accuracy of the expert portrait and user satisfaction. The specific data comparison is shown in Table 1 below.

[0148] Table 1 Annual comparison table of the recommendation effects of the expert portrait system of the present invention and the traditional method

[0149]

[0150] It can be seen from the data in Table 1 above that the method for depicting and dynamically updating the expert portrait proposed by the present invention has achieved a significant improvement in the recommendation task response rate. The expert response rate has increased from 73.2% of the traditional system to 92.5%; the matching score between the recommended experts and the tasks has also increased significantly, from the original 3.8 points to 4.7 points, and the overall user satisfaction score has even increased from 80.2 points to 95.3 points, reflecting an overall improvement in the accuracy and matching degree of the expert portrait. In addition, the success rate of the present invention in new label prediction is as high as 88.6%. The average monthly label correction frequency of the expert portrait has also decreased from the traditional 1.2 times to 0.15 times, and the average update cycle of the expert portrait has been shortened from the past 35 days to 7 days. This fully demonstrates that the label dynamic evolution and expert behavior feedback mechanism of the present invention effectively solves the static and lagging problems of the traditional expert portrait system, truly and accurately reflects the behavior changes and cognitive state evolution of experts, and can effectively improve the quality, real-time performance, and interpretability of expert recommendation services, with strong practical application value and obvious technical advantages.

[0151] In this embodiment, by deploying the method for depicting and dynamically updating expert portraits proposed by the present invention in an actual scientific research institution, the problems existing in the traditional expert portrait system, such as static structure, lagging update, redundant tags, and low recommendation accuracy, are significantly solved. During the implementation process, an expert cognitive representation tensor is constructed by integrating structural, behavioral, and semantic data. The portrait structure is adaptively evolved by combining a variable-dimensional mimicry generation network, and a confidence-guided and tag conflict discrimination mechanism is introduced to optimize the candidate portrait, greatly improving the accuracy of portrait expression and the rationality of the structure. Experimental data shows that the present invention can achieve weekly updates of expert portraits, the response rate of the recommendation task is increased to 92.5%, the success rate of new tag prediction reaches 88.6%, and the frequency of tag correction is significantly reduced, demonstrating strong feedback response ability and evolutionary optimization effect. Overall, this embodiment verifies that the present invention has the advantages of real-time modeling, accurate expression, behavior perception, and continuous optimization, and can be widely applied to actual application scenarios such as expert recommendation, scientific research management, and talent mining.

[0152] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.

Claims

1. A method for characterizing and dynamically updating expert portraits based on reinforcement learning, characterized in that: The steps include: S1, collect the structural data, behavioral data and semantic data of experts, and pre-process them to generate knowledge vectors, behavioral vectors and semantic vectors; S2, fuse the knowledge vector, behavior vector and semantic vector to construct the expert cognitive representation tensor with potential cognitive labels as the core; S3, inputting the expert cognitive representation tensor into a variable-dimensional mimicry generation network, wherein the variable-dimensional mimicry generation network includes a label growth unit, a structure compression unit, and a dimension deformation unit to generate an expert portrait candidate; S4. Use multi-strategy enhanced whale optimization algorithm to optimize the expert portrait candidates, and generate optimized expert portraits with structural confidence, consistency of potential cognitive labels and compatibility of expert historical behavior as constraints; S5. Generate an expert portrait evolution path graph based on the optimized expert portrait, where the nodes represent the optimized expert portrait status and the edges represent growth, contraction and replacement relationships; S6. Collect the real behavior feedback data of the expert portrait, and combine it with the expert portrait evolution path map, generate a joint reward signal through the reward function, and use the joint reward signal to drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, so as to realize the dynamic update of the expert portrait and the optimized evolution of the potential cognitive labels.

2. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S2 specifically includes: S21, transforming the knowledge vector, behavior vector and semantic vector into a unified feature dimension through linear mapping; S22, taking the behavior vector as the main factor, calculating the interaction weights between the behavior vector, the knowledge vector and the semantic vector and performing weighted processing, and concatenating the weighted knowledge vector, the semantic vector and the behavior vector to form a fusion feature vector; S23, constructing a potential cognitive label embedding matrix, wherein the potential cognitive label embedding matrix is ​​obtained by embedding coding a set of labels, wherein the set of labels is derived from long-term accumulated expert behavior logs, portrait label history, and semantic topic analysis; S24. Perform a tensor product calculation on the fused feature vector and the potential cognitive label embedding matrix to generate an expert cognitive representation tensor. The expert cognitive representation tensor is a third-order tensor, including a label dimension, a feature dimension, and a time dimension, which respectively represent the feature combination under the potential cognitive label dimension, the label response degree, and the time change characteristics during the evolution of the expert behavior.

3. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S3 specifically includes: S31. Slice the expert cognitive representation tensor in the label dimension direction to extract the potential cognitive label vector set, denoted as ,in Indicates potential cognitive label vectors, represents the total number of potential cognitive label vectors; S32. Input each potential cognitive label vector into the label growth unit and calculate the growth probability: ; in, Indicates The growth probability of potential cognitive label vectors, represents the activation function, and represents the weight matrix, represents the Gaussian error linear unit activation function, and represents the bias vector, represents the label amplitude adjustment factor, express The L2 norm of ; like If the threshold is exceeded, the label expansion operation is triggered to add new labels to the label dimension of the expert cognitive representation tensor; S33. Input the updated expert cognitive representation tensor into the structure compression unit and calculate the retention importance score of each potential cognitive label vector: ; in, Indicates The preservation importance score of the latent cognitive label vector, and represents the weight coefficient, Indicates The average response value of the latent cognitive label vector in the attention mechanism, Indicates Potential cognitive label vectors are in front Activation variance in wheel structure; like If it is lower than the set threshold, it is removed from the expert cognitive representation tensor structure; S34. Input the compressed tensor into the dimension deformation unit, and perform affine perturbation and nonlinear transformation on each label vector: ; in, Indicates candidate label vectors, and represents the structural perturbation matrix, represents the disturbance amplitude adjustment coefficient, represents the offset vector, represents the Gaussian noise perturbation tensor; S35. Calculate the structural difference index of the candidate label vectors and screen them to obtain the expert portrait candidates: ; in, Represents the comprehensive structural difference of the candidate label vector, represents the historical reference label vector, represents the adjustment coefficient of the KL divergence term, Represents the candidate label vector distribution Vector with historical reference labels The KL divergence between like If it falls into the preset structural deviation range, the corresponding expert cognitive representation tensor is retained as the expert portrait candidate.

4. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S4 specifically includes: S41, assigning a confidence evaluation module to the expert portrait candidate, wherein the confidence evaluation module performs fusion judgment based on the coverage of the label dimension of the expert cognitive representation tensor in the expert's historical behavior trajectory, the participation frequency of completed tasks, and the matching degree with the expert's semantic description information to form a confidence input source; S42. Calculate the structural confidence score of the expert portrait candidate based on the weighted sum method: ; in, Indicates The structural confidence scores of the expert portrait candidates, represents the total number of potential cognitive label vectors, , and represents the weighting coefficient, Indicates The coverage of potential cognitive label vectors in the expert's historical behavior trajectory, Indicates The frequency of participation of latent cognitive label vectors in completed tasks, Indicates The matching degree between the potential cognitive label vector and the expert semantic description information; S43, using the structure confidence score to regulate the search direction control parameters of individuals in the multi-strategy enhanced whale optimization algorithm, wherein the expert portrait candidates whose structure confidence score is greater than the set score threshold will preferentially approach the current optimal structure during the search process; the expert portrait candidates whose structure confidence score is less than or equal to the set score threshold will delay participating in the aggregation process during the search process; S44. Perform consistency check on the potential cognitive label vectors in the expert profile candidates. The conflict judgment conditions include: Whether there is duplication, inclusion, or hyponymy in the semantics of the first potential cognitive label vector; Whether there is functional overlap or logical inconsistency in the task capabilities mapped by the second latent cognitive label vector; S45. When there are two or more potential cognitive label vectors in the expert portrait candidate that meet any conflict judgment condition, they are marked as label conflict bodies and will not be used as reference for updating the position of the next generation of whale individuals to avoid the spread of erroneous structures; S46. In each iteration of the multi-strategy enhanced whale optimization algorithm, only the expert portrait candidates that meet the following two conditions are retained to obtain the optimized expert portrait: The first structure confidence score is not less than the median of the structure confidence scores of all expert portrait candidates; The second is not a label conflict body.

5. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S5 specifically includes: S51, sorting the optimized expert portraits according to the iteration order, and taking each optimized expert portrait as a node in the graph, wherein the node data includes the potential cognitive label vector and the arrangement order corresponding to the optimized expert portrait; S52, compare potential cognitive label vectors of any two sequentially adjacent nodes, identify evolutionary behaviors based on differences, and mark a potential cognitive label vector as growth when it is added to the latter node, and mark it as contraction when it is missing; if the original potential cognitive label vector no longer appears and is replaced by a functionally equivalent or semantically identical potential cognitive label vector, it is marked as replacement; S53. Based on the results of differential recognition evolutionary behavior, directed edges are established between adjacent nodes, and the corresponding potential cognitive label vector evolution type is marked for each edge, including growth, contraction and replacement relationships, to generate an expert portrait evolution path map.

6. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S6 specifically includes: S61, collecting real behavior feedback data of expert portraits, wherein the real behavior feedback data includes the user's click situation on the recommended expert portraits, the completion rate of the matching task, and the user's acceptance or correction operation on the portrait label; S62, using the expert portrait evolution path map to determine the structural change mode between each expert portrait and the previous state, and for different evolution types, statistically analyzing the corresponding behavioral feedback data; S63. Construct a reward function to calculate the joint reward signal of the expert portrait: ; in, represents the joint reward signal, , , and represents the weight coefficient, Indicates the click-through rate of expert portrait recommendations, represents the task completion rate, represents the label acceptance rate, Indicates the frequency of users correcting portrait labels; S64. Use the generated joint reward signal to update the structural parameters of the variable-dimensional mimicry generation network: ; in, represents the updated structural parameters, represents the structural parameters before updating, represents the structural parameter learning rate, Represents the joint reward signal For the structural parameters before updating The gradient of S65. Simultaneously use the joint reward signal to update the search parameters of the multi-strategy enhanced whale optimization algorithm: ; in, Represents the updated search parameters, Indicates the search parameters before updating. represents the update step size; S66. Repeat steps S61 to S65. After each round of iteration, after the dynamic update of the expert portrait is completed, the next round of behavioral feedback data is collected again to continuously drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, thereby realizing the dynamic update of the expert portrait and the optimized evolution of potential cognitive labels.

Citation Information

Patent Citations

  • Rolling bearing fault diagnosis method based on IWOA-ELM

    CN115096590A

  • Method and apparatus for generating user portrait, and device and medium

    WO2021189922A1