Expert portrait description and dynamic updating method based on reinforcement learning
By introducing reinforcement learning methods in the expert portrait system, building a variable structure network and using a two-way optimization mechanism driven by behavior feedback, the problem of difficulty in dynamic update of expert portraits and fusion of multimodal features in the existing technology is solved, and efficient and adaptive expert portrait modeling is achieved.
Patent Information
- Application Number
- CN202510422163.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing expert portrait methods are difficult to truly reflect the changes in the expert's dynamic behavior patterns and cognitive state, cannot meet the needs of continuous modeling and predictive recommendation, and lack the understanding of the deep semantic relationships between multimodal features and the adaptive deformation ability of the portrait structure.
Using reinforcement learning-based methods, integrating reinforcement learning and label structure change-dimensional methods, building a variable structure network, introducing a two-way optimization mechanism driven by behavior feedback, and realizing dynamic portrayal and evolutionary updates of expert portraits.
It realizes the accurate expression of expert portraits, structural adaptability, optimization closed-loop and controllability of the expert portrait system in complex tasks, and improves the adaptability, real-timeness and interpretability of the expert portrait system in complex tasks.
Smart Images

Figure CN119918642A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reinforcement learning, and in particular to a method for characterizing and dynamically updating expert portraits based on reinforcement learning. Background Art
[0002] With the development of artificial intelligence, knowledge graphs and big data analysis technologies, expert portraits, as a means of modeling the multi-dimensional characteristics of individual experts in specific fields, such as knowledge, ability, behavioral habits and research interests, have been widely used in tasks such as scientific research talent recommendation, team intelligent formation, expert scheduling and academic decision support. The existing expert portrait system is mainly based on the idea of static feature modeling, that is, by analyzing structural data (such as papers, projects, titles, research fields) and semantic data (such as academic profiles, research interests, etc.), a relatively fixed set of labels and weight representations are constructed to describe the knowledge and ability of experts. However, in real scenarios, the behavior patterns, research directions and cross-domain migration behaviors of experts have obvious dynamic evolution characteristics. The traditional static portrait method is difficult to truly reflect the changes in the cognitive state of experts and cannot meet the actual needs of continuous modeling and predictive recommendation of experts.
[0003] From a methodological perspective, traditional expert portrait methods mostly use vector space models or expert representations based on topic models. Their fusion methods often rely on simple splicing or linear weighting, and lack an understanding of the deep semantic relationship between multimodal features. Especially when processing behavioral data (such as task responses, click paths, and cross-domain migration), their time series and behavior-driven characteristics are often ignored, and they are only used as auxiliary information for scoring calculations. This results in the inability of the portrait model to extract dynamic label signals with cognitive value from the expert's behavioral trajectory, further limiting the depth of modeling of the label evolution mechanism. In addition, most existing methods do not consider the variability and adaptive evolution of the expert portrait structure itself. Once the portrait label set is determined, it is difficult to expand or shrink the structure as the behavior evolves, which is not conducive to the continuous characterization and discovery of the expert's potential capabilities.
[0004] In terms of optimization mechanism, the parameter learning of current portrait models mostly relies on supervised learning or heuristic rule matching based on feature similarity, lacking a clear feedback loop and self-evolution capabilities. Although some work has attempted to introduce reinforcement learning strategies to optimize the expert recommendation process, it often only reinforces the scores at the output level, without linking with the intrinsic evolution mechanism of the portrait structure, resulting in low feedback utilization efficiency, slow learning convergence, and unclear evolution path. At the same time, the semantic conflicts and redundant label problems in label combinations have not been effectively solved. In the absence of fine structural control, semantically repeated or logically inconsistent label structures are often generated, affecting the interpretability and credibility of the final portrait.
[0005] In addition, the lack of a path structure recording mechanism in expert portrait modeling is also a major flaw in current technology. Since expert portraits may undergo multiple updates and evolutions at multiple stages, if the structural change process between portraits is not explicitly modeled, the system will not be able to provide a basis for the evolution of the portraits, and will not be able to track the logical process of the generation, transfer, or disappearance of a certain label. This poses a serious obstacle to the transparent interpretation and behavior tracing of expert portraits in application systems, and also limits the actual application capabilities of expert portrait models in high-confidence recommendations, expert behavior prediction, and long-term learning tasks.
[0006] Based on this, the existing expert portrait methods still have obvious defects in the following aspects: First, there is a lack of modeling of the deep fusion mechanism between structural, behavioral and semantic multi-source data, and it is impossible to construct a high-dimensional cognitive representation with unified semantic expression capabilities; second, there is a lack of adaptive deformation and generation capabilities at the portrait structure level, and it is impossible to support the evolution of the label structure as the expert behavior changes; third, there is a lack of portrait confidence guidance, label conflict suppression and behavior compatibility control mechanism in the optimization process, which affects the optimization effect and rationality of the portrait structure; fourth, a complete portrait evolution trajectory map has not been established, and there is a lack of modeling and recording of the structural evolution relationship within the life cycle of the expert portrait; fifth, there is a lack of a joint reward-driven mechanism based on behavioral feedback, and it is impossible to achieve two-way iterative updates of the portrait structure and optimization strategy, and there is a lack of effective dynamic learning capabilities.
[0007] Therefore, how to provide an expert portrait characterization and dynamic update method based on reinforcement learning is an urgent problem that technicians in this field need to solve. Summary of the invention
[0008] One purpose of the present invention is to propose a method for characterizing and dynamically updating expert portraits based on reinforcement learning. The present invention integrates reinforcement learning and label structure dimensionality change methods, constructs a variable structure network including label growth, structure compression and dimensional deformation, and introduces a bidirectional optimization mechanism driven by behavioral feedback to achieve dynamic characterization and evolutionary update of expert portraits. It has the advantages of accurate expression, structural adaptability, optimization closed loop and controllable evolution, and can effectively improve the adaptability, real-time performance and interpretability of the expert portrait system in complex tasks.
[0009] A method for characterizing and dynamically updating expert portraits based on reinforcement learning according to an embodiment of the present invention includes the following steps: S1, collect the structural data, behavioral data and semantic data of experts, and pre-process them to generate knowledge vectors, behavioral vectors and semantic vectors; S2, fuse the knowledge vector, behavior vector and semantic vector to construct the expert cognitive representation tensor with latent cognitive labels as the core; S3, inputting the expert cognitive representation tensor into a variable-dimensional mimicry generation network, wherein the variable-dimensional mimicry generation network includes a label growth unit, a structure compression unit, and a dimension deformation unit to generate an expert portrait candidate; S4. Use multi-strategy enhanced whale optimization algorithm to optimize the expert portrait candidates, and generate optimized expert portraits with portrait confidence, potential cognitive label consistency and expert historical behavior compatibility as constraints; S5. Generate an expert portrait evolution path graph based on the optimized expert portrait, where the nodes represent the optimized expert portrait status and the edges represent the growth, contraction and replacement relationships; S6. Collect the real behavior feedback data of the expert portrait, and combine it with the expert portrait evolution path map, generate a joint reward signal through the reward function, and use the joint reward signal to drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, so as to realize the dynamic update of the expert portrait and the optimized evolution of the potential cognitive labels.
[0010] Optionally, the S2 specifically includes: S21, transforming the knowledge vector, behavior vector and semantic vector into a unified feature dimension through linear mapping; S22, taking the behavior vector as the main factor, calculating the interaction weights between the behavior vector, the knowledge vector and the semantic vector and performing weighted processing, and concatenating the weighted knowledge vector, the semantic vector and the behavior vector to form a fusion feature vector; S23, constructing a potential cognitive label embedding matrix, wherein the potential cognitive label embedding matrix is obtained by embedding coding a set of labels, wherein the set of labels is derived from long-term accumulated expert behavior logs, portrait label history, and semantic topic analysis; S24. Perform a tensor product calculation on the fused feature vector and the potential cognitive label embedding matrix to generate an expert cognitive representation tensor. The expert cognitive representation tensor is a third-order tensor, including a label dimension, a feature dimension, and a time dimension, which respectively represent the feature combination under the potential cognitive label dimension, the label response degree, and the time change characteristics during the evolution of the expert behavior.
[0011] Optionally, the S3 specifically includes: S31. Slice the expert cognitive representation tensor in the label dimension direction to extract the potential cognitive label vector set, denoted as ,in Indicates potential cognitive label vectors, represents the total number of potential cognitive label vectors; S32. Input each potential cognitive label vector into the label growth unit and calculate the growth probability: ; in, Indicates The growth probability of potential cognitive label vectors, represents the activation function, and represents the weight matrix, represents the Gaussian error linear unit activation function, and represents the bias vector, represents the label amplitude adjustment factor, express of norm; like If it exceeds the set threshold, the label expansion operation is triggered to add new labels to the label dimension of the expert cognitive representation tensor; S33. Input the updated expert cognitive representation tensor into the structure compression unit and calculate the retention importance score of each potential cognitive label vector: ; in, Indicates The preservation importance score of the latent cognitive label vector, and represents the weight coefficient, Indicates The average response value of the latent cognitive label vector in the attention mechanism, Indicates Potential cognitive label vectors are in front Activation variance in wheel structure; like If it is lower than the set threshold, it is removed from the expert cognitive representation tensor structure; S34. Input the compressed tensor into the dimension deformation unit, and perform affine perturbation and nonlinear transformation on each label vector: ; in, Indicates candidate label vectors, and represents the structural perturbation matrix, represents the disturbance amplitude adjustment coefficient, represents the offset vector, represents the Gaussian noise perturbation tensor; S35. Calculate the structural difference index of the candidate label vectors and screen them to obtain the expert portrait candidates: ; in, Represents the comprehensive structural difference of the candidate label vector, represents the historical reference label vector, express The adjustment coefficient of the divergence term, Represents the candidate label vector distribution Vector with historical reference labels Between Divergence; like If it falls into the preset structural deviation range, the corresponding expert cognitive representation tensor is retained as the expert portrait candidate.
[0012] Optionally, the S4 specifically includes: S41, assigning a confidence evaluation module to the expert portrait candidate, wherein the confidence evaluation module performs fusion judgment based on the coverage of the label dimension of the expert cognitive representation tensor in the expert's historical behavior trajectory, the participation frequency of completed tasks, and the matching degree with the expert's semantic description information to form a confidence input source; S42. Calculate the structural confidence score of the expert portrait candidate based on the weighted sum method: ; in, Indicates The structural confidence scores of the expert portrait candidates, represents the total number of potential cognitive label vectors, , and represents the weighting coefficient, Indicates The coverage of potential cognitive label vectors in the expert's historical behavior trajectory, Indicates The frequency of participation of latent cognitive label vectors in completed tasks, Indicates The matching degree between the potential cognitive label vector and the expert semantic description information; S43, using the structure confidence score to regulate the search direction control parameters of individuals in the multi-strategy enhanced whale optimization algorithm, wherein the expert portrait candidates whose structure confidence score is greater than the set score threshold will preferentially approach the current optimal structure during the search process; the expert portrait candidates whose structure confidence score is less than or equal to the set score threshold will delay participating in the aggregation process during the search process; S44: Check the consistency of potential cognitive label vectors in the expert profile candidates. The conflict judgment conditions include: Whether there is duplication, inclusion, or hyponymy in the semantics of the first latent cognitive label vector; Whether there is functional overlap or logical inconsistency in the task capabilities mapped by the second latent cognitive label vector; S45. When there are two or more potential cognitive label vectors in the expert portrait candidate that meet any conflict judgment condition, they are marked as label conflict bodies and will not be used as reference for updating the position of the next generation of whale individuals to avoid the spread of erroneous structures; S46. In each iteration of the multi-strategy enhanced whale optimization algorithm, only the expert portrait candidates that meet the following two conditions are retained to obtain the optimized expert portrait: The first structure confidence score is not less than the median of the structure confidence scores of all expert portrait candidates; The second is not a label conflict body.
[0013] Optionally, the S5 specifically includes: S51, sorting the optimized expert portraits according to the iteration order, and taking each optimized expert portrait as a node in the graph, wherein the node data includes the potential cognitive label vector and the arrangement order corresponding to the optimized expert portrait; S52, compare potential cognitive label vectors of any two sequentially adjacent nodes, identify evolutionary behaviors based on differences, and mark a potential cognitive label vector as growth when it is added to the latter node, and mark it as contraction when it is missing; if the original potential cognitive label vector no longer appears and is replaced by a functionally equivalent or semantically identical potential cognitive label vector, it is marked as replacement; S53. Based on the results of differential recognition evolutionary behavior, directed edges are established between adjacent nodes, and the corresponding potential cognitive label vector evolution type is marked for each edge, including growth, contraction and replacement relationships, to generate an expert portrait evolution path map.
[0014] Optionally, the S6 specifically includes: S61, collecting real behavior feedback data of expert portraits, wherein the real behavior feedback data includes the user's click situation on the recommended expert portraits, the completion rate of the matching task, and the user's acceptance or correction operation on the portrait label; S62, using the expert portrait evolution path map to determine the structural change mode between each expert portrait and the previous state, and for different evolution types, statistically analyzing the corresponding behavioral feedback data; S63. Construct a reward function to calculate the joint reward signal of the expert portrait: ; in, represents the joint reward signal, , , and represents the weight coefficient, Indicates the click-through rate of expert portrait recommendations, represents the task completion rate, represents the label acceptance rate, Indicates the frequency of users correcting portrait labels; S64. Use the generated joint reward signal to update the structural parameters of the variable-dimensional mimicry generation network: ; in, represents the updated structural parameters, represents the structural parameters before updating, represents the structural parameter learning rate, Represents the joint reward signal For the structural parameters before updating The gradient of S65. Simultaneously use the joint reward signal to update the search parameters of the multi-strategy enhanced whale optimization algorithm: ; in, Represents the updated search parameters, Indicates the search parameters before updating. represents the update step size; S66. Repeat steps S61 to S65. After each round of iteration, after the dynamic update of the expert portrait is completed, the next round of behavioral feedback data is collected again to continuously drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, thereby realizing the dynamic update of the expert portrait and the optimized evolution of potential cognitive labels.
[0015] The beneficial effects of the present invention are: First, the present invention generates knowledge vectors, behavior vectors and semantic vectors respectively through unified collection and preprocessing of structural data, behavior data and semantic data, and on this basis, constructs an expert cognitive representation tensor with potential cognitive labels as the core through a unified fusion mechanism, achieving deep fusion and unified expression between multimodal heterogeneous features, and effectively improving the semantic integrity and cognitive representation ability of expert portraits. In particular, the behavior vector participates in the fusion as the leading feature, which enables the expert portrait to have behavioral perception ability and provides a plasticity basis for subsequent structural evolution.
[0016] Secondly, the variable-dimensional mimicry generation network proposed in the present invention has structural variability and active evolution capabilities. The network integrates label growth units, structural compression units and dimensional deformation units. It can dynamically expand, shrink and adjust the structure of labels according to the internal information of the expert cognitive representation tensor, generate multiple structurally differentiated portrait candidates, fully demonstrate the label combination forms of experts in different cognitive states, and significantly improve the portrait model's characterization accuracy of potential cognitive abilities and structural expression flexibility.
[0017] In addition, the present invention introduces a multi-strategy enhanced whale optimization algorithm to optimize the structure of the portrait candidates, and introduces a structural confidence evaluation mechanism and a label conflict discrimination mechanism during the optimization process. By calculating the historical behavior coverage, task participation frequency and semantic matching of the labels in each candidate portrait, the credibility of the portrait structure is dynamically controlled; at the same time, the semantic redundancy, logical conflict and functional duplication between candidate labels are identified, and candidates with unreasonable structures are filtered out, effectively avoiding the problems of label misuse and semantic overlap, and improving the label rationality and cognitive expression consistency of the optimized portrait.
[0018] Furthermore, the present invention designs a mechanism for generating expert portrait evolution path maps. The system can automatically identify the label growth, contraction and replacement relationship between portraits at different stages, establish label evolution paths in the form of directed edges, and form a structural map with time series attributes. This map not only records the change trajectory of the portrait structure during the evolution process, but also provides a traceable structural basis for subsequent portrait causal analysis, trend judgment and evolution prediction, enhancing the interpretability and trustworthiness of the portrait system in practical applications.
[0019] Finally, by introducing the idea of reinforcement learning, this paper designs a joint reward mechanism based on behavioral feedback, which converts the click rate, task completion rate, label acceptance rate, and correction frequency of the expert portrait in the real system into a quantitative reward signal, and then synchronously drives the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm to be jointly updated. This mechanism realizes the full-cycle closed-loop evolution of the expert portrait model from generation to deployment to feedback to optimization, significantly improving the dynamic update capability and online adaptability of the expert portrait system. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is an overall flow chart of the expert portrait characterization and dynamic updating method based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION
[0021] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0022] refer to Figure 1 , a method for characterizing and dynamically updating expert portraits based on reinforcement learning, comprising the following steps: S1, collect the structural data, behavioral data and semantic data of experts, and pre-process them to generate knowledge vectors, behavioral vectors and semantic vectors; S2, fuse the knowledge vector, behavior vector and semantic vector to construct the expert cognitive representation tensor with latent cognitive labels as the core; S3, inputting the expert cognitive representation tensor into a variable-dimensional mimicry generation network, wherein the variable-dimensional mimicry generation network includes a label growth unit, a structure compression unit, and a dimension deformation unit to generate an expert portrait candidate; S4. Use multi-strategy enhanced whale optimization algorithm to optimize the expert portrait candidates, and generate optimized expert portraits with portrait confidence, potential cognitive label consistency and expert historical behavior compatibility as constraints; S5. Generate an expert portrait evolution path graph based on the optimized expert portrait, where the nodes represent the optimized expert portrait status and the edges represent the growth, contraction and replacement relationships; S6. Collect the real behavior feedback data of the expert portrait, and combine it with the expert portrait evolution path map, generate a joint reward signal through the reward function, and use the joint reward signal to drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, so as to realize the dynamic update of the expert portrait and the optimized evolution of the potential cognitive labels.
[0023] The present invention realizes for the first time a full-process closed-loop portrait modeling from multi-source data collection, deep fusion, structure generation, optimization and regulation to behavior feedback-driven by constructing a complete expert portrait characterization and dynamic update method with adaptive evolution capabilities. This method not only breaks the limitations of traditional expert portrait static modeling and fixed structure output, but also integrates reinforcement learning strategies to achieve the continuous evolution of expert portraits and cognitive label structures under the influence of actual behaviors. Through the collaborative design of variable-dimensional mimicry generation network and multi-strategy enhanced whale optimization algorithm, the flexible adjustment of the portrait structure in the two stages of generation and optimization is achieved, and the portrait is dynamically corrected in combination with the click behavior, task completion status and user feedback collected in the actual system. This method has strong plasticity, high expressiveness and behavioral responsiveness, effectively improving the accuracy, timeliness and system adaptability of expert portraits, and has good practical application and promotion value.
[0024] In this implementation, S2 specifically includes: S21, transforming the knowledge vector, behavior vector and semantic vector into a unified feature dimension through linear mapping; S22, taking the behavior vector as the main factor, calculating the interaction weights between the behavior vector, the knowledge vector and the semantic vector and performing weighted processing, and concatenating the weighted knowledge vector, the semantic vector and the behavior vector to form a fusion feature vector; S23, constructing a potential cognitive label embedding matrix, wherein the potential cognitive label embedding matrix is obtained by embedding coding a set of labels, wherein the set of labels is derived from long-term accumulated expert behavior logs, portrait label history, and semantic topic analysis; S24. Perform a tensor product calculation on the fused feature vector and the potential cognitive label embedding matrix to generate an expert cognitive representation tensor. The expert cognitive representation tensor is a third-order tensor, including a label dimension, a feature dimension, and a time dimension, which respectively represent the feature combination under the potential cognitive label dimension, the label response degree, and the time change characteristics during the evolution of the expert behavior.
[0025] Through the expert cognitive representation tensor construction process of step S2, the problems of inconsistent multimodal feature fusion dimensions and unclear weak semantic expression of behavioral features in existing expert portrait methods are solved. The present invention maps knowledge vectors, behavioral vectors and semantic vectors to a unified feature space through linear transformation, and uses an attention mechanism dominated by behavioral vectors to extract key interactive relationships, thereby significantly improving the degree of integration between label semantics and behavioral logic. At the same time, the present invention introduces a latent cognitive label embedding matrix, so that portrait construction not only depends on the explicit label set, but also can mine potential cognitive intentions from historical behaviors and semantic themes, and construct a third-order expert cognitive tensor with a clear structure and distinct expression levels, which provides a highly plastic input basis for subsequent structure generation and dynamic optimization, and enhances the model's adaptability and expression accuracy to the evolution of cognitive structures.
[0026] In this implementation, S3 specifically includes: S31. Slice the expert cognitive representation tensor in the label dimension direction to extract the potential cognitive label vector set, denoted as ,in Indicates potential cognitive label vectors, represents the total number of potential cognitive label vectors; S32. Input each potential cognitive label vector into the label growth unit and calculate the growth probability: ; in, Indicates The growth probability of potential cognitive label vectors, represents the activation function, and represents the weight matrix, represents the Gaussian error linear unit activation function, and represents the bias vector, represents the label amplitude adjustment factor, express of norm; like If it exceeds the set threshold, the label expansion operation is triggered to add new labels to the label dimension of the expert cognitive representation tensor; S33. Input the updated expert cognitive representation tensor into the structure compression unit and calculate the retention importance score of each potential cognitive label vector: ; in, Indicates The preservation importance score of the latent cognitive label vector, and represents the weight coefficient, Indicates The average response value of the latent cognitive label vector in the attention mechanism, Indicates Potential cognitive label vectors are in front Activation variance in wheel structure; like If it is lower than the set threshold, it is removed from the expert cognitive representation tensor structure; S34. Input the compressed tensor into the dimension deformation unit, and perform affine perturbation and nonlinear transformation on each label vector: ; in, Indicates candidate label vectors, and represents the structural perturbation matrix, represents the disturbance amplitude adjustment coefficient, represents the offset vector, represents the Gaussian noise perturbation tensor; S35. Calculate the structural difference index of the candidate label vectors and screen them to obtain the expert portrait candidates: ; in, Represents the comprehensive structural difference of the candidate label vector, represents the historical reference label vector, express The adjustment coefficient of the divergence term, Represents the candidate label vector distribution Vector with historical reference labels Between Divergence; like If it falls into the preset structural deviation range, the corresponding expert cognitive representation tensor is retained as the expert portrait candidate.
[0027] By structurally decomposing the variable-dimensional mimetic generation network in step S3, and constructing a label growth unit, a structure compression unit, and a dimension deformation unit respectively, the controllable transformation and multi-morphic reconstruction of the expert cognitive representation tensor structure are achieved, effectively solving the problem of rigid label structure and lack of adaptive generation capability in expert portraits. The label growth unit can trigger the growth of new potential labels according to the current cognitive state of the expert, the structure compression unit effectively removes low-frequency weakly associated labels, and the dimension deformation unit provides the deformation capability of label representation in multiple spaces, so that the portrait has nonlinear structural migration and tensor space reconstruction capabilities. The network can generate multiple candidates with different structures, thereby improving the coverage of the portrait on the expert's diverse cognitive expressions, and also providing a structural selection space for subsequent optimization. It is the basic module for achieving the core goal of cognitive label evolution in the present invention.
[0028] In this implementation, S4 specifically includes: S41, assigning a confidence evaluation module to the expert portrait candidate, wherein the confidence evaluation module performs fusion judgment based on the coverage of the label dimension of the expert cognitive representation tensor in the expert's historical behavior trajectory, the participation frequency of completed tasks, and the matching degree with the expert's semantic description information to form a confidence input source; S42. Calculate the structural confidence score of the expert portrait candidate based on the weighted sum method: ; in, Indicates The structural confidence scores of the expert portrait candidates, represents the total number of potential cognitive label vectors, , and represents the weighting coefficient, Indicates The coverage of potential cognitive label vectors in the expert's historical behavior trajectory, Indicates The frequency of participation of latent cognitive label vectors in completed tasks, Indicates The matching degree between the potential cognitive label vector and the expert semantic description information; S43, using the structure confidence score to regulate the search direction control parameters of individuals in the multi-strategy enhanced whale optimization algorithm, wherein the expert portrait candidates whose structure confidence score is greater than the set score threshold will preferentially approach the current optimal structure during the search process; the expert portrait candidates whose structure confidence score is less than or equal to the set score threshold will delay participating in the aggregation process during the search process; S44: Check the consistency of potential cognitive label vectors in the expert profile candidates. The conflict judgment conditions include: Whether there is duplication, inclusion, or hyponymy in the semantics of the first latent cognitive label vector; Whether there is functional overlap or logical inconsistency in the task capabilities mapped by the second latent cognitive label vector; S45. When there are two or more potential cognitive label vectors in the expert portrait candidate that meet any conflict judgment condition, they are marked as label conflict bodies and will not be used as reference for updating the position of the next generation of whale individuals to avoid the spread of erroneous structures; S46. In each iteration of the multi-strategy enhanced whale optimization algorithm, only the expert portrait candidates that meet the following two conditions are retained to obtain the optimized expert portrait: The first structure confidence score is not less than the median of the structure confidence scores of all expert portrait candidates; The second is not a label conflict body.
[0029] Step S4 introduces portrait confidence guidance and label conflict discrimination mechanism in the optimization process, which significantly improves the accuracy of expert portrait candidates and the rationality of label combination in the structural optimization process. The confidence module performs weighted evaluation on the label historical behavior coverage, task participation frequency and semantic matching of each portrait, so that the optimization algorithm can give priority to retaining candidates with credible structures and avoid misleading convergence of high-deviation structures. The label conflict discrimination module can identify semantic redundancy, upper and lower conflicts and functional overlap between labels, filter semantically inconsistent label combinations from the source, and effectively improve the logical consistency and semantic clarity of the optimized portrait. The two work together to enhance the convergence robustness of the whale optimization algorithm in complex label structure space, reduce structural redundancy, offset and ambiguity in the optimization process, and ensure that the final output expert portrait is more accurate, stable and interpretable.
[0030] In this implementation manner, S5 specifically includes: S51, sorting the optimized expert portraits according to the iteration order, and taking each optimized expert portrait as a node in the graph, wherein the node data includes the potential cognitive label vector and the arrangement order corresponding to the optimized expert portrait; S52, compare potential cognitive label vectors of any two sequentially adjacent nodes, identify evolutionary behaviors based on differences, and mark a potential cognitive label vector as growth when it is added to the latter node, and mark it as contraction when it is missing; if the original potential cognitive label vector no longer appears and is replaced by a functionally equivalent or semantically identical potential cognitive label vector, it is marked as replacement; S53. Based on the results of differential recognition evolutionary behavior, directed edges are established between adjacent nodes, and the corresponding potential cognitive label vector evolution type is marked for each edge, including growth, contraction and replacement relationships, to generate an expert portrait evolution path map.
[0031] Step S5 achieves complete modeling of the structural evolution trajectory of the portrait during the dynamic generation process by constructing an expert portrait evolution path map, filling the gap in the lack of structural evolution recording and tracking mechanism in traditional expert portrait methods. The present invention uses the portrait structure state as a node and the growth, contraction and replacement of potential cognitive labels as edges to establish a directed path map, which not only enables the system to track the introduction, elimination and replacement process of labels, but also assists in portrait causal interpretation, label evolution prediction and structural traceability analysis. In addition, the introduction of the graph structure also provides the system with label structure version management capabilities, supports the visual display and time series modeling of expert portraits, enhances the controllability and explanatory power of the model, and provides a strong evolutionary evidence basis for the intelligent decision-making recommendation system.
[0032] In this implementation manner, S6 specifically includes: S61, collecting real behavior feedback data of expert portraits, wherein the real behavior feedback data includes the user's click situation on the recommended expert portraits, the completion rate of the matching task, and the user's acceptance or correction operation on the portrait label; S62, using the expert portrait evolution path map to determine the structural change mode between each expert portrait and the previous state, and for different evolution types, statistically analyzing the corresponding behavioral feedback data; S63. Construct a reward function to calculate the joint reward signal of the expert portrait: ; in, represents the joint reward signal, , , and represents the weight coefficient, Indicates the click-through rate of expert portrait recommendations, represents the task completion rate, represents the label acceptance rate, Indicates the frequency of users correcting portrait labels; S64. Use the generated joint reward signal to update the structural parameters of the variable-dimensional mimicry generation network: ; in, represents the updated structural parameters, represents the structural parameters before updating, represents the structural parameter learning rate, Represents the joint reward signal For the structural parameters before updating The gradient of S65. Simultaneously use the joint reward signal to update the search parameters of the multi-strategy enhanced whale optimization algorithm: ; in, Represents the updated search parameters, Indicates the search parameters before updating. represents the update step size; S66. Repeat steps S61 to S65. After each round of iteration, after the dynamic update of the expert portrait is completed, the next round of behavioral feedback data is collected again to continuously drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, thereby realizing the dynamic update of the expert portrait and the optimized evolution of potential cognitive labels.
[0033] Step S6 collects real behavioral feedback of expert portraits in actual applications, and constructs a joint reward function based on the expert portrait evolution path map, thereby realizing two-way dynamic adjustment of the portrait model structure and optimization parameters. The reward function comprehensively considers multiple key behavioral indicators such as user click rate, task completion rate, label acceptance and correction frequency, and constitutes a quantitative signal that accurately reflects the quality of expert portraits; through this reward signal, the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the whale optimization algorithm are simultaneously updated, so that the model has behavioral responsiveness and self-adjustment capabilities, breaking through the problems of broken feedback chains and separation of parameter updates and structural evolution in traditional model optimization. This mechanism realizes real-time closed-loop linkage between the portrait generation process and behavioral performance, so that the expert portrait can continue to evolve adaptively, and ultimately achieve the unified goal of optimal structure, semantic accuracy and behavioral consistency.
[0034] Embodiment 1: In order to verify the feasibility of the present invention in implementation, the present invention is applied to the scenario of an expert intelligent recommendation system of a certain scientific research institute to solve the long-standing problems of strong static nature of expert portraits, inability of dynamic evolution and update of portrait labels, and low portrait accuracy. The scientific research institute has a traditional expert portrait system, which has long relied on simple structural data (experts’ published papers, undertaken projects, and titles) and semantic data (expert research direction description texts) to model expert portraits, and is unable to capture expert behavior changes and cross-domain migrations in real time. Therefore, in scenarios such as recommending experts to participate in scientific research projects, review tasks, and conference reports, problems such as insufficient recommendation accuracy and low expert response rates often occur, and user satisfaction is not high overall. In addition, due to the inability to dynamically update expert portraits in real time, the system’s manual cycle for updating expert portraits is at least one month each time, which seriously lags behind changes in the actual behavior status of experts, reducing the practical value of the recommendation service.
[0035] In view of the above problems, the present invention has been specifically applied and practiced in the expert intelligent recommendation system of the scientific research institute. Specifically, in this embodiment, the structural data (papers, projects, titles, research fields), behavioral data (click expert recommended tasks, response records of accepting tasks, cross-domain task migration) and semantic data (research direction description, abstract, speech) of a total of 300 experts in the institute are first collected to generate three types of feature vectors: knowledge, behavior and semantics. Subsequently, the fusion method of the present invention is used to construct an expert cognitive representation tensor with potential cognitive tags as the core, and the complex heterogeneous features are deeply integrated to form a unified portrait structure.
[0036] Furthermore, the variable-dimensional mimetic generation network of the present invention is applied to the expert system, and the label structure is updated once a week based on the newly collected expert data. During the application process, the label growth unit in the network automatically generates new potential cognitive labels based on the expert's activity in new fields or hot issues; the structure compression unit deletes the corresponding label when the expert's activity in a certain field is significantly reduced; the dimension deformation unit flexibly adjusts the portrait structure dimension according to the actual situation, so that each expert's portrait always accurately reflects his current status.
[0037] At the same time, the present invention introduces a multi-strategy enhanced whale optimization algorithm, which calculates the credibility and label conflict of the expert portrait based on the historical behavior coverage, task participation frequency and semantic matching degree of the expert portrait, dynamically optimizes and screens the expert portrait every week, eliminates redundant labels and unreasonable combinations, and ensures that the final expert portrait structure is logically rigorous and semantically clear. Then, through the expert portrait evolution path map, the historical changes of each expert's portrait are visualized and recorded to ensure that the change process of the expert portrait can be traced, providing a more powerful basis for recommendation decisions.
[0038] Statistical data from one year of practical application show that the present invention significantly improves the accuracy of expert portraits and user satisfaction. The specific data comparison is shown in Table 1 below.
[0039] Table 1 Annual comparison of the recommendation effects of the expert portrait system of the present invention and the traditional method
[0040] According to the data in Table 1 above, it can be seen that the expert portrait characterization and dynamic update method proposed in the present invention has achieved a significant improvement in the response rate of recommendation tasks, and the expert response rate has been increased from 73.2% of the traditional system to 92.5%; the matching score of recommended experts and tasks has also been significantly improved, from the original 3.8 points to 4.7 points, and the overall user satisfaction score has been increased from 80.2 points to 95.3 points, reflecting the overall improvement of the accuracy and matching degree of expert portraits. In addition, the success rate of the present invention in new label prediction is as high as 88.6%, the average monthly label correction frequency of expert portraits has also been reduced from the traditional 1.2 times to 0.15 times, and the average update cycle of expert portraits has been shortened from the previous 35 days to 7 days. This fully demonstrates that the dynamic evolution of labels and the expert behavior feedback mechanism of the present invention effectively solve the static and lagging problems of the traditional expert portrait system, truly and accurately reflects the behavioral changes and cognitive state evolution of experts, and can effectively improve the quality, real-time and interpretability of expert recommendation services, and has strong practical application value and obvious technical advantages.
[0041] This embodiment significantly solves the problems of static structure, update lag, label redundancy and low recommendation accuracy in traditional expert portrait systems by deploying the expert portrait characterization and dynamic update method proposed by the present invention in actual scientific research institutions. During the implementation process, the expert cognitive representation tensor is constructed by integrating structural, behavioral and semantic data, the portrait structure is adaptively evolved in combination with the variable-dimensional mimetic generation network, and the confidence guidance and label conflict discrimination mechanism are introduced to optimize the candidate portraits, which greatly improves the accuracy of the portrait expression and the structural rationality. Experimental data show that the present invention can achieve weekly updates of expert portraits, the response rate of recommendation tasks is increased to 92.5%, the success rate of new label prediction is 88.6%, and the frequency of label correction is significantly reduced, reflecting a strong feedback response capability and evolutionary optimization effect. On the whole, this embodiment verifies that the present invention has the advantages of real-time modeling, precise expression, behavior perception and continuous optimization, and can be widely used in practical application scenarios such as expert recommendation, scientific research management, and talent mining.
[0042] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for characterizing and dynamically updating expert portraits based on reinforcement learning, characterized in that: The steps include: S1, collect the structural data, behavioral data and semantic data of experts, and pre-process them to generate knowledge vectors, behavioral vectors and semantic vectors; S2, fuse the knowledge vector, behavior vector and semantic vector to construct the expert cognitive representation tensor with latent cognitive labels as the core; S3, inputting the expert cognitive representation tensor into a variable-dimensional mimicry generation network, wherein the variable-dimensional mimicry generation network includes a label growth unit, a structure compression unit, and a dimension deformation unit to generate an expert portrait candidate; S4. Use multi-strategy enhanced whale optimization algorithm to optimize the expert portrait candidates, and generate optimized expert portraits with portrait confidence, potential cognitive label consistency and expert historical behavior compatibility as constraints; S5. Generate an expert portrait evolution path graph based on the optimized expert portrait, where the nodes represent the optimized expert portrait status and the edges represent the growth, contraction and replacement relationships; S6. Collect the real behavior feedback data of the expert portrait, and combine it with the expert portrait evolution path map, generate a joint reward signal through the reward function, and use the joint reward signal to drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, so as to realize the dynamic update of the expert portrait and the optimized evolution of the potential cognitive labels.
2. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S2 specifically includes: S21, transforming the knowledge vector, behavior vector and semantic vector into a unified feature dimension through linear mapping; S22, taking the behavior vector as the main factor, calculating the interaction weights between the behavior vector, the knowledge vector and the semantic vector and performing weighted processing, and concatenating the weighted knowledge vector, the semantic vector and the behavior vector to form a fusion feature vector; S23, constructing a potential cognitive label embedding matrix, wherein the potential cognitive label embedding matrix is obtained by embedding coding a set of labels, wherein the set of labels is derived from long-term accumulated expert behavior logs, portrait label history, and semantic topic analysis; S24. Perform a tensor product calculation on the fused feature vector and the potential cognitive label embedding matrix to generate an expert cognitive representation tensor. The expert cognitive representation tensor is a third-order tensor, including a label dimension, a feature dimension, and a time dimension, which respectively represent the feature combination under the potential cognitive label dimension, the label response degree, and the time change characteristics during the evolution of the expert behavior.
3. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S3 specifically includes: S31. Slice the expert cognitive representation tensor in the label dimension direction to extract the potential cognitive label vector set, denoted as ,in Indicates potential cognitive label vectors, represents the total number of potential cognitive label vectors; S32. Input each potential cognitive label vector into the label growth unit and calculate the growth probability: ; in, Indicates The growth probability of potential cognitive label vectors, represents the activation function, and represents the weight matrix, represents the Gaussian error linear unit activation function, and represents the bias vector, represents the label amplitude adjustment factor, express of norm; like If it exceeds the set threshold, the label expansion operation is triggered to add new labels to the label dimension of the expert cognitive representation tensor; S33. Input the updated expert cognitive representation tensor into the structure compression unit, and calculate the retention importance score of each potential cognitive label vector: ; in, Indicates The preservation importance score of the latent cognitive label vector, and represents the weight coefficient, Indicates The average response value of the latent cognitive label vector in the attention mechanism, Indicates Potential cognitive label vectors are in front Activation variance in wheel structure; like If it is lower than the set threshold, it is removed from the expert cognitive representation tensor structure; S34. Input the compressed tensor into the dimension deformation unit, and perform affine perturbation and nonlinear transformation on each label vector: ; in, Indicates candidate label vectors, and represents the structural perturbation matrix, represents the disturbance amplitude adjustment coefficient, represents the offset vector, represents the Gaussian noise perturbation tensor; S35. Calculate the structural difference index of the candidate label vectors and screen them to obtain the expert portrait candidates: ; in, Represents the comprehensive structural difference of the candidate label vector, represents the historical reference label vector, express The adjustment coefficient of the divergence term, Represents the candidate label vector distribution Vector with historical reference labels Between Divergence; like If it falls into the preset structural deviation range, the corresponding expert cognitive representation tensor is retained as the expert portrait candidate.
4. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S4 specifically includes: S41, assigning a confidence evaluation module to the expert portrait candidate, wherein the confidence evaluation module performs fusion judgment based on the coverage of the label dimension of the expert cognitive representation tensor in the expert's historical behavior trajectory, the participation frequency of completed tasks, and the matching degree with the expert's semantic description information to form a confidence input source; S42. Calculate the structural confidence score of the expert portrait candidate based on the weighted sum method: ; in, Indicates The structural confidence scores of the expert portrait candidates, represents the total number of potential cognitive label vectors, , and represents the weighting coefficient, Indicates The coverage of potential cognitive label vectors in the expert's historical behavior trajectory, Indicates The frequency of participation of latent cognitive label vectors in completed tasks, Indicates The matching degree between the potential cognitive label vector and the expert semantic description information; S43, using the structure confidence score to regulate the search direction control parameters of individuals in the multi-strategy enhanced whale optimization algorithm, wherein the expert portrait candidates whose structure confidence score is greater than the set score threshold will preferentially approach the current optimal structure during the search process; the expert portrait candidates whose structure confidence score is less than or equal to the set score threshold will delay participating in the aggregation process during the search process; S44: Check the consistency of potential cognitive label vectors in the expert profile candidates. The conflict judgment conditions include: Whether there is duplication, inclusion, or hyponymy in the semantics of the first latent cognitive label vector; Whether there is functional overlap or logical inconsistency in the task capabilities mapped by the second latent cognitive label vector; S45. When there are two or more potential cognitive label vectors in the expert portrait candidate that meet any conflict judgment condition, they are marked as label conflict bodies and will not be used as reference for updating the position of the next generation of whale individuals to avoid the spread of erroneous structures; S46. In each iteration of the multi-strategy enhanced whale optimization algorithm, only the expert portrait candidates that meet the following two conditions are retained to obtain the optimized expert portrait: The first structure confidence score is not less than the median of the structure confidence scores of all expert portrait candidates; The second is not a label conflict body.
5. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S5 specifically includes: S51, sorting the optimized expert portraits according to the iteration order, and taking each optimized expert portrait as a node in the graph, wherein the node data includes the potential cognitive label vector and the arrangement order corresponding to the optimized expert portrait; S52, compare potential cognitive label vectors of any two sequentially adjacent nodes, identify evolutionary behaviors based on differences, and mark a potential cognitive label vector as growth when it is added to the latter node, and mark it as contraction when it is missing; if the original potential cognitive label vector no longer appears and is replaced by a functionally equivalent or semantically identical potential cognitive label vector, it is marked as replacement; S53. Based on the results of differential recognition evolutionary behavior, directed edges are established between adjacent nodes, and the corresponding potential cognitive label vector evolution type is marked for each edge, including growth, contraction and replacement relationships, to generate an expert portrait evolution path map.
6. The method for characterizing and dynamically updating expert portraits based on reinforcement learning according to claim 1, characterized in that: The S6 specifically includes: S61, collecting real behavior feedback data of expert portraits, wherein the real behavior feedback data includes the user's click situation on the recommended expert portraits, the completion rate of the matching task, and the user's acceptance or correction operation on the portrait label; S62, using the expert portrait evolution path map to determine the structural change mode between each expert portrait and the previous state, and for different evolution types, statistically analyzing the corresponding behavioral feedback data; S63. Construct a reward function to calculate the joint reward signal of the expert portrait: ; in, represents the joint reward signal, , , and represents the weight coefficient, Indicates the click-through rate of expert portrait recommendations, represents the task completion rate, represents the label acceptance rate, Indicates the frequency of users correcting portrait labels; S64. Use the generated joint reward signal to update the structural parameters of the variable-dimensional mimicry generation network: ; in, represents the updated structural parameters, represents the structural parameters before updating, represents the structural parameter learning rate, Represents the joint reward signal For the structural parameters before updating The gradient of S65. Simultaneously use the joint reward signal to update the search parameters of the multi-strategy enhanced whale optimization algorithm: ; in, Represents the updated search parameters, Indicates the search parameters before updating. represents the update step size; S66. Repeat steps S61 to S65. After each round of iteration, after the dynamic update of the expert portrait is completed, the next round of behavioral feedback data is collected again to continuously drive the bidirectional update of the structural parameters of the variable-dimensional mimicry generation network and the search parameters of the multi-strategy enhanced whale optimization algorithm, thereby realizing the dynamic update of the expert portrait and the optimized evolution of potential cognitive labels.
Citation Information
Patent Citations
Rolling bearing fault diagnosis method based on IWOA-ELM
CN115096590A
Security propaganda and education recommendation method and system based on demand portrait and content label
CN118797173A
Traveling salesman problem solving method combining deep reinforcement learning and heuristic algorithm
CN119106778A
Model training method, product recommendation method, equipment, medium and program product
CN119558886A
A system of oriented medical insurance prediction with machine learning integrated IoT
IN201941041623A
Cited By
Dynamic updating method for expert portrait description
CN120106198A
A dynamic update method for expert profiling
CN120106198B
AI personalized tutoring method based on cognitive portrait
CN120407950A
AI personalized tutoring method based on cognitive portrait
CN120407950B
E-commerce user portrait construction method based on reinforcement and increase learning
CN120687908A