Talent pool modeling method and system based on large language model
Patent Information
- Application Number
- CN202611240764.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-17
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明的主要目的在于提供一种基于大语言模型的人才库建模方法及系统,旨在解决现有人才库建模方法因依赖静态标签和粗糙的关键词提取,缺乏对多源异构数据中细粒度行为原子的语义解构及动态逻辑依赖关系的精准捕捉,导致难以全面量化人才的深层胜任力与演化特征的技术问题
[0062]本发明提供了一种基于大语言模型的人才库建模方法,该方法通过融合历史工作文本、行为轨迹与实时交互反馈等多源异构数据,利用大语言模型对人才工作流进行细粒度的任务原子化解析,实现了从非结构化文本中提取包含属性特征与执行逻辑的任务节点,显著提升了人才能力单元的可解释性与结构化表达能力;通过构建行为原子语义序列并引入注意力机制建模不同行为间的逻辑依赖关系,实现了在语义域的动态关联与认知效能的最大化覆盖,有效捕捉了人才行为模式的时序演化特征;在此基础上,结合任务原子与行为序列进行动态匹配度计算,提出累积胜任力指数与协同技能映射分布双维度评估体系,能够精准预测个体在具体任务节点上的胜任水平,并揭示多技能协同作用的内在机制;最终通过人才特征数据对人才库结构进行偏差识别与优化迭代,实现了人才画像的持续精化与组织人才资源配置的智能化升级。整体方案突破了传统人才建模中静态、粗粒度、语义缺失的局限,显著提高了人才-岗位匹配的准确性、动态适应性与系统智能化水平,为高潜人才识别、岗位适配推荐和组织能力建设提供了科学、可计算的数字模型支撑。
Smart Images

Figure CN122820153A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of talent pool modeling technology, and in particular to a talent pool modeling method and system based on a large language model. Background Technology
[0002] With the rapid development of artificial intelligence and big data technologies, enterprises are increasingly demanding more refined and intelligent talent management. Traditional talent pool modeling methods mainly rely on resume analysis, static tagging systems, and manual experience assessment, which are insufficient to comprehensively and dynamically depict the actual abilities and behavioral characteristics of talents. Especially in scenarios such as complex job matching, high-potential talent identification, and organizational capability planning, existing methods generally suffer from problems such as coarse granularity, insufficient semantic understanding, and lack of behavioral evolution modeling, resulting in low accuracy in matching talents with positions and insufficient exploration of talent potential. In recent years, large language models have demonstrated powerful capabilities in natural language understanding, semantic representation, and contextual reasoning, providing a new technical path for the deep analysis of talent data. However, directly applying large language models to talent modeling still faces many challenges: on the one hand, raw text data (such as work experience and project descriptions) is usually highly abstract and diverse in expression, making it difficult to directly map to quantifiable task units; on the other hand, a person's abilities are not only reflected in the content of the tasks they complete, but also in their behavioral patterns, response efficiency, contextual adaptability, and other dynamic characteristics, which are often ignored in traditional modeling processes.
[0003] Existing research attempts to construct talent profiles through keyword extraction, skill tag classification, or simple sequence modeling, but lacks a systematic modeling mechanism for fine-grained semantic relationships between tasks, behaviors, and abilities. Especially when dealing with cross-domain, multimodal talent data (such as text descriptions, operation logs, and interaction feedback), it is difficult to achieve accurate segmentation and semantic alignment of behavioral units, and it is also unable to effectively capture the logical dependencies between behaviors and their synergistic impact on overall competence.
[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this invention is to provide a talent pool modeling method and system based on a large language model. This aims to solve the technical problem that existing talent pool modeling methods rely on static labels and coarse keyword extraction, lacking the ability to accurately capture the semantic deconstruction of fine-grained behavioral atoms and dynamic logical dependencies in multi-source heterogeneous data, thus making it difficult to comprehensively quantify the deep competence and evolutionary characteristics of talents.
[0006] To achieve the above objectives, this invention provides a talent pool modeling method based on a large language model, the method comprising:
[0007] Acquire candidate candidates' historical work text data, behavioral trajectory data, and real-time interactive feedback data;
[0008] The talent workflow is discretized based on historical work text data to obtain task atomic data containing task node attribute information and execution logic information;
[0009] Based on real-time interactive feedback data, dynamic modeling of talent behavior patterns is performed to obtain behavioral atom semantic sequence data. Specifically, the changes in behavioral atoms are achieved by setting attention weight combinations with different logical dependencies between behavioral atoms to perform association operations of multiple behaviors in the semantic domain and maximum effectiveness coverage in the cognitive domain.
[0010] Based on task atomic data and behavioral atomic semantic sequence data, dynamic matching degree calculation is performed on talent competency and skill distribution to obtain talent characteristic data including cumulative competency index and collaborative skill mapping distribution. Specifically, the dynamic matching degree calculation is to predict the competency of each task node through job requirement constraints and skill distribution characteristics.
[0011] The talent pool structure is evaluated and optimized based on talent characteristic data to obtain a refined digital model of the talent pool.
[0012] Optionally, the discretization of the talent workflow based on historical work text data to obtain task atomic data containing task node attribute information and execution logic information includes:
[0013] The talent workflow is divided into tasks based on historical work text data to obtain task data, where the task density is set to no less than 5 behavior atoms per thousand words of text.
[0014] Based on the task data, attribute values are assigned and logical relationships are calculated for each task atom to obtain atomic topology attribute data containing atomic attribute information and logical connection information.
[0015] Based on the atomic topology attribute data, unique semantic identifiers are assigned to the task atoms to obtain task atom data with index identifiers.
[0016] Optionally, the step of dynamically modeling talent behavior patterns based on real-time interactive feedback data to obtain behavioral atomic semantic sequence data includes:
[0017] Extract the operation response latency, contextual relevance, semantic position, and logical evolution direction of each behavioral atom from real-time interactive feedback data, and establish a structured parameter table for each behavioral atom;
[0018] The logical weight value of each behavior is determined by the total number of atoms in the parameter table, and a logical configuration list is generated by assigning different logical starting points to each atom by equally dividing the complete semantic cycle.
[0019] Based on the parameter table and the logical configuration list, generate a set of behavioral atom sequences containing personalized logical dependency descriptions for each behavioral atom;
[0020] Based on the set of behavioral atom sequences, the entire interaction cycle is sampled according to a preset semantic interval, and the logical connection state of all behavioral atoms at each semantic point is calculated, thereby forming a behavioral state sequence table arranged in semantic order.
[0021] By using a behavior state sequence table, a fast query correspondence between timestamps and behavioral semantic information is established, thereby obtaining behavioral atomic semantic sequence data.
[0022] Optionally, the step of dynamically calculating the matching degree of talent competency and skill distribution based on task atomic data and behavioral atomic semantic sequence data to obtain talent characteristic data including cumulative competency index and collaborative skill mapping distribution includes:
[0023] The connectivity calculation of behavior and task atoms is performed based on the semantic sequence data of behavior atoms and the data of task atoms. By calculating the semantic association distance between each behavior atom at each semantic point and each task atom in the talent pool, the behavior-task connectivity matrix data is obtained.
[0024] Based on the behavioral task connectivity matrix data and the preset job competency threshold, the effective skill range is determined. By comparing connectivity, task atoms within the effective skill range of each behavioral atom are selected to obtain behavioral skill coverage relationship data.
[0025] Based on the behavioral skill coverage relationship data, the instantaneous competence value of each task atom is calculated using a skill allocation distribution model based on an exponential decay function, thus obtaining the instantaneous competence data of the atom.
[0026] Based on the instantaneous competence data of atoms, a semantic sequence-based cumulative calculation process is performed to obtain the cumulative competence index data of atoms. Specifically, the cumulative calculation process is performed by semantic integration on the competence of each task atom during the entire interaction cycle.
[0027] The atomic cumulative competence index data is used to superimpose and calculate the effect of multi-behavior atomic synergy to obtain atomic synergy skill mapping distribution data;
[0028] Based on the atomic cumulative competence index data and the atomic collaborative skill mapping distribution data, the cumulative competence information and the collaborative skill mapping distribution information are matched according to the task atomic index identifier to obtain talent characteristic data.
[0029] Optionally, the effective skill range determination process based on the behavioral task connectivity matrix data and a preset job competency threshold, filtering out task atoms within the effective skill range of each behavioral atom through connectivity comparison, and obtaining behavioral skill coverage relationship data includes:
[0030] The correlation determination results are obtained by comparing the behavioral task connectivity matrix data and the preset job competency thresholds one by one.
[0031] Based on the correlation determination results, behavioral task pairs that meet the skill range conditions are screened and extracted. By retaining behavioral task combinations with correlation values less than or equal to the competence threshold, a list of effective behavioral task pairs is obtained.
[0032] Based on the list of valid behavioral task pairs, the skill task set of each behavioral atom is grouped according to the atom identifier to obtain behavioral group atom set data;
[0033] Based on the behavior grouping atomic set data, the skill task set of each behavior atom is traversed and the number of covered behaviors of each atom is counted to obtain the list of covered behaviors of the atom.
[0034] By constructing a bidirectional index mapping table from behavior to atom and from atom to behavior using behavior grouping atom set data and a list of behaviors covered by atoms, behavior skill coverage relationship data can be obtained.
[0035] Optionally, the instantaneous competence value calculation for each task atom based on the behavioral skill coverage relationship data using a skill allocation distribution model with an exponential decay function, to obtain the atom instantaneous competence data, includes:
[0036] Extract the semantic association distance information between each behavioral atom and the task atoms within its skill range from the behavioral skill coverage relationship data to obtain an effective semantic distance value data set;
[0037] Normalized distance ratio data is obtained by using a set of effective semantic distance values and a preset exponential decay parameter to perform normalized distance calculation.
[0038] Based on the normalized semantic distance ratio data, the normalized semantic distance ratio is multiplied by a negative exponential decay coefficient and the exponential function is taken to obtain the exponential decay factor data.
[0039] The skill value is calculated based on the exponential decay factor data and the preset maximum skill capacity parameter of the atom to obtain the basic skill allocation value of the task atom.
[0040] Based on the assigned values of basic skills of the task atoms, the data is processed by correcting and solving the results based on the direction of behavior atomic operations and the phase angle of task semantics, thereby obtaining the instantaneous competence data of atoms.
[0041] Optionally, the step of performing correction and result solving based on the behavioral atomic operation direction and task semantic phase angle according to the task atomic basic skill allocation value to obtain atomic instantaneous competence data includes:
[0042] Based on the basic skill allocation values of task atoms, the phase difference between behavioral atoms and task semantics is calculated and the angle correction is performed. The cosine value of the phase angle between the operation direction of the behavioral atom and the task semantics is calculated and used as the correction coefficient to obtain the skill allocation value after angle correction.
[0043] Based on the skill allocation value and the behavioral atom response speed information of the current semantic point, the speed influence factor is calculated and processed to obtain behavioral atom speed modulation skill data;
[0044] Based on the behavioral atom velocity modulation skill data, the effects of multiple behavioral atoms on the same task atom at the current semantic point are superimposed and summed to obtain the instantaneous competence data of the atom.
[0045] Optionally, the step of using atomic cumulative competence index data to superimpose and calculate the multi-behavior atomic synergy effect to obtain atomic synergy skill mapping distribution data includes:
[0046] Based on the atomic cumulative competence index data, the neighborhood relationship of each task atom in the talent pool is established, the direct neighboring atoms of each atom are identified and the topological connection relationship between atoms is established, and the atomic neighborhood topology data is obtained.
[0047] The local skill intensity gradient of each atom is calculated based on the atomic neighborhood topology data and the atomic cumulative competence index data. The atomic skill gradient data is obtained by calculating the competence difference between the current atom and its neighboring atoms and dividing it by the atomic semantic distance.
[0048] Based on the atomic skill gradient data, the synergistic enhancement effect between multi-behavioral atoms is quantitatively calculated based on gradient smoothness to obtain atomic synergistic enhancement coefficient data.
[0049] Synergistic effect modulation calculations are performed using atomic synergistic enhancement coefficient data and atomic cumulative competence index data. The original skill allocation is optimized and modulated by multiplying the cumulative competence with the synergistic enhancement coefficient to obtain atomic synergistic modulation skill data.
[0050] By performing spatial distribution statistical processing on the overall synergistic effect of the talent pool using atomic synergistic modulation skill data, and rearranging and normalizing the synergistic modulation skills of each atom according to their semantic positions, atomic synergistic skill mapping distribution data is obtained.
[0051] Optionally, the step of evaluating and optimizing the talent pool structure based on talent characteristic data to obtain refined talent pool digital model data includes:
[0052] Based on talent characteristic data, statistical analysis is performed on the skill allocation distribution of the talent pool to calculate the mean, variance, and standard deviation of the global skill allocation, thereby obtaining statistical characteristic data of skill allocation.
[0053] The balance deviation of each region of the talent pool is quantitatively assessed by using statistical characteristic data of skill allocation. The relative deviation between the allocation of atomic skills for each task and the global mean is calculated and a histogram of deviation distribution is established to obtain balance deviation assessment data.
[0054] Based on the balance deviation assessment data, areas with substandard talent pool structure are identified and marked. By setting a deviation threshold range and filtering out task atom sets that exceed the threshold, abnormal area marking data of talent profile is obtained.
[0055] The behavior atomic parameters were adjusted and the talent pool feature mapping effect was optimized based on the abnormal area marking data of the talent profile. The convergence of the optimization effect was verified, thereby obtaining refined digital model data of the talent pool.
[0056] Furthermore, to achieve the above objectives, the present invention also provides a talent pool modeling system based on a large language model, the system comprising:
[0057] The multi-source acquisition module is used to acquire candidates' historical work text data, behavioral trajectory data, and real-time interactive feedback data.
[0058] The task discretization module is used to discretize the talent workflow based on historical work text data to obtain task atomic data containing task node attribute information and execution logic information.
[0059] The behavior modeling module is used to dynamically model talent behavior patterns based on real-time interactive feedback data, and obtain behavioral atom semantic sequence data. Specifically, the changes in behavioral atoms are achieved by setting attention weight combinations with different logical dependencies between behavioral atoms to perform semantic domain association operations and maximum efficiency coverage in the cognitive domain.
[0060] The dynamic matching module is used to perform dynamic matching degree calculation on talent competency and skill distribution based on task atomic data and behavior atomic semantic sequence data, and obtain talent characteristic data including cumulative competency index and collaborative skill mapping distribution. Specifically, the dynamic matching degree calculation process is to predict the competency of each task node through job requirement constraints and skill distribution characteristics.
[0061] The talent pool structure optimization module is used to evaluate and optimize the talent pool structure based on talent characteristic data, resulting in refined digital model data of the talent pool.
[0062] This invention provides a talent pool modeling method based on a large language model. This method integrates multi-source heterogeneous data such as historical work texts, behavioral trajectories, and real-time interactive feedback. It utilizes a large language model to perform fine-grained task atomization analysis of talent workflows, extracting task nodes containing attribute features and execution logic from unstructured text. This significantly improves the interpretability and structured expression of talent competency units. By constructing semantic sequences of behavioral atoms and introducing an attention mechanism to model the logical dependencies between different behaviors, it achieves dynamic association in the semantic domain and maximizes the coverage of cognitive efficacy, effectively capturing the temporal evolution characteristics of talent behavior patterns. Based on this, it combines task atoms and behavioral sequences for dynamic matching degree calculation, proposing a dual-dimensional evaluation system of cumulative competence index and collaborative skill mapping distribution. This system can accurately predict an individual's competence level at specific task nodes and reveal the intrinsic mechanism of multi-skill synergy. Finally, it uses talent feature data to identify and iteratively optimize the talent pool structure, achieving continuous refinement of talent profiles and intelligent upgrading of organizational talent resource allocation. The overall solution breaks through the limitations of traditional talent modeling, which is static, coarse-grained, and lacks semantics. It significantly improves the accuracy, dynamic adaptability, and system intelligence of talent-job matching, and provides scientific and computable digital model support for high-potential talent identification, job matching recommendation, and organizational capability building. Attached Figure Description
[0063] Figure 1 This is a flowchart illustrating an embodiment of the talent pool modeling method based on a large language model according to the present invention;
[0064] Figure 2 This is a structural block diagram of an embodiment of the talent pool modeling system based on a large language model according to the present invention.
[0065] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0066] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0067] Reference Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the talent pool modeling method based on a large language model according to the present invention.
[0068] In one embodiment, the talent pool modeling method based on a large language model includes:
[0069] Step S100: Obtain the candidate's historical work text data, behavior trajectory data, and real-time interactive feedback data.
[0070] The candidate's historical work text data can be unstructured natural language text describing the candidate's past professional experience, including projects, responsibilities, and achievements. This text can serve as the raw semantic input for talent competency modeling, extracting task execution content and logical structure. In this embodiment, the candidate's historical work text data can be collected through resume parsing, internal system logs, project documents, or third-party platform scraping. Behavioral trajectory data can be digital behavioral logs recording the candidate's operation sequences, response times, and path selections during work or interaction. This data can reflect the candidate's behavioral patterns and contextual adaptability in actual tasks. For example, behavioral trajectory data can be automatically generated by recording user operation behaviors through an internal enterprise system. Real-time interactive feedback data can be the candidate's immediate responses, Q&A content, evaluation results, or system feedback signals generated in the current interactive scenario. This data can be used to dynamically update the behavioral model and capture the candidate's cognitive performance and response characteristics in new situations. In an exemplary embodiment, real-time interactive feedback data can be collected in real-time through online assessments, interview dialogues, task simulation platforms, etc., to collect interactive content and system scores.
[0071] Step S200: Discretize the talent workflow based on historical work text data to obtain task atomic data containing task node attribute information and execution logic information.
[0072] The talent workflow can be the complete process by which a candidate completes a task in a specific position or project, including task sequence, dependencies, and execution context. It can serve as the basic structural unit for task atomization and parsing. In one specific embodiment, the talent workflow can be a task flow graph reconstructed from historical work text through semantic understanding. Task node attribute information can be static features describing the skill types, domain knowledge, tool usage, and output formats involved in the task atoms. This can be used to characterize the task's capability requirements and support subsequent competency mapping. Furthermore, task node attribute information can include, but is not limited to, professional skill attributes, soft skill attributes, and tool proficiency attributes. Execution logic information can be dynamic execution rules describing the temporal relationships, conditional branches, parallel structures, or causal dependencies between task atoms. This can be used to reconstruct the contextual logic of the task in the real workflow, enhancing the interpretability of task atoms. For example, execution logic information can employ linear execution logic, conditional judgment logic, and iterative loop logic. Task atom data can be the smallest interpretable task unit discretized from the talent workflow, containing task node attribute information and execution logic information. This can be used to achieve fine-grained expression of talent capability units, supporting task-level matching and capability traceability. In this embodiment, task atomic data can utilize a large language model to perform semantic segmentation and structured extraction on historical work texts, identifying independent and semantically complete task fragments. Furthermore, task atomic data, together with behavioral atomic semantic sequence data, can form the input basis for dynamic matching degree calculation; its attribute information provides dimensional basis for collaborative skill mapping distribution.
[0073] Step S300: Dynamically model talent behavior patterns based on real-time interactive feedback data to obtain behavioral atom semantic sequence data. Specifically, behavioral atom changes are achieved by setting attention weight combinations with different logical dependencies between behavioral atoms to perform semantic domain association operations and maximize the effectiveness coverage of the cognitive domain.
[0074] Talent behavior patterns can be stable behavioral tendencies exhibited by candidates in multiple scenarios, such as decision-making style, response rhythm, and collaboration preferences, which can be used as target representations for behavioral atom semantic sequence modeling. In an exemplary embodiment, talent behavior patterns can be formed through pattern mining and clustering of behavioral trajectory data and real-time interactive feedback data. Behavioral atoms can be the smallest semantic behavioral units constituting a behavioral pattern, representing a behavioral event with a clear intent and context, and can be used as basic elements for constructing behavioral sequences, supporting temporal modeling and logical dependency analysis. Furthermore, behavioral atoms can include, but are not limited to, information query behavior, collaboration initiation behavior, and problem-solving behavior. Behavioral atom semantic sequence data can be a structured sequence formed by arranging multiple behavioral atoms in chronological or logical order, embedding semantic vector representations, which can be used to model the temporal evolution and contextual dependencies of behavior, supporting dynamic competency assessment. In a specific embodiment, behavioral atom semantic sequence data can be based on real-time interactive feedback data, using a large language model to map the original behavioral logs into semantic behavioral atoms, and organizing them into sequences according to context. Furthermore, behavioral atom semantic sequence data can be jointly input with task atom data into a dynamic matching degree calculation module; its sequence structure provides a modeling basis for attention mechanisms.
[0075] The attention weight combination for logical dependencies can be a set of correlation strength parameters between different behavioral atoms calculated through an attention mechanism. This can be used to quantify the logical connections between behavioral atoms, achieving dynamic association within the semantic domain and optimizing cognitive effectiveness. In this embodiment, the attention weight combination for logical dependencies can apply self-attention or cross-attention mechanisms to the semantic sequences of behavioral atoms to learn the contextual dependency weights between behaviors. Semantic domain association operations can be operations such as vector fusion, relational reasoning, or context alignment of multiple behavioral atoms in a unified semantic space. This can be used to achieve semantic integration across behaviors, improving the overall representational ability of behavioral patterns. Maximizing the effectiveness coverage of the cognitive domain can be achieved by optimizing the distribution of attention weights, ensuring that the representation of behavioral sequences reflects the cognitive processing capabilities of individuals to the greatest extent possible. This can be used to ensure that behavioral modeling covers not only explicit operations but also implicit decision-making and adaptive thinking.
[0076] Step S400: Based on the task atomic data and behavior atomic semantic sequence data, perform dynamic matching degree calculation on talent competency and skill distribution to obtain talent characteristic data including cumulative competency index and collaborative skill mapping distribution. Specifically, the dynamic matching degree calculation process is to predict the competency of each task node through job requirement constraints and skill distribution characteristics.
[0077] Among these, talent competence can be the comprehensive ability level of a candidate to successfully execute a task at a specific task node, serving as one of the core output targets for dynamic matching degree calculation. Skill distribution can be the distribution of a candidate's ability intensity in a multi-dimensional skill space, reflecting the breadth and depth of skills. It can be used to reveal the mechanism of multi-skill synergy and support the construction of collaborative skill mapping distribution. Job requirement constraints can be a set of limiting conditions regarding the abilities, experience, and behavioral characteristics required for task execution by the target job. They can be used as external constraints for dynamic matching degree calculation, guiding the direction of competence prediction. Skill distribution characteristics can be the statistical or topological features of skill distribution in terms of structure, such as sparsity, clustering, and complementarity. They can be used to adjust the calculation strategy of collaborative skill mapping and reflect the inherent laws of skill combinations. Competency prediction for task nodes can be the prediction of a candidate's success probability or ability score for completing a specific task atom, providing basic unit values for the cumulative competence index.
[0078] The cumulative competency index is a comprehensive capability indicator formed by weighting and accumulating the competency prediction results of each node in the task's atomic sequence according to execution logic. It can be used to quantify the overall competency level of talent in a complete workflow, supporting high-potential identification and job matching. Furthermore, the cumulative competency index can employ path-weighted accumulation, key-node-driven accumulation, and fault-tolerant correction accumulation. The collaborative skill mapping distribution can be a probability distribution describing the joint activation intensity and interaction patterns of multiple skills when completing a specific task. It can be used to reveal skill collaboration mechanisms and avoid the one-sidedness of single skill labels. Talent characteristic data can be structured talent capability representation data containing the cumulative competency index and the collaborative skill mapping distribution. It can be used as the core data carrier for talent profiling, driving talent pool structure optimization. In a specific embodiment, talent characteristic data can be generated by dynamic matching degree calculation, integrating task and behavioral dual-dimensional information. Furthermore, talent characteristic data can be directly input into the evaluation and optimization processing module to identify talent pool structure deviations.
[0079] Step S500: Evaluate and optimize the talent pool structure based on talent characteristic data to obtain refined talent pool digital model data.
[0080] The talent pool structure can be the organization of talent data within the talent pool, including classification systems, indexing mechanisms, and relational networks. It determines the efficiency and accuracy of talent retrieval, recommendation, and allocation. Evaluation and optimization processing involves detecting biases and adjusting the structure of the talent pool based on talent characteristic data. This allows for continuous refinement of talent profiles and adaptive evolution of the pool structure. In an exemplary embodiment, evaluation and optimization processing compares the classification or clustering results of talent characteristic data with those of the existing pool structure, identifies label biases, coverage blind spots, or redundant clusters, and adjusts indexing or classification rules. Furthermore, evaluation and optimization processing can be implemented by using clustering stability indices to detect structural biases and by evaluating feature distribution shifts based on adversarial verification methods to trigger retraining. The refined talent pool digital model data can be a high-precision, dynamically updated structured data model of the talent pool formed after evaluation and optimization processing. This model can support intelligent talent resource allocation, job recommendation, and organizational capability building. In this embodiment, the refined talent pool digital model data can be output by the evaluation and optimization processing, integrating the latest talent characteristic data with optimized organizational logic.
[0081] Taking the identification of high-potential talents in R&D positions in high-tech enterprises as an example, the talent pool modeling method based on a large language model in this embodiment can be as follows: The system collects project documents, Git commit records, and online coding assessment interaction logs of an algorithm engineer; the system discretizes the engineer's workflow into task atoms such as data cleaning, model design, debugging, and deployment using a large language model, and labels each atom with the required skills and logical dependencies; at the same time, the system converts the engineer's coding behavior, debugging response, and team collaboration text into a sequence of behavioral atoms, and through attention mechanism modeling, it finds that the engineer exhibits strong context adaptability in complex problems; combined with the requirements of distributed training and model compression for AI researcher positions, the system predicts that the engineer's competence in model deployment task nodes is at a high level, and generates a cumulative competence index and collaborative skill mapping showing that the engineer's PyTorch and Kubernetes skills have a strong complementary effect; finally, the talent characteristic data triggers the structural adjustment of the algorithm engineer sub-library in the talent pool, classifying the engineer into a high-potential distributed AI talent cluster, supporting subsequent key project team recommendations.
[0082] In one embodiment, the talent workflow is discretized based on historical work text data to obtain task atomic data containing task node attribute information and execution logic information, including:
[0083] The talent workflow is divided into tasks based on historical work text data to obtain task data, where the task density is set to no less than 5 behavior atoms per thousand words of text.
[0084] In this context, task data can be a set of coarse-grained task units initially segmented from the talent workflow after task partitioning. This data can serve as intermediate input for attribute assignment and logical relationship calculation, supporting the generation of atomic topology attribute data. In this embodiment, task data can be based on historical work text data, utilizing a large language model to identify semantic boundaries and action subjects, dividing continuous narratives into independent task segments. Task density can be an indicator of the number of behavioral atoms parsed per unit text length, used to control the fineness of task partitioning. It forces an increase in parsing granularity by setting a lower threshold, avoiding sparsity of capability units. A behavioral atom can be the smallest behavioral unit with a clear action intent and execution subject identified during task partitioning. It can be used as the basic counting unit for task density calculation, reflecting the fineness of task partitioning. A task atom can be a task unit with complete semantic and functional boundaries, formed by the aggregation of one or more behavioral atoms. It can be used as the basic object for attribute assignment and logical modeling, differing from task atoms directly generated end-to-end from text in previous schemes; here, it is emphasized that it is aggregated from behavioral atoms.
[0085] Based on historical work text data, the talent workflow is divided into tasks to obtain task data. This can be achieved by using a large language model to identify action verbs, roles, and task boundaries in the text, segmenting continuous work descriptions into preliminary task units. Further, this operation can be implemented by using a dependency syntax-based action-object extraction framework to divide tasks, guiding the large language model to output a task list through prompting engineering, and then clustering and merging the results. This generates structured task data, providing a foundation for subsequent refined atom construction. The task density is set to no less than 5 behavioral atoms per thousand characters of text. This can be achieved by introducing a density monitoring mechanism during task division; if the current division result does not reach the threshold, fine-grained re-parsing is triggered. Further, this operation can be achieved by dynamically adjusting the segmentation granularity parameters of the large language model until the density requirement is met, splitting long sentences into clauses and extracting behavioral atoms layer by layer, thus forcibly increasing the parsing granularity and avoiding the loss of capability units due to text abstraction.
[0086] Based on the task data, attribute values are assigned and logical relationships are calculated for each task atom to obtain atomic topology attribute data containing atomic attribute information and logical connection information.
[0087] The process of assigning attributes can be a process of labeling task atoms with static features such as their skill domain, tool type, and output format. This can be used to generate atomic attribute information to support subsequent competency mapping. Logical relationship calculation and processing can analyze the dependencies between task atoms in terms of time, causality, or condition, and quantify their connection strength. This can be used to generate logical connection information and construct the execution topology between tasks. Atomic attribute information can be a set of static capability tags carried by task atoms, such as programming languages, management methods, and industry knowledge, which can be used to characterize the capability requirements of a task. For example, atomic attribute information can include hard skill attributes, soft skill attributes, and domain knowledge attributes. Logical connection information can be structured data describing the dependencies between task atoms, including sequential, parallel, and conditionally triggered relationship types, which can be used to reconstruct the execution context in a real workflow. In a specific embodiment, logical connection information can include temporal dependency connections, conditional branch connections, and resource contention connections. Atomic topological attribute data can be a structured task atom representation formed by fusing atomic attribute information and logical connection information, reflecting its position and characteristics in the task network. It can be used to transform task nodes from isolated labels into knowledge graph units, supporting topological reasoning and path analysis. In this embodiment, atomic topological attribute data can jointly encapsulate the results of attribute assignment and logical relationship calculation to form a task atom description with a graph structure. Furthermore, atomic topological attribute data can serve as the input basis for unique semantic identifier allocation processing, determining the semantic distinctiveness of the identifier.
[0088] Based on the task data, attribute assignment and logical relationship calculation are performed on each task atom. This can be achieved by performing semantic classification (attribute assignment) and relational reasoning (logical calculation) in parallel for each task atom, outputting structured attributes and connection information. Furthermore, this operation can be implemented by using a multi-task large language model to simultaneously predict attribute labels and adjacent tasks, and by training attribute classifiers and graph neural networks separately for logical relationship modeling. This allows for the simultaneous construction of semantic features and topological relationships of task atoms, forming atomic topological attribute data. The atomic topological attribute data, which contains atomic attribute information and logical connection information, can be obtained by aligning the attribute assignment results and logical relationship calculation results by task atom ID and encapsulating them into a unified data structure. This enables structured knowledge representation of task atoms, supporting graph-based storage and reasoning.
[0089] Based on the atomic topology attribute data, unique semantic identifiers are assigned to the task atoms to obtain task atom data with index identifiers.
[0090] The unique semantic identifier allocation process can be a process of generating globally unique, semantically parsable numerical identifiers for each task atom based on atomic topological attribute data. This ensures that task atoms can be accurately aligned and traced across different personnel, positions, and timeframes. The index identifier can be a unique string or hash code appended to the task atom, encoding its semantic content and topological location information. This supports efficient retrieval, deduplication, and semantic matching. Task atom data with index identifiers can be standardized task atom data formed after unique semantic identifier allocation, containing attributes, logic, and a unique index. This can be used as a high-fidelity, computable basic data unit for subsequent behavior modeling and matching evaluation. In an exemplary embodiment, the task atom data with index identifiers is the final output form of this embodiment, which adds topological structure and unique identifier constraints compared to previous solutions.
[0091] Assigning unique semantic identifiers to task atoms based on their topological attribute data can be achieved by generating irreversible semantic hashes or path codes based on the combination of task atom attributes and their topological location. Furthermore, this operation can be implemented by using content hashing (such as SHA3) combined with topological paths to generate composite identifiers, or by constructing identifiers using semantic embedding cluster center IDs plus local offsets. This gives task atoms globally unique and semantically interpretable identities, supporting cross-domain alignment. Obtaining task atom data with indexed identifiers can be achieved by appending unique semantic identifiers to the atomic topological attribute data, forming the final standardized output. This results in task atom data with high structure, traceability, and computability, providing high-quality input for subsequent processes.
[0092] Taking the modeling of a financial risk control analyst position as an example, the talent pool modeling method based on a large language model in this embodiment can be as follows: The system processes a 2,000-word text describing the three-year work experience of a risk control expert. During the task division stage, 12 behavioral atoms (such as designing scorecards, backtesting, regulatory reporting, etc.) are identified, meeting the density requirement of no less than 5 atoms per thousand words. Subsequently, the relevant behavioral atoms are aggregated into 4 task atoms (such as model development, compliance verification, system deployment). Each task atom is assigned attribute values (such as Python, Basel Accords, SQL) and its logical connections are calculated (such as model development must precede compliance verification). Based on this topological attribute data, a unique semantic identifier is generated, for example, by fusing skill tags and prior task IDs with hash encoding. The final output task atom data with index identifiers can be accurately matched to the newly established anti-money laundering model position. The system recognizes that the individual has high competence in the model development task node, and their compliance verification experience can be transferred to new regulatory scenarios.
[0093] In one embodiment, talent behavior patterns are dynamically modeled based on real-time interactive feedback data to obtain behavioral atomic semantic sequence data, including:
[0094] Extract the operation response latency, contextual relevance, semantic position, and logical evolution direction of each behavioral atom from real-time interactive feedback data, and establish a structured parameter table for each behavioral atom;
[0095] Operational response latency refers to the time interval between the triggering of a behavioral atom and the completion of system recording, reflecting the efficiency of human reaction and decision-making rhythm. Operational response latency can serve as a key indicator for measuring the dynamic characteristics of behavior, distinguishing between efficient response and sluggish behavior patterns. Contextual relevance refers to the strength of the semantic or task-objective correlation between the current behavioral atom and its preceding or subsequent behaviors. Contextual relevance can quantify the contextual dependence of behavior in the overall interaction flow, supporting logical evolution modeling. Semantic position refers to the relative semantic stage of a behavioral atom within its complete semantic cycle. Semantic position can identify the functional role of a behavior in the task semantic flow, assisting in the allocation of logical starting points.
[0096] The logical evolution direction can be the logical tendency of behavioral atoms to drive the task state forward in the sequence. The logical evolution direction can be used to characterize the driving effect of behavior on the overall process and support the generation of personalized logical dependency descriptions. In a specific embodiment, the logical evolution direction can include progressive evolution direction, backtracking correction direction, parallel expansion direction, etc. The structured parameter table of each behavioral atom can be a data structure that organizes the multi-dimensional features (including operation response latency, contextual relevance, semantic position, and logical evolution direction) of each behavioral atom in tabular form. The structured parameter table of each behavioral atom can be stored aligned by behavioral atom ID after extracting features from real-time interactive feedback data. Extracting the operation response latency, contextual relevance, semantic position, and logical evolution direction of each behavioral atom from real-time interactive feedback data can be achieved by using a large language model to perform multi-dimensional semantic parsing of the interaction log and simultaneously outputting four types of behavioral features. Furthermore, this operation can be achieved by using multi-head prompts to extract the four types of features separately, then fusing them, and training a multi-task sequence labeling model to predict all dimensions end-to-end, thereby transforming abstract interactive behavior into structured, quantifiable dynamic features. Establishing a structured parameter table for each behavioral atom can be achieved by organizing the extracted four types of features into a matrix data table based on the behavioral atom ID, thereby forming standardized behavioral feature inputs that facilitate subsequent calculation and processing.
[0097] The logical weight value of each behavior is determined by the total number of atoms in the parameter table, and a logical configuration list is generated by assigning different logical starting points to each atom by equally dividing the complete semantic cycle.
[0098] The total number of atoms can be the total number of behavioral atoms extracted within a single interaction cycle. The total number of atoms can be used to normalize the logical weight values of each behavior, ensuring that the total weight is controllable. The logical weight values can be importance coefficients assigned to each behavioral atom based on the total number of atoms and individual behavioral characteristics. Logical weight values can reflect the differences in contribution of different behaviors to the overall competence assessment, avoiding treating all behaviors equally. The complete semantic cycle can be the entire semantic process covered by a complete interaction task from start to finish. The complete semantic cycle can be used as a unified temporal-semantic framework for equal segmentation to assign logical starting points. Logical starting points can be logical time anchors assigned to each behavioral atom within the equally segmented semantic cycle segments. Logical starting points can be used to assign relative positions to behavioral atoms in a unified semantic temporal sequence, supporting sequence alignment. The logical configuration list can be a configuration file recording the logical starting points and logical weight values corresponding to each behavioral atom. The logical configuration list can be jointly generated by determining the weights from the total number of atoms and determining the starting points from the complete semantic cycle segmentation.
[0099] The logical weight of each behavior is determined using the total number of atoms in the parameter table. This can be achieved by normalizing based on the total number of atoms and adjusting individual weights in conjunction with latency or relevance. Furthermore, this operation can be implemented by using a 1 / N base weight with a latency penalty factor and using contextual relevance as a weight amplification coefficient, thus enabling differentiated modeling of behavior importance and avoiding equalization. Different logical starting points are assigned to each atom by equally dividing the complete semantic cycle. This can be achieved by dividing the interaction cycle into a preset number of segments and mapping the semantic position of the behavior atom to the nearest segmentation point. Further, this operation can be achieved by using linear interpolation to assign non-integer starting points and dynamically adjusting the segmentation granularity in conjunction with semantic density, thereby assigning relative positions to behavior atoms within a unified semantic temporal framework and supporting sequence alignment. Generating a logical configuration list can be achieved by combining the logical weight of each behavior atom with its logical starting point into configuration entries, thus outputting the temporal and importance configuration metadata required for behavior modeling.
[0100] Based on the parameter table and the logical configuration list, generate a set of behavioral atom sequences containing personalized logical dependency descriptions for each behavioral atom;
[0101] The personalized logical dependency description can be a customized dependency statement generated by combining the characteristics of the behavioral atom itself and its position in the logical configuration list. This personalized logical dependency description replaces the general dependency template, enabling precise characterization of individual behavioral patterns. The behavioral atom sequence set can be a structured set containing each behavioral atom and its personalized logical dependency description. This set can serve as the basic data source for semantic sampling, supporting stateful modeling. Generating a behavioral atom sequence set containing personalized logical dependency descriptions for each behavioral atom based on the parameter table and logical configuration list can be achieved by fusing behavioral features and configuration information, and generating customized dependency description text through a large language model. Furthermore, this operation can be implemented by generating descriptions by filling key parameters based on templates and directly outputting natural language dependency statements using a fine-tuned model, thereby achieving individualized behavioral dependency modeling and overcoming the expressive limitations of general sequence models.
[0102] Based on the set of behavioral atom sequences, the entire interaction cycle is sampled according to a preset semantic interval, and the logical connection state of all behavioral atoms at each semantic point is calculated, thereby forming a behavioral state sequence table arranged in semantic order.
[0103] The preset semantic interval can be a fixed semantic step size set manually within a complete semantic cycle for periodic sampling. The preset semantic interval can control the granularity and coverage density of the behavior state sequence list, balancing computational overhead and representational accuracy. The interaction cycle can be the time span for a candidate to complete a full interaction in a specific task scenario. A semantic point can be a discrete semantic sampling moment defined under the preset semantic interval. A semantic point can be used as the basic time unit for calculating logical connection states. A logical connection state can be the set of dependencies activated between all behavioral atoms at a specific semantic point and their intensity distribution. The logical connection state can reflect the cooperative state of the behavioral system at that moment, constituting the core content of the behavior state sequence list. The behavior state sequence list can be a structured sequence arranged semantically, containing logical connection states at each semantic point. The behavior state sequence list can be generated by sampling the set of behavioral atom sequences at each semantic point and aggregating logical connection states. Furthermore, the behavior state sequence list can be directly used to establish a fast query correspondence between timestamps and behavioral semantic information.
[0104] Sampling the entire interaction cycle based on a set of behavioral atom sequences at preset semantic intervals can be achieved by scanning the set of behavioral atom sequences at equidistant semantic points to collect active behaviors and their dependencies, thereby discretizing the continuous behavior flow into computable state sampling points. Calculating the logical connection states of all behavioral atoms at each semantic point can be achieved by aggregating the dependencies between the behavioral atoms involved in that semantic point to form a state snapshot. Furthermore, this operation can be implemented by using graph neural networks to aggregate adjacency relationships and using a rule engine to parse dependent activation states, thereby generating an instantaneous state representation reflecting the co-evolution of behaviors. Forming a sequence of behavioral states arranged semantically can be achieved by concatenating the logical connection states of each semantic point into a sequence in chronological order, thereby constructing a stateful and temporally sequential high-order behavioral representation.
[0105] By using a behavior state sequence table, a fast query correspondence between timestamps and behavioral semantic information is established, thereby obtaining behavioral atomic semantic sequence data.
[0106] The timestamp can be the system's absolute time stamp corresponding to each semantic point in the behavior state sequence table. The timestamp can be used to align the semantic sequence with the real timeline, supporting real-time queries. Behavioral semantic information can be the comprehensive semantic content carried by the behavior state at a specific timestamp (such as collaborative verification, independent debugging, etc.). Behavioral semantic information can be used as the target output for fast queries, for matching calculations. The fast query correspondence can be a mapping mechanism from timestamps to behavioral semantic information based on an index structure. The fast query correspondence can be achieved by binding timestamps to corresponding semantic states using hash tables or inverted indexes. Establishing a fast query correspondence between timestamps and behavioral semantic information using the behavior state sequence table can be achieved by attaching a timestamp to each state entry and building an index structure to support O(1) queries. Furthermore, this operation can be implemented by using a Redis hash table to store timestamp-semantic pairs and building an Elasticsearch inverted index to support range queries, thereby achieving efficient retrieval and real-time inference of behavioral semantics. Obtaining behavioral atomic semantic sequence data can be achieved by encapsulating an indexed behavioral state sequence table into a standard output format, thereby outputting behavioral semantic sequences with stateful, temporal, and queryable characteristics, supporting dynamic matching.
[0107] Taking engineer behavior modeling in online programming interviews as an example, the talent pool modeling method based on a large language model in this embodiment can be as follows: The system collects a candidate's keyboard input, compilation logs, and Q&A records during a 90-minute coding interview; extracts 28 behavioral atoms (such as defining functions, debugging breakpoints, and querying interface documentation), and records features such as operation response latency (average 1.2 seconds) and contextual relevance (such as debugging behavior being highly correlated with preceding code); calculates the logical weight of each behavior based on the total number of atoms (28), and divides the complete 90-minute semantic cycle into 10 segments, assigning logical weights to each behavior. Starting point; after generating the logical configuration list, output personalized logical dependency descriptions in conjunction with the parameter table (e.g., the debugging behavior depends on the three preceding coding behaviors and has a high correction intent); sample at a semantic interval of 9 minutes, calculate the activation state between behaviors at each semantic point (e.g., strong connection between coding and debugging behaviors in the 3rd interval); finally form a behavior state sequence table, and establish a fast query relationship from timestamp to semantic information (e.g., 14:05-14:14 is in the "module integration verification" state), output behavior atomic semantic sequence data, which is used for subsequent dynamic matching with backend development task atoms.
[0108] In one embodiment, dynamic matching degree calculation is performed on talent competency and skill distribution based on task atomic data and behavioral atomic semantic sequence data to obtain talent characteristic data including cumulative competency index and collaborative skill mapping distribution, including:
[0109] The connectivity calculation of behavior and task atoms is performed based on the semantic sequence data of behavior atoms and the data of task atoms. By calculating the semantic association distance between each behavior atom at each semantic point and each task atom in the talent pool, the behavior-task connectivity matrix data is obtained.
[0110] The connectivity calculation of behavior and task atoms can be a process of measuring the degree of matching between each behavior in the semantic sequence of behavior atoms and task atoms in the talent pool in the semantic space. This can provide a basic logical framework for subsequent semantic association distance calculation. In this embodiment, the connectivity calculation of behavior and task atoms can encode behavior atoms and task atoms into a unified semantic vector space and calculate their semantic association distance. Further, this process can be achieved by using contrastive learning to align the behavior and task embedding spaces, or by directly outputting association scores using a cross-modal large language model, thereby achieving cross-modal semantic alignment between behavior and task. The semantic association distance can be a measure of the vector distance between behavior atoms and task atoms in the embedded semantic space, reflecting their semantic similarity. It can be used to quantify the support strength of behavior for a specific task, serving as a basis for connectivity judgment. The behavior-task connectivity matrix data can be structured data recording the semantic association distance between each behavior atom and all task atoms at each semantic point in matrix form. This can be used to achieve cross-modal semantic alignment of behavior and task, supporting the determination of the effective skill range. In one exemplary embodiment, the behavior-task connectivity matrix data can be mapped to a unified semantic space using a large language model, and then the pairwise distances can be calculated and organized into a matrix.
[0111] Based on the behavioral task connectivity matrix data and the preset job competency threshold, the effective skill range is determined. By comparing connectivity, task atoms within the effective skill range of each behavioral atom are selected to obtain behavioral skill coverage relationship data.
[0112] The preset job competency threshold can be a semantic association distance upper limit set according to job requirements. This threshold defines whether a behavior effectively supports a task and can filter weakly related task atoms, ensuring the assessment focuses on the core competency domain of the job. The effective skill range determination process, based on the behavior-task connectivity matrix data and the job competency threshold, filters out task atoms with a semantic association distance below the threshold. This dynamically defines the effective task set that each behavior atom can cover. This process compares each distance value in the matrix with the threshold, marking items below the threshold as effective coverage, thus dynamically filtering irrelevant tasks and focusing on the core competency domain. Behavioral skill coverage relationship data can be structured data recording the mapping relationship between each behavior atom and its effectively covered task atoms. This clarifies the boundary of a behavior's ability to support a task, avoiding interference from irrelevant tasks. Furthermore, behavioral skill coverage relationship data can serve as the input basis for calculating instantaneous competency values. This process extracts effective coverage items and constructs a behavior-task binary relationship list, thereby clarifying the boundary of a behavior's ability to support a task.
[0113] Based on the behavioral skill coverage relationship data, the instantaneous competence value of each task atom is calculated using a skill allocation distribution model based on an exponential decay function, thus obtaining the instantaneous competence data of the atom.
[0114] The skill allocation distribution model using the exponential decay function can be a weighting function that decays exponentially over time. This model models the time-sensitive contribution of behavior to current competence, allowing recent behaviors to have a greater impact on competence and reflecting the dynamic evolution of capabilities. Atomic instantaneous competence data can be the immediate competence value of a single behavior on a task atom, calculated based on effective coverage relationships and the decay model at a specific semantic point. This can be used as the basic unit for cumulative calculation, reflecting the capability contribution of a behavior at a given moment. Furthermore, the atomic instantaneous competence data can be input into a semantic sequence-based cumulative calculation processing module. This calculation applies an exponential decay weight to each effective coverage relationship based on the distance of the behavior's occurrence time from the current semantic point, calculating a weighted competence value. This calculation can be further implemented using an exponential function with a fixed half-life or by adaptively adjusting the decay coefficient based on job type, thereby introducing time-sensitive modeling and reflecting the dynamic evolution of capabilities.
[0115] Based on the instantaneous competence data of atoms, a semantic sequence-based cumulative calculation process is performed to obtain the cumulative competence index data of atoms. Specifically, the cumulative calculation process is performed by semantic integration on the competence of each task atom during the entire interaction cycle.
[0116] The semantic sequence-based cumulative computation can be a process of aggregating the instantaneous competence values of all atoms within the entire interaction cycle. This can be used to integrate the impact of discrete behaviors into a continuous and comparable comprehensive index. Semantic integration can be a mathematical operation that performs a weighted integration of instantaneous competence values on the semantic sequence timeline. The weights are determined by both semantic density and timeliness, and can be used to achieve a smooth transition from instantaneous values to a cumulative index, preserving the temporal information of the behavior. In an exemplary embodiment, semantic integration can include, but is not limited to, Riemann summation, kernel smoothing integration, and attention-weighted integration. This operation can perform a weighted summation of instantaneous competence values on the semantic timeline, with the weights determined by both semantic density and timeliness. Furthermore, this operation can be implemented by using the trapezoidal rule for approximate integration or by automatically learning the integration weights using an attention mechanism, thereby preserving temporal information while achieving smooth accumulation. The atomic cumulative competence index data can be a comprehensive index obtained after semantic integration, representing the candidate's long-term ability accumulation on each task atom. This can be used to quantify an individual's overall competence level at specific task nodes, supporting high-potential identification. Furthermore, atomic cumulative competency index data can be used simultaneously to overlay collaborative skill mapping and final talent characteristic construction.
[0117] The atomic cumulative competence index data is used to superimpose and calculate the effect of multi-behavior atomic synergy to obtain atomic synergy skill mapping distribution data;
[0118] The synergistic effect of multiple behavioral atoms can be described as a nonlinear capability enhancement or complementary effect generated when multiple behavioral atoms cover the same task atom. This can be used to reveal the joint contribution mechanism of skill combinations to complex tasks. The superposition calculation process can be a computational process that nonlinearly fuses the cumulative competence of multiple behavioral atoms to the same task atom, and can be used to generate a mapping distribution reflecting the synergistic effect of multiple skills. This process can nonlinearly fuse the cumulative values of multiple behaviors covering the same task atom (e.g., product, max pooling, or gated weighting). Furthermore, this process can be implemented by using neural networks to learn cooperative functions or by combining complementary / substitution logic based on rule definitions, thereby revealing the nonlinear enhancement effect of multi-skill synergy on complex tasks. Atomic synergistic skill mapping distribution data can be a probability distribution describing the joint activation intensity and interaction mode of multiple skills when completing a specific task atom. This can be used to reveal multi-skill synergistic mechanisms and avoid the one-sidedness of single-skill evaluation. Furthermore, atomic synergistic skill mapping distribution data can be combined with atomic cumulative competence index data to constitute talent characteristic data.
[0119] Based on the atomic cumulative competence index data and the atomic collaborative skill mapping distribution data, the cumulative competence information and the collaborative skill mapping distribution information are matched according to the task atomic index identifier to obtain talent characteristic data.
[0120] The task atom index identifier can be a unique semantic identifier assigned to a task atom in previous schemes, serving as an anchor point for aligning cumulative competency information with collaborative skill mapping distribution information. Furthermore, the task atom index identifier ensures precise matching of the two types of data along the same task dimension. Cumulative competency information can be the capability accumulation representation content carried by the atomic cumulative competency index data, reflecting an individual's long-term performance stability on tasks. Collaborative skill mapping distribution information can be the multi-skill interaction pattern content carried by the atomic collaborative skill mapping distribution data, revealing the capability combination structure required for complex tasks. This matching, using the task atom index identifier as the key, aligns the two types of data along the task dimension and encapsulates them into a unified structure, thereby outputting a structured, two-dimensional talent capability representation to support intelligent matching and recommendation.
[0121] Taking the SaaS product director job suitability assessment as an example, the talent pool modeling method based on a large language model in this embodiment can be as follows: The system obtains the semantic sequence of behavioral atoms (including behaviors such as requirement clarification, prototype feedback, and priority ranking) of a candidate's online product review meeting and the "develop product roadmap" task atom in the talent pool; calculates the semantic association distance between each behavior and the task, and finds that the distance of the requirement clarification behavior is 0.32 (lower than the job threshold of 0.4), and includes it in the effective coverage; based on the fact that this behavior occurs 15 minutes before the meeting, the instantaneous competence value is obtained by applying the exponential decay function (half-life of 10 minutes); semantic integration is performed on all relevant behaviors in the entire 90-minute interaction period to obtain the cumulative competence index of the task atom as 0.75; at the same time, it is found that the two behaviors of requirement clarification and priority ranking jointly cover the task, and the superposition calculation shows that their collaborative skill mapping distribution shows a strong complementary effect (skill combination score increases by 22%); finally, using the unique index identifier of the task atom as the key, the cumulative index of 0.75 and the complementary collaborative distribution are integrated into talent feature data for matching senior product positions.
[0122] In one embodiment, effective skill range determination is performed based on behavioral task connectivity matrix data and a preset job competency threshold. Task atoms within the effective skill range of each behavioral atom are selected through connectivity comparison to obtain behavioral skill coverage relationship data, including:
[0123] The correlation determination results are obtained by comparing the behavioral task connectivity matrix data and the preset job competency thresholds one by one.
[0124] The correlation value can be a numerical value in the behavior-task connectivity matrix representing the semantic correlation strength between behavior atoms and task atoms. It can be used as the original basis for comparison with the competency threshold to determine whether to include a pair within the effective coverage scope. The competency threshold can be a specific numerical instance of a preset job competency threshold in this operation. It can be used to control the strictness of skill coverage, ensuring that only behavior-task pairs with strong job relevance are retained. The correlation determination result data can be a set of Boolean or labeled determination results generated after comparing each correlation value in the behavior-task connectivity matrix with the job competency threshold. It can be used to identify whether each behavior-task pair meets the effective skill coverage conditions, serving as a screening criterion.
[0125] Based on the correlation determination results, behavioral task pairs that meet the skill range conditions are screened and extracted. By retaining behavioral task combinations with correlation values less than or equal to the competence threshold, a list of effective behavioral task pairs is obtained.
[0126] Among them, behavior-task pairs that meet the skill range criteria can be pairs of behavior atoms and task atoms with an association value less than or equal to the competence threshold, which can be used as the basic unit to construct the list of effective behavior-task pairs. Behavior-task combinations can be tuples consisting of a behavior atom and a task atom, which can be used as the basic data structure unit for filtering and grouping. The list of effective behavior-task pairs can be an ordered list of all behavior-task combinations that meet the skill range criteria after filtering, which can be used to clarify which behaviors support which tasks. For example, the list of effective behavior-task pairs can be obtained by extracting valid items from the association determination result data and storing them in a structured manner. In a specific embodiment, the list of effective behavior-task pairs can be input into a grouping processing module based on atom identifiers.
[0127] Based on the list of valid behavioral task pairs, the skill task set of each behavioral atom is grouped according to the atom identifier to obtain behavioral group atom set data;
[0128] Here, the atom identifier can be a unique semantic identifier for a task atom, which can be used as a grouping key to ensure accurate aggregation at the task atom level. The skill task set can be the collection of all task atoms covered by a single behavior atom, reflecting the capability reach of that behavior. The behavior grouping atom set data can be structured data organized at the behavior atom level, containing its corresponding skill task set, which can be used to implement a task coverage view from a behavior perspective, supporting the construction of behavior capability profiles. In this embodiment, the behavior grouping atom set data can be obtained by traversing the list of valid behavior task pairs and aggregating task atoms by behavior atom ID. Furthermore, the behavior grouping atom set data can be used to generate a list of behaviors covered by atoms and a bidirectional index mapping table.
[0129] Based on the behavior grouping atomic set data, the skill task set of each behavior atom is traversed and the number of covered behaviors of each atom is counted to obtain the list of covered behaviors of the atom.
[0130] The number of covered behaviors can be a count of how many different behavior atoms cover each task atom, which can be used to quantify the multi-behavior support of a task and reflect skill redundancy and synergy potential. In an exemplary embodiment, the number of covered behaviors can include single-point coverage, multi-point coverage, and high-redundancy coverage. The list of covered behaviors can be a list recording each task atom and its corresponding number of covered behaviors, which can be used to provide behavioral support strength information from a task perspective and assist in synergy modeling. For example, the list of covered behaviors can be obtained by traversing the behavior group atom set data and counting the frequency of occurrence of each task atom. In this embodiment, the list of covered behaviors can be used together with the behavior group atom set data to construct a bidirectional index.
[0131] By constructing a bidirectional index mapping table from behavior to atom and from atom to behavior using behavior grouping atom set data and a list of behaviors covered by atoms, behavior skill coverage relationship data can be obtained.
[0132] The bidirectional index mapping table from behavior to atom and from atom to behavior can be an index structure that simultaneously supports forward queries from behavior to task and reverse tracing from task to behavior. This allows for efficient bidirectional retrieval, supporting instantaneous competency calculation and collaborative overlay analysis. In one specific embodiment, this bidirectional index mapping table can be obtained by building a forward hash table based on behavior-grouped atom set data and a reverse inverted index based on the list of behaviors covered by atoms. Furthermore, the bidirectional index mapping table from behavior to atom and from atom to behavior can be the final representation of behavior skill coverage relationship data.
[0133] Taking the job suitability analysis for AI algorithm engineers as an example, the talent pool modeling method based on large language models in this embodiment can be as follows: The system processes the behavioral sequence of a candidate in an online coding assessment and 200 task atoms in the talent pool; the behavioral task connectivity matrix shows that the correlation value between the "debug model" behavior and the "hyperparameter tuning" task is 0.35, which is lower than the job threshold of 0.4, and is marked as valid; after screening, 128 valid behavioral task pairs are obtained; after grouping by behavioral atoms, it is found that the "read document" behavior covers 5 tasks and the "submit code" behavior covers 8 tasks; statistics show that the "model deployment" task is covered by 7 different behaviors, indicating that it has a high degree of collaborative support; the finally constructed bidirectional index allows the system to quickly query which tasks the "debug model" behavior can support, and can also reversely find out which behaviors jointly support the "model deployment" task. This behavioral skill coverage relationship data is then used to calculate the cumulative competence and collaborative skill distribution of the candidate on MLOps-related tasks.
[0134] In one embodiment, based on behavioral skill coverage relationship data, the instantaneous competence value of each task atom is calculated using a skill allocation distribution model based on an exponential decay function, resulting in atom instantaneous competence data, including:
[0135] Extract the semantic association distance information between each behavioral atom and the task atoms within its skill range from the behavioral skill coverage relationship data to obtain an effective semantic distance value data set;
[0136] Normalized distance ratio data is obtained by using a set of effective semantic distance values and a preset exponential decay parameter to perform normalized distance calculation.
[0137] Based on the normalized semantic distance ratio data, the normalized semantic distance ratio is multiplied by a negative exponential decay coefficient and the exponential function is taken to obtain the exponential decay factor data.
[0138] The skill value is calculated based on the exponential decay factor data and the preset maximum skill capacity parameter of the atom to obtain the basic skill allocation value of the task atom.
[0139] Based on the assigned values of basic skills of the task atoms, the data is processed by correcting and solving the results based on the direction of behavior atomic operations and the phase angle of task semantics, thereby obtaining the instantaneous competence data of atoms.
[0140] The semantic association distance information between each behavior atom and its task atoms within its skill range can be the original semantic distance values between valid behavior-task pairs extracted from behavior skill coverage relationship data and filtered through a threshold. This can be used as the original similarity basis for instantaneous competence calculation, reflecting the semantic matching strength between behavior and task. In this embodiment, the semantic association distance information between each behavior atom and its task atoms within its skill range can be obtained by traversing valid pairings in the bidirectional index mapping table and querying the original behavior-task connectivity matrix to obtain the corresponding distance values. The valid semantic distance value dataset can be a structured set composed of the semantic association distance values of all valid behavior-task pairs, which can be used to provide the input data basis for normalization processing. In an exemplary embodiment, the valid semantic distance value dataset can be obtained by traversing valid pairings from the bidirectional index mapping table and extracting the corresponding distance values.
[0141] The preset exponential decay parameter can be a system configuration parameter that controls the steepness of the exponential decay function, determining the rate at which semantic distance affects competence. It can be used to adjust the model's sensitivity to semantic deviations and avoid excessive interference from weakly correlated behaviors at long distances in the evaluation. Normalized distance calculation can be the process of mapping the original semantic distance value to the [0, 1] interval, typically scaled based on the maximum possible distance. This can be used to eliminate dimensional differences between different tasks or behaviors, making distance values comparable. The normalized semantic distance ratio data can be the normalized semantic distance value, representing the relative degree of deviation. It can be used as an input variable for the exponential decay function; the smaller the value, the closer the semantics. The negative exponential decay coefficient can be a negative real constant used to control the decay rate, typically equal to -1 divided by the decay time constant. This ensures that the decay factor monotonically decreases as the distance increases, conforming to cognitive patterns. The exponential decay factor data can be a weight coefficient obtained by applying an exponential function to the normalized semantic distance ratio. It reflects the contribution of semantic matching to competence and can be used to realize a non-linear mapping from semantic distance to competence weight, reflecting the decay characteristic of "nearer is stronger, farther is weaker". In this embodiment, the exponential decay factor data can be obtained by multiplying the normalized semantic distance ratio by a negative exponential decay coefficient and then taking the exponential function.
[0142] The preset maximum skill capacity parameter for atoms can be the upper limit of the maximum skill allocation that each task atom can accept, reflecting the capability boundary of the task. This can be used to prevent over-emphasis on competence due to excessive weighting of a single behavior, ensuring that the assessment conforms to realistic constraints. The basic skill allocation value of a task atom can be the initial competence value before considering direction correction, obtained by multiplying the exponential decay factor by the maximum skill capacity, and can be used as a benchmark input for direction correction. In an exemplary embodiment, the basic skill allocation value of a task atom can be input to a correction processing module based on operation direction and phase angle. The operation direction of a behavior atom can be the execution intention of the behavior atom in the semantic vector space, such as the semantic direction of actions like verification, construction, and optimization, and can be used to determine whether the behavioral intention is consistent with the task goal. The task semantic phase angle can be the required behavioral target direction of the task atom in the semantic space, representing the current semantic stage of the task and the expected behavior type, and can be used as a reference benchmark for behavior direction alignment. The correction and result solving process can be a calculation process that adjusts the basic skill allocation value based on the cosine value of the angle between the direction of the behavior atomic operation and the phase angle of the task semantics. It can be used to ensure that skills can be effectively converted into competencies only when the behavior intention is aligned with the semantics of the task objective.
[0143] Taking online collaborative review by front-end development engineers as an example, the talent pool modeling method based on a large language model in this embodiment can be as follows: The system analyzes an engineer's behavior of "suggesting UI optimization" in a code review meeting. The semantic association distance between this behavior and the task of "refactoring component interaction logic" is 0.28. After normalization, the distance is 0.28 (assuming the maximum distance is 1.0). Let the attenuation coefficient λ = 2.0, then the attenuation factor is exp(-2.0 × 0.28) = 0.57. The maximum atomic skill capacity of this task is set to 0.9, resulting in a basic skill allocation value of 0.513. Further analysis shows that the operation direction of "suggesting" is verification-type, while the task is currently in the verification phase, with a phase angle cosine similarity of 0.92. The final atomic instantaneous competence is 0.513 × 0.92 ≈ 0.472. If the same behavior is applied to a task in the requirement understanding phase, the direction is mismatched (cosine = 0.3), and the competence drops to 0.154, demonstrating the key role of direction correction.
[0144] In one embodiment, the data is processed by correcting and solving the results based on the task atomic basic skill allocation values and the behavioral atomic operation direction and task semantic phase angle, thereby obtaining atomic instantaneous competence data, including:
[0145] Based on the basic skill allocation values of task atoms, the phase difference between behavioral atoms and task semantics is calculated and the angle is corrected. The cosine of the phase angle between the operation direction of the behavioral atom and the task semantics is calculated and used as the correction coefficient to obtain the skill allocation values after angle correction.
[0146] The phase difference can be the angle between the action atom operation direction vector and the task semantic phase angle vector, used to measure their alignment in the semantic space. In this embodiment, the phase difference can serve as the geometric basis for calculating the cosine correction coefficient. The cosine value can be the output of the cosine function of the phase difference angle, ranging from -1 to +1, reflecting the strength of directional consistency. The cosine value can be directly used as a correction coefficient to quantify the semantic alignment quality. The correction coefficient can be a multiplicative weight directly applied by the cosine value to adjust the basic skill allocation value. In a specific embodiment, the correction coefficient can be applied to the task atom basic skill allocation value to generate an angle-corrected skill allocation value. The angle-corrected skill allocation value can be a skill value modulated by the cosine correction coefficient, reflecting the effective capability contribution after semantic direction alignment. In this embodiment, the angle-corrected skill allocation value can be obtained by multiplying the task atom basic skill allocation value by the cosine value.
[0147] Based on the skill allocation value and the behavioral atom response speed information of the current semantic point, the speed influence factor is calculated and processed to obtain the behavioral atom speed modulation skill data.
[0148] The current semantic point can be a discrete semantic sampling moment defined in previous schemes, divided within a preset semantic interval. In this embodiment, the current semantic point can limit the effective time window of the response speed information to ensure timeliness. The behavioral atom response speed information can be a time efficiency indicator of the behavioral atom from triggering the task context to actual execution completion, usually in seconds. Furthermore, the behavioral atom response speed information can reflect the agility and execution efficiency of talent in a dynamic environment. The speed influence factor calculation process can be a process of constructing a nonlinear gain function based on the response speed information to modulate the skill value. In a specific embodiment, the speed influence factor calculation process can give high-efficiency response behaviors additional weight, reflecting the value of dynamic adaptability. The behavioral atom speed-modulated skill data can be the final single-behavior skill contribution value after response speed weighting. In this embodiment, the behavioral atom speed-modulated skill data can be obtained by multiplying the angle-corrected skill allocation value by the speed influence factor. Furthermore, the behavioral atom speed-modulated skill data can integrate the dual dynamic features of semantic alignment and response efficiency, and use them as the basic unit for superposition and summation.
[0149] Based on the behavioral atom velocity modulation skill data, the effects of multiple behavioral atoms on the same task atom at the current semantic point are superimposed and summed to obtain the instantaneous competence data of the atom.
[0150] In this context, the same task atom can be a specific task atom that is covered by multiple behavioral atoms at the current semantic point. The effects of multiple behavioral atoms can be the set of modulated skill values of several behavioral atoms that simultaneously influence the same task atom at the current semantic point. Furthermore, the effects of multiple behavioral atoms can constitute an input source for superposition and summation, reflecting multi-source behavioral synergy. Superposition and summation can be a linear or nonlinear aggregation operation on the velocity-modulated skill data of multiple behavioral atoms affecting the same task atom. In an exemplary embodiment, superposition and summation can generate the comprehensive instantaneous competence of the task atom at the current semantic point. The instantaneous competence data of the atom is obtained by superposition and summing the effects of multiple behavioral atoms on the same task atom at the current semantic point based on the velocity-modulated skill data of the behavioral atoms. This can be achieved by traversing all behaviors covering the task atom at the current semantic point and accumulating their velocity-modulated skill values. Furthermore, this process can be achieved by dynamically weighting the contributions of each behavior using an attention mechanism and introducing a synergistic gain term to add extra points to complementary behavior combinations, thereby aggregating multi-source behavioral contributions, initially reflecting the multi-skill synergy effect, and generating the final instantaneous competence.
[0151] Taking the real-time problem handling assessment of online customer service specialists as an example, the talent pool modeling method based on a large language model in this embodiment can be as follows: When a customer service representative is handling the task of "order anomaly", they perform two actions in succession: "querying logistics status" (response speed 8 seconds) and "initiating a refund process" (response speed 12 seconds). The semantic phase angle of this task is "execution resolution", the operation direction of the former is "information acquisition" (cosine = 0.6), and the latter is "operation execution" (cosine = 0.92). The basic skill allocation values are 0.5 and 0.7 respectively. After angle correction, they are 0.3 and 0.644. Let the speed influence factor be log(1 + 10 / t), then the former factor = log(1 + 10 / 8) = 0.92, and the latter = log(1 + 10 / 12) = 0.87. The speed modulation skill data are 0.276 and 0.560. The summation yields that the instantaneous competence of the task atom at the current semantic point is 0.836. If another customer service representative performs the same action but responds more slowly (20 seconds), their speed modulation value is significantly reduced, reflecting the actual impact of response efficiency on competence.
[0152] In one embodiment, atomic cumulative competence index data is used to superimpose and calculate the multi-behavior atomic synergy effect to obtain atomic synergy skill mapping distribution data, including:
[0153] Based on the atomic cumulative competence index data, the neighborhood relationship of each task atom in the talent pool is established, the direct neighboring atoms of each atom are identified and the topological connection relationship between atoms is established, and the atomic neighborhood topology data is obtained.
[0154] The local skill intensity gradient of each atom is calculated based on the atomic neighborhood topology data and the atomic cumulative competence index data. The atomic skill gradient data is obtained by calculating the competence difference between the current atom and its neighboring atoms and dividing it by the atomic semantic distance.
[0155] Based on the atomic skill gradient data, the synergistic enhancement effect between multi-behavioral atoms is quantitatively calculated based on gradient smoothness to obtain atomic synergistic enhancement coefficient data.
[0156] Synergistic effect modulation calculations are performed using atomic synergistic enhancement coefficient data and atomic cumulative competence index data. The original skill allocation is optimized and modulated by multiplying the cumulative competence with the synergistic enhancement coefficient to obtain atomic synergistic modulation skill data.
[0157] By performing spatial distribution statistical processing on the overall synergistic effect of the talent pool using atomic synergistic modulation skill data, and rearranging and normalizing the synergistic modulation skills of each atom according to their semantic positions, atomic synergistic skill mapping distribution data is obtained.
[0158] The neighborhood relationship establishment process can be a process of identifying and constructing direct neighboring atomic connections based on the semantic or logical relevance of task atoms, which can be used to provide a local topological structure foundation for skill gradient calculation. In this embodiment, the neighborhood relationship establishment process can combine the execution logic information and semantic embedding similarity of task atoms to identify directly neighboring atoms. For example, the neighborhood relationship establishment process can construct adjacency relationships based on logical connection information in the task atom topological attribute data and use semantic K-nearest neighbors (K=3~5) to supplement logically missing connections, thereby achieving the technical effect of constructing a local topological structure of the task semantic space. Directly neighboring atoms can be other task atoms in the task atom topological network that have a direct logical dependency or semantic association with the current atom, which can be used to form local neighborhoods for gradient calculation and synergy effect analysis. The topological connection relationship between atoms can be a structured edge relationship established between task atoms through execution logic, shared skills, or semantic similarity, which can be used to form the skeleton of the task knowledge graph to support neighborhood propagation and gradient modeling. Atom neighborhood topological data can be structured graph data that records each task atom and its set of directly neighboring atoms, which can be used to provide local task context for calculating skill intensity gradients. Furthermore, atomic neighborhood topology data can be obtained by jointly constructing an adjacency list based on the execution logic information of task atoms and semantic embedding similarity.
[0159] The local skill intensity gradient can be seen as the rate of change in the capability distribution of a task atom within its local neighborhood, reflecting skill balance and transfer potential. It can be used as a core criterion for synergistic enhancement effects; a flatter gradient indicates stronger synergy. The competency difference between the current atom and its neighboring atoms can be seen as the numerical difference in the cumulative competency index between the central task atom and its neighboring atoms. It can be used to measure the degree of local capability imbalance and is the numerator of the gradient calculation. The atomic semantic distance, previously defined, is the distance metric between two task atoms in a unified semantic vector space. It can be used as the denominator in the gradient calculation to normalize the competency difference and eliminate scale effects. Atom skill gradient data can be the local gradient vector or scalar set obtained by dividing the competency difference by the atomic semantic distance. It can be used to quantify the steepness or flatness of the capability distribution in the task semantic space. The synergistic enhancement effect between multi-behavior atoms can be seen as the overall capability gain phenomenon generated by multiple behaviors covering complementary task atoms. It can be used to reflect the nonlinear enhancement effect of skill combinations on complex tasks. Gradient smoothness can be the inverse of the absolute value of the local skill gradient or a variance statistic. A larger value indicates a more balanced ability distribution, which can be used as the basis for calculating the synergistic enhancement coefficient, with smooth regions yielding higher gains. In a specific embodiment, gradient smoothness may include, but is not limited to, the inverse of the gradient magnitude, the neighborhood gradient variance, and the proportion of the maximum gradient threshold. Atomic synergistic enhancement coefficient data can be multiplicative gain factors generated based on gradient smoothness, used to modulate the original competence, and can be used to explicitly model the amplification effect of multi-skill synergy on individual abilities.
[0160] Synergistic effect modulation calculation can be a computational process of multiplying cumulative competence with synergistic enhancement coefficient to optimize the original skill allocation. This can be used to dynamically correct the individual ability assessment through synergistic mechanisms. The original skill allocation can be the ability value represented by the atomic cumulative competence index data before synergistic modulation, used as the baseline input for modulation calculation. Optimized modulation can be an adjustment operation that nonlinearly amplifies the original skill allocation through the synergistic enhancement coefficient, which can be used to make the ability assessment reflect the actual effectiveness gain of the skill combination. Atomic synergistic modulation skill data can be the task atomic competence value after synergistic effect modulation, integrating multi-skill synergistic gains, which can be used as the basic data for synergistic distribution statistics. In a specific embodiment, atomic synergistic modulation skill data can be input to the spatial distribution statistical processing module. The overall synergistic effect can be the global capability enhancement pattern formed by all task atoms in the talent pool after synergistic modulation, which can be used to reflect the comprehensive synergistic potential of an organization or individual in complex tasks. Spatial distribution statistical processing can be an aggregation analysis of the distribution pattern of atomic synergistic modulation skill data in the task semantic space, which can be used to reveal high synergy areas and capability gaps, supporting organizational capability building. Rearrangement and normalization can be operations that sort co-modulation skill values by semantic position and scale them to a uniform numerical range, which can be used to generate comparable, interpretable, and visualized standard output formats.
[0161] Taking the composite competency assessment of intelligent driving system engineers as an example, the talent pool modeling method based on a large language model in this embodiment can be as follows: The system analyzes that an engineer's cumulative competency in the three modules of perception, decision-making, and control is 0.82, 0.65, and 0.78, respectively. The neighborhood topology is constructed to find that "sensor fusion" and "path planning" are adjacent atoms with a semantic distance of 0.35, a competency difference of 0.17, and a gradient of 0.49. Meanwhile, "path planning" and "vehicle control" have a semantic distance of 0.28, a difference of 0.13, and a gradient of 0.46. The variance of the neighborhood gradient of each atom is calculated, revealing that the "path planning" region is the smoothest, with a collaborative enhancement coefficient of 1.35. After modulation, its competency increases from 0.65 to 0.878. Finally, the data is arranged and normalized according to semantic position (perception, decision-making, control) to generate atomic collaborative skill mapping distribution data, showing that the engineer has a significant enhancement effect in cross-module collaborative tasks, making them particularly suitable for system integration positions.
[0162] In one embodiment, the talent pool structure is evaluated and optimized based on talent characteristic data to obtain refined digital model data of the talent pool, including:
[0163] Based on talent characteristic data, statistical analysis is performed on the skill allocation distribution of the talent pool to calculate the mean, variance, and standard deviation of the global skill allocation, thereby obtaining statistical characteristic data of skill allocation.
[0164] The balance deviation of each region of the talent pool is quantitatively assessed by using statistical characteristic data of skill allocation. The relative deviation between the allocation of atomic skills for each task and the global mean is calculated and a histogram of deviation distribution is established to obtain balance deviation assessment data.
[0165] Based on the balance deviation assessment data, areas with substandard talent pool structure are identified and marked. By setting a deviation threshold range and filtering out task atom sets that exceed the threshold, abnormal area marking data of talent profile is obtained.
[0166] The behavior atomic parameters were adjusted and the talent pool feature mapping effect was optimized based on the abnormal area marking data of the talent profile. The convergence of the optimization effect was verified, thereby obtaining refined digital model data of the talent pool.
[0167] The skill allocation distribution can be seen as the numerical distribution of all task atoms in the talent pool across the dimensions of cumulative competence or collaborative modulation skills. This can reflect the resource allocation pattern of an organization or group across various capabilities. Statistical analysis can be a mathematical process of measuring the central tendency and dispersion of the skill allocation distribution. This can be used to establish a global capability benchmark and support deviation diagnosis. Furthermore, statistical analysis of the skill allocation distribution in the talent pool based on talent characteristic data can involve extracting the skill values of all task atoms and calculating their first and second-order statistics. For example, this operation can be achieved by using streaming statistical algorithms to process large-scale data, calculating independently for each job cluster, and then weighted and summarizing, thereby establishing an objective, data-driven capability benchmark reference system. The mean of the global skill allocation can be the arithmetic mean of the skill allocation values of all task atoms in the talent pool, which can be used as a central reference point to measure balance. The variance can be the squared average of the deviations of each task atom's skill allocation value from the mean, reflecting the overall dispersion and quantifying the degree of imbalance in the capability distribution of the talent pool. Standard deviation can be the square root of variance, which has the same dimension as the original skill value, making it easy to interpret. It can be used to provide an interpretable range of fluctuation indicators and assist in threshold setting.
[0168] Skill allocation statistical characteristic data can be a structured dataset containing statistics such as mean, variance, and standard deviation, which can be used to provide benchmark parameters for assessing balance deviation. In a specific embodiment, skill allocation statistical characteristic data can be calculated by traversing all task atom skill values in the talent characteristic data. Calculating the mean, variance, and standard deviation of the global skill allocation to obtain the skill allocation statistical characteristic data can be done by applying standard statistical formulas: the mean equals the sum of all skill values divided by the total number of task atoms; the variance equals the sum of the squares of the differences between each skill value and the mean divided by the total number; and the standard deviation equals the square root of the variance. This allows for the quantification of the centrality and dispersion characteristics of the overall ability distribution of the talent pool. The balance deviation in different regions of the talent pool can be a systematic deviation between the skill allocation of local task atom sets and the global mean, which can be used to identify structurally problematic areas of ability overload or deficiency.
[0169] Quantitative assessment can be the process of converting balance deviations into numerical indicators, enabling a shift from qualitative judgment to quantitative diagnosis. Task atom skill allocation can be the cumulative competence or collaborative modulation skill value corresponding to each task atom in talent characteristic data, serving as the basic unit for deviation calculation. The global mean can be the mean term in the skill allocation statistical characteristic data, serving as a benchmark for relative deviation calculation. Relative deviation can be the difference between a single task atom skill allocation value and the global mean divided by the standard deviation or a normalized indicator of the mean, used to eliminate the influence of dimensions and support cross-task comparisons. A deviation distribution histogram can be a statistical chart with relative deviation on the horizontal axis and the number of task atoms on the vertical axis, used to visualize the overall balance of the talent pool and identify anomalous tails. Balance deviation assessment data can be a structured assessment result containing the relative deviations of each task atom and histogram statistics, used to accurately locate areas of capability imbalance. In an exemplary embodiment, balance deviation assessment data can be calculated and generated based on skill allocation statistical characteristic data.
[0170] Quantitatively assessing the imbalance deviation in different regions of the talent pool using statistical characteristic data on skill allocation can be achieved by calculating the deviation of the skill value from the mean for each task atom, thus transforming the abstract imbalance problem into a calculable indicator. Calculating the relative deviation between the skill allocation of each task atom and the global mean, and constructing a histogram of the deviation distribution, yields the imbalance deviation assessment data. This can be achieved by calculating the difference between the skill value and the mean of each task atom, divided by the standard deviation, as the relative deviation, and generating a histogram based on interval statistical frequency. For example, this operation can be achieved by using kernel density estimation instead of a histogram and dynamically adjusting the number of bins to adapt to the distribution pattern, thereby visualizing and structurally recording the balance status.
[0171] Regions with substandard talent pool structures can be clusters of task atoms whose skill allocation significantly deviates from the global equilibrium level, and can be used as targets for optimization intervention. Identification markers can be the result of attaching anomaly labels to substandard task atoms, and can be used to structurally record areas requiring optimization. Deviation threshold ranges can be preset upper and lower limits of relative deviation, used to define normal and abnormal regions, and can be used to control optimization sensitivity and avoid over-adjusting noise points. Sets of task atoms exceeding the threshold can be a list of task atom IDs whose absolute relative deviation value is greater than the upper threshold limit, and can be used to constitute the core members of abnormal regions in the talent profile.
[0172] The talent profile anomaly area labeling data can be a structured set of labels recording all anomalous task atoms and their degree of deviation, which can be used to guide the direction and scope of subsequent parameter adjustments. In an exemplary embodiment, the talent profile anomaly area labeling data can be generated by filtering out items exceeding the threshold in the balance deviation assessment data. Identifying and labeling non-compliant areas of the talent pool structure based on the balance deviation assessment data can be done by setting a deviation threshold (e.g., the absolute value of the relative deviation is greater than 1.5), labeling task atoms that exceed the limit, thereby automatically locating the anomalous areas of capabilities that need optimization. By setting a deviation threshold range and filtering out the set of task atoms that exceed the threshold, the talent profile anomaly area labeling data can be obtained through execution condition filtering: selecting task atom identifiers with a relative deviation absolute value greater than the threshold, thereby outputting a structured list of anomalous areas. Behavioral atom parameters can be key parameters defined in previous schemes that affect behavioral modeling, such as logical weights, response speed factors, phase angles, etc., which can be used as direct objects for optimization and parameter tuning to achieve underlying model calibration. The talent pool feature mapping effect can be the accuracy of the talent feature data in representing real capabilities after the behavioral atom parameters are adjusted, which can be used as an optimization target and needs to be improved iteratively.
[0173] Adjusting behavioral atomic parameters and optimizing talent pool feature mapping based on talent profile anomaly region marker data can involve adjusting the logical weights, velocity factors, or phase angle parameters of the behaviors associated with anomalous task atoms via callbacks. Furthermore, this operation can be achieved by fine-tuning parameters using gradient backpropagation and triggering preset tuning strategies based on a rule engine, thus enabling targeted calibration of underlying modeling parameters. The convergence of the optimization effect can be achieved by ensuring the equilibrium deviation index stabilizes during continuous iterations, guaranteeing overfitting and reliable results. Verifying the convergence of the optimization effect to obtain refined talent pool digital model data can involve repeated evaluation-parameter tuning loops until the deviation index changes less than a preset tolerance in two consecutive rounds. For example, this operation can be achieved by monitoring the decrease curve of the number of anomalous atoms and employing an early stopping mechanism to prevent overfitting, ensuring stable and reliable optimization results and outputting the final refined model.
[0174] For example, in the scenario of quarterly optimization of the R&D talent pool in an internet company, the talent pool modeling method based on a large language model in this embodiment can be as follows: The system analyzes the collaborative modulation skill values of 2000 task atoms in the current talent pool, calculates the global mean of 0.62 and the standard deviation of 0.18; it finds that the skill value of the "large model fine-tuning" task atom reaches 0.95, with a relative deviation z = (0.95-0.62) / 0.18≈1.83, exceeding the threshold of 1.5, and is marked as an abnormally high-load area; while the atom value of "traditional service operation and maintenance" is only 0.31 (z = -1.72), and is marked as a capability depression; the system automatically calls back the relevant behavioral atom parameters, reduces the logical weight of high-frequency fine-tuning behavior, and improves the response speed factor of operation and maintenance behavior; after two rounds of iteration, the z values of the two regions are reduced to 1.2 and -1.1 respectively, the deviation variance decreases by 23%, and the convergence verification is passed; the final output refined talent pool digital model more evenly reflects the organization's real capability distribution and supports the precise allocation of talent resources during the AI transformation period.
[0175] In addition, refer to Figure 2 To achieve the above objectives, the present invention also provides a talent pool modeling system based on a large language model, the system comprising:
[0176] The multi-source acquisition module 10 is used to acquire the candidate's historical work text data, behavioral trajectory data, and real-time interactive feedback data.
[0177] The task discretization module 20 is used to discretize the talent workflow based on historical work text data to obtain task atomic data containing task node attribute information and execution logic information.
[0178] The behavior modeling module 30 is used to dynamically model talent behavior patterns based on real-time interactive feedback data to obtain behavioral atom semantic sequence data. Specifically, the behavioral atom changes are achieved by setting attention weight combinations with different logical dependencies between behavioral atoms to perform association operations of multiple behaviors in the semantic domain and maximum effectiveness coverage in the cognitive domain.
[0179] The dynamic matching module 40 is used to perform dynamic matching degree calculation on talent competence and skill distribution based on task atomic data and behavior atomic semantic sequence data, and obtain talent characteristic data including cumulative competence index and collaborative skill mapping distribution. Specifically, the dynamic matching degree calculation process is to predict the competence of each task node through job requirement constraints and skill distribution characteristics.
[0180] The talent pool structure optimization module 50 is used to evaluate and optimize the talent pool structure based on talent characteristic data to obtain refined talent pool digital model data.
[0181] Other embodiments or specific implementations of the talent pool modeling system based on a large language model described in this invention can be referred to the above-described method embodiments, and will not be repeated here.
[0182] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A talent pool modeling method based on a large language model, characterized in that, The method includes: Acquire candidate candidates' historical work text data, behavioral trajectory data, and real-time interactive feedback data; The talent workflow is discretized based on historical work text data to obtain task atomic data containing task node attribute information and execution logic information; Based on real-time interactive feedback data, dynamic modeling of talent behavior patterns is performed to obtain behavioral atom semantic sequence data. Specifically, the changes in behavioral atoms are achieved by setting attention weight combinations with different logical dependencies between behavioral atoms to perform association operations of multiple behaviors in the semantic domain and maximum effectiveness coverage in the cognitive domain. Based on task atomic data and behavioral atomic semantic sequence data, dynamic matching degree calculation is performed on talent competency and skill distribution to obtain talent characteristic data including cumulative competency index and collaborative skill mapping distribution. Specifically, the dynamic matching degree calculation is to predict the competency of each task node through job requirement constraints and skill distribution characteristics. The talent pool structure is evaluated and optimized based on talent characteristic data to obtain a refined digital model of the talent pool.
2. The talent pool modeling method based on a large language model as described in claim 1, characterized in that, The process of discretizing the talent workflow based on historical work text data yields task atomic data containing task node attribute information and execution logic information, including: The talent workflow is divided into tasks based on historical work text data to obtain task data, where the task density is set to no less than 5 behavior atoms per thousand words of text. Based on the task data, attribute values are assigned and logical relationships are calculated for each task atom to obtain atomic topology attribute data containing atomic attribute information and logical connection information. Based on the atomic topology attribute data, unique semantic identifiers are assigned to the task atoms to obtain task atom data with index identifiers.
3. The talent pool modeling method based on a large language model as described in claim 2, characterized in that, The process of dynamically modeling talent behavior patterns based on real-time interactive feedback data to obtain behavioral atomic semantic sequence data includes: Extract the operation response latency, contextual relevance, semantic position, and logical evolution direction of each behavioral atom from real-time interactive feedback data, and establish a structured parameter table for each behavioral atom; The logical weight value of each behavior is determined by the total number of atoms in the parameter table, and a logical configuration list is generated by assigning different logical starting points to each atom by equally dividing the complete semantic cycle. Based on the parameter table and the logical configuration list, generate a set of behavioral atom sequences containing personalized logical dependency descriptions for each behavioral atom; Based on the set of behavioral atom sequences, the entire interaction cycle is sampled according to a preset semantic interval, and the logical connection state of all behavioral atoms at each semantic point is calculated, thereby forming a behavioral state sequence table arranged in semantic order. By using a behavior state sequence table, a fast query correspondence between timestamps and behavioral semantic information is established, thereby obtaining behavioral atomic semantic sequence data.
4. The talent pool modeling method based on a large language model as described in claim 3, characterized in that, The process of dynamically calculating the matching degree of talent competency and skill distribution based on task atomic data and behavioral atomic semantic sequence data yields talent characteristic data including cumulative competency index and collaborative skill mapping distribution, including: The connectivity calculation of behavior and task atoms is performed based on the semantic sequence data of behavior atoms and the data of task atoms. By calculating the semantic association distance between each behavior atom at each semantic point and each task atom in the talent pool, the behavior-task connectivity matrix data is obtained. Based on the behavioral task connectivity matrix data and the preset job competency threshold, the effective skill range is determined. By comparing connectivity, task atoms that are within the effective skill range of each behavioral atom are selected, and behavioral skill coverage relationship data is obtained. Based on the behavioral skill coverage relationship data, the instantaneous competence value of each task atom is calculated using a skill allocation distribution model based on an exponential decay function, thus obtaining the instantaneous competence data of the atom. Based on the instantaneous competence data of atoms, a semantic sequence-based cumulative calculation process is performed to obtain the cumulative competence index data of atoms. Specifically, the cumulative calculation process is performed by semantic integration on the competence of each task atom during the entire interaction cycle. The atomic cumulative competence index data is used to superimpose and calculate the effect of multi-behavior atomic synergy to obtain atomic synergy skill mapping distribution data; Based on the atomic cumulative competence index data and the atomic collaborative skill mapping distribution data, the cumulative competence information and the collaborative skill mapping distribution information are matched according to the task atomic index identifier to obtain talent characteristic data.
5. The talent pool modeling method based on a large language model as described in claim 4, characterized in that, The process of determining the effective skill range based on the behavioral task connectivity matrix data and a preset job competency threshold, and filtering out task atoms within the effective skill range of each behavioral atom through connectivity comparison, yields behavioral skill coverage relationship data, including: The correlation determination results are obtained by comparing the behavioral task connectivity matrix data and the preset job competency thresholds one by one. Based on the correlation determination results, behavioral task pairs that meet the skill range conditions are screened and extracted. By retaining behavioral task combinations with correlation values less than or equal to the competence threshold, a list of effective behavioral task pairs is obtained. Based on the list of valid behavioral task pairs, the skill task set of each behavioral atom is grouped according to the atom identifier to obtain behavioral group atom set data; Based on the behavior grouping atom set data, the skill task set of each behavior atom is traversed and the number of covered behaviors of each atom is counted to obtain the list of covered behaviors of the atom. By constructing a bidirectional index mapping table from behavior to atom and from atom to behavior using behavior grouping atom set data and atom covered behavior list data, behavior skill coverage relationship data can be obtained.
6. The talent pool modeling method based on a large language model as described in claim 5, characterized in that, The instantaneous competency value of each task atom is calculated based on a skill allocation distribution model with an exponential decay function, using behavioral skill coverage relationship data, to obtain atom instantaneous competency data, including: Extract the semantic association distance information between each behavioral atom and the task atoms within its skill range from the behavioral skill coverage relationship data to obtain an effective semantic distance value data set; Normalized distance ratio data is obtained by using a set of effective semantic distance values and a preset exponential decay parameter to perform normalized distance calculation. Based on the normalized semantic distance ratio data, the normalized semantic distance ratio is multiplied by a negative exponential decay coefficient and the exponential function is taken to obtain the exponential decay factor data. The skill value is calculated based on the exponential decay factor data and the preset maximum skill capacity parameter of the atom to obtain the basic skill allocation value of the task atom. Based on the assigned values of basic skills of the task atoms, the data is processed by correcting and solving the results based on the direction of behavior atomic operations and the phase angle of task semantics, thereby obtaining the instantaneous competence data of atoms.
7. The talent pool modeling method based on a large language model as described in claim 6, characterized in that, The process of correcting and solving the data based on the behavioral atomic operation direction and task semantic phase angle according to the assigned numerical values of the task's basic skills, thereby obtaining the atomic instantaneous competence data, includes: Based on the basic skill allocation values of task atoms, the phase difference between behavioral atoms and task semantics is calculated and the angle correction is performed. The cosine value of the phase angle between the operation direction of the behavioral atom and the task semantics is calculated and used as the correction coefficient to obtain the skill allocation value after angle correction. Based on the skill allocation value and the behavioral atom response speed information of the current semantic point, the speed influence factor is calculated and processed to obtain behavioral atom speed modulation skill data; Based on the behavioral atom velocity modulation skill data, the effects of multiple behavioral atoms on the same task atom at the current semantic point are superimposed and summed to obtain the instantaneous competence data of the atom.
8. The talent pool modeling method based on a large language model as described in claim 7, characterized in that, The process of superimposing and calculating the multi-behavior atomic synergy effect using atomic cumulative competence index data to obtain atomic synergy skill mapping distribution data includes: Based on the atomic cumulative competence index data, the neighborhood relationship of each task atom in the talent pool is established, the direct neighboring atoms of each atom are identified and the topological connection relationship between atoms is established, and the atomic neighborhood topology data is obtained. The local skill intensity gradient of each atom is calculated based on the atomic neighborhood topology data and the atomic cumulative competence index data. The atomic skill gradient data is obtained by calculating the competence difference between the current atom and its neighboring atoms and dividing it by the atomic semantic distance. Based on the atomic skill gradient data, the synergistic enhancement effect between multi-behavioral atoms is quantitatively calculated based on gradient smoothness to obtain atomic synergistic enhancement coefficient data. Synergistic effect modulation calculations are performed using atomic synergistic enhancement coefficient data and atomic cumulative competence index data. The original skill allocation is optimized and modulated by multiplying the cumulative competence with the synergistic enhancement coefficient to obtain atomic synergistic modulation skill data. By performing spatial distribution statistical processing on the overall synergistic effect of the talent pool using atomic synergistic modulation skill data, and rearranging and normalizing the synergistic modulation skills of each atom according to their semantic positions, atomic synergistic skill mapping distribution data is obtained.
9. The talent pool modeling method based on a large language model as described in claim 8, characterized in that, The process of evaluating and optimizing the talent pool structure based on talent characteristic data to obtain refined digital model data for the talent pool includes: Based on talent characteristic data, statistical analysis is performed on the skill allocation distribution of the talent pool to calculate the mean, variance, and standard deviation of the global skill allocation, thereby obtaining statistical characteristic data of skill allocation. The balance deviation of each region of the talent pool is quantitatively assessed by using statistical characteristic data of skill allocation. The relative deviation between the allocation of atomic skills for each task and the global mean is calculated and a histogram of deviation distribution is established to obtain balance deviation assessment data. Based on the balance deviation assessment data, areas with substandard talent pool structure are identified and marked. By setting a deviation threshold range and filtering out task atom sets that exceed the threshold, abnormal area marking data of talent profile is obtained. The behavior atomic parameters were adjusted and the talent pool feature mapping effect was optimized based on the abnormal area marking data of the talent profile. The convergence of the optimization effect was verified, thereby obtaining refined digital model data of the talent pool.
10. A talent pool modeling system based on a large language model, characterized in that, The system includes: The multi-source acquisition module is used to acquire candidates' historical work text data, behavioral trajectory data, and real-time interactive feedback data. The task discretization module is used to discretize the talent workflow based on historical work text data to obtain task atomic data containing task node attribute information and execution logic information. The behavior modeling module is used to dynamically model talent behavior patterns based on real-time interactive feedback data, and obtain behavioral atom semantic sequence data. Specifically, the changes in behavioral atoms are achieved by setting attention weight combinations with different logical dependencies between behavioral atoms to perform semantic domain association operations and maximum efficiency coverage in the cognitive domain. The dynamic matching module is used to perform dynamic matching degree calculation on talent competency and skill distribution based on task atomic data and behavior atomic semantic sequence data, and obtain talent characteristic data including cumulative competency index and collaborative skill mapping distribution. Specifically, the dynamic matching degree calculation process is to predict the competency of each task node through job requirement constraints and skill distribution characteristics. The talent pool structure optimization module is used to evaluate and optimize the talent pool structure based on talent characteristic data, resulting in refined digital model data of the talent pool.