Human resource information file management system based on human resource demand analysis
Patent Information
- Application Number
- CN202611006764.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-29
AI Technical Summary
该方式难以在全局层面同时平衡多个岗位的子需求,当一个员工的技能满足多个岗位的部分要求时,仅依靠局部匹配指标容易造成能力资源浪费或关键岗位的能力覆盖度不足
通过构建需求-能力异构图并采用图卷积网络进行分层聚合,将每个岗位的实时需求向量作为需求节点、每个员工的实时能力向量作为能力节点,在需求节点与能力节点之间依据岗位所需技能和员工所持技能的余弦匹配度建立边连接,保留匹配度高于预设阈值的边;同时在需求节点之间依据岗位间业务协作频率建立加权无向边,在能力节点之间依据员工间合作紧密系数和发起方任务贡献占比建立有向边。以此构建的异构图内有三种关系同时作用,使孤立的需求向量和能力向量被组织进一个统一的交互结构中。图卷积网络对需求节点聚合其邻域能力节点特征时,通过基于实时需求向量与能力向量的内积计算注意力系数并归一化加权,使聚合特征侧重于与该需求高度匹配的能力信息;再将聚合向量与需求节点自身向量拼接后经非线性变换,得到需求侧聚合特征;对能力节点同样采用需求邻域加权聚合得到供给侧聚合特征。这种聚合方式将跨类型节点的匹配强度、同类型节点间的业务关系与协作历史均编码进节点特征表示之中,相比直接将原始需求与能力向量做拼接或简单统计特征,大幅增强了特征对组织内部人才供需结构和动态依赖的刻画能力,使得后续时序模型能够感知到需求变化在协作网络中的传导效应以及能力侧组合潜力,最终输出的预测需求向量在多岗位协同情景下具有更高的合理性和一致性。
Smart Images

Figure CN122840531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human resource information management technology, specifically a human resource information file management system based on human resource demand analysis. Background Technology
[0002] Managing organizational human resource information files based on human resource needs requires the ability to accurately predict future changes in job requirements and dynamically adjust employee groupings. Existing human resource information file management systems typically process needs analysis and file grouping as two separate processes.
[0003] In demand forecasting, existing technologies mostly rely on single time-series indicators such as the number of historical job vacancies and turnover rates, employing statistical models like autoregressive moving averages or simple machine learning regression methods for inference. These methods treat job requirements and employee capabilities as isolated time-series variables, neglecting inter-job collaboration, the degree of matching between employee capabilities and job requirements, and the impact of implicit capability networks formed through past collaborations on demand evolution. Due to the lack of modeling these rich interactions, the forecast results lag in responding to structural changes within organizations and the dynamic skill set requirements under project-based operations. The forecast vectors are prone to inconsistencies between dimensions and deviations from actual staffing needs.
[0004] In terms of file grouping, rule-based or locally greedy strategies are typically used for grouping. For example, employees are assigned in descending order of job matching degree until each position meets the basic number of employees required. This approach makes it difficult to balance the sub-requirements of multiple positions at a global level. When an employee's skills meet some of the requirements of multiple positions, relying solely on local matching indicators can easily lead to a waste of capability resources or insufficient capability coverage for key positions.
[0005] Therefore, when faced with the complex heterogeneous dependency between human resource demand and capacity supply, the system needs to address how to generate features that fully represent the interaction between the demand and supply sides to improve prediction accuracy, and how to achieve optimal configuration of global capacity coverage for file grouping under multidimensional sub-demand constraints. Summary of the Invention
[0006] The present invention aims to provide a human resources information archive management system based on human resources demand analysis, which can extract temporal features from the heterogeneous interaction between demand and capability and perform global optimization of archive grouping.
[0007] The objective of this invention can be achieved through the following technical solutions: This invention provides a human resource information file management system based on human resource demand analysis. The system includes a demand acquisition module, a graph feature aggregation module, a time-series prediction module, a dynamic grouping module, and a blockchain storage module. The demand acquisition module acquires real-time demand vectors for each position within the organization and real-time capability vectors from employee personnel files. The graph feature aggregation module constructs requirements based on the real-time demand vectors and real-time capability vectors. The system generates demand-side and supply-side features by aggregating node features in a heterogeneous capability graph using a graph convolutional network. The time-series prediction module inputs these features into a pre-defined time-series prediction model and outputs a predicted demand vector for the next time window. The dynamic grouping module regroups employee personnel files using dynamic programming based on the predicted demand vector and real-time capability vector, generating target file grouping results. The blockchain storage module writes the target file grouping results to a blockchain storage node and updates the personnel information file index table. Through the collaborative work of these modules, the system can deeply characterize and predict human resource needs and capabilities, and achieve dynamic and traceable file reorganization, improving the foresight and accuracy of organizational human resource allocation.
[0008] As a preferred technical solution of the present invention, the graph feature aggregation module meets the construction requirements. In the heterogeneous capability graph, each job's real-time demand vector is treated as a demand node, and each employee's real-time capability vector as a capability node. Edges are established between demand nodes and capability nodes based on the matching degree between the skills required for the job and the skills held by the employee. This matching degree is calculated using cosine similarity, and only edges with a matching degree greater than a preset similarity threshold are retained; edges below the threshold are discarded as noise edges. This makes the graph structure more focused on strong relationships and reduces redundant information interference. Between demand nodes, edges are established based on the frequency of business collaboration between jobs. Specifically, the number of times each pair of jobs jointly participated in the same project is extracted from the organization's historical project allocation records, and this number is divided by the total number of projects to obtain the business collaboration frequency. When this frequency exceeds a set first frequency threshold, an undirected edge is established between the corresponding two demand nodes, and the edge weight is assigned the business collaboration frequency. This integrates collaborative dependency information between jobs into the demand side. Between capability nodes, edge connections are established based on the project collaboration history between employees. The number of tasks jointly completed by each pair of employees is extracted from the historical project collaboration records and divided by the geometric mean of the total number of tasks each employee participated in to obtain the collaboration tightness coefficient between employees. When this coefficient is greater than the set second coefficient threshold, a directed edge is established between the corresponding two capability nodes, with the direction from the collaboration initiator to the collaboration recipient, and the edge weight is the collaboration tightness coefficient multiplied by the initiator's task contribution ratio, thereby characterizing the tightness and directionality of employee collaboration on the supply side.
[0009] Furthermore, the graph feature aggregation module uses hierarchical aggregation rules in graph convolutional networks to generate demand-side and supply-side features. For each demand node, the real-time capability vectors of all its neighboring capability nodes are collected. The attention coefficient of each neighboring capability node is calculated based on the inner product of its real-time demand vector and the real-time capability vector of the capability node. After normalization, the normalized attention weights are used to weightedly sum the real-time capability vectors of the neighboring capability nodes to obtain the neighboring aggregation vector of the demand node. The real-time demand vector of the demand node and the neighboring aggregation vector are concatenated and passed through a nonlinear transformation layer to obtain the demand-side aggregated feature. For each capability node, the real-time demand vectors of all its neighboring demand nodes are collected. A second attention coefficient is calculated based on the dot product of the capability node's real-time capability vector and the real-time demand vector of the demand node. After normalization, the real-time demand vectors of the neighboring demand nodes are weightedly summed to obtain the second neighboring aggregation vector of the capability node. The real-time capability vector of the capability node and the second neighboring aggregation vector are concatenated and passed through another nonlinear transformation layer to obtain the supply-side aggregated feature. Finally, the aggregated features of the demand side are mapped to demand-side features through a fully connected layer, and the aggregated features of the supply side are mapped to supply-side features through another fully connected layer. This attention-weighted hierarchical aggregation method can adaptively capture the matching strength between demand and capability, as well as the correlation influence between nodes of the same type, making the generated features more expressive.
[0010] As a preferred embodiment of the present invention, the time-series prediction module arranges the demand-side and supply-side features of multiple consecutive historical time windows in chronological order to form a demand-side time series and a supply-side time series. The demand-side and supply-side time series are then input into two independent long short-term memory (LSTM) network encoders to obtain a demand-side hidden state sequence and a supply-side hidden state sequence, respectively. The two hidden state sequences are concatenated at each time step to obtain a fused hidden state sequence. This fused hidden state sequence is then input into an attention mechanism layer to calculate the attention weights for each historical time step, and the fused hidden state sequence is weighted and summed to obtain a context vector. Finally, the context vector is input into a decoder employing a gated recurrent unit (GRU) structure to progressively generate the predicted demand vector for the next time window. Specifically, the attention mechanism layer initializes a learnable query vector, calculates the dot product of the query vector and the fused hidden state at each time step to obtain a raw score, normalizes all raw scores to obtain the attention weights for each time step, and uses these weights to weight and sum the fused hidden states to form the context vector. By employing a dual-path encoding and attention decoding structure, it is possible to fully capture the temporal evolution patterns of both the demand and supply sides, as well as their interactive dynamics on the timeline, thereby effectively improving the accuracy of human resource demand forecasting.
[0011] When generating target file grouping results based on the predicted demand vector and real-time capability vector, the dynamic grouping module treats each employee personnel file as a file item, with each file item containing a real-time capability vector and an employee identifier. The predicted demand vector is decomposed into multiple sub-demand vectors, each corresponding to a target capability combination for a file group. A state table is initialized for dynamic programming search. This state table is designed as a three-dimensional structure: the first dimension is the file item index order, the second dimension is the satisfaction level vector of each sub-demand vector, and the third dimension is the maximum capability coverage when reaching the current state. For the first row, the satisfaction level vector is initialized to an all-zero vector, and the maximum capability coverage is zero. Each file item is traversed sequentially. For the current file item, all recorded states in the previous row are traversed, attempting to superimpose the real-time capability vector of the current file item onto the satisfaction level vector of the state, generating a new candidate satisfaction level vector. Each dimension of this new candidate satisfaction level vector is truncated to ensure that each dimension does not exceed the capability threshold of the corresponding sub-demand vector. The superimposed capability coverage is calculated, which is the sum of the ratios of each dimension value of the new candidate satisfaction level vector to the corresponding sub-demand vector value. If the new candidate fulfillment level vector is not yet recorded in the status table of the current row, or if the capability coverage is greater than the recorded maximum capability coverage, then the maximum capability coverage at the corresponding position in the status table is updated and the selection path is recorded. After traversing all file items, the target state that makes the fulfillment level of all sub-requirement vectors reach the threshold is searched in the status table, and the file item selection path corresponding to the target state is traced back. Based on the traced file item selection path, the selected human file items are classified according to their corresponding sub-requirement vectors, generating the target file grouping results. Preferably, when the same file item simultaneously meets the requirements of multiple sub-requirement vectors, it is uniquely assigned according to the belonging weight of the file item's historical grouping record to avoid personnel allocation conflicts. This dynamic grouping module can achieve optimal matching between file items and requirements while meeting multi-dimensional and multi-grouping requirements, maximizing overall capability coverage.
[0012] When writing the target file grouping results to the blockchain storage node, the blockchain notarization module serializes the target file grouping results into binary data blocks and calculates the hash value of these binary data blocks. It selects the current leader node responsible for record-keeping from the blockchain network, sends the binary data block and hash value to this leader node, and receives transaction confirmation information from the leader node, which includes the block height and block timestamp. Based on the block height and block timestamp, a record is added to the human resources file index table, including the file group identifier, a list of employee identifiers, the block height, and the block timestamp. The updated human resources file index table is then synchronously stored in the local cache and the distributed database. Preferably, the human resources file index table uses a skip list structure, with employee identifiers as keys and block heights and block timestamps as values, supporting fast querying of file grouping history by time range. The combination of blockchain notarization and skip list indexing ensures the immutability and efficient traceability of file grouping results, providing a reliable data trust foundation for human resources information management.
[0013] The beneficial effects of this invention are: By constructing a heterogeneous demand-capability graph and employing a graph convolutional network for hierarchical aggregation, the real-time demand vector of each position is used as a demand node, and the real-time capability vector of each employee is used as a capability node. Edges are established between demand nodes and capability nodes based on the cosine similarity between the skills required for the position and the skills held by the employee, retaining edges with a similarity higher than a preset threshold. Simultaneously, weighted undirected edges are established between demand nodes based on the frequency of business collaboration between positions, and directed edges are established between capability nodes based on the degree of cooperation between employees and the contribution ratio of the initiator's task. This heterogeneous graph incorporates three relationships simultaneously, organizing isolated demand and capability vectors into a unified interactive structure. When aggregating the features of neighboring capability nodes for a demand node, the graph convolutional network calculates an attention coefficient based on the inner product of the real-time demand vector and the capability vector, and then normalizes and weights the coefficients, ensuring that the aggregated features emphasize capability information highly matched to the demand. The aggregated vector is then concatenated with the demand node's own vector and subjected to a nonlinear transformation to obtain the demand-side aggregated features. Similarly, demand-neighborhood weighted aggregation is used for capability nodes to obtain the supply-side aggregated features. This aggregation method encodes the matching strength of cross-type nodes, the business relationships and collaboration history between nodes of the same type into the node feature representation. Compared with directly concatenating the original demand and capability vectors or simple statistical features, it greatly enhances the feature's ability to characterize the talent supply and demand structure and dynamic dependence within the organization. This enables subsequent time series models to perceive the transmission effect of demand changes in the collaboration network and the potential for capability combination. The final output predicted demand vector has higher rationality and consistency in multi-position collaboration scenarios.
[0014] During dynamic grouping, the predicted demand vector is decomposed into multiple sub-demand vectors. Each sub-demand vector corresponds to a target capability combination for a file group. A three-dimensional state table is initialized, with each dimension representing the file item index order, the satisfaction level vector of each sub-demand vector, and the maximum capability coverage when reaching the current state. When sequentially traversing each employee file item, its real-time capability vector is attempted to be superimposed onto the satisfaction level vector of the previous row to generate a new candidate satisfaction level vector. Dimensions exceeding the sub-demand capability threshold are truncated, and the sum of the satisfaction level ratios of each dimension is calculated as the capability coverage. The state table is updated and the selection path is recorded only when a candidate vector is not recorded or the capability coverage is greater. After traversal, the file item selection path that achieves the highest capability coverage and satisfies all sub-demands is obtained through backtracking. These paths are then categorized according to their corresponding sub-demands, and when multiple sub-demands are satisfied by the same file item, they are uniquely assigned based on historical affiliation weights. This dynamic programming method effectively escapes local optima from an exponentially possible combination of employee allocations by using state pruning and maximum coverage selection. It obtains the grouping results that meet the capability requirements of multiple grouping objectives from a global perspective and have the most sufficient capability coverage. This avoids the imbalance problem caused by the traditional successive allocation, which results in the over-accumulation of capabilities in some positions and the lack of capabilities in other positions. Attached Figure Description
[0015] The invention will now be further described with reference to the accompanying drawings.
[0016] Figure 1 This is a schematic diagram of the structure of a human resources information file management system based on human resources demand analysis; Figure 2 This is a flowchart of the demand-capability heterogeneous graph attention aggregation network processing flow. Figure 3 This is a schematic diagram of the structure of a dual encoder-attention-decoder temporal prediction model; Figure 4 This is a flowchart of the blockchain notarization and skip list index update process for file grouping results. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] See Figure 1This invention provides a human resource information file management system based on human resource demand analysis, including a demand acquisition module, a graph feature aggregation module, a time series prediction module, a dynamic grouping module, and a blockchain notarization module. The demand acquisition module acquires real-time demand vectors for each position within the organization and real-time capability vectors from employee personnel files. The demand acquisition module interfaces with the organization's internal human resource management system, enterprise resource planning system, and project management system, automatically acquiring real-time demand data for each position and real-time capability data from employee personnel files at preset intervals (e.g., every morning), and converting it into a structured vector representation.
[0019] Specifically, for generating real-time demand vectors, the demand acquisition module calls the job creation interface of the HRMS system to retrieve the number of employees on duty, the number of vacancies, and the expected turnover rate within a future time window for each job, forming job manpower gap data. Simultaneously, it calls the project management system interface to retrieve the skill types and skill level requirements for each job in future project plans. This data is aligned and normalized according to the system's pre-configured skill dimension dictionary, generating a real-time demand vector Ri with each dimension taking values in the range [0,1]. Each dimension corresponds to a score indicating the urgency of a business skill requirement. For generating real-time capability vectors, the demand acquisition module accesses the distributed database of employee personnel files, reading each employee's skill certification records, historical project completion quality scores, and supervisor evaluation levels. Through a pre-trained skill-capability mapping model, unstructured text evaluations and certificate levels are mapped to standardized capability scores. A weighted average of the multi-source capability data for the same employee is then used to generate a real-time capability vector Aj with each dimension taking values in the range [0,1]. The dimension definitions are completely consistent with the real-time demand vectors, and the dimension order is uniformly managed by the system's global configuration table. During the vector generation process, the demand acquisition module also records the data collection timestamp and data source version number simultaneously, and writes the generated real-time demand vector and real-time capability vector into the local memory cache, which is then read in batches by the image feature aggregation module according to the time window.
[0020] The graph feature aggregation module constructs a demand-capacity heterogeneous graph based on real-time demand and capability vectors, and uses a graph convolutional network to aggregate node features from the heterogeneous graph, generating demand-side and supply-side features. The time-series prediction module inputs the demand-side and supply-side features into a preset time-series prediction model and outputs the predicted demand vector for the next time window. The dynamic grouping module regroups employee personnel files using dynamic programming based on the predicted demand and real-time capability vectors, generating target file grouping results. The blockchain evidence storage module writes the target file grouping results to the blockchain storage node and updates the personnel information file index table.
[0021] In practice, the process of constructing the demand-capability heterogeneous graph includes two stages: node generation and edge connection establishment. The demand acquisition module uses the real-time demand vector of each position as a demand node and the real-time capability vector of each employee as a capability node. The real-time demand vector is a multi-dimensional numerical vector, with each dimension corresponding to the demand level score of a business skill; the real-time capability vector is also a multi-dimensional numerical vector, with each dimension corresponding to the employee's capability score in the corresponding business skill. The two vectors have the same dimensions, and the order of the dimensions is pre-configured and synchronized by the system.
[0022] When establishing edge connections between demand nodes and capability nodes, the matching degree between the skills required for the position and the skills possessed by the employee is calculated. This matching degree is obtained through cosine similarity calculation, using the following formula: Mi,j=Ri·Aj‖Ri‖·‖Aj‖ Where Mi,j represents the matching degree between the i-th demand node and the j-th capability node; Ri represents the real-time demand vector corresponding to the i-th demand node; Aj represents the real-time capability vector corresponding to the j-th capability node; · represents the vector dot product operation; ||Ri|| represents the Euclidean norm of vector Ri; ||Aj|| represents the Euclidean norm of vector Aj. The value range of Mi,j is [ [1,1], the closer the value is to 1, the more consistent the demand vector and the capability vector are in direction. After calculation, each matching degree Mi,j is compared with a preset similarity threshold. The preset similarity threshold is a constant set during system initialization, used to define the minimum acceptable matching degree of highly correlated connections between demand and capability. When Mi,j is greater than the preset similarity threshold, an edge is established between the i-th demand node and the j-th capability node, and the edge weight is set to Mi,j. When Mi,j is less than or equal to the preset similarity threshold, no edge is established, and the corresponding demand node-capability node pair is regarded as a noise edge and discarded. The subsequent graph feature aggregation process does not use noise edges to transmit information.
[0023] In some embodiments, the preset similarity threshold is configured to be 0.7. The basis for setting the preset similarity threshold to 0.7 is that, in a skill demand and supply matching scenario, pairings with a cosine similarity below 0.7 exhibit low correlation in actual job matching history, while pairings with a cosine similarity above 0.7 show a significant positive correlation. Thus, low-correlation connections are filtered out during the graphing stage, while high-correlation connections are retained for feature aggregation.
[0024] When establishing connections between demand nodes, historical project allocation records within the organization are retrieved. These records store project identifiers and all job identifiers involved in the project. The number of times each pair of different jobs jointly participates in the same project is extracted from these records. Specifically, for job u and job v, the number of projects in which both jobs appear simultaneously is counted, denoted as Cu,v. The number of times Cu,v jointly participates in the same project is divided by the total number of projects executed by the organization, Ntotal, to obtain the inter-job collaboration frequency Fu,v between job u and job v: Fu,v = Cu,vNtotal. A first frequency threshold is set, which is a preset minimum collaboration frequency limit used to determine whether two jobs have a close business collaboration relationship. When Fu,v is greater than the first frequency threshold, an undirected edge is established between the demand node corresponding to job u and the demand node corresponding to job v, and the edge weight of this undirected edge is assigned to Fu,v. When Fu,v is less than or equal to the first frequency threshold, no edge is established between the corresponding demand nodes.
[0025] In some embodiments, the first frequency threshold is set to 0.2. The basis for setting the first frequency threshold to 0.2 is that, through analysis of the density of the organizational collaboration network, when the proportion of two positions jointly participating in a project exceeds 20% of the total number of projects, the business dependency relationship between the two positions begins to show a stable interaction pattern. Collaboration between positions below this proportion is considered to be more accidental and is not included in the graph structure relationship constraints.
[0026] When establishing edge connections between capability nodes, historical project collaboration records between employees are retrieved. These records contain task breakdown information for each project and the employee's role in each task. The number of tasks jointly completed by each pair of employees is extracted from these records; specifically, for employees x and y, the number of tasks they are jointly assigned to the same task is counted, denoted as Tx,y. The total number of tasks each employee x and y participates in, Tx, and the total number of tasks each employee y participates in, Ty, are obtained, and their geometric mean, Tx·Ty, is calculated. The number of jointly completed tasks, Tx,y, is divided by this geometric mean to obtain the employee collaboration tightness coefficient Kx,y between employees x and y: Kx,y = Tx,yTx·Ty. A second coefficient threshold is set, which is a preset minimum threshold for determining whether two employees have a significant collaborative dependency. When Kx,y is greater than the second coefficient threshold, a directed edge is established between the capability node corresponding to employee x and the capability node corresponding to employee y. The direction of directed edges is determined as follows: The collaboration pattern between two employees in a joint task is identified from historical task records. The party initiating or leading the collaboration is designated as the collaboration initiator, and the other party as the collaboration recipient. The direction of the directed edge points from the collaboration initiator to the collaboration recipient. If both sides have initiation and acceptance records, a bidirectional edge is established, or a dominant direction is selected based on the strength of dominance. After determining the edge direction, a direction coefficient is assigned to the directed edge. The direction coefficient is equal to the collaboration closeness coefficient Kx,y between employees multiplied by the initiator's task contribution percentage (Pinitiator). Here, the initiator's task contribution percentage (Pinitiator) is the proportion of the initiator's workload to the total workload of both employees in the task completed jointly. When Kx,y is less than or equal to the second coefficient threshold, no edge is established between the corresponding capability nodes.
[0027] In some embodiments, the second coefficient threshold is set to 0.15. The basis for setting the second coefficient threshold to 0.15 is that, through cluster analysis of the employee collaboration network, when the collaboration tightness coefficient exceeds 0.15, the relationship between two employees enters the range of strong collaboration subgroups. Loose collaboration relationships below this threshold do not contribute enough to the prediction of capability synergy, and removing them can reduce graph noise and highlight the core collaboration structure.
[0028] Through the above process, after determining the edges and their weights between demand nodes and capability nodes, the undirected edges and their weights between demand nodes, and the directed edges and their direction coefficients between capability nodes, all nodes and edges are combined to generate the complete demand-capability heterogeneous graph. This heterogeneous graph contains two types of heterogeneous nodes and three types of heterogeneous edges, which are used by subsequent graph convolutional networks for node feature aggregation.
[0029] In specific implementation, please refer to Figure 2 The constructed demand-capacity heterogeneous graph is subjected to hierarchical aggregation rules in graph convolutional networks for node feature aggregation. The graph convolutional network consists of a demand-side attention aggregation subnetwork for demand node aggregation and a supply-side attention aggregation subnetwork for capacity node aggregation. The two subnetworks are symmetrical in structure but independent in parameters.
[0030] For the demand-side attention aggregation subnetwork, the processing is performed separately for each demand node. An arbitrary demand node is selected as the target demand node, and all neighboring capability nodes of the target demand node are collected. These neighboring capability nodes are those that have edges connected to the target demand node in the demand-capability heterogeneous graph. The real-time capability vector corresponding to each neighboring capability node is obtained, and simultaneously, the real-time demand vector corresponding to the target demand node is obtained.
[0031] For the k-th neighboring capability node of the target demand node, calculate the attention coefficient between the real-time demand vector of the demand node and the real-time capability vector of the capability node. The attention coefficient is calculated based on the inner product of the two vectors, as follows: ekdem=Dtarget·Sk Where ekdem represents the attention coefficient between the target demand node and the k-th neighboring capability node, which is unnormalized; Dtarget represents the real-time demand vector corresponding to the target demand node; Sk represents the real-time capability vector corresponding to the k-th neighboring capability node; and · represents the inner product operation of the vectors. The range of the attention coefficient ekdem is determined by the component value ranges of the two vectors. When all vector components are normalized to the interval [0,1], the value of ekdem is proportional to the dimension of the two vectors. The attention coefficients e1dem, e2dem, ..., eKdem corresponding to all K neighboring capability nodes of the target demand node are collected, and all attention coefficients in the set are normalized. The normalization process uses the softmax function, and the calculation method is as follows: for the k-th neighboring capability node, the normalized attention weight αkdem is equal to the exponent value of ekdem divided by the sum of the exponent values of the attention coefficients of all neighboring capability nodes. After normalization, we obtain normalized attention weights α1dem, α2dem, ..., αKdem that correspond one-to-one with the K neighboring capability nodes. The sum of all normalized attention weights is 1.
[0032] The real-time capability vectors of neighboring capability nodes are weighted and summed using normalized attention weights. The real-time capability vector Sk of the k-th neighboring capability node is multiplied by the corresponding normalized attention weight αkdem to obtain the weighted contribution vector of that neighboring capability node. The weighted contribution vectors of all K neighboring capability nodes are summed according to their element positions to obtain the neighborhood aggregation vector Ndem of the target demand node.
[0033] The real-time demand vector Dtarget of the target demand node is concatenated with the neighborhood aggregation vector Ndem. The concatenation operation links the two vectors end-to-end along their dimensional direction, resulting in a concatenated vector with a length equal to the sum of the dimensions of Dtarget and Ndem. This concatenated vector is then input into a nonlinear transformation layer. This layer contains a fully connected matrix multiplication operation and a nonlinear activation function operation. The fully connected matrix multiplication operation maps the concatenated vector to an intermediate vector of a preset dimension, and the nonlinear activation function applies a nonlinear transformation to each component of the intermediate vector. The vector output after the nonlinear transformation layer represents the demand-side aggregation feature of the target demand node.
[0034] The above process is performed on each demand node to obtain the demand-side aggregated features for each demand node. The demand-side aggregated features of all demand nodes constitute the demand-side aggregated feature set.
[0035] For the supply-side attention aggregation sub-network, the processing is performed separately for each capability node. An arbitrary capability node is selected as the target capability node, and all neighboring demand nodes are collected. These neighboring demand nodes are those that have edges connected to the target capability node in the demand-capacity heterogeneous graph. The real-time demand vector corresponding to each neighboring demand node is obtained, and simultaneously, the real-time capability vector corresponding to the target capability node is obtained.
[0036] For the m-th neighboring demand node of the target capability node, calculate the second attention coefficient between the real-time capability vector of the capability node and the real-time demand vector of the demand node. The second attention coefficient is calculated based on the dot product of the two vectors. The dot product is the sum of the products of the corresponding components of the real-time capability vector of the capability node and the real-time demand vector of the demand node, denoted as emsup. Collect the second attention coefficients e1sup, e2sup, ..., eMsup corresponding to all M neighboring demand nodes of the target capability node. Normalize all the second attention coefficients in the set using the softmax function, and calculate the normalized second attention weights α1sup, α2sup, ..., αMsup corresponding one-to-one with the M neighboring demand nodes. Multiply the real-time demand vector of the m-th neighboring demand node by the corresponding normalized second attention weight αmsup to obtain the weighted contribution vector; sum the weighted contribution vectors of all M neighboring demand nodes according to their element positions to obtain the second neighborhood aggregation vector Nsup of the target capability node. The real-time capability vector of the target capability node is concatenated with the second-neighborhood aggregation vector Nsup to obtain a concatenated vector whose length is equal to the sum of the dimensions of the real-time capability vector and the second-neighborhood aggregation vector. This concatenated vector is then input into another nonlinear transformation layer. This other nonlinear transformation layer has the same structure as the demand-side nonlinear transformation layer but with independent parameters, containing a fully connected matrix multiplication operation and a nonlinear activation function operation. The vector output after this second nonlinear transformation layer represents the supply-side aggregation feature of the target capability node.
[0037] The above process is performed on each capability node to obtain the supply-side aggregation characteristics of each capability node. The supply-side aggregation characteristics of all capability nodes constitute the supply-side aggregation characteristic set.
[0038] Each demand-side aggregated feature in the demand-side aggregated feature set is mapped to a demand-side feature through a fully connected layer. The fully connected layer contains a weight matrix and a bias vector. The number of rows in the weight matrix represents the target dimension of the demand-side feature, and the number of columns represents the dimension of the demand-side aggregated feature. The dimension of the bias vector is equal to the target dimension of the demand-side feature. For each demand-side aggregated feature, it is multiplied by the weight matrix of the fully connected layer, and then the bias vector is added to obtain the demand-side feature.
[0039] Each supply-side aggregated feature in the supply-side aggregated feature set is mapped to a supply-side feature through another fully connected layer. This other fully connected layer has the same structure as the fully connected layer used for the demand side but with independent parameters. The number of rows in the weight matrix corresponds to the target dimension of the supply-side feature, and the number of columns corresponds to the dimension of the supply-side aggregated feature. For each supply-side aggregated feature, it is multiplied by the weight matrix of the other fully connected layer, and then a bias vector is added to obtain the supply-side feature.
[0040] In some embodiments, the target dimension of the demand-side feature and the target dimension of the supply-side feature are set to the same value to ensure that the dimensions of the two features are aligned with the requirements in the subsequent time series forecasting model.
[0041] In the hierarchical aggregation process of the graph convolutional network described above, the parameters involved in the demand-side attention aggregation subnetwork and the supply-side attention aggregation subnetwork include the fully connected weight matrices and bias vectors in two nonlinear transformation layers, and the weight matrices and bias vectors in two mapping fully connected layers. These parameters are determined through end-to-end training. During training, the real-time demand vector and real-time capacity vector of the historical time window are used as inputs, and the actual demand vector of the corresponding subsequent time window is used as the supervision label. Mean squared error is used as the loss function, and gradient descent algorithm is used for parameter updates. The training of the graph convolutional network is completed offline before system deployment. After training, the parameters are fixed, and the demand-side features and supply-side features are computed forward during system runtime.
[0042] In implementation, the pre-defined time-series prediction model employs a dual-encoder-attention-decoder architecture. The dual encoders consist of a first long short-term memory (LSTM) network encoder and a second LSM network encoder; the two LSM network encoders are structurally independent and do not share parameters. The decoder uses a gated recurrent unit (GRU) structure. (See also...) Figure 3 The input data comes from the demand-side and supply-side features generated by the graph feature aggregation module in multiple consecutive historical time windows.
[0043] Demand-side features from multiple consecutive historical time windows are arranged chronologically to form a demand-side time series. Each time step in the demand-side time series corresponds to a demand-side feature vector for a historical time window, with the time step index from smallest to largest indicating the time period from oldest to most recent. Supply-side features from multiple consecutive historical time windows are arranged chronologically to form a supply-side time series, with the time steps of the supply-side time series aligned one-to-one with those of the demand-side time series.
[0044] The demand-side time series is fed into the first Long Short-Term Memory (LSTM) network encoder. The LSM encoder consists of multiple LSM units cascaded along the time axis. Each LSM unit receives the demand-side feature vector at the current time step, as well as the hidden state and cell state passed from the previous time step, outputs the demand-side hidden state at the current time step, and updates the cell state. After forward propagation through all time steps, the LSM encoder outputs the demand-side hidden state sequence. The hidden state vector at each time step in the demand-side hidden state sequence corresponds to a demand-side encoded representation of a historical time window. The dimension of the hidden state vector is determined by the number of hidden layer units in the LSM encoder.
[0045] The supply-side time series data is fed into a second long short-term memory (LSM) network encoder. The internal structure of the second LSM encoder is the same as that of the first LSM encoder, but its network weight parameters are independent. The second LSM encoder receives the supply-side feature vector at each time step and outputs the supply-side hidden state for that time step. After forward propagation through all time steps, the second LSM encoder outputs a sequence of supply-side hidden states, where the hidden state vector at each time step corresponds to the supply-side encoded representation of a historical time window.
[0046] The demand-side and supply-side hidden state sequences are concatenated at each time step to obtain a fused hidden state sequence. The concatenation operation links the two hidden state vectors at the same time step end-to-end in the dimensional direction. If the dimension of the demand-side hidden state vector is ddem and the dimension of the supply-side hidden state vector is dsup, then the dimension of the fused hidden state vector at that time step is ddem + dsup. The concatenation results of all historical time steps constitute the fused hidden state sequence. This fused hidden state sequence is then input into the attention mechanism layer. The attention mechanism layer maintains a learnable query vector with the same dimension as the hidden state dimension at each time step in the fused hidden state sequence, i.e., ddem + dsup. The query vector serves as a trainable parameter for the attention mechanism layer and is updated through backpropagation during the training of the time series prediction model. For the t-th historical time step, the dot product of the query vector and the fused hidden state vector at that time step is calculated to obtain the original score for the t-th historical time step. The dot product is the sum of the products of the corresponding components of the two vectors, and the original score for the t-th historical time step is denoted as st. Iterate through all T historical time steps to obtain the original score sequence s1, s2, ..., sT. Input all T original scores into the normalized exponential function. The normalized exponential function outputs the attention weight for each historical time step, calculated using the following formula: wt = exp(st)τ = 1Texp(sτ) Where wt represents the attention weight at the t-th historical time step; st represents the raw score at the t-th historical time step; sτ represents the raw score at the τ-th historical time step, and the summation symbol accumulates the exponential values of the raw scores of all historical time steps from 1 to T; exp represents the exponential function with the natural constant e as the base. The attention weight wt ranges from (0,1), and the sum of the attention weights at all time steps is 1.
[0047] After obtaining the attention weights for each historical time step, the fused hidden state vector for each historical time step is multiplied by the corresponding attention weight to obtain the weighted hidden state vector for that time step. The weighted hidden state vectors of all T historical time steps are then summed element-wise to obtain the context vector. The dimension of the context vector is the same as the dimension of a single fused hidden state vector.
[0048] The context vector is input into the decoder. The decoder employs a gated recurrent unit (GRU) structure, which contains two gating mechanisms: an update gate and a reset gate. The decoder uses the context vector as the initial hidden state and generates the output sequence step by step, starting from a preset initial symbol vector. In each generation step, the GRU receives the output vector from the previous step and the current hidden state. After processing the update gate, reset gate, and candidate hidden state, it outputs a new hidden state. This new hidden state is then mapped to the output vector of the current step through a fully connected output layer. The number of generation steps set by the decoder corresponds to the dimension of the predicted demand vector. Each step generates one dimension component of the predicted demand vector. After all generation steps are completed, the predicted demand vector for the next time window is obtained.
[0049] In some embodiments, the number of hidden layer units in the Long Short-Term Memory network encoder, the number of hidden layer units in the gated recurrent unit decoder, and the dimension of the query vector are all set to the same value to simplify dimension alignment and reduce the complexity of hyperparameter tuning.
[0050] In the time-series prediction model, the weight parameters of the first and second Long Short-Term Memory (LSTM) encoders, the query vector values, the weight parameters of the gated recurrent unit (GRU) decoder, and the weight and bias parameters of the fully connected output layer are all trained using supervised learning. The training data consists of demand-side feature sequences and supply-side feature sequences from multiple consecutive time windows in historical periods, along with the corresponding actual demand vectors for the next time window. During training, the mean squared error loss function is used to measure the difference between the predicted and actual demand vectors. An adaptive moment estimation optimizer is used to update all trainable parameters. After training iterations until the loss function converges, the model parameters are saved for inference.
[0051] In practice, the dynamic grouping module receives the predicted demand vector from the time-series forecasting module and the real-time capability vector from all employee personnel files from the demand acquisition module. Each employee personnel file is defined as a file item, and each file item contains a real-time capability vector and an employee identifier. The real-time capability vector is a multi-dimensional numerical vector extracted from the employee's personnel information file, with each dimension corresponding to a capability score for a business skill; the employee identifier is a string or code that uniquely distinguishes different employees.
[0052] The predicted demand vector is decomposed into multiple sub-demand vectors. Each predicted demand vector is a multi-dimensional numerical vector, with each dimension corresponding to the predicted demand level of a business skill in the next time window. The decomposition operation is performed based on a system-preset business job template. Each business job template specifies a subset of skills and a corresponding weight configuration. For each business job template, the dimension components belonging to that template's skill subset are extracted from the predicted demand vector to form the corresponding sub-demand vector. The dimension of the sub-demand vector equals the number of skills contained in that template's skill subset. Each sub-demand vector corresponds to a target capability combination for a file group, determined by the product of the sub-demand vector's dimension values and the template's weight configuration.
[0053] Initialize a state table. The state table is designed as a three-dimensional structure. The first dimension is the file item index, with values ranging from 1 to the total number of file items, I. Each row corresponds to the index order of a file item. The second dimension is the satisfaction level vector of each sub-demand vector. The satisfaction level vector is a fixed-length vector whose dimension is equal to the sum of the dimensions of all sub-demand vectors. The third dimension is the maximum capability coverage when reaching the current state. The state table is stored using a key-value pair structure. The key consists of the row index and the satisfaction level vector, and the values are the corresponding maximum capability coverage and the selected path record. The selected path record is a list that stores the file item indices of the selected file items in order.
[0054] For the first row of the status table, the corresponding file item index is 1. The satisfied level vector is initialized to an all-zero vector. The dimension of the all-zero vector is equal to the sum of the dimensions of all sub-requirement vectors, and each component value of the vector is 0. The maximum capability coverage is initialized to 0.
[0055] In practical implementation, capability coverage represents the overall coverage level of the satisfied requirement vector across all sub-requirement vectors. The formula for calculating capability coverage is: O=p=1Pcpqp Where O represents the calculated capability coverage, a non-negative real number; P represents the total number of dimensions of all sub-demand vectors, i.e., the sum of the dimensions of each sub-demand vector; p represents the dimension index, traversing all dimensions from 1 to P; cp represents the value of the current satisfied level vector in the p-th dimension; qp represents the capability threshold corresponding to the p-th dimension, which is directly taken from the demand level score of the corresponding sub-demand vector in the corresponding dimension, i.e., the value of each component of the sub-demand vector is the capability threshold of its respective dimension. In the actual calculation of capability coverage O, if qp is 0 in a certain dimension, the ratio term for that dimension is set to 0 to avoid division by zero.
[0056] Starting from the second row of the state table, iterate through each file item sequentially. For the current file item, let its corresponding file item index be i, and obtain the real-time capability vector of the i-th file item. Iterate through all the recorded states in the previous row, i.e., the file item index is i. The state table row corresponding to 1. For each recorded state in the previous row, obtain the satisfaction level vector and maximum capability coverage corresponding to that state.
[0057] For a recorded state, attempt to overlay the real-time capability vector of the current file item onto the satisfaction level vector of that state. The overlay operation involves adding the satisfaction level vector and the real-time capability vector component-by-component according to their dimensional correspondence. The dimensional correspondence between the real-time capability vector and the satisfaction level vector is determined by a system-preset dimensional mapping table, which records which component position of which sub-demand vector within the satisfaction level vector corresponds to each skill dimension of the real-time capability vector. For component positions with existing mappings, the value of the real-time capability vector component is added to the value of the corresponding component in the satisfaction level vector; for component positions without mappings, the corresponding component of the satisfaction level vector remains unchanged. After overlay, a new candidate satisfaction level vector is generated.
[0058] Check whether the value of each dimension in the new candidate satisfaction vector exceeds the capability threshold of the corresponding sub-demand vector in the same dimension. For the p-th dimension, if the value of the candidate component is greater than qp, then truncate the value of the candidate component to qp, i.e., set it equal to qp; if the value of the candidate component is less than or equal to qp, then retain the original value. The truncation operation ensures that each dimension of the satisfaction vector does not exceed the upper limit of the capability threshold.
[0059] After truncation, the superimposed capability coverage is calculated using the capability coverage calculation formula described above. Each dimension of the new candidate satisfiedness vector is used as cp and substituted into the formula to calculate the capability coverage Ocandidate corresponding to the candidate state. After obtaining the capability coverage Ocandidate, it is determined whether the state table needs to be updated. In the i-th row of the state table, it is checked whether there exists a satisfiedness vector record that is exactly the same as the new candidate satisfiedness vector. If no identical record exists, the new candidate satisfiedness vector and its corresponding capability coverage Ocandidate are written together into the i-th row of the state table, and the selection path to that state is recorded. The selection path is recorded by copying the selection path list of the previously recorded state and appending the current file item index i to the end of the list. If a record with the same satisfiedness vector as the candidate already exists in the i-th row of the state table, the maximum recorded capability coverage is compared with the size of Ocandidate. When Ocandidate is strictly greater than the recorded maximum capability coverage, the recorded maximum capability coverage is replaced by Ocandidate, and the original selection path record is replaced by the currently constructed selection path; when Ocandidate is less than or equal to the recorded maximum capability coverage, the original record at that position in the state table remains unchanged.
[0060] In some embodiments, to control the size of the state table and reduce the space complexity of dynamic programming, an upper limit is set on the number of states in the i-th row of the state table when traversing each file item. When the number of states recorded in the i-th row exceeds the preset upper limit, they are sorted from highest to lowest according to their maximum capability coverage, and the first preset upper limit of states are retained while the remaining states are discarded. The preset upper limit is determined based on the available storage space capacity, taking into account both solution quality and computational resource consumption. After traversing all I file items, the target state that makes the satisfaction level of all sub-demand vectors reach the threshold is found from the last row of the state table, i.e., the i-th row. The determination method for the satisfaction level reaching the threshold is: the component value of each dimension of the satisfaction level vector corresponding to the target state is greater than or equal to the capability threshold qp of the corresponding sub-demand vector, that is, for any dimension p, cp≥qp is satisfied. If there are multiple target states that meet the condition, the one with the highest maximum capability coverage is selected as the final target state. The selection path list corresponding to the final target state is backtracked, and the file item index contained in the list is the selected file item.
[0061] Based on the backtracked file item selection path, the selected HR file items are categorized according to their corresponding sub-requirement vectors. For each selected file item in the selection path, its assigned group is determined based on the matching degree between its real-time capability vector and each sub-requirement vector. The matching degree is calculated as the sum of the component values of the real-time capability vector on the corresponding skill dimension of the sub-requirement vector. The file item is assigned to the group corresponding to the sub-requirement vector with the highest matching degree. If the same file item simultaneously meets the requirements of multiple sub-requirement vectors, i.e., the real-time capability vector has the highest matching degree with multiple sub-requirement vectors, then a unique assignment is made based on the assignment weight of the file item's historical group records. The assignment weight of historical group records is obtained by querying the HR information file index table, using the employee identifier as the key to retrieve the frequency ratio of the employee being assigned to each group in past file groups, and selecting the group with the highest frequency ratio as the unique assignment target; if there is no historical record, the assignment is made according to the priority order of the sub-requirement vectors, and the priority order is determined by the arrangement order of the business position templates in the system configuration. After classification, target file grouping results are generated. The target file grouping results include group identifiers, sub-requirement vectors corresponding to the group, and a list of employee identifiers assigned to the group.
[0062] In specific implementation, please refer to Figure 4 The blockchain evidence storage module receives the target file grouping results generated by the dynamic grouping module. The target file grouping results include a group identifier, the corresponding sub-requirement vector, and a list of employee identifiers assigned to the group. The blockchain evidence storage module serializes the target file grouping results into binary data blocks. The serialization operation uses a preset data exchange format protocol, converting the group identifier in the target file grouping results into a fixed-length string encoding, converting the values of each dimension of the sub-requirement vector into a fixed-width floating-point binary representation, converting each employee identifier in the employee identifier list into a fixed-length string encoding, and inserting separators between the data fields. After the conversion, a continuous binary data block is obtained.
[0063] The hash value is calculated for the serialized binary data block. The hash value calculation uses a cryptographic hash algorithm. The binary data block is taken as input, and through iterative steps such as padding, block division, and compression, a fixed-length hash value is generated. The length of the hash value depends on the version of the hash algorithm used; when using the SHA-256 algorithm, the hash value length is 256 bits.
[0064] A leader node responsible for record-keeping is selected from the blockchain network. The blockchain network's node list maintains the leader node identifier for each consensus cycle. The selection method involves querying the leader node identifier corresponding to the current consensus cycle and obtaining the leader node's communication address through network addressing. The serialized binary data block and the calculated hash value are packaged into transaction data, which is then sent to the leader node via a secure transmission protocol. The secure transmission protocol encrypts and verifies the integrity of the transaction data at the application layer.
[0065] After receiving transaction data, the leader node verifies the consistency between the hash value and the binary data block. If the verification is successful, it packages the transaction data into a new block, executes the consensus process, and appends the block to the blockchain. The blockchain notarization module receives transaction confirmation information returned by the leader node. This confirmation information includes the block height and the block timestamp. The block height represents the sequential number of the transaction within the blockchain, a monotonically increasing integer counting from the genesis block. The block timestamp represents the Unix time at the time the block was generated, accurate to the second.
[0066] Based on the received block height and block timestamp, add a record to the personnel information file index table. The added record includes the file group identifier, employee identifier list, block height, and block timestamp. The file group identifier is directly taken from the group identifier in the target file grouping result. The employee identifier list is directly taken from the employee identifier list in the target file grouping result. The block height and block timestamp are taken from the block height and block timestamp in the transaction confirmation information, respectively.
[0067] The personnel information file index table uses a skip list structure for storage. The skip list structure consists of multiple levels of ordered linked lists. The bottom-level linked list contains all record nodes, and each upper-level linked list is a subset of the indexes of the lower-level linked lists. Each record node uses a composite key consisting of the employee identifier and the block height, with the block height and the block timestamp as values. The keys are sorted lexicographically by employee identifier. The number of skip list levels is determined by a preset probability factor during index table creation, controlling the probability of each node being promoted to a higher level. When inserting a new record into the skip list, the insertion position is searched level by level starting from the top level, creating a new node in the corresponding level's linked list. Nodes at different levels are connected by pointers. When searching the skip list by employee identifier prefix, a range scan is performed starting from the top level based on the employee identifier prefix of the composite key to quickly locate all record nodes corresponding to that employee. The query operation using employee identifier as the key supports retrieving archive grouped historical records by time range. Specifically, after locating the corresponding record node based on the employee identifier, iterate through all record nodes linked in chronological order under that employee identifier, and filter out records whose block timestamps fall within the specified time range.
[0068] The updated personnel information file index table is synchronously stored in both the local cache and the distributed database. When synchronizing to the local cache, newly added record nodes in the personnel information file index table are cached in a skip list data structure in memory, with the local cache using a least recently used strategy to manage its capacity. When synchronizing to the distributed database, newly added records in the personnel information file index table are converted into key-value pairs, with the key being the employee identifier and the value being a combination of the serialized block height and the block timestamp. Storage operations are submitted through the distributed database's batch write interface. The distributed database employs a multi-replica replication mechanism to ensure data persistence and availability; the number of replicas is configured during system initialization. A strategy combining periodic synchronization and write penetration is used between the local cache and the distributed database to ensure that query operations prioritize hitting the local cache for low-latency responses, while simultaneously guaranteeing the complete and persistent storage of data in the distributed database.
[0069] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A human resource information file management system based on human resource demand analysis, characterized in that: include: The requirements acquisition module acquires real-time requirements vectors for each position within the organization and real-time capability vectors from employee personnel files. The graph feature aggregation module constructs a demand-capacity heterogeneous graph based on the real-time demand vector and the real-time capability vector, and uses a graph convolutional network to aggregate node features of the heterogeneous graph to generate demand-side features and supply-side features. The time-series forecasting module inputs the demand-side features and supply-side features into a preset time-series forecasting model and outputs the forecasted demand vector for the next time window. The dynamic grouping module regroups employee personnel files using dynamic programming based on the predicted demand vector and the real-time capability vector, generating target file grouping results. The blockchain evidence storage module writes the target file grouping results into the blockchain storage node and updates the human resources information file index table.
2. The human resource information file management system based on human resource demand analysis according to claim 1, characterized in that, Based on the real-time demand vector and the real-time capability vector, a demand-capacity heterogeneous graph is constructed, and a graph convolutional network is used to aggregate node features of the heterogeneous graph to generate demand-side features and supply-side features, including: The real-time demand vector of each position is taken as a demand node, and the real-time capability vector of each employee is taken as a capability node. Between the demand node and the capability node, edge connections are established based on the matching degree between the skills required for the position and the skills held by the employee. The matching degree is calculated using cosine similarity. Between the required nodes, edge connections are established based on the frequency of business collaboration between positions; Edge connections are established between the capability nodes based on the project collaboration history between employees; For the constructed demand-capacity heterogeneous graph, the hierarchical aggregation rule in graph convolutional networks is adopted to first aggregate the features of the neighboring capacity nodes of each demand node to obtain the demand-side aggregation features, and then aggregate the features of the neighboring demand nodes of each capacity node to obtain the supply-side aggregation features. The demand-side aggregated features are mapped to the demand-side features through a fully connected layer, and the supply-side aggregated features are mapped to the supply-side features through another fully connected layer.
3. The human resource information file management system based on human resource demand analysis according to claim 2, characterized in that, When establishing edge connections between the demand node and the capability node based on the matching degree, only edges with a matching degree greater than a preset similarity threshold are retained, and edges with a matching degree lower than the threshold are discarded as noise edges.
4. The human resource information file management system based on human resource demand analysis according to claim 1, characterized in that, The demand-side and supply-side characteristics are input into a preset time-series forecasting model, and the predicted demand vector for the next time window is output, including: The demand-side characteristics and supply-side characteristics of multiple consecutive historical time windows are arranged in chronological order to form a demand-side time series and a supply-side time series. The demand-side time series and the supply-side time series are respectively input into two independent long short-term memory network encoders to obtain the demand-side hidden state sequence and the supply-side hidden state sequence. The demand-side hidden state sequence and the supply-side hidden state sequence are spliced together at each time step to obtain the fused hidden state sequence. The fused hidden state sequence is input into the attention mechanism layer, the attention weight of each historical time step is calculated, and the fused hidden state sequence is weighted and summed to obtain the context vector. The context vector is input into the decoder, which employs a gated recurrent unit structure to progressively generate the prediction demand vector for the next time window.
5. The human resource information file management system based on human resource demand analysis according to claim 1, characterized in that, Based on the predicted demand vector and the real-time capability vector, a dynamic programming method is used to regroup employee personnel files, generating target file grouping results, including: Each employee's personnel file is defined as a file item, and each file item contains the real-time capability vector and the employee identifier; The predicted demand vector is decomposed into multiple sub-demand vectors, each sub-demand vector corresponding to a target capability combination of a file group; Initialize a three-dimensional state table, wherein the first dimension of the state table corresponds to the index order of the file items, the second dimension corresponds to the satisfaction degree vector of each sub-requirement vector, and the third dimension corresponds to the maximum capability coverage when the current state is reached; Each file item is traversed sequentially. For the current file item, the state table is updated based on the feasibility of superimposing the real-time capability vector of the file item with the satisfaction level vector of each state in the current state table, and the file item selection path when each new state is reached is recorded. After traversing all file items, search the state table for the target state that makes the satisfaction level of all sub-requirement vectors reach the threshold, and backtrack the file item selection path corresponding to the target state. Based on the path of the archive item selection obtained from the backtracking, the selected human resources archive items are classified according to the corresponding sub-demand vectors to generate the target archive grouping results.
6. The human resource information file management system based on human resource demand analysis according to claim 5, characterized in that, When selecting the human resource archive items according to the path obtained from the backtracking and classifying the selected human resource archive items according to the corresponding sub-demand vectors, if the same archive item meets the requirements of multiple sub-demand vectors at the same time, it will be uniquely assigned according to the belonging weight of the archive item's historical grouping records.
7. The human resource information file management system based on human resource demand analysis according to claim 1, characterized in that, Write the target file grouping results to the blockchain storage node and update the human resources information file index table, including: The target file grouping results are serialized into binary data blocks, and the hash value of the binary data blocks is calculated; Select the current leader node responsible for bookkeeping in the blockchain network, and send the binary data block and the hash value to the leader node; Receive transaction confirmation information returned by the leader node, the transaction confirmation information including block height and block timestamp; Based on the block height and block timestamp, add a record to the human resources information file index table. The record includes a file group identifier, an employee identifier list, the block height, and the block timestamp. The updated human resources information file index table is synchronously stored in the local cache and distributed database.
8. The human resource information file management system based on human resource demand analysis according to claim 7, characterized in that, The human resources information file index table is stored in a skip list structure, with the employee identifier as the key and the block height and block timestamp as the value, supporting querying file group history records by time range.
9. The human resource information file management system based on human resource demand analysis according to claim 2, characterized in that, Between the required nodes, edge connections are established based on the frequency of business collaboration between roles, including: Obtain historical project allocation records from within the organization, and extract the number of times each pair of positions jointly participated in the same project from the historical project allocation records; Divide the number of times the two parties jointly participated in the same project by the total number of projects to obtain the frequency of business collaboration between positions. A first frequency threshold is set. When the business collaboration frequency between two positions is greater than the first frequency threshold, an undirected edge is established between the two corresponding requirement nodes, and a weight value is assigned to the edge, which is equal to the business collaboration frequency.
10. The human resource information file management system based on human resource demand analysis according to claim 2, characterized in that, Between the capability nodes, edge connections are established based on the project collaboration history between employees, including: Obtain historical project collaboration records among employees, and extract the number of tasks jointly completed by each pair of employees from the historical project collaboration records; The employee collaboration tightness coefficient is obtained by dividing the number of tasks jointly completed by the geometric mean of the total number of tasks each employee participated in. A second coefficient threshold is set. When the cooperation closeness coefficient between two employees is greater than the second coefficient threshold, a directed edge is established between the two corresponding capability nodes. The direction of the directed edge is from the cooperation initiator to the cooperation recipient, and a direction coefficient is assigned to the edge. The direction coefficient is equal to the cooperation closeness coefficient multiplied by the task contribution ratio of the initiator.