Industrial talent big data analysis method based on retrieval enhancement generation

By constructing an industrial talent intelligence map and nonlinear diffusion equation, combining multi-objective optimization and projection tracking forest algorithm, the problem of failure to comprehensively evaluate the impact of industrial talent loss and collaborative interruption in traditional methods is solved, and more scientific management strategy formulation and risk prediction are achieved.

CN120542952APending Publication Date: 2025-08-26JIANGSU SHUANGCHUANG TALENT UNITED CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510607543.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

When evaluating the impact of industrial talent loss and collaborative disruption on employers, the existing technology fails to fully consider the indirect and long-term impact, and lacks scientific basis for strategy formulation. Traditional methods only focus on single indicators and empirical judgments.

Method used

The industrial talent big data analysis method based on search enhancement is adopted to build an industrial talent intelligence map, quantify the impact of nonlinear diffusion equations, combine multi-objective optimization framework and projection tracking forest algorithm to search for the optimal talent management strategy.

Benefits of technology

More comprehensively assess the impact of talent loss and collaborative disruption, provide more scientific management strategies, balance short-term costs and long-term value, and improve the effectiveness and robustness of management strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542952A_ABST
    Figure CN120542952A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial talent big data analysis method based on retrieval enhancement generation, and relates to the field of talent data processing, and the method comprises the steps: building an industrial talent information graph based on industrial talent data and in combination with a retrieval enhancement generation technology, and enabling the industrial talent information graph to be used for quantifying talent flow data and cooperation density data; according to the industrial talent intelligence map and the constructed nonlinear diffusion equation, the influence of talent loss and cooperative interruption on the employer is quantified; evaluating the centrality of talent nodes in the industrial talent intelligence graph by using an evaluation algorithm; on the basis of a multi-objective optimization framework, an optimal talent management strategy is searched in combination with an influence quantification result of talent loss and cooperative interruption on an employer and a talent node centrality evaluation result; and predicting the effectiveness and potential risk of the optimal talent management strategy by using a projection tracking forest algorithm and influence factor data. By analyzing industrial talent data, scientific and effective talent management strategy suggestions are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of talent data processing, and in particular to an industrial talent big data analysis method based on retrieval enhancement generation. Background Art

[0002] Industrial talent refers to individuals who possess the specialized knowledge, skills, and experience within a specific industry or sector, driving technological innovation, management optimization, and business development within that sector. This talent encompasses not only those directly engaged in technology research and development and manufacturing, but also those with comprehensive expertise in operations management, market development, financial analysis, and other areas. They are a key force in promoting industrial upgrading and enhancing corporate competitiveness.

[0003] By analyzing industry talent big data, we can comprehensively assess talent value and develop more scientific and effective talent management strategies to reduce the risks of talent loss and collaboration disruption, ultimately improving the overall competitiveness of the organization. However, there are still the following deficiencies in analyzing industry talent big data:

[0004] (1) How to more comprehensively evaluate the impact of talent loss and collaboration disruption on employers, while traditional analysis methods usually only consider direct impacts and ignore indirect and long-term impacts.

[0005] (2) How to evaluate talent value more effectively, while traditional analytical methods usually only focus on a single indicator, such as performance or cost.

[0006] (3) How to formulate more scientific and effective talent management strategies. Traditional strategy formulation is based on experience and judgment and lacks scientific basis.

[0007] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention

[0008] In response to the problems in the related technology, the present invention proposes an industrial talent big data analysis method based on retrieval enhancement generation to overcome the above-mentioned technical problems existing in the existing related technology.

[0009] To this end, the specific technical solutions adopted in the present invention are as follows:

[0010] A method for analyzing industrial talent big data based on retrieval-enhanced generation, comprising:

[0011] Based on industry talent data and combined with search-enhanced generation technology, we build an industry talent intelligence map to quantify talent flow data and collaboration density data;

[0012] Based on the industry talent intelligence map and the constructed nonlinear diffusion equation, the impact of talent loss and collaboration disruption on employers is quantified. The centrality of talent nodes in the industry talent intelligence map is evaluated using an evaluation algorithm.

[0013] Based on a multi-objective optimization framework, we search for the optimal talent management strategy by combining the quantitative results of the impact of talent loss and collaboration disruption on employers and the results of talent node centrality assessment.

[0014] The projection pursuit forest algorithm and influencing factor data are used to predict the effectiveness and potential risks of the optimal talent management strategy.

[0015] Furthermore, based on industry talent data and combined with search-enhanced generation technology, an industry talent intelligence map is constructed to quantify talent flow data and collaboration density data, including:

[0016] Obtain standardized and structured industry talent data; use knowledge graph construction tools to convert industry talent data into an industry talent intelligence graph, where nodes represent entities such as talent, institutions, skills, and projects, and edges represent relationships between entities;

[0017] Utilize retrieval-enhanced generation technology to transform the information in the industry talent intelligence map into a computable vector representation and learn the relationship patterns between nodes to enhance the industry talent intelligence map;

[0018] Quantify talent flow data and collaboration density data through industrial talent intelligence maps.

[0019] Furthermore, the talent flow data and collaboration density data quantified through the industry talent intelligence map include:

[0020] Based on the employment history information of talents in the industry talent intelligence map, calculate the mobility frequency of each talent, the talent mobility rate of employers, and the talent mobility data between specific employers;

[0021] Based on the talent cooperation relationship data in the industrial talent intelligence map, the number of collaborations between talents is calculated, and weighted calculation is performed according to the strength of the cooperation relationship to obtain the collaboration density data.

[0022] Furthermore, based on the industry talent intelligence map and the constructed nonlinear diffusion equation, the impact of talent loss and collaboration disruption on employers is quantified, including:

[0023] Using pre-set rules for judging talent loss and collaboration interruption, relevant node and edge information is extracted from the industry talent intelligence map to form a data set of talent loss and collaboration interruption events;

[0024] According to the characteristics and impact range of brain drain and collaboration disruption events, the parameters of the nonlinear diffusion equation are set, including the diffusion coefficient and the attenuation coefficient;

[0025] Use numerical methods to solve the nonlinear diffusion equation, discretize time, and gradually calculate the degree of influence of each node at each time step;

[0026] The impact levels of nodes related to employers are aggregated, and the quantitative impact of talent loss and collaboration disruption on employers is calculated.

[0027] Furthermore, the expression of the nonlinear diffusion equation is:

[0028]

[0029] Where D ij represents the collaboration density weight between node i and node j;

[0030] w ij Represents the weight of edge (i, j) in the industrial talent intelligence graph;

[0031] N(i) represents the set of neighbor nodes of node i;

[0032] C i (t) represents the impact degree of node i at time t, C j (t) represents the impact value of node j at time t;

[0033] α i represents the talent loss rate of the employer where node i is located;

[0034] S i Represents a source term.

[0035] Furthermore, according to the characteristics and impact range of the brain drain and collaboration interruption events, the parameters of the nonlinear diffusion equation are set as follows:

[0036] Based on the feature vectors of brain drain and collaboration disruption events, talent flow data, collaboration density data, and industry talent intelligence maps, and using network analysis methods combined with expert knowledge, the impact scope of brain drain and collaboration disruption events is assessed;

[0037] A machine learning model is used to learn the relationship between the feature vectors and impact ranges of talent loss and collaboration disruption events and the parameters of the nonlinear diffusion equation, and output the diffusion coefficient and attenuation coefficient of each talent loss and collaboration disruption event.

[0038] Furthermore, the centrality of talent nodes in the industry talent intelligence map is evaluated using an evaluation algorithm, including:

[0039] Using the selected centrality evaluation algorithm, calculate the weighted degree centrality, weighted closeness centrality, and weighted betweenness centrality of each node in the industrial talent intelligence map; then sum them up to obtain the weighted centrality score of each node;

[0040] Normalize the weighted centrality score of each node and store the nodes and their corresponding weighted centrality scores in a table;

[0041] Among them, when calculating weighted degree centrality, the collaboration density of the node is used as the weight of the edge in the industrial talent intelligence map, and the higher the collaboration density, the greater the weight of the edge;

[0042] When calculating weighted proximity centrality, the talent flow data between nodes is used as the weight of the distance. The more frequent the talent flow, the closer the distance.

[0043] When calculating weighted betweenness centrality, the sum of the collaboration densities of the nodes on the shortest path through the node is used as the weight of the path. The higher the collaboration density, the greater the weight of the path.

[0044] Normalize the edge weights, distance weights, and path weights.

[0045] Furthermore, based on a multi-objective optimization framework, combined with the quantitative results of the impact of talent loss and collaboration disruption on employers and the results of talent node centrality assessment, the optimal talent management strategy is searched, including:

[0046] The calculated quantitative impact of talent loss and collaboration disruption on employers and the weighted centrality scores of talent nodes are normalized and used as benchmarks for the total loss value caused by talent loss and collaboration disruption and the total value of the talent team; this also sets a benchmark for the total cost of talent management.

[0047] Construct the objective function of the multi-objective optimization problem and use resource efficiency as a constraint. By setting the parameters of the multi-objective optimization algorithm, a configured multi-objective optimization algorithm is obtained, and the decision variables are defined, including the standardized number of employees to be hired, the training investment for each employee, and the incentive investment for each employee.

[0048] Use the configured multi-objective optimization algorithm to search the solution space of the multi-objective optimization problem. At each iteration, combine the values ​​of the current decision variables to calculate the total loss value caused by the current talent loss and collaboration interruption, the total cost of current talent management, and the total value of the current talent team.

[0049] By substituting the total loss value caused by current talent loss and collaboration interruption, the total cost of current talent management, and the total value of the current talent team into the objective function, the pros and cons of the current solution are evaluated;

[0050] Select the optimal solution and convert the values ​​of the corresponding decision variables into the optimal talent management strategy, including recruitment strategy, training strategy and incentive strategy.

[0051] Furthermore, the objective function and constraints of the multi-objective optimization problem are:

[0052] Minimize L=Lo+Co-Va

[0053] Co≤B;

[0054] Where, L represents the total loss;

[0055] Lo represents the total loss value caused by talent loss and collaboration interruption;

[0056] Co represents the total cost of talent management;

[0057] Va represents the total value of the talent team;

[0058] B represents the budget value;

[0059] Among them, by establishing linear relationships between the values ​​of decision variables and the total loss value caused by talent loss and collaboration interruption, the total cost of talent management and the total value of the talent team, the total loss value caused by the current talent loss and collaboration interruption, the total cost of current talent management and the total value of the current talent team are calculated.

[0060] Furthermore, using the projection pursuit forest algorithm and influencing factor data, we predict the effectiveness and potential risks of the optimal talent management strategy, including:

[0061] Based on the optimal talent management strategy, the optimal strategy feature vector is obtained; based on the influencing factor data, the influencing factor feature vector is obtained;

[0062] Use historical talent management strategy data, corresponding influencing factor data, and strategy effectiveness results as training data sets;

[0063] Use the training dataset to train the projection pursuit forest model to learn the nonlinear relationship between policy features, influencing factors and policy effectiveness;

[0064] Use the trained projection pursuit forest model to predict the effectiveness of the current optimal talent management strategy under given influencing factors and output the predicted effectiveness probability;

[0065] When the probability of effectiveness is lower than the preset threshold, the most important influencing factors are found and set as potential risks based on the feature importance analysis technology of the projection pursuit forest model.

[0066] The beneficial effects of the present invention are:

[0067] (1) Using retrieval-enhanced generation technology to construct an industry talent intelligence map and quantify talent mobility and collaboration density. Retrieval-enhanced generation technology automatically learns the relationship patterns between nodes from large-scale data and integrates them into the map, thereby constructing a more complete and accurate industry talent intelligence map.

[0068] (2) Quantify the impact of talent loss and collaboration disruption on employers based on the nonlinear diffusion equation. The nonlinear diffusion equation simulates the process by which the impact of talent loss and collaboration disruption propagates through organizational networks, thereby more comprehensively assessing their impact on employers and gaining a deeper understanding of the potential risks of talent loss and collaboration disruption.

[0069] (3) Using the evaluation algorithm, the centrality of talent nodes in the industry talent intelligence map is evaluated, and the quantitative value of the impact of talent loss and collaboration interruption on employers and the centrality evaluation value of talent nodes are used as the benchmark for multi-objective optimization. Combined with the total cost factor of talent management, the value of talent is evaluated more comprehensively and effectively, the short-term cost and long-term value are balanced, and the effectiveness of talent management strategies is improved.

[0070] (4) Combine a multi-objective optimization framework and the projection pursuit forest algorithm to search and predict the effectiveness and potential risks of the optimal talent management strategy. A multi-objective optimization framework is used to simultaneously consider multiple objectives and search for the optimal talent management strategy. Combined with the projection pursuit forest algorithm, the effectiveness and potential risks of the optimal talent management strategy under different influencing factors are predicted, thereby helping employers select more robust and effective strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0072] Figure 1 This is a flowchart of an industrial talent big data analysis method based on retrieval enhancement generation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0073] To further illustrate each embodiment, the present invention provides drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. By referring to these contents, ordinary technicians in this field should be able to understand other possible implementation methods and advantages of the present invention. The components in the figures are not drawn to scale, and similar component symbols are generally used to represent similar components.

[0074] According to an embodiment of the present invention, a method for analyzing industrial talent big data based on retrieval enhancement generation is provided.

[0075] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1 As shown, the industrial talent big data analysis method based on retrieval enhancement generation according to an embodiment of the present invention includes:

[0076] S1. Based on industry talent data and combined with search enhancement generation technology, an industry talent intelligence map is constructed to quantify talent flow data and collaboration density data.

[0077] In one embodiment, based on industry talent data and combined with search-enhanced generation technology, an industry talent intelligence map is constructed to quantify talent flow data and collaboration density data, including:

[0078] Obtain standardized and structured industrial talent data (including basic information, skill tags, work experience, project experience, cooperative relationships, etc.); use knowledge graph construction tools to convert industrial talent data into industrial talent intelligence maps, where nodes represent entities including talents, institutions, skills and projects, and edges represent relationships between entities, such as employment relationships between talents and institutions, cooperative relationships between talents, and mastery relationships between talents and skills; use retrieval enhancement generation technology to convert information in the industrial talent intelligence map into a computable vector representation, and learn the relationship patterns between nodes to enhance the industrial talent intelligence map; quantify talent flow data and collaboration density data through the industrial talent intelligence map.

[0079] We obtain industry talent data through recruitment websites, company websites, industry reports, and talent databases provided by partner companies. We use natural language processing techniques to process unstructured data, such as textual descriptions of job responsibilities and project descriptions, extract key information, and convert it into structured data. For example, we extract information such as project name, project time, project members, and project achievements from project descriptions. We use the Neo4j graph database to construct an industry talent intelligence graph, and use a pre-trained language model based on Transformer to perform representation learning on the industry talent intelligence graph, incorporating the learned latent relationships into the graph.

[0080] In one embodiment, quantifying talent flow data and collaboration density data through the industry talent intelligence map includes:

[0081] Based on the work experience information of talents in the industrial talent intelligence map, the mobility frequency of each talent, the talent mobility rate of employers and the talent mobility data between specific employers are calculated; based on the talent cooperation relationship data in the industrial talent intelligence map, the number of collaborations between talents is calculated, and a weighted calculation is performed according to the strength of the cooperation relationship to obtain the collaboration density data.

[0082] The frequency of talent turnover is determined by calculating the number of times talent has changed jobs over the past three years. The turnover rate for each employer is calculated by including both voluntary and involuntary resignations. Talent mobility data between specific employers is obtained by counting the number of talent moves between competitors within a specific industry. Joint participation in a project and a key role within it constitutes a collaborative relationship. The strength of the collaborative relationship is determined by the duration and outcomes of the collaboration. For example, a joint project participation of one year is weighted 1; a joint project participation of three years is weighted 3; a joint publication of a high-impact paper is weighted 2; and a joint publication of a standard paper is weighted 1.

[0083] S2. Based on the industry talent intelligence map and the constructed nonlinear diffusion equation, quantify the impact of talent loss and collaboration disruption on employers; use the evaluation algorithm to evaluate the centrality of talent nodes in the industry talent intelligence map.

[0084] In one embodiment, based on the industry talent intelligence map and the constructed nonlinear diffusion equation, quantifying the impact of talent loss and collaboration disruption on employers includes:

[0085] Using pre-set talent loss and collaboration interruption judgment rules, relevant node and edge information is extracted from the industry talent intelligence map to form a data set of talent loss and collaboration interruption events; according to the characteristics and impact range of talent loss and collaboration interruption events, the parameters of the nonlinear diffusion equation are set, including the diffusion coefficient and attenuation coefficient; the nonlinear diffusion equation is solved using numerical methods (such as the finite difference method), and the time is discretized to gradually calculate the degree of influence of each node at each time step; the influence degree of nodes related to employers is aggregated, and the quantitative value of the impact of talent loss and collaboration interruption on employers is calculated.

[0086] In specific applications, we use employee voluntary resignation as the criterion, examining resignations within the past year. We also use the termination of collaborative projects as the criterion, setting a threshold for collaboration interruption when the frequency falls below once a month. We consider all employees of an employer as nodes related to the employer, averaging the impact of all employees to determine the final impact on the employer.

[0087] In one embodiment, the nonlinear diffusion equation is expressed as:

[0088]

[0089] Where D ij represents the collaboration density weight (diffusion coefficient) between node i and node j. The collaboration density is calculated by weighting the number of collaborations and project duration. ij represents the weight of edge (i, j) in the industrial talent intelligence map; N(i) represents the set of neighboring nodes of node i (directly cooperating talents or related institutions); C i (t) represents the impact value of node i at time t, C j (t) represents the impact value of node j at time t; α i represents the talent loss rate (attenuation coefficient) of the employer where node i is located; S i Represents the source term, which is determined by the node centrality score.

[0090] Among them, D ij ·∑ j∈N(i) w ij [C j (t)-C i (t)] represents the diffusion term, and its impact propagation rate is positively correlated with the collaboration density. The more frequent the collaboration (D ij The larger the α, the faster the influence spreads between nodes. i C i (t) represents the attenuation term, and the unit with high talent loss rate (α i When core talent leaves, the source value is larger and the initial impact is stronger.

[0091] In one embodiment, according to the characteristics and impact range of the brain drain and collaboration interruption events, setting the parameters of the nonlinear diffusion equation includes:

[0092] Based on the characteristic vectors of talent loss and collaboration interruption events, talent flow data, collaboration density data and industrial talent intelligence maps, and using network analysis methods (such as calculating indicators such as degree centrality and closeness centrality of event nodes) combined with expert knowledge, the impact range of talent loss and collaboration interruption events is evaluated; a machine learning model (such as a regression model or a classification model) is used to learn the relationship between the characteristic vectors and impact range of talent loss and collaboration interruption events and the parameters of the nonlinear diffusion equation, and output the diffusion coefficient and attenuation coefficient of each talent loss and collaboration interruption event.

[0093] Use numbers 1-5 to represent job levels, with 1 representing the lowest level and 5 representing the highest level. Use multi-label classification to represent skill types: use numbers to represent years of work experience. Use numbers to represent the number of people involved; use numbers 1-3 to represent the importance of the project, with 1 representing low importance and 3 representing high importance. Calculate the degree centrality of the event node. The higher the degree centrality, the more connections the node has in the network and the wider the impact. In conjunction with domain experts, evaluate the impact range of the event and use the expert evaluation results as one of the input features of the model. Specifically, use the support vector regression model to predict the diffusion coefficient and attenuation coefficient. Collect event data that occurred in the past year to form a training dataset.

[0094] In one embodiment, using an evaluation algorithm to evaluate the centrality of talent nodes in an industry talent intelligence graph includes:

[0095] Using the selected centrality evaluation algorithm, the weighted degree centrality, weighted proximity centrality and weighted betweenness centrality of each node in the industrial talent intelligence map are calculated; after summing up, the weighted centrality score of each node is obtained; the weighted centrality score of each node is normalized, and the node and the corresponding weighted centrality score are stored in a table; among them, when calculating the weighted degree centrality, the collaboration density of the node is used as the weight of the edge in the industrial talent intelligence map, and the higher the collaboration density, the greater the weight of the edge, which affects the centrality score of the node; when calculating the weighted proximity centrality, the talent flow data between nodes is used as the weight of the distance, the more frequent the talent flow, the closer the distance, which affects the centrality score of the node; when calculating the weighted betweenness centrality, the sum of the collaboration density of the nodes on the shortest path through the node is used as the weight of the path, the higher the collaboration density, the greater the weight of the path, which affects the centrality score of the node; the edge weight, distance weight and path weight are normalized.

[0096] Among them, the centrality evaluation algorithm is used to identify important nodes in the graph, that is, to quantify the importance of nodes through different indicators, such as the number of connections of the node, the position of the node in the network, and the influence of the node on information dissemination. The weighted degree centrality calculation formula is:

[0097]

[0098] Where, WDC(v i ) represents node v i The weighted degree centrality of n represents the total number of nodes in the graph, that is, the total number of talents, and c′ ij Represents node v i and v j The normalized weight of the edge between them represents the collaboration density between them.

[0099] The formula for calculating weighted closeness centrality is:

[0100]

[0101] Where, WCC(v i ) represents node v i The weighted closeness centrality of ij Represents node v i and v j The normalized distance between them is calculated based on the talent flow data. The more frequent the talent flow, the closer the distance.

[0102] The formula for calculating weighted betweenness centrality is:

[0103]

[0104] Where, WBC(v i ) represents node v i The weighted betweenness centrality of , s and t represent the weighted betweenness centrality of the graph except v i For any two nodes other than st (v i ) represents the number of shortest paths from node s to node t. The path length is calculated by the sum of the normalized distances of all edges on the path, and the normalized distance of the edges is obtained from the talent flow data, σ st Indicates that the shortest path from node s to node t passes through node v i The number of paths.

[0105] S3. Based on a multi-objective optimization framework, we search for the optimal talent management strategy by combining the quantitative results of the impact of talent loss and collaboration disruption on employers and the results of talent node centrality assessment.

[0106] In one embodiment, based on a multi-objective optimization framework, combined with the quantitative results of the impact of talent loss and collaboration disruption on employers and the talent node centrality evaluation results, the search for the optimal talent management strategy includes:

[0107] The calculated quantitative impact of talent loss and collaboration disruption on the employer and the weighted centrality scores of talent nodes are normalized and used as benchmarks for the total loss from talent loss and collaboration disruption and the total value of the talent team. A benchmark for the total cost of talent management is set. An objective function for the multi-objective optimization problem is constructed, with resource efficiency as a constraint. The algorithm parameters are set to obtain a configured multi-objective optimization algorithm, and decision variables are defined, including the standardized number of employees to be hired, the training investment for each employee, and the incentive investment for each employee. The configured multi-objective optimization algorithm (such as NSGA-II) is used to search the solution space for the multi-objective optimization problem. At each iteration, the total loss from talent loss and collaboration disruption, the total cost of talent management, and the total value of the talent team are calculated based on the current values ​​of the decision variables. The current solution is evaluated by substituting the total loss from talent loss and collaboration disruption, the total cost of talent management, and the total value of the talent team into the objective function. The optimal solution is selected, and the corresponding decision variable values ​​are converted into the optimal talent management strategy, including the recruitment strategy, training strategy, and incentive strategy. Recruitment strategies include recruitment channels, number of recruits, recruitment budget, etc.; training strategies include training duration, training content, training methods, etc.; incentive strategies include salary levels, promotion mechanisms, reward systems, etc.

[0108] In one embodiment, the objective function and constraints of the multi-objective optimization problem are:

[0109] Minimize L=Lo+Co-Va

[0110] Co≤B;

[0111] In the formula, L represents the total loss, and the goal is to minimize the total loss; Lo represents the total loss value caused by talent loss and collaboration interruption; Co represents the total cost of talent management; Va represents the total value of the talent team; B represents the budget value; Among them, by establishing linear relationships between the values ​​of decision variables and the total loss value caused by talent loss and collaboration interruption, the total cost of talent management, and the total value of the talent team, the current total loss value caused by talent loss and collaboration interruption, the current total cost of talent management, and the total value of the current talent team are calculated.

[0112] Furthermore, weights are assigned based on the importance of the total losses from talent loss and disrupted collaboration, the total cost of talent management, and the total value of the talent team. For example, the objective function is L = 0.5Lo + 0.2Co - 0.3Va. Budget B is determined based on the company's annual budget and human resources plan. When establishing linear relationships between the decision variables and the total losses from talent loss and disrupted collaboration, the total cost of talent management, and the total value of the talent team, the corresponding coefficients in the linear relationships are calculated. Specifically, the coefficients are obtained through statistical analysis of historical data.

[0113] S4. Use the projection pursuit forest algorithm and influencing factor data to predict the effectiveness and potential risks of the optimal talent management strategy.

[0114] In one embodiment, using the projection pursuit forest algorithm and influencing factor data, the effectiveness and potential risks of the optimal talent management strategy are predicted, including:

[0115] Based on the optimal talent management strategy, the optimal strategy feature vector is obtained; based on the influencing factor data, the influencing factor feature vector is obtained; the historical talent management strategy data and the corresponding influencing factor data and strategy effectiveness results are used as training data sets; the training data set is used to train the projection tracking forest model to learn the nonlinear relationship between strategy characteristics, influencing factors and strategy effectiveness; the trained projection tracking forest model is used to predict the effectiveness of the current optimal talent management strategy under given influencing factors, and the predicted effectiveness probability is output; when the effectiveness probability is lower than the preset threshold, the most important influencing factors are found according to the feature importance analysis technology of the projection tracking forest model, and set as potential risks.

[0116] The optimal talent management strategy is converted into numerical features using one-hot encoding. Influencing factors also include the macroeconomic environment, industry competition (e.g., number of competitors, market share), and internal company factors (e.g., corporate culture, employee satisfaction). These factors are also converted into numerical features. The effectiveness of relevant strategies is evaluated using indicators such as talent turnover rate, employee performance, and team innovation capabilities.

[0117] By combining projection pursuit technology with the concept of random forests, a projection pursuit forest is constructed. This model improves prediction accuracy by finding the projection direction that best reflects the data structure. The training parameters of the projection pursuit forest model include the number of trees and the dimensions of the projection matrix. A validity probability threshold is also set. If the predicted validity probability falls below the threshold, the most important influencing factors are identified as potential risks based on the results of feature importance analysis.

[0118] In order to facilitate understanding of the above technical solutions of the present invention, the working principle of the present invention in actual process is described in detail below.

[0119] The basic information, skill tags, work experience, project experience, and cooperative relationships of 50 employees of a technology company were obtained from recruitment websites, corporate websites, industry reports, and talent databases provided by partner companies. For textual descriptions of job responsibilities and project introductions, natural language processing techniques (such as named entity recognition and relationship extraction) were used to extract key information and convert it into structured data. The Neo4j graph database was used to construct an industrial talent intelligence map, where nodes represent talents, institutions, skills, and projects, and edges represent the relationships between these entities. For example, it was found that there were cooperative relationships between 30 pairs of talents, and the average cooperation time for each pair of talents was 2 years. The Transformer-based pre-trained model BERT was combined with a graph neural network to perform representation learning on the industrial talent intelligence map, obtaining a vector representation of each node.

[0120] Over the past three years, the tech company has had 10 employees leave, 7 voluntarily and 3 involuntarily, resulting in a calculated talent turnover rate of 20%. Collaboration within the core team is frequent, averaging more than once a month, while cross-departmental collaboration is less frequent, averaging only once a year. The intensity of collaboration is weighted based on the duration and outcomes of the collaboration to generate an overall collaboration density score. A regression model is used to learn the relationship between the characteristic vectors and impact ranges of talent loss and collaboration disruption events and the parameters of the nonlinear diffusion equation. The model then outputs the diffusion coefficient and attenuation coefficient for each talent loss and collaboration disruption event. Based on the nonlinear diffusion equation parameters, the finite difference method is used to gradually calculate the degree of impact of each node at each time step, which is then aggregated to obtain the overall impact on the employer.

[0121] The total losses from talent loss and disrupted collaboration, the total cost of talent management, and the total value of the talent team are normalized and used as benchmarks. After setting an annual talent management budget, the NSGA-II algorithm is used to search for the optimal talent management strategy, including the selection of recruitment channels, the determination of the number of hires, training investment, and incentives. For example, it recommends increasing the number of hires to 15, increasing the training period for new hires to six months per person, and improving the salary and reward system in the incentive system to reduce talent turnover and improve team collaboration efficiency.

[0122] Based on the optimal talent management strategy, we obtain a strategic feature vector and collect data on influencing factors such as the macroeconomic environment, industry competition, and corporate culture. We use historical data to train a projection tracking forest model to predict the effectiveness of the current optimal talent management strategy and identify potential risk points.

[0123] In practical applications, through the innovative use of large models, event graphs, and search-enhanced generation technologies, we have developed industrial talent big data analysis software. This software intelligently analyzes massive amounts of talent industry activity data. By digitizing, mapping, and modeling the potential collaborative relationships between employers and talent across resources, capabilities, and markets, and through automated intelligent processing and matching, it forms an industrial talent intelligence map. This software directly provides comprehensive technical support and assurance for human resources departments in industrial enterprises and research institutes in areas such as talent evaluation, talent recruitment, and optimized human resource allocation.

[0124] In summary, we leverage retrieval-enhanced generative technology to construct an industrial talent intelligence graph and use it to quantify talent mobility and collaboration density. This technology automatically learns relationship patterns between nodes from large-scale data and incorporates them into the graph, thereby constructing a more complete and accurate industrial talent intelligence graph. The impact of talent loss and collaboration disruption on employers is quantified using a nonlinear diffusion equation. This equation models the propagation of the effects of talent loss and collaboration disruption within organizational networks, enabling a more comprehensive assessment of their impact on employers and a deeper understanding of the potential risks of talent loss and collaboration disruption. An evaluation algorithm is used to assess the centrality of talent nodes in the industrial talent intelligence graph. The quantified impact of talent loss and collaboration disruption on employers and the assessed centrality of talent nodes serve as benchmarks for a multi-objective optimization process. This, combined with the total cost of talent management, allows for a more comprehensive and effective assessment of talent value, balancing short-term costs and long-term value, and improving the effectiveness of talent management strategies. A multi-objective optimization framework and the projection pursuit forest algorithm are combined to search for and predict the effectiveness and potential risks of optimal talent management strategies. This multi-objective optimization framework simultaneously considers multiple objectives and searches for the optimal talent management strategy. Combined with the projection pursuit forest algorithm, the effectiveness and potential risks of the optimal talent management strategy under different influencing factors are predicted, thereby helping employers choose more robust and effective strategies.

[0125] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for analyzing industrial talent big data based on retrieval enhancement generation, characterized in that: include: Based on industry talent data and combined with search-enhanced generation technology, we build an industry talent intelligence map to quantify talent flow data and collaboration density data; Based on the industry talent intelligence map and the constructed nonlinear diffusion equation, the impact of talent loss and collaboration disruption on employers is quantified; Using the evaluation algorithm, the centrality of talent nodes in the industry talent intelligence map is evaluated; Based on a multi-objective optimization framework, we search for the optimal talent management strategy by combining the quantitative results of the impact of talent loss and collaboration disruption on employers and the results of talent node centrality assessment. The projection pursuit forest algorithm and influencing factor data are used to predict the effectiveness and potential risks of the optimal talent management strategy.

2. The method for analyzing industrial talent big data based on retrieval enhancement generation according to claim 1 is characterized in that: The above-mentioned industrial talent intelligence map is constructed based on industrial talent data and combined with search enhancement generation technology to quantify talent flow data and collaboration density data, including: Obtain standardized and structured industry talent data; use knowledge graph construction tools to convert industry talent data into an industry talent intelligence graph, where nodes represent entities such as talent, institutions, skills, and projects, and edges represent relationships between entities; Utilize retrieval-enhanced generation technology to transform the information in the industry talent intelligence map into a computable vector representation and learn the relationship patterns between nodes to enhance the industry talent intelligence map; Quantify talent flow data and collaboration density data through industrial talent intelligence maps.

3. The method for analyzing industrial talent big data based on retrieval enhancement generation according to claim 2 is characterized in that: The talent flow data and collaboration density data quantified through the industry talent intelligence map include: Based on the employment history information of talents in the industry talent intelligence map, calculate the mobility frequency of each talent, the talent mobility rate of employers, and the talent mobility data between specific employers; Based on the talent cooperation relationship data in the industrial talent intelligence map, the number of collaborations between talents is calculated, and weighted calculation is performed according to the strength of the cooperation relationship to obtain the collaboration density data.

4. The method for analyzing industrial talent big data based on search-enhanced generation according to claim 1 is characterized in that: The impact of talent loss and collaboration disruption on employers, quantified based on the industry talent intelligence map and the constructed nonlinear diffusion equation, includes: Using pre-set rules for judging talent loss and collaboration interruption, relevant node and edge information is extracted from the industry talent intelligence map to form a data set of talent loss and collaboration interruption events; According to the characteristics and impact range of brain drain and collaboration disruption events, the parameters of the nonlinear diffusion equation are set, including the diffusion coefficient and the attenuation coefficient; Use numerical methods to solve the nonlinear diffusion equation, discretize time, and gradually calculate the degree of influence of each node at each time step; The impact levels of nodes related to employers are aggregated, and the quantitative impact of talent loss and collaboration disruption on employers is calculated.

5. The method for analyzing industrial talent big data based on retrieval enhancement generation according to claim 4 is characterized in that: The expression of the nonlinear diffusion equation is: Where D ij represents the collaboration density weight between node i and node j; w ij Represents the weight of edge (i, j) in the industrial talent intelligence graph; N(i) represents the set of neighbor nodes of node i; C i (t) represents the impact degree of node i at time t, C j (t) represents the impact value of node j at time t; α i represents the talent loss rate of the employer where node i is located; S i Represents a source term.

6. The method for analyzing industrial talent big data based on search-enhanced generation according to claim 4 is characterized in that: The parameters of the nonlinear diffusion equation are set according to the characteristics and impact range of the brain drain and collaboration interruption events, including: Based on the feature vectors of brain drain and collaboration disruption events, talent flow data, collaboration density data, and industry talent intelligence maps, and using network analysis methods combined with expert knowledge, the impact scope of brain drain and collaboration disruption events is assessed; A machine learning model is used to learn the relationship between the feature vectors and impact ranges of talent loss and collaboration disruption events and the parameters of the nonlinear diffusion equation, and output the diffusion coefficient and attenuation coefficient of each talent loss and collaboration disruption event.

7. The method for analyzing industrial talent big data based on search-enhanced generation according to claim 1 is characterized in that: The evaluation algorithm used to evaluate the centrality of talent nodes in the industry talent intelligence map includes: Using the selected centrality evaluation algorithm, calculate the weighted degree centrality, weighted closeness centrality, and weighted betweenness centrality of each node in the industrial talent intelligence map; then sum them up to obtain the weighted centrality score of each node; Normalize the weighted centrality score of each node and store the nodes and their corresponding weighted centrality scores in a table; Among them, when calculating weighted degree centrality, the collaboration density of the node is used as the weight of the edge in the industrial talent intelligence map, and the higher the collaboration density, the greater the weight of the edge; When calculating weighted proximity centrality, the talent flow data between nodes is used as the weight of the distance. The more frequent the talent flow, the closer the distance. When calculating weighted betweenness centrality, the sum of the collaboration densities of the nodes on the shortest path through the node is used as the weight of the path. The higher the collaboration density, the greater the weight of the path. Normalize the edge weights, distance weights, and path weights.

8. The method for analyzing industrial talent big data based on search-enhanced generation according to claim 1 is characterized in that: Based on the multi-objective optimization framework, combined with the quantitative results of the impact of talent loss and collaboration disruption on employers and the results of talent node centrality evaluation, the search for the optimal talent management strategy includes: The calculated impact of talent loss and collaboration disruption on employers and the weighted centrality scores of talent nodes are normalized and used as benchmarks for the total loss from talent loss and collaboration disruption and the total value of the talent team; this also sets a benchmark for the total cost of talent management. Construct the objective function of the multi-objective optimization problem and use resource efficiency as a constraint. By setting the parameters of the multi-objective optimization algorithm, a configured multi-objective optimization algorithm is obtained, and the decision variables are defined, including the standardized number of employees to be hired, the training investment for each employee, and the incentive investment for each employee. Use the configured multi-objective optimization algorithm to search the solution space of the multi-objective optimization problem. At each iteration, combine the current values ​​of the decision variables to calculate the total loss caused by the current talent loss and collaboration interruption, the total cost of current talent management, and the total value of the current talent team. By substituting the total loss value caused by current talent loss and collaboration interruption, the total cost of current talent management, and the total value of the current talent team into the objective function, the pros and cons of the current solution are evaluated; Select the optimal solution and convert the values ​​of the corresponding decision variables into the optimal talent management strategy, including recruitment strategy, training strategy and incentive strategy.

9. The method for analyzing industrial talent big data based on search-enhanced generation according to claim 8 is characterized in that: When constructing the objective function of the multi-objective optimization problem, a linear relationship is established between the values ​​of the decision variables and the total loss value caused by talent loss and collaboration interruption, the total cost of talent management, and the total value of the talent team, so as to calculate the total loss value caused by the current talent loss and collaboration interruption, the total cost of the current talent management, and the total value of the current talent team.

10. The method for analyzing industrial talent big data based on search-enhanced generation according to claim 1, characterized in that: The use of the projection pursuit forest algorithm and influencing factor data to predict the effectiveness and potential risks of the optimal talent management strategy includes: Based on the optimal talent management strategy, the optimal strategy feature vector is obtained; based on the influencing factor data, the influencing factor feature vector is obtained; Use historical talent management strategy data, corresponding influencing factor data, and strategy effectiveness results as training data sets; Use the training dataset to train the projection pursuit forest model to learn the nonlinear relationship between policy features, influencing factors and policy effectiveness; Use the trained projection pursuit forest model to predict the effectiveness of the current optimal talent management strategy under given influencing factors and output the predicted effectiveness probability; When the probability of effectiveness is lower than the preset threshold, the most important influencing factors are found and set as potential risks based on the feature importance analysis technology of the projection pursuit forest model.