Talent flow influence factor identification method and device, storage medium and electronic equipment
By combining causal machine learning models and linear regression models, the problem of inaccurate identification of influencing factors on the mobility of scientific and technological innovation talents was solved, and a systematic analysis and accurate identification of multiple factors were achieved.
Patent Information
- Application Number
- CN202510882200.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies lack a systematic approach in analyzing the factors influencing the flow of scientific and technological innovation talents, making it difficult to comprehensively consider the combined effects of multiple factors, resulting in inaccurate identification.
We use a multi-effects model based on causal machine learning to estimate the approximate hidden confounding factors, and combine it with a pre-defined linear regression model to identify the key factors in the flow of scientific and technological innovation talents.
By establishing a multiple causal effect estimation method for alternative confounding factors, we can effectively control confounding bias, identify the impact of single factors on the flow of scientific and technological innovation talents, and provide a systematic analysis reference.
Smart Images

Figure CN120996328A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, storage medium, and electronic device for identifying factors influencing talent mobility. Background Technology
[0002] A scientific analysis of the factors influencing the flow of scientific and technological innovation talent is fundamental to understanding the patterns of this flow and systematically constructing macro-management strategies. Current research on the factors influencing the flow of scientific and technological innovation talent can be divided into three levels: macro, meso, and micro. Most macro-level studies focus on summarizing and generalizing the influencing factors, lacking systematic analysis. Existing analyses often start from single-domain theories, rarely analyzing the combined impact of various factors, and are also insufficient in terms of comprehensive benefit analysis. The influencing factors and internal mechanisms of the flow of scientific and technological innovation talent are complex. Traditional descriptive statistical methods for talent distribution and flow have played a role in understanding talent flow trends, but they are insufficient for analyzing the complex internal influencing mechanisms of talent flow at a systemic level. The flow of scientific and technological innovation talent is influenced by multiple factors. Identifying the key factors influencing the flow of scientific and technological innovation talent will facilitate systematic analysis of the influencing mechanisms by management departments.
[0003] Therefore, it is urgent to introduce new methods and perspectives to systematically analyze the influencing factors of the flow of scientific and technological innovation talents, so as to provide effective reference for policy formulation and decision-making. Summary of the Invention
[0004] In view of this, the present invention provides a method, device, storage medium and electronic device for identifying factors influencing talent mobility, the main purpose of which is to solve the problem of inaccurate identification of factors influencing the mobility of scientific and technological talents.
[0005] To address the aforementioned issues, this application provides a method for identifying factors influencing talent mobility, including:
[0006] Obtain observational data on the factors influencing the mobility of the target talent and observational data on the mobility outcomes of the target talent;
[0007] An approximate hidden confounding factor estimate is obtained by using a multiple-effects model based on causal machine learning to estimate the observed data of the influencing factors;
[0008] Based on the observed data of the influencing factors, the observed data of the flow results, and the estimated value of the substitution and confounding factors, a preset linear regression model is used to identify the influencing factors of talent flow corresponding to the target talent.
[0009] Optionally, the step of using a multi-effects model based on causal machine learning to approximate the hidden confounding factor estimation of the observed data of the influencing factors, and obtaining alternative confounding factor estimates, specifically includes:
[0010] The influencing factors are identified using a pre-defined latent variable model, resulting in confounding factors.
[0011] The maximum likelihood estimation (EM) iterative algorithm is used to calculate the confusion factor to obtain an estimated alternative confusion factor.
[0012] Optionally, before identifying the observed data of the influencing factors using a preset latent variable model, the method further includes: constructing a preset latent variable model;
[0013] The construction of the pre-defined latent variable model specifically includes:
[0014] Obtain multiple sample data on talent mobility;
[0015] The sample data of multiple talent mobility samples are divided into an observation dataset and a validation dataset.
[0016] The initial latent variable model is trained based on the observed dataset to obtain a latent variable model that meets the preset conditions.
[0017] The latent variable model is validated using the validation dataset. When the predicted test value of the latent variable model is greater than a preset threshold, the latent variable model is determined as the preset latent variable model.
[0018] Optionally, the step of training the initial latent variable model based on the observed dataset to obtain a latent variable model that meets preset conditions specifically includes:
[0019] Step 1: Randomly initialize the model parameters of the initial latent variable model. The model parameters include a weight matrix and noise variance, which characterize the linear influence of the latent variables on the causal variables.
[0020] Step 2: Calculate the posterior distribution of latent variables for each of the observation datasets based on the model parameters, and obtain the posterior covariance and posterior mean corresponding to each of the observation datasets;
[0021] Step 3: Update the model parameters based on the posterior covariance and the posterior mean of each parameter;
[0022] Step 4: If the current iteration number is less than the preset threshold, repeat steps 2 to 3 based on the updated model parameters until the current iteration number equals the preset threshold. Then, construct the model based on the updated model parameters and the initial latent variable model to obtain the latent variable model.
[0023] Optionally, the method further includes:
[0024] When the noise variance in the updated model parameters is less than a preset threshold, the model is constructed based on the updated model parameters and the initial latent variable model to obtain the latent variable model.
[0025] Optionally, the step of validating the latent variable model using the validation dataset, and determining the latent variable model as the preset latent variable model when the predicted test value of the latent variable model is greater than a preset threshold, specifically includes:
[0026] Obtain a validation dataset with the same confusion factor;
[0027] Based on the confusion factor and the model parameters of the latent variable model, multiple prediction datasets corresponding to the validation dataset are obtained.
[0028] The Monte Carlo averaging algorithm is used to calculate and process the validation dataset and each of the prediction datasets respectively to obtain a first statistic corresponding to the validation dataset and a second statistic corresponding to each of the prediction datasets.
[0029] Based on the first statistic and each of the second statistics, calculations are performed to obtain the predicted test value corresponding to the same confounding factor, so as to obtain the predicted test value corresponding to each confounding factor respectively.
[0030] When each of the predicted test values is greater than a preset threshold, the latent variable model is determined as the preset latent variable model.
[0031] Optionally, the identification of talent mobility influencing factors based on the observed data of the influencing factors, the observed data of the mobility outcomes, and the estimated value of the substitution confounding factor, using a preset linear regression model, to obtain the identification results of talent mobility influencing factors corresponding to the target talent, specifically includes:
[0032] Based on the observed data of the influencing factors, the observed data of the flow results, and the estimated value of the substitution confounding factor, the preset linear regression result model is used to calculate and process the data to obtain the causal effect weight coefficient vector corresponding to the observed data of the influencing factors and the weight coefficient vector corresponding to the estimated value of the substitution confounding factor.
[0033] By filtering each causal effect weight coefficient in the causal effect weight coefficient vector, the target causal effect weight coefficient is obtained.
[0034] Based on the target causal effect weight coefficient, the target influencing factors affecting talent mobility are determined to obtain the identification results of talent mobility influencing factors corresponding to the target talent.
[0035] To address the aforementioned issues, this application provides a device for identifying factors influencing talent mobility, comprising:
[0036] The acquisition module is used to acquire observational data on the factors influencing the mobility of the target talent and observational data on the mobility results of the target talent.
[0037] The identification module is used to perform approximate hidden confounding factor estimation on the observed data of the influencing factors using a multi-effect model based on causal machine learning, so as to obtain the estimated value of the alternative confounding factor.
[0038] The identification module is used to identify the talent mobility influencing factors corresponding to the target talent by using a preset linear regression model based on the observed data of the influencing factors, the observed data of the mobility results, and the estimated value of the substitution confounding factor.
[0039] To address the aforementioned issues, this application provides a storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned method for identifying factors influencing talent mobility.
[0040] To address the aforementioned problems, this application provides an electronic device, comprising at least a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program in the memory, implements the steps of the aforementioned method for identifying factors influencing talent mobility.
[0041] The beneficial effects of this application are as follows: This application analyzes the influencing factors of the flow of scientific and technological innovation talents from a system level, establishes a multiple causal effect estimation method with alternative confounding factors, effectively controls confounding bias, and the model is simple, easy to implement and conforms to prediction verification. A linear regression model is established on the enhanced dataset to effectively consider the comprehensive influence of influencing factors. The model can identify the impact of single factors on the flow of scientific and technological innovation talents and effectively discover the factors that truly affect the flow of scientific and technological innovation talents.
[0042] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0043] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0044] Figure 1 A flowchart illustrating a method for identifying factors influencing talent mobility provided in an embodiment of this application is shown.
[0045] Figure 2 A flowchart illustrating a method for identifying factors influencing talent mobility, as provided in another embodiment of this application, is shown.
[0046] Figure 3 A structural block diagram of a talent mobility influencing factor identification device provided in another embodiment of this application is shown. Detailed Implementation
[0047] Various embodiments and features of this application are described herein with reference to the accompanying drawings.
[0048] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this application will be apparent to those skilled in the art.
[0049] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.
[0050] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.
[0051] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application.
[0052] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.
[0053] Specific embodiments of this application are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this application, which can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure the application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in a variety of substantially any suitable detailed structures.
[0054] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.
[0055] This application provides a method for identifying factors influencing talent mobility, such as... Figure 1 As shown, it includes:
[0056] Step S101: Obtain observational data on the influencing factors affecting the mobility of the target talent and observational data on the mobility results of the target talent;
[0057] In the specific implementation process of this step, the observed data of influencing factors include data on the economy, policies, culture, social environment and network evolution characteristics; the observed data of flow results include observed data on talent inflow and talent outflow.
[0058] Step S102: Use a multi-effects model based on causal machine learning to estimate the hidden confounding factor in the observed data of the influencing factors to obtain the estimated value of the alternative confounding factor;
[0059] In this step, a pre-defined latent variable model is used to identify the observed data of the influencing factors to obtain the confounding factors; the maximum likelihood estimation (EM) iterative algorithm is used to calculate the confounding factors to obtain the estimated values of the alternative confounding factors.
[0060] Step S103: Based on the observed data of the influencing factors, the observed data of the flow results, and the estimated value of the substitution and confusion factors, a preset linear regression result model is used to identify the talent flow influencing factors corresponding to the target talent.
[0061] In this step, the observed data of influencing factors, the observed data of flow results, and the estimated value of the substitution confounding factor are used to calculate and process the results using the preset linear regression model to obtain the causal effect weight coefficient vector corresponding to the observed data of influencing factors and the weight coefficient vector corresponding to the estimated value of the substitution confounding factor. The causal effect weight coefficients in the causal effect weight coefficient vector are then filtered to obtain the target causal effect weight coefficient. Based on the target causal effect weight coefficient, the target influencing factors affecting talent flow are determined to obtain the identification result of talent flow influencing factors corresponding to the target talent.
[0062] This application analyzes the influencing factors of the flow of scientific and technological innovation talents from a systemic perspective. By establishing a multiple causal effect estimation method with alternative confounding factors, it effectively controls confounding bias. Furthermore, the model is simple, easy to implement, and conforms to prediction verification. A linear regression model is established on the enhanced dataset to effectively consider the comprehensive impact of influencing factors. The model can identify the impact of single factors on the flow of scientific and technological innovation talents and effectively discover the factors that truly affect the flow of scientific and technological innovation talents.
[0063] Another embodiment of this application provides a different method for identifying factors influencing talent mobility, such as... Figure 2 As shown, it includes:
[0064] Step S201: Obtain observational data on the influencing factors affecting the mobility of the target talent and observational data on the mobility results of the target talent;
[0065] In the specific implementation of this step, the observed data of influencing factors include data on the economy, policies, culture, social environment, and network evolution characteristics; the observed data of flow outcomes include data on talent inflow and outflow. Given N outcomes of the flow of scientific and technological innovation talents, each outcome is represented by a vector, where m represents the number of influencing factors (e.g., economy, policies, culture, social environment, and network evolution characteristics), and each influencing factor corresponds to both positive and negative flow states, denoted by α. j+ α j- Representing α j Positive and negative flow results in this aspect; therefore, α = (α 1+ α 1- , ..., α m+ α m- Suppose Y is an outcome variable of the flow of science and technology innovation talent to a certain region. The goal of this model is to estimate the average causal effect on a specific outcome variable Y. Within the latent outcome framework, this is equivalent to estimating the following latent outcome equation:
[0066] y(α):R 2m →R
[0067] y(α) represents the potential outcome at treatment level α, R 2m R represents the input space, and R represents the output space, which is the space in which the result variable Y can take values.
[0068] Step S202: Construct a pre-defined latent variable model;
[0069] In the specific implementation process, the construction of the preset latent variable model includes the following steps:
[0070] Step 1: Randomly initialize the model parameters of the initial latent variable model. The model parameters include a weight matrix and noise variance, which characterize the linear influence of the latent variables on the causal variables.
[0071] A latent variable model p(z,α) is constructed for multiple influencing factors α. 1+ α 1- , ..., α m+ α m- The mathematical expression of the initial latent variable model, z∈Z, can be expressed as follows: (1)
[0072]
[0073] Where: β represents the distribution parameter of the substitution confusion factor; θ j The model parameters represent the latent variable model; the causal variable α is the factor influencing the model. j The distribution parameters, z i W is a latent variable representing a confounding factor. j The weight matrix used to characterize the linear influence of latent variables on causal variables is as follows: Noise variance. Random initialization. The value of .
[0074] Step 2: Calculate the posterior distribution of latent variables for each of the observation datasets based on the model parameters, and obtain the posterior covariance and posterior mean corresponding to each of the observation datasets;
[0075] θ is estimated using the EM (Expectation Maximization Algorithm) iterative algorithm based on maximum likelihood estimation. j and inference z i The maximum likelihood estimation (EM) algorithm consists of an E-step (Expection-Step) computation and an M-step (Maximization-Step) computation. First, for each observation dataset, the latent variable z for each sample is calculated. i The posterior covariance and posterior mean; the mathematical formula for the E-step calculation can be shown in the following formula (2):
[0076]
[0077] The posterior covariance ∑ z The mathematical formula for calculating it can be shown in the following formula (3):
[0078] Σ z =(W T Σ -1 W+I) -1 (3)
[0079] The posterior mean The mathematical formula for calculating it can be shown in the following formula (4):
[0080]
[0081] Among them, A e Let be the observation vector corresponding to the e-th observation dataset, consisting of causal variables α of 2m influencing factors. j Composition, A ej Let z be the observed value of the j-th causal variable in the observation vector of the e-th sample, and z be the latent variable. e (One for each sample) follows a Gaussian distribution z e ~N(0,I); Let α be the posterior mean of the e-th sample; α is the causal variable of the influencing factors. j The mathematical expression for can be expressed by the following formula (5).
[0082]
[0083] Wherein, the weight matrix W j ∈R 1ⅹK , represents the latent variable z e The linear effect on the j-th causal variable, ∈ j For noise, This represents the noise variance. Observation vector A e The mathematical expression for can be represented by the following formula (6):
[0084] A e =W*z e +∈,∈~N(0,∑) (6)
[0085] Where W∈R 2mⅹK By W j Stacked together, where ∑ is the diagonal covariance matrix, and the diagonal is...
[0086] Step 3: Update the model parameters based on the posterior covariance and the posterior mean of each parameter;
[0087] The maximum likelihood estimation (EM) algorithm takes M steps to process the causal variable α for each influencing factor. j Update parameter W j ,
[0088] For W j The mathematical expression for updating can be represented by the following formula (7):
[0089]
[0090] right The updated mathematical expression can be represented by the following formula (8):
[0091]
[0092] Where M is the number of observation datasets, This is the updated weight matrix of influencing factors; This is the updated noise variance.
[0093] Step 4: If the current iteration number is less than the preset threshold, repeat steps 2 to 3 based on the updated model parameters until the current iteration number equals the preset threshold. Then, construct the model based on the updated model parameters and the initial latent variable model to obtain the latent variable model.
[0094] When the number of iterations reaches the target number, the iteration stops, and the latent variable model is obtained. Alternatively, when the noise variance in the updated model parameters is less than a preset threshold, a model is constructed based on the updated model parameters and the initial latent variable model to obtain the latent variable model.
[0095] Step 5: Validate the latent variable model using the validation dataset. When the predicted test value of the latent variable model is greater than a preset threshold, the latent variable model is determined as the preset latent variable model. This specifically includes the following steps:
[0096] Step 1: Obtain the validation dataset with the same confusion factor;
[0097] Before the predictive test, the dataset is prepared by randomly selecting a subset of multi-faceted flow factor values for each region as the validation set α. i,held The remainder serves as the observation set α. i,obs Principal component analysis for fitting probability utilizes only the observation set, while the predictive test uses the validation set. The predictive test value is calculated by comparing the observed multifactor values with the multifactor values drawn from the fitted predictive distribution. The observation set is used. A probabilistic principal component analysis model is fitted to obtain the latent variable model, thereby obtaining the corresponding model parameters and the confusion factor corresponding to the observation set; the model parameters include the weight matrix W. obs Noise ∈ obs The confounding factor is a latent variable z. obs Multiple random sampling of latent variable z obs samples Samples are randomly drawn from the distribution of latent variables, and these samples are used to generate prediction data. The prediction data is generated by substituting the latent variable samples into the model formula (combining the weight matrix and noise), and is used to simulate real data.
[0098] Step 2: Perform calculations based on the confusion factor and the model parameters of the latent variable model to obtain multiple prediction datasets corresponding to the validation dataset;
[0099] Specifically, based on the formula Generate validation set α i,held Multiple prediction datasets The validation set refers to actual observation data used to verify the predictive ability of the model.
[0100] Step 3: Calculate the data using the Monte Carlo averaging algorithm based on the validation dataset and each of the prediction datasets to obtain a first statistic corresponding to the validation dataset and a second statistic corresponding to each of the prediction datasets.
[0101] Specifically, for the validation set α i,held Calculate the log-conditional probability density of the observed data of the target influencing factors based on the target's latent variables. The mathematical expression can be represented by the following formula (9):
[0102]
[0103] Among them, ||.|| 2 W represents the squared Euclidean distance between vectors, where d is the dimension of the observed data; X , These are the latent variables of the corresponding samples. The weights and noise are calculated based on the log-conditional probability density of the observed data of each influencing factor to obtain the first statistic of the validation dataset; the mathematical formula for calculating the first statistic is as follows (10):
[0104]
[0105] The same steps are used to calculate the second statistic for each prediction dataset. The Monte Carlo method is a technique for approximating statistical measures using random sampling. These measures can be the mean, variance, distance, etc., and are used to measure the characteristics of the data.
[0106] Step 4: Perform calculations based on the first statistic and each of the second statistics to obtain the predicted test value corresponding to the same confusion factor, so as to obtain the predicted test value corresponding to each confusion factor respectively.
[0107] Specifically, after obtaining the statistic t(α) of the true validation set... i,held ) and statistics of multiple forecast data Compare the second statistic of each predicted data point with the first statistic of the true validation set, and calculate the probability that the statistic of the predicted data is less than the statistic of the true validation set, i.e., the predicted test value p. c The mathematical formula for calculating the predicted test value can be shown in the following formula (11):
[0108]
[0109] p c The value is greater than 0 and less than 1.
[0110] Step 5: When each of the predicted test values is greater than the preset threshold, the latent variable model is determined as the preset latent variable model.
[0111] Specifically, the preset threshold p is determined based on the actual situation and different application scenarios. * If the predictive test value p c Greater than the preset threshold p * This indicates that the latent variable model can generate multi-factor values that are similar to the true values in the validation set, thus passing the predictive test.
[0112] Step S203: Identify the observed data of the influencing factors using a preset latent variable model to obtain the confounding factors;
[0113] In this step, the observed data of influencing factors are input into a preset latent variable model. The preset latent variable model identifies the observed data of influencing factors and obtains confounding factors. The confounding factors obtained by the preset latent variable model can be one or more. When there are multiple confounding factors, the maximum likelihood estimation (EM) iterative algorithm is used to calculate the estimated values of alternative confounding factors corresponding to the multiple confounding factors.
[0114] Step S204: The maximum likelihood estimation (EM) iterative algorithm is used to calculate the confusion factor to obtain the estimated value of the alternative confusion factor;
[0115] In this step, the latent variable posterior distribution of the observed data of each influencing factor is calculated based on the model parameters of the latent variable model to obtain the target posterior covariance and target posterior mean corresponding to the observed data of the influencing factors. The target posterior mean is then determined as the estimated value of the alternative confounding factor corresponding to the confounding factor.
[0116] Step S205: Based on the observed data of the influencing factors, the observed data of the flow results, and the estimated value of the substitution confounding factor, the preset linear regression result model is used for calculation and processing to obtain the causal effect weight coefficient vector corresponding to the observed data of the influencing factors and the weight coefficient vector corresponding to the estimated value of the substitution confounding factor.
[0117] In the specific implementation process of this step, the mathematical expression of the preset linear regression result model can be represented by the following formula (12):
[0118] f(a,z)=β T ar T z(12)
[0119] The observed data of influencing factors, the observed data of flow outcomes, and the estimated value of the alternative confounding factor constitute the target dataset for increasing the confounding factor. Based on the target dataset, the preset linear regression result model is substituted into it for calculation and processing to obtain the causal effect weight coefficient vector β corresponding to the observed data of influencing factors and the weight coefficient vector r corresponding to the estimated value of the substitution confounding factor. β records the causal effect of a single factor on talent mobility.
[0120] Step S206: Filter each causal effect weight coefficient in the causal effect weight coefficient vector to obtain the target causal effect weight coefficient;
[0121] In the specific implementation process, this step involves filtering each causal effect weight coefficient in the causal effect weight coefficient vector to obtain the maximum causal effect weight coefficient, and then determining the maximum causal effect weight coefficient as the target causal effect weight coefficient.
[0122] Step S207: Determine the target influencing factors affecting talent mobility based on the target causal effect weight coefficient, so as to obtain the identification results of talent mobility influencing factors corresponding to the target talent.
[0123] In the specific implementation process of this step, the influencing factors corresponding to the target causal effect weight coefficient are determined as target influencing factors. The target influencing factors are the most important factors affecting the target talent, so as to obtain the identification results of talent mobility influencing factors corresponding to the target talent.
[0124] This application obtains observational data on influencing factors affecting the mobility of target talent and observational data on the mobility outcomes of the target talent; constructs a pre-defined latent variable model; uses the pre-defined latent variable model to identify the influencing factor observational data to obtain confounding factors; uses the maximum likelihood estimation (EM) iterative algorithm to calculate the confounding factors to obtain estimated values of alternative confounding factors; uses the pre-defined linear regression model to process the influencing factor observational data, the mobility outcome observational data, and the estimated values of alternative confounding factors to obtain a causal effect weight coefficient vector corresponding to the influencing factor observational data and a weight coefficient vector corresponding to the estimated values of alternative confounding factors; filters each causal effect weight coefficient in the causal effect weight coefficient vector to obtain target causal effect weight coefficients; and determines the target influencing factors affecting talent mobility based on the target causal effect weight coefficients to obtain the talent mobility influencing factor identification results corresponding to the target talent. This application analyzes the influencing factors of the flow of scientific and technological innovation talents from a systemic perspective. By establishing a multiple causal effect estimation method with alternative confounding factors, it effectively controls confounding bias. Furthermore, the model is simple, easy to implement, and conforms to prediction verification. A linear regression model is established on the enhanced dataset to effectively consider the comprehensive impact of influencing factors. The model can identify the impact of single factors on the flow of scientific and technological innovation talents and effectively discover the factors that truly affect the flow of scientific and technological innovation talents.
[0125] Another embodiment of this application provides a device for identifying factors influencing talent mobility, such as... Figure 3 As shown, it includes:
[0126] Module 1 is used to acquire observational data on factors influencing the mobility of target talent and observational data on the mobility results of the target talent.
[0127] Identification module 2 is used to perform approximate hidden confounding factor estimation on the observed data of the influencing factors using a multi-effect model based on causal machine learning, and to obtain the estimated value of the alternative confounding factor;
[0128] The identification module 3 is used to identify the talent mobility influencing factors corresponding to the target talent by using a preset linear regression result model based on the observed data of the influencing factors, the observed data of the mobility results, and the estimated value of the substitution and confusion factors.
[0129] In the specific implementation process, the identification module 2 is specifically used to: identify the observation data of the influencing factors using a preset latent variable model to obtain the confusion factor; and calculate the confusion factor using the maximum likelihood estimation EM iterative algorithm to obtain the estimated value of the alternative confusion factor.
[0130] In specific implementation, the device further includes a model building module, which is specifically used for: acquiring multiple talent mobility sample data; dividing the multiple talent mobility sample data into a sample set to obtain an observation dataset and a validation dataset; training an initial latent variable model based on the observation dataset to obtain a latent variable model that meets preset conditions; validating the latent variable model using the validation dataset, and determining the latent variable model as the preset latent variable model when the predicted test value of the latent variable model is greater than a preset threshold.
[0131] In the specific implementation process, the model building module is also used for: training the initial latent variable model based on the observation dataset to obtain a latent variable model that meets preset conditions, specifically including: Step 1, randomly initializing the model parameters of the initial latent variable model, the model parameters including a weight matrix and noise variance used to characterize the linear influence of latent variables on causal variables; Step 2, calculating the posterior distribution of latent variables for each observation dataset based on the model parameters, to obtain the posterior covariance and posterior mean corresponding to each observation dataset; Step 3, updating the model parameters based on each posterior covariance and each posterior mean; Step 4, if the current iteration number is less than a preset threshold, repeating steps 2 to 3 based on the updated model parameters until the current iteration number is equal to the preset threshold, then building the model based on the updated model parameters and the initial latent variable model to obtain the latent variable model.
[0132] In specific implementation, the device further includes an update module, which is specifically used to: when the noise variance in the updated model parameters is less than a preset threshold, construct a model based on the updated model parameters and the initial latent variable model to obtain the latent variable model.
[0133] In specific implementation, the device further includes a model validation module, which is specifically used to validate the latent variable model using the validation dataset. When the predicted test value of the latent variable model is greater than a preset threshold, the latent variable model is determined as the preset latent variable model. Specifically, this includes: obtaining a validation dataset with the same confounding factor; performing calculations based on the confounding factor and the model parameters of the latent variable model to obtain multiple prediction datasets corresponding to the validation dataset; performing calculations based on the validation dataset and each prediction dataset using the Monte Carlo averaging algorithm to obtain a first statistic corresponding to the validation dataset and a second statistic corresponding to each prediction dataset; performing calculations based on the first statistic and each second statistic to obtain a predicted test value corresponding to the same confounding factor, thus obtaining predicted test values corresponding to each confounding factor; and determining the latent variable model as the preset latent variable model when each predicted test value is greater than the preset threshold.
[0134] In the specific implementation process, the identification module 3 is specifically used to: calculate and process the influencing factor observation data, the flow result observation data, and the estimated value of the substitution confounding factor using the preset linear regression result model to obtain the causal effect weight coefficient vector corresponding to the influencing factor observation data and the weight coefficient vector corresponding to the estimated value of the substitution confounding factor; filter each causal effect weight coefficient in the causal effect weight coefficient vector to obtain the target causal effect weight coefficient; and determine the target influencing factor affecting talent flow based on the target causal effect weight coefficient to obtain the talent flow influencing factor identification result corresponding to the target talent.
[0135] This application analyzes the influencing factors of the flow of scientific and technological innovation talents from a systemic perspective. By establishing a multiple causal effect estimation method with alternative confounding factors, it effectively controls confounding bias. Furthermore, the model is simple, easy to implement, and conforms to prediction verification. A linear regression model is established on the enhanced dataset to effectively consider the comprehensive impact of influencing factors. The model can identify the impact of single factors on the flow of scientific and technological innovation talents and effectively discover the factors that truly affect the flow of scientific and technological innovation talents.
[0136] Another embodiment of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the following method steps:
[0137] Step 1: Obtain observational data on the factors influencing the mobility of the target talent and observational data on the mobility outcomes of the target talent;
[0138] Step 2: Use a multi-effects model based on causal machine learning to estimate the hidden confounding factors in the observed data of the influencing factors, and obtain the estimated values of the alternative confounding factors;
[0139] Step 3: Based on the observed data of the influencing factors, the observed data of the flow results, and the estimated value of the substitution and confounding factors, a preset linear regression model is used to identify the influencing factors of talent flow corresponding to the target talent.
[0140] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0141] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0142] The specific implementation process of the above method steps can be found in the embodiments of the above-mentioned method for identifying factors influencing talent mobility, which will not be repeated here.
[0143] This application analyzes the influencing factors of the flow of scientific and technological innovation talents from a systemic perspective. By establishing a multiple causal effect estimation method with alternative confounding factors, it effectively controls confounding bias. Furthermore, the model is simple, easy to implement, and conforms to prediction verification. A linear regression model is established on the enhanced dataset to effectively consider the comprehensive impact of influencing factors. The model can identify the impact of single factors on the flow of scientific and technological innovation talents and effectively discover the factors that truly affect the flow of scientific and technological innovation talents.
[0144] Another embodiment of this application provides an electronic device, which can be a server. The electronic device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the program is executed by the processor, it implements the functions or steps of a server-side method for identifying factors influencing talent mobility.
[0145] In one embodiment, an electronic device is provided, which can be a client. The electronic device includes a processor, memory, a network interface, a display screen, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When the program of the electronic device is executed by the processor, it implements the functions or steps of a client-side method for identifying factors influencing talent mobility.
[0146] Another embodiment of this application provides an electronic device, including at least a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program in the memory, performs the following method steps:
[0147] Step 1: Obtain observational data on the factors influencing the mobility of the target talent and observational data on the mobility outcomes of the target talent;
[0148] Step 2: Use a multi-effects model based on causal machine learning to estimate the hidden confounding factors in the observed data of the influencing factors, and obtain the estimated values of the alternative confounding factors;
[0149] Step 3: Based on the observed data of the influencing factors, the observed data of the flow results, and the estimated value of the substitution and confounding factors, a preset linear regression model is used to identify the influencing factors of talent flow corresponding to the target talent.
[0150] The specific implementation process of the above method steps can be found in the embodiments of the above-mentioned method for identifying factors influencing talent mobility, which will not be repeated here.
[0151] This application analyzes the influencing factors of the flow of scientific and technological innovation talents from a systemic perspective. By establishing a multiple causal effect estimation method with alternative confounding factors, it effectively controls confounding bias. Furthermore, the model is simple, easy to implement, and conforms to prediction verification. A linear regression model is established on the enhanced dataset to effectively consider the comprehensive impact of influencing factors. The model can identify the impact of single factors on the flow of scientific and technological innovation talents and effectively discover the factors that truly affect the flow of scientific and technological innovation talents.
[0152] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A method for identifying talent flow influencing factors, characterized in that, The method comprises the following steps: Obtaining observation data of influencing factors affecting the flow of target talents and observation data of the flow results of the target talents; Performing approximate hidden confounder estimation on the observation data of the influencing factors by using a multiple-effect model based on causal machine learning to obtain an estimated value of a substitute confounder; Performing identification by using a preset linear regression result model based on the observation data of the influencing factors, the observation data of the flow results and the estimated value of the substitute confounder to obtain an identification result of the talent flow influencing factors corresponding to the target talents.
2. The method of claim 1, wherein, The approximate hidden confounder estimation on the observation data of the influencing factors by using the multiple-effect model based on causal machine learning to obtain the estimated value of the substitute confounder specifically comprises: Performing identification on the observation data of the influencing factors by using a preset hidden variable model to obtain a confounder; Performing calculation on the confounder by using an EM iterative algorithm of maximum likelihood estimation to obtain the estimated value of the substitute confounder.
3. The method of claim 2, wherein, Before the identification on the observation data of the influencing factors by using the preset hidden variable model, the method further comprises the following step: Constructing a preset hidden variable model; The construction of the preset hidden variable model specifically comprises the following steps: Obtaining multiple talent flow sample data; Dividing the multiple talent flow sample data into a sample set to obtain an observation data set and a verification data set; Performing model training on an initial hidden variable model based on the observation data set to obtain a hidden variable model satisfying a preset condition; 4. The method of claim 3, wherein, Verifying the hidden variable model by using the verification data set, and determining the hidden variable model as the preset hidden variable model when a prediction test value of the hidden variable model is greater than a preset threshold. The model training on the initial hidden variable model based on the observation data set to obtain the hidden variable model satisfying the preset condition specifically comprises the following steps: Step 1, randomly initializing model parameters of the initial hidden variable model, wherein the model parameters comprise a weight matrix used for representing a linear influence degree of a hidden variable on a cause variable and a noise variance; Step 2, performing hidden variable posterior distribution calculation on each observation data set based on the model parameters to obtain a posterior covariance and a posterior mean value corresponding to each observation data set; Step 3, updating the model parameters based on each posterior covariance and each posterior mean value; 5. The method of claim 4, wherein, Step 4, when a current iteration number is less than a preset number threshold, repeatedly performing steps 2 to 3 based on the updated model parameters until, when the current iteration number is equal to the preset number threshold, performing model construction based on the updated model parameters and the initial hidden variable model to obtain the hidden variable model. The method further comprises the following steps:
6. The method of claim 3, wherein, When the noise variance in the updated model parameters is less than a preset threshold, performing model construction based on the updated model parameters and the initial hidden variable model to obtain the hidden variable model. The verification of the hidden variable model by using the verification data set, and the determination of the hidden variable model as the preset hidden variable model when the prediction test value of the hidden variable model is greater than a preset threshold, specifically comprises the following steps: Obtaining a verification data set of the same confounder; Performing calculation processing based on the confounding factors and model parameters of the latent variable model to obtain a plurality of predicted data sets corresponding to the validation data set; Performing calculation processing based on the validation data set and each of the predicted data sets using a Monte Carlo average algorithm to obtain a first statistic corresponding to the validation data set and a second statistic corresponding to each of the predicted data sets; Performing calculation processing based on the first statistic and each of the second statistics to obtain a predicted test value corresponding to the same confounding factor, so as to obtain a predicted test value corresponding to each confounding factor; When each of the predicted test values is greater than a preset threshold value, determining the latent variable model as the preset latent variable model.
7. The method of claim 1, wherein, The preset linear regression result model is used to identify the talent flow influence factor identification result corresponding to the target talent based on the influence factor observation data, the flow result observation data, and the estimated value of the alternative confounding factor, specifically including: The preset linear regression result model is used to perform calculation processing based on the influence factor observation data, the flow result observation data, and the estimated value of the alternative confounding factor to obtain a causal effect weight coefficient vector corresponding to the influence factor observation data and a weight coefficient vector corresponding to the estimated value of the alternative confounding factor; Each causal effect weight coefficient in the causal effect weight coefficient vector is screened to obtain a target causal effect weight coefficient. The target influence factor affecting talent flow is determined based on the target causal effect weight coefficient, so as to obtain the talent flow influence factor identification result corresponding to the target talent. 8.A talent flow influence factor identification device, characterized in that, It includes: An acquisition module is configured to acquire influence factor observation data affecting target talent flow and flow result observation data of the target talent; An identification module is configured to approximate a hidden confounding factor estimation of the influence factor observation data using a multiple effect model based on causal machine learning to obtain an estimated value of an alternative confounding factor; An identification module is configured to identify a talent flow influence factor identification result corresponding to the target talent based on the influence factor observation data, the flow result observation data, and the estimated value of the alternative confounding factor using a preset linear regression result model.
9. A storage medium, characterized by The storage medium stores a computer program, and the computer program is executed by the processor to realize the steps of the talent flow influence factor identification method in any one of claims 1-7.
10. An electronic device, comprising: At least including a memory and a processor, the memory stores a computer program, and the processor realizes the steps of the talent flow influence factor identification method in any one of claims 1-7 when executing the computer program on the memory.