Student subsidization project effect analysis method based on tendency score matching and double difference
Through the combination of propensity score matching and double-differential method, the problems of parallel trend assumptions and sample self-selecting bias in scholar-sponsored project evaluation are solved, and more accurate analysis of the effect of scholar-sponsored project is achieved, providing high-reliability evaluation results.
Patent Information
- Application Number
- CN202510496340.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-22
AI Technical Summary
The existing method of evaluation of the effect of scholar-funded projects relies on parallel trend assumptions and instrumental variables, making it difficult to effectively solve the problems of sample self-selecting bias and insufficient data quality, resulting in inaccurate estimation results.
A comprehensive analysis method based on propensity score matching and double difference was adopted, and a multi-level causal effect evaluation framework was constructed through data preprocessing, academic index calculation, sample screening, deep learning model training and weighted regression analysis to ensure the comparability and dynamic feature matching between the processing and control groups, and weaken the influence of unobserved confounders.
It significantly improves the accuracy and reliability of the effect analysis of scholarly funded projects, breaks through the dependence on parallel trend assumptions and exogeneity of instrumental variables, provides a highly credible evaluation system, strong adaptability, and can more accurately identify the net effect of funded projects.
Smart Images

Figure CN120355262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an evaluation technique using statistics and causal inference, and particularly to a method for analyzing the effects of scholar funding projects based on propensity score matching and difference-in-differences. Background Art
[0002] Evaluating the effects of scholar funding projects is of great value to academic research, scientific research management, and social development. At the academic level, evaluating the funding effects can reveal the specific impacts of funding on scholars' scientific research output, innovation ability, and career development, help understand how financial support is transformed into academic achievements, and provide a basis for optimizing resource allocation. Secondly, the evaluation results provide scientific references for funding agencies, helping to adjust the funding scale, optimize the screening criteria, and improve the project design, thereby improving the efficiency of fund use and promoting the healthy development of the scientific research ecosystem.
[0003] Current evaluation methods still mainly rely on causal inference. The instrumental variable method solves the endogeneity problem by introducing external variables related to funding allocation but not directly affecting the results, providing a consistent estimate. However, instrumental variables rely on exogeneity and correlation assumptions and are difficult to find. The fixed effects model reduces the omitted variable bias by controlling the time-invariant heterogeneity at the individual level and is suitable for panel data analysis, but it cannot handle time-varying unobserved factors and has high data requirements.
[0004] The difference-in-differences method is an econometric method widely used in the evaluation field to evaluate the causal effects of intervention measures. Its core idea is to eliminate the interference of time trends and individual heterogeneity on the results by comparing the change differences between the treatment group and the control group before and after the intervention, so as to more accurately estimate the intervention effect. Compared with the traditional single-difference method, the difference-in-differences method can better solve the endogeneity problem and improve the accuracy of estimation. However, its effectiveness depends on the parallel trends assumption, that is, the trends of the treatment group and the control group should be consistent before the intervention. If this assumption does not hold, it may lead to estimation bias. In addition, the difference-in-differences method is highly sensitive to data quality and model specification, and it is necessary to carefully select control variables and test hypotheses to ensure the reliability of the results. With the in-depth research, extended models of the difference-in-differences method in aspects such as dynamic treatment effects and heterogeneous treatment effects have emerged continuously, further enhancing its flexibility and interpretability in practical applications. However, the application of the difference-in-differences method also has certain limitations and challenges. First, this method depends on the parallel trends assumption, that is, the trends of the treatment group and the control group should be consistent before the intervention. If this assumption does not hold, the estimation results may be biased. Secondly, the selection of the treatment group and the control group may be affected by the sample self-selection problem, resulting in inaccurate estimation results. In addition, the time point, duration of the intervention, and changes in the external environment may also interfere with the estimation results.
[0005] Propensity score matching aims to reduce confounding bias by constructing comparability between the treatment group and the control group, thereby more accurately estimating the causal effect of the intervention. Its core idea is to use the propensity score to match individuals with similar characteristics in the treatment group and the control group to simulate the environment of a randomized experiment. The difference-in-differences method eliminates the influence of time trends and confounding factors by comparing the differences between the treatment group and the control group before and after the intervention. However, its effectiveness depends on the parallel trends assumption and may be interfered by sample self-selection bias; propensity score matching reduces confounding bias by constructing comparability between the treatment group and the control group, but is highly sensitive to unobservable confounding factors and matching quality.
[0006] It should be noted that the information disclosed in the above background art section is only used for understanding the background of the present application, and thus may include information that does not constitute prior art known to those of ordinary skill in the art. Summary of the Invention
[0007] The main object of the present invention is to overcome the defects existing in the above background art and provide a method for analyzing the effects of scholar funding projects based on propensity score matching and difference-in-differences.
[0008] To achieve the above object, the present invention adopts the following technical solutions: A method for analyzing the effects of scholar funding projects based on propensity score matching and difference-in-differences, comprising the following steps: S1. Data acquisition and preprocessing: Obtain scholar data from an academic database and perform data preprocessing; S2. Academic indicator calculation: According to the unique identifier of the scholar, analyze their academic publication status in different years, calculate and summarize a series of academic indicators, and the academic indicators are used to reflect the academic output and influence of the scholar; S3. Sample screening: Screen out the scholar samples of the experimental group and the control group, where the experimental group is the scholars who have received project funding, and the control group is the scholars who have not received project funding but have the same parallel trend as the experimental group before the funding; S4. Difference-in-differences regression analysis: Construct a difference-in-differences regression model, by comparing the change differences between the treatment group and the control group before and after the intervention, test the significance of the interaction term coefficient to determine whether the parallel trends assumption holds, so as to evaluate the preliminary estimate of the effect of the funding project on the scholars; S5. Deep propensity score matching model training: Use a deep learning model to predict the propensity scores of all control group samples, and the training set includes experimental group samples and randomly selected non-experimental group samples; S6. Weighted regression analysis: Assign fixed weights to the experimental group and weights to the control group according to the propensity scores. Use weighted least squares to fit a difference-in-differences regression model that includes the treatment variable, time variable, and their interaction terms. Determine the final estimate of the effect of the funded project on scholars through the statistical results of the regression analysis. If the coefficient of the interaction term is significant after the intervention, it indicates that the funded project has had a significant academic impact on scholars.
[0009] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method for analyzing the effect of a scholar-funded project based on propensity score matching and difference-in-differences.
[0010] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method for analyzing the effect of a scholar-funded project based on propensity score matching and difference-in-differences.
[0011] The present invention has the following beneficial effects: The present invention provides a method for analyzing the effect of a scholar-funded project based on propensity score matching and difference-in-differences. By integrating propensity score matching and the difference-in-differences method, a multi-level and complementary causal effect evaluation framework is constructed, significantly improving the accuracy and reliability of the analysis of the effect of the scholar-funded project. First, with the help of propensity score matching technology, dynamic feature matching is performed on the treatment group and the control group before the intervention, effectively eliminating sample selection bias caused by observable confounding factors, and ensuring high comparability between the two groups in key dimensions such as academic fields and academic outputs. Second, the difference-in-differences method is used to capture the dynamic changes before and after the intervention. While controlling for time trends and individual heterogeneity, it further weakens the interference of unobserved confounding factors on the estimation results, thereby more accurately identifying the net effect of the funded project. In addition, a deep learning model is used to optimize the prediction of propensity scores, enhancing the ability to capture non-linear features in the matching process, and combining weighted regression analysis to dynamically adjust the sample weights, solving the estimation bias problem caused by insufficient data quality or overly strong model assumptions in traditional methods. The method of the present invention breaks through the excessive dependence of single causal inference techniques on conditions such as the parallel trend assumption and the exogeneity of instrumental variables. Through multi-stage data cleaning, model verification, and sensitivity analysis, a scientific, rigorous, and highly adaptable evaluation system is formed, providing high-confidence technical support for effect quantification in complex scientific research management scenarios.
[0012] Other beneficial effects in the embodiments of the present invention will be further described below. Description of the Drawings
[0013] Figure 1 It is a flowchart of the method for analyzing the effect of a scholar-funded project based on propensity score matching and difference-in-differences according to an embodiment of the present invention.
[0014] Figure 2 The annual cumulative impact factor curve graph of 20 candidate scholars in the control group and scholars in the experimental group.
[0015] Figure 3 The trend graph of the change in the interaction term coefficient of scholars before and after being funded by the National Science Fund for Distinguished Young Scholars. Specific implementation manners
[0016] The following makes a detailed description of the implementation manners of the present invention. It should be emphasized that the following description is merely exemplary and not intended to limit the scope of the present invention and its applications.
[0017] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0018] Refer to Figure 1 , the embodiments of the present invention provide a method for analyzing the effects of scholar funding projects based on propensity score matching and difference-in-differences, including the following steps: Step S1, data acquisition and preprocessing: Obtain scholar data from an academic database and perform data preprocessing. Specifically, it includes removing samples with a large amount of missing data and complementing samples with a small amount of missing data to ensure the integrity and availability of the data.
[0019] Step S2, academic indicator calculation: According to the unique identifier of the scholar, analyze his / her academic publication situation in different years, calculate and summarize a series of academic indicators, and the academic indicators are used to reflect the academic output and influence of the scholar.
[0020] In some embodiments, the academic indicators include the annual cumulative number of published papers, impact factor, journal division, and number of co-authors.
[0021] In some embodiments, step S2 specifically includes: loading the JCR partition, Chinese Academy of Sciences partition, and impact factor data of the journals from the specified CSV file; traversing the publication records of the scholars, counting the cumulative number of papers published annually by year, and calculating the annual cumulative, average, and highest impact factor (IF) and journal citation indicator (JCI) based on the JCR data; according to the JCR partition determination rule, accumulating the number of papers in Q1 and Q2 partitions published by the scholars annually; based on the major category partition and Top mark of the Chinese Academy of Sciences partition table, counting the number of papers in the first and second zones and top journals published by the scholars annually; calculating the annual cumulative number of co-authors by de-duplicating the author ID list in the publications; storing the obtained indicators as a dictionary structure by year, where the dictionary key is the year and the value is a nested dictionary containing the values of each indicator.
[0022] Step S3, sample screening: Screen out the scholar samples of the experimental group and the control group. The experimental group is the scholars who have received project funding, and the control group is the scholars who have not received project funding but have the same parallel trend as the experimental group before the funding. By comparing the trend consistency of the two groups before the intervention, the comparability of the samples can be ensured.
[0023] In some embodiments, the screening of the control group specifically includes: Based on the similarity of the research fields of the scholars, preliminarily screen the potential control group samples that match the characteristics of the scholars in the experimental group from the candidate scholar pool; randomly sample the potential control group samples to generate a preset number of candidate control group sample groups; pre-screen the candidate control group sample groups with the experimental group for parallel trends, and output the control group that meets the preset comparability conditions.
[0024] Step S4, difference-in-differences regression analysis: Construct a difference-in-differences regression model, and by comparing the change differences between the treatment group and the control group before and after the intervention, test the significance of the interaction term coefficient to determine whether the parallel trend hypothesis holds, so as to evaluate the preliminary estimate of the effect of the funded project on the scholars; In some embodiments, step S4 specifically includes: For each candidate control group sample group, construct a difference-in-differences regression model, and verify the parallel trend consistency between the experimental group and the candidate control group before the intervention by testing the statistical significance of the interaction term between the treatment group and time; screen out the candidate control group sample groups whose interaction term coefficients do not pass the significance test during the pre-intervention time period as the final control group that meets the parallel trend hypothesis.
[0025] Step S5, training of the deep propensity score matching model: Use a deep learning model to predict the propensity scores of all control group samples. The training set includes the experimental group samples and randomly selected non-experimental group samples. Through propensity score matching, the differences in observable characteristics between the treatment group and the control group can be further balanced, and the impact of sample self-selection bias on the results can be reduced.
[0026] In some embodiments, step S5 specifically includes: constructing a neural network model, which sequentially includes a feature expansion module, a temporal context information embedding module, an encoding module based on a multi-head attention mechanism, a feature dimensionality reduction module, and a sequence of fully connected layers; marking the experimental group samples as the treatment group and the non-experimental group samples as the control group, and optimizing the parameters of the neural network model based on the marked training set data using a classification loss function; predicting the propensity scores of the control group samples through the neural network model to generate probability scores within the range of 0 to 1; sorting the control group samples based on the propensity scores, screening and removing the samples with scores lower than a preset threshold to form an optimized control group.
[0027] Step S6, weighted regression analysis: Assign a fixed weight to the experimental group and assign weights to the control group according to the propensity scores. Use weighted least squares to fit a difference-in-differences regression model that includes a treatment variable, a time variable, and their interaction term. Through the statistical results of the regression analysis, determine the final estimate of the effect of the funding project on scholars. If the interaction term coefficient is significant after the intervention, it indicates that the funding project has had a significant academic impact on scholars.
[0028] In some embodiments, step S6 specifically includes: dynamically generating a comparison data set between the experimental group and the control group according to the intervention time parameter input by the user; loading a data set containing treatment status, time variable, and dependent variable based on the individual unique identifier, assigning a fixed weight to the experimental group samples, and assigning a dynamic weight based on the propensity scores to the control group samples; using weighted least squares to fit a difference-in-differences regression model that includes a treatment variable, a time variable, and their interaction term; determining the statistical effect of the interaction term through a significance test; determining the final estimate of the effect of the funding project on scholars based on the significance test of the interaction term in the statistical results of the regression analysis.
[0029] In some embodiments, step S6 further includes performing at least one of the following operations: calling a data visualization module to generate a time series trend chart; performing a difference-in-differences analysis and outputting the results of a parallel trends test; outputting a statistical summary of the regression analysis, including coefficient estimates and significance test information.
[0030] The method for analyzing the effect of a scholar funding project based on propensity score matching and difference-in-differences combines the difference-in-differences method and propensity score matching. It can use propensity score matching to match the samples of the treatment group and the control group before the intervention to ensure comparability in key features between the two groups, and then further eliminate the influence of time trends and unobserved confounding factors through the difference-in-differences method, so as to more accurately estimate the net effect of the intervention. This comprehensive analysis method based on propensity score matching and difference-in-differences provides a scientific and reliable technical support for the evaluation of the effect of scholar funding projects, can effectively solve the endogeneity and selection bias problems existing in traditional evaluation methods, and provides a more accurate decision-making basis for project optimization.
[0031] The specific embodiments of the present invention and experimental verification are further described below.
[0032] A method for analyzing the effect of a scholar funding project based on propensity score matching and difference-in-differences, comprising the following steps: Crawl scholar data from the Scopus database; clean the scholar data in Scopus, remove scholar samples with a large amount of missing data, and appropriately complete the data of scholar samples with less missing data.
[0033] Analyze the academic publications of scholars in different years according to the given Scopus ID of the scholars. Calculate and summarize a series of indicators, including the annual cumulative number of published papers, the annual average / cumulative / highest impact factor (IF) and JCI (Journal Citation Indicator), the number of papers in JCR categories (Q1, Q2) and Chinese Academy of Sciences categories (category 1, category 2, top journals), and the annual cumulative number of co-authors. The code loads journal information from a specified CSV file, traverses the publication records of scholars, counts the relevant indicators by year, and stores the results in a dictionary.
[0034] Select scholar samples for the experimental group and the control group. The scholars in the experimental group are those who have received project funding, and the scholars in the control group are those who have not received project funding but had the same parallel trend as the scholars in the experimental group before the project funding.
[0035] Regression example model of the difference-in-differences method:
[0036] where is the outcome variable of individual at time ; is the treatment group dummy variable, taking the value of 1 indicating that the individual belongs to the treatment group and 0 indicating the control group; is the time dummy variable, taking the value of 1 indicating after the intervention and 0 indicating before the intervention. is the interaction term between the treatment group and time, used to capture the causal effect of the intervention. Conduct a regression analysis to test whether the coefficient of the interaction term is significant. If the coefficient of the interaction term is not significant, it supports the parallel trend hypothesis. If a group of scholar samples satisfies the parallel trend for a certain experimental group of scholar samples, it is the control group sample for the scholars in that experimental group.
[0037] Train a deep propensity score matching model. Use a deep learning model to predict the propensity scores of all control group samples. The training set includes samples from the experimental group and randomly selected non-experimental group samples.
[0038] The depth propensity score matching model is a neural network that combines the structures of feature expansion, positional encoding, Transformer encoding, attention mechanism, and fully connected layers. It first doubles the input feature dimension through a linear layer. Then, temporal context information is added to the expanded features according to their positions in the sequence. Next, the sequence is processed by a two-layer TransformerEncoder, with each layer having four attention heads. After dimension adjustment, the AttentionLayer calculates weighted features in the temporal dimension and reduces the output dimension to the representation of each sample. Subsequently, three fully connected layers process the data in sequence, finally mapping it to a single output and generating a prediction score between 0 and 1 through the sigmoid activation function.
[0039] Perform a difference-in-differences regression analysis to determine significance. Load data containing individual IDs, years, treatment status, and dependent variables through the Scopus ID, then assign a fixed weight of 1 to the experimental group and assign weights of propensity scores to the control group according to individual IDs. Use weighted least squares to fit a model containing treatment variables, time variables, and their interaction terms, and return the fitting results.
[0040] According to the regression example model of the difference-in-differences method, if the interaction term coefficient is significant after the intervention, it indicates that the funded project has brought academic help to scholars.
[0041] The example is as follows.
[0042] Use Scopus as the core academic database for research. Scopus is an interdisciplinary literature database developed by Elsevier Publishing Company, covering more than 24,000 journals, books, and conference papers globally, providing rich content including full text or abstracts, citation information, and bibliometric indicators. Its coverage spans multiple fields such as natural sciences, social sciences, humanities, and medicine. It is worth mentioning that Scopus not only automatically generates records for each scholar with literature inclusion but also allows authors without literature inclusion to register and obtain Scopus IDs by themselves.
[0043] In the research process, a search query was first constructed to extract all papers published by authors affiliated with institutions located in Shenzhen from the Scopus database during a specific time period. For example, through the search query AFFILCITY(Shenzhen) AND PUBDATETXT("January 2022"), all papers with author affiliations in Shenzhen in January 2022 were obtained. The set time range was from January 1990 to December 2022, and a total of 152,359 documents were retrieved, involving 419,978 authors (identified by Scopus ID). Subsequently, based on the JCR journal classification report for 2021, papers published in journals indexed by Web of Science were selected, totaling 88,662 papers involving 279,129 authors. On this basis, scholars who had published 7 or more papers were defined as the "scholar pool", totaling 14,050 people. The reason for choosing 7 as the threshold was to focus on scholars who had already made some accumulation in the academic field and regarded scientific research as their career direction. In universities or research institutions in Shenzhen, many doctoral students may publish 3 to 4 journal papers upon graduation, and some outstanding doctoral students may even publish 5 to 6 papers during their studies, but they may not continue to engage in academic research after graduation. By setting a paper quantity threshold, it was aimed to exclude these student groups studying in Shenzhen and focus on researchers who had been engaged in scientific research in Shenzhen for a long time.
[0044] Finally, the basic information of these 14,050 scholars, including their affiliated institutions and publication records, etc., was extracted from the Scopus database, and cases of missing or incomplete data were removed through data cleaning. Eventually, 13,150 scholars were retained, and their publication records covered 10,796 journals. Then, the scholar data obtained from Scopus was cleaned, which included two sub-steps: First, scholar samples with a large amount of missing data were removed, usually those lacking major category information, such as those lacking publication journal or institution information, etc.; Second, appropriate data supplementation was carried out for scholar samples with less missing data, while ensuring that no significant bias was introduced during the supplementation process.
[0045] Then, a well-written script program is used to read the CSV files of scholars and calculate the corresponding academic metrics. Each scholar contains five files: affiliation_current, affiliation_history, author_info, publications, and subject_areas. This code calculates the academic metrics of the scholar with the input Scopus ID by reading the above five files through the get_all_metrics function, extracts information from the stored publication data based on the Scopus ID, and conducts statistics in combination with the JCR and Chinese Academy of Sciences (CAS) division data. The function first reads the journal data of JCR 2021 and CAS 2022, and then for the publication file (publications.csv) of the specified scholar, traverses each paper by year (2012 - 2022) and calculates the following metrics: 1) The annual cumulative number of published papers is obtained by counting the number of publications each year; 2) The annual cumulative / average / maximum IF and JCI are extracted from the JCR data for the impact factor and JCI value respectively, and the sum, average, and maximum values are calculated; 3) The number of papers in JCR Q1 / Q2 is judged according to the JIF division and accumulated; 4) The number of papers in the first / second zone and top journals of the CAS division is counted according to the major division and Top mark in the CAS division table; 5) The annual cumulative number of co-authors is calculated by deduplicating the author ID list in the publications. Finally, a list of years and a dictionary containing all metrics are returned.
[0046] Next, taking a scholar in the experimental group as an example, the control group is selected. Based on the information obtained by the crawler, scholar samples with similar research fields are selected, and then 15 samples are randomly sampled from these samples each time. A regression analysis is performed on each group of samples and the samples in the experimental group, and a difference-in-differences regression model is fitted:
[0047] where is the outcome variable of individual at time ; is the treatment group dummy variable, taking the value of 1 indicating that the individual belongs to the treatment group and 0 indicating the control group; is the time dummy variable, taking the value of 1 indicating after the intervention and 0 indicating before the intervention. The interaction term between the treatment group and time is used to capture the causal effect of the intervention.
[0048] Check whether the p-value of the interaction term coefficient after the regression analysis is greater than 0.05. If so, it means that is not significant, and this group of samples satisfies the parallel trends assumption and can be used as the control group.
[0049] All the scholar samples of the experimental group and a random sample of the scholar samples of the non-experimental group were selected as the training set of the deep propensity score matching model. The scholar samples of the experimental group were marked as 1, and the scholar samples of the non-experimental group were marked as 0, and the cross entropy loss function was used as the loss function for training. Then the deep learning model was used to predict the propensity scores of all control group samples, and the five scholar samples with the lowest scores were eliminated.
[0050] Perform regression analysis again and fit the double difference regression model:
[0051] Here, weighted regression is used to assign a fixed weight of 1 to the experimental group and a weight of the propensity score to the control group based on the individual id.
[0052] In the actual program, the specific year of project funding can be entered by the user, and the train_and_contrast function is used to generate the contrast data set new_df based on train_id and contrast_id, and save it as a CSV file. Subsequently, different operations are selected through the mode (1, 2, 3) entered by the user: mode=1 calls the drawing_all_data function to draw a line graph of all data and save it; mode=2 executes the DID function to perform double difference analysis and draw a parallel trend graph; mode=3 calls the DID_after function to output the statistical results of the regression analysis and print a detailed summary, including coefficient estimation, significance test and other information.
[0053] The example analysis is as follows.
[0054] Through the crawled scholar data, a scholar with a certain ID in Scopus was selected as the experimental group. This scholar received funding from a certain academic fund in 2020. It is hoped to investigate whether the number of collaborators of this scholar has changed significantly after receiving project funding.
[0055] Next, we calculated the four indicators of the cumulative number of papers published by all scholars, impact factors, journal classification quality, and number of collaborators. We randomly sampled 20 scholars to conduct parallel hypothesis tests with the experimental scholars, and selected co_authors (the cumulative number of collaborators in the year) as the outcome variable. We selected scholars with coefficients close to 0 and P values close to 1 as the basic control group. After the above data cleaning, a total of 528 scholars passed the parallel hypothesis test. According to the similarity matrix of scholars' journals, we continued to screen the 528 scholars to obtain the top 20 scholars with similarity rankings for further data analysis. Figure 2 The graph shows the annual cumulative impact factor of 20 candidate scholars in the control group and the scholars in the experimental group.
[0056] Predict the propensity scores of 20 candidate scholars in the control group through the trained deep propensity score prediction model, and screen out the 5 scholars with the lowest scores. Perform weighted regression analysis on the remaining 15 scholars and the scholars in the experimental group.
[0057] As can be seen from Figure 3 , before the implementation of the project, the confidence interval basically covered 0. Therefore, before the implementation of the funded project, there was no significant difference between the treatment group and the experimental group. After the funding of the National Science Fund for Distinguished Young Scholars, the coefficient of the interaction term was significantly not 0. The summary statistical results are shown in Table 1 below.
[0058] Table 1 Model Statistical Results
[0059] Among them, after the interaction term is "treated x treated", it represents the net effect of the treatment (that is, the additional effect of the treatment group relative to the control group in the time change). The coefficient is 55.3553, but the p-value is 0.000, which is much less than 0.05 and is significant. Therefore, it can be proved that the funded project has brought a significant improvement in the academic impact of scholars.
[0060] An embodiment of the present invention further provides a storage medium for storing a computer program, which when executed, at least executes the method described above.
[0061] An embodiment of the present invention further provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is used to execute the computer program to at least execute the method described above.
[0062] An embodiment of the present invention further provides a processor, which executes a computer program and at least executes the method described above.
[0063] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but not limited to, these and any other suitable types of memories.
[0064] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0065] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0066] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above integrated unit can be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.
[0067] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the aforementioned storage medium includes: various media that can store program codes, such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0068] Alternatively, if the above integrated units are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the aforementioned storage medium includes: various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0069] The methods disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0070] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0071] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0072] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention belongs, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for analyzing the effect of scholar funding projects based on propensity score matching and difference-in-differences, characterized in that, It includes the following steps: S1. Data acquisition and preprocessing: Obtain scholar data from academic databases and perform data preprocessing; S2. Academic indicator calculation: According to the unique identifier of the scholar, analyze their academic publication situation in different years, calculate and summarize a series of academic indicators, which are used to reflect the academic output and influence of the scholar; S3. Sample screening: Screen out the scholar samples of the experimental group and the control group, where the experimental group is the scholars who have received project funding, and the control group is the scholars who have not received project funding but have the same parallel trend as the experimental group before the funding; S4. Difference-in-differences regression analysis: Construct a difference-in-differences regression model, by comparing the change differences between the treatment group and the control group before and after the intervention, test the significance of the interaction term coefficient, to judge whether the parallel trend hypothesis holds, so as to evaluate the preliminary estimate of the effect of the funding project on the scholar; S5. Deep propensity score matching model training: Use a deep learning model to predict the propensity scores of all control group samples, and the training set contains experimental group samples and randomly selected non-experimental group samples; S6. Weighted regression analysis: Assign fixed weights to the experimental group, assign weights to the control group according to the propensity scores, use weighted least squares to fit a difference-in-differences regression model containing the treatment variable, time variable and their interaction terms, and determine the final estimate of the effect of the funding project on the scholar through the statistical results of the regression analysis. If the interaction term coefficient is significant after the intervention, it indicates that the funding project has had a significant academic impact on the scholar.
2. The method according to claim 1, wherein In step S2, the academic indicators include the annual cumulative number of published papers, impact factor, journal partition and the number of collaborators.
3. The method according to claim 1, characterized in that, Step S2 specifically includes: Load the JCR partition, Chinese Academy of Sciences partition and impact factor data of the journals from the specified CSV file; Traverse the scholar publication records, count the annual cumulative number of published papers by year, and calculate the annual cumulative, average and highest impact factor (IF) and journal citation index (JCI) based on the JCR data; According to the JCR partition determination rules, accumulate the number of papers in Q1 and Q2 partitions published by the scholar annually; Based on the major category partition and Top mark of the Chinese Academy of Sciences partition table, count the number of papers in the first zone, second zone and top journals published by the scholar annually; Calculate the annual cumulative number of collaborators by de-duplicating the author ID list in the publications; Store the obtained indicators by year as a dictionary structure, where the dictionary key is the year and the value is a nested dictionary containing the values of each indicator.
4. The method according to claim 1, wherein In step S3, the screening of the control group specifically includes: Based on the similarity of the scholar research fields, preliminarily screen the potential control group samples that match the characteristics of the experimental group scholars from the candidate scholar pool; Randomly sample the potential control group samples to generate a preset number of candidate control group sample groups; Perform parallel trend pre-screening on the candidate control group sample groups and the experimental group, and output the control group that meets the preset comparability conditions.
5. The method according to claim 1, wherein Step S4 specifically includes: For each candidate control group sample group, construct a difference-in-differences regression model, and verify the consistency of the parallel trend between the experimental group and the candidate control group before the intervention by testing the statistical significance of the treatment group and time interaction term; Select candidate control group samples whose interaction term coefficients do not pass the significance test during the pre-intervention time period as the final control group that meets the parallel trend assumption.
6. The method according to claim 1, characterized in that, Step S5 specifically includes: Construct a neural network model, which sequentially includes a feature expansion module, a time context information embedding module, an encoding module based on the multi-head attention mechanism, a feature dimensionality reduction module, and a sequence of fully connected layers; Mark the experimental group samples as the treatment group and the non-experimental group samples as the control group. Based on the marked training set data, optimize the parameters of the neural network model using the classification loss function; Predict the propensity scores of the control group samples through the neural network model to generate probability scores in the range of 0 to 1; Sort the control group samples based on the propensity scores, screen and remove samples with scores lower than the preset threshold to form an optimized control group.
7. The method according to claim 1, characterized in that, Step S6 specifically includes: Dynamically generate a comparison data set between the experimental group and the control group according to the intervention time parameter input by the user; Load the data set containing the treatment status, time variable, and dependent variable based on the individual unique identifier, assign a fixed weight to the experimental group samples, and assign a dynamic weight based on the propensity score to the control group samples; Use the weighted least squares method to fit the difference-in-differences regression model, which includes a treatment variable, a time variable, and their interaction term; Determine the statistical effect of the interaction term through a significance test; Based on the significance test of the interaction term in the statistical results of the regression analysis, determine the final estimate of the effect of the funded project on scholars.
8. The method according to claim 7, characterized in that, Step S6 also includes performing at least one of the following operations: Call the data visualization module to generate a time series trend chart; Perform a difference-in-differences analysis and output the parallel trend test results; Output the statistical summary of the regression analysis, including coefficient estimates and significance test information.
9. A computer-readable storage medium storing a computer program, characterized in that, When executed by a processor, the computer program implements the method for analyzing the effect of a funded project for scholars based on propensity score matching and difference-in-differences as described in any one of claims 1 to 8.
10. A computer program product comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method for analyzing the effect of a funded project for scholars based on propensity score matching and difference-in-differences as described in any one of claims 1 to 8.
Citation Information
Patent Citations
An academic influence prediction method
CN109840616A
Policy effect evaluation method and device, electronic equipment and readable storage medium
CN117593109A
Project value evaluation method and program product
CN118172080A
Critical event influence assessment method and system based on multi-stage double difference model
CN118735132A
Intervention effect analysis system, and intervention effect analysis method
JP2023076082A
Cited By
Standard system applicability evaluation method and system based on multi-source data
CN121480982A