A scientific research strength evaluation method based on synergy quantification of strength and potential
Patent Information
- Application Number
- CN202610897760.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-25
AI Technical Summary
由于缺乏对发展潜力的量化,现有评估结果难以有效预测科研主体未来的竞争力和发展空间,导致资源分配往往倾向于已有优势的机构或人才,而那些虽当前实力一般但增长迅速、后劲充足的科研主体无法被及时识别和支持
[0015]本发明中,首次在科研评估中融合了实力与潜力两个时间维度。通过采集项目数据、创新平台数据和学位点数据,并基于这些数据统计出涵盖活跃度增长、强度增长、人才潜力、国际化程度等潜力相关的二级指标原始值,再经数据变换与归一化处理后加权聚合,最终输出实力得分和潜力得分。这使得评估结果既能反映科研主体当前的科研规模与水平,又能刻画其未来的发展态势,有效解决了现有技术普遍忽视发展潜力评估的问题,从而能够更早识别和扶持增长迅速、后劲充足的科研主体。本发明通过计算各学科的资源配置比例获得相对偏离度,以及通过相对熵计算特色性指数,实现了跨学科可比性和学科布局特色性的量化,填补了现有评估体系在学科特色量化方面的空白。
Smart Images

Figure CN122819979A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and in particular to a method for evaluating scientific research capabilities based on the synergistic quantification of strength and potential. Background Technology
[0002] Objective and quantitative evaluation of the research strength of universities, research institutions, disciplines, and research talents is a crucial foundation for research management, resource allocation, and talent evaluation. Existing technologies include various research evaluation methods and indicator systems, primarily including: bibliometric evaluation methods that utilize paper databases to statistically analyze output and impact indicators; statistical methods based on project funding data that use single-dimensional indicators such as the number of projects and total funding amount for ranking; evaluation methods within university or discipline ranking systems that integrate multiple indicators such as teaching, research, and internationalization through weighted summation; and statistical methods based on talent titles, using the number of various talent projects as a direct measure of a scholar's influence.
[0003] However, the aforementioned existing technologies generally lack a systematic assessment of development potential. Specifically, most existing methods focus only on static indicators of a research entity's current "strength," such as the total number of projects, total funding, and number of paper citations, while neglecting "potential" indicators that reflect future growth trends, such as the average annual growth rate of the number of projects, the growth trend of funding, the reserve of emerging talents, the degree of international cooperation, and the researchers' research sustainability and rapid development potential. Due to the lack of quantification of development potential, existing assessment results are difficult to effectively predict the future competitiveness and development space of research entities, leading to resource allocation often favoring institutions or talents with existing advantages, while those research entities that, although currently of average strength, are growing rapidly and have sufficient potential for future development cannot be identified and supported in a timely manner.
[0004] In addition, existing technologies for evaluation generally suffer from problems such as a lack of standardized processing for skewed data in indicator calculation, a lack of relative indicators to eliminate scale effects, failure to quantify the institutional and disciplinary characteristics from the perspective of resource distribution, crude talent evaluation, and a lack of robustness in data transformation and growth rate calculation. Summary of the Invention
[0005] To address the technical problems existing in the background art, this invention proposes a scientific research strength evaluation method based on the synergistic quantification of strength and potential.
[0006] This invention proposes a scientific research strength assessment method based on the synergistic quantification of strength and potential, comprising: S1. Obtain project data, innovation platform data, and degree program data from research entities. After cleaning and unifying the format of the three types of data, merge them to form a structured dataset. S2. Extract fields from the dataset according to the three levels of institution, discipline and talent, perform statistical calculations on the extracted fields, and generate original values of multiple secondary indicators under the strength dimension and potential dimension. S3. Based on the original value type of each secondary indicator, perform corresponding data transformation and normalization to obtain the score of each secondary indicator; S4. The scores of each secondary indicator are weighted and aggregated to calculate the strength score and potential score of the research subject. The relative deviation is calculated based on the resource allocation ratio of each discipline, and the characteristic index is calculated based on the relative entropy between the discipline resource distribution and the reference distribution. S5. Combine the strength score, potential score, relative deviation, and distinctiveness index into structured data and output it as the evaluation result.
[0007] Preferably, S1 specifically involves: extracting project records of research entities from the project funding management system, with each project record including project name, project leader, funding amount, and project year, forming a project data set; extracting platform identification records of research entities from the platform management system, with each platform record including platform name, platform level, and platform category, forming a platform data set; extracting degree program records of research entities from the degree authorization management system, with each degree program record including the number of master's programs and doctoral programs, forming a degree program data set; performing duplicate record removal and invalid record filtering on the project data set, platform data set, and degree program data set respectively, and then associating and merging the three types of data sets according to the unified identifier of the research entity to form a structured dataset.
[0008] Preferably, the strength dimension includes the institutional strength dimension, the discipline strength dimension, and the talent strength dimension, and the potential dimension includes the institutional potential dimension, the discipline potential dimension, and the talent potential dimension.
[0009] Preferably, step S3 specifically involves: inputting the original value of each secondary indicator, calculating the data distribution characteristics of the original value, determining whether its data distribution type is continuous skewed or discrete based on the calculation result, and outputting the distribution type determination result; inputting the original value of the secondary indicator and its distribution type determination result into the data transformation processing stage; if the distribution type is continuous skewed, using Box-Cox transformation or Yeo-Johnson transformation to convert the original value into a continuous value with an approximate normal distribution; if the distribution type is discrete, using binning discretization to map the original value into a level value; and outputting the first processed value after transformation processing; inputting the first processed value and determining its data type; if the data type is continuous data, using cumulative distribution function normalization to convert the first processed value into a cumulative probability value; if the data type is binary discrete data, using direct mapping normalization to map the first processed value into a binary score value; and if the data type is multivariate discrete data, using empirical cumulative distribution function normalization to convert the first processed value into an empirical cumulative proportion value; and outputting the second processed value as the score of the secondary indicator after normalization processing.
[0010] Preferably, the step of calculating the strength score and potential score of the research subject by weighted aggregation of the scores of each secondary indicator is as follows: input the scores of each secondary indicator, read the weight value corresponding to each secondary indicator score from the preset weight database, and output the pairing set of secondary indicator scores and weight values; select the secondary indicator scores and their corresponding weight values under the strength dimension from the pairing set, perform weighted summation calculation, and output the calculation result as the strength score of the research subject; select the secondary indicator scores and their corresponding weight values under the potential dimension from the pairing set, perform weighted summation calculation, and output the calculation result as the potential score of the research subject.
[0011] Preferably, step S4 further includes weighting and aggregating the strength score and potential score according to a preset ratio to obtain a comprehensive score for measuring the overall level of the research subject.
[0012] Preferably, the calculation of relative deviation based on the resource allocation ratio of each discipline specifically involves: inputting the resource allocation data of the target research subject in each discipline, calculating the resource quantity of each discipline, dividing the resource quantity of each discipline by the total resource quantity of all disciplines to generate the subject discipline ratio distribution; inputting the resource allocation data of all research subjects within the reference range in each discipline, calculating the resource quantity of each discipline, dividing the resource quantity of each discipline by the total resource quantity of all disciplines to generate the reference discipline ratio distribution; inputting the subject discipline ratio distribution and the reference discipline ratio distribution in a one-to-one correspondence for each discipline, performing a division operation of the subject ratio value divided by the reference ratio value for each discipline, taking the logarithm of the division result, and outputting the logarithmic result as the relative deviation of that discipline.
[0013] Preferably, the calculation of the distinctiveness index based on the relative entropy between the disciplinary resource distribution and the reference distribution specifically involves: inputting the resource allocation data of the target research subject in each discipline, calculating the resource quantity of each discipline, dividing the resource quantity of each discipline by the total resource quantity of all disciplines to generate a subject probability distribution; inputting the resource allocation data of all research subjects within the reference range in each discipline, calculating the resource quantity of each discipline, dividing the resource quantity of each discipline by the total resource quantity of all disciplines to generate a reference probability distribution; inputting the subject probability distribution and the reference probability distribution one-to-one by discipline, performing a division operation of the subject probability value divided by the reference probability value for each discipline, taking the logarithm of the division result, multiplying the logarithm of each discipline by the subject probability value of that discipline, summing the product results of all disciplines, and outputting the summed result as the distinctiveness index of the research subject.
[0014] Preferably, step S5 specifically involves: combining the strength score and potential score into a strength-potential result set; combining the relative deviation and distinctiveness index into a distinctiveness evaluation result set; merging the strength-potential result set and the distinctiveness evaluation result set to form a complete evaluation result set containing the strength score, potential score, relative deviation, and distinctiveness index; and outputting this complete evaluation result set as the evaluation result.
[0015] This invention, for the first time, integrates two time dimensions—strength and potential—in scientific research evaluation. By collecting project data, innovation platform data, and degree program data, and statistically analyzing these data to derive raw values for potential-related secondary indicators such as activity growth, intensity growth, talent potential, and internationalization, the data is then weighted and aggregated after data transformation and normalization to ultimately output strength and potential scores. This allows the evaluation results to reflect both the current scale and level of research and its future development trend, effectively addressing the problem of existing technologies generally neglecting the assessment of development potential. This enables earlier identification and support of rapidly growing and high-potential research entities. Furthermore, this invention achieves cross-disciplinary comparability and quantification of disciplinary layout characteristics by calculating the relative deviation of resource allocation ratios across disciplines and by calculating a distinctiveness index through relative entropy, filling the gap in the quantification of disciplinary characteristics in existing evaluation systems. Attached Figure Description
[0016] Figure 1 This is a flowchart of the multi-dimensional evaluation method for scientific research strength proposed in this invention. Detailed Implementation
[0017] Reference Figure 1 This invention proposes a scientific research strength assessment method based on the synergistic quantification of strength and potential, comprising: S1. Obtain project data, innovation platform data, and degree program data from research entities, specifically including the following steps: First, project data is collected. The project data comes from the project funding management system of research funding agencies, such as data on projects funded by the National Natural Science Foundation of China. Each project record contains the following fields: project approval number, project name, project leader, application code, direct funding cost, affiliated institution name, project start date, project end date, project year, basic type, supplementary type, affiliated academic department, first-level discipline name, and second-level discipline name.
[0018] Secondly, data on innovation platforms is collected. This data originates from platform identification information managed by relevant departments such as science and technology, development and reform, industry and information technology, and education within the platform management system. Each platform record includes the following fields: platform name, implementing unit, platform level, platform category, city where the platform is located, and the relevant management department. Platform levels are categorized into national and provincial / ministerial levels, and platform categories include key laboratories, engineering research centers, enterprise technology centers, and industrial technology infrastructure public service platforms, among others.
[0019] Next, the degree program data is collected. This data comes from the statistical information on degree-granting programs published by the education authorities in the degree authorization management system. Each record contains the following fields: institution name, number of master's programs in the first-level discipline, and number of doctoral programs in the first-level discipline.
[0020] Finally, the collected project data, innovation platform data, and degree program data underwent data cleaning and format standardization. Specifically, this included: removing duplicate records (deleting identical records resulting from repeated submissions or imports); removing invalid records (deleting data with missing key fields or obvious errors); and standardizing data formats, such as unifying all funding unit to ten thousand yuan, date fields to year-month-day format, and institution names to standard full names. The cleaned and format-standardized data was then integrated to form a structured dataset for subsequent steps.
[0021] S2. Based on the dataset, and according to the evaluation needs of research entities at the institutional, disciplinary, and talent levels, fields corresponding to the respective levels are extracted from the dataset. Statistical calculations are then performed on these extracted fields to generate raw values for multiple secondary indicators under the strength and potential dimensions. The strength and potential dimensions are: Institutional Strength Dimension, Institutional Potential Dimension, Disciplinary Strength Dimension, Disciplinary Potential Dimension, Talent Strength Dimension, and Talent Potential Dimension. The following will explain each dimension: (1) Statistics of the original values of secondary indicators under the institutional strength dimension The institutional strength dimension measures the current research scale and level of a research institution, comprising five secondary indicators: research activity, research intensity, talent cultivation, innovation platform, and talent strength. The specific statistical process for each secondary indicator under this dimension is as follows: Extract all project records from the dataset for the institution, and count the number of projects as the raw value for "Institutional Research Activity". Extract the funding field for all projects and sum them up to obtain the raw value for "Institutional Research Intensity". Extract the institution's degree program records from the dataset, obtaining the number of master's programs and doctoral programs respectively. Weight the number of master's programs and doctoral programs according to a preset coefficient and sum them to obtain the raw value for "Institutional Talent Cultivation" (the preset coefficient is used to balance the difference in talent cultivation capabilities between master's programs and doctoral programs). Extract the institution's innovation platform records from the dataset, assigning different weights according to the platform level (e.g., national, provincial / ministerial), and sum the weights of each platform to obtain the raw value for "Innovation Platform". Extract the institution's talent project records from the dataset, including Young Scientists Fund projects, Excellent Young Scientists Fund projects, National Science Fund for Distinguished Young Scholars projects, etc. Assign different weights to different types of talent projects, and sum the weighted numbers to obtain the raw value for "Institutional Talent Intensity".
[0022] (2) Statistics on the original values of secondary indicators under the institutional potential dimension The institutional potential dimension measures the development trend of research institutions and includes four secondary indicators: activity growth, intensity growth, talent potential, and internationalization level. The specific statistical process for each secondary indicator under this dimension is as follows: The dataset is used to extract the project quantity sequence of the institution over the years, calculate the average annual growth rate of this sequence, and obtain the raw value of "Institutional Activity Growth". Boundary cases such as no data or insufficient data in the starting year need to be handled during the calculation. The dataset is also used to extract the total funding amount sequence of the institution over the years, calculate the average annual growth rate of this sequence, and obtain the raw value of "Institutional Strength Growth". The dataset is then used to extract the number of Young Scientists Fund projects and the number of project leaders of Innovative Research Groups affiliated with the institution, and the weighted sum of these two values yields the raw value of "Institutional Talent Potential". Finally, the dataset is used to search for whether the institution has undertaken any international collaborative projects; if so, it is marked as "1", otherwise marked as "0", yielding the raw value of "Internationalization Level".
[0023] (3) Statistics on the original values of secondary indicators under the discipline strength dimension The discipline strength dimension is used to measure the current level of a specific discipline in a research institution, and includes four secondary indicators: research activity, research intensity, research quality, and talent strength.
[0024] The specific statistical process for the original values of each secondary indicator is as follows: the project records of the institution under the discipline are screened from the dataset and the number of projects is counted as the original value of "Discipline Research Activity"; the total amount of funding for all projects is counted as the original value of "Discipline Research Intensity"; the average amount of funding per project is calculated as the original value of "Research Quality"; and the weighted number of talent projects under the discipline is counted as the original value of "Discipline Talent Intensity".
[0025] (4) Statistics on the original values of secondary indicators under the discipline potential dimension The subject potential dimension is used to measure the development trend of a specific subject, and includes three secondary indicators: activity growth, intensity growth, and talent potential.
[0026] The specific statistical process for the original values of each secondary indicator under this dimension is as follows: Extract the sequence of the number of projects of this discipline in the institution over the years from the dataset, calculate the average annual growth rate, and obtain the original value of "Discipline Activity Growth"; extract the sequence of the total amount of funding over the years, calculate the average annual growth rate, and obtain the original value of "Discipline Intensity Growth"; count the number of leaders of the Young Scientists Fund projects and innovative research groups under this discipline, as the original value of "Discipline Talent Potential".
[0027] The formula used to calculate the average annual growth rate is as follows: ,in This is the initial value for the first period (e.g., when calculating the growth of subject activity). (Number of projects in the first year). This is the final value at the end of the nth period (e.g., when calculating the growth of subject activity). (This refers to the number of projects in year n).
[0028] (5) Statistics on the original values of secondary indicators under the talent strength dimension The talent strength dimension is used to measure the current research capabilities of researchers, and includes four secondary indicators: project activity, funding strength, project quality, and the level of the affiliated institution.
[0029] The specific statistical process for the original values of each secondary indicator under this dimension is as follows: extract all project records of the researcher from the dataset, and count the number of projects as the original value of "project activity"; count the total amount of funding for all projects as the original value of "financial strength"; calculate the average amount of funding per project as the original value of "project quality"; obtain the institutional strength score of the institution to which the researcher belongs from the dataset as the original value of "affiliated institution level".
[0030] (6) Statistics on the original values of secondary indicators under the talent potential dimension The talent potential dimension is used to measure the growth potential of researchers and includes four secondary indicators: research diversity, research sustainability, emerging index, and rapid development potential.
[0031] The specific statistical process for the raw values of each secondary indicator under this dimension is as follows: Analyze the subject codes of the projects undertaken by the researcher from the dataset. If multiple different primary disciplines are involved, it is considered to have "research diversity" and marked as "1"; otherwise, it is marked as "0". Extract the start and end dates of each project from the dataset, calculate the duration (in months) of each project, and take the average of all projects as the raw value of "research sustainability". Find the year in which the researcher first received project funding from the dataset, calculate the difference between the current year and the year of the first funding, and take the inverse logarithmic transformation of the reciprocal of the difference to obtain the raw value of the "emergence index". The larger this value, the more recently the scholar has entered the research field. The calculation formula is as follows: The algorithm finds the number of years between the year a researcher first receives funding and the year they receive funding for the second time from the dataset. The inverse of this interval is then transformed using an inverse logarithmic transformation to obtain the original value of "rapid development potential." A larger value indicates a faster rate of success from initial funding to continuous grant approvals. The calculation formula is as follows: .
[0032] S3. Based on the original value type of each secondary indicator, perform corresponding data transformation and normalization processing to finally output the score of the secondary indicator in the range of 0-1. The specific process is as follows: First, input the original value of each secondary indicator, calculate the data distribution characteristics of the original value, determine whether the data distribution type is continuous skewed or discrete based on the calculation results, and output the distribution type determination result.
[0033] Based on the statistical characteristics of the original values of secondary indicators, they are divided into two types: continuously skewed and discrete. Continuously skewed data refers to indicators whose values vary continuously within the real number range and whose distribution exhibits a clear long tail or skewed characteristic, such as the total number of projects, total funding, average project funding, and average annual growth rate. A common characteristic of this type of data is that most entities have small values, while a few entities have extremely large values, leading to poor discriminative power when used directly for comparison. Discrete data refers to indicators whose values are a finite number of discrete integers, such as the number of master's and doctoral programs, the number of platforms, the number of talent projects, binary markers of research diversity, the number of months of project duration (rounded down), and intermediate results in the calculation of emerging indices and rapid development potential. This type of data is essentially a counting or classification result and does not satisfy the assumption of a continuous normal distribution.
[0034] Then, based on the above determination results, the original values of different types are processed using the corresponding data transformation methods to obtain the first processed value.
[0035] For continuously skewed data, the Box-Cox or Yeo-Johnson transform is used. The transformed data distribution approximates a normal distribution, eliminating the effects of skewness and outliers.
[0036] The Box-Cox transform is applicable to non-negative data, as shown in equation (1).
[0037]
[0038] The Yeo-Johnson transform is applicable to any data containing 0 or negative numbers, as shown in equation (2).
[0039] for : or ; for : or .
[0040] In equations (1) and (2), This refers to a data sample, primarily the project amount. These are deformation parameters, which are adjusted according to the actual effect.
[0041] For discrete-state data, binning discretization is used. This embodiment primarily employs equal-width and equal-frequency binning strategies. Depending on the research direction and data distribution, the binning method will also vary.
[0042] For example, the average annual growth rate of project amount can be binned using an equal-width binning strategy, and the binning mapping rules are shown in Table 1 below:
[0043] The average number of months of talent research sustainability was binned using an equal-width binning strategy, and the binning mapping rules are shown in Table 2 below:
[0044] The funding amount is binned using an equal-frequency strategy, and the binning mapping rules are shown in Table 3 below:
[0045] The bin numbers (0,1,2,3,4) and (0,1,2,3,4,5) mentioned above are the first processed values. Bin discretization converts continuous or discrete raw values into graded values, enhancing the robustness of the evaluation.
[0046] After the above transformation, each secondary index's original value corresponds to a first processed value. For continuous skewed types, the first processed value is the transformed continuous value (approximately normally distributed); for discrete types, the first processed value is the bin number (multivariate discrete value, taking integer values from 0 to 4 or 5).
[0047] After obtaining the first processed value, its data type is further determined in order to select the subsequent normalization method. The data types of the first processed value are divided into three types: continuous data, binary discrete data, and multivariate discrete data.
[0048] Continuous data refers to values that have undergone Box-Cox or Yeo-Johnson transformations, with values continuously distributed within the real number range and theoretically approximating a normal distribution. Binary discrete data refers to data that takes only two values (usually 0 and 1). This type of data mainly originates from two scenarios: first, the original indicator itself is a binary indicator (e.g., whether or not it has undertaken international collaborative projects), and it may still retain its binary nature after binning and discretization; second, some discrete data, after binning, has exactly two bins (e.g., most entities are concentrated in bin 0, and a few in bin 1). Multivariate discrete data refers to data that takes more than two discrete levels. It mainly originates from bin numbers (0, 1, 2, 3, 4, etc.) generated after binning and discretization, as well as multi-level levels formed after binning of certain original count indicators.
[0049] Next, based on the data type of the first processed value, the corresponding normalization method is selected for processing to obtain the second processed value.
[0050] For continuous data The sample set ,if , follows the mean The sum and variance are If the Gaussian distribution is used, the cumulative distribution function can be used for normalization calculation. The calculation result represents the quantile position of the entity's index value among all entities, with a value range of (0,1). The cumulative distribution function used is shown in equation (3):
[0051] In equation (3), and They are based on sample sets The sample mean and sample standard deviation, .
[0052] For the probability of successfully obeying is binary discrete data with binary discrete distribution The sample set Equation (4) is used for direct mapping normalization.
[0053]
[0054] in, For sample set Frequency of success events: .
[0055] When the first processing value is 0, the normalized score is 0; when the first processing value is 1, the normalized score is 1. This processing preserves the original semantics of the binary attribute without additional scaling.
[0056] For those that obey the probability mass function, Discrete probability distribution data The sample set Then, the empirical cumulative distribution function shown in equation (5) is used for normalization. The score represents the cumulative percentage of the entity's level among all entities, with a value range of [0,1].
[0057]
[0058] in, In the sample set middle Frequency of occurrence , .
[0059] The second processed value obtained after the above normalization process is the score of the secondary indicator, and its value range is uniformly [0,1].
[0060] Additionally, during transformation and normalization, if missing values are encountered (e.g., the average annual growth rate of an organization cannot be calculated due to insufficient data), the score of that indicator is marked as missing, and the median score of that indicator across all entities is used to fill the missing values during subsequent weighted aggregation. For data transformation... In extreme cases such as estimation failure, a fixed method can be used instead. Alternatively, a robust estimation method can be used.
[0061] S4. The scores of each secondary indicator are weighted and aggregated to calculate the strength score and potential score of the research entity. The relative deviation is calculated based on the resource allocation ratio of each discipline, and the distinctiveness index is calculated based on the relative entropy between the discipline's resource distribution and the reference distribution. This includes the following sub-steps: 1) Calculate the strength score and potential score First, retrieve the preset weights corresponding to the scores of each secondary indicator from the preset weight database. Preset weights The importance of each secondary indicator within its respective dimension is predetermined, for example, through the analytic hierarchy process (AHP) or expert surveys. Weights are fixed before the evaluation, and the same indicator uses the same weight across different evaluation subjects.
[0062] Then, each input secondary indicator score is paired and bound to its corresponding weight value to form a dataset containing all "score-weight" pairs. From the pairing set, based on the dimension to which the indicator belongs, all secondary indicator scores and their corresponding weight values belonging to the strength dimension (i.e., institutional strength, discipline strength, talent strength) are selected. Then, the secondary indicator scores under the strength dimension and the potential dimension are weighted and summed according to equation (6).
[0063]
[0064] in, For the first Scores for each secondary indicator These are the corresponding preset weights. For example, the weights of the secondary indicators of institutional strength are: research activity 0.241, research intensity 0.227, talent cultivation 0.047, innovation platform 0.247, and talent intensity 0.238.
[0065] For the strength dimension, the scores of the corresponding secondary indicators under the institutional strength dimension, discipline strength dimension, and talent strength dimension are assigned to their respective preset weights. After multiplying and summing, the strength score of the research entity is obtained. Specifically, the institutional strength score is obtained by weighted summing of the scores of five secondary indicators: institutional research activity, institutional research intensity, institutional talent cultivation, innovation platform, and institutional talent intensity; the discipline strength score is obtained by weighted summing of the scores of four secondary indicators: discipline research activity, discipline research intensity, research quality, and discipline talent intensity; and the talent strength score is obtained by weighted summing of the scores of four secondary indicators: project activity, funding strength, project quality, and the level of the affiliated institution.
[0066] Similarly, for the potential dimension, the scores of the corresponding secondary indicators under the institutional potential dimension, disciplinary potential dimension, and talent potential dimension are weighted and summed with their corresponding preset weights to obtain the potential score of the research subject. The institutional potential dimension includes institutional activity growth, institutional intensity growth, institutional talent potential, and internationalization level; the disciplinary potential dimension includes disciplinary activity growth, disciplinary intensity growth, and disciplinary talent potential; the talent potential dimension includes research diversity, research sustainability, emerging index, and rapid development potential.
[0067] After calculating the strength score and potential score, the strength score and potential score can be weighted and aggregated according to a preset ratio to obtain a comprehensive score for measuring the overall level of the research subject. The preset ratio can be adjusted according to the evaluation focus. For example, if the evaluation focuses on the current strength, the strength weight can be set to 0.7 and the potential weight to 0.3, while if the evaluation focuses on the development trend, the potential weight can be increased.
[0068] 2) Calculate the relative deviation of each subject. Relative deviation is used to measure the research subject In a certain discipline The degree of resource allocation deviation relative to the reference level is determined by the normalized logarithm relative resource index.
[0069] First, input the resource allocation data of the target research entity in each discipline, and calculate the resource quantity for each discipline. Then, calculate the target research entity's resource allocation according to equation (7). In each discipline Resource allocation ratio In other words, the proportion of major disciplines. Resources can be the number of projects or the total amount of funding.
[0070]
[0071] Secondly, according to formula (8), all research entities within the reference range are calculated in each discipline. Resource allocation ratio That is, referencing the proportion of disciplines. The reference scope can be the entire province, the entire country, or a collection of institutions of a certain type.
[0072]
[0073] Then, the relative deviation (NLRR) is calculated according to equation (9):
[0074] In equations (7), (8), and (9) As the main body of scientific research In a certain discipline The amount of resources invested.
[0075] This ratio reflects the target research entity's position in the discipline. The ratio of resource concentration on a given level to a reference level. When The value indicates that the school has a relative advantage in that subject; the larger the absolute value of NLRR, the more obvious the advantage or disadvantage.
[0076] 3) Calculate the distinctiveness index The Distinctiveness Index (IDSI) measures the overall difference between the distribution of disciplinary resources of a research subject and a reference distribution. First, it inputs the resource allocation data of the target research subject across various disciplines. Statistics on the resource volume of each discipline across all research entities Then, the relative entropy calculation formula is used for calculation, as shown in equation (10):
[0077] The Distinctiveness Index (IDSI) is suitable for measuring the distribution of subject resources in schools. Average distribution compared to a reference range (e.g., the entire province) The distance between them. The greater the distribution difference, the stronger the school's distinctiveness; if it is consistent with the distribution of the reference range, then IDSI = 0.
[0078] 4) Output the strength score of the research subject (and optional comprehensive score), the potential score of the research subject, the relative deviation of each discipline (NLRR), and the distinctiveness index (IDSI) of the research subject.
[0079] S5. Combine the input strength score and potential score according to a preset data structure. The combination method is to construct a structured data object containing two fields: one field stores the strength score and is labeled "Strength," and the other field stores the potential score and is labeled "Potential." Output a strength-potential result set containing both the strength score and the potential score. This result set directly reflects the quantitative assessment results of the research subject in two dimensions: current research scale (strength) and future development trend (potential).
[0080] The input relative deviation and distinctiveness index of each discipline are combined according to a pre-defined data structure. The relative deviation of each discipline is organized into a mapping structure based on discipline identifiers (such as discipline codes or discipline names), with each discipline identifier corresponding to a relative deviation value, forming a set of "discipline-deviation" key-value pairs. Positive values indicate that the discipline has a relative advantage, while negative values indicate a relative disadvantage. The distinctiveness index is stored separately as a scalar value and labeled with the "distinctiveness" identifier, reflecting the overall distinctiveness of the research entity's discipline layout. The output is a distinctiveness assessment result set containing the relative deviation and distinctiveness index of each discipline. This result set reveals the structural characteristics and uniqueness of the research entity's discipline resource allocation.
[0081] The two input result sets are merged, integrating all fields from the strength and potential result set with all fields from the characteristic assessment result set into a single data structure. This results in a complete assessment result set containing strength scores, potential scores, relative deviation sets for each discipline, and a characteristic index. The merged complete assessment result set is then output. The output format can be a structured data format (such as JSON or XML), a database record, or a visual report. This assessment result set encompasses the overall strength level of the research entity, its development potential, the relative strength and weakness distribution of each discipline, and the overall disciplinary layout characteristics, achieving a multi-dimensional and comprehensive presentation of assessment information.
[0082] The aforementioned evaluation results encompass the overall strength and development potential of the research entity, the relative strengths and weaknesses of each discipline, and the overall characteristics of its disciplinary structure, presenting multi-dimensional and comprehensive evaluation information. Through these results, users can intuitively obtain a comprehensive evaluation profile of the research entity, supporting subsequent research management decisions, resource allocation, talent recruitment, and disciplinary development planning.
[0083] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A scientific research strength evaluation method based on the synergistic quantification of strength and potential, characterized in that, include: S1. Obtain project data, innovation platform data, and degree program data from research entities. After cleaning and unifying the format of the three types of data, merge them to form a structured dataset. S2. Extract fields from the dataset according to the three levels of institution, discipline and talent, perform statistical calculations on the extracted fields, and generate original values of multiple secondary indicators under the strength dimension and potential dimension. S3. Based on the original value type of each secondary indicator, perform corresponding data transformation and normalization to obtain the score of each secondary indicator; S4. The scores of each secondary indicator are weighted and aggregated to calculate the strength score and potential score of the research subject. The relative deviation is calculated based on the resource allocation ratio of each discipline, and the characteristic index is calculated based on the relative entropy between the discipline resource distribution and the reference distribution. S5. Combine the strength score, potential score, relative deviation, and distinctiveness index into structured data and output it as the evaluation result.
2. The scientific research strength evaluation method based on the synergistic quantification of strength and potential as described in claim 1, characterized in that, S1 specifically involves: extracting project records of research entities from the project funding management system. Each project record includes the project name, project leader, funding amount, and project year, forming a project data set. Platform identification records of research entities are extracted from the platform management system. Each platform record includes the platform name, platform level, and platform category, forming a platform dataset. Degree program records of research entities are extracted from the degree authorization management system. Each degree program record includes the number of master's programs and the number of doctoral programs, forming a degree program dataset. Duplicate record removal and invalid record filtering are performed on the project dataset, platform dataset, and degree program dataset, respectively. The three types of datasets are then linked and merged according to the unified identifier of the research entity to form a structured dataset.
3. The scientific research strength evaluation method based on the synergistic quantification of strength and potential as described in claim 1, characterized in that, The strength dimension includes institutional strength, disciplinary strength, and talent strength, while the potential dimension includes institutional potential, disciplinary potential, and talent potential.
4. The scientific research strength evaluation method based on the synergistic quantification of strength and potential as described in claim 1, characterized in that, S3 specifically involves: inputting the original value of each secondary indicator, calculating the data distribution characteristics of the original value, determining whether the data distribution type is continuous skewed or discrete based on the calculation result, and outputting the distribution type determination result; inputting the original value of the secondary indicator and the distribution type determination result into the data transformation processing stage; if the distribution type is continuous skewed, using Box-Cox transformation or Yeo-Johnson transformation to convert the original value into a continuous value with an approximate normal distribution; if the distribution type is discrete, using binning discretization to map the original value into a level value; and outputting the first processed value after transformation processing; inputting the first processed value and determining its data type; if the data type is continuous data, using cumulative distribution function normalization to convert the first processed value into a cumulative probability value; if the data type is binary discrete data, using direct mapping normalization to map the first processed value into a binary score value; and if the data type is multivariate discrete data, using empirical cumulative distribution function normalization to convert the first processed value into an empirical cumulative proportion value; and outputting the second processed value as the score of the secondary indicator after normalization processing.
5. The scientific research strength evaluation method based on the synergistic quantification of strength and potential as described in claim 1, characterized in that, The specific steps for calculating the strength score and potential score of the research subject by weighting and aggregating the scores of each secondary indicator are as follows: input the scores of each secondary indicator, read the weight value corresponding to each secondary indicator score from the preset weight database, and output the pairing set of secondary indicator scores and weight values. The scores of secondary indicators under the strength dimension and their corresponding weight values are selected from the pairing set, and a weighted summation calculation is performed. The calculation result is output as the strength score of the research subject. The scores of secondary indicators under the potential dimension and their corresponding weight values are selected from the pairing set, and a weighted summation calculation is performed. The calculation result is output as the potential score of the research subject.
6. The scientific research strength evaluation method based on the synergistic quantification of strength and potential as described in claim 1, characterized in that, S4 also includes weighting and aggregating the strength score and potential score according to a preset ratio to obtain a comprehensive score for measuring the overall level of the scientific research subject.
7. The scientific research strength evaluation method based on the synergistic quantification of strength and potential as described in claim 1, characterized in that, The calculation of relative deviation based on the resource allocation ratio of each discipline is specifically as follows: Input the resource allocation data of the target research subject in each discipline, calculate the resource quantity of each discipline, divide the resource quantity of each discipline by the total resource quantity of all disciplines, and generate the subject discipline ratio distribution; Input the resource allocation data of all research subjects within the reference range in each discipline, calculate the resource quantity of each discipline, divide the resource quantity of each discipline by the total resource quantity of all disciplines, and generate the reference discipline ratio distribution; Input the subject discipline ratio distribution and the reference discipline ratio distribution in a one-to-one correspondence for each discipline, perform a division operation of the subject ratio value by the reference ratio value for each discipline, take the logarithm of the division result, and output the logarithmic result as the relative deviation of that discipline.
8. The scientific research strength evaluation method based on the synergistic quantification of strength and potential as described in claim 1, characterized in that, The specific steps for calculating the distinctiveness index based on the relative entropy between the disciplinary resource distribution and the reference distribution are as follows: Input the resource allocation data of the target research subject in each discipline, calculate the resource quantity of each discipline, divide the resource quantity of each discipline by the total resource quantity of all disciplines, and generate the subject probability distribution; Input the resource allocation data of all research subjects in each discipline within the reference range, calculate the resource quantity of each discipline, divide the resource quantity of each discipline by the total resource quantity of all disciplines, and generate the reference probability distribution; Input the subject probability distribution and the reference probability distribution one-to-one by discipline, perform a division operation of the subject probability value divided by the reference probability value for each discipline, take the logarithm of the division result, multiply the logarithm of each discipline by the subject probability value of that discipline, sum the product results of all disciplines, and output the sum as the distinctiveness index of the research subject.
9. The scientific research strength evaluation method based on the synergistic quantification of strength and potential as described in claim 1, characterized in that, Specifically, S5 involves: combining the strength score and potential score into a strength-potential result set; combining the relative deviation and distinctiveness index into a distinctiveness assessment result set; merging the strength-potential result set and the distinctiveness assessment result set to form a complete assessment result set containing the strength score, potential score, relative deviation, and distinctiveness index; and outputting this complete assessment result set as the assessment result.