Methods and apparatus for constructing institutional evaluation models

By constructing an institutional evaluation model and combining structured and unstructured data evaluation models with weight ratios, the irrationality caused by the single indicator in the traditional scientific research institution evaluation system is solved, and a more accurate evaluation of scientific research institutions is achieved.

CN120494620BActive Publication Date: 2025-12-02SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510580133.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-12-02
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Traditional evaluation systems for research institutions use single-indicator metadata, resulting in evaluation results that are not reasonable or reliable and cannot fully reflect the comprehensive capabilities of research institutions.

Method used

An institutional evaluation model is constructed by acquiring the basic elements, scientific research capabilities, and carrying capacity elements of sample institutions, separating them into structured and unstructured datasets, and converting them into feature vectors and floating-point vectors. The evaluation models of unstructured and structured data are used to predict scores, and the final evaluation model is constructed by combining the weight ratios.

Benefits of technology

It improved the accuracy of institutional evaluations, enabling rapid and accurate prediction of evaluation scores for the institutions being evaluated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494620B_ABST
    Figure CN120494620B_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for constructing an institutional evaluation model. The method includes: acquiring basic elements, research capability elements, and carrying capacity elements of sample institutions; extracting structured and unstructured datasets; converting the unstructured datasets into floating-point vectors and the structured datasets into feature vectors; inputting the floating-point vectors and evaluation score labels into the unstructured data evaluation model for score prediction, and inputting the feature vectors and evaluation score labels into the structured data evaluation model for score prediction; selecting target weight ratios; and constructing an institutional evaluation model based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratios. This invention improves the accuracy of institutional evaluation models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of institutional evaluation technology, and in particular to a method and apparatus for constructing an institutional evaluation model. Background Technology

[0002] Traditional research institution evaluation systems typically use single-indicator metadata to evaluate research institutions. For example, for universities, the evaluation result is obtained by combining the total number of representative academic papers published and the total number of citations of those papers. Traditional research institution evaluation systems that rely on single-indicator metadata consider only a limited range of factors and cannot provide a reasonable and reliable evaluation of research institutions. Summary of the Invention

[0003] This invention provides a method and apparatus for constructing an institutional evaluation model, which can improve the accuracy of the institutional evaluation model and thus enable rapid and accurate prediction of the evaluation score of the institution to be evaluated.

[0004] On the one hand, the present invention provides a method for constructing an institutional evaluation model, the method comprising:

[0005] The basic elements, scientific research capabilities, and carrying capacity of the sample institutions are obtained as the sample source data; the sample source data is labeled with the sample evaluation score tags of the sample institutions.

[0006] Based on the basic elements of the sample institutions and the elements of the sample scientific research capabilities, numerical data and character data of the samples are extracted as a sample structured dataset.

[0007] Extract the sample long text from the sample scientific research capability element and the sample institutional carrying capacity element to form the sample unstructured dataset;

[0008] The unstructured dataset is converted into a floating-point vector, and the structured dataset is converted into a feature vector.

[0009] The sample floating-point vector and the sample evaluation score label are input into the unstructured data evaluation model for score prediction to obtain the first sample score; and the sample feature vector and the sample evaluation score label are input into the structured data evaluation model for score prediction to obtain the second sample score.

[0010] Traverse the combined weight set of the model, and select the target weight ratio from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label;

[0011] An institutional evaluation model is constructed based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio.

[0012] On the other hand, an apparatus for constructing an institutional evaluation model is provided, the apparatus comprising:

[0013] The sample data acquisition module is used to acquire the basic elements, scientific research capabilities, and carrying capacity of the sample institutions as the sample source data; the sample source data is labeled with the sample evaluation score tags of the sample institutions.

[0014] The sample structured data determination module is used to extract sample numerical data and sample character data based on the sample institution's basic elements and the sample's scientific research capability elements, as a sample structured dataset.

[0015] The unstructured data determination module is used to extract long text from the scientific research capability elements and institutional carrying capacity elements of the samples, as the unstructured dataset of the samples.

[0016] The sample vector conversion module is used to convert the unstructured sample dataset into a sample floating-point vector and the structured sample dataset into a sample feature vector.

[0017] The sample score prediction module is used to input the sample floating-point vector and the sample evaluation score label into the unstructured data evaluation model to predict the score and obtain the first sample score; and to input the sample feature vector and the sample evaluation score label into the structured data evaluation model to predict the score and obtain the second sample score.

[0018] The target weight ratio determination module is used to traverse the combined weight set of the model and select the target weight ratio from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label.

[0019] The evaluation model construction module is used to construct an institutional evaluation model based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio.

[0020] On the other hand, an electronic device is provided, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the method for constructing the institutional evaluation model as described above.

[0021] On the other hand, a computer storage medium is provided that stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the method for constructing the institutional evaluation model as described above.

[0022] On the other hand, a computer program product is provided, including a computer program that is loaded and executed by a processor to implement the method for constructing the institutional evaluation model as described above.

[0023] The method and apparatus for constructing an institutional evaluation model provided by this invention have the following technical effects:

[0024] This invention acquires the basic elements, research capability elements, and carrying capacity elements of sample institutions as sample source data. The sample source data is labeled with the sample evaluation score tags of the sample institutions. Based on the basic elements and research capability elements, numerical and character data are extracted as a structured dataset. Long text is extracted from the research capability and carrying capacity elements as an unstructured dataset, thus dividing the sample source data into two main categories. The unstructured dataset is converted into floating-point vectors, and the structured dataset is converted into feature vectors. The floating-point vectors and the evaluation score tags are input into an unstructured data evaluation model. A score prediction is performed to obtain a first sample score; then, the sample feature vector and the sample evaluation score label are input into a structured data evaluation model to predict the score, resulting in a second sample score; thus, the evaluation scores of the sample institutions predicted by the two types of data are obtained respectively; next, the combined weight set of the model is traversed, and a target weight ratio is selected from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label; an institution evaluation model is constructed based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio; the institution evaluation model constructed using the method of this invention has higher accuracy, thereby enabling rapid and accurate prediction of the evaluation score of the institution to be evaluated. Attached Figure Description

[0025] To more clearly illustrate the technical solutions and advantages in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1This is a schematic diagram of the application environment of a method for constructing an institutional evaluation model provided in the embodiments of this specification;

[0027] Figure 2 This is a flowchart illustrating a method for constructing an institutional evaluation model provided in the embodiments of this specification;

[0028] Figure 3 This is a flowchart illustrating a training method for an unstructured data evaluation model provided in the embodiments of this specification.

[0029] Figure 4 This is a flowchart illustrating a training method for a structured data evaluation model provided in the embodiments of this specification;

[0030] Figure 5 This is a flowchart illustrating a method for selecting a target weight ratio from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label, as provided in an embodiment of this specification.

[0031] Figure 6 This is a flowchart illustrating a method for determining the combined evaluation index value of the model corresponding to each weight ratio based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label, as provided in an embodiment of this specification.

[0032] Figure 7 This is a flowchart illustrating a method for determining the evaluation score of an organization to be evaluated, as provided in the embodiments of this specification.

[0033] Figure 8 This is a schematic diagram of the structure of a system for constructing an institutional evaluation model provided in the embodiments of this specification;

[0034] Figure 9 This is a schematic diagram illustrating a method for calculating a first sample score provided in an embodiment of this specification;

[0035] Figure 10 This is a schematic diagram of the structure of a device for constructing an organizational evaluation model provided in the embodiments of this specification;

[0036] Figure 11 This is a schematic diagram of the structure of a server provided in the embodiments of this specification. Detailed Implementation

[0037] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0039] The following describes a method for constructing an institutional evaluation model according to the present invention. Figure 1 This is a flowchart illustrating a method for constructing an institutional evaluation model provided in the embodiments of this specification. This specification provides the operational steps described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or server product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown in the embodiments or drawings... Figure 1 As shown, the method may include:

[0040] S101: Obtain the basic elements, scientific research capabilities, and carrying capacity elements of the sample institutions as sample source data; the sample source data is labeled with the sample evaluation score tags of the sample institutions.

[0041] In the embodiments of this specification, the method of this embodiment can be applied to a system for constructing an institutional evaluation model. This system may include an unstructured data evaluation model, a structured data evaluation model, and a weight calculation layer. The sample institutions can be research institutions, including but not limited to research departments of universities, independent research institutes, and R&D departments within enterprises. The sample source data of the sample institutions may include at least one of the following: basic elements of the sample institution, research capability elements of the sample institution, and carrying capacity elements of the sample institution. For example, the sample source data is multi-source data, including basic elements of the sample institution, research capability elements of the sample institution, and carrying capacity elements of the sample institution. The sample source data may include one or more institutional element libraries.

[0042] The basic elements of the sample institutions are set at two levels, including the project source category and the basic information of the institution. The project source category is classified into the number of local tasks, the number of national tasks, the number of domestic commissioned projects, the number of tasks independently deployed by the research institute, the number of tasks of the academy units, and the number of other tasks. The basic information of the institution is set at three levels, including the institution team, the institution budget, and the number of projects. The institution team is classified into the number of full professors, associate professors, intermediate-level staff, junior staff, ordinary staff, doctoral staff, master's staff, and bachelor's staff.

[0043] The research capabilities of the sample institutions are categorized into five levels: invention patents, published papers, monographs, scientific and technological reports, and technology transfer. Invention patents include the number of domestic patents granted, the number of foreign patents granted, the total number of invention patents, the number of utility model patents, the number of design patents, and the total number of citations. Published papers include the total number of published papers and the total number of citations. Monographs include the number of monographs. Scientific and technological reports include the number of reports and details of breakthroughs in key core technologies. Breakthroughs in key core technologies and technology transfer will be described in long-form text.

[0044] The sample organization's capacity elements are set at five levels, including project level, material resources, information resources, technical resources, and human resources. Among them, material resources, information resources, technical resources, and human resources will be described in long text format. The project level is classified into project, sub-project, and task.

[0045] S102: Based on the basic elements of the sample organization and the elements of the sample scientific research capabilities, extract the sample numerical data and sample character data to form a sample structured dataset.

[0046] In the embodiments of this specification, sample numerical data and sample character data can be extracted from the sample institution's basic elements and sample research capability elements. The sample character data and sample numerical data can be a set of data that are related; for example, if the institution includes 30 undergraduate students, and the sample character data represents the number of undergraduate students, then the sample numerical data is 30.

[0047] S103: Extract the sample long text from the sample scientific research capability element and the sample institutional carrying capacity element, and use it as the sample unstructured dataset.

[0048] In the embodiments of this specification, the sample long text is text with a length greater than a preset threshold; both the sample scientific research capability element and the sample institutional carrying capacity element include long text content, and the sample long text such as the breakthrough content of key core technologies and the content of achievement transformation in the sample scientific research capability element can be extracted as the sample unstructured dataset; the material resources, information resources, technical resources, human resources, etc. in the sample institutional carrying capacity element can be extracted as the sample unstructured dataset.

[0049] S104: Convert the unstructured dataset of samples into a floating-point vector of samples, and convert the structured dataset of samples into a feature vector of samples.

[0050] In the embodiments of this specification, the unstructured dataset is converted into a sample floating-point vector for inputting into the unstructured data evaluation model, and the structured dataset is converted into a sample feature vector for inputting into the structured data evaluation model.

[0051] S105: Input the sample floating-point vector and the sample evaluation score label into the unstructured data evaluation model to predict the score and obtain the first sample score; and input the sample feature vector and the sample evaluation score label into the structured data evaluation model to predict the score and obtain the second sample score.

[0052] In the embodiments of this specification, the unstructured data evaluation model is a model trained based on the unstructured data of historical sample institutions and the historical sample evaluation score labels, and the structured data evaluation model is a model trained based on the structured data of historical sample institutions and the historical sample evaluation score labels. The evaluation score of the sample institution, i.e., the first sample score, can be obtained through the sample floating-point vector and the unstructured data evaluation model; the evaluation score of the sample institution, i.e., the second sample score, can be obtained through the sample feature vector and the structured data evaluation model. This facilitates subsequent adjustment of the weights of the two models based on the first and second sample scores.

[0053] S106: Traverse the combined weight set of the model, and select the target weight ratio from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label.

[0054] In the embodiments of this specification, after obtaining the evaluation scores corresponding to the two models, a combined weight set can be input into the weight calculation layer to adjust the weights of the two models and obtain the optimal target weight ratio. The combined weight set of the models is a set of multiple preset weight ratios set for the unstructured data evaluation model and the structured data evaluation model. The weights corresponding to the unstructured data evaluation model and the structured data evaluation model can be determined according to each preset weight ratio in the combined weight set, wherein the sum of the weights of the two models is 1. The optimal target weight ratio can be selected from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label.

[0055] S107: Construct an institutional evaluation model based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio.

[0056] In the embodiments of this specification, the first weight of the prediction result of the unstructured data evaluation model and the second weight of the preset result of the structured data evaluation model can be determined according to the unstructured data evaluation model and the target weight ratio corresponding to the structured data evaluation model, thereby constructing an institution evaluation model for predicting institution evaluation scores.

[0057] In some embodiments, the step of extracting sample numerical data and sample character data as a sample structured dataset based on the sample institution's basic elements and the sample's scientific research capability elements includes:

[0058] Extract the number of projects, the number of tasks per project, the budget for each project, and the number of people at each level in the organization's basic elements from the sample organization to obtain the first sample of structured data;

[0059] The number of research achievements is extracted from the research capability elements of the sample to obtain the structured data of the second sample.

[0060] The sample structured dataset is constructed based on the first sample structured data and the second sample structured data.

[0061] In the embodiments of this specification, the number of projects, the number of tasks per project, the budget for each project, and the number of people at each level in the institutional team can be extracted from the basic elements of the sample institution to obtain the first sample structured data. For example, the first sample structured data includes the number of local tasks, national tasks, domestic commissioned projects, research institute self-deployment projects, CAS unit tasks, and other tasks corresponding to the project source categories in the basic elements of the sample institution; and the number of institutional teams classified into senior professional, associate senior professional, intermediate, junior, ordinary, doctoral, master's, and bachelor's degree holders.

[0062] The number of research achievements in the research capability elements of the sample can be extracted to obtain the second sample structured data. For example, the second sample structured data includes the number of domestic patents granted, the number of foreign patents granted, the number of invention patents, the number of utility model patents, the number of design patents, and the total number of citations in the research capability elements of the sample; the total number of published papers and the total number of citations in the published papers; the number of monographs; and the number of scientific and technological reports in the scientific and technological reports.

[0063] Finally, the set formed by the first sample structured data and the second sample structured data is determined as the sample structured dataset, thereby realizing the accurate extraction of the sample structured dataset from the sample source data.

[0064] In some embodiments, such as Figure 2 As shown, the process of converting the unstructured dataset into a floating-point vector and the structured dataset into a feature vector includes:

[0065] S1041: Input the unstructured dataset of the sample into the unstructured data processing layer for data cleaning to obtain the first standard dataset of the sample;

[0066] S1042: Input the first sample standard dataset into the first vector representation layer to extract text semantics, and obtain the sample floating-point vector;

[0067] S1043: Input the sample structured dataset into the structured data processing layer for data cleaning to obtain the second sample standard dataset;

[0068] S1044: Input the second sample standard dataset into the second vector representation layer for data correlation analysis to obtain the sample feature vector.

[0069] In the embodiments of this specification, the standard dataset consisting of the first sample standard dataset and the second sample standard dataset can also be transformed into a standardized dataset usable by the model through a series of data preprocessing steps, including text cleaning, word segmentation, stop word removal, normalization, and missing value imputation. The institutional evaluation model construction system of this embodiment also includes a data processing layer, which may include an unstructured data processing layer and a structured data processing layer. The data processing layer mainly performs data extraction and preprocessing on the collected institutional data to output a dataset that meets the model input requirements, laying the foundation for subsequent construction of the evaluation model. The data preprocessing process performs different preprocessing operations according to the differences in data categories. For unstructured data, text segmentation, stop word removal, and stemming are performed; for structured data, data cleaning, missing value imputation, and outlier handling are performed.

[0070] In the process of transforming the unstructured dataset into floating-point vectors, the unstructured dataset can first be input into an unstructured data processing layer for data cleaning to obtain a first standard dataset. Then, the first standard dataset is input into a first vector representation layer for text semantic extraction to obtain the floating-point vector (high-dimensional vector). The high-dimensional vector is obtained by processing long text inputs through a multilingual model, mapping sentences and paragraphs to a dense vector space, thus obtaining a floating-point vector representation of the text. The unstructured data processing layer can be a pre-trained structure layer for cleaning unstructured datasets; the first vector representation layer can be a pre-trained structure layer for extracting text semantic features from the first standard dataset, and the extracted text semantic features are the floating-point vectors. These floating-point vectors conform to the input data rules of the unstructured data evaluation model, thus facilitating the unstructured data evaluation model to further predict institutional evaluation scores based on the floating-point vectors.

[0071] In the process of transforming the structured dataset into sample feature vectors, the structured dataset can first be input into a structured data processing layer for data cleaning to obtain a second standard dataset. Then, the second standard dataset is input into a second vector representation layer for data correlation analysis to obtain sample feature vectors (feature vectors) that characterize data correlation. These feature vectors are extracted and transformed by analyzing the inherent correlations within the structured data. The structured data processing layer can be a pre-trained structured layer used for cleaning the structured dataset; the second vector representation layer can be a pre-trained structured layer used to extract correlation features from the second standard dataset, and the extracted correlation features are the sample feature vectors. These sample feature vectors are vectors that conform to the input data rules of the structured data evaluation model, thus facilitating the structured data evaluation model to further predict institutional evaluation scores based on the sample feature vectors.

[0072] In some embodiments, such as Figure 3 As shown, the training method for the unstructured data evaluation model includes:

[0073] S301: Input the sample floating-point vector into the unstructured data evaluation network to predict the evaluation score and obtain the first predicted score;

[0074] S302: Determine the first loss data based on the difference between the first predicted score and the sample evaluation score label;

[0075] S303: Adjust the parameters of the unstructured data processing layer, the first vector representation layer, and the unstructured data evaluation network according to the first loss data until the training termination condition is met, and determine the unstructured data evaluation network at the end of training as the unstructured data evaluation model.

[0076] In the embodiments of this specification, the unstructured data evaluation model is a model generated by training an unstructured data evaluation network based on sample floating-point vectors and sample target vectors corresponding to sample evaluation score labels. The unstructured data evaluation network can be a deep learning algorithm network. Specifically, the sample floating-point vectors and sample evaluation score labels can be input into the unstructured data evaluation network to obtain a first vector corresponding to a first predicted score and a sample target vector corresponding to the sample evaluation score label, respectively. Then, the difference between the first vector and the sample target vector is calculated as the difference between the first predicted score and the sample evaluation score label, resulting in first loss data. The parameters of the unstructured data processing layer, the first vector representation layer, and the unstructured data evaluation network are then adjusted based on the first loss data until the training termination condition is met. The unstructured data evaluation network at the end of training is then determined as the unstructured data evaluation model. The training termination condition can be determined based on the first loss data and / or the number of training iterations.

[0077] In some embodiments, the unstructured data processing layer at the end of training, the first vector representation layer, and the unstructured data evaluation network can be combined to form a first score prediction model; in the application process, only the unstructured dataset of the institution to be evaluated needs to be input into the first score prediction model to obtain the first evaluation score.

[0078] In some embodiments, such as Figure 4 As shown, the training method for the structured data evaluation model includes:

[0079] S401: Input the sample feature vector into the structured data evaluation network to predict the evaluation score and obtain the second predicted score;

[0080] S402: Determine the second loss data based on the difference between the second predicted score and the sample evaluation score label;

[0081] S403: Adjust the parameters of the structured data processing layer, the second vector representation layer, and the structured data evaluation network according to the second loss data until the training termination condition is met, and determine the structured data evaluation network at the end of training as the structured data evaluation model.

[0082] In the embodiments of this specification, the structured data evaluation model is a model generated by training a structured data evaluation network based on sample feature vectors and sample target vectors corresponding to sample evaluation score labels. The structured data evaluation network can be a linear model. Specifically, the sample feature vectors and sample evaluation score labels can be input into the structured data evaluation network to obtain a second vector corresponding to the second predicted score and a sample target vector corresponding to the sample evaluation score label, respectively. Then, the difference between the second vector and the sample target vector is calculated as the difference between the second predicted score and the sample evaluation score label, resulting in second loss data. The parameters of the structured data processing layer, the second vector representation layer, and the structured data evaluation network are then adjusted based on the second loss data until the training termination condition is met. The structured data evaluation network at the end of training is then determined as the structured data evaluation model. The training termination condition can be determined based on the second loss data and / or the number of training iterations.

[0083] In some embodiments, the structured data processing layer, the second vector representation layer, and the structured data evaluation network at the end of training can be combined into a second score prediction model; in the application process, only the structured dataset of the institution to be evaluated needs to be input into the second score prediction model to obtain the second evaluation score.

[0084] In some embodiments, such as Figure 5 As shown, the combined weight set of the traversal model, based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label, filters out the target weight ratio from multiple weight ratios, including:

[0085] S1061: Traverse the combined weight set of the model, and determine the combined evaluation index value of the model corresponding to each weight ratio based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label.

[0086] S1062: The optimal weight ratio that satisfies the preset conditions for the combined evaluation index values ​​is determined as the target weight ratio.

[0087] In the embodiments of this specification, the combined evaluation index value can be the numerical value corresponding to a single evaluation index, or it can be the average or weighted average of multiple evaluation indices. Different institutions can choose different evaluation indices based on their actual circumstances. Evaluation indices can include, but are not limited to, accuracy and recall. Specifically, the combined weight set can be input into the weight calculation layer, and then the combined weight set of the model can be traversed. Based on the weight ratio of each weight in the combined weight set, the first sample score, and the second sample score, the final sample evaluation score is determined. Then, based on the difference between the sample evaluation score and the sample evaluation score label, the target loss data is determined. Finally, the model parameters of the unstructured data evaluation model and the structured data evaluation model are fine-tuned based on the target loss data. For example, if the combined evaluation index value is accuracy, the termination condition for model fine-tuning training can be determined as the target loss data being less than a preset threshold, and the weight ratio at the end of fine-tuning training can be determined as the optimal target weight ratio. The unstructured data evaluation model and the structured data evaluation model at the end of fine-tuning training can then be used as the models in application.

[0088] In some embodiments, such as Figure 6 As shown, the step of traversing the combined weight set of the model and determining the combined evaluation index value of the model corresponding to each weight ratio based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label includes:

[0089] S10611: Input the combined weight set of the model into the weight calculation layer, traverse the combined weight set, and adjust the weight ratio between the unstructured data evaluation model and the structured data evaluation model;

[0090] S10612: Determine the sample prediction evaluation score of the sample institution under each weight ratio based on the first sample score, the second sample score, and each weight ratio;

[0091] S10613: Determine the combined evaluation index value of the model corresponding to each weight ratio based on the difference between the sample predicted evaluation score and the sample evaluation score label under each weight ratio.

[0092] In the embodiments of this specification, in the process of determining the target weight ratio, it is necessary to calculate the product of the first sample score and the first weight ratio to obtain the first product, and calculate the product of the second sample score and the second weight ratio to obtain the second product; then calculate the sum of the first product and the second product to obtain the sample predicted evaluation score; finally, based on the difference between the sample predicted evaluation score under each weight ratio and the sample evaluation score label, the combined evaluation index value of the model corresponding to each weight ratio can be determined.

[0093] In some embodiments, after constructing the institutional evaluation model based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio, as follows: Figure 7 As shown, the method includes:

[0094] S701: Obtain the basic elements, scientific research capabilities, and institutional carrying capacity of the institution to be evaluated as the source data for evaluation;

[0095] S702: Extract structured datasets based on the basic elements and research capability elements of the institution; and extract long texts from the research capability elements and the institution's carrying capacity elements as unstructured datasets.

[0096] S703: Convert the unstructured dataset into a floating-point vector and the structured dataset into a feature vector;

[0097] S704: Input the floating-point vector into the unstructured data evaluation model to predict the score and obtain the first evaluation score;

[0098] S705: Input the feature vector into the structured data evaluation model to predict the score and obtain the second evaluation score;

[0099] S706: Based on the first evaluation score, the second evaluation score, and the target weight ratio, the evaluation score of the organization to be evaluated is obtained.

[0100] In the embodiments of this specification, the classification and preprocessing methods of the input data during the application of the model are the same as those of the sample source data. The model can predict the first evaluation score of the organization to be evaluated based on the unstructured dataset and the second evaluation score based on the structured dataset. Then, based on the first evaluation score, the second evaluation score, and the target weight ratio, the evaluation score of the organization to be evaluated is obtained. Thus, the evaluation score of the organization to be evaluated can be predicted quickly and accurately through the two models.

[0101] In some embodiments, obtaining the evaluation score of the organization to be evaluated based on the first evaluation score, the second evaluation score, and the target weight ratio includes:

[0102] The first weight of the unstructured data evaluation model and the second weight of the structured data evaluation model are determined based on the target weight ratio.

[0103] Calculate the first product of the first evaluation score and the first weight, and the second product of the second evaluation score and the second weight;

[0104] The sum of the first product and the second product is calculated to obtain the evaluation score of the organization to be evaluated.

[0105] In the embodiments of this specification, after determining the optimal target weight ratio of the two models, the first product of the first evaluation score and the first weight, and the second product of the second evaluation score and the second weight can be calculated; then the sum of the first product and the second product is calculated to obtain the evaluation score of the organization to be evaluated. Thus, the evaluation score of the organization to be evaluated is further improved by the weighted summation score calculation method.

[0106] Specifically, in the embodiments of this specification, such as Figure 8 As shown, Figure 8 This is a schematic diagram of the structure of a system for constructing an institutional evaluation model according to this embodiment. It includes an input layer, a data processing layer, a vector representation layer, a model layer, a weight calculation layer, and an output layer. The input layer is used to input long text data, character-based data, and numerical data. The data processing layer is used to preprocess the long text, numerical, and character-based data in the source data to output a standard dataset that meets the input criteria of the model layer. The vector representation layer is used to extract the semantic content of long text in the unstructured dataset and convert it into a high-dimensional vector; analyze the correlation of the structured dataset and convert it into a feature vector; and convert historical institutional ratings into target vectors. The model layer receives the high-dimensional vectors and feature vectors output by the vector representation layer, uses deep learning and machine learning algorithms to train the model, and constructs both unstructured and structured data evaluation models. Simultaneously, the accuracy of the generated models is evaluated to ensure their effectiveness and reliability. The weight calculation layer calculates the weight of each model based on the importance of each model in the evaluation model. The output layer outputs the final evaluation score.

[0107] The data processing layer primarily performs data extraction and preprocessing on the collected institutional data to output a dataset that meets the model's input requirements, laying the foundation for subsequent evaluation model construction. The data preprocessing process executes different operations based on the data type. For unstructured data, it performs text segmentation, stop word removal, and stemming; for structured data, it performs data cleaning, missing value imputation, and outlier handling.

[0108] The purpose of the vector representation layer is to transform data into vectors suitable for artificial intelligence algorithms. Standard datasets include floating-point vectors, feature vectors, and target vectors.

[0109] For example, based on the basic elements, scientific research capabilities, and carrying capacity of the sample institutions, the extracted structured and unstructured datasets can also be processed using a classification model. A classification model for the sample source data can be pre-trained, and then multiple sample source data, such as the basic elements, scientific research capabilities, and carrying capacity of the sample institutions, can be input into the model to automatically divide the sample source data into structured and unstructured datasets. This can further improve the classification efficiency of structured and unstructured data and enhance the construction efficiency of the institution evaluation model.

[0110] For example, it is possible to Figure 8 The input layer, data processing layer, vector representation layer, model layer, weight calculation layer, and output layer are jointly trained. The loss is constructed based on the difference between the sample evaluation score output by the output layer and the sample evaluation score label of the sample institution. Multiple structural layers are trained synchronously, and the combination of each structural layer at the end of training is used as the institution evaluation model. In the application process, the unstructured dataset and structured dataset of the institution to be evaluated can be directly input into the two corresponding input layers of the institution evaluation model for data processing, so that the final evaluation score of the institution evaluation model can be directly output through the output layer.

[0111] For example, a method for constructing an institutional evaluation model based on this system includes:

[0112] S1. Obtain the source data and the number of combination rules, then proceed to S2;

[0113] S2. The source data is divided into unstructured datasets and structured datasets. Output the unstructured dataset and structured dataset, and continue to S3.

[0114] S3. Process the unstructured dataset, structured dataset, and historical institutional ratings to output floating-point vectors, feature vectors, and target vectors, and continue to S4.

[0115] S3.1. The Sentence-BERT method is used to process unstructured datasets, mapping sentences and paragraphs to a high-dimensional dense vector space and outputting floating-point vectors.

[0116] S3.2 Perform normalization, one-hot encoding, and other operations on the structured dataset and historical institutional ratings, and output feature vectors and target vectors;

[0117] S4. Determine whether the standard dataset conforms to the model input rules. If it does not conform, return to S3. If it does conform, continue to S5.

[0118] S5. Train the DNN algorithm based on the floating-point vector and the target vector to generate the DL_model evaluation model (unstructured data evaluation model), train the linear algorithm based on the feature vector and the target vector to generate the ML_model evaluation model (structured data evaluation model), output DL_model and ML_model, and continue to S6;

[0119] S6. Output the combination weight set according to the preset number of combination rules;

[0120] S7. Traverse the combined weight set, adjust the weight ratio of DL_model and ML_model, and calculate the combined evaluation index value of the model.

[0121] S8. Select the combined model with the lowest combined evaluation index value, and output the evaluation model and combined weight value.

[0122] The vector representation layer outputs a standard dataset, including a floating-point vector, a feature vector, and a target vector corresponding to the sample evaluation score label.

[0123] Specifically, a floating-point vector T can be generated for each piece of text in the sample source data. i The calculation formula is as follows:

[0124]

[0125] The floating-point vectors of all the data are combined into a floating-point vector set T as follows:

[0126] T = [T1 T2 … T] m ]

[0127] Where i represents the long text information of the i-th data, and v and n represent the size of the floating-point vector space of the long text;

[0128] Specifically, after removing highly correlated attributes, the remaining attributes are used as features, and the feature dimension F is constructed as follows:

[0129]

[0130] Where d is the number of features; m is the number of data records; f ij Let be the j-th dimension and the i-th data point in the feature dimensions.

[0131] Specifically, the historical institutional ratings (sample evaluation score labels) constitute the target dimension S as follows:

[0132]

[0133] Among them, s i Give each data point a historical score.

[0134] Specifically, the evaluation model is trained based on a floating-point vector set and a target vector training algorithm, iterating through T and continuously adjusting its weight parameters w. [1] w [2] and bias parameter b [1] The calculation formulas for each parameter are as follows:

[0135]

[0136] Among them, w [1] w [2] To train the weight parameters of different network layers in an unstructured model, network layers can be added as needed, thus increasing the weight parameter space. [1] S is the bias parameter. [1i] Let i be the evaluation score of the i-th text passed by the model, such as Figure 9 As shown, Figure 9 S is a first sample score [1] A schematic diagram of the calculation method.

[0137] Specifically, the evaluation model is trained using a training algorithm based on the feature dimension vector and the target vector, and the calculation formulas for the parameters are as follows:

[0138]

[0139] F·w [3] +b [2] =S [2]

[0140] Among them, w [3] To train the weight parameters of the structured model, b [2] S is the bias parameter. [2] The second sample score

[0141] The final score is S = k1S [1] +k2S [2] That is, output = k1DL_model + k2ML_model, and k1 + k2 = 1.

[0142] As can be seen from the technical solutions provided in the embodiments of this specification above, the embodiments of this specification obtain the basic elements, scientific research capability elements, and carrying capacity elements of the sample institutions as sample source data; the sample source data is labeled with the sample evaluation score tags of the sample institutions; based on the basic elements and scientific research capability elements of the sample institutions, sample numerical data and sample character data are extracted as sample structured datasets; sample long text is extracted from the scientific research capability elements and the carrying capacity elements of the sample institutions as sample unstructured datasets; thus, the sample source data is divided into two major categories of data; the sample unstructured dataset is converted into sample floating-point number vectors, and the sample structured dataset is converted into sample feature vectors; the sample floating-point number vectors and the sample evaluation score tags are... An unstructured data evaluation model is input to predict scores, resulting in a first sample score. The sample feature vector and the sample evaluation score label are then input into a structured data evaluation model to predict scores, resulting in a second sample score. This yields the evaluation scores of the sample institutions predicted by the two types of data. The combined weight set of the model is then traversed, and a target weight ratio is selected from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label. An institution evaluation model is constructed based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio. The institution evaluation model constructed using this invention has higher accuracy, enabling rapid and accurate prediction of the evaluation scores of the institutions to be evaluated.

[0143] This specification also provides an apparatus for constructing an institutional evaluation model, such as... Figure 10 As shown, the device includes:

[0144] The sample data acquisition module 1010 is used to acquire the basic elements of the sample institution, the scientific research capability elements, and the carrying capacity elements of the sample institution as the sample source data; the sample source data is labeled with the sample evaluation score tag of the sample institution.

[0145] The sample structured data determination module 1020 is used to extract sample numerical data and sample character data based on the sample institution basic elements and sample scientific research capability elements, as a sample structured dataset.

[0146] The sample unstructured data determination module 1030 is used to extract the sample long text from the sample scientific research capability element and the sample institutional carrying capacity element as the sample unstructured dataset.

[0147] The sample vector conversion module 1040 is used to convert the unstructured sample dataset into a sample floating-point vector and the structured sample dataset into a sample feature vector.

[0148] The sample score prediction module 1050 is used to input the sample floating-point vector and the sample evaluation score label into an unstructured data evaluation model to predict the score and obtain a first sample score; and to input the sample feature vector and the sample evaluation score label into a structured data evaluation model to predict the score and obtain a second sample score.

[0149] The target weight ratio determination module 1060 is used to traverse the combined weight set of the model and select the target weight ratio from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label.

[0150] The evaluation model construction module 1070 is used to construct an institutional evaluation model based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio.

[0151] In some embodiments, the target weight ratio determination module includes:

[0152] The evaluation index value determination unit is used to traverse the combined weight set of the model and determine the combined evaluation index value of the model corresponding to each weight ratio based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label.

[0153] The target weight ratio determination unit is used to determine the optimal weight ratio of the combined evaluation index values ​​that meet preset conditions as the target weight ratio.

[0154] In some embodiments, the evaluation index value determination unit includes:

[0155] The weight ratio adjustment subunit is used to input the combined weight set of the model into the weight calculation layer, traverse the combined weight set, and adjust the weight ratio between the unstructured data evaluation model and the structured data evaluation model.

[0156] The sample score determination subunit is used to determine the sample prediction evaluation score of the sample organization under each weight ratio based on the first sample score, the second sample score, and each weight ratio.

[0157] The evaluation index value determination sub-unit is used to determine the combined evaluation index value of the model corresponding to each weight ratio based on the difference between the sample predicted evaluation score and the sample evaluation score label under each weight ratio.

[0158] In some embodiments, the sample vector transformation module includes:

[0159] The first data processing unit is used to input the sample unstructured dataset into the unstructured data processing layer for data cleaning to obtain the first sample standard dataset.

[0160] The sample floating-point vector extraction unit is used to input the first sample standard dataset into the first vector representation layer for text semantic extraction to obtain the sample floating-point vector.

[0161] The second data processing unit is used to input the sample structured dataset into the structured data processing layer for data cleaning to obtain the second sample standard dataset.

[0162] The sample feature vector determination unit is used to input the second sample standard dataset into the second vector representation layer for data correlation analysis to obtain the sample feature vector.

[0163] In some embodiments, the apparatus further includes:

[0164] The first score prediction module is used to input the sample floating-point vector into the unstructured data evaluation network to predict the evaluation score and obtain the first predicted score.

[0165] The first loss determination module is used to determine the first loss data based on the difference between the first predicted score and the sample evaluation score label;

[0166] The first training module is used to adjust the parameters of the unstructured data processing layer, the first vector representation layer, and the unstructured data evaluation network according to the first loss data until the training termination condition is met, and to determine the unstructured data evaluation network at the end of training as the unstructured data evaluation model.

[0167] In some embodiments, the training method for the structured data evaluation model includes:

[0168] The second score prediction module is used to input the sample feature vector into the structured data evaluation network to predict the evaluation score and obtain the second predicted score.

[0169] The second loss determination module is used to determine the second loss data based on the difference between the second predicted score and the sample evaluation score label;

[0170] The second training module is used to adjust the parameters of the structured data processing layer, the second vector representation layer, and the structured data evaluation network according to the second loss data until the training termination condition is met, and to determine the structured data evaluation network at the end of training as the structured data evaluation model.

[0171] In some embodiments, the sample structured data determination module includes:

[0172] The first sample data extraction unit is used to extract the number of projects, the number of tasks for each project, the budget for each project, and the number of people at each level in the organization's basic elements to obtain the first sample structured data.

[0173] The second sample data extraction unit is used to extract the number of scientific research achievements in the scientific research capability elements of the sample to obtain the second sample structured data.

[0174] The sample structured dataset construction unit is used to construct the sample structured dataset based on the first sample structured data and the second sample structured data.

[0175] In some embodiments, the apparatus further includes:

[0176] The module for acquiring source data to be evaluated is used to acquire the basic elements, scientific research capabilities, and institutional carrying capacity of the institution to be evaluated, as source data to be evaluated.

[0177] The unstructured dataset determination module is used to extract structured datasets based on the institution's basic elements and research capability elements; and to extract long texts from the research capability elements and the institution's carrying capacity elements as unstructured datasets.

[0178] The feature vector conversion module is used to convert the unstructured dataset into a floating-point vector and the structured dataset into a feature vector.

[0179] The first evaluation score prediction module is used to input the floating-point number vector into the unstructured data evaluation model to predict the score and obtain the first evaluation score.

[0180] The second evaluation score prediction module is used to input the feature vector into the structured data evaluation model to predict the score and obtain the second evaluation score.

[0181] The evaluation score determination module is used to obtain the evaluation score of the organization to be evaluated based on the first evaluation score, the second evaluation score, and the target weight ratio.

[0182] In some embodiments, the evaluation score determination module includes:

[0183] The weight determination unit is used to determine the first weight of the unstructured data evaluation model and the second weight of the structured data evaluation model according to the target weight ratio.

[0184] The product calculation unit is used to calculate the first product of the first evaluation score and the first weight, and the second product of the second evaluation score and the second weight;

[0185] The evaluation score calculation unit is used to calculate the sum of the first product and the second product to obtain the evaluation score of the organization to be evaluated.

[0186] The apparatus and method embodiments described herein are based on the same inventive concept.

[0187] In this embodiment of the invention, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0188] This specification provides an electronic device including a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the method for constructing an institutional evaluation model as provided in the above method embodiments.

[0189] Embodiments of the present invention also provide a computer storage medium, which can be disposed in a terminal to store at least one instruction or at least one program related to the construction method of an institutional evaluation model in the method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the construction method of the institutional evaluation model provided in the above method embodiment.

[0190] Embodiments of the present invention also provide a computer program product or computer program comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method for constructing the mechanism evaluation model provided in the above-described method embodiments.

[0191] Optionally, in the embodiments of this specification, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0192] The memory described in the embodiments of this specification can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for the functions, etc.; the data storage area may store data created according to the use of the device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.

[0193] The embodiments of the mechanism evaluation model construction method provided in this specification can be executed on a mobile terminal, computer terminal, server, or similar computing device. Taking running on a server as an example, Figure 11 This is a hardware structure block diagram of a server for a method of constructing an institutional evaluation model provided in the embodiments of this specification. For example... Figure 11As shown, the server 1100 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1110 (CPUs 1110 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 1130 for storing data, and one or more storage media 1120 (e.g., one or more mass storage devices) for storing application programs 1123 or data 1122. The memory 1130 and storage media 1120 may be temporary or persistent storage. The program stored in the storage media 1120 may include one or more modules, each module including a series of instruction operations on the server. Furthermore, the CPU 1110 may be configured to communicate with the storage media 1120 and execute the series of instruction operations stored in the storage media 1120 on the server 1100. Server 1100 may also include one or more power supplies 1160, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1140, and / or one or more operating systems 1121, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0194] The input / output interface 1140 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 1100. In one example, the input / output interface 1140 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 1140 may be a radio frequency (RF) module for wireless communication with the Internet.

[0195] Those skilled in the art will understand that Figure 11 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 1100 may also include... Figure 11 The more or fewer components shown, or having the same Figure 11 The different configurations shown.

[0196] As can be seen from the embodiments of the construction method, apparatus, equipment, or storage medium of the institutional evaluation model provided by the present invention, the present invention obtains the basic elements, scientific research capability elements, and carrying capacity elements of the sample institutions as sample source data; the sample source data is labeled with the sample evaluation score tags of the sample institutions; based on the basic elements and scientific research capability elements of the sample institutions, sample numerical data and sample character data are extracted as sample structured datasets; sample long text is extracted from the scientific research capability elements and the carrying capacity elements of the sample institutions as sample unstructured datasets; thus, the sample source data is divided into two major categories of data; the sample unstructured dataset is converted into sample floating-point number vectors, and the sample structured dataset is converted into sample feature vectors; the sample floating-point number vectors and the sample... The evaluation score label is input into the unstructured data evaluation model for score prediction to obtain the first sample score; then the sample feature vector and the sample evaluation score label are input into the structured data evaluation model for score prediction to obtain the second sample score; thus, the evaluation scores of the sample institutions predicted by the two types of data are obtained respectively; then, the combined weight set of the model is traversed, and a target weight ratio is selected from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label; based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio, an institution evaluation model is constructed; the institution evaluation model constructed using the method of this invention has higher accuracy, thus enabling rapid and accurate prediction of the evaluation score of the institution to be evaluated.

[0197] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0198] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0199] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer storage medium, such as a read-only memory, a disk, or an optical disk.

[0200] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing an institutional evaluation model, characterized in that, The method includes: The basic elements, scientific research capabilities, and carrying capacity of the sample institutions are obtained as the sample source data; the sample source data is labeled with the sample evaluation score tags of the sample institutions. Based on the basic elements of the sample institutions and the elements of the sample scientific research capabilities, numerical data and character data of the samples are extracted as a sample structured dataset. Extract the sample long text from the sample scientific research capability element and the sample institutional carrying capacity element to form the sample unstructured dataset; The unstructured dataset is converted into a floating-point vector, and the structured dataset is converted into a feature vector. The sample floating-point vector and the sample evaluation score label are input into the unstructured data evaluation model for score prediction to obtain the first sample score; and the sample feature vector and the sample evaluation score label are input into the structured data evaluation model for score prediction to obtain the second sample score. Traverse the combined weight set of the model, and select the target weight ratio from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label; An institutional evaluation model is constructed based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio.

2. The method according to claim 1, characterized in that, The combined weight set of the traversal model, based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label, filters out the target weight ratio from multiple weight ratios, including: Traverse the combined weight set of the model, and determine the combined evaluation index value of the model corresponding to each weight ratio based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label; The optimal weight ratio that satisfies the preset conditions for the combined evaluation index values ​​is determined as the target weight ratio.

3. The method according to claim 2, characterized in that, The process of traversing the combined weight set of the model, and determining the combined evaluation index value of the model corresponding to each weight ratio based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label, includes: The combined weight set of the model is input into the weight calculation layer, the combined weight set is traversed, and the weight ratio between the unstructured data evaluation model and the structured data evaluation model is adjusted. Based on the first sample score, the second sample score, and each weight ratio, the sample organization's sample prediction evaluation score under each weight ratio is determined; Based on the difference between the sample predicted evaluation score and the sample evaluation score label under each weight ratio, the combined evaluation index value of the model corresponding to each weight ratio is determined.

4. The method according to claim 1, characterized in that, The step of converting the unstructured dataset into a floating-point vector and the structured dataset into a feature vector includes: The unstructured dataset of the sample is input into the unstructured data processing layer for data cleaning to obtain the first standard dataset of the sample. The first sample standard dataset is input into the first vector representation layer for text semantic extraction to obtain the sample floating-point vector; The sample structured dataset is input into the structured data processing layer for data cleaning to obtain the second sample standard dataset. The second sample standard dataset is input into the second vector representation layer for data correlation analysis to obtain the sample feature vector.

5. The method according to claim 4, characterized in that, The training methods for the unstructured data evaluation model include: The sample floating-point vector is input into an unstructured data evaluation network to predict the evaluation score, and a first predicted score is obtained. The first loss data is determined based on the difference between the first predicted score and the sample evaluation score label; The parameters of the unstructured data processing layer, the first vector representation layer, and the unstructured data evaluation network are adjusted based on the first loss data until the training termination condition is met, and the unstructured data evaluation network at the end of training is determined as the unstructured data evaluation model.

6. The method according to claim 4, characterized in that, The training method for the structured data evaluation model includes: The sample feature vector is input into a structured data evaluation network to predict the evaluation score, thus obtaining a second predicted score. The second loss data is determined based on the difference between the second predicted score and the sample evaluation score label; The parameters of the structured data processing layer, the second vector representation layer, and the structured data evaluation network are adjusted according to the second loss data until the training termination condition is met, and the structured data evaluation network at the end of training is determined as the structured data evaluation model.

7. The method according to claim 1, characterized in that, The step of extracting numerical and character data from the sample institutions and their research capabilities to form a structured dataset includes: Extract the number of projects, the number of tasks per project, the budget for each project, and the number of people at each level in the organization's basic elements from the sample organization to obtain the first sample of structured data; The number of research achievements is extracted from the research capability elements of the sample to obtain the structured data of the second sample. The sample structured dataset is constructed based on the first sample structured data and the second sample structured data.

8. The method according to claim 1, characterized in that, After constructing the institutional evaluation model based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio, the method includes: Obtain the basic elements, research capabilities, and carrying capacity of the institution to be evaluated as the source data for evaluation. Based on the institutional basic elements and scientific research capability elements, a structured dataset is extracted; and long texts are extracted from the scientific research capability elements and institutional carrying capacity elements as unstructured datasets. The unstructured dataset is converted into a floating-point vector, and the structured dataset is converted into a feature vector; The floating-point vector is input into the unstructured data evaluation model to predict the score, and a first evaluation score is obtained. The feature vector is input into the structured data evaluation model to predict the score, thereby obtaining a second evaluation score; The evaluation score of the organization to be evaluated is obtained based on the first evaluation score, the second evaluation score, and the target weight ratio.

9. The method according to claim 8, characterized in that, The step of obtaining the evaluation score of the organization to be evaluated based on the first evaluation score, the second evaluation score, and the target weight ratio includes: The first weight of the unstructured data evaluation model and the second weight of the structured data evaluation model are determined based on the target weight ratio. Calculate the first product of the first evaluation score and the first weight, and the second product of the second evaluation score and the second weight; The sum of the first product and the second product is calculated to obtain the evaluation score of the organization to be evaluated.

10. An apparatus for constructing an institutional evaluation model, characterized in that, The device includes: The sample data acquisition module is used to acquire the basic elements, scientific research capabilities, and carrying capacity of the sample institutions as the sample source data; the sample source data is labeled with the sample evaluation score tags of the sample institutions. The sample structured data determination module is used to extract sample numerical data and sample character data based on the sample institution's basic elements and the sample's scientific research capability elements, as a sample structured dataset. The unstructured data determination module is used to extract long text from the scientific research capability elements and institutional carrying capacity elements of the samples, as the unstructured dataset of the samples. The sample vector conversion module is used to convert the unstructured sample dataset into a sample floating-point vector and the structured sample dataset into a sample feature vector. The sample score prediction module is used to input the sample floating-point vector and the sample evaluation score label into the unstructured data evaluation model to predict the score and obtain the first sample score; and to input the sample feature vector and the sample evaluation score label into the structured data evaluation model to predict the score and obtain the second sample score. The target weight ratio determination module is used to traverse the combined weight set of the model and select the target weight ratio from multiple weight ratios based on each weight ratio in the combined weight set, the first sample score, the second sample score, and the sample evaluation score label. The evaluation model construction module is used to construct an institutional evaluation model based on the unstructured data evaluation model, the structured data evaluation model, and the target weight ratio.

Citation Information

Patent Citations

  • Scientific research institution integration evaluation method and system thereof

    CN105447633A

  • Determination method and device for implementation project, electronic equipment and storage medium

    CN112200463A