Scoring method for scientific innovation grade
By automatically identifying and mapping the data source table structure, dynamically generating the table structure and processing abnormal data, the problems of multi-source data fusion and dynamic indicator weighting in scientific research management are solved, a unified quantitative scoring of scientific innovation levels is achieved, and the objectivity and comparability of the evaluation are improved.
Patent Information
- Application Number
- CN202510753428.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies in scientific research management lack the ability to automatically access multi-source data, dynamically weight indicators, and perform standardized processing, resulting in highly subjective evaluation results, insufficient horizontal comparability, and insufficient vertical traceability.
By automatically identifying the table structure and field types of relational business data sources, performing metadata mapping and dynamic table structure generation, and combining abnormal data processing and weight configuration, a comprehensive scoring of scientific innovation levels can be achieved.
It realizes unified quantitative processing of data sources, improves the objectivity and accuracy of evaluation, supports dynamic adjustment and rapid response, and enhances the horizontal fairness of evaluation results and the sensitivity of vertical trend analysis.
Smart Images

Figure CN120634798A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scientific research information management and data processing, and in particular to a scoring method for scientific innovation levels. Background Art
[0002] In the field of scientific research management, the rating evaluation of entity objects (including university laboratories, research institutes and innovative enterprises) has long relied on a multi-level indicator system and highly subjective expert reviews; traditional methods mostly use offline questionnaires, static financial indicators or single dimensions such as the number of achievements for comprehensive scoring. The indicator caliber varies due to differences between institutions, and different dimensions are difficult to directly compare, resulting in the inability to accurately quantify some key indicators, large subjective fluctuations in evaluation results, and insufficient horizontal comparability and vertical traceability.
[0003] With the popularization of big data and artificial intelligence technologies in scientific research statistics, how to build a set of innovation level quantification methods that are oriented to multi-source business data, can be iterated in real time and have unified measurement standards has become an issue that needs to be urgently addressed by science and technology management departments and various scientific research entities.
[0004] CN101676932A discloses an electronic supervision method for government investment projects based on a shared database, which can conduct process monitoring and real-time level assessment of each link of the project. However, its core focus is on investment progress and management compliance, and it lacks an in-depth quantitative mechanism for scientific research and innovation output and capacity growth.
[0005] CN118863670A discloses a project review business management system and method, which improves post-review efficiency through modular processes and expert online review. However, its indicator system is still mainly based on manual data filling, lacking the ability to automatically extract, map and standardize original operational data distributed in different business systems and heterogeneous databases, and is unable to achieve dynamic weight adjustment and normalized scoring for cross-year, multi-dimensional innovation indicators.
[0006] In summary, the existing technology still has limitations in aspects such as automatic retrieval of data sources, consistency of field mapping, elimination of abnormal data, and normalization and weighted model construction of multidimensional indicators. Therefore, there is an urgent need for a method that can automatically access relational business data sources, perform multi-field mapping and dynamic table structure generation, combine annual statistical parameters to achieve standardization and normalization processing, and output comprehensive grade scores through configurable weights. The present invention solves the problems of inconsistent quantitative scales and strong subjectivity of traditional evaluation methods through a closed loop from metadata capture to comprehensive score output, and fills the gaps in existing technologies in multi-source data fusion and dynamic indicator weighting. Summary of the Invention
[0007] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract of the specification and the title of the invention of this application to avoid blurring the purpose of this section, the abstract of the specification and the title of the invention, and such simplifications or omissions cannot be used to limit the scope of the invention.
[0008] In view of the above existing problems, the present invention is proposed.
[0009] To solve the above technical problems, the present invention provides the following technical solutions: Step 1: Establish a connection with a relational business data source, automatically retrieve and identify its table structure, field type and association relationship, and obtain metadata;
[0010] Step 2: Map the business fields in the metadata to the fields required for the innovation level evaluation indicators in a one-to-one correspondence according to the preset mapping rules to generate a data set to be processed;
[0011] Step 3: Based on the field mapping result, a dynamic table structure generation mechanism is used to construct a pre-processing table adapted to the data set to be processed;
[0012] Step 4: writing the data set to be processed into the pre-processing table in batches according to a preset time interval, and performing field integrity verification and abnormal data processing during the writing process;
[0013] Step 5: Calculate the mean and standard deviation of the indicator data for each first-level innovation dimension according to the statistical year, and standardize each indicator value to a standard value of mean 0 and standard deviation 1;
[0014] Step 6: Sum all standard values under the same first-level innovation dimension to obtain the total standard value of the first-level innovation dimension, and perform normalization processing on it to obtain a normalized score in the range of 0 to 1;
[0015] Step 7: weighting the normalized score according to a preset weight to obtain a weighted score;
[0016] Step 8. Add up the weighted scores of all first-level innovation dimensions and output the comprehensive scientific innovation grade score of the target entity object.
[0017] As a preferred embodiment of the scoring method for scientific innovation level of the present invention, in step 1, obtaining metadata includes:
[0018] Automatically identify the data source type and call the corresponding communication driver;
[0019] Call the system data directory interface to read the data dictionary and form a field name-field type correspondence table;
[0020] Parse the identified primary key-foreign key relationships and generate a logical dependency graph between fields;
[0021] The field name-field type correspondence table and the logical dependency graph between fields are summarized into a machine-readable metadata list and stored in a memory buffer.
[0022] As a preferred solution of the scoring method for scientific innovation level of the present invention, according to the preset mapping rules, the business fields in the metadata are mapped one-to-one with the fields required for the innovation level evaluation indicators, including:
[0023] Read the mapping rule library and load the rule set corresponding to the current business scenario;
[0024] Automatically match metadata fields based on field name similarity, data type compatibility, and business meaning comparison;
[0025] When the automatic matching confidence is lower than a first threshold, triggering a manual correction interface for manual adjustment;
[0026] The final matching relationship is written into the mapping result table.
[0027] As a preferred solution of the scoring method for scientific innovation level described in the present invention, a dynamic table structure generation mechanism is used to construct a pre-processing table adapted to the dataset to be processed, including:
[0028] Automatically generate column definitions and index strategies based on field mapping results;
[0029] Determine whether a table with the same name already exists in the target database. If so, synchronize column additions, deletions, and modifications and retain historical versions.
[0030] Update the column order according to the field sorting rules;
[0031] The updated table structure information is recorded in the structure version control library, which is the pre-processing table.
[0032] As a preferred solution of the scoring method for scientific innovation level of the present invention, the batch writing of data in step 4 specifically includes:
[0033] Perform non-null constraint check on the data to be written. If the field value is null, it is marked as an exception.
[0034] Perform data type consistency check on numeric fields. If the type does not match, mark it as an exception.
[0035] Extreme outliers are identified through statistical distribution thresholds, and if they deviate from the annual mean by more than three standard deviations, they are marked as abnormal;
[0036] Write the exception record into the exception data queue, record the batch number and error cause, and exclude subsequent writes;
[0037] Write the validated data in batches to the pre-processing table and return the write result log.
[0038] As a preferred embodiment of the scoring method for scientific innovation level of the present invention, the step 5 specifically includes:
[0039] The indicator data are grouped using the statistical year and the first-level innovation dimension as double grouping keys;
[0040] Missing values are filled using the mean strategy of the same dimension in the same year;
[0041] A boundary truncation strategy is used for the detected extreme values, replacing them with the 95% quantile value within the group;
[0042] After completing missing value processing and extreme value truncation, the mean and standard deviation were calculated for each group;
[0043] Based on the calculation results, each indicator data in the group was converted into a standardized value with 0 as the mean and 1 as the standard deviation.
[0044] As a preferred embodiment of the scoring method for scientific innovation level of the present invention, in step 6, obtaining the normalized score includes:
[0045] Retrieve the total standard value set of all first-level innovation dimensions in the same statistical year;
[0046] Determine the minimum and maximum values of the set;
[0047] According to the linear scaling rule, each total standard value is mapped to a unified interval of 0-1 to obtain a normalized score;
[0048] The normalized scores are stored in a normalized result table, and the entity objects are associated with the statistical years.
[0049] As a preferred embodiment of the scoring method for scientific innovation level of the present invention, the normalized score is weighted according to a preset weight to obtain a weighted score, including:
[0050] Load the weight set corresponding to the current statistical year from the weight configuration table;
[0051] Check whether the sum of the loaded weights is equal to one. If not, normalize them proportionally;
[0052] The normalized score of each first-level innovation dimension is multiplied by the corresponding weight to generate a weighted score.
[0053] As a preferred embodiment of the scoring method for scientific innovation level of the present invention, in step eight, outputting the comprehensive scientific innovation level score of the target entity object includes:
[0054] The weighted scores of all first-level innovation dimensions of the same entity object in the same statistical year are summed to obtain the comprehensive grade score;
[0055] According to the position of the comprehensive grade score in the interval [0,1], the entity object is divided into four grade labels: excellent, good, qualified, and needs improvement;
[0056] Within the same industry dimension, the comprehensive ratings of all entity objects are sorted in descending order to generate ranking results;
[0057] The comprehensive grade score, grade label and ranking results are written into the scoring database, and a visual analysis report is generated.
[0058] Beneficial effects of the present invention:
[0059] 1. Complete field semantic unification and storage structure optimization before data enters the analysis phase, avoiding repeated cleaning and format migration due to inconsistent calibers in the later stage, reducing manual sorting costs, shortening data preparation cycles, and significantly improving scalability for large-scale horizontal comparisons.
[0060] 2. Utilize threshold and distribution detection algorithms to isolate null values, type conflicts, and outlier records in real time, ensuring that subsequent statistical calculations are based only on highly reliable samples, reducing the risk of error propagation to a negligible level and improving model robustness.
[0061] 3. Completely eliminate the dimensional differences and scale effects between indicators, allowing different entities to be directly compared within the same evaluation coordinate system, significantly improving evaluation fairness and enhancing the sensitivity of longitudinal trend analysis;
[0062] 4. The introduction of a configurable weighting and weighted result accumulation mechanism enables the dynamic adaptation of the evaluation model to industry characteristics or policy orientations, allowing decision makers to adjust dimension weights in real time based on strategic priorities and instantly obtain new comprehensive scores, supporting differentiated resource allocation, precise level incentives, and rapid early warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:
[0064] Figure 1This is a schematic diagram of the overall process of a scoring method for scientific innovation level according to the present invention;
[0065] Figure 2 A schematic diagram of the process flow for constructing a pre-processing table shown in the present invention;
[0066] Figure 3 This is a schematic diagram of the comprehensive grade scoring process shown in the present invention. DETAILED DESCRIPTION
[0067] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.
[0068] Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without making any creative work should fall within the scope of protection of the present invention.
[0069] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0070] According to an embodiment of the present invention, Figure 1 The flowchart shown is a scoring method for scientific innovation level, which specifically includes the following steps:
[0071] Step 1: Establish a connection to the relational business data source, automatically retrieve and identify its table structure, field types and relationships, and obtain metadata;
[0072] Step 2: According to the preset mapping rules, the business fields in the metadata are mapped one-to-one with the fields required for the innovation level evaluation indicators to generate a data set to be processed;
[0073] Step 3: Based on the field mapping results, use the dynamic table structure generation mechanism to build a pre-processing table that is compatible with the dataset to be processed;
[0074] Step 4: Write the dataset to be processed into the pre-processing table in batches according to the preset time interval, and perform field integrity verification and abnormal data processing during the writing process;
[0075] Step 5: Calculate the mean and standard deviation of the indicator data for each first-level innovation dimension according to the statistical year, and standardize each indicator value to a standard value of mean 0 and standard deviation 1;
[0076] Step 6: Sum all standard values under the same first-level innovation dimension to obtain the total standard value of the first-level innovation dimension, and perform normalization processing on it to obtain a normalized score in the range of 0 to 1;
[0077] Step 7: Weight the normalized score according to the preset weight to obtain a weighted score;
[0078] Step 8. Add up the weighted scores of all first-level innovation dimensions and output the comprehensive scientific innovation grade score of the target entity object.
[0079] It should be noted that the embodiment of the present invention adopts a dynamic table structure generation mechanism to automatically generate an adaptive data table according to the actual mapping fields, thereby solving the problem of insufficient business flexibility caused by the traditional static table structure design; the traditional fixed table structure requires manual adjustment when adding or deleting indicator fields, which is a time-consuming process and prone to structural inconsistencies; the present invention automatically adjusts the data storage structure by real-time perception of changes in field mapping rules, so that data from different sources and different structures can be quickly integrated, shortening the data processing cycle and improving the system's ability to respond quickly to changes in innovation evaluation indicators.
[0080] Furthermore, when writing data in batches, this embodiment introduces a field integrity check and abnormal data processing process. Unlike the simple batch import in the existing technology, step four pre-defines non-empty rules, type consistency rules, business logic boundary rules and statistical outlier detection rules to filter abnormal data in real time, avoid the accumulation of abnormal data and interference with subsequent evaluation results, thereby reducing the cost of later error correction and manual verification, and improving the accuracy and stability of the evaluation model calculation results.
[0081] Preferably, the present invention combines the standardization of mean and standard deviation, summation within the indicator dimension, and minimum-maximum normalization processing operations to solve the scale distortion problem caused by the simple weighted accumulation of the original data in the prior art. Since the magnitude of innovation data of various scientific research entities varies greatly, simple aggregation will cause the contribution of indicators to be masked or over-amplified. The standardization and normalization steps of the present invention enable cross-dimensional and cross-year indicators to be compared with unified standards, effectively improving the horizontal fairness and comparability of innovation level evaluation results between different scientific research entities, eliminating the errors caused by inconsistent dimensions, and greatly enhancing the objectivity and accuracy of the evaluation model.
[0082] The following combination Figures 2 and 3 The flowcharts shown and some preferred or optional examples of the present invention more specifically describe the implementation process and / or effects of certain examples of the present invention.
[0083] Generate a dataset to be processed
[0084] Parse the JDBC / ODBC connection string provided in the access configuration file;
[0085] Automatically determine the data source type based on the connection string prefix or the server identifier returned by the handshake. If the identifier contains keywords such as "MySQL, PostgreSQL, Oracle, SQLServer, MariaDB", the corresponding communication driver is loaded respectively;
[0086] Establish a session-level connection and record the successful connection event in the connection log table;
[0087] Call the target data source system table (such as INFORMATION_SCHEMA.COLUMNS, ALL_TAB_COLUMNS);
[0088] Read the field name, field type, null value flag, and length information line by line, and generate a "field name-field type" mapping list;
[0089] Write the list into the memory column cache for subsequent fast indexing;
[0090] Query the constraint views in the target data source system tables (such as INFORMATION_SCHEMA.KEY_COLUMN_USAGE or ALL_CONSTRAINTS);
[0091] For each primary key or foreign key constraint, extract the table name, constraint name, constraint column sequence, and reference table information;
[0092] Construct the vertex set and edge set of the field dependency directed graph, form a logical dependency graph structure and save it in memory;
[0093] Using fields as granularity, associate the "name-type" entry with its node identifier in the logical dependency graph;
[0094] Generate a sort key based on the field's sequence number and constraint type in the table, which is used for priority judgment in subsequent mapping algorithms;
[0095] Construct a metadata object collection according to the following formula and cache it in the metadata buffer:
[0096]
[0097] Among them, f i Indicates the field name, T i Indicates the field data type, K i Indicates the key attribute (primary key, foreign key or ordinary column), r i The reference field set to which this field is connected.
[0098] It should be further explained that, in the embodiments of the present invention, the relational business data source refers to a production-level database instance that is organized using a relational model (Relational Model) and interacts in structured query language (SQL), including enterprise-level commercial databases (such as comprehensive scientific research management of large universities, special financial fund management), open source community databases (such as provincial project review systems), cloud-native relational databases (such as cross-campus scientific research cloud platforms, results transformation services), embedded relational databases (such as stand-alone scientific research archive collection terminals) and data warehouse-type relational databases (provincial scientific research data marts, annual statistics of innovation indexes).
[0099] In an optional implementation manner, the method for identifying the data source type is specifically:
[0100] Read the JDBC / ODBC URL provided by the user;
[0101] Use regular expressions to match protocol prefixes (such as jdbc:oracle:thin:@), port number constants (such as 1521, 3306, 5432), and database name fragments;
[0102] Establish a socket-level handshake and capture the server banner;
[0103] Attempt to query product-unique system views (such as Oracle's DUAL and SQL Server's sysobjects) using a minimally privileged account.
[0104] If the returned structure matches the expected one, the type is confirmed, otherwise it falls back to the next candidate;
[0105] After the type identification is completed, the corresponding Driver Class is retrieved from the driver registry, reflectively loaded and the connection pool is initialized.
[0106] For example, the field name-field type correspondence table is as follows:
[0107] Table 1. Field name-field type table
[0108]
[0109] Exemplarily, generating a logical dependency graph between fields specifically includes:
[0110] Vertex: database column, uniquely identified as<table_name> .<column_name> ;
[0111] Directed edges: primary key to foreign key PK→FK (reference), foreign key to reference column FK→PK (referenced);
[0112] Cascading dependency: If the foreign key column is referenced by other tables, a multi-level directed path is formed;
[0113] edge_type: referential, unique_constraint, check_constraint;
[0114] update_rule / delete_rule: CASCADE, SET NULL, RESTRICT;
[0115] Weight: Automatically assigned based on reference frequency or query statistics, used for subsequent mapping priority sorting;
[0116] Traverse the system constraint view and extract all PK / FK tuples;
[0117] For each constraint, add a new vertex and insert a directed edge into the graph;
[0118] Perform strong connectivity detection based on the Tarjan algorithm to ensure that the dependency graph has no self-loop deadlocks, perform hierarchical topological sorting on the dependency graph, and output the table dependency hierarchy. This hierarchy information is used for parallelism planning during subsequent batch loading;
[0119] The dependency graph is serialized in JSON Graph structure. The following example shows it:
[0120]
[0121] It should be noted that this graph guides automatic join path inference in the mapping phase, reduces manual specification, determines the partition key and index strategy when the dynamic table is generated, and improves subsequent query performance.
[0122] Furthermore, according to the preset mapping rules, the business fields in the metadata are mapped one-to-one with the fields required for the innovation level evaluation indicators to generate a data set to be processed, including the following steps:
[0123] Parse the business scenario identifier in the request message (such as enterprise technology project application);
[0124] Using the identifier as the primary key, the corresponding rule set is called from the mapping rule library. Rule set = (target indicator field ID, synonymous keyword set, allowed data type, semantic category label, weight parameter);
[0125] Write the rule set into the memory rule cache for subsequent fast matching;
[0126] Furthermore, the following three-dimensional scoring is performed on each metadata field f:
[0127] Name similarity S1: using the improved cosine n-gram + Jaro–Winkler combination algorithm;
[0128] Data type compatibility S2: If the metadata field type belongs to the allowed type in the rule set, it is recorded as 1, otherwise it is recorded as 0;
[0129] Semantic category consistency S3: Use pre-trained Word2Vec to generate word vectors and calculate the cosine similarity between the field description and the target indicator label;
[0130] The comprehensive confidence score is generated according to the formula Score = w1S1 + w2S2 + w3S3, where w1, w2, and w3 are the weights of name similarity, data type compatibility, and semantic category consistency, respectively;
[0131] If Score ≥ the first threshold (e.g., 0.80), the match is automatically confirmed;
[0132] If the score is less than the first threshold and greater than or equal to the second threshold (e.g., 0.60), the patient enters the manual correction queue;
[0133] If Score is less than the second threshold, the match is rejected directly;
[0134] Write the determined correspondence into the mapping result table. The core columns include: business field, target indicator field, data type, matching source (automatic / manual), and timestamp.
[0135] When multiple indicators conflict for the same business field, one record with the highest confidence level is retained, and the remaining records are marked as pending conflict status;
[0136] Trigger the mapping result event to notify the dynamic table creation module to generate the pre-processing table structure;
[0137] For the confirmed matching field pairs (B i ,I j ), extract the original value V from the business source table i ;
[0138] According to the field mapping, a column → value key-value pair sequence is assembled in memory to form a data set to be processed. Its mathematical expression formula is:
[0139] D={(I j ,V i )|M(B i )=I j ,i=1,…,n}
[0140] Among them, M represents the mapping function, I j is the innovation level indicator field, V iTo correspond to the business field value, D is written into the temporary cache table for the batch write module to call.
[0141] It should be noted that, through the above-mentioned process operation, the embodiment of the present invention completes the field semantic unification and storage structure optimization before the data enters the analysis stage, avoiding repeated cleaning and format migration due to inconsistent caliber in the later stage, reducing manual sorting costs, shortening the data preparation cycle, and significantly improving the scalability during large-scale horizontal comparison.
[0142]
Build preprocessing table
[0143] Based on the above field mapping results, a dynamic table structure generation mechanism is used to build a preprocessing table that is adapted to the dataset to be processed. The following steps are included:
[0144] Read the field mapping result set {business field B i , indicator field I j , data type T i , key attribute K i};
[0145] For each mapping record, construct column definition D i =<Column Name=I j ,Type=T i , whether empty = N i , default value = Def i >
[0146] Generate indexing strategies based on key attribute rules:
[0147] If K i =PK, then it is a composite primary key candidate;
[0148] If K i =FK, a B-tree foreign key index is generated;
[0149] If the column cardinality is ≥ 80% of the number of table rows, then a bitmap index is added;
[0150] D i Write it into the temporary DDL cache queue together with the corresponding index strategy;
[0151] Call the metadata database interface to query whether there is a table P with the same name in the target schema;
[0152] If it does not exist, create a new table and proceed to the next step;
[0153] If it exists, enter the column comparison process:
[0154] Compare the column list of the existing table P with D i , generate Δ = new column set ∪ deleted column set ∪ modified column set;
[0155] For the newly added columns in Δ, execute ALTER TABLE ADD COLUMN;
[0156] For the deleted columns in Δ, execute ALTER TABLE DROP COLUMN and archive the column definition;
[0157] To change the data type of the column in Δ, execute ALTER TABLE MODIFY COLUMN;
[0158] Write the complete column change script and execution results into the structure migration log table;
[0159] Arrange according to the column sorting rules, where:
[0160] The primary key column is placed at the top;
[0161] Columns that are not empty and have the top 50% access frequency are sorted in descending order of frequency;
[0162] The remaining columns are grouped by business module label and arranged in lexicographical order;
[0163] Call the database online reordering function to update the physical order of columns in lock-free mode;
[0164] Record column order change mapping table to maintain the consistency of field order between upper-level ORM and report script, and generate structure version number V k =SHA-256 (such as DDL script + timestamp);
[0165] Write the table name, version number, DDL script summary, executor, execution time, and column metadata snapshot into the structure version control repository;
[0166] If it is a new table, mark status = INITIAL; if it is a change, mark status =
[0167] MIGRATION and record the previous version number;
[0168] Trigger the pre-processing table ready event to notify the batch write module to enable the latest structure.
[0169] It should be noted that through the above processing steps, the method of the present invention can quickly construct a preprocessing table based on the real-time field mapping results without human intervention, automatically generate column definitions and indexes, ensure the consistency of storage structure and analysis requirements, and optimize the column order while maintaining logical consistency while improving I / O continuity and query efficiency.
[0170] Furthermore, the data set to be processed is written into the pre-processing table in batches according to a preset time interval, and field integrity verification and abnormal data processing are performed during the writing process; specifically, the following steps are included:
[0171] Preset time interval Δt=15 minutes;
[0172] The scheduler triggers the write task according to Δt, pulling the record set formed in the pending buffer area in the previous cycle;
[0173] Bucket the record set according to the data source system and statistical year to generate several batch buckets B k (k=1…m), the number of records in each bucket does not exceed 50,000 to ensure that the volume of a single transaction log is controllable;
[0174] Barrel B k For each record r(i), execute field by field:
[0175] Not null constraint: If the preprocessed table metadata declares nullable = false and r(i).f i If it is empty, set the flag ERR_NULL;
[0176] Data type consistency: Call the database's native IS OF TYPE function. If r(i).f i If it is incompatible with the column data type, the flag ERR_TYPE is set;
[0177] Length / precision truncation check: If the string length or numeric precision overflows the column definition, the flag ERR_LEN is set;
[0178] Perform outlier identification on a set of numeric fields that pass integrity checks:
[0179] Preload the mean μ and standard deviation σ of the same field in the corresponding statistical year;
[0180] If |r(i).f i -μ|>3σ, then set the flag ERR_OUTLIER;
[0181] For time series fields, the upper limit of the box plot Q3 + 1.5IQR is used as an alternative threshold, and the intersection of the two is taken to reduce misjudgment;
[0182] The records whose collection mark is not empty constitute the abnormal batch E k ;
[0183] E k Write to the exception data queue, with additional fields batch_id, error_code, and error_msg;
[0184] If the same source field has a cumulative ERR_TYPE of more than 5% in three consecutive batches, the field will be added to the field blacklist;
[0185] For the normal record set G after filtering out the exceptions k =B k \E k , using client streaming batch writing, where:
[0186] Start a single batch transaction and use the INSERT...ON CONFLICT DO UPDATE statement to commit the data in batches.
[0187] Set the batch size to 1,000 rows / pack and enable JDBC batch parameterized statements to reduce network round trips.
[0188] If the write is not completed within 5 seconds, it will be automatically submitted in batches recursively to prevent table locks, and a write receipt R will be generated for each batch of transactions. k =<batch_id,total_cnt,success_cnt,error_cnt,duration_ms,commit_ts> ;
[0189] R k Write the result log table and push it to the monitoring platform via Webhook;
[0190] If error_cnt / total_cnt ≥ 10%, automatically adjust the next cycle Δt to 5 minutes.
[0191] It should be noted that the method of the present invention completes the high-speed database filling of the data to be processed within 15 minutes through the implementation of the above-mentioned process operations, and uses multiple verification and queue isolation methods to ensure data quality, reduce lock contention and log expansion caused by frequent small transactions, quickly locate high-risk fields and block the spread of errors, and through streaming transaction batch writing, take into account both throughput and rollback safety, ensuring the continuous consistency of the pre-processing table structure and content.
[0192]
Comprehensive Rating of Computational Science Innovation
[0193] Read the table T of indicators to be processed after exception filtering, with the core columns: entity_id, year, dim_lv1, metric_code, and metric_value;
[0194] The composite grouping key G is composed of year (statistical year) and dim_lv1 (first-level innovation dimension).<year,dim_lv1> ;
[0195] Execute GROUP BY year,dim_lv1 to generate group segment T(G);
[0196] Perform missing detection on the metric_value in each group T(G);
[0197] If the missing rate is ≤ 20%, the mean of the currently available samples in the group is μ G fill;
[0198] If the missing rate is >20%, an integrity alarm is triggered and the standardization of this group is temporarily suspended until data is supplemented;
[0199] Calculate the 95% quantile Q for metric_value in T(G) 95 ;
[0200] If record r satisfies r.metric_value>Q 95 , then perform boundary truncation: r.metric_value←Q 95 , and record the truncated entries in the trim_log table;
[0201] After missing value filling and extreme value truncation are completed, T(G) is re-aggregated;
[0202] Calculate the within-group mean μ G =Σmetric_value / N;
[0203] Calculate within-group standard deviation
[0204] (μ G ,σ G ) is written to the stat_param table with the key G;
[0205] For each record r in T(G), execute z=(r.metric_value-μ G ) / σ G ;
[0206] Store z in the new column std_value with four decimal places;
[0207] If σ G =0 (the indicator has no variance), then z=0 and the zero-var flag is recorded;
[0208] Write the result set containing entity_id, year, dim_lv1, metric_code, and std_value into the annual standardized table;
[0209] Generate a version number V = hash (such as year + processing timestamp) for this process and record it in proc_history;
[0210] Trigger the standardization completion event for subsequent normalization and weighting processing.
[0211] As an example, for the first-level innovation dimension R&D investment in the 2024 statistical year, there are five entities with original funding data (in 10,000 yuan): {450, --, 780, 9250, 500}, where -- indicates missing;
[0212] Then the average value (excluding missing values) = (450 + 780 + 9250 + 500) / 4 = 2 745 → fill in the missing values with 2745;
[0213] Extreme value cutoff: 95% quantile Q 95 ≈8 582→9 250 is truncated to 8 582;
[0214] The new data set {450, 2 745, 780, 8 582, 500} is obtained, with a mean μ = 2 611.4 and a standard deviation σ ≈ 3385.2;
[0215] Standardization: 450 → (450-2 611.4) / 3 385.2 ≈ -0.64; 2 745 → 0.04; 780 → -0.54; 8 582 → 1.76; 500 → -0.55.
[0216] Furthermore, all standard values under the same first-level innovation dimension are summed to obtain the total standard value of the first-level innovation dimension, and normalized to obtain a normalized score in the range of 0 to 1; including:
[0217] For all first-level innovation dimensions of the same entity object E in a single statistical year Y, retrieve its total standard value set S Y ={S1,S2,…,S p};
[0218] Calculate the minimum value min(S Y ) and the maximum value max(S Y ), perform a linear mapping on any dimension j:
[0219]
[0220] Will <E,Y,j,N j >Write to the normalized result table and stamp the version;
[0221] Among them, S j is the total standard value of the first-level dimension j obtained through statistics, N j is the normalized score;
[0222] Furthermore, the normalized scores are weighted according to the preset weights to obtain weighted scores, and then the weighted scores of all the first-level innovation dimensions are accumulated to output the comprehensive scientific innovation grade score of the target entity object, which specifically includes:
[0223] Read the weight set w corresponding to year Y from the weight configuration table Y ={w1,w2,…,w p};
[0224] like Normalize by proportion
[0225] For each dimension j, calculate:
[0226] P j =w j ·N j
[0227] and record <E,Y,j,P j >To weighted score table;
[0228] Among them, p is the number of first-level innovation dimensions, w Y is the annual weight set, w j For policy focus, P j weighted scores for the dimension level;
[0229] In the same year Y, sum the weighted scores of all first-level dimensions of entity E:
[0230]
[0231] Map entity levels by rating interval:
[0232] Table 2. Level mapping table
[0233] interval Level Label <![CDATA[0.85≤P E ≤1]]> excellence <![CDATA[0.65≤P E <0.85]]> good <![CDATA[0.45≤P E <0.65]]> qualified <![CDATA[0≤P E <0.45]]> Needs improvement
[0234] Aggregate all entities within the same industry classification {P E}, press P E Generate a yearly ranking list in descending order and output <E,Y,P E ,Level,Rank> to the scoring database.
[0235] It should be noted that the embodiment of the present invention eliminates cross-dimensional dimensional differences through the above calculation method, maps the emphasis, and generates the dimension score P j , forming a single-valued index P E , and then assign grades and rankings, ultimately forming a comprehensive scientific innovation rating framework for decision-making, which effectively improves the horizontal fairness and comparability of innovation grade evaluation results between different scientific research entities, eliminates the errors caused by inconsistent dimensions, and greatly enhances the objectivity and accuracy of the evaluation model.
[0236] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A scoring method for scientific innovation level, characterized in that: include: Step 1: Establish a connection to the relational business data source, automatically retrieve and identify its table structure, field types and relationships, and obtain metadata; Step 2: Map the business fields in the metadata to the fields required for the innovation level evaluation indicators in a one-to-one correspondence according to the preset mapping rules to generate a data set to be processed; Step 3: Based on the field mapping result, a dynamic table structure generation mechanism is used to construct a pre-processing table adapted to the data set to be processed; Step 4: writing the data set to be processed into the pre-processing table in batches according to a preset time interval, and performing field integrity verification and abnormal data processing during the writing process; Step 5: Calculate the mean and standard deviation of the indicator data for each first-level innovation dimension according to the statistical year, and standardize each indicator value to a standard value of mean 0 and standard deviation 1; Step 6: Sum all standard values under the same first-level innovation dimension to obtain the total standard value of the first-level innovation dimension, and perform normalization processing on it to obtain a normalized score in the range of 0 to 1; Step 7: weighting the normalized score according to a preset weight to obtain a weighted score; Step 8. Add up the weighted scores of all first-level innovation dimensions and output the comprehensive scientific innovation grade score of the target entity object.
2. The scoring method for scientific innovation level according to claim 1, characterized in that: In the step 1, obtaining metadata includes: Automatically identify the data source type and call the corresponding communication driver; Call the system data directory interface to read the data dictionary and form a field name-field type correspondence table; Parse the identified primary key-foreign key relationships and generate a logical dependency graph between fields; The field name-field type correspondence table and the logical dependency graph between fields are summarized into a machine-readable metadata list and stored in a memory buffer.
3. The scoring method for scientific innovation level according to claim 1 or 2, characterized in that: According to the preset mapping rules, the business fields in the metadata are mapped one-to-one with the fields required for the innovation level evaluation indicators, including: Read the mapping rule library and load the rule set corresponding to the current business scenario; Automatically match metadata fields based on field name similarity, data type compatibility, and business meaning comparison; When the automatic matching confidence is lower than a first threshold, triggering a manual correction interface for manual adjustment; The final matching relationship is written into the mapping result table.
4. The scoring method for scientific innovation level according to claim 3, characterized in that: Use the dynamic table structure generation mechanism to build a pre-processing table that adapts to the dataset to be processed, including: Automatically generate column definitions and index strategies based on field mapping results; Determine whether a table with the same name already exists in the target database. If so, synchronize column additions, deletions, and modifications and retain historical versions. Update the column order according to the field sorting rules; The updated table structure information is recorded in the structure version control library, which is the pre-processing table.
5. The scoring method for scientific innovation level according to claim 4, characterized in that: The data batch writing in step 4 specifically includes: Perform non-null constraint check on the data to be written. If the field value is null, it is marked as an exception. Perform data type consistency check on numeric fields. If the type does not match, mark it as an exception. Extreme outliers are identified through statistical distribution thresholds, and if they deviate from the annual mean by more than three standard deviations, they are marked as abnormal; Write the exception record into the exception data queue, record the batch number and error cause, and exclude subsequent writes; Write the validated data in batches to the pre-processing table and return the write result log.
6. The scoring method for scientific innovation level according to claim 1, characterized in that: The step five specifically includes: The indicator data are grouped using the statistical year and the first-level innovation dimension as double grouping keys; Missing values are filled using the mean strategy of the same dimension in the same year; A boundary truncation strategy is used for the detected extreme values, replacing them with the 95% quantile value within the group; After completing missing value processing and extreme value truncation, the mean and standard deviation were calculated for each group; Based on the calculation results, each indicator data in the group was converted into a standardized value with 0 as the mean and 1 as the standard deviation.
7. The scoring method for scientific innovation level according to claim 1, characterized in that: In step 6, obtaining a normalized score includes: Retrieve the total standard value set of all first-level innovation dimensions in the same statistical year; Determine the minimum and maximum values of the set; According to the linear scaling rule, each total standard value is mapped to a unified interval of 0-1 to obtain a normalized score; The normalized scores are stored in a normalized result table, and the entity objects are associated with the statistical years.
8. The scoring method for scientific innovation level according to claim 1 or 7, characterized in that: The normalized score is weighted according to a preset weight to obtain a weighted score, including: Load the weight set corresponding to the current statistical year from the weight configuration table; Check whether the sum of the loaded weights is equal to one. If not, normalize them proportionally; The normalized score of each first-level innovation dimension is multiplied by the corresponding weight to generate a weighted score.
9. The scoring method for scientific innovation level according to claim 8, characterized in that: In step eight, outputting the scientific innovation comprehensive grade score of the target entity object includes: The weighted scores of all first-level innovation dimensions of the same entity object in the same statistical year are summed to obtain the comprehensive grade score; According to the position of the comprehensive grade score in the interval [0,1], the entity object is divided into four grade labels: excellent, good, qualified, and needs improvement; Within the same industry dimension, the comprehensive ratings of all entity objects are sorted in descending order to generate ranking results; The comprehensive grade score, grade label and ranking results are written into the scoring database, and a visual analysis report is generated.
Citation Information
Patent Citations
Electronic monitoring method and system of government-invested construction project
CN101676932A
Table data quality exploration method and device
CN111209538A
Scientific research institution rating system, method and equipment and storage medium
CN114169731A
Quantitative evaluation method for scientific and technological innovation
CN118411069A
Integrated metadata management method and system based on data source expandability
CN119576863A