Innovation ability assessment method based on large language model and multi-modal semantic understanding

By employing large language models and multimodal semantic understanding, this approach addresses the shortcomings of existing innovation capability assessment methods, such as insufficient modeling of inter-enterprise relationships, rigid assessment standards, and poor cross-industry comparability. It enables the scientific quantification and dynamic assessment of enterprise innovation capabilities, thereby enhancing the precision and reliability of the assessment.

CN122020264APending Publication Date: 2026-05-12JIANGSU PRODUCTIVITY PROMOTION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU PRODUCTIVITY PROMOTION CENT
Filing Date
2026-04-14
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing innovation capability assessment methods are unable to model the relationships and group positions among enterprises. The assessment standards are static and rigid, lack dynamic adaptability, have poor cross-industry comparability, and are difficult to achieve fair comparisons across fields and time periods.

Method used

Based on large language models and multimodal semantic understanding, this method collects heterogeneous data from multiple sources both inside and outside enterprises, generates a unified feature matrix using BERT and RoBERTa-large models, constructs a Transformer encoder, and designs a gating network using graph structure techniques to evaluate the innovation capabilities of technical topics. By using BERT and RoBERTa-large models to process classification and generate a unified feature matrix, adversarial training is used to extract domain-invariant features and dynamically generate domain-specific features to generate an enterprise innovation capability evaluation method.

Benefits of technology

It enables the scientific quantification of enterprise innovation capabilities and fair comparison across industries and time periods. It can integrate multi-source data, jointly characterize the relationships among enterprise groups, and has domain-adaptive capabilities and dynamic benchmark generation capabilities, thereby improving the precision and reliability of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020264A_ABST
    Figure CN122020264A_ABST
Patent Text Reader

Abstract

The invention discloses an innovation ability assessment method based on a large language model and multi-modal semantic understanding, and belongs to the technical field of information processing. The method comprises the steps of collecting multi-source heterogeneous data inside and outside an enterprise, and generating a unified feature matrix; realizing global feature interaction among enterprises through a multi-head self-attention mechanism, and outputting an enterprise feature matrix; designing a gating network to perform element-by-element dynamic modulation on the enterprise feature matrix; domain invariant features are extracted through adversarial training, domain specific features are dynamically generated through a super network, and final adaptive features are obtained through gating aggregation; and obtaining an enterprise innovation dynamic score. According to the invention, domain invariance constraint is carried out on enterprise features through a DANN structure, so that features in different technical fields are aligned in a bottleneck space; according to the method, HyperNetwork is combined with features to guide soft domain routing, and a domain dissimilatory feature transformation matrix is dynamically generated for each enterprise, so that the final features comprehensively reflect enterprise innovation input, output and potential value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information processing technology, specifically relating to a method for evaluating innovation capabilities based on large language models and multimodal semantic understanding. Background Technology

[0002] In the era of digital economy and innovation-driven development, corporate innovation capability has become an important benchmark for measuring industrial competitiveness and productivity. Government departments need to categorize and classify enterprises based on their innovation capabilities when formulating science and technology support policies, financial subsidies, and tax reduction standards. Financial institutions also need to quantitatively characterize enterprises' technological innovation level and continuous innovation capability when conducting science and technology financial lending, equity investment, and mergers and acquisitions. Enterprises themselves hope to identify weaknesses and optimize resource allocation through quantitative assessment results to achieve continuous improvement in innovation efficiency. Therefore, how to conduct scientific, objective, and dynamic quantitative assessment of corporate innovation capability has become a common need in policymaking, resource allocation, and corporate management.

[0003] Existing methods for assessing innovation capabilities primarily employ traditional multi-indicator evaluation techniques such as expert scoring, the Analytic Hierarchy Process (AHP), and fuzzy comprehensive evaluation. These methods typically rely on a limited number of structured indicators, such as R&D intensity, number of patents, and revenue growth rate, and then assign fixed weights and static scoring standards. These methods have the following limitations:

[0004] 1. Difficulty in modeling inter-firm relationships and group positions: Traditional methods often treat firms as independent evaluation objects, scoring them based solely on their static indicators, ignoring the network structures formed between firms in terms of technological similarity, patent citations, upstream and downstream collaboration in the industrial chain, and joint R&D. In practical applications, a firm's relative position within a technology group (such as leader, follower, or potential hidden champion) is crucial for judging innovation capabilities, but existing methods lack systematic modeling of such group interaction information;

[0005] 2. The evaluation criteria are static and rigid, lacking dynamic adaptability: Existing evaluation systems often use pre-set fixed scoring standards or threshold grading, making it difficult to reflect changes in the macroeconomic environment, industry cycles, technological iterations, and overall innovation levels in a timely manner. The actual meaning of the same score may drift significantly at different times, resulting in a lack of reliability in longitudinal comparisons across years and stages.

[0006] 3. Poor cross-industry comparability, with industry differences and innovation capabilities mixed together: Simply adopting a unified indicator system and linear standardization method can easily mistake "inherent industry attributes" for "innovation quality", resulting in a lack of comparability in the scores of companies in different industries, making it difficult to support cross-industry resource allocation and horizontal benchmarking between regions.

[0007] With the development of pre-trained large language models and the maturity of multimodal semantic understanding technology, large models are used to perform unified semantic encoding on large-scale heterogeneous data. Combined with graph structure modeling, adversarial learning, and dynamic scoring mechanisms, a new technical path for assessing enterprise innovation capabilities has been provided. However, existing publicly available technologies lack an integrated method that simultaneously considers multi-source data fusion, enterprise group interaction modeling, domain adaptation, and dynamic benchmark generation. This makes it difficult to achieve a detailed characterization of enterprise innovation capabilities and a fair, cross-domain quantitative comparison in complex real-world scenarios.

[0008] Therefore, there is an urgent need for a new method for evaluating innovation capabilities that can integrate multi-source heterogeneous data, jointly characterize the relationships among enterprise groups, and possess both domain-adaptive capabilities and dynamic benchmark generation capabilities, so as to achieve scientific quantification of enterprise innovation levels and fair comparisons across industries and time periods. Summary of the Invention

[0009] The purpose of this invention is to provide a method for evaluating innovation capabilities based on large language models and multimodal semantic understanding, so as to solve the problems mentioned in the background art.

[0010] The objective of this invention is achieved as follows: a method for evaluating innovation capability based on large language models and multimodal semantic understanding, comprising the following steps:

[0011] Step S1: Collect multi-source heterogeneous data from both inside and outside the enterprise and classify it. Then, use the BERT-base model and the RoBERTa-large model to process the classified multi-source heterogeneous data and generate a unified feature matrix.

[0012] Step S2: Construct a Transformer encoder with adjacency matrix attention constraints, realize global feature interaction between enterprises through multi-head self-attention mechanism, and output enterprise feature matrix;

[0013] Step S3: Design a gating network to dynamically modulate the enterprise feature matrix element by element, and then aggregate it with the unified feature matrix through residual connection and layer normalization to form a deep fusion feature;

[0014] Step S4: Extract domain-invariant features through adversarial training, dynamically generate domain-specific features through a hypernetwork, and obtain the final fitting features through gated aggregation;

[0015] Step S5: Obtain the dynamic score of enterprise innovation by suppressing outliers, dynamically generating a baseline value after hybrid correction, and performing piecewise nonlinear mapping.

[0016] Preferably, the process of collecting and classifying multi-source heterogeneous data from both inside and outside the enterprise, and then processing the classified multi-source heterogeneous data using both the BERT-base model and the RoBERTa-large model, specifically involves:

[0017] Step S1-1: Define the sample enterprise set, obtain multi-source heterogeneous data from the sample enterprise set, and classify the multi-source heterogeneous data, specifically as follows:

[0018] Define the sample enterprise set , The total number of companies participating in the assessment; each company The acquired data was divided into two categories: structured data and... and text data Structured data is a collection of structured information carried by numerical values ​​or categories, while text data is a collection of unstructured information carried by natural language.

[0019] Step S1-2: Process text data using the RoBERTa-large model:

[0020] Step S1-2-1: For each enterprise raw text data set The text data set is preprocessed by sequentially performing cleaning, deduplication, and quality filtering operations. ;in, For enterprises The number of text documents;

[0021] Step S1-2-2: Perform multi-document aggregation on the preprocessed text data set:

[0022] Sort by information priority: Chinese documents are arranged according to a predefined fixed priority;

[0023] Concatenation and Truncation: The sorted documents are concatenated sequentially, separated by RoberTa delimiters, and a start marker is added to the beginning of the sequence. <s>The total length is truncated at the token level to the maximum input length. tokens, expressed as:

[0024] ;in, It is an aggregated sequence;

[0025] Step S1-2-3: Aggregate the sequence Feature extraction using the RoBERTa-large model:

[0026] The RoBERTa-large model serves as a fixed feature extractor, and a parameter freezing strategy is employed to freeze the training parameters of the RoBERTa-large model to a fixed number. The RoBERTa-large model extracts features, specifically as follows:

[0027] Extract sequence start marker <s>The corresponding hidden state vector is used as the aggregate semantic representation of the entire text sequence to obtain the text feature line vector:

[0028] ;

[0029] in, Represents a sequence matrix;

[0030] All The text feature row vectors of each company are stacked row by row to form a text feature matrix, expressed as:

[0031] ;

[0032] in, ;

[0033] Steps S1-3: Process structured data using BERT-base;

[0034] Steps S1-4: The two features are fused into a unified feature matrix through a cross-modal attention weighted aggregation mechanism, and an enterprise association adjacency matrix is ​​constructed based on the fused enterprise-level feature matrix as auxiliary structural information.

[0035] Preferably, in steps S1-3, BERT-base is used to process structured data, specifically as follows:

[0036] Each company Structured field records , This represents the total number of structured fields. For the Mth key-value pair field value; perform unified preprocessing on the structured data, which is done through a predefined field configuration table. Drive; among which, For field type, For unit text, The number of significant digits to retain; The name of the Mth key-value pair field;

[0037] Step S1-3-2: Convert the preprocessed field records into a text input sequence acceptable to BERT-base according to a predefined template and fixed order:

[0038] Field arrangement order: Fields are arranged in a fixed order according to the indicator hierarchy of innovation capability assessment;

[0039] Serialization template: Each field is converted into a text fragment in the format "field name: field value", with fragments connected by [SEP], and a [CLS] marker is added to the beginning of the sequence. The expression is:

[0040] ;

[0041] Step S1-3-3: Perform dimensional projection using the BERT-base model:

[0042] The BERT-base model is used as a word segmenter. It employs a parameter freezing mode to freeze all training parameters and extracts the final hidden state at the [CLS] position as the aggregate representation.

[0043] ;

[0044] in, Dimensional alignment is achieved through learnable linear projection layers:

[0045] ;in, The weight matrix of the linear projection layer. The offset vector of the linear projection layer;

[0046] All The structured features of each company are stacked row by row into a structured feature matrix:

[0047] ;

[0048] In steps S1-4, the text feature matrix and the structured feature matrix are fused into a unified enterprise-level feature matrix through a cross-modal attention weighted aggregation mechanism, specifically as follows:

[0049] The text feature matrix and the structured feature matrix are fed into a shared-parameter attention scoring network, and first projected onto... Hidden space, through After activation, it is compressed into a scalar fraction;

[0050] The original scores are converted into fusion weights that sum to 1 by Softmax normalization.

[0051] Attention weights are used to perform a weighted summation of the text feature matrix and the structured feature matrix, and the fusion ratio of the two modalities is adaptively adjusted according to the specific situation of each enterprise.

[0052] The fusion features are further refined by using a projection layer with non-linear activation, enabling it to capture the non-linear interaction patterns between the two features. The weighted features before fusion and the projected features after fusion are added element by element to alleviate information decay during the fusion process and ensure that the original fusion information can be retained through the direct path even if the learning effect of the projection layer is not good.

[0053] All The integrated characteristics of individual enterprises are stacked row by row to form a unified enterprise-level characteristic matrix:

[0054] ;in, Indicates enterprise After cross-modal attention-weighted aggregation, it integrates the combined features of textual semantic information and structured quantitative information;

[0055] In the unified feature matrix Based on this, an adjacency matrix reflecting the relationship structure between enterprises is constructed using cosine similarity calculation. ;in, .

[0056] Preferably, in step S3, a gating network is designed to dynamically modulate the enterprise feature matrix element by element, and then aggregated with the unified feature matrix through residual connections and layer normalization to form a deep fusion feature, specifically:

[0057] Step S3-1: Generate gating weights:

[0058] The input to the gating network is formed by concatenating the enterprise feature matrix and the unified feature matrix along the feature dimensions.

[0059] ;in, For the enterprise feature matrix, To unify the feature matrix; The total number of companies participating in the assessment;

[0060] The unified feature matrix carries the original multi-source fusion information, forming a complementary perspective with the enterprise feature matrix, enabling the gating network to comprehensively encode the two information levels before and after to judge the importance of each dimension;

[0061] The concatenated features are mapped to element-wise gated weights using a fully connected layer and a sigmoid activation function.

[0062] ;

[0063] in, The weight matrix is ​​a learnable weight matrix; A learnable bias vector; The sigmoid function compresses each output element to... ; Indicates enterprise The The relevance of 3D deep encoding features to the evaluation task A value close to 1 indicates that the dimension carries a key signal and should be retained, while a value close to 0 indicates that the dimension is noise or redundancy and should be suppressed.

[0064] Step S3-2: Element-wise modulation is performed on the enterprise feature matrix using gating weights, and then aggregated into a deep fusion feature matrix through residual connections and layer normalization, specifically as follows:

[0065] Using gating weights Enterprise feature matrix Each element performs on / off control independently:

[0066] ;in, For element-wise multiplication, This is the enterprise feature matrix after gated refinement;

[0067] Gated modulated features With the unified feature matrix Element-wise summation, followed by layer normalization, results in a deep fusion feature matrix:

[0068] ;

[0069] in, For residual connections; For layer normalization operation; This is the feature matrix for deep fusion.

[0070] Preferably, in step S4, domain-invariant features are extracted through adversarial training, domain-specific features are dynamically generated through a hypernetwork, and the final fitting features are obtained through gated aggregation. Specifically:

[0071] Step S4-1: Construct a DANN model to obtain domain-invariant features that generalize across domains:

[0072] The DANN model includes a feature encoder. Gradient Reversal Layer (GRL) and Neighborhood Discriminator Using feature encoders Extracting common feature representations shared across domains:

[0073] Feature encoder For a single-layer fully connected network, 3D deep features compressed to Bottleneck space: ;

[0074] in, The weight matrix of the feature encoder. This is the bias vector of the feature encoder; For enterprises Encoding representation in the bottleneck space;

[0075] The behavior of the gradient inversion layer (GRL) is asymmetric in forward and backward propagation, with the forward propagation being an identity mapping. If the gradient direction is reversed, then the gradient direction is reversed and... Controlling the intensity of reversal: ;in, For gradient tensors, The gradient reversal coefficient;

[0076] Gradient reversal coefficient The sigmoid scheduling function proposed by DANN is adopted:

[0077] ;

[0078] in, This represents the percentage of training progress. For the current iteration step, This represents the total number of iterations.

[0079] Domain Discriminator It is a two-layer fully connected network that receives encoded features passed through the gradient inversion layer (GRL). And predict the domain origin of each sample, the first layer will 3D encoded features are projected onto nonlinear transformations. Dimensional discriminative subspace: ;

[0080] in, This is the weight matrix of the first layer of the discriminator. This is the bias vector for the first layer of the discriminator;

[0081] The second layer linearly projects the discriminant subspace features to... Each domain category is normalized to a probability distribution using softmax:

[0082] ;

[0083] in, This is the weight matrix for the second layer of the discriminator. This is the bias vector for the second layer of the discriminator; This represents the total number of domain categories. For enterprises Predicted probability distributions belonging to various fields, ; The hidden layer activation feature matrix of the discriminator;

[0084] Domain-Discrimination Cross-Entropy Loss: Cross-entropy loss measures the deviation between the discriminator's predicted domain probability distribution and the true domain labels. The expression is:

[0085] ;

[0086] in, The one-hot encoding matrix for real-world domain labels; It means if and only if the enterprise Belongs to the One field;

[0087] Define the encoder output as a neighborhood-invariant feature:

[0088] ;

[0089] Step S4-2: Based on HyperNetwork and feature-guided soft domain routing (FSDR), dynamically generate domain-specific features;

[0090] Step S4-3: Through a gated network and dual-path projection, the domain-invariant features and domain-specific features are adaptively aggregated into the final fitting features.

[0091] Preferably, in step S4-2, the domain-specific features are dynamically generated based on HyperNetwork and feature-guided soft domain routing (FSDR).

[0092] Step S4-2-1: Calculate feature projection and soft route weights using Feature-Guided Soft Domain Routing (FSDR):

[0093] Will The domain-invariant feature of dimension is projected to A domain embedding space of dimension maps enterprise features to the same semantic space as the domain embedding to perform similarity calculations:

[0094] ;

[0095] in, The learnable route projection matrix; For enterprises The domain-invariant eigenvectors; For enterprises Routing query vectors in the domain embedding space;

[0096] Using route query vectors For querying, embed tables with learnable domains Each row is a key, and the enterprise's soft routing weights for each domain are calculated using scaled dot product attention:

[0097] ;

[0098] in, For the first Learnable embedding vectors for each domain; This is the scaling factor for the software router weight; Indicates enterprise Belonging to the domain The soft weights, softmax guarantees and ;

[0099] By weighting and mixing the domain embedding table with soft route weights, personalized domain embeddings are synthesized for each enterprise:

[0100] ;

[0101] in, For enterprises Personalized domain embedding; when When it degenerates into a one-hot vector, ;

[0102] To prevent route weights from deviating completely from known neighborhood labels, a regularization loss is introduced to provide a weak anchoring constraint for soft routes:

[0103] ;

[0104] in, Indicates loss or punishment;

[0105] Step S4-2-2: Embedding Personalization Domain The parameter generation module of the input supernetwork dynamically synthesizes the enterprise-specific feature transformation matrix, generating a low-rank transformation matrix:

[0106] Will The complete transformation matrix is ​​decomposed into left factors With right factor The product;

[0107] The generation structure of the two factors is symmetrical and can be expressed in a unified form: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] ,set up This represents the total number of elements corresponding to the factor. Let be the target shape corresponding to the factor, where Let be the rank of the low-rank decomposition. Domain-specific feature dimensions:

[0108] ; ;

[0109] in, The factor generation matrix of the hypernetwork; For hypernetworks targeting enterprises A flat parameter vector embedded in the personalized domain;

[0110] Combining two low-rank factors into a complete transformation matrix: ;

[0111] in, For enterprises The left factor of the specific transformation matrix captures from dimensional invariant space to Compressed mapping of intermediate space in dimensionality; For the right factor, capture from dimensional invariant space to Compressed mapping of intermediate space in dimensionality; For enterprises A dedicated complete transformation matrix;

[0112] Step S4-2-3: Project the domain-invariant features onto the domain-specific space using a dynamic transformation matrix to extract domain-specific features:

[0113] The domain-invariant features of each sample are re-expressed in a domain-specific form according to their domain-specific rules:

[0114] ;

[0115] in, For enterprises The domain-invariant eigenvectors; For enterprises Domain-specific feature vectors; Encoded the enterprise A cross-domain universal model The rules for transforming the characteristics of the enterprise's domain are encoded. Multiplying the two together will re-express the general pattern as a domain-specific form according to the domain rules.

[0116] For all Each sample undergoes the above projection and is stacked along the batch dimension:

[0117] ;in, Indicates domain-specific characteristics.

[0118] Preferably, in step S4-3, the domain-invariant features and domain-specific features are adaptively aggregated into the final fitting features through a gated network and dual-path projection, specifically as follows:

[0119] Step S4-3-1: Design the aggregation weights of the adaptive gating network:

[0120] A scalar gating weight is generated for each sample using a gating network. This quantifies the relative dependence of the sample on domain-invariant and domain-specific features:

[0121] ;

[0122] in, This represents the domain-invariant feature after gated modulation. Broadcast along the feature dimension to Later and Element-by-element multiplication This is element-wise multiplication; For the gating general feature projection matrix, This is the feature projection matrix for dual-path projection stitching. This is the dual-path projection bias vector; To ultimately adapt the feature dimensions;

[0123] Step S4-3-2: Dual-path projection and feature aggregation: First, concatenate the domain-invariant features and domain-specific features along the feature dimension: ;

[0124] Then, the final adaptation features are generated through dual-path projection:

[0125] ;

[0126] in, This represents the domain-invariant feature after gated modulation. Broadcast along the feature dimension to Later and Element-by-element multiplication This is element-wise multiplication; For the gating general feature projection matrix, This is the feature projection matrix for dual-path projection stitching. This is the dual-path projection bias vector; To ultimately adapt the feature dimensions.

[0127] Preferably, in step S5, a dynamic score for enterprise innovation is obtained by suppressing outliers, dynamically generating a baseline value corrected for promiscuity, and performing piecewise nonlinear mapping. Specifically:

[0128] Step S5-1: Using the scoring prediction head, compress the high-dimensional fitting features into a predicted value of the enterprise's continuous innovation capability:

[0129] The scoring prediction head uses a two-layer fully connected network. The first layer will... 3D adaptation features are projected onto the nonlinear transformation. Dimensional subspace: ;

[0130] in, The hidden layer weight matrix is... This is the hidden layer bias vector; For final adaptation features;

[0131] The second layer compresses the hidden features into unconstrained one-dimensional real values ​​through linear projection, serving as a continuous prediction of each company's innovation capability by the model: ;

[0132] in, To output the weight vector, To output the bias scalar; For the model of enterprises Continuous predicted values ​​of innovation capability; the score prediction vector is:

[0133] ;in, It is the transpose symbol;

[0134] Step S5-2: Eliminate dimensional differences and suppress the perturbation of population statistics by extreme values ​​through robust standardization based on quartiles and adaptive decay of outliers;

[0135] Step S5-3: Generate a dynamic benchmark value with confounding correction based on the current population distribution, and convert it into the final score through piecewise nonlinear mapping.

[0136] Preferably, in step S5-2, robust standardization based on quartiles and adaptive decay of outliers eliminate dimensional differences and suppress the disturbance of extreme values ​​to population statistics, specifically manifested as follows:

[0137] Step S5-2-1: Calculate the population quantile statistic:

[0138] Group median: ;

[0139] Interquartile range: ;

[0140] in, Used as a location reference instead of the mean. As a scale reference, it can be used as an alternative to standard deviation.

[0141] Step S5-2-2: Calculate the adaptive penalty intensity coefficient for outliers:

[0142] The degree of anomaly of each sample is quantified based on how much it exceeds the Tukey boxplot criterion, providing an adaptive penalty for subsequent decay:

[0143] ;

[0144] in, It is the numerical stability constant. For indicator functions, the penalty intensity coefficient Determined by the degree to which the sample deviates from the median:

[0145] Step S5-2-3: Robust normalization and adaptive decay of outliers, the expression is:

[0146] ;

[0147] When not outliers Furthermore, when the indicator function is set to 0, the normalization degenerates into the classic robust form; for outliers... And the indicator function is 1, the molecule is Compression increases with the extreme deviation from the target.

[0148] For all The robust standardized scoring vector is formed by the following companies performing the above operations:

[0149] ;

[0150] Each element Reflecting on enterprises The relative position of the predicted innovation capacity value in the current population has eliminated the difference in dimensions and suppressed the amplitude disturbance of outliers.

[0151] Preferably, in step S5-3, a dynamic benchmark value with confounding correction is generated based on the current population distribution, and then transformed into a final score through piecewise nonlinear mapping, specifically as follows:

[0152] Step S5-3-1: Construct a feature matrix of confounding factors and remove systematic biases using OLS regression:

[0153] Construct a feature vector of confounding factors for each enterprise. ,in, Unique hot coding in the industry, To evaluate the yearly unique heat encoding, all samples are stacked into a confounding feature matrix. ;

[0154] The regression coefficients of each confounding factor were estimated by fitting a linear model using OLS:

[0155] ;

[0156] in, This represents the baseline offset of the score when all confounding factors are zero. For the first Regression coefficients of each confounding factor; Number of year categories; The first characteristic matrix of confounding factors Line number Column elements; The impact of industry classification and assessment year on enterprises was comprehensively quantified. The systematic impact of standardized scoring;

[0157] Remove bias from standardized scoring: ;

[0158] in, The corrected standardized score;

[0159] Step S5-3-2: Calculate the relative reference position index:

[0160] Relative benchmark indicators include the median deviation indicator and the head-to-head distance indicator. The median deviation indicator measures a company's... The relative degree of deviation from the group median is expressed as: ;

[0161] Head distance index measures enterprises The relative distance from the head threshold is expressed as: ;

[0162] in, Indicates the measurement of enterprises The relative degree of deviation from the group median; This indicates the relative distance from the head threshold; Indicates the median level of the current group. This indicates the threshold for the leading level of the current group;

[0163] Step S5-3-3: Calculate the domain semantic calibration term and synthesize the comprehensive dynamic benchmark value:

[0164] Embedding personalization domain With routing query vector The degree of similarity between the two is measured by cosine similarity:

[0165] ;

[0166] A high value indicates that the company's characteristics are highly consistent with its industry profile, and the benchmark is adjusted appropriately; a low value indicates that the company deviates from its industry profile, and the calibration item reduces the impact of this deviation on the benchmark.

[0167] The three dimensions of indicators are linearly combined with fixed weights to synthesize a comprehensive dynamic benchmark value:

[0168] ;

[0169] in, , , ; Comprehensive reflection of enterprises After adjusting for confounding factors, the performance relative to the current group median and head level was calibrated for domain similarity.

[0170] Step S5-3-4: Segmentation using a nonlinear mapping function:

[0171] when At this point, for the lower-score segment, a superlinear mapping is used, with slower penalty for lagging companies, thus enhancing the differentiation of bottom-tier companies. The expression is: ;

[0172] when For the middle segment, using a strict linear mapping, the scores of the main enterprises are evenly distributed in the range of 40–60 points to maintain the fairness of the assessment. The expression is: ;

[0173] when At that time, a sublinear mapping is used to suppress the head saturation effect, and the expression is: ;

[0174] Final evaluation set: , ;

[0175] As a benchmark, To upgrade, It is a level of attention; the score is generated based on the current data, and the benchmark is dynamically adjusted according to the composition and performance of the assessed group, which overcomes the lag of static standards and can sensitively reflect the latest trend of productivity development.

[0176] Compared with the prior art, the present invention has the following improvements and advantages:

[0177] 1. The invention uses a DANN structure to impose domain invariant constraints on enterprise features, aligning features from different technology fields within the bottleneck space; at the same time, it utilizes HyperNetwork combined with feature-guided soft domain routing (FSDR) to dynamically generate domain-specific feature transformation matrices for each enterprise, characterizing the different mapping rules for innovation performance in different fields without adding a large number of parameters.

[0178] 2. By adopting a cross-modal attention weighted fusion mechanism to adaptively allocate the weights of text and structured information in a unified vector space, it can automatically identify the focus of technology-intensive enterprises and finance-driven enterprises, so that the final features comprehensively reflect the enterprise's innovation input, output and potential value.

[0179] 3. This invention introduces robust standardization based on the median and interquartile range and adaptive outlier decay under the Tukey criterion, effectively suppressing the distortion of the overall distribution by extreme values. It uses OLS regression to explicitly model the systematic influence of industry category and year on the score, and removes this bias from the standardized score to obtain a confounding-corrected relative position index. The score generated by this invention can adaptively adjust with the overall level of the current assessment group, ensuring fair comparison between different companies at the same point in time and improving the stability and interpretability of cross-year longitudinal comparisons. Attached Figure Description

[0180] Figure 1 This is a flowchart illustrating the method of the present invention.

[0181] Figure 2 This is a schematic diagram of the Transformer encoder structure.

[0182] Figure 3 This is a schematic diagram of the DANN model structure.

[0183] Figure 4 This is a schematic diagram showing the MSE convergence curves of each scheme on the validation set.

[0184] Figure 5 This is a schematic diagram showing the domain-specific MSE results of each scheme in different fields.

[0185] Figure 6 This is a bar chart comparing two indicators in the ablation experiment. Detailed Implementation

[0186] The invention will be further summarized below with reference to the accompanying drawings.

[0187] like Figure 1 As shown, step S1: Collect multi-source heterogeneous data from inside and outside the enterprise and classify it. Use BERT-base model and RoBERTa-large model to process the classified multi-source heterogeneous data and generate a unified feature matrix.

[0188] Step S1-1: Define the sample enterprise set, obtain multi-source heterogeneous data from the sample enterprise set, and classify the multi-source heterogeneous data:

[0189] Define the sample enterprise set , The total number of companies participating in the assessment; for each company Raw data is collected from the following multi-source heterogeneous data source systems: patent databases provide patent abstracts, full text of claims, IPC classification numbers, and number of applications / grants;

[0190] Corporate financial statements provide structured indicators such as R&D investment amount, revenue growth rate, R&D personnel ratio, and debt-to-equity ratio;

[0191] The policy document repository provides policy summaries, descriptions of supported areas, and funding amounts; the science and technology public opinion platform provides media reports and industry analyst reviews; corporate annual reports provide management discussion and analysis sections and strategic planning descriptions; and industry-academia-research cooperation data provides the number of collaborating institutions, the number of joint patents, and project funding.

[0192] Each enterprise The acquired data was divided into two categories: structured data and... and text data ,

[0193] Text data is a collection of unstructured information in natural language, including patent abstracts and claims, MD&A sections of annual reports, narrative paragraphs of policy documents, and texts of public opinion reports;

[0194] Structured data is a collection of structured information carried by numerical values ​​or categories, including R&D investment amount, number of patents granted, revenue growth rate, employee education distribution, and number of industry-academia-research cooperation projects; each company's structured data is organized as follows: Records with 1 field ,in, For field names, This is the value of the Mth key-value pair field;

[0195] Step S1-2: Process text data using the RoBERTa-large model:

[0196] Each enterprise raw text data set The cleaning, deweighting, and mass filtration operations are performed sequentially, as follows:

[0197] Cleaning: Remove HTML tags, special control characters, and redundant whitespace; remove formatted legal jargon from patent texts, retaining only substantive technical descriptions;

[0198] Deduplication: At the enterprise level, similar duplicate documents are detected and removed based on SimHash text fingerprints to avoid feature bias caused by information redundancy;

[0199] Quality filtering: Removes fewer characters than the threshold Short documents, with low information density, are prone to introducing noise.

[0200] Obtain the preprocessed text data set ;in, For enterprises The number of text documents.

[0201] Step S1-2-2: Perform multi-document aggregation on the preprocessed text data set, and perform the following operations:

[0202] Each company typically has multiple text documents, which need to be aggregated into a single input sequence and sorted according to information priority: Chinese documents are arranged according to a predefined fixed priority: invention patent abstracts > annual report MD&A > policy document descriptions > public opinion report texts, with high-priority documents placed first to ensure that the most critical information is retained when truncation is performed;

[0203] Concatenation and Truncation: Concatenates the sorted documents sequentially, using Robertia delimiters between them.< / s> Separate, add a start marker to the beginning of the sequence. <s>The total length is truncated at the token level to the maximum input length. tokens, expressed as:

[0204] ;in, It is an aggregated sequence.

[0205] Step S1-2-3: Aggregate the sequence Feature extraction is performed using the RoBERTa-large model, and the following steps are taken:

[0206] The RoBERTa-large model is used as a fixed feature extractor, and the parameter freezing strategy is adopted to freeze the training parameters of the RoBERTa-large model to a fixed number. There are three reasons for parameter freezing: (1) For the enterprise data scale in this scenario, full fine-tuning of 355M parameters has a serious risk of overfitting; (2) The scale pre-training has given the model sufficient general semantic encoding capabilities; (3) Task-level feature adaptation is undertaken by subsequent operations.

[0207] Feature extraction from the RoBERTa-large model: Extracting sequence start markers <s>The corresponding hidden state vector is used as the aggregate semantic representation of the entire text sequence, and the resulting text feature row vector is: ;

[0208] in, Represents a sequence matrix. Sequence matrix The final hidden layer output through the forward propagation of the RoBERTa-large model;

[0209] All The text feature row vectors of each company are stacked row by row to form a text feature matrix, expressed as:

[0210] ;in, Represents the text feature matrix;

[0211] Step S1-3: Process structured data using BERT-base:

[0212] Step S1-3-1: Preprocess the structured fields:

[0213] Each company Structured field records , This represents the total number of structured fields; the structured data undergoes unified preprocessing, which is performed through a predefined field configuration table. Drive; among which, For field type, For unit text, The number of significant digits to retain; The name of the Mth key-value pair field;

[0214] Numeric field formatting: Convert raw numerical values ​​into natural language expressions that the language model can understand according to field type rules. Add order units to absolute quantities, convert ratios to percentages, add quantifiers to counts, and use name text directly for categories.

[0215] Missing value handling: For fields with missing values, replace them with placeholder text "Not disclosed" instead of deleting the field to maintain a consistent field sequence structure across all companies and avoid representation discrepancies caused by different numbers of fields;

[0216] Step S1-3-2: Convert the preprocessed field records into a text input sequence acceptable to BERT-base according to a predefined template and fixed order:

[0217] Field arrangement order: Fields are arranged in a fixed order according to the indicator hierarchy of innovation capability assessment: core financial indicators → R&D investment indicators → intellectual property indicators → human resource indicators → collaborative innovation indicators; the fixed arrangement ensures that fields with the same semantic category are in the same position in the sequence, which is conducive to positional encoding to capture the structural patterns between fields;

[0218] Serialization template: Each field is converted into a text fragment in the format "field name: field value", with fragments connected by [SEP], and a [CLS] marker is added to the beginning of the sequence. The expression is:

[0219] ;

[0220] Step S1-3-3: Perform dimensional projection using the BERT-base model:

[0221] Receive the serialized text output in step S1-3-2 The goal is to encode the serialized structured data into an aggregate representation, and then use a learnable linear projection layer to scale its dimensions from the output dimensions of BERT-based datasets. Align to RoBERTa-large output dimension This eliminates the dimensionality difference between the two feature streams, providing a unified dimensional input for subsequent cross-modal fusion. A parameter-frozen BERT-base model is used as the encoding tool, while a learnable linear projection layer is introduced to perform dimensionality mapping.

[0222] Semantic encoding: For a BERT-base model with frozen input parameters, extract the final hidden state at the [CLS] position as the aggregate representation: ;

[0223] in, ; For enterprises Structured data serialization text; Represents aggregate semantic information;

[0224] Dimensional projection: Dimensional alignment is achieved through learnable linear projection layers, thus... The aggregation representation of dimensions is mapped to Dimensionality allows structured features and textual features to reside in the same vector space, eliminating fusion barriers caused by dimensionality mismatch:

[0225] ;in, The weight matrix of the linear projection layer. The offset vector of the linear projection layer;

[0226] All The structured features of each company are stacked row by row into a structured feature matrix:

[0227] ;in, For the structured feature matrix, the first... OK Encoded the enterprise Structured field information.

[0228] In steps S1-4, the text feature matrix and the structured feature matrix are fused into a unified enterprise-level feature matrix through a cross-modal attention weighted aggregation mechanism, specifically as follows:

[0229] A cross-modal attention-weighted aggregation mechanism is used to fuse the text feature matrix and the structured feature matrix into a unified enterprise-level feature matrix, and an enterprise association and adjacency matrix is ​​constructed based on the fused enterprise-level feature matrix as auxiliary structural information.

[0230] The text feature matrix and structured feature matrix encode different information aspects of an enterprise. The relative importance of these two types of information varies among different enterprises. Technology-intensive enterprises have higher text feature content, while mature manufacturing enterprises, known for their financial indicators, have more discriminative structured features. Simple equal-weighted concatenation or summation cannot reflect this enterprise-level heterogeneity. This step employs a cross-modal attention-weighted aggregation mechanism to adaptively calculate the fusion weights of the two features for each enterprise. A unified feature matrix is ​​generated through five stages: attention scoring, Softmax normalization, weighted summation, nonlinear projection, and residual connection. A sparse adjacency matrix is ​​then constructed based on this matrix.

[0231] Attention Scoring: An attention scoring network with shared parameters is used to calculate scalar importance scores for each company's textual and structured features, quantifying the relative contribution of these two types of information to the assessment of the company's innovation capability.

[0232] ; ;

[0233] in, For the shared attention projection weight vector, Multidimensional features are compressed into scalar fractions; For shared bias; The activation function constrains the score to A range is defined to prevent numerical overflow. and For enterprises The original importance scores of textual features and structured features.

[0234] The two original scores are converted into fusion weights that sum to 1, making the weighted summation result interpretable on the same scale:

[0235] ;

[0236] ;

[0237] satisfy , ; This indicates the importance of the company's textual information to the assessment task. The larger the value, the more important the company's textual information is to the evaluation task; Indicate the importance of structured information, The larger the value, the more important the structured information.

[0238] Weighted summation: Attention weights are used to weight and aggregate textual and structured features, and the fusion ratio of the two modalities is adaptively adjusted according to the specific circumstances of each company.

[0239] ;

[0240] in, For enterprises Weighted features before fusion.

[0241] Further transformations are performed on the weighted fused features through a fully connected layer with non-linear activation to capture the non-linear interaction patterns between the two feature paths:

[0242] ;

[0243] in, To fuse the projection weight matrix, To fuse the projection bias vector, To modify the activation function of the linear unit; This represents the fused projection features; both the projection output dimension and the input dimension are 1. To support residual connections.

[0244] The weighted features before fusion and the projected features after fusion are added element-wise to mitigate information attenuation during the fusion process and ensure that the original fusion information can be retained through the direct path even if the learning effect of the projection layer is poor.

[0245] ;

[0246] All The integrated features of the companies are stacked row by row to form a unified feature matrix:

[0247] ;

[0248] In the unified feature matrix Based on this, an adjacency matrix reflecting the relationship structure between enterprises is constructed using cosine similarity calculation:

[0249] Cosine similarity calculation: for any enterprise pair Cosine similarity is calculated based on a unified feature vector:

[0250] ;in, for Norm, ;

[0251] Introducing a threshold Perform edge filtering to retain only highly similar connections. Take the 70th quantile of the similarity distribution across all enterprises to ensure that approximately 30% of the strongest connections are retained:

[0252] ;

[0253] Finally, a sparse symmetric adjacency matrix is ​​obtained. It can be viewed as a graph of enterprise relationships with enterprises as nodes and innovation similarity as edge weights.

[0254] Step S2: Construct a Transformer encoder with adjacency matrix attention constraints, and achieve global feature interaction between enterprises through a multi-head self-attention mechanism to obtain the enterprise feature matrix, specifically:

[0255] like Figure 2 As shown, the number of encoding layers in the Transformer encoder These are: a multi-head self-attention sublayer and a feedforward network sublayer;

[0256] The multi-head self-attention sublayer performs the following operations:

[0257] The Transformer encoder's self-attention mechanism is insensitive to the order of the inputs. It injects discriminative information into each position through learnable positional encoding, expressed as: ;

[0258] in, A learnable position encoding matrix, The maximum enterprise capacity preset by the system. This serves as the initial input after the location information is injected.

[0259] Multi-head self-attention is defined as including layer, Each attention head takes the output of the previous layer as its input. The input is projected into three sets of representations: Query, Key, and Value. In the enterprise evaluation scenario, each enterprise simultaneously acts as both the queryer and the queryee: the Query encodes the enterprise. Information retrieval requests initiated to a group are represented by a Key that encodes the matching feature signatures exposed to the outside world. The dot product of the two measures the degree of matching between enterprises. The Value carries the actual feature content of the enterprise and is weighted and aggregated according to the degree of matching to form the output.

[0260] ; ; ;

[0261] in, The first Layer The query, key, and value projection matrix of each attention head will 3D feature space mapping to 3D attention subspace; For the first Layer The query matrix of the attention head, the first Line coding enterprise Information retrieval needs within this head subspace; Let be the key matrix, the first... Line coding enterprise Externally exposed matchable signature features; Let be a value matrix, the first... Enterprises carrying out operations The actual features and contents within this head subspace;

[0262] Obtain the attention weight matrix for multi-head self-attention:

[0263] ;

[0264] in, For the first Layer Attention weight matrix for each attention head. The scaling factor is used to suppress the variance growth of high-dimensional dot products and prevent the softmax output from approaching a one-hot distribution, thus preventing gradient vanishing. For adjacency matrix Constructed attention bias matrix;

[0265] ;

[0266] For oneself ( ) and in the adjacency matrix There are connected enterprise pairs in the adjacency matrix, with a bias of 0, and no intervention is applied; Unconnected enterprises exert pressure The negative bias significantly suppresses, but does not completely reduce, the attention weights after softmax; a soft penalty is used here. Instead of hard mask Hard masks completely block the flow of information between companies with low similarity, potentially losing valuable signals contained in weak cross-domain connections; soft penalties strike a balance between focus and coverage. In all Size and The attention weights are shared within the layer and only need to be computed once; to prevent overfitting, Dropout is applied to the attention weights after softmax normalization, and during training, they are calculated probabilistically. Randomly set to zero and then renormalize.

[0267] The output of the multi-head self-attention sublayer is obtained by using weighted aggregation and multi-head concatenation:

[0268] ;

[0269] ;

[0270] in, For the first Layer The output of each attention head is obtained by weighted aggregation of the attention weight matrix and the value matrix. For the first The output of a multi-head self-attention sublayer. ; To concatenate along the feature dimensions, the concatenated dimensions are... , To output the projection matrix; Each entity captures different aspects of inter-enterprise relationships within its respective subspace, and the output projection integrates multi-perspective information into a unified representation.

[0271] The multi-head self-attention output and the sub-layer input are added together via residual connection, and then the layer is normalized.

[0272] ;

[0273] in, Holding learnable parameter scaling parameters For the first The residual connections of the multi-head self-attention sublayer and the intermediate output after layer normalization are used to connect the layers. ; For regularization operations, during the training phase, probability is used. Randomly set elements to zero and scale the remaining elements. To maintain the expected value unchanged, the reasoning stage involves an identity transformation.

[0274] The feedforward network sublayer performs the following operations:

[0275] Two-layer feedforward transform: ;

[0276] in, For the first The output of the feedforward network sublayer, This is the first-level up-dimensional weight matrix. This is the first-level up-dimensional weight bias vector; This is the second-layer dimensionality reduction weight matrix. This is the second-layer dimensionality reduction weight bias vector; all bias vectors are broadcast row-by-row to OK;

[0277] Residual connectivity and layer normalization: ;

[0278] in, For the first The final output of the Transformer encoding layer, At the same time, as the first Layer input; It has independent learnable scaling parameters With offset parameter ;

[0279] Output after encoding by the Transformer encoder: ;in, This is the enterprise characteristic matrix.

[0280] Step S3 involves designing a gating network mechanism to dynamically modulate the enterprise feature matrix element-wise, and then aggregating it with the unified feature matrix through residual connections and layer normalization to form a deep fusion feature. Specifically:

[0281] The enterprise feature matrix contains rich information on inter-enterprise interactions, but not all feature dimensions are equally important to the innovation capability assessment task. Some dimensions may encode redundant patterns that are irrelevant to the assessment objective. By using a gating network mechanism, the relevance of each feature dimension to the assessment task is automatically learned, key dimensions are amplified, noisy dimensions are suppressed, and the refined results are aggregated with the original fusion information of the unified feature matrix through residual connections.

[0282] Step S3-1: Generate gating weights:

[0283] Enterprise feature matrix With the unified feature matrix The input to the gating network is formed by concatenating the features along their respective dimensions. ;

[0284] in, This indicates concatenation along the feature dimension. The enterprise feature matrix output by the Transformer encoder. To unify the feature matrix; the dimensions after concatenation are from Expand to This provides richer criteria for gating networks;

[0285] The two are combined rather than just used The reason for using it as a gating input is: Carrying the original multi-source fusion information from step S1, Carrying inter-enterprise interaction information after multi-layer self-attention transformation by the Transformer encoder, the two form a complementary perspective, enabling the gating network to comprehensively encode the two information layers before and after to judge the importance of each dimension.

[0286] Element-wise gated weights are generated using a fully connected layer and a Sigmoid activation function:

[0287] ;in, The weight matrix is ​​a learnable weight matrix; A learnable bias vector; The sigmoid function compresses each output element to... ; Indicates enterprise The The relevance of 3D deep encoding features to the evaluation task A value close to 1 indicates that the dimension carries a key signal and should be retained, while a value close to 0 indicates that the dimension is noise or redundancy and should be suppressed.

[0288] Step S3-2: Gated modulation and residual output:

[0289] Each element acts as a scalar coefficient on The corresponding elements independently control the on / off states for each feature dimension:

[0290] ;in, This is element-wise multiplication; This represents the enterprise feature matrix after gating refinement;

[0291] After gating modulation With enterprise feature matrix By adding element-wise, we obtain the deep fusion feature matrix, expressed as:

[0292] ;

[0293] in, Represents the deep fusion feature matrix; This represents the layer normalization operation; LayerNorm balances the numerical distribution of the summed residuals.

[0294] The deeply fused feature matrix encodes each firm's relative position within the group and its gated key features; however, when the training data covers multiple heterogeneous domains, Inevitably, there are domain-specific distribution biases involved. For example, there are systematic differences in the distribution of R&D investment between "technology" companies and manufacturing companies. These differences reflect domain attributes rather than innovation capabilities themselves; if we directly... When inputting into downstream evaluation modules, the model may over-rely on these domain biases for prediction, leading to a decline in cross-domain generalization performance. Simultaneously, completely stripping away domain information introduces another contradiction: certain discriminative information crucial for assessing innovation capabilities manifests differently across domains. For example, "technology transfer efficiency" in the biomedical field manifests as the clinical translation of patents and the progress of new drug approvals, while in software technology it is reflected in code contribution and product iteration speed. While purely domain-invariant features capture the general concept of "efficiency," they may lose the specific manifestations characteristic of each domain.

[0295] Step S4: Construct a DANN model and use the DANN model to obtain the final fitting features;

[0296] like Figure 3 As shown, the DANN model includes a feature encoder. Gradient Reversal Layer (GRL) and Neighborhood Discriminator Using feature encoders Extracting common feature representations shared across domains:

[0297] Feature encoder For a single-layer fully connected network, 3D deep features compressed to Bottleneck space: ;

[0298] in, The weight matrix of the feature encoder. This is the bias vector of the feature encoder; For enterprises Encoding representation in the bottleneck space;

[0299] The behavior of the gradient inversion layer (GRL) is asymmetric in forward and backward propagation, with the forward propagation being an identity mapping. If the gradient direction is reversed, then the gradient direction is reversed and... Controlling the intensity of reversal: ;in, For gradient tensors, The gradient reversal coefficient;

[0300] Gradient reversal coefficient The sigmoid scheduling function proposed by DANN is adopted:

[0301] ;

[0302] in, This represents the percentage of training progress. For the current iteration step, This represents the total number of iterations.

[0303] Domain Discriminator It is a two-layer fully connected network that receives encoded features passed through the gradient inversion layer (GRL). And predict the domain origin of each sample, the first layer will 3D encoded features are projected onto nonlinear transformations. Dimensional discriminative subspace: ;

[0304] in, This is the weight matrix of the first layer of the discriminator. This is the bias vector for the first layer of the discriminator;

[0305] The second layer linearly projects the discriminant subspace features to... Each domain category is normalized to a probability distribution using softmax:

[0306] ;

[0307] in, This is the weight matrix for the second layer of the discriminator. This is the bias vector for the second layer of the discriminator; This represents the total number of domain categories. For enterprises Predicted probability distributions belonging to various fields, ; The hidden layer activation feature matrix of the discriminator;

[0308] Domain-Discrimination Cross-Entropy Loss: Cross-entropy loss measures the deviation between the discriminator's predicted domain probability distribution and the true domain labels. The expression is:

[0309] ;

[0310] in, The one-hot encoding matrix for real-world domain labels; It means if and only if the enterprise Belongs to the One field;

[0311] Define the encoder output as a neighborhood-invariant feature:

[0312] ;

[0313] Adversarial training for DANNs is automatically performed within each iteration step using a gradient inversion (GRL) layer, eliminating the need for manual switching of optimization targets; discriminator parameters take over The regular gradient is minimized; encoder parameters Receive the gradient after GRL inversion, and perform maximization equivalently. .

[0314] Define the enterprise feature matrix output by the Transformer encoder as a domain-invariant feature: ;

[0315] Domain-invariant features It has the ability to generalize across domains, but may lose key discriminative information that varies from domain to domain. The loss of domain-invariant feature information can be made up by restoring and enhancing these domain-specific information.

[0316] The HyperNetwork is introduced, whose core idea is to "generate the parameters of another network from one network". It dynamically synthesizes feature transformation matrices specific to each domain based on the domain embedding vector, instead of statically storing a set of independent parameters for each domain. All domains share the same set of parameter generation network, and the differences between domains are entirely driven by the input domain embedding. Adding a new domain only requires adding an embedding vector without adding an independent parameter set.

[0317] HyperNetwork has two improvements. The first improvement is that the transformation matrix is ​​decomposed into two low-rank factors. The number of parameters was reduced from approximately 16.8M to approximately 4.2M; the second improvement is to calculate the attention weight of the embedding of all domains based on the enterprise characteristics to generate personalized hybrid embeddings, so that cross-domain enterprises can obtain embeddings that reflect their multi-domain characteristics, while single-domain enterprises automatically degenerate into hard lookup table behavior.

[0318] The original hypernetic design uses hard labels Domain embeddings are obtained by looking up tables. However, in the scenario of assessing enterprise innovation capabilities, the hard label assumption deviates from reality: enterprises operating across industries are forced to use single domain embeddings, losing their cross-domain characteristics; enterprises with huge differences in subcategories within the same domain are forced to use the same embeddings, ignoring fine-grained differences.

[0319] We propose an improved Feature-Guided Soft Domain Routing (FSDR) approach. This approach uses the domain-invariant features of each enterprise as the query and performs attention weighting on the domain embedding table to generate a personalized domain embedding hybrid vector for each enterprise.

[0320] Step S4-2: Generate domain-specific features based on HyperNetwork and feature-guided soft domain routing (FSDR):

[0321] Step S4-2-1: Calculate feature projection and soft route weights: The domain-invariant feature of dimension is projected to A dimensional domain embedding space is defined, with the projection result as the query and a learnable domain embedding table as the query. Each row is a key; calculate the scaled dot product attention: ; ;

[0322] in, The learnable route projection matrix; The OK For the first Learnable embedding vectors for each domain, initialized according to... normal distribution; This is the scaling factor for the software router weight; Indicates enterprise Belonging to the domain The soft weights, softmax guarantees and ;

[0323] Interpretability of routing weights: Naturally formed enterprises Domain profile, for example The model suggests that the company is 60% technology and 30% manufacturing, reflecting its cross-industry characteristics.

[0324] Weighted hybridization generates personalized domain embeddings: ;

[0325] in, For enterprises Personalized domain embedding; degenerate relationship with the original hard lookup table: when When it degenerates into a one-hot vector, The strict degradation to the standard HyperNetwork's hard lookup table scheme indicates that FSDR is a strict generalization of the original.

[0326] Domain label prior regularization: To prevent routing weights from deviating completely from known domain labels, a regularization loss is introduced.

[0327] ;

[0328] The penalty for the loss However, due to the weight The weight of the loss is much smaller than that of the main task, serving only as a soft anchor and encouraging... It is not zero, but it is not forced to be 1, allowing the model to allocate some weights to other related domains.

[0329] Step S4-2-2: Dynamic generation of low-rank transformation matrix: embedding personalized domain The parameter generation module of the input hypernetwork dynamically synthesizes two low-rank factors of the enterprise-specific transformation matrix; the low-rank decomposition will... The complete transformation matrix is ​​decomposed into left factors With right factor The product of these factors significantly reduces the number of parameters that the supernetwork needs to output.

[0330] The generation structure of the two factors is symmetrical and can be expressed in a unified form: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] ,set up This represents the total number of elements corresponding to the factor. Let be the target shape corresponding to the factor, where Let be the rank of the low-rank decomposition. Domain-specific feature dimensions:

[0331] ; ;

[0332] in, The factor generation matrix of the hypernetwork; For hypernetworks targeting enterprises A flat parameter vector embedded in the personalized domain; This indicates that HyperNetworks is for enterprises Dynamically generated A domain-specific weight matrix with low-rank factors; , ; ;

[0333] Combined into a complete transformation matrix: ;in, For enterprises The left factor of the specific transformation matrix captures from dimensional invariant space to Compressed mapping of intermediate space in dimensionality; For the right factor, capture from dimensional invariant space to Compressed mapping of intermediate space in dimensionality; For enterprises Dedicated complete transformation matrix

[0334] Step S4-2-3: Project the domain-invariant features onto the domain-specific space using a dynamic transformation matrix to extract domain-specific features:

[0335] By using dynamically generated transformation matrices, the domain-invariant features of each sample are projected onto the characteristic feature space of its domain; Encoded the enterprise Domain-independent general pattern The code encodes the feature transformation rules of the enterprise's domain. Multiplying the two results in re-expressing the general pattern in a domain-specific form according to the domain rules.

[0336] ;

[0337] in, For enterprises The domain-invariant eigenvectors; For enterprises Domain-specific feature vectors;

[0338] For all Each sample undergoes the above projection and is stacked along the batch dimension:

[0339] ;

[0340] Step S4-3: Adaptively aggregate domain-invariant and domain-specific features into the final fitting features using a gated network and dual-path projection.

[0341] and The innovation characteristics of each company are described from two complementary perspectives. Different samples rely on the two types of information to varying degrees: companies with highly focused businesses may benefit more from domain-specific characteristics, while diversified companies with cross-industry businesses may rely more on general characteristics.

[0342] Step S4-3-1: Calculate the adaptive aggregation weights:

[0343] A scalar gating weight is generated for each sample using a gating network. This quantifies the degree to which the sample depends on general features; gating networks use... For input, As the common source of both characteristics, its own pattern is sufficient to reflect the domain affiliation characteristics of an enterprise, avoiding the introduction of... The resulting circular dependency has the following aggregate weight:

[0344] ;

[0345] in, For the learnable weight vector of the gating network, For the learnable bias scalar of the gated network; The Sigmoid function compresses the output to... interval; For enterprises The gating weight scalar, A value close to 1 indicates that the sample should focus on domain-invariant features, while a value close to 0 indicates that the sample should focus on domain-specific features.

[0346] Step S4-3-2: Concatenate the two types of features along the feature dimension and achieve information fusion through two parallel projection paths. By performing indiscriminate linear combination of spliced ​​features, cross-type feature interaction patterns can be captured. The general features are subjected to sample-level intensity adjustment by gating weights, and the sum of the two is then activated by ReLU to obtain the final adapted features:

[0347] ;

[0348] in, , Broadcast along the feature dimension and Element-by-element multiplication; The projection matrix of the gated general feature; This is the feature projection matrix for dual-path projection stitching. This is the dual-path projection bias vector; To ultimately adapt the feature dimensions.

[0349] The total loss function is: ;

[0350] in, The primary objective of the downstream innovation capability assessment task; This is an adversarial constraint on domain invariance. For FSDR, use the domain label prior regularization; The discriminator parameters are subjected to conventional minimization, while the encoder parameters are subjected to equivalent maximization after GRL inversion. This minimax game is automatically implemented by GRL in the computation graph, and the entire network can complete the joint update of all parameters in a single forward-backward propagation.

[0351] Final adaptation features Transformed into an interpretable and comparable score for enterprise innovation capabilities;

[0352] In step S5, outliers are suppressed, a hybridized baseline value is dynamically generated, and a piecewise nonlinear mapping is performed to obtain a dynamic score for enterprise innovation. Specifically:

[0353] Step S5-1: Use the scoring prediction head to obtain the predicted value of the company's continuous innovation capability:

[0354] The scoring prediction head uses a two-layer fully connected network. The first layer will... 3D adaptation features are projected onto the nonlinear transformation. Dimensional subspace: ;

[0355] in, The hidden layer weight matrix is... This is the hidden layer bias vector; For final adaptation features;

[0356] The second layer compresses the hidden features into unconstrained one-dimensional real values ​​through linear projection, serving as a continuous prediction of each company's innovation capability by the model: ;

[0357] in, To output the weight vector, To output the bias scalar; For the model of enterprises Continuous predicted values ​​of innovation capability; the score prediction vector is:

[0358] ;in, It is the transpose symbol;

[0359] In step S5-2, robust standardization based on quartiles and adaptive decay of outliers are used to eliminate dimensional differences and suppress the disturbance of extreme values ​​on population statistics. Specifically, this is manifested as follows:

[0360] Step S5-2-1: Calculate the population quantile statistic:

[0361] Group median: ;

[0362] Interquartile range: ;

[0363] in, Used as a location reference instead of the mean. As a scale reference, it can be used as an alternative to standard deviation.

[0364] Step S5-2-2: Calculate the adaptive penalty intensity coefficient for outliers:

[0365] The degree of anomaly of each sample is quantified based on how much it exceeds the Tukey boxplot criterion, providing an adaptive penalty for subsequent decay:

[0366] ;

[0367] in, It is the numerical stability constant. For indicator functions, the penalty intensity coefficient Determined by the degree to which the sample deviates from the median:

[0368] Step S5-2-3: Robust normalization and adaptive decay of outliers, the expression is:

[0369] ;

[0370] When not outliers Furthermore, when the indicator function is set to 0, the normalization degenerates into the classic robust form; for outliers... And the indicator function is 1, the molecule is Compression increases with the extreme deviation from the target.

[0371] For all The robust standardized scoring vector is formed by the following companies performing the above operations:

[0372] ;

[0373] Each element Reflecting on enterprises The relative position of the predicted innovation capacity value in the current population has eliminated the difference in dimensions and suppressed the amplitude disturbance of outliers.

[0374] In step S5-3, a dynamic baseline value with confounding correction is generated based on the current population distribution, and then transformed into the final score through piecewise nonlinear mapping, specifically as follows:

[0375] Step S5-3-1: Construct a feature matrix of confounding factors and remove systematic biases using OLS regression:

[0376] Construct a feature vector of confounding factors for each enterprise. ,in, Unique hot coding in the industry, To evaluate the yearly unique heat encoding, all samples are stacked into a confounding feature matrix. ;

[0377] The regression coefficients of each confounding factor were estimated by fitting a linear model using OLS:

[0378] ;

[0379] in, This represents the baseline offset of the score when all confounding factors are zero. For the first Regression coefficients of each confounding factor; Number of year categories; The first characteristic matrix of confounding factors Line number Column elements; The impact of industry classification and assessment year on enterprises was comprehensively quantified. The systematic impact of standardized scoring;

[0380] Remove bias from standardized scoring: ;

[0381] in, The corrected standardized score;

[0382] Step S5-3-2: Calculate the relative reference position index:

[0383] Relative benchmark indicators include the median deviation indicator and the head-to-head distance indicator. The median deviation indicator measures a company's... The relative degree of deviation from the group median is expressed as: ;

[0384] Head distance index measures enterprises The relative distance from the head threshold is expressed as: ;

[0385] in, Indicates the measurement of enterprises The relative degree of deviation from the group median; This indicates the relative distance from the head threshold; Indicates the median level of the current group. This indicates the threshold for the leading level of the current group;

[0386] Step S5-3-3: Calculate the domain semantic calibration term and synthesize the comprehensive dynamic benchmark value:

[0387] Embedding personalization domain With routing query vector The degree of similarity between the two is measured by cosine similarity:

[0388] ;

[0389] A high value indicates that the company's characteristics are highly consistent with its industry profile, and the benchmark is adjusted appropriately; a low value indicates that the company deviates from its industry profile, and the calibration item reduces the impact of this deviation on the benchmark.

[0390] The three dimensions of indicators are linearly combined with fixed weights to synthesize a comprehensive dynamic benchmark value:

[0391] ;

[0392] in, , , ; Comprehensive reflection of enterprises After adjusting for confounding factors, the performance relative to the current group median and head level was calibrated for domain similarity.

[0393] Step S5-3-4: Segmentation using a nonlinear mapping function:

[0394] when At this point, for the lower-score segment, a superlinear mapping is used, with slower penalty for lagging companies, thus enhancing the differentiation of bottom-tier companies. The expression is: ;

[0395] when For the middle segment, using a strict linear mapping, the scores of the main enterprises are evenly distributed in the range of 40–60 points to maintain the fairness of the assessment. The expression is: ;

[0396] when At that time, a sublinear mapping is used to suppress the head saturation effect, and the expression is: ;

[0397] Final evaluation set: , ;

[0398] As a benchmark, To upgrade, It is a level of attention; the score is generated based on the current data, and the benchmark is dynamically adjusted according to the composition and performance of the assessed group, which overcomes the lag of static standards and can sensitively reflect the latest trend of productivity development.

[0399] To verify the effectiveness of the technical solution of this invention, the following experimental simulation environment was built. The hardware platform uses a high-performance computing server equipped with an NVIDIA A100 graphics processor, specifically an Intel Xeon Gold 6330 (2.0GHz, 28 cores), and 256GB of system memory. The software environment consists of an Ubuntu 20.04 operating system, the deep learning framework PyTorch 2.0.1, and CUDA 11.8.

[0400] The experimental data were derived from publicly available invention patent bibliographic data, full-text claims, abstracts, and citation data from the State Intellectual Property Office. After cleaning, a total of 127,536 valid patent samples were retained, covering six technical fields: manufacturing, information technology, new energy, high-end equipment, biomedicine, and new materials. A comprehensive value score was constructed using patent maintenance period, citation frequency, and licensing / transfer records as monitoring labels, and the values ​​were normalized to [value missing]. Intervals. Datasets are categorized by... The dataset is divided into training, validation, and test sets, and stratified sampling is used to ensure that the sample proportions are consistent across different domains.

[0401] The model was trained using the AdamW optimizer with an initial learning rate of [missing information]. Weight decay The cosine annealing scheduling strategy was used, with a total of 200 training rounds and a batch size of 256. All experiments were independently repeated three times, and the average value was taken. The evaluation metrics were mean squared error (MSE, the lower the better) and Spearman rank correlation coefficient (SCR). (The higher the better).

[0402] To comprehensively evaluate the contribution of each technical aspect of this invention, the following five comparison schemes are set up: The method of this invention: includes all technical aspects, namely multi-source feature extraction and fusion, Transformer-based swarm interaction coding, and a domain-adaptive evaluation mechanism composed of domain adversarial networks, super-network domain-specific parameter generation, and few-shot domain regularization; MLP baseline method: uses only the same original feature input as the method of this invention, directly mapping it to value scoring via a fully connected multilayer perceptron, without including swarm interaction coding and the domain-adaptive mechanism, serving as a performance lower bound benchmark; Ablation variant: removes the domain-adaptive mechanism): in this... Based on the method of the invention, three techniques—domain adversarial network, hypernetwork parameter generation, and few-sample domain regularization—are removed simultaneously, and all domains share the same evaluation head parameters; Ablation variant (removal of group interaction coding): Based on the method of the invention, the Transformer group interaction coding module is removed, and relative ranking and association information in the historical patent group of the same applicant are no longer used, and evaluation is based only on the features of a single patent; Ablation variant (removal of hypernetwork domain specificity): Based on the method of the invention, only the hypernetwork parameter generation submodule is removed, and the domain adversarial network is retained for feature space alignment, but all domains share the evaluation head parameters.

[0403] like Figure 4 As shown, in terms of convergence speed, the method of this invention enters a stable convergence range around the 60th round, while the MLP baseline method still shows a significant decrease until the 100th round, indicating that the Transformer population encoding and domain adaptation mechanism can effectively accelerate model fitting; in terms of final state accuracy, the validation set MSE of the method of this invention is stable at 0.0308, while the MLP baseline method is stuck at 0.0517, and the method of this invention achieves a 40.4% reduction in MSE.

[0404] The performance of the two ablation variants further reveals the differences in the contributions of each technical component. After removing the swarm interaction encoding, the final-state MSE is 0.0372, with a convergence speed close to the complete system but slightly lower final-state accuracy, indicating that swarm ranking information mainly affects prediction accuracy rather than learning efficiency. After removing the domain adaptation mechanism, the final-state MSE increases to 0.0453, and the convergence speed also slows significantly, indicating that the domain adversarial network begins to play a cross-domain feature alignment role early in training. Throughout the training process, the four curves consistently maintain the order of superiority over the inventive method (superior to the method removing the swarm interaction encoding), which is superior to the method removing the domain adaptation mechanism, which is superior to the MLP baseline method, with no overlap, indicating that the gains of each module remain stable at different training stages.

[0405] like Figure 5 As shown, the MSE histograms for each scheme in the five technical fields of information technology, new energy, high-end equipment, biomedicine, and new materials are presented. The error bars represent the standard deviation of three independent runs.

[0406] The method of this invention achieved the lowest MSE in all five domains, verifying its cross-domain generalization ability. The gain magnitudes in different domains exhibited significant differences: in the biomedical and new materials domains, where data heterogeneity is high, the method of this invention achieved MSE reductions of 49.3% and 52.1% relative to the baseline, respectively; while in the information technology field, where patterns are more consistent, the reduction was 18.2%. This difference indicates that the greater the data distribution shift between different domains, the more significant the correction effect brought about by the domain adaptation mechanism composed of domain adversarial networks and hypernetworks.

[0407] Ablation analysis provides a more detailed explanation of this phenomenon. After removing all domain adaptation mechanisms, the MSE in the biomedical domain reaches a high of 0.0631, close to the baseline level of 0.0672, indicating that the model has almost lost its adaptability to this domain. However, by only removing the hypernetwork domain-specification method and retaining the domain adversarial network, the gap between the model and the complete system is only 4.6% in the information technology domain, but the gap widens to 41.0% in the new materials domain. This result suggests that in relatively homogeneous domains, feature space alignment by the domain adversarial network is sufficiently effective; however, in domains with significant distributional differences, it is still necessary to rely on the hypernetwork to generate customized evaluation head parameters for each domain to fully bridge the inter-domain gap.

[0408] Furthermore, the standard deviation of the five-domain MSE of the method of the present invention is approximately 0.0036, while that of the baseline method is 0.0168, the latter being 4.7 times that of the former. This indicates that the method of the present invention not only reduces the overall error but also significantly narrows the performance gap between different domains.

[0409] like Figure 6 As shown, the complete system of this invention significantly outperforms the baseline method in two metrics: MSE is reduced from 0.0517 to 0.0308 (a reduction of 40.4%). The value increased from 0.8936 to 0.9531 (an increase of 6.66%), demonstrating that the method of the present invention has a substantial improvement in both absolute error and prediction ranking accuracy.

[0410] The contributions of each technical component exhibit a clear hierarchical structure. After removing group interaction coding, the MSE increases to 0.0372. The MSE dropped to 0.9382, a relatively manageable degradation, as the model can still be evaluated based on individual patent features. Further removal of the supernetwork domain specificity increased the MSE to 0.0421. The MSE dropped to 0.9214, indicating a significant worsening of the degradation; after completely removing the domain adaptation mechanism, the MSE rose to 0.0453. The value dropped to 0.9073, showing the largest degradation. Quantifying the contributions of each module: the domain adaptation mechanism's MSE contribution accounted for 69.4% of the total gain, and the group interaction coding accounted for 30.6%, with the sum of the two being 100%, indicating good additivity of gains between modules. All five schemes were completely consistent in their ranking on both metrics, and there was no overlap in error bars between adjacent schemes, indicating that the gains of each component were statistically significant.

[0411] Experimental results show that the method of the present invention improves overall prediction accuracy (MSE reduction of 40.4%) and ranking correlation ( The method significantly outperforms traditional baseline methods in both core metrics (6.66% improvement) and the advantage remains stable throughout the training process. The method maintains a leading position in different technical fields, with MSE reductions of up to 49%–52% in challenging domains with high data heterogeneity. Ablation experiments quantitatively confirm that the domain adaptation mechanism contributes about 70% of the gain and the group interaction coding contributes about 30% of the gain, with good complementarity between the two.

[0412] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.< / s> < / s> < / s>

Claims

1. A method for assessing innovation capabilities based on large language models and multimodal semantic understanding, characterized by: The method includes the following steps: Step S1: Collect multi-source heterogeneous data from both inside and outside the enterprise and classify it. Then, use the BERT-base model and the RoBERTa-large model to process the classified multi-source heterogeneous data and generate a unified feature matrix. Step S2: Construct a Transformer encoder with adjacency matrix attention constraints, realize global feature interaction between enterprises through multi-head self-attention mechanism, and output enterprise feature matrix; Step S3: Design a gating network to dynamically modulate the enterprise feature matrix element by element, and then aggregate it with the unified feature matrix through residual connection and layer normalization to form a deep fusion feature; Step S4: Extract domain-invariant features through adversarial training, dynamically generate domain-specific features through a hypernetwork, and obtain the final fitting features through gated aggregation; Step S5: Obtain the dynamic score of enterprise innovation by suppressing outliers, dynamically generating a baseline value after hybrid correction, and performing piecewise nonlinear mapping.

2. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 1, characterized in that: In step S1, multi-source heterogeneous data from both inside and outside the enterprise are collected and classified. The classified multi-source heterogeneous data are processed using the BERT-base model and the RoBERTa-large model, respectively. Specifically: Step S1-1: Define the sample enterprise set, obtain multi-source heterogeneous data from the sample enterprise set, and classify the multi-source heterogeneous data, specifically as follows: Define the sample enterprise set , The total number of companies participating in the assessment; each company The acquired data was divided into two categories: structured data and... and text data Structured data is a collection of structured information carried by numerical values ​​or categories, while text data is a collection of unstructured information carried by natural language. Step S1-2: Process text data using the RoBERTa-large model: Step S1-2-1: For each enterprise raw text data set The text data set is preprocessed by sequentially performing cleaning, deduplication, and quality filtering operations. ;in, For enterprises The number of text documents; Step S1-2-2: Perform multi-document aggregation on the preprocessed text data set: Sort by information priority: Chinese documents are arranged according to a predefined fixed priority; Concatenation and Truncation: The sorted documents are concatenated sequentially, separated by RoberTa delimiters, and a start marker is added to the beginning of the sequence. <s>The total length is truncated at the token level to the maximum input length. tokens, expressed as:< / s> <s> ;in, It is an aggregated sequence; Step S1-2-3: Aggregate the sequence Feature extraction using the RoBERTa-large model: The RoBERTa-large model serves as a fixed feature extractor, and a parameter freezing strategy is employed to freeze the training parameters of the RoBERTa-large model to a fixed number. The RoBERTa-large model extracts features, specifically as follows: Extract sequence start marker <s> The corresponding hidden state vector is used as the aggregate semantic representation of the entire text sequence to obtain the text feature line vector:< / s> <s> ; in, Represents a sequence matrix; All The text feature row vectors of each company are stacked row by row to form a text feature matrix, expressed as: ; in, ; Steps S1-3: Process structured data using BERT-base; Steps S1-4: The two features are fused into a unified feature matrix through a cross-modal attention weighted aggregation mechanism, and an enterprise association adjacency matrix is ​​constructed based on the fused enterprise-level feature matrix as auxiliary structural information.

3. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 2, characterized in that: In steps S1-3, BERT-base is used to process structured data, specifically as follows: Each company Structured field records , This represents the total number of structured fields. This is the value of the Mth key-value pair field; Structured data undergoes unified preprocessing, which is accomplished through a predefined field configuration table. Drive; among which, For field type, For unit text, The number of significant digits to retain; The name of the Mth key-value pair field; Step S1-3-2: Convert the preprocessed field records into a text input sequence acceptable to BERT-base according to a predefined template and fixed order: Field arrangement order: Fields are arranged in a fixed order according to the indicator hierarchy of innovation capability assessment; Serialization template: Each field is converted into a text fragment in the format "field name: field value", with fragments connected by [SEP], and a [CLS] marker is added to the beginning of the sequence. The expression is: ; Step S1-3-3: Perform dimensional projection using the BERT-base model: The BERT-base model is used as a word segmenter. It employs a parameter freezing mode to freeze all training parameters and extracts the final hidden state at the [CLS] position as the aggregate representation. ; in, ; For enterprises Structured data serialization text; Represents aggregate semantic information; Semantic encoding: Dimension alignment is achieved through a learnable linear projection layer. ;in, The weight matrix of the linear projection layer. The offset vector of the linear projection layer; All The structured features of each company are stacked row by row into a structured feature matrix: ; In steps S1-4, the text feature matrix and the structured feature matrix are fused into a unified enterprise-level feature matrix through a cross-modal attention weighted aggregation mechanism, specifically as follows: The text feature matrix and the structured feature matrix are fed into a shared-parameter attention scoring network, and first projected onto... Hidden space, through After activation, it is compressed into a scalar fraction; The original scores are converted into fusion weights that sum to 1 by Softmax normalization. Attention weights are used to perform a weighted summation of the text feature matrix and the structured feature matrix, and the fusion ratio of the two modalities is adaptively adjusted according to the specific situation of each enterprise. The fusion features are further refined by using a projection layer with non-linear activation, enabling it to capture the non-linear interaction patterns between the two features. The weighted features before fusion and the projected features after fusion are added element by element to alleviate information decay during the fusion process and ensure that the original fusion information can be retained through the direct path even if the learning effect of the projection layer is not good. All The integrated features of the companies are stacked row by row to form a unified feature matrix: ;in, Indicates enterprise After cross-modal attention-weighted aggregation, it integrates the combined features of textual semantic information and structured quantitative information; In the unified feature matrix Based on this, an adjacency matrix reflecting the relationship structure between enterprises is constructed using cosine similarity calculation. ;in, .

4. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 1, characterized in that: In step S3, a gated network is designed to dynamically modulate the enterprise feature matrix element by element, and then aggregate it with the unified feature matrix through residual connections and layer normalization to form a deep fusion feature. Specifically: Step S3-1: Generate gating weights: The input to the gating network is formed by concatenating the enterprise feature matrix and the unified feature matrix along the feature dimensions. ;in, For the enterprise feature matrix, To unify the feature matrix; The total number of companies participating in the assessment; The unified feature matrix carries the original multi-source fusion information, forming a complementary perspective with the enterprise feature matrix, enabling the gating network to comprehensively encode the two information levels before and after to judge the importance of each dimension; The concatenated features are mapped to element-wise gated weights using a fully connected layer and a sigmoid activation function. ; in, The weight matrix is ​​a learnable weight matrix; A learnable bias vector; The sigmoid function compresses each output element to... ; Indicates enterprise The The relevance of 3D deep encoding features to the evaluation task A value close to 1 indicates that the dimension carries a key signal and should be retained, while a value close to 0 indicates that the dimension is noise or redundancy and should be suppressed. Step S3-2: Element-wise modulation is performed on the enterprise feature matrix using gating weights, and then aggregated into a deep fusion feature matrix through residual connections and layer normalization, specifically as follows: Using gating weights Enterprise feature matrix Each element performs on / off control independently: ;in, For element-wise multiplication, This is the enterprise feature matrix after gated refinement; Gated modulated features With the unified feature matrix Element-wise summation, followed by layer normalization, results in a deep fusion feature matrix: ; in, For residual connections; For layer normalization operation; This is the feature matrix for deep fusion.

5. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 1, characterized in that: In step S4, domain-invariant features are extracted through adversarial training, domain-specific features are dynamically generated through a hypernetwork, and the final adaptive features are obtained through gated aggregation. Specifically: Step S4-1: Construct a DANN model to obtain domain-invariant features that generalize across domains: The DANN model includes a feature encoder. Gradient Reversal Layer (GRL) and Neighborhood Discriminator Using feature encoders Extracting common feature representations shared across domains: Feature encoder For a single-layer fully connected network, 3D deep features compressed to Bottleneck space: ; in, The weight matrix of the feature encoder. This is the bias vector of the feature encoder; For enterprises Encoding representation in the bottleneck space; The behavior of the gradient inversion layer (GRL) is asymmetric in forward and backward propagation, with the forward propagation being an identity mapping. If the gradient direction is reversed, then the gradient direction is reversed and... Controlling the intensity of reversal: ;in, For gradient tensors, The gradient reversal coefficient; Gradient reversal coefficient The sigmoid scheduling function proposed by DANN is adopted: ; in, This represents the percentage of training progress. For the current iteration step, This represents the total number of iterations. Domain Discriminator It is a two-layer fully connected network that receives encoded features passed through the gradient inversion layer (GRL). And predict the domain origin of each sample, the first layer will 3D encoded features are projected onto nonlinear transformations. Dimensional discriminative subspace: ; in, This is the weight matrix of the first layer of the discriminator. This is the bias vector for the first layer of the discriminator; The second layer linearly projects the discriminant subspace features to... Each domain category is normalized to a probability distribution using softmax: ; in, This is the weight matrix for the second layer of the discriminator. This is the bias vector for the second layer of the discriminator; This represents the total number of domain categories. For enterprises Predicted probability distributions belonging to various fields, ; The hidden layer activation feature matrix of the discriminator; Domain-Discrimination Cross-Entropy Loss: Cross-entropy loss measures the deviation between the discriminator's predicted domain probability distribution and the true domain labels. The expression is: ; in, The one-hot encoding matrix for real-world domain labels; It means if and only if the enterprise Belongs to the One field; Define the encoder output as a neighborhood-invariant feature: ; Step S4-2: Based on HyperNetwork and feature-guided soft domain routing (FSDR), dynamically generate domain-specific features; Step S4-3: Through a gated network and dual-path projection, the domain-invariant features and domain-specific features are adaptively aggregated into the final fitting features.

6. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 5, characterized in that: In step S4-2, based on the HyperNetwork and feature-guided soft domain routing (FSDR), domain-specific features are dynamically generated as follows: Step S4-2-1: Calculate feature projection and soft route weights using Feature-Guided Soft Domain Routing (FSDR): Will The domain-invariant feature of dimension is projected to A domain embedding space of dimension maps enterprise features to the same semantic space as the domain embedding to perform similarity calculations: ; in, The learnable route projection matrix; For enterprises The domain-invariant eigenvectors; For enterprises Routing query vectors in the domain embedding space; Using route query vectors For querying, embed tables with learnable domains Each row is a key, and the enterprise's soft routing weights for each domain are calculated using scaled dot product attention: ; in, For the first Learnable embedding vectors for each domain; This is the scaling factor for the software router weight; Indicates enterprise Belonging to the domain The soft weights, softmax guarantees and ; By weighting and mixing the domain embedding table with soft route weights, personalized domain embeddings are synthesized for each enterprise: ; in, For enterprises Personalized domain embedding; when When it degenerates into a one-hot vector, ; To prevent route weights from deviating completely from known neighborhood labels, a regularization loss is introduced to provide a weak anchoring constraint for soft routes: ; in, Indicates loss or punishment; Step S4-2-2: Embedding Personalization Domain The parameter generation module of the input supernetwork dynamically synthesizes the enterprise-specific feature transformation matrix, generating a low-rank transformation matrix: Will The complete transformation matrix is ​​decomposed into left factors With right factor The product; The generation structure of the two factors is symmetrical and can be expressed in a unified form: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] ,set up This represents the total number of elements corresponding to the factor. Let be the target shape corresponding to the factor, where Let be the rank of the low-rank decomposition. Domain-specific feature dimensions: ; ; in, The factor generation matrix of the hypernetwork; For hypernetworks targeting enterprises A flat parameter vector embedded in the personalized domain; This indicates that HyperNetworks is for enterprises Dynamically generated A domain-specific weight matrix with low-rank factors; Combining two low-rank factors into a complete transformation matrix: ; in, For enterprises The left factor of the specific transformation matrix captures from dimensional invariant space to Compressed mapping of intermediate space in dimensionality; For the right factor, capture from dimensional invariant space to Compressed mapping of intermediate space in dimensionality; For enterprises A dedicated complete transformation matrix; Step S4-2-3: Project the domain-invariant features onto the domain-specific space using a dynamic transformation matrix to extract domain-specific features: The domain-invariant features of each sample are re-expressed in a domain-specific form according to their domain-specific rules: ; in, For enterprises The domain-invariant eigenvectors; For enterprises Domain-specific feature vectors; Encoded the enterprise A cross-domain universal model The rules for transforming the characteristics of the enterprise's domain are encoded. Multiplying the two together will re-express the general pattern as a domain-specific form according to the domain rules. For all Each sample undergoes the above projection and is stacked along the batch dimension: ;in, Indicates domain-specific characteristics.

7. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 5, characterized in that: In step S4-3, the domain-invariant features and domain-specific features are adaptively aggregated into the final fitting features through a gated network and dual-path projection. Specifically: Step S4-3-1: Design the aggregation weights of the adaptive gating network: A scalar gating weight is generated for each sample using a gating network. This quantifies the relative dependence of the sample on domain-invariant and domain-specific features: ; in, For the learnable weight vector of the gating network, For the learnable bias scalar of the gated network; The Sigmoid function compresses the output to... interval; For enterprises The gating weight scalar, A value close to 1 indicates that the sample should focus on domain-invariant features, while a value close to 0 indicates that the sample should focus on domain-specific features. Step S4-3-2: Dual-path projection and feature aggregation: First, concatenate the domain-invariant features and domain-specific features along the feature dimension: ; Then, the final adaptation features are generated through dual-path projection: ; in, This represents the domain-invariant feature after gated modulation. Broadcast along the feature dimension to Later and Element-by-element multiplication This is element-wise multiplication; For the gating general feature projection matrix, This is the feature projection matrix for dual-path projection stitching. This is the dual-path projection bias vector; To ultimately adapt the feature dimensions.

8. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 1, characterized in that: In step S5, a dynamic score for enterprise innovation is obtained by suppressing outliers, dynamically generating a base value corrected for promiscuity, and performing piecewise nonlinear mapping. Specifically: Step S5-1: Using the scoring prediction head, compress the high-dimensional fitting features into a predicted value of the enterprise's continuous innovation capability: The scoring prediction head uses a two-layer fully connected network. The first layer will... 3D adaptation features are projected onto the nonlinear transformation. Dimensional subspace: ; in, The hidden layer weight matrix is... This is the hidden layer bias vector; For final adaptation features; The second layer compresses the hidden features into unconstrained one-dimensional real values ​​through linear projection, serving as a continuous prediction of each company's innovation capability by the model: ; in, To output the weight vector, To output the bias scalar; For the model of enterprises Continuous predicted values ​​of innovation capability; the score prediction vector is: ;in, It is the transpose symbol; Step S5-2: Eliminate dimensional differences and suppress the perturbation of population statistics by extreme values ​​through robust standardization based on quartiles and adaptive decay of outliers; Step S5-3: Generate a dynamic benchmark value with confounding correction based on the current population distribution, and convert it into the final score through piecewise nonlinear mapping.

9. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 8, characterized in that: In step S5-2, robust standardization based on quartiles and adaptive decay of outliers are used to eliminate dimensional differences and suppress the disturbance of extreme values ​​to population statistics. Specifically, this is manifested as follows: Step S5-2-1: Calculate the population quantile statistic: Group median: ; Interquartile range: ; in, Used as a location reference instead of the mean. As a scale reference, it can be used as an alternative to standard deviation. Step S5-2-2: Calculate the adaptive penalty intensity coefficient for outliers: The degree of anomaly of each sample is quantified based on how much it exceeds the Tukey boxplot criterion, providing an adaptive penalty for subsequent decay: ; in, It is the numerical stability constant. For indicator functions, the penalty intensity coefficient Determined by the degree to which the sample deviates from the median: Step S5-2-3: Robust normalization and adaptive decay of outliers, the expression is: ; When not outliers Furthermore, when the indicator function is set to 0, the normalization degenerates into the classic robust form; for outliers... And the indicator function is 1, the molecule is Compression increases with the extreme deviation from the target. For all The robust standardized scoring vector is formed by the following companies performing the above operations: ; Each element Reflecting on enterprises The relative position of the predicted innovation capacity value in the current population has eliminated the difference in dimensions and suppressed the amplitude disturbance of outliers.

10. The innovation capability assessment method based on large language models and multimodal semantic understanding according to claim 8, characterized in that: In step S5-3, a dynamic benchmark value with confounding correction is generated based on the current population distribution, and then transformed into the final score through piecewise nonlinear mapping, specifically as follows: Step S5-3-1: Construct a feature matrix of confounding factors and remove systematic biases using OLS regression: Construct a feature vector of confounding factors for each enterprise. ,in, Unique hot coding in the industry, To evaluate the yearly unique heat encoding, all samples are stacked into a confounding feature matrix. ; The regression coefficients of each confounding factor were estimated by fitting a linear model using OLS: ; in, This represents the baseline offset of the score when all confounding factors are zero. For the first Regression coefficients of each confounding factor; Number of year categories; The first characteristic matrix of confounding factors Line number Column elements; The impact of industry classification and assessment year on enterprises was comprehensively quantified. The systematic impact of standardized scoring; Remove bias from standardized scoring: ; in, The corrected standardized score; Step S5-3-2: Calculate the relative reference position index: Relative benchmark indicators include the median deviation indicator and the head-to-head distance indicator. The median deviation indicator measures a company's... The relative degree of deviation from the group median is expressed as: ; Head distance index measures enterprises The relative distance from the head threshold is expressed as: ; in, Indicates the measurement of enterprises The relative degree of deviation from the group median; This indicates the relative distance from the head threshold; Indicates the median level of the current group. This indicates the threshold for the leading level of the current group; Step S5-3-3: Calculate the domain semantic calibration term and synthesize the comprehensive dynamic benchmark value: Embedding personalization domain With routing query vector The degree of similarity between the two is measured by cosine similarity: ; A high value indicates that the company's characteristics are highly consistent with its industry profile, and the benchmark is adjusted appropriately; a low value indicates that the company deviates from its industry profile, and the calibration item reduces the impact of this deviation on the benchmark. The three dimensions of indicators are linearly combined with fixed weights to synthesize a comprehensive dynamic benchmark value: ; in, , , ; Comprehensive reflection of enterprises After adjusting for confounding factors, the performance relative to the current group median and head level was calibrated for domain similarity. Step S5-3-4: Segmentation using a nonlinear mapping function: when At this stage, for the lower-score segment, a superlinear mapping is used, with slower penalty for lagging companies, thus enhancing the differentiation of bottom-tier companies. The expression is: ; when At this point, representing the middle segment, a strict linear mapping is used to ensure that the scores of the main enterprises are evenly distributed between 40 and 60 points, maintaining the fairness of the assessment. The expression is: ; when At that time, a sublinear mapping is used to suppress the head saturation effect, and the expression is: ; Final evaluation set: , ; As a benchmark, To upgrade, The rating is at the attention level; the score is generated based on the current data, and the benchmark is dynamically adjusted according to the composition and performance of the assessed group, which overcomes the lag of static standards and can sensitively reflect the latest trends in productivity development. < / s> < / s>