Scientific and technological achievement data fusion method based on big data

By performing heterogeneous modeling, differential feature extraction, and dynamic weight adjustment on multi-source scientific and technological achievement data, the problem of insufficient semantic consistency in multi-source scientific and technological achievement data fusion was solved, and adaptive association and stability improvement between semantic nodes were achieved.

CN121524943APending Publication Date: 2026-02-13NANJING DATA ASSOCIATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511707234.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing multi-source scientific and technological achievement data fusion methods have failed to effectively address the dynamic changes in semantic hierarchical structure differences and contextual dependencies, resulting in insufficient semantic consistency and imbalance of semantic node weights.

Method used

By collecting multi-source scientific and technological achievement data, heterogeneous modeling and data format unification are carried out, differential features are extracted, differential feature vectors are generated, lightweight differential feature blocks are obtained by combining time window compression, multi-layer semantic relationship matrix is ​​constructed, semantic similarity and context relevance are fused, semantic node fusion weights are dynamically adjusted, and semantically consistent fusion results are generated.

Benefits of technology

It achieves an adaptive distribution of the association strength between semantic nodes, maintains the overall stability of the semantic relationship network, accurately reflects the real dependency relationship between different semantic levels, and improves semantic consistency and fusion accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524943A_ABST
    Figure CN121524943A_ABST
Patent Text Reader

Abstract

The invention discloses a scientific and technological achievement data fusion method based on big data, and relates to the technical field of data processing, and the method comprises the following steps: collecting multi-source scientific and technological achievement data, carrying out heterogeneous modeling according to labels, field types and content formats in the multi-source scientific and technological achievement data, and forming an original multi-source scientific and technological achievement data set; based on the original multi-source scientific and technological achievement data set, differential features are extracted through a big data distributed architecture, a multi-source scientific and technological achievement data change difference value is calculated, and a differential feature vector is generated; based on the differential feature vector, combining the differential feature weight and time window compression to obtain a lightweight differential feature block set; a multi-dimensional semantic feature set is extracted from the lightweight differential feature block set, a multilayer semantic relation matrix is constructed through semantic analysis and feature labeling, and a unified semantic feature space is formed. According to the method, dynamic adjustment and constraint optimization are carried out on the semantic node fusion weight, and self-adaptive distribution of the association strength between semantic nodes is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for data fusion of scientific and technological achievements based on big data. Background Technology

[0002] With the continuous growth in the quantity and sources of scientific and technological achievements, scientific and technological achievement data is gradually exhibiting characteristics of multi-source, heterogeneous, and dynamic nature. Differences exist among different research institutions, platforms, and databases in data structure, semantic expression, and time update cycles, prompting traditional achievement management to gradually evolve towards integrated and intelligent analysis based on big data. In recent years, utilizing big data technology to conduct multi-source achievement data modeling, feature extraction, and semantic association analysis has become an important means of scientific and technological intelligence mining and scientific research achievement management.

[0003] Existing multi-source scientific and technological achievement data fusion methods mostly remain at the level of static feature alignment or simple semantic aggregation, failing to fully address the dynamic changes in semantic hierarchical structure differences and contextual dependencies. Especially in the semantic fusion stage, the correlation strength and fusion weights between semantic nodes from different sources often employ fixed ratios or static normalization strategies, leading to insufficient semantic consistency, weight imbalances between semantic nodes, and semantic distortion. Therefore, there is an urgent need for a semantic fusion method that can dynamically optimize weights by combining semantic similarity, contextual relevance, and hierarchical structure differences to improve the semantic consistency and fusion accuracy of cross-source scientific and technological achievement data. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a data fusion method for scientific and technological achievements based on big data to solve the problem of inconsistent hierarchical relationships and unbalanced distribution of semantic node weights in the semantic feature fusion process of multi-source heterogeneous scientific and technological achievements data, which leads to insufficient semantic consistency expression.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a data fusion method for scientific and technological achievements based on big data. The method includes: collecting multi-source scientific and technological achievement data; performing heterogeneous modeling based on tags, field types, and content formats within the multi-source data to form an original multi-source data set; extracting differential features from the original multi-source data set using a big data distributed architecture, calculating the difference in changes between the multi-source data, and generating a differential feature vector; obtaining a lightweight differential feature block set based on the differential feature vector, combined with differential feature weights and time window compression; extracting a multi-dimensional semantic feature set from the lightweight differential feature block set, constructing a multi-layer semantic relationship matrix through semantic parsing and feature annotation to form a unified semantic feature space; performing semantic similarity and contextual relevance fusion matching based on the multi-dimensional semantic feature set in the unified semantic feature space, dynamically adjusting the semantic node fusion weights to generate a semantically consistent fusion result; and inputting the semantically consistent fusion result into a visualization analysis terminal to generate a semantic association path and differential evolution view of the multi-source data.

[0007] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method of the present invention, the specific steps for forming the original multi-source scientific and technological achievement dataset are as follows: parsing and merging the label fields of the multi-source scientific and technological achievement data, establishing a unified standard label dictionary, and performing field type identification, semantic alignment, and unified standardization of data format and units; based on the standardized multi-source scientific and technological achievement data, constructing a heterogeneous data model by combining the mapping relationship between labels and fields, and correcting and outputting the original multi-source scientific and technological achievement dataset through consistency and integrity detection.

[0008] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method of the present invention, the method for extracting differential features based on the original multi-source scientific and technological achievement dataset through a big data distributed architecture includes the following specific steps: the original multi-source scientific and technological achievement dataset is partitioned and loaded in parallel within the big data distributed architecture according to the data source and time dimension of the multi-source scientific and technological achievement; within each partition, a time series differential benchmark set is established based on the timestamp; by comparing data at adjacent time nodes, differences in additions, deletions, and modifications are identified to form differential reference pairs; based on the differential reference pairs, the magnitude and direction of content changes are calculated at the field level to extract differential features.

[0009] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method of the present invention, wherein: the calculation of the difference in changes of multi-source scientific and technological achievement data is obtained by performing differential operations on the corresponding fields of adjacent time node data in the differential reference pair and combining them with field weights for weighted fusion; the differential feature vector is generated by standardizing, weighted fusion, and vectorized encoding the difference in changes of each field in the differential reference pair.

[0010] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method of the present invention, the specific steps of obtaining a lightweight differential feature block set by combining differential feature weights and time window compression are as follows: the differential feature vectors are sorted and grouped according to timestamps and multi-source scientific and technological achievement data sources, and a label system is established by combining a standard label dictionary to generate a time window; within the time window, feature weights are obtained based on the changes in the differential feature vectors; weighted filtering and aggregation are performed on the differential feature vectors based on the feature weights, and features with repeated changes in the time series are compressed into feature fragments, and the lightweight differential feature block set is output by recombining them in chronological order.

[0011] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method of the present invention, the specific process of extracting a multidimensional semantic feature set from the lightweight differential feature block set is as follows: the lightweight differential feature block set is expanded in chronological order, the feature segments are sorted by timestamp and multi-source scientific and technological achievement data sources to form a continuous feature sequence, and the feature segments from different sources are compared and analyzed under the same time series to calculate the cross-source correlation and establish a feature association mapping table; the feature segments with correlation are aggregated according to the feature association mapping table, the aggregation results are normalized and dimensionality compressed to remove redundant features and unify the feature scale, and a multidimensional semantic feature set is extracted.

[0012] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method described in this invention, the formation of a unified semantic feature space involves the following steps: A semantic parsing engine is constructed based on a multi-dimensional semantic feature set and a tag system. The multi-dimensional semantic feature set is then input into the semantic parsing engine for word segmentation, part-of-speech tagging, and dependency analysis to extract semantic entities and semantic relationships between nodes. Semantic entities are matched with the tag system, and each semantic node is assigned a semantic category and node weight to form a semantic annotation result. Based on the semantic annotation result, the semantic nodes are hierarchically divided according to semantic categories, node weights, and semantic relationships between nodes, resulting in a hierarchical division consisting of a theoretical method layer, an application layer, and an achievement layer. A multi-layer semantic relationship matrix is ​​constructed based on the hierarchical division result. The multi-dimensional semantic feature set is then fused and normalized through semantic parsing, feature annotation, and the multi-layer semantic relationship matrix to unify the semantic expressions at different levels, forming a unified semantic feature space.

[0013] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method of the present invention, the specific process of performing semantic similarity and context relevance fusion matching is as follows: semantic nodes in a unified semantic feature space are transformed into semantic vectors; embedding calculations are performed between nodes to obtain semantic similarity; a context window is constructed for each semantic node based on semantic similarity, semantic neighborhoods are extracted, and context relevance is calculated; semantic similarity and context relevance are fused and matched to obtain fusion matching results; after normalizing the fusion matching results, a semantic consistency matrix is ​​generated based on the normalized results, and then fused and matched using semantic similarity and context relevance.

[0014] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method of the present invention, the specific process of dynamically adjusting the semantic node fusion weights to generate semantically consistent fusion results is as follows: The matching values ​​of each semantic node are extracted from the semantic consistency matrix and associated with semantic vectors in a unified semantic feature space; the semantic node fusion weights are initialized based on the association mapping; based on the initial semantic node fusion weights, the semantic offset relationship between nodes is calculated according to the matching difference between semantic similarity and contextual relevance, and hierarchical association parameters are generated by combining the hierarchical division results in the multi-layer semantic relationship matrix; the semantic offset relationship and hierarchical association parameters are used as constraints to dynamically adjust the semantic node fusion weights; normalization and constraint optimization processing are performed on the dynamically adjusted semantic node fusion weight distribution to generate semantically consistent fusion results.

[0015] As a preferred embodiment of the big data-based scientific and technological achievement data fusion method of the present invention, wherein: the semantic association path refers to a directed semantic connection sequence constructed based on the semantic similarity, contextual relationship and fusion weight changes between semantic nodes in the semantic consistency fusion result; the differential evolution view refers to an evolution graph formed by visually mapping the differential changes of semantic attributes, fusion weights and mutual relationships of semantic nodes with time and version iteration based on the semantic association path.

[0016] The beneficial effects of this invention are as follows: By dynamically adjusting and constraining the fusion weights of semantic nodes, an adaptive distribution of the association strength between semantic nodes is achieved. Using the numerical difference between semantic similarity and contextual relevance as the basis for weight correction, the fusion weights of semantic nodes are continuously updated, enhancing the weights of highly associated nodes and suppressing the weights of low-associated nodes, thereby maintaining the overall stability of the semantic relationship network. After normalization and relevance constraint processing, the semantic node weights form a balanced distribution across levels, accurately reflecting the true dependencies between different semantic levels. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a data fusion method for scientific and technological achievements based on big data.

[0019] Figure 2 This is a flowchart for differential feature extraction.

[0020] Figure 3 A flowchart for constructing a multi-level semantic relation matrix.

[0021] Figure 4 This is a flowchart for semantic consistency fusion. Detailed Implementation

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides a data fusion method for scientific and technological achievements based on big data, including the following steps: S1: Collect multi-source scientific and technological achievement data, and perform heterogeneous modeling based on the labels, field types and content formats within the multi-source scientific and technological achievement data to form the original multi-source scientific and technological achievement dataset.

[0026] S1.1: Parse and merge the tag fields of multi-source scientific and technological achievement data, establish a unified standard tag dictionary, and perform field type identification, semantic alignment, and unified standardization of data format and units.

[0027] Specifically, the label fields in multi-source scientific and technological achievement data are extracted and semantically merged. Labels from different sources but with the same or similar meanings are unified into standardized expressions, and a unified standard label dictionary is established. Based on the standard label dictionary, the fields of multi-source scientific and technological achievement data are identified and semantically aligned. The data format and units of each field are standardized, and the units of measurement for numerical fields are unified, and the encoding format and time representation method for text fields are unified.

[0028] S1.2: Based on the standardized multi-source scientific and technological achievements data, a heterogeneous data model is constructed by combining the mapping relationship between labels and fields, and the original multi-source scientific and technological achievements dataset is corrected and output through consistency and integrity detection.

[0029] Specifically, the standardized multi-source scientific and technological achievements data are used as input. Based on the mapping relationship between labels and fields, multi-source scientific and technological achievements data from different sources and with different structures are matched and corresponded according to field semantics to construct a heterogeneous data model. Consistency checks are performed on the constructed heterogeneous data model, and integrity corrections are made, supplements are added, and abnormal data is adjusted. After the heterogeneous data model is corrected, the corrected multi-source scientific and technological achievements data are exported and summarized according to the mapping relationship to form the original multi-source scientific and technological achievements dataset.

[0030] It should be noted that the heterogeneous data model is constructed through four steps: field semantic recognition, mapping group generation, field index establishment, and structural reorganization of multi-source scientific and technological achievement data.

[0031] S2: Based on the original multi-source scientific and technological achievement dataset, differential features are extracted through a big data distributed architecture, the difference in changes of multi-source scientific and technological achievement data is calculated, and differential feature vectors are generated.

[0032] S2.1: The original multi-source scientific and technological achievement dataset is partitioned and loaded in parallel within the big data distributed architecture according to the data source and time dimension of the multi-source scientific and technological achievement dataset.

[0033] Specifically, the original multi-source scientific and technological achievement dataset is imported into the big data distributed architecture environment, and partitioning rules are set according to the source and time dimension of the multi-source scientific and technological achievement data. The multi-source scientific and technological achievement data is divided according to the partitioning rules, and the multi-source scientific and technological achievement data from different sources is allocated to the corresponding node storage area. Data blocks are organized in chronological order within each node. Then, the multi-source scientific and technological achievement data of each partition is read in parallel using a distributed loading mechanism.

[0034] S2.2: Within each partition, timestamps are used as the basis to establish a time series difference benchmark set. By comparing data from adjacent time nodes, differences in additions, deletions, and modifications are identified to form difference reference pairs.

[0035] Specifically, within each multi-source scientific and technological achievement data partition, timestamps are used as the basis to sort multi-source scientific and technological achievement data from the same source in chronological order, constructing a time series data chain; adjacent time nodes are used as comparisons to compare the content of multi-source scientific and technological achievement data at different time nodes field by field, and differences in multi-source scientific and technological achievement data are identified by detecting the addition, absence, and change of field values; differences in each pair of adjacent time nodes are recorded in pairs to generate differential reference pairs.

[0036] S2.3: Based on the differential reference pair, calculate the magnitude and direction of content changes at the field level and extract differential features.

[0037] Specifically, the differential reference pair is used as input. The numerical values ​​and text content of the corresponding fields in the multi-source scientific and technological achievement data are compared item by item. The degree of change of the value of the same field in the multi-source scientific and technological achievement data between adjacent time nodes is calculated. The change is represented by the numerical difference. The growth and decrease trend types are distinguished according to the direction of change. The degree of change of each field is integrated to form the field-level differential result. The differential result is standardized and vectorized to extract the differential features that can be expressed quantitatively.

[0038] S2.4: The difference in changes of multi-source scientific and technological achievement data is calculated by performing differential operations on the corresponding fields of adjacent time nodes in the differential reference pair and then combining the field weights for weighted fusion.

[0039] Specifically, the differential reference pair is used as input to extract the corresponding field values ​​of multi-source scientific and technological achievements data at adjacent time points; differential operations are performed on each corresponding field to calculate the difference in changes of the field between the two time points.

[0040] It should be noted that the process involves performing differential operations on the fields of multi-source scientific and technological achievements data at adjacent time nodes in the differential reference pair, and then normalizing the data after weighted fusion based on the importance of the fields. According to the importance of the fields in the overall multi-source scientific and technological achievements data structure, weight coefficients are assigned to different fields, and the differential results of each field are weighted and fused. The fused weighted difference is then normalized to obtain the change difference value of the multi-source scientific and technological achievements data.

[0041] S2.5: The differential feature vector is generated by standardizing, weighting and fusion, and vectorizing the difference in changes of each field in the differential reference pair.

[0042] Specifically, based on the differential reference pair, the difference in changes of each field at adjacent time points is read; the difference in changes of all fields is standardized; the field weight coefficients generated by comprehensively calculating the historical change frequency, the number of times the field appears, the correlation between the field and the core indicators, and information entropy of each field in the multi-source scientific and technological achievement data are used to perform weighted fusion on the standardized differences to highlight the change characteristics of the fields; the weighted fusion result is vectorized and encoded to express the difference characteristics of each field in a fixed-dimensional numerical form, generating a difference feature vector.

[0043] S3: Based on the differential feature vector, a lightweight differential feature block set is obtained by combining differential feature weights and time window compression.

[0044] S3.1: Sort and group the differential feature vectors according to timestamps and the sources of multi-source scientific and technological achievements data, and establish a label system in conjunction with a standard label dictionary to generate a time window; within the time window, obtain feature weights based on the changes in the differential feature vectors.

[0045] Specifically, the differential feature vectors are sorted and grouped according to timestamps and the sources of multi-source scientific and technological achievements data; a corresponding label system is established for the differential feature vectors by combining the field labels and category information in the standard label dictionary, and time windows are generated based on temporal continuity; within each time window, the numerical fluctuation of the differential feature vectors is statistically analyzed, and the activity level in the time dimension is reflected by calculating the magnitude and frequency of changes.

[0046] It should be noted that by sorting and grouping the differential feature vectors according to timestamps and the sources of multi-source scientific and technological achievements data, and establishing a label system, the magnitude of numerical changes and frequency of occurrence are statistically analyzed within the time window to obtain feature weights.

[0047] S3.2: Perform weighted filtering and aggregation on the differential feature vectors based on feature weights, compress features with repeated changes in the time series into feature fragments, and reassemble them in chronological order to output a lightweight differential feature block set.

[0048] Specifically, the differential feature vectors are weighted and sorted according to feature weights; the changing trend of feature values ​​is analyzed in the time series dimension; differential features with small numerical fluctuations and repeated occurrences in different time windows are compressed and integrated; differential features with similar continuous changes are aggregated into feature segments; feature segments are reorganized in time order to output a lightweight differential feature block set.

[0049] S4: Extract a multi-dimensional semantic feature set from the lightweight differential feature block set, construct a multi-layer semantic relationship matrix through semantic parsing and feature annotation, and form a unified semantic feature space.

[0050] S4.1: Expand the lightweight differential feature block set in chronological order, sort the feature segments by timestamp and multi-source scientific and technological achievement data sources to form a continuous feature sequence, and compare and analyze feature segments from different sources under the same time series to calculate cross-source correlation and establish a feature association mapping table.

[0051] Specifically, based on the timestamp information in each lightweight differential feature block set, all feature segments are sorted by time and classified by source. After sorting, the timestamp is used as the main index to arrange the feature segments from different multi-source scientific and technological achievement data sources in a continuous time order, forming a continuous feature sequence covering the entire time interval.

[0052] It should be noted that in continuous feature sequences, based on a unified time node, feature fragments from different sources at the same time node are compared field by field. By calculating the difference, similarity and rate of change between each field, the consistency and differences between features from different sources are identified. The cross-source correlation values ​​of each field in the comparison are recorded together with the corresponding timestamp, the source of multi-source scientific and technological achievements data and the feature fragment identifier, and summarized to form a feature association mapping table.

[0053] Let the set of field weight coefficients be defined by the formula: ; In the formula, Represents the field weight vector. Representation field Weighting coefficients Initial values ​​were obtained by performing statistical analysis of information gain, variance, and frequency of occurrence on historical samples of multi-source scientific and technological achievements data. After normalization, the values ​​ranged as follows: And satisfy The example values ​​range from 0.01 to 0.15. This represents the total number of feature dimensions; Timestamp Data source below With data source The cross-source relevance is calculated using the following formula: ; In the formula, Represents timestamp Multiple sources of scientific and technological achievements data and Cross-source relevance, An identifier indicating the first source of multi-source scientific and technological achievement data. An identifier indicating the second source of multi-source scientific and technological achievement data. Indicates the index of the field number. Represents timestamp Down, Corresponding fields The difference characteristic values, Represents timestamp Down, Corresponding fields The difference characteristic values, Indicates from field number To field number The accumulation operator.

[0054] S4.2: Aggregate the feature fragments of relevance based on the feature association mapping table, normalize and compress the aggregation results to remove redundant features and unify the feature scale, and extract a multi-dimensional semantic feature set.

[0055] Specifically, the feature fragments in the feature association mapping table are matched and grouped. The values ​​of each feature fragment in the same field position are superimposed one by one and the average value is calculated to generate an aggregation result. The aggregation result is normalized to unify the values ​​of different dimensions to the same value range, which is between 0 and 1.

[0056] After normalization, the aggregation results are subjected to dimensionality compression. The number of redundant features is reduced by calculating feature correlation. By calculating the correlation coefficient between feature segments and comparing the magnitude of the correlation coefficient, highly correlated feature segments are identified and merged to reduce the number of redundant features and compress feature dimensions. The numerical scale of the feature segments is also uniformly processed to extract a multidimensional semantic feature set.

[0057] S4.3: Construct a semantic parsing engine based on the multidimensional semantic feature set and tag system, and input the multidimensional semantic feature set into the semantic parsing engine for word segmentation, part-of-speech recognition and dependency analysis to extract semantic entities and semantic relationships between nodes.

[0058] Specifically, the multidimensional semantic feature set is structured according to source type and field attributes; the multidimensional semantic feature set is then sequentially input into the semantic parsing engine for word segmentation, which identifies basic semantic segments by dividing the text content into words; part-of-speech tagging is performed on each semantic segment to determine its grammatical attributes; after part-of-speech tagging, dependency analysis is performed on the multidimensional semantic feature set to identify dependency relationships between semantic segments by analyzing syntactic structure; and through joint analysis of word segmentation results, part-of-speech tagging results, and dependency relationships, semantic associations between semantic entities and nodes are extracted.

[0059] S4.4: Match semantic entities with the label system, assign semantic categories and node weights to each semantic node, and form semantic annotation results.

[0060] Specifically, semantic content recognition is performed on semantic entities by comparing their text features, contextual positions, and semantic attributes with the field labels and category labels defined in the label system.

[0061] It should be noted that during the matching process, by comparing the semantic similarity values ​​between the semantic entity and each tag in the tag system, the tag with the closest semantic expression is selected as the corresponding tag of the semantic entity; based on the completed matching, a semantic category is assigned to each semantic node, and the semantic category level to which the semantic node belongs is determined according to the classification structure in the tag system; combining the frequency of occurrence and semantic importance of the semantic entity in the multidimensional semantic feature set, a node weight is assigned to each semantic node; based on the completed semantic category and node weight, the correspondence between semantic nodes, semantic categories and node weights is recorded as a structured result to form the semantic annotation result.

[0062] S4.5: Based on the semantic annotation results, the semantic nodes are hierarchically divided according to semantic categories, node weights and semantic relationships between nodes, resulting in a hierarchical division consisting of a theoretical method layer, an application layer and a result layer.

[0063] Specifically, based on the category attributes and node weight values ​​of semantic nodes in the semantic annotation results, hierarchical division is carried out, and semantic nodes are divided into three initial sets according to their semantic categories: technology category, application category, and achievement category. After the classification is completed, the node weight recorded in the semantic annotation results is used as the priority basis for hierarchical division. Semantic nodes with higher node weights are classified into higher levels, reflecting their core role in the multidimensional semantic feature set.

[0064] By combining the semantic relationships between nodes, the subordinate relationships between different semantic nodes are analyzed. By identifying the dependency direction and semantic connection strength between semantic nodes, the hierarchical position of each semantic node in the semantic structure is clarified. Semantic nodes of the technology category are assigned to the technology layer, semantic nodes of the application category are assigned to the application layer, and semantic nodes of the result category are assigned to the result layer, forming a hierarchical division result consisting of the theoretical method layer, the application layer, and the result layer.

[0065] S4.6: Construct a multi-level semantic relationship matrix based on the hierarchical division results. The multi-dimensional semantic feature set is then fused and normalized through semantic parsing, feature annotation, and the multi-level semantic relationship matrix to unify the semantic expression of different levels and form a unified semantic feature space.

[0066] Specifically, based on the hierarchical division results, semantic nodes of the technical layer, semantic nodes of the application layer, and semantic nodes of the result layer are used as row and column indices to construct semantic relationship sub-matrices of the technical layer, application layer, and result layer, respectively. According to the correspondence between semantic nodes in different layers, the multi-layer semantic relationship sub-matrices are aligned in rows and columns. When semantic nodes are repeated or cross-layer connections exist, the corresponding semantic relationship values ​​are weighted and averaged to obtain the fused multi-layer semantic relationship matrix.

[0067] Normalization is performed on the numerical values ​​of the multi-level semantic relation matrix along the row and column directions respectively, mapping the semantic relation values ​​to a unified numerical range, with values ​​between 0 and 1, so that the semantic expressions of different levels can be compared on a unified scale; using the index position of the semantic node in the matrix and the semantic relation value as the semantic feature dimensions, the semantic relations of the technical layer, the semantic relations of the application layer, and the semantic relations of the result layer are jointly mapped to a unified semantic representation space, forming a unified semantic feature space.

[0068] A superior approach, compared to using static matrix fusion and simple global normalization methods to process semantic relation matrices, constructs multi-layer semantic relation sub-matrices based on the semantic node hierarchy partitioning results. It utilizes the contextual dependencies between semantic nodes and semantic intensity differences for inter-layer alignment, and maintains the hierarchical balance of semantic information distribution through summation and weighted average fusion. Simultaneously, it combines row and column bidirectional normalization to eliminate numerical biases between levels, achieving a unified dimensional mapping of semantic relation strength. This enables the multi-layer semantic relation matrix to accurately represent the nonlinear interactions and semantic dependencies between the technical layer, application layer, and result layer.

[0069] S5: In a unified semantic feature space, semantic similarity and contextual relevance are fused and matched based on a multi-dimensional semantic feature set, and the semantic node fusion weights are dynamically adjusted to generate semantically consistent fusion results.

[0070] S5.1: Transform semantic nodes in the unified semantic feature space into semantic vectors, perform embedding calculations between nodes, and obtain semantic similarity.

[0071] Specifically, semantic relationship values ​​are extracted based on the position index of semantic nodes in the multi-layer semantic relationship matrix, and these semantic relationship values ​​are used as the feature dimension input of the semantic vector; embedding calculation is performed on the semantic vector of each semantic node.

[0072] It should be noted that the similarity between semantic nodes is obtained by calculating the cosine similarity between each semantic vector, which measures the proximity of semantic vectors in a unified semantic feature space. During the embedding calculation, the semantic vectors are mapped in a high dimension using the semantic coordinate relationship in the unified semantic feature space, so that the distance distribution between semantic nodes can reflect the semantic proximity. The semantic similarity between semantic nodes is obtained by comparing the similarity values ​​between different semantic nodes.

[0073] S5.2: Based on semantic similarity, construct a context window for each semantic node, extract semantic neighborhood, and calculate context relevance.

[0074] Specifically, based on the similarity ranking results between semantic nodes in the unified semantic feature space, several semantic nodes with the highest similarity value to the target semantic node are selected as context neighbor nodes; with the target semantic node as the center, the window boundary is determined in the similarity distribution according to the time order or semantic position relationship, thereby forming a context window containing the target semantic node and semantic neighborhood.

[0075] It should be noted that within the context window, the semantic relationship values ​​between the target semantic node and each neighboring semantic node are paired and compared. The semantic association strength between semantic nodes is measured by calculating the correlation coefficient or cosine similarity between the semantic feature vectors of adjacent semantic nodes. Finally, the correlation coefficients of all paired nodes within the context window are weighted and averaged to obtain the context relevance of the target semantic node.

[0076] S5.3: Combine semantic similarity with contextual relevance to obtain the combined matching result.

[0077] Specifically, in the unified semantic feature space, the semantic similarity value and context relevance value of each semantic node are paired, and the two types of values ​​of the same semantic node are matched one-to-one according to the corresponding index of the semantic node; based on the hierarchical weight of the semantic node in the multi-level semantic relationship matrix, the semantic similarity value and context relevance value are weighted and fused, according to the proportion of the influence of global semantic association and local context relationship.

[0078] Based on the weighted fusion after completion, the fusion result is normalized to adjust all fusion matching values ​​to a uniform numerical scale range, with values ​​between 0 and 1. By comparing the magnitude of the fusion matching values, the comprehensive semantic association between semantic nodes is obtained, forming the fusion matching result.

[0079] S5.4: After normalizing the fusion matching results, a semantic consistency matrix is ​​generated based on the normalized results, and the matching is fused by semantic similarity and context relevance.

[0080] Specifically, the normalized fusion matching results are matrixed and generated using semantic nodes as row and column indices. The fusion matching values ​​are then filled into the matrix to represent the semantic consistency strength between each semantic node.

[0081] Based on the constructed semantic consistency matrix, symmetry processing is performed on the row and column directions of the semantic consistency matrix to maintain the balance of the matching relationship between semantic nodes. The semantic consistency matrix is ​​used as the basis for the fusion of semantic similarity and context relevance to calculate the consistency score between nodes and complete the fusion matching of semantic similarity and context relevance.

[0082] S5.5: Extract the matching value of each semantic node from the semantic consistency matrix, and perform association mapping with the semantic vector in the unified semantic feature space. Initialize the semantic node fusion weight based on the association mapping.

[0083] Specifically, the row and column indices of semantic nodes in the semantic consistency matrix are used as the retrieval basis, and the matching values ​​corresponding to semantic nodes are read one by one; the matching values ​​of semantic nodes extracted from the semantic consistency matrix are compared with the semantic vectors in the unified semantic feature space, and the association mapping between semantic nodes and semantic vectors is obtained by matching the index positions of semantic nodes in the semantic feature space; based on the association mapping, the initial semantic node fusion weights are generated according to the similarity relationship between the semantic node matching values ​​and semantic vectors.

[0084] S5.6: Based on the initial semantic node fusion weights, calculate the semantic offset relationship between nodes according to the matching difference of semantic similarity and context relevance, and generate hierarchical association parameters by combining the hierarchical division results in the multi-layer semantic relationship matrix.

[0085] Specifically, the matching values ​​of each semantic node in the semantic consistency matrix are used as a reference. The difference between semantic similarity and contextual relevance is calculated for each pair of semantic nodes. The difference between semantic similarity and contextual relevance is used to obtain the semantic offset difference. The number ratio and connection density of semantic node levels in the multi-level semantic relationship matrix are combined with hierarchical statistics and structural analysis to calculate the difference between the number ratio and connection density of different levels and generate hierarchical association parameters.

[0086] S5.7: Use semantic offset relationship and hierarchical association parameters as constraints to dynamically adjust the semantic node fusion weights.

[0087] Specifically, the offset direction of semantic nodes in the unified semantic feature space is determined based on the change magnitude of semantic offset relationship, the weight update ratio of semantic nodes at different levels is controlled according to the hierarchical strength of hierarchical association parameters, and the fusion weight of semantic nodes is dynamically adjusted through iterative update.

[0088] S5.8: Perform normalization and constraint optimization on the dynamically adjusted semantic node fusion weight distribution to generate semantically consistent fusion results.

[0089] Specifically, the fusion weights of all semantic nodes are numerically normalized to unify the weight range of different semantic nodes to the same numerical range, with values ​​between 0 and 1, thereby eliminating the impact of differences in weight distribution among semantic nodes; and constraint optimization processing is performed on the fusion weight distribution of semantic nodes.

[0090] It should be noted that by calculating the correlation coefficient and deviation measure between the fusion weights, while maintaining the overall semantic consistency structure, the weight differences of different semantic nodes are balanced. The optimized semantic node fusion weight distribution is used as the basis to reconstruct the semantic connection relationship between semantic nodes and generate a semantically consistent fusion result.

[0091] In contrast, existing semantic fusion methods typically employ static normalization and fixed-ratio adjustments during the weight distribution optimization stage. The method described above, however, not only normalizes the semantic node fusion weights numerically but also dynamically optimizes the constraints using correlation coefficients and deviation metrics, achieving adaptive balance adjustment of the coupling relationships between semantic nodes. By dynamically adjusting the fusion weight distribution while maintaining overall semantic consistency and structural stability, the weight allocation between semantic nodes becomes more reasonable, and the association expression more accurate, thereby improving the structural balance and the ability to accurately reflect semantic associations in the semantic consistency fusion results. Compared to traditional methods, this invention can achieve self-adjusting optimization of semantic weights under multi-layer semantic association conditions, significantly enhancing the fine-grained expressiveness and overall consistency of semantic feature fusion.

[0092] S6: Input the semantic consistency fusion results into the visualization analysis terminal to generate semantic association paths and differential evolution views of multi-source scientific and technological achievement data.

[0093] S6.1: Semantic association path refers to the directed semantic connection sequence constructed based on the semantic similarity, contextual relationship and fusion weight changes between semantic nodes in the semantic consistency fusion result.

[0094] Specifically, the semantic similarity values, contextual relationships, and semantic node fusion weight changes among all semantic nodes are extracted; semantic similarity is used as the starting point for path establishment, and the connection direction between semantic nodes is obtained according to the contextual relationships; the magnitude and direction of the semantic node fusion weight changes are used as path weight parameters to identify the strength and flow of semantic relationships between semantic nodes; the directed connections between semantic nodes are arranged sequentially to form a directed semantic connection sequence composed of semantic nodes and weight relationships, thus generating a semantic association path.

[0095] S6.2: The differential evolution view refers to the evolution graph formed by visually mapping the differential changes of semantic attributes, fusion weights and interrelationships of semantic nodes with time and version iterations, based on the semantic association path.

[0096] Specifically, based on semantic association paths, we extract the semantic attributes, fusion weights, and association values ​​between semantic nodes at different times or versions; we perform differential calculations on the semantic attributes, fusion weights, and association values ​​of semantic nodes at adjacent time points and versions to extract and quantify the changing characteristics of semantic nodes and their connections.

[0097] The continuous feature sequence formed by sorting semantic nodes by timestamp and data source is mapped to the corresponding difference value. The changes in semantic attributes, fusion weights and relationships are displayed through visualization methods such as color gradient, node size and line thickness. All semantic nodes and their changes over time are arranged in chronological order to generate an evolution graph.

[0098] In summary, this invention achieves an adaptive distribution of the association strength between semantic nodes by dynamically adjusting and constraining the fusion weights of semantic nodes. Using the numerical difference between semantic similarity and contextual relevance as the basis for weight correction, the fusion weights of semantic nodes are continuously updated, enhancing the weights of highly associated nodes and suppressing the weights of low-associated nodes, thereby maintaining the overall stability of the semantic relationship network. After normalization and relevance constraint processing, the semantic node weights form a balanced distribution across levels, accurately reflecting the true dependencies between different semantic levels.

[0099] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A data fusion method for scientific and technological achievements based on big data, characterized in that: include, Collect multi-source scientific and technological achievements data, and perform heterogeneous modeling based on the tags, field types and content formats within the multi-source scientific and technological achievements data to form the original multi-source scientific and technological achievements dataset. Based on the original multi-source scientific and technological achievement dataset, differential features are extracted through a big data distributed architecture, the difference in changes of multi-source scientific and technological achievement data is calculated, and a differential feature vector is generated. A lightweight differential feature block set is obtained by combining differential feature vectors with differential feature weights and time window compression. A multi-dimensional semantic feature set is extracted from the lightweight differential feature block set, and a multi-layer semantic relationship matrix is ​​constructed through semantic parsing and feature annotation to form a unified semantic feature space. In a unified semantic feature space, semantic similarity and context relevance are fused and matched based on a multi-dimensional semantic feature set, and the fusion weights of semantic nodes are dynamically adjusted to generate semantically consistent fusion results. Input the semantic consistency fusion results into the visualization analysis terminal to generate semantic association paths and differential evolution views of multi-source scientific and technological achievement data.

2. The data fusion method for scientific and technological achievements based on big data as described in claim 1, characterized in that: The specific steps for forming the original multi-source scientific and technological achievement dataset are as follows. The label fields of multi-source scientific and technological achievement data are parsed and merged to establish a unified standard label dictionary, and field type identification, semantic alignment and data format and unit standardization are performed. Based on the standardized multi-source scientific and technological achievements data, a heterogeneous data model is constructed by combining the mapping relationship between labels and fields, and the original multi-source scientific and technological achievements dataset is corrected and output through consistency and integrity detection.

3. The data fusion method for scientific and technological achievements based on big data as described in claim 2, characterized in that: The extraction of differential features based on the original multi-source scientific and technological achievement dataset using a big data distributed architecture involves the following specific steps. The original multi-source scientific and technological achievement datasets are partitioned and loaded in parallel within a big data distributed architecture according to the data source and time dimension of the multi-source scientific and technological achievement datasets. Within each partition, timestamps are used as the basis to establish a time series difference benchmark set. By comparing data from adjacent time nodes, differences in additions, deletions, and modifications are identified to form difference reference pairs. Based on differential reference pairs, the magnitude and direction of content changes are calculated at the field level, and differential features are extracted.

4. The data fusion method for scientific and technological achievements based on big data as described in claim 3, characterized in that: The calculation of the difference in changes in multi-source scientific and technological achievement data is obtained by performing differential operations on the corresponding fields of adjacent time node data in the differential reference pair and combining them with field weights for weighted fusion. The differential feature vector is generated by standardizing, weighting and fusion, and vectorizing the difference in changes of each field in the differential reference pair.

5. The data fusion method for scientific and technological achievements based on big data as described in claim 4, characterized in that: The specific steps for obtaining a lightweight differential feature block set by combining differential feature weights and time window compression are as follows: The differential feature vectors are sorted and grouped according to timestamps and the sources of multi-source scientific and technological achievements data, and a label system is established in conjunction with a standard label dictionary to generate time windows; Within the time window, feature weights are obtained based on changes in the difference feature vector; Based on feature weights, the differential feature vectors are weighted, filtered and aggregated, and features with repeated changes in the time series are compressed into feature fragments and reassembled in chronological order to output a lightweight differential feature block set.

6. The data fusion method for scientific and technological achievements based on big data as described in claim 5, characterized in that: The specific process for extracting a multidimensional semantic feature set from the lightweight differential feature block set is as follows: The lightweight differential feature block set is expanded in chronological order, and the feature segments are sorted by timestamp and multi-source scientific and technological achievement data sources to form a continuous feature sequence. Under the same time series, feature segments from different sources are compared and analyzed to calculate cross-source correlation and establish a feature association mapping table. Based on the feature association mapping table, feature fragments with relevance are aggregated. The aggregation results are normalized and dimensionality compressed to remove redundant features and unify feature scale, and a multidimensional semantic feature set is extracted.

7. The data fusion method for scientific and technological achievements based on big data as described in claim 6, characterized in that: The specific process for forming a unified semantic feature space is as follows. A semantic parsing engine is constructed based on a multidimensional semantic feature set and a tag system. The multidimensional semantic feature set is then input into the semantic parsing engine for word segmentation, part-of-speech recognition, and dependency analysis to extract semantic relationships between semantic entities and nodes. The semantic entities are matched with the tag system, and each semantic node is assigned a semantic category and node weight to form a semantic annotation result; Based on the semantic annotation results, the semantic nodes are hierarchically divided according to semantic category, node weight and semantic relationship between nodes, resulting in a hierarchical division result consisting of theoretical method layer, application layer and result layer; Based on the hierarchical division results, a multi-level semantic relationship matrix is ​​constructed. The multi-dimensional semantic feature set is then fused and normalized through semantic parsing, feature annotation, and the multi-level semantic relationship matrix to unify the semantic expression of different levels and form a unified semantic feature space.

8. The data fusion method for scientific and technological achievements based on big data as described in claim 7, characterized in that: The specific process of performing semantic similarity and contextual relevance fusion matching is as follows. The semantic nodes in the unified semantic feature space are transformed into semantic vectors, and embedding calculations are performed between the nodes to obtain semantic similarity. Based on semantic similarity, a context window is constructed for each semantic node to extract semantic neighborhood and calculate context relevance. The semantic similarity and contextual relevance are fused and matched to obtain the fused matching result; After normalizing the fusion matching results, a semantic consistency matrix is ​​generated based on the normalized results, and the matching is fused by semantic similarity and context relevance.

9. The data fusion method for scientific and technological achievements based on big data as described in claim 8, characterized in that: The specific process for dynamically adjusting the semantic node fusion weights to generate a semantically consistent fusion result is as follows. Extract the matching value of each semantic node from the semantic consistency matrix and associate it with the semantic vector in the unified semantic feature space. Initialize the semantic node fusion weight based on the association mapping. Based on the initial semantic node fusion weights, the semantic offset relationship between nodes is calculated according to the matching difference of semantic similarity and context relevance, and the hierarchical association parameters are generated by combining the hierarchical partitioning results in the multi-layer semantic relationship matrix. Using semantic offset relationships and hierarchical association parameters as constraints, the semantic node fusion weights are dynamically adjusted. Normalization and constraint optimization are performed on the dynamically adjusted semantic node fusion weight distribution to generate semantically consistent fusion results.

10. The data fusion method for scientific and technological achievements based on big data as described in claim 9, characterized in that: The semantic association path refers to the directed semantic connection sequence constructed based on the semantic similarity, contextual relationship and fusion weight changes between semantic nodes in the semantic consistency fusion result; The differential evolution view refers to the evolution graph formed by visually mapping the differential changes of semantic attributes, fusion weights, and interrelationships of semantic nodes over time and version iterations, based on the semantic association path.