Methods and Systems for Heterogeneous Resource Virtualization Labeling and Knowledge Base Creation

By extracting cross-modal semantic features and constructing a semantic association network, the problem of fixing deep semantic associations between heterogeneous resources is solved, and adaptive updating and efficient management of heterogeneous resource knowledge bases are achieved.

CN120892945BActive Publication Date: 2026-01-06GUIYANG RUIBAO TECHNOLOGY SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511369624.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-06
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture deep semantic relationships between resources of different modalities, and these relationships are fixed and cannot adapt to resource evolution, resulting in low annotation efficiency, poor data consistency, and slow knowledge base updates.

Method used

By extracting cross-modal semantic features, a set of cross-modal resource semantic features is generated. A semantic association network containing semantic nodes, associated edges, and evolutionary weights is constructed. Cross-modal virtualization mapping rules are mined to generate a set of virtualized labeled data. Finally, an evolutionary heterogeneous resource knowledge base is constructed.

Benefits of technology

It enables adaptive adjustment of semantic relationships between heterogeneous resources, improves the accuracy and timeliness of labeled data, ensures that the knowledge base reflects the latest status and semantic changes of resources, and improves the level of knowledge updating and application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892945B_ABST
    Figure CN120892945B_ABST
Patent Text Reader

Abstract

The application provides a heterogeneous resource virtualization labeling and knowledge base creation method and system, which comprises the following steps: obtaining a plurality of types of heterogeneous scientific project resource sets, performing cross-modal semantic feature extraction operation, and generating a cross-modal resource semantic feature set; constructing a semantic association network based on the cross-modal resource semantic feature set, performing cross-modal virtualization mapping rule mining operation through the semantic association network, generating a cross-modal virtualization mapping rule set, and performing virtualization labeling processing on the heterogeneous scientific project resource set through the cross-modal virtualization mapping rule set to generate a virtualization labeling data set; and constructing an evolutionary heterogeneous resource knowledge base based on the virtualization labeling data set, wherein the evolutionary heterogeneous resource knowledge base comprises a concept unit, a relationship unit and an attribute unit. The method provided by the application can effectively improve the organization, management and application level of heterogeneous resource knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more specifically, to a method and system for virtualizing and labeling heterogeneous resources and creating a knowledge base. Background Technology

[0002] With the development of science and technology project management techniques, the effective integration and knowledge management of heterogeneous resources has become an important research direction. Heterogeneous resources refer to various science and technology project resources that differ in type, structure, format, or access method. Virtualizing and annotating these resources and creating knowledge bases are key technologies for achieving efficient resource utilization. Currently, the common approach is to extract features and convert formats for different types of resources separately, then annotate the resources based on preset rules, and store the annotated data in a structured database to build a knowledge base. However, this approach struggles to capture the deep semantic relationships between different modalities of resources, and once established, these relationships remain fixed, failing to adapt to the evolving nature of resources. Furthermore, the preset rules lack flexibility and struggle to cope with diverse resource formats and access interface differences, resulting in low annotation efficiency and poor data consistency. Consequently, the constructed knowledge base fails to effectively integrate the semantic information of heterogeneous resources, exhibits slow knowledge updates, and fails to meet the needs of science and technology projects for resource knowledge application. Summary of the Invention

[0003] This invention provides a method and system for heterogeneous resource virtualization annotation and knowledge base creation.

[0004] In a first aspect, embodiments of the present invention provide a method for heterogeneous resource virtualization annotation and knowledge base creation. The method includes: acquiring a set of heterogeneous scientific and technological project resources of various types; performing cross-modal semantic feature extraction on the heterogeneous scientific and technological project resource sets to generate a cross-modal resource semantic feature set; constructing a semantic association network based on the cross-modal resource semantic feature set, wherein the semantic association network includes semantic nodes, association edges, and evolutionary weights; performing cross-modal virtualization mapping rule mining on the semantic association network to generate a cross-modal virtualization mapping rule set, wherein the cross-modal virtualization mapping rule set includes text mapping rules, structure mapping rules, and interface mapping rules; performing virtualization annotation processing on the heterogeneous scientific and technological project resource sets using the cross-modal virtualization mapping rule set to generate a virtualized annotation data set; and constructing an evolutionary heterogeneous resource knowledge base based on the virtualized annotation data set, wherein the evolutionary heterogeneous resource knowledge base includes concept units, relation units, and attribute units.

[0005] Secondly, embodiments of the present invention provide a computer system, including: a memory storing a computer program; and a processor for loading the computer program to implement the heterogeneous resource virtualization annotation and knowledge base creation method as described above.

[0006] The heterogeneous resource virtualization annotation and knowledge base creation method provided by this invention extracts cross-modal semantic features from a set of heterogeneous technology project resources to generate a set of cross-modal resource semantic features. Based on this set of cross-modal resource semantic features, a semantic association network is constructed. Then, cross-modal virtualization mapping rule mining is performed through the semantic association network to generate a set of cross-modal virtualization mapping rules. This set of cross-modal virtualization mapping rules is then used to perform virtualization annotation processing on the heterogeneous technology project resource set to generate a set of virtualized annotation data. Finally, an evolutionary heterogeneous resource knowledge base is constructed based on the set of virtualized annotation data. This method can comprehensively capture the deep semantic information of heterogeneous resources of different types, structures, formats, and access methods at the text, structure, and interface levels. It provides multi-dimensional and high-value semantic input for the construction of the semantic association network. By constructing a semantic association network containing semantic nodes, association edges, and evolutionary weights, it can accurately characterize the relationships between heterogeneous resources. The changing nature of semantic relationships allows the connections between resources to adaptively adjust as resources are updated and evolve, improving the accuracy and timeliness of semantic association expression. Based on semantic association networks, cross-modal virtualization mapping rule mining can uncover potential dynamic mapping patterns between resources of different modalities. The generated set of cross-modal virtualization mapping rules can effectively guide the virtualization annotation of heterogeneous resources, improving the applicability of the rules and the accuracy of the mapping. Virtualization annotation processing using the set of cross-modal virtualization mapping rules enables unified and standardized annotation of heterogeneous resources, ensuring that the annotated data reflects the latest state and semantic changes of the resources. This provides high-quality structured annotation data for the construction of the knowledge base. The final evolved heterogeneous resource knowledge base, containing conceptual units, relational units, and attribute units, enables knowledge updates and continuous evolution, effectively improving the organization, management, and application of heterogeneous resource knowledge. Attached Figure Description

[0007] Figure 1 This is a flowchart of a method for virtualizing and labeling heterogeneous resources and creating a knowledge base, provided in an embodiment of the present invention.

[0008] Figure 2 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation

[0009] Please see Figure 1 , Figure 1 A flowchart illustrating a method for heterogeneous resource virtualization annotation and knowledge base creation, provided in an embodiment of the present invention, is included. This method can be executed by a computer system and includes the following steps:

[0010] Step S100: Obtain a collection of heterogeneous technology project resources of various types, perform cross-modal semantic feature extraction on the collection of heterogeneous technology project resources, and generate a cross-modal resource semantic feature set, which includes text semantic features, structural semantic features and interface semantic features.

[0011] A heterogeneous collection of science and technology project resources encompasses a variety of resources related to science and technology projects, in different types, formats, and from different sources. These resources may originate from various databases, literature repositories, project documents, etc., such as research reports, academic papers, and project outcome presentation documents published by different research institutions. Cross-modal semantic feature extraction is the process of processing different modal information within the heterogeneous collection of science and technology project resources to extract semantically meaningful features. Textual semantic features are features extracted from the text content of the resources that express semantic information, such as the semantics implied in the text of the science and technology project name and key technology description. Structural semantic features are semantic features related to the resource structure, such as the semantics reflected in the chapter structure of a document and the row and column relationships of a data table. Interface semantic features are semantic features related to the resource interface, such as the meaning of interface parameters and return value types.

[0012] When acquiring heterogeneous technology project resource sets, relevant resources can be crawled from publicly available information such as major scientific research databases and academic websites using web crawlers. Alternatively, with authorization, existing project documents can be collected from internal project management systems. For cross-modal semantic feature extraction, for textual semantic features, word embedding models from natural language processing, such as the Word2Vec model, can be used to convert words in the text into vector representations, thereby extracting the text's semantic features. For structural semantic features, graph structure analysis methods can be employed, representing the resource structure as a graph, with nodes representing different elements and edges representing relationships between elements. Graph embedding algorithms, such as the DeepWalk algorithm, can then be used to extract structural semantic features. For interface semantic features, information such as parameter names, parameter types, and return value types can be extracted by parsing the interface documentation. Classification algorithms from machine learning, such as support vector machines, can then be used to classify this information and extract features.

[0013] Step S200: Construct a semantic association network based on the cross-modal resource semantic feature set. The semantic association network includes semantic nodes, association edges, and evolution weights.

[0014] Semantic association networks are network structures used to represent the associations between semantic features of cross-modal resources. Semantic nodes represent entities with corresponding semantics, such as a technical concept or research object in a science and technology project. Association edges represent the associations between semantic nodes, such as the reference relationship between technical concepts or the causal relationship between research objects. Evolutionary weights reflect how the associations change over time and reflect the importance of the associations at different times.

[0015] Constructing a semantic association network based on a cross-modal resource semantic feature set involves comprehensively utilizing textual semantic features, structural semantic features, and interface semantic features to identify semantic relationships and represent them in network form. The construction process begins by analyzing and processing different types of semantic features to determine the specific content of semantic nodes and association edges. For determining evolutionary weights, historical data analysis can be used to statistically analyze the frequency and strength of associations at different time points, thereby assigning appropriate evolutionary weights to the association edges.

[0016] As one implementation method, step S200 can be specifically implemented as the following steps S210~S250:

[0017] Step S210: Perform concept extraction processing on the text semantic features in the cross-modal resource semantic feature set, extract the concept units that evolve over time through the temporal semantic entity recognition algorithm, perform semantic drift detection processing on the concept units, and generate a set of concept nodes with temporal tags.

[0018] Concept extraction processing identifies concepts with explicit semantics from text semantic features. Temporal semantic entity recognition algorithms are algorithms that consider the time factor and can identify semantic entities (i.e., concept units) that change over time. Semantic drift detection processing detects the semantic changes of concept units at different times, because the meaning of some concepts may change over time. A set of concept nodes with temporal tags is a set formed by tagging extracted concept units with time information; each concept node is associated with a corresponding time. In concept extraction processing, Named Entity Recognition (NER) technology can be used, combined with pre-trained language models, such as the BERT model, to identify concept units from text semantic features. For temporal semantic entity recognition algorithms, methods based on time series analysis and machine learning can be used to divide the text according to time sequence, perform semantic entity recognition within each time interval, and then identify concept units that evolve over time by comparing the recognition results of different time intervals. For semantic drift detection processing, the semantic similarity of concept units at different times can be calculated, such as using cosine similarity to calculate the similarity between word vectors. If the similarity is lower than a pre-set threshold, semantic drift has occurred.

[0019] Step S220: Perform relation identification processing on the structural semantic features in the cross-modal resource semantic feature set, determine the association type and direction attribute between concept units through spatiotemporal dependency parsing, calculate the temporal occurrence frequency of association type to generate relation strength value, and generate an association edge set based on association type, direction attribute and relation strength value.

[0020] Relationship identification is the process of identifying the relationships between conceptual units from structural semantic features. Spatiotemporal dependency parsing combines temporal and spatial information to analyze the dependency relationships between words in a sentence or structure. This method can determine the type of relationship between conceptual units, such as causal, inclusion, and parallel relationships, as well as the directional attribute of the relationship—that is, the direction from one concept to another. The relationship strength value is a numerical value that measures the tightness of the relationship between conceptual units, determined by calculating the frequency of occurrence of the relationship type at different times. The set of relationship edges is a collection of multiple relationship edges, each containing information such as the relationship type, directional attribute, and relationship strength value.

[0021] When performing relation identification processing, for spatiotemporal dependency parsing, dependency parsing tools in natural language processing, such as Stanford Core NLP, can be used to analyze structural semantic features in conjunction with temporal information. For example, for the chapter structure of a document, by analyzing the citation and sequence relationships between chapters, the association type and direction attributes between conceptual units can be determined. When calculating the temporal frequency of association types, the structural semantic features at different times can be classified and statistically analyzed, recording the occurrence frequency of each association type in each time window, and then the relation strength value can be calculated based on these frequencies.

[0022] As one implementation method, step S220 can be specifically implemented as the following steps S221~S225:

[0023] Step S221: Perform temporal syntactic dependency analysis on the structural semantic features, identify the structure of relational phrases through the temporal dependency relation tree, extract the temporal change sequence of the core relational verbs, and generate a set of relational verbs.

[0024] Temporal syntactic dependency parsing is a process of performing syntactic dependency analysis on structural semantic features by incorporating temporal factors. A temporal dependency tree is a tree-like structure that combines syntactic dependency relations with temporal information. It can identify relational phrase structures within structural semantic features. Relational phrase structures are phrases containing relational information, such as phrases expressing relationships like "lead to" or "promote." The core relational verb is the verb that plays a key role in the relational phrase structure, reflecting the association between conceptual units. Extracting the temporal change sequence of the core relational verb records the occurrence and trends of these verbs at different times, generating a set of relational verbs that includes these core relational verbs.

[0025] When performing temporal dependency parsing, deep learning-based dependency parsing models, such as the BiLSTM-CRF model, can be used. Structural semantic features are input into the model in chronological order to generate a temporal dependency tree. By traversing and analyzing the temporal dependency tree, the structure of relational phrases can be identified. To extract the temporal variation sequence of the core relational verbs, the core relational verbs identified within each time window can be recorded to form a sequence.

[0026] Step S222: Perform temporal semantic expansion processing on each relational verb in the relational verb set, obtain the synonym set for different time windows through temporal word vector similarity calculation, and generate a standardized relation type set based on the synonym set and the same relation type.

[0027] Temporal semantic expansion is the process of semantically expanding relational verbs along the time dimension. Temporal word vector similarity calculation assesses the similarity between word vectors at different times, identifying words semantically similar to a given relational verb at different times, thus obtaining a thesaurus. Merging similar relation types combines relation types represented by semantically similar relational verbs to form more unified and standardized relation types. The standardized relation type set is the set containing standardized relation types obtained after the merging process.

[0028] When performing temporal semantic expansion, for calculating temporal word vector similarity, a pre-trained word vector model, such as the GloVe model, can be used to extract the word vectors of relational verbs at different times, and then calculate the cosine similarity between them. If the similarity is higher than a pre-set threshold, these words are considered synonyms. For example, "promote" and "push" may have similar meanings at different times, and by calculating their word vector similarity, they can be classified as synonyms. Based on the obtained synonym set, the relation types represented by relational verbs with the same or similar semantics are merged. For example, the relation types represented by "promote" and "push" are merged into a standardized "promote / push" relation type.

[0029] Step S223: Perform temporal direction attribute determination processing on each relation type in the standardized relation type set. Identify and determine the temporal flow of the starting and ending concepts of the relation by the temporal changes in the position of the arguments, and generate a directed relation graph containing temporal direction markers.

[0030] The process of determining the temporal direction attribute is the process of determining the directional attribute of a relation type at different times. The temporal change of argument position refers to the change in position of arguments (i.e., conceptual units) in a relation at different times. By analyzing this change, the temporal flow of the starting and ending concepts of the relation can be determined. A directed relation graph with temporal direction markings represents the relation type using a directed graph, and marks the graph with time information. Each directed edge represents a relation, the direction of the edge indicates the flow of the relation, and the edge also carries a time marker.

[0031] When determining the temporal direction attribute, the relation instances within each time window can be analyzed to observe the position of the arguments. For example, in the relation "A causes B", "A" is the starting concept and "B" is the ending concept. By comparing the changes in the position of arguments in relation instances across different time windows, the temporal flow of the relation can be determined. If, during a certain time period, the relation points from "A" to "B", and during another time period, the relation may change from "B" to "A", this change is recorded. After organizing the temporal flow information for each relation type, a directed relation graph is constructed, and the time information is marked on the graph.

[0032] Step S224: Calculate the temporal strength of each directed edge in the relational directed graph, count the rate of change of the frequency of occurrence of the same relation type in resource sets in different time windows, and generate a relation strength value by combining the temporal importance weight of the associated concept.

[0033] Temporal strength calculation is the process of calculating the strength of directed edges over time. The frequency of occurrence change rate refers to the proportion of change in the frequency of occurrence of the same relation type across different time windows; statistically analyzing this rate of change can reflect the activity and stability of the relation type. The temporal importance weight of associated concepts is a weight assigned based on the importance of associated concepts at different times; importance can be determined based on factors such as the number of times a concept is cited and its influence. The relation strength value is a numerical value representing the relation strength obtained by comprehensively considering both the frequency of occurrence change rate and the temporal importance weight of associated concepts.

[0034] When calculating temporal importance, the frequency of occurrence of each relation type is first counted across different time windows, and the rate of change in frequency between adjacent time windows is calculated. For the temporal importance weight of related concepts, the weights can be determined using methods such as the Analytic Hierarchy Process (AHP) by calculating indicators such as the number of times the concept is cited and referenced at different times. Then, the rate of change in frequency and the temporal importance weight of the related concepts are weighted and summed to obtain the relation strength value.

[0035] Step S225: Construct relation feature tuples based on standardized relation types, temporal direction labels, and relation strength values, and combine multiple relation feature tuples according to time windows to generate a set of associated edges.

[0036] A relation feature tuple is a tuple containing information such as standardized relation type, temporal direction marker, and relation strength value; it provides a comprehensive description of the relation. Combining multiple relation feature tuples by time window means grouping relation feature tuples within the same time window together to form a set. An association edge set is a set composed of relation feature tuples combined by time window, with each association edge corresponding to one relation feature tuple.

[0037] When constructing relation feature tuples, the standardized relation type, temporal direction label, and relation strength value are combined to form a tuple. For example, for a directed edge representing a "causal" relation, with a standardized relation type of "causal," a temporal direction label of "from cause concept to effect concept," and a relation strength value of 0.7, the relation feature tuple can be represented as ("causal," from cause to effect, 0.7). All relation feature tuples within each time window are collected and organized, grouped by time window, and a set of associated edges is generated.

[0038] Step S230: Perform attribute extraction processing on the interface semantic features in the cross-modal resource semantic feature set, and extract the attribute names and value range sequences corresponding to the concept units through the adaptive key-value pair recognition algorithm to generate an attribute label set containing temporal constraints.

[0039] Attribute extraction is the process of extracting attribute information from the semantic features of an interface into conceptual units. Adaptive key-value pair recognition algorithms are algorithms that can adaptively identify key-value pairs based on different interface semantic features. These algorithms can extract the attribute name and value range sequence corresponding to a conceptual unit. The attribute name describes a specific attribute of the conceptual unit, and the value range sequence is the sequence of possible values ​​for the attribute over time. The attribute tag set containing temporal constraints is a collection of attribute tags formed by combining the attribute name, value range sequence, and temporal constraints. The temporal constraints represent the restrictions on the attribute's values ​​at different times.

[0040] When performing attribute extraction, an adaptive key-value pair recognition algorithm can employ a combination of rule-based and machine learning approaches. First, rules are established based on the interface document format and common key-value pair patterns. Then, machine learning models, such as Support Vector Machines (SVM), are used to identify cases that do not conform to the rules. For example, interface documents may contain formats like "parameter name: value range." Rules can identify most key-value pairs, while machine learning models are used to supplement the identification of some special formats. After extracting the attribute names and value range sequences, temporal constraints are determined based on business rules and data statistics at different times. For example, an attribute may only take positive values ​​during a certain time period, while its value range expands to include both positive and negative values ​​during another time period. This information is combined into attribute labels, ultimately generating a set of attribute labels containing temporal constraints.

[0041] As one implementation method, step S230 can be specifically implemented as the following steps S231~S235:

[0042] Step S231: Perform time-series key-value pair recognition on the interface semantic features, and use time-series regular expressions to match the combination patterns of attribute names and values ​​in different time windows to extract the original set of attribute key-value pairs.

[0043] Temporal key-value pair recognition is the process of identifying key-value pairs in the semantic features of an interface along a time dimension. Temporal regular expressions are regular expressions that consider the time factor, allowing matching of combinations of attribute names and values ​​across different time windows. The original attribute key-value pair set is a collection of key-value pairs extracted through matching, containing both the attribute name and its value.

[0044] When performing time-series key-value pair recognition, different regular expression patterns can be designed for time-series regular expressions based on the format of the API documentation and the changing patterns over different time periods. For example, within a certain time window, the attribute names and values ​​in the API documentation may appear in the format "attribute name:value", while in another time window they may appear in the format "value (attribute name)". Corresponding regular expressions are designed for different formats. The semantic features of the API in each time window are matched with the corresponding regular expressions to extract the attribute names and values, forming the original set of attribute key-value pairs.

[0045] Step S232: Perform time-series standardization of attribute names on the original set of attribute key-value pairs. By matching the time sequence of the attribute dictionary, synonymous attribute names in different time windows are unified into standard attribute names, generating standardized attribute key-value pairs.

[0046] Attribute name temporal standardization is the process of standardizing attribute names along a time dimension. Temporal matching of the attribute dictionary involves matching the attribute names in the original attribute key-value pairs against the attribute dictionary, considering the time factor to identify synonymous attribute names from different times. Standard attribute names are the attribute names determined after unified processing, and standardized attribute key-value pairs are the key-value pairs obtained by replacing the attribute names in the original attribute key-value pairs with the standard attribute names. When performing attribute name temporal standardization, an attribute dictionary is first constructed, containing synonymous attribute names from different times and their corresponding standard attribute names. Then, each attribute name in the original set of attribute key-value pairs is matched against the attribute dictionary, finding the corresponding standard attribute name based on the time information. For example, in interface documents from different times, there might be different names for the same attribute, such as "parameter A" and "parameter 1." Through temporal matching of the attribute dictionary, these are unified as "standard parameter A." The attribute names in the original attribute key-value pairs are replaced with the standard attribute names to generate standardized attribute key-value pairs. This makes attribute names more consistent, facilitating subsequent attribute analysis and management.

[0047] Step S233: Perform time-series data type identification processing on the attribute values ​​in the standardized attribute key-value pairs. Determine the time-series changes of data types such as numeric, text, and date through time-series analysis of value formats, and generate attribute description tuples containing data type time-series tags.

[0048] Time-series data type identification is the process of identifying and analyzing the data type of attribute values ​​along the time dimension. Time-series analysis of value formats analyzes how the format of attribute values ​​changes over different times, thus determining how the data type of the attribute value changes over time. Data type time signatures are tags that represent the data type of attribute values ​​at different times. An attribute description tuple containing data type time signatures is a tuple formed by combining the attribute name, attribute value, and data type time signature.

[0049] When performing time-series data type identification, a combination of regular expressions and machine learning can be used for time-series analysis of value formats. For common data types, such as numeric, text, and date types, corresponding regular expressions are designed for initial identification. For some complex formats or data types that are difficult to identify using regular expressions, machine learning models, such as decision trees, are used for classification. For example, an attribute value might be in date format like "2023-01-01" at one time, which can be identified as a date type using regular expressions; at another time, it might be in text format like "abc", which can be accurately identified as text type using a machine learning model. The identified data types are then labeled with time information and combined with the attribute name and attribute value to form an attribute description tuple.

[0050] Step S234: Perform time-series inference on the value range of the attribute description tuple, calculate the time-series changes of the extreme values ​​and mean values ​​of numerical attributes in different time windows through statistical analysis, identify the time-series evolution of the value set of text attributes through enumeration and induction, and generate attribute constraints.

[0051] Value range time-series inference is the process of inferring the value range of an attribute over time. For numerical attributes, statistical analysis calculates the temporal changes of extreme values ​​(maximum and minimum values) and mean values ​​across different time windows, revealing the dynamic changes in the value range. For textual attributes, enumeration and induction methods are used to collect and organize textual values ​​appearing in different time windows, identifying the temporal evolution of the value set—that is, how the value set changes over different times. Attribute constraints are the restrictions on attribute values ​​at different times, determined based on the results of value range time-series inference.

[0052] When inferring the range of values ​​over time, for numerical attributes, statistical analysis methods are used, such as calculating the maximum, minimum, and mean values ​​of sample data in different time windows. By comparing the extreme values ​​and mean values ​​in different time windows, the changes in the value range are determined. For textual attributes, the text values ​​appearing in each time window are enumerated. For example, the text values ​​appearing in time window T1 are {"apple", "banana"}, and the text values ​​appearing in time window T2 are {"apple", "banana", "orange"}, showing that the set of values ​​is expanding. Based on these analysis results, attribute constraints are generated, such as the value range of numerical attributes within a certain time period being [minimum, maximum], and the value set of textual attributes within a certain time period being {specific values}. These attribute constraints provide a basis for data validity checks and usage.

[0053] Step S235: Combine the standard attribute name, data type timing mark and attribute constraint to generate attribute label. Group multiple attribute labels according to their respective concept nodes and time windows to generate an attribute label set.

[0054] An attribute label is a tag that combines a standard attribute name, data type timing marker, and attribute constraints. It can comprehensively describe the characteristics of a specific attribute of a conceptual unit in the time dimension. Grouping multiple attribute labels by their respective conceptual node and time window involves placing attribute labels belonging to the same conceptual node and within the same time window together to form a group. The attribute label set is a collection of these grouped attribute labels.

[0055] When generating attribute tags, the standard attribute name, data type timing marker, and attribute constraints are combined according to a specific format. For example, for an attribute with the standard attribute name "Parameter A", a data type timing marker of "numeric in time window T1, text in time window T2", and attribute constraints of "value range [1,10] in time window T1, and value set {"Option 1", "Option 2"} in time window T2", the attribute tags ("Parameter A", "numeric in time window T1, text in time window T2", "value range [1,10] in time window T1, and value set {"Option 1", "Option 2"}") are generated. The attribute tags for each concept node in different time windows are collected and grouped to generate an attribute tag set.

[0056] Step S240: Input the set of concept nodes, the set of associated edges, and the set of attribute labels into the network to build the model. Establish the time-varying mapping relationship between concept nodes and associated edges through the temporal node connection algorithm. Bind attribute labels to concept nodes in a temporal order through the temporal label attachment algorithm to generate the initial semantic association network.

[0057] The network construction model is used to build semantic association networks. It takes a set of concept nodes, a set of associated edges, and a set of attribute labels as input, processes them through a pre-defined algorithm, and generates the semantic association network. The temporal node connection algorithm is an algorithm that considers the time factor and establishes a time-varying mapping relationship between concept nodes and associated edges. This relationship reflects the changes in the association between concept nodes at different times. The temporal label attachment algorithm is an algorithm that binds attribute labels to concept nodes in the time dimension, so that each concept node is associated with a corresponding attribute label at the corresponding time. The initial semantic association network is a preliminary semantic association network generated after processing by the network construction model.

[0058] After inputting the set of concept nodes, the set of associated edges, and the set of attribute labels into the network to construct the model, a graph neural network (GNN)-based approach can be used for temporal node connection algorithms. Concept nodes and associated edges are represented as nodes and edges in a graph, and a GNN model, such as Graph SAGE, is used to learn the time-varying mapping relationship between concept nodes and associated edges. By processing the graph structure at different times, the features of nodes and edges are updated, establishing the time-varying mapping relationship. For temporal label attachment algorithms, a hash table can be used to map attribute labels to concept nodes according to time. For example, for each concept node, its corresponding attribute labels at different times are stored in a hash table, and lookup is performed using time as the key, achieving temporal binding between attribute labels and concept nodes.

[0059] Step S250: Perform network structure optimization on the initial semantic association network. Identify closely connected concept node groups under different time windows using the temporal community discovery algorithm, delete redundant association edges with time-varying connection strength below the threshold within the group, and retain core association edges with time-varying connection strength above the threshold across groups to obtain the optimized semantic association network.

[0060] Network structure optimization involves adjusting and optimizing the structure of the initial semantic association network to improve its performance and understandability. The temporal community detection algorithm identifies tightly connected clusters of concept nodes across different time windows. These clusters form communities due to the strong connections between their internal nodes. Time-varying connection strength reflects the importance of each edge over time. Redundant edges are those with intra-cluster time-varying connection strength below a threshold; these edges contribute little to the connectivity within the community and can be removed. Core edges are those with cross-cluster time-varying connection strength above a threshold; these are critical paths for information transfer between different communities and should be retained.

[0061] When performing network structure optimization, dynamic graph-based community detection algorithms, such as the DynGNMF algorithm, can be used for time-series community detection. This algorithm incorporates the time factor into the community detection process, identifying closely connected concept node groups by analyzing the graph structure at different times. For calculating time-varying connection strength, the relationship strength values ​​calculated in previous steps can be used in conjunction with the time factor for comprehensive consideration. Specifically, time-varying connection strength reflects the connection strength of associated edges at different times, reflecting their importance in the time dimension. First, the relationship strength values ​​of associated edges in different time windows are obtained from previous steps. These relationship strength values ​​are calculated by combining the rate of change of the frequency of occurrence of the relationship represented by the associated edge in the resource set of different time windows with the temporal importance weight of the associated concepts. Then, to reflect the time-varying characteristics, a time decay factor is introduced. This factor is set according to the proximity of the time window; the closer the time window is to the current time, the larger its corresponding time decay factor, and vice versa. The relationship strength value of each time window is multiplied by the corresponding time decay factor, and then the results of all time windows are weighted and summed to finally obtain the time-varying connection strength of the associated edge. A threshold is set, and edges with time-varying connectivity strength below this threshold within a group are marked as redundant edges and removed from the network. For cross-group edges, their inter-community connectivity between communities (reflecting the importance of the edge in information transfer between different communities) are calculated, and edges with inter-community connectivity between communities above another threshold are retained as core edges.

[0062] As one implementation method, step S250 can be specifically implemented as the following steps S251~S255:

[0063] Step S251: The initial semantic association network is divided into communities using a temporal community detection algorithm. The temporal community detection algorithm combines the temporal evolution characteristics of network topology and node attributes to identify closely connected concept node groups under different time windows.

[0064] Community partitioning is the process of dividing the initial semantic association network into different communities. Temporal community detection algorithms comprehensively consider the network's topology (i.e., the connections between nodes and edges) and the evolutionary characteristics of node attributes over time, enabling them to more accurately identify closely connected clusters of concept nodes across different time windows. These clusters of concept nodes have relatively tight connections, similar semantics or functions, and form a relatively independent community.

[0065] When performing community segmentation, for temporal community detection algorithms, deep learning-based methods can be employed, such as attention-based graph neural network (GNN) models. These models can learn the temporal evolution characteristics of network topology and node attributes, calculating the similarity and connection strength between nodes by processing the graph structure at different times. For example, using the T-GCN (Temporal Graph Convolutional Network) model, convolution operations are performed on the graph for each time window to update node features, and then clustering is performed based on these features to identify tightly connected conceptual node groups.

[0066] Step S252: Perform intra-group association edge analysis on each concept node group, calculate the mean and variance of the time-varying connection strength of each association edge in the group within a preset time window, and mark the association edges with a mean lower than the first threshold and a variance higher than the second threshold as potential redundant edges.

[0067] Intra-group correlation edge analysis is the process of analyzing and evaluating the correlation edges within each concept node group. The mean of the time-varying connection strength within a preset time window reflects the average connection strength of the correlation edges within that time window, while the variance reflects the fluctuation of the connection strength. The first threshold and the second threshold are pre-set criteria used to determine whether a correlation edge is a potentially redundant edge. Potentially redundant edges are correlation edges whose mean time-varying connection strength within the group is lower than the first threshold and whose variance is higher than the second threshold. These edges have weak connection strength and large fluctuations, and their contribution to the connection of nodes within the community is unstable; therefore, they can be considered as edges to be deleted.

[0068] When performing intra-group edge analysis, for each edge, the mean and variance of its time-varying connection strength within a preset time window are calculated. For example, assuming the preset time window is [T1, T2], for edge E, its time-varying connection strength value at each time point within that time window is recorded, and then the mean and variance of these values ​​are calculated. If the mean is lower than a first threshold, it indicates that the average connection strength of the edge is weak; if the variance is higher than a second threshold, it indicates that the connection strength fluctuates greatly. Edges that simultaneously meet both conditions are marked as potentially redundant edges.

[0069] Step S253: Perform temporal contribution evaluation on the marked potentially redundant edges, calculate the comprehensive contribution score by combining the historical cumulative contribution value and the current contribution value of the edge, and delete redundant related edges whose comprehensive contribution score is lower than the third threshold.

[0070] The temporal contribution evaluation of associated edges is a process of assessing the contribution of potentially redundant edges to network connectivity over time. The historical cumulative contribution value is the accumulated contribution of associated edges to network connectivity over past periods, while the current contribution value is the contribution of associated edges to network connectivity at the current time. The comprehensive contribution score is a score calculated by combining the historical cumulative contribution value and the current contribution value, used to measure the overall contribution of associated edges. The third threshold is a pre-defined standard for determining whether an associated edge is redundant; associated edges with a comprehensive contribution score below the third threshold are considered redundant and need to be removed from the network.

[0071] When evaluating the temporal contribution of associated edges, the historical cumulative contribution value can be calculated by recording the connection strength and number of connections of the associated edge over past times. For example, the connection strength at each past time point is multiplied by the number of connections at that time point, and then summed to obtain the historical cumulative contribution value. For the current contribution value, the connection strength at the current time can be used as the current contribution value. The historical cumulative contribution value and the current contribution value are then weighted and summed according to certain weights to obtain the comprehensive contribution score. If the comprehensive contribution score is lower than the third threshold, the associated edge is marked as a redundant edge and removed from the network.

[0072] Step S254: Perform bridging importance assessment on the associated edges of cross-concept node groups, calculate the inter-community connectivity centrality of the edges, and retain the core associated edges with a centrality higher than the fourth threshold. The core associated edges serve as the key paths for information transmission between communities.

[0073] Inter-community connectivity between-intermediation centrality is an indicator that measures the mediating role of an edge in information transfer between different communities, reflecting the criticality of the edge in connecting different communities. The fourth threshold is a pre-defined standard for judging whether an edge is a core interconnected edge. Interconnected edges with a between-intermediation centrality higher than the fourth threshold are considered core interconnected edges. These edges are the key paths for information transfer between different communities and need to be retained.

[0074] When performing bridging importance assessment, for each cross-concept node group's associated edge, its inter-community connectivity between-edges is calculated. Classical between-edge centrality calculation methods can be used, such as determining between-edge centrality by calculating the frequency of its occurrence across all shortest paths. For example, for a network, the shortest paths between all node pairs are calculated, the number of times each cross-community associated edge appears on these shortest paths is counted, and the number of occurrences is divided by the total number of shortest paths to obtain the between-edge's between-edge centrality. If the between-edge centrality is higher than a fourth threshold, the associated edge is marked as a core associated edge and retained. This ensures smooth information transfer between different communities and improves the overall network performance.

[0075] Step S255: Perform multiple rounds of iterative optimization on the optimized semantic association network, repeatedly performing community partitioning, redundant edge deletion, and core edge retention operations until the network structure change rate of two consecutive iterations is lower than the fifth threshold, thus obtaining the final optimized semantic association network.

[0076] Multi-round iterative optimization involves repeatedly performing optimization operations on the optimized semantic association network to further improve its performance and stability. Community partitioning, redundant edge removal, and core edge retention operations are the same operations described in the previous steps: partitioning the network into communities, removing redundant association edges, and retaining core association edges. The network structure change rate is an indicator that measures the degree of change in the network structure between two consecutive iterations. It can be determined by calculating factors such as the increase or decrease in the number of nodes and edges, and changes in connection relationships. The fifth threshold is a pre-set standard for determining whether the iteration should stop. When the network structure change rate is lower than the fifth threshold for two consecutive iterations, it indicates that the network structure has stabilized, the iteration can stop, and the final optimized semantic association network is obtained.

[0077] In the multi-round iterative optimization process, each round optimizes the network according to the community partitioning, redundant edge removal, and core edge retention operations described earlier. The rate of change of the network structure after each iteration is calculated and compared to a fifth threshold. If the rate of change of the network structure is lower than the fifth threshold for two consecutive iterations, it indicates that the network structure has basically stabilized and no further optimization is needed. The network obtained at this point is the final optimized semantic association network. Through this multi-round iterative optimization method, the network structure can be continuously adjusted to achieve optimal performance.

[0078] Step S300: Perform cross-modal virtualization mapping rule mining operation through semantic association network to generate a cross-modal virtualization mapping rule set, which includes text mapping rules, structure mapping rules and interface mapping rules.

[0079] Cross-modal virtualization mapping rule mining is the process of extracting mapping rules between different modalities from a semantic association network. Cross-modality refers to different types of resources, such as text, structure, and interface. Virtualization mapping rules are rules used to associate and transform resources from different modalities. Text mapping rules are mapping rules between text resources, such as mapping different expressions of textual concepts to the same standard concept. Structure mapping rules are mapping rules between structural resources, such as mapping different structural relationships to standardized structural relationships. Interface mapping rules are mapping rules between interface resources, such as mapping different interface parameters to a unified parameter type. The cross-modal virtualization mapping rule set is a collection of text mapping rules, structure mapping rules, and interface mapping rules.

[0080] When mining cross-modal virtualization mapping rules, graph mining and machine learning methods can be used. The features of nodes and edges in the semantic association network are analyzed to identify association patterns between different modalities. For example, for text mapping rules, the semantic similarity of concept nodes can be analyzed to identify concept nodes with similar semantics, and the mapping relationships between them can be used as text mapping rules. For structural mapping rules, the type and temporal changes of associated edges can be analyzed to find the transformation rules between different structural relationships, generating structural mapping rules. For interface mapping rules, the relationship between attribute labels and concept nodes can be analyzed to find mapping rules for interface parameters. Through these operations, a set of cross-modal virtualization mapping rules is generated, which can provide a rule foundation for the virtualization annotation of heterogeneous technology project resources.

[0081] As one implementation method, step S300 can be specifically implemented as the following steps S310~S350:

[0082] Step S310: Extract temporal feature vectors from concept nodes in the semantic association network, map concept nodes in different time windows into low-dimensional dense temporal vectors using a graph embedding algorithm, and calculate the cosine similarity between temporal vectors to generate a concept similarity matrix.

[0083] Temporal feature vector extraction is the process of extracting vectors with temporal features from concept nodes in a semantic association network. Graph embedding algorithms can map concept nodes in different time windows into low-dimensional dense temporal vectors, which can more effectively represent the features and temporal information of concept nodes. Cosine similarity is an indicator that measures the similarity between two vectors. By calculating the cosine similarity between temporal vectors, the similarity between concept nodes can be obtained. The concept similarity matrix is ​​a matrix composed of the similarity values ​​between all concept nodes, where each element represents the similarity between two concept nodes.

[0084] When performing temporal feature vector extraction, the Node2Vec algorithm can be used for graph embedding algorithms. This algorithm samples node sequences in the graph through random walks and then uses a Skip-Gram model to learn the vector representation of the nodes. For graph structures with different time windows, the Node2Vec algorithm is applied separately to obtain low-dimensional dense temporal vectors of concept nodes at different times. Then, the cosine similarity formula is used to calculate the cosine similarity between the temporal vectors of any two concept nodes. The cosine similarity values ​​between all concept nodes are stored in a matrix to generate a concept similarity matrix.

[0085] Step S320: Based on the concept similarity matrix, perform evolutionary similarity concept group identification processing, filter concept node clusters with similarity higher than the threshold under different time windows through temporal density clustering algorithm, perform temporal hierarchical relationship analysis on the nodes within the cluster to determine the core evolutionary concept, and generate text mapping rules containing the core evolutionary concept and subordinate evolutionary concepts.

[0086] Evolutionary similarity concept group identification is the process of identifying groups of concept nodes with evolutionary similarity from a concept similarity matrix. Temporal density clustering algorithms can filter out concept node clusters with similarity exceeding a threshold at different time windows; nodes within these clusters have high similarity. Temporal hierarchical relationship analysis analyzes the hierarchical relationship of nodes within a cluster over time to determine the core evolutionary concept and subordinate evolutionary concepts. The core evolutionary concept is the concept that plays a dominant role in the evolutionary process, while subordinate evolutionary concepts are related to and dependent on the core evolutionary concept. Text mapping rules are rules that map subordinate evolutionary concepts to core evolutionary concepts. The antecedent of the rule includes the subordinate evolutionary concept and the temporal similarity threshold condition, the consequent is the core evolutionary concept, and the rule confidence is set as the weighted average of the cluster temporal consistency index.

[0087] When identifying evolutionarily similar concept groups, for temporal density clustering algorithms, a time-extended version of the DBSCAN algorithm, such as the TD-DBSCAN algorithm, can be used. This algorithm incorporates the time factor into the density clustering process, selecting concept node clusters with similarity higher than a threshold under different time windows based on the concept similarity matrix. For the temporal hierarchical relationship analysis of nodes within a cluster, temporal centrality calculation methods can be used, such as calculating the connectivity centrality of nodes at different times. Nodes with high connectivity centrality are designated as core evolutionary concepts, and the remaining nodes are designated as subordinate evolutionary concepts. For example, in a concept node cluster, node A has high connectivity centrality at different times, so it is designated as the core evolutionary concept, and nodes B, C, etc., are designated as subordinate evolutionary concepts. The generated text mapping rule can be expressed as follows: if the temporal similarity between a subordinate evolutionary concept (such as node B) and a core evolutionary concept (such as node A) is higher than a certain threshold, then the subordinate evolutionary concept is mapped to the core evolutionary concept. The rule confidence is determined based on the weighted average of the cluster temporal consistency index, which is the average similarity value of all nodes in the cluster across each time window.

[0088] As one implementation method, step S320 can be specifically implemented as the following steps S321~S325:

[0089] Step S321: Perform time-series threshold filtering on the concept similarity matrix, set a similarity lower limit threshold that changes with the time window, retain concept node pairs in the matrix whose similarity values ​​are higher than the corresponding threshold under different time windows, and generate an initial set of similar node pairs.

[0090] Time-series threshold filtering is a process of threshold filtering of the concept similarity matrix over time. Setting a lower similarity threshold that varies with the time window involves setting different similarity thresholds for each time window based on business needs and data characteristics at different times. Retaining concept node pairs with similarity values ​​higher than the corresponding threshold in different time windows involves filtering out those concept node pairs with high similarity to form an initial set of similar node pairs.

[0091] When performing time-series threshold filtering, a lower similarity threshold is set based on the semantic stability and data distribution of concept nodes at different times. For example, in early time windows, due to limited data, the semantics of concepts may be less stable, so a lower similarity threshold is set; in later time windows, with more data and clearer semantics, a higher similarity threshold is set. For each element in the concept similarity matrix, its corresponding time window is compared with the corresponding lower similarity threshold. If the similarity value is higher than the threshold, the corresponding concept node pair is retained, generating an initial set of similar node pairs.

[0092] Step S322: Perform temporal clustering on the initial set of similar node pairs. Merge node pairs from different time windows into concept node clusters using a temporal hierarchical clustering algorithm. Calculate the average similarity value of all nodes in each time window within the cluster as the cluster temporal consistency index.

[0093] Temporal clustering is the process of clustering an initial set of similar node pairs along the time dimension. Temporal hierarchical clustering algorithms can merge node pairs from different time windows into concept node clusters. The cluster temporal consistency index measures the consistency of nodes within a cluster along the time dimension, and is obtained by calculating the average similarity value of all nodes within the cluster across each time window.

[0094] When performing temporal clustering, a time-extended version of the Agglomerative Clustering algorithm can be used for hierarchical temporal clustering. This algorithm starts with each pair of nodes and progressively merges highly similar pairs to form larger clusters of concept nodes. During the merging process, the time factor is considered to ensure temporal consistency among the merged node pairs. For example, for two pairs of nodes in different time windows, if their time interval is short and their similarity is high, they are merged. For each formed cluster of concept nodes, the average similarity value of all nodes within the cluster is calculated across all time windows. For example, for a cluster containing nodes A, B, and C, the similarity values ​​between them are calculated in time windows T1, T2, and T3, and then the average is obtained to obtain the cluster temporal consistency index.

[0095] Step S323: Perform temporal consistency verification on the concept node clusters, select node clusters whose cluster temporal consistency index is higher than the verification threshold in each time window as candidate evolutionary similar concept groups, perform temporal noise node detection on the candidate evolutionary similar concept groups, and delete abnormal nodes whose similarity in consecutive time windows within the group is lower than the cluster mean.

[0096] Temporal consistency verification is the process of verifying the consistency of concept node clusters over time. The verification threshold is a pre-set standard used to determine whether a node cluster possesses temporal consistency. Node clusters whose temporal consistency index is higher than the verification threshold in each time window are selected as candidate evolutionarily similar concept groups. Nodes within these clusters exhibit high temporal consistency. Temporal noise node detection is the process of detecting noisy nodes within the candidate evolutionarily similar concept groups. Noisy nodes are anomalous nodes whose similarity is lower than the cluster mean in consecutive time windows. These nodes may affect the consistency of the cluster and need to be removed.

[0097] During temporal consistency verification, the cluster temporal consistency index of each concept node cluster is compared with the verification threshold in each time window. If the value in all time windows is higher than the verification threshold, the node cluster is selected as a candidate evolutionarily similar concept group. For the candidate evolutionarily similar concept group, the mean similarity of all nodes within the cluster in each time window is calculated. Then, it is checked whether the similarity of each node in consecutive time windows is lower than the cluster mean. If it is lower, the node is marked as an abnormal node and deleted.

[0098] Step S324: Perform temporal hierarchical relationship analysis on the purified evolutionary similar concept group. Through temporal centrality calculation, identify the concept node with the highest connectivity in different time windows within the group as the core evolutionary concept, and the remaining nodes as subordinate evolutionary concepts to generate a core-subordinate temporal relationship structure.

[0099] Temporal hierarchical relationship analysis is the process of analyzing the hierarchical relationships of nodes within a purified group of evolutionarily similar concepts along the time dimension. Temporal centrality calculation calculates the centrality index of nodes along the time dimension, identifying the concept nodes with the highest connectivity across different time windows. The core evolutionary concept is the concept that plays a dominant role in the evolutionary process, while subordinate evolutionary concepts are related to and dependent on the core evolutionary concept. The core-subordinate temporal relationship structure represents the temporal relationship between the core evolutionary concept and subordinate evolutionary concepts.

[0100] When performing temporal hierarchical relationship analysis, various centrality indices, such as time-extended versions of degree centrality and betweenness centrality, can be used for temporal centrality calculation. For example, using the time-extended version of degree centrality, the connectivity of each node is calculated in different time windows. For a purified group of evolutionarily similar concepts, the connectivity of nodes is calculated in each time window, and the node with the highest connectivity is identified as the core evolutionary concept. For example, node A has the highest connectivity in time window T1, and node B has the highest connectivity in time window T2. The core evolutionary concept is determined comprehensively based on the situation in different time windows. The remaining nodes are considered as subordinate evolutionary concepts, forming a core-subordinate temporal relationship structure.

[0101] Step S325: Generate text mapping rules based on the core-subordinate temporal relationship structure. The antecedent of the text mapping rule includes subordinate evolutionary concepts and temporal similarity threshold conditions. The consequent of the text mapping rule is the core evolutionary concept. The rule confidence is set as the weighted average of the cluster temporal consistency index.

[0102] Text mapping rules are used to map subordinate evolutionary concepts to core evolutionary concepts. The antecedent of the rule includes the subordinate evolutionary concept and a temporal similarity threshold condition; that is, the rule is triggered when the subordinate evolutionary concept meets a certain temporal similarity threshold. The consequent of the rule is the core evolutionary concept, i.e., the target concept for mapping. Rule confidence is an indicator that measures the reliability of the rule, set as a weighted average of the cluster temporal consistency index, which reflects the consistency of nodes within a cluster over time.

[0103] When generating text mapping rules, subordinate evolutionary concepts are associated with the core evolutionary concept based on the core-subordinate temporal relationship structure. For example, in a core-subordinate temporal relationship structure, the core evolutionary concept is "artificial intelligence technology," and the subordinate evolutionary concepts are "machine learning technology," "deep learning technology," etc. A temporal similarity threshold is set; for example, when the temporal similarity between "machine learning technology" and "artificial intelligence technology" is higher than 0.8, the rule is triggered. The consequent of the rule is "artificial intelligence technology," meaning that "machine learning technology" is mapped to "artificial intelligence technology." The rule confidence is determined based on the weighted average of the cluster temporal consistency index. For example, if the cluster temporal consistency index values ​​in different time windows are 0.8, 0.9, and 0.8, and the weighted average is 0.85, then the rule confidence is set to 0.85.

[0104] Step S330: Perform temporal type matching processing on the associated edges in the semantic association network, extract the temporal type identifier of the associated edges and the sequence of associated node pairs, identify the isomorphic relationship patterns under different time windows through the relation path mining algorithm, and generate structural mapping rules containing the temporal transformation direction of the relations by combining the node temporal attribute labels.

[0105] Temporal type matching is the process of matching the types of associated edges in a semantic association network along the time dimension. Extracting the temporal type identifiers of associated edges and the sequence of associated node pairs involves extracting the type identifiers of associated edges and the sequence of their associated node pairs along the time dimension. Relational path mining algorithms can identify isomorphic relation patterns, i.e., relation patterns with the same structure and type. Node temporal attribute labels are the attribute labels of nodes along the time dimension; combining these labels can generate structural mapping rules that include the direction of temporal transitions in relations. Structural mapping rules are mapping rules between structural resources, and these rules include the direction of temporal transitions in relations, i.e., the direction of relation transitions at different times.

[0106] When performing temporal type matching, a hash table can be used to store the type identifiers of associated edges and the sequence of associated node pairs according to time. For relation path mining algorithms, graph pattern matching algorithms, such as the GraphMatch algorithm, can be used. This algorithm searches for relation paths with the same structure and type in the semantic association network in different time windows, identifying isomorphic relation patterns. Combined with the temporal attribute labels of nodes, the direction of relation transformation at different times is analyzed. For example, for an associated edge representing a "causal" relationship, it may point from node A to node B in time window T1 and from node B to node A in time window T2. By analyzing the temporal attribute labels of nodes, the cause and direction of this transformation can be determined. The antecedent of the generated structure mapping rule includes temporal constraints on the original relation type and node type, the consequent of the rule is the standardized relation type, and the scope of application of the rule is set to the temporal scenario matching the relation pattern template.

[0107] As one implementation method, step S330 can be specifically implemented as the following steps S331~S335:

[0108] Step S331: Extract temporal features from the association edges in the semantic association network, extracting the relation type identifier, the start node ID sequence, the end node ID sequence, and the relation strength value sequence, and generate a relation feature vector.

[0109] Temporal feature extraction is the process of extracting features from the edges of a semantic association network along the time dimension. Relation type identifiers represent the type of the associated edge, such as "causal" or "containment." The start node ID sequence is the sequence of the IDs of the start nodes of the associated edge along the time dimension, and the end node ID sequence is the sequence of the IDs of the end nodes of the associated edge along the time dimension. The relation strength value sequence is the sequence of the relation strength values ​​of the associated edge along the time dimension. The relation feature vector is a vector formed by combining the relation type identifiers, the start node ID sequence, the end node ID sequence, and the relation strength value sequence; this vector comprehensively represents the features of the associated edge along the time dimension.

[0110] During temporal feature extraction, for each associated edge, its relation type identifier, starting node ID, ending node ID, and relation strength value are recorded in chronological order. For example, for an associated edge, in time window T1, the relation type identifier is "causal," the starting node ID is A, the ending node ID is B, and the relation strength value is 0.8; in time window T2, the relation type identifier is "causal," the starting node ID is B, the ending node ID is A, and the relation strength value is 0.7. This information is arranged in chronological order to generate a relation feature vector.

[0111] Step S332: Perform temporal type clustering on the relation feature vectors. Use the temporal K-means algorithm to cluster relation edges of the same type in different time windows into one class, calculate the temporal variance of the relation strength within the class, and delete abnormal relation edges with temporal variance values ​​higher than the threshold.

[0112] Temporal clustering is the process of clustering relation feature vectors along the time dimension. The temporal K-means algorithm can cluster relation edges of the same type across different time windows into one class. The temporal variance of the intra-class relation strength is an indicator of the degree of fluctuation in the relation strength of intra-class edges along the time dimension. By calculating this variance, abnormal relation edges can be identified. Abnormal relation edges are those with a temporal variance value higher than a threshold. These edges have large fluctuations in relation strength, which may be due to data noise or anomalies, and need to be deleted.

[0113] When performing temporal clustering, the temporal K-means algorithm uses relation feature vectors as input and sets the number of clusters K, typically determined by the number of relation types. The algorithm clusters relation feature vectors of the same type across different time windows into one class. For each class, it calculates the temporal variance of the relation strength within that class. For example, for a relation edge within a class, it records its relation strength values ​​across different time windows and calculates the variance of these values. If the variance exceeds a threshold, the relation edge is marked as an abnormal relation edge and deleted.

[0114] Step S333: Perform temporal isomorphic pattern recognition processing on the clustered relation edge classes, identify relation paths with the same node type sequence under different time windows through graph pattern matching algorithm, and generate relation pattern templates.

[0115] Temporal isomorphic pattern recognition is the process of identifying isomorphic relation patterns in clustered relation edge classes over time. Graph pattern matching algorithms are used to search for patterns with the same structure and type in a graph. These algorithms can identify relation paths with sequences of the same node type across different time windows. Relation pattern templates are templates representing isomorphic relation patterns, containing information about the structure and type of the relation.

[0116] When performing temporal isomorphic pattern recognition, the GraphMatch algorithm can be used for graph pattern matching. This algorithm searches for relational paths with the same sequence of node types in the semantic association network across different time windows. For example, in time window T1, there is a relational path from node A (type: technical concept) to node B (type: application scenario), and in time window T2, there is a relational path from node C (type: technical concept) to node D (type: application scenario). If the sequence of node types is the same, and the relational types are also the same, then these two relational paths can be considered isomorphic. The structure and type information of these isomorphic relational paths are then organized to generate relational pattern templates.

[0117] Step S334: Perform a temporal direction consistency check on the relation schema template to ensure that the start node type and end node type of all relation edges in the template are consistent in different time windows, thereby generating a directed relation schema.

[0118] Temporal direction consistency verification is the process of validating the direction of relation edges in a relation schema template along the time dimension. The purpose is to ensure that the start and end node types of all relation edges in the template maintain a consistent direction across different time windows, avoiding directional confusion. A directed relation schema, after undergoing temporal direction consistency verification, has a clearly defined direction.

[0119] When performing temporal direction consistency checks, the start and end node types of each relation edge in the relation schema template are examined across different time windows. For example, for a "causal" relation edge in a relation schema template, if it points from the technical concept node to the application scenario node in time window T1, it should also maintain the direction from the technical concept node to the application scenario node in time window T2. If inconsistencies occur, the cause is analyzed, which may be due to data recording errors or changes in the relation itself over time. For data recording errors, corrections are made; for changes in the relation itself, the specific business logic determines whether template adjustments are needed or whether it should be treated as a special case.

[0120] Step S335: Generate structure mapping rules based on the directed relation pattern. The antecedent of the structure mapping rule includes the temporal constraints of the original relation type and node type. The consequent of the structure mapping rule is the standardized relation type. The scope of application of the rule is set to the temporal scenario of matching the relation pattern template.

[0121] Generating structure mapping rules based on directed relation schemas is the process of transforming directed relation schemas into specific mapping rules. The antecedent of a structure mapping rule includes temporal constraints on the original relation type and node type, meaning that triggering the rule requires satisfying the temporal constraints of the original relation type and node type. The consequent of a structure mapping rule is a standardized relation type, i.e., a relation type that has been unified and standardized. The scope of application of the rule is set to the temporal scenarios that match the relation schema template, clearly defining the time scenarios in which the rule applies.

[0122] When generating structural mapping rules, the temporal constraints of the original relation types and node types are analyzed based on the structure and direction of the directed relation pattern. For example, for a directed relation pattern representing a "causal" relationship from a technical concept node to an application scenario node, the original relation type may be various similar causal expressions, such as "cause" or "initiate," with node types being technical concept and application scenario, respectively. These temporal constraints of the original relation types and node types are combined into rule antecedents, such as "when the original relation type is 'cause' or 'initiate,' and the starting node type is technical concept, and the ending node type is application scenario, within the time window T1-T3." The rule consequent is a standardized relation type, such as "causal." The rule's scope is set to the temporal scenarios matching the directed relation pattern; that is, the rule applies when a relation conforming to the directed relation pattern appears in the semantic association network.

[0123] Step S340: Perform temporal feature matching processing on the attribute tags in the semantic association network, identify the attribute set that can be reused across resource types by calculating the temporal similarity of attribute names, analyze the temporal distribution pattern of attribute values, and generate interface mapping rules containing temporal conversion rules of data types.

[0124] Temporal feature matching is the process of matching attribute labels in a semantic association network along the time dimension. Attribute name temporal similarity calculation calculates the similarity between attribute names at different times. This calculation identifies semantically similar attribute names across different resource types, thus finding attribute sets that can be reused across resource types. Attribute value temporal distribution patterns describe the distribution of attribute values ​​over time. Analyzing these patterns generates interface mapping rules that include data type temporal conversion rules. Interface mapping rules are mapping rules between interface resources, containing data type conversion rules at different times.

[0125] When performing temporal feature matching, for calculating the temporal similarity of attribute names, edit distance algorithms or word vector similarity calculation methods can be used. For example, the Word2Vec model can be used to convert attribute names into word vectors, and then the cosine similarity between the word vectors of attribute names at different times can be calculated. If the similarity is higher than a pre-set threshold, the two attribute names are considered semantically similar and can be included as part of an attribute set that can be reused across resource types. For analyzing the temporal distribution patterns of attribute values, statistical analysis methods can be used, such as calculating the mean, variance, and distribution histogram of attribute values ​​in different time windows. For example, for a numerical attribute, the changes in its mean and variance in different time windows can be analyzed to determine whether there is a data type conversion trend. If it is found that the attribute value changes from an integer to a decimal in a certain time window, or the value range changes significantly, data type temporal conversion rules can be generated based on these patterns. The antecedent of the generated interface mapping rule includes the temporal constraints of the original attribute name and data type, and the consequent of the rule is the standardized attribute name and the converted data type and value range. Through these rules, attributes of different interfaces can be uniformly mapped and converted, improving the compatibility and usability of interface resources.

[0126] Step S350: Integrate text mapping rules, structure mapping rules, and interface mapping rules; identify contradictory rule entries under different time windows using a rule conflict detection algorithm; resolve conflicts based on the temporal priority of the rule application scenario; and generate a cross-modal virtualization mapping rule set.

[0127] Integrating text mapping rules, structure mapping rules, and interface mapping rules involves merging these three different types of mapping rules into a unified rule set. The rule conflict detection algorithm is used to detect contradictions between different rule entries in the rule set. Contradictory rule entries refer to rules whose antecedents are similar or identical, but whose consequents differ, leading to uncertainty when applying the rules. The temporal priority of rule applicability scenarios is determined based on factors such as the rule's scope of application, accuracy, and importance at different times. Conflict resolution is then performed based on this priority, i.e., resolving contradictions between rules. The cross-modal virtualization mapping rule set is a collection containing all valid rules after integration and conflict resolution.

[0128] During integration, text mapping rules, structure mapping rules, and interface mapping rules are stored in a rule base according to a specific format. For rule conflict detection algorithms, a similarity calculation method based on rule feature vectors can be used. Each rule is converted into a multi-dimensional feature vector containing rule type, applicable conditions, confidence level, and time window label. Then, the similarity between any two rule feature vectors is calculated. If the similarity exceeds a pre-set threshold and the consequents of the rules are different, the two rules are considered to conflict. For example, for two rules, one rule's antecedent is "when concept A and concept B have a similarity greater than 0.8, map to concept C," and the other rule's antecedent is "when concept A and concept B have a similarity greater than 0.8, map to concept D." These two rules conflict. When resolving conflicts based on the temporal priority of rule application scenarios, the accuracy, coverage, and temporal decay factor of the rule's application frequency within the historical time window are comprehensively considered. High-priority rules are retained, while low-priority rules are marked as rules to be adjusted. For rules to be adjusted, conflict resolution can be achieved by narrowing the time window range of the rule's applicable conditions or lowering the rule confidence threshold. The corrected rules are then added back to the rule set. Finally, the rule synergy effect is evaluated on the rule set after conflict resolution. The correlation support between different types of rules is calculated, and rules with correlation support higher than the synergy threshold are grouped into rule groups to generate a cross-modal virtualization mapping rule set containing rule group information.

[0129] As one implementation method, step S350 can be specifically implemented as the following steps S351-S356:

[0130] Step S351: Construct a rule feature vector space, and convert text mapping rules, structure mapping rules and interface mapping rules into multi-dimensional feature vectors containing rule type, applicable conditions, confidence level and time window label.

[0131] The rule feature vector space is a space used to store rule feature vectors, representing different types of rules in vector form, which facilitates similarity calculation and conflict detection. Text mapping rules, structure mapping rules, and interface mapping rules are converted into multi-dimensional feature vectors. Each vector contains information such as rule type (e.g., text mapping, structure mapping, interface mapping), applicable conditions (e.g., concept similarity threshold, node type constraints), confidence level (the reliability of the rule), and time window marker (the time range within which the rule applies).

[0132] When constructing the rule feature vector space, the meaning and value range of each dimension are first defined. For example, rule type can be represented by different integer codes, applicable conditions can be numericalized, confidence level can be represented by a value between 0 and 1, and time window markers can be represented by start and end times. For each rule, its information is converted into a vector according to the defined dimensions. After converting all rules into such multi-dimensional feature vectors, they are stored in the rule feature vector space. This space provides the basic data for subsequent rule conflict detection, enabling rules to be compared and analyzed through vector operations.

[0133] Step S352: Calculate the similarity of rule entries in the feature vector space using a rule conflict detection algorithm to identify explicit conflict rule pairs where the antecedents are the same but the consequents are different, as well as implicit conflict rule pairs where the antecedents have an implication relationship but the consequents are contradictory.

[0134] Rule conflict detection algorithms are used to detect whether there are conflicts between rules. Similarity calculation calculates the similarity between rule entries in the feature vector space, and this calculation can identify potentially conflicting rule pairs. Explicitly conflicting rule pairs are those where the antecedents are the same but the consequents are different; these conflicts are relatively obvious and easy to identify. Implicitly conflicting rule pairs are those where the antecedents of rules have an implication relationship (e.g., the antecedent of one rule contains the antecedent of another rule), but the consequents contradict each other; these conflicts are relatively difficult to detect.

[0135] When performing rule conflict detection, the cosine similarity algorithm can be used to calculate the similarity between rule entries. For any two rule vectors in the feature vector space, the cosine similarity between them is calculated. If the similarity is higher than a pre-set threshold, the antecedent and consequent of the rule are further examined. For explicit conflict rule pairs, the antecedent and consequent of the rule are directly compared. If the antecedent is the same but the consequent is different, it is marked as an explicit conflict rule pair. For example, rule R1: "When concept A and concept B have a similarity greater than 0.8, map to concept C", and rule R2: "When concept A and concept B have a similarity greater than 0.8, map to concept D". These two rules are explicit conflict rule pairs. For implicit conflict rule pairs, the implication relationship of the antecedent of the rule is analyzed. For example, rule R3: "When concept A and concept B have a similarity greater than 0.7, map to concept E", and rule R4: "When concept A and concept B have a similarity greater than 0.8, map to concept F". Since 0.8 is included in the range of 0.7 and the consequent is different, these two rules have an implicit conflict. This method identifies all explicit and implicit conflicting rule pairs.

[0136] Step S353: Perform conflict intensity quantification on the identified conflict rule pairs, calculate the conflict intensity value based on the difference in rule confidence, the overlap of applicable scenarios and the overlap rate of time windows, and sort the conflict rule pairs in descending order of conflict intensity value.

[0137] Conflict intensity quantification is the process of quantifying the degree of conflict among identified conflicting rule pairs. Rule confidence difference refers to the difference in confidence levels between the two rules in a conflicting rule pair, reflecting the difference in their reliability. Applicable scenario overlap refers to the degree of overlap in the applicable scenarios of the conflicting rule pairs, and time window overlap rate refers to the proportion of overlap in the time windows of the conflicting rule pairs. The conflict intensity value is a numerical value calculated by comprehensively considering rule confidence difference, applicable scenario overlap, and time window overlap rate, used to measure the severity of the conflict. Conflicting rule pairs are sorted in descending order of conflict intensity value to facilitate prioritizing the processing of rule pairs with high conflict intensity.

[0138] When quantifying conflict intensity, for the difference in rule confidence, the absolute value of the difference between the confidence levels of two rules can be directly calculated. For the overlap of applicable scenarios, the proportion of samples satisfying the applicable conditions of two rules to the total number of samples can be calculated by analyzing the applicable conditions of the rules. For the overlap rate of time windows, the proportion of the overlap time length of two time windows to the total time length is calculated. For example, if the time window of rule R1 is T1-T2 and the time window of rule R2 is T3-T4, the ratio of the overlap time length (if there is overlap) to (T2-T1+T4-T3) is calculated. Then, these three factors are weighted and summed according to preset weights to obtain the conflict intensity value. After calculating the conflict intensity values ​​of all conflicting rule pairs, they are sorted in descending order, and rule pairs with high conflict intensity are processed first to improve the efficiency of conflict resolution.

[0139] Step S354: Conflict resolution is performed based on the temporal priority of the applicable scenario of the rule. The temporal priority comprehensively considers the accuracy, coverage and temporal decay factor of the rule in the historical time window and the applicable frequency. Rules with high priority are retained and rules with low priority are marked as rules to be adjusted.

[0140] The temporal priority of rule application scenarios is determined based on the rule's applicability at different times. It comprehensively considers the rule's accuracy (the proportion of correctly applied rules), coverage (the proportion of samples to which the rule applies out of the total sample size), and application frequency (the number of times the rule is applied) within a historical time window, along with a temporal decay factor. The temporal decay factor is a weighted factor applied to historical data, taking into account that the effectiveness of rules may decrease over time. Conflict resolution is performed based on this priority, retaining high-priority rules and marking low-priority rules as requiring adjustment.

[0141] When resolving conflicts, the temporal priority of each rule is first calculated. For the accuracy, coverage, and frequency of application within a historical time window, different weights are assigned to each window according to chronological order, with the weights gradually decreasing over time, forming a temporal decay factor. For example, the weight for data from the last three months is 0.6, for data from the last year it is 0.3, and for even older data it is 0.1. The accuracy, coverage, and frequency of application of each rule in different time windows are calculated, then multiplied by the corresponding temporal decay factor and weighted summed to obtain the rule's overall score, which serves as its temporal priority. For conflicting rule pairs, the temporal priorities of the two rules are compared; the rule with the higher priority is retained, and the rule with the lower priority is marked as a rule to be adjusted. For example, if rule R1 has a temporal priority of 0.8 and rule R2 has a temporal priority of 0.6, rule R1 is retained during conflict resolution, and rule R2 is marked as a rule to be adjusted. This method ensures that the retained rules are more effective in the time dimension, improving the overall performance of the rule set.

[0142] Step S355: Perform rule correction processing on the marked rules to be adjusted. Conflict resolution is achieved by narrowing the time window range of the rule's applicable conditions or lowering the rule confidence threshold. The corrected rules are then added back to the rule set.

[0143] Rule correction processing involves modifying marked rules to eliminate rule conflicts. Narrowing the time window for rule application reduces the timeframe in which the rule applies, decreasing the likelihood of conflicts with other rules. Lowering the rule confidence threshold reduces the rule's reliability, preventing it from being preferentially applied in case of conflict.

[0144] When revising rules, the time window range for narrowing the applicability of a rule can be adjusted based on the overlap of time windows between conflicting rules. For example, if rule R1 and rule R2 conflict and their time windows overlap, the time window of rule R2 (the rule to be revised) can be narrowed to a period that does not overlap with that of rule R1. The rule confidence threshold can be lowered based on the intensity of the conflict and the rule's historical performance. If the conflict intensity is high and the rule's historical accuracy and coverage are not particularly high, the rule confidence threshold can be appropriately lowered. For example, the confidence of rule R2 can be lowered from 0.9 to 0.7. The revised rule then undergoes a new feature vector transformation and is added to the rule feature vector space.

[0145] Step S356: Evaluate the rule synergy effect of the rule set after conflict resolution, calculate the association support between different types of rules, group the rules with association support higher than the synergy threshold into a rule group, and generate a cross-modal virtualization mapping rule set containing rule group information.

[0146] Rule synergy evaluation is the process of assessing the synergistic effect between different types of rules in a rule set after conflict resolution. Association support is an indicator that measures the degree of association between two rules, representing the proportion of samples that simultaneously satisfy both rules out of the total sample size. The synergy threshold is a pre-defined threshold used to determine whether rules have a synergistic effect; rule combinations with association support higher than the synergy threshold are called rule groups. The cross-modal virtualization mapping rule set is a rule set containing rule group information obtained after conflict resolution and rule grouping.

[0147] When evaluating the synergistic effect of rules, for each rule, its association support with other rules is calculated. This can be achieved by statistically analyzing the sample data to which the rules apply, identifying the number of samples that simultaneously satisfy two rules. For example, for rules R1 and R2, the number of samples that simultaneously satisfy both rules R1 and R2 is counted, and then divided by the total number of samples to obtain the association support. Rules with association support higher than the synergistic threshold are grouped into rule sets. For example, if the association support of rules R1, R2, and R3 are all higher than the synergistic threshold, they are grouped into a rule set. The resulting cross-modal virtualization mapping rule set not only includes individual rules that have undergone conflict resolution but also information on rule sets with synergistic effects. Such a rule set enables more effective virtualization annotation of heterogeneous technology project resources, improving the accuracy and efficiency of annotation.

[0148] Step S400: Perform virtualization annotation processing on the heterogeneous technology project resource set through the cross-modal virtualization mapping rule set to generate a virtualized annotation data set.

[0149] Virtualized annotation processing is the process of annotating heterogeneous sets of scientific and technological project resources using a cross-modal virtualized mapping rule set. It unifies the annotation of resources from different modalities according to the rules, giving them clearer semantics and structure. The virtualized annotated dataset is a dataset containing annotation information generated after annotation processing, transforming heterogeneous scientific and technological project resources into a more manageable and analyzable form.

[0150] During virtualization annotation, for each resource unit in the heterogeneous technology project resource set, rules from the cross-modal virtualization mapping rule set are applied sequentially. For text resources, original conceptual fields are replaced with standard concepts according to text mapping rules; for structural resources, original relation types are converted into standardized relation types according to structural mapping rules; for interface resources, original attributes are mapped to attribute labels containing temporal constraints according to interface mapping rules. During the annotation process, the identifiers of successfully matched rules and the mapping results are recorded for subsequent inspection and traceability.

[0151] As one implementation method, step S400 can be specifically implemented as the following steps S410-S450:

[0152] Step S410: Perform time-series rule matching processing on each resource unit in the heterogeneous technology project resource set, sequentially matching text mapping rules, structure mapping rules and interface mapping rules of different time windows, and recording the rule identifiers and mapping results of successful matches.

[0153] Temporal rule matching is the process of matching rules for each resource unit in a heterogeneous technology project resource set along the time dimension. It involves sequentially matching text mapping rules, structure mapping rules, and interface mapping rules across different time windows, matching resource units one by one according to the rule's time window and type. During temporal rule matching, for each resource unit, rules for the corresponding time window are selected from the cross-modal virtualization mapping rule set based on its time information. For text resources, the conceptual fields in the resource are matched against text mapping rules to check if the rule's antecedent conditions are met. For structure resources, the resource's relational structure is analyzed and matched against structure mapping rules. For interface resources, the resource's attribute information is checked and matched against interface mapping rules. Rules across different time windows are matched sequentially, and the identifiers and mapping results of all successfully matched rules are recorded.

[0154] Step S420: Based on the mapping results, perform virtualization feature generation processing, replace the original concept fields of the resource unit with the core evolution concept specified by the rule consequent, convert the original relation type into a standardized relation type, and map the original attributes into attribute labels containing temporal constraints.

[0155] Virtualization feature generation is the process of generating virtualization features based on the mapping results. It transforms the original information of resource units according to rules to generate virtualization features with unified semantics and structure. Replacing the original conceptual fields of resource units with the core evolutionary concepts specified by the rule consequent unifies different expressions of concepts into standard concepts; converting original relation types into standardized relation types unifies different relation structures into standardized relation types; and mapping original attributes to attribute labels containing temporal constraints normalizes and temporally constrains attribute information.

[0156] During virtualization feature generation, for text resources, records are matched according to the text mapping rules in the mapping results, and the original concept fields are replaced with the core evolutionary concepts specified by the rule consequents. For example, "machine learning algorithm" is replaced with "artificial intelligence algorithm". For structure resources, records are matched according to the structure mapping rules, and the original relation types are converted into standardized relation types. For example, "cause" relation is converted into "causal" relation. For interface resources, records are matched according to the interface mapping rules, and the original attribute key-value pairs are converted into attribute labels containing temporal constraints.

[0157] As one implementation method, step S420 can be specifically implemented as the following steps S421-S425:

[0158] Step S421: Parse the text mapping rule matching records in the mapping results, replace the original concept field of the resource unit with the core evolution concept specified by the rule consequent, and retain the values ​​of the original concept in different time windows as the time series annotation of the mapping source.

[0159] Parsing the text mapping rule matching records in the mapping results is the process of analyzing and processing the text mapping rule matching information of the records. Replacing the original concept fields of resource units with the core evolutionary concepts specified by the rule consequents unifies the concepts expressed differently into standard concepts, thereby improving the consistency of concepts.

[0160] During parsing, for each text mapping rule matching record, the core evolutionary concept specified by the rule consequent is extracted. For example, for the record "Rule identifier R1: When the similarity between the concepts 'machine learning algorithm' and 'artificial intelligence algorithm' is higher than 0.8, map to 'artificial intelligence algorithm'", "machine learning algorithm" in the resource unit is replaced with "artificial intelligence algorithm". At the same time, the values ​​of "machine learning algorithm" in different time windows are retained as time-series annotations of the mapping source.

[0161] Step S422: Parse the structure mapping rule matching records in the mapping results, convert the original relation type of the resource unit into the standardized relation type specified by the rule consequent, and adjust the relation direction to meet the requirements of the temporal direction attribute.

[0162] Parsing the structure mapping rule matching records in the mapping results involves analyzing and processing the structure mapping rule matching information of the records. Converting the original relation type of resource units into the standardized relation type specified by the rule consequent unifies different relation structures into a standardized relation type, facilitating unified management and analysis of resource structures. Adjusting the relation direction to conform to the temporal direction attribute requirements ensures that the direction of the relation at different times conforms to the rule definition, avoiding directional confusion.

[0163] During parsing, for each structure mapping rule matching record, the standardized relation type specified by the rule consequent is extracted. For example, for the record "Rule identifier R2: When the original relation type is 'cause', the starting node type is 'technical concept', and the ending node type is 'application scenario', convert it to a 'causal' relation," the "cause" relation in the resource unit is converted to a "causal" relation. Simultaneously, the direction of the relation is adjusted according to the rule's temporal direction attribute requirements. If the rule specifies that a "causal" relation points from a technical concept node to an application scenario node, and the relation direction in the resource unit is the opposite, then adjustments are made.

[0164] Step S423: Parse the interface mapping rule matching records in the mapping results, convert the original attribute key-value pairs into attribute labels containing time constraints, perform time-series unification processing on the units of numerical attributes, and perform time-series evolution processing on the standardized expression of text attributes.

[0165] Parsing the interface mapping rule matching records in the mapping results involves analyzing and processing the interface mapping rule matching information of the records. Converting the original attribute key-value pairs into attribute labels with temporal constraints standardizes and temporally constrains the attribute information, making the meaning of the attributes clearer. Time-series unification of units for numeric attributes unifies numeric attributes with different units into the same unit, facilitating data comparison and analysis. Time-series evolution processing of standardized expressions for textual attributes unifies textual attributes with different expressions into a standard expression, considering their changes over time.

[0166] During parsing, for each interface mapping rule matching record, the standardized attribute name, data type, and value range specified by the rule posterior are extracted.

[0167] Step S424: Perform cross-rule temporal consistency checks on the transformed concepts, relationships, and attributes to ensure that there are no logical contradictions between different mapping results of the same resource unit in different time windows, and generate intermediate virtualization features.

[0168] Cross-rule temporal consistency checking is the process of verifying whether the mapping results of transformed concepts, relationships, and attributes are consistent across different time windows, ensuring that there are no logical contradictions in the information of the same resource unit at different times. Intermediate virtualization features are preliminary virtualization features generated after cross-rule temporal consistency checking, which to a certain extent guarantee the consistency and logicality of information.

[0169] When performing cross-rule temporal consistency checks, the logical relationships between concepts, relations, and attributes are examined for mapping results across different time windows of the same resource unit. For example, for the concept "artificial intelligence algorithm" within a resource unit, which maps to "advanced artificial intelligence algorithm" in time window T1 and to "basic artificial intelligence algorithm" in time window T2, it's necessary to check whether this mapping conforms to the logical development over time. For relations, the direction and type of relations are checked for consistency across different time windows. For attributes, the reasonableness of attribute value ranges and temporal constraints is checked. If logical contradictions are found, the causes are analyzed and corrected. For example, if the contradiction is caused by rule conflicts, rule matching and mapping are re-performed. After checking and correction, intermediate virtualized features are generated to provide a reliable data foundation for subsequent temporal integrity supplementation.

[0170] Step S425: Perform temporal integrity supplementation processing on intermediate virtualization features. Assign values ​​to missing attributes based on the default value temporal rules in the rule set, supplement missing relationships through temporal inference of association rules, and generate a complete set of virtualization features.

[0171] Temporal integrity supplementation is the process of supplementing intermediate virtualization features to make them more complete in the time dimension. Assigning values ​​to missing attributes based on default values ​​in the rule set involves assigning default values ​​to missing attributes in resource units according to the rules, taking into account the time factor. Supplementing missing relationships through temporal inference of association rules involves inferring missing relationships based on existing relationship information and association rules, considering changes in relationships over different times. The complete virtualization feature set is the set of virtualization features containing complete information obtained after temporal integrity supplementation.

[0172] When performing temporal integrity supplementation, for missing attributes in intermediate virtualization features, the default value temporal rule in the rule set is searched. For example, if the attribute "parameter B" is missing in a resource unit, and the rule set has a default value temporal rule "In time window T1-T2, the default value of parameter B is 5", then the attribute is assigned a value of 5. For missing relationships, temporal inference is performed based on association rules. For example, if the relationship between "technical concept A" and "application scenario B" exists in a resource unit, and the association rule indicates that "technical concept A usually leads to the occurrence of application scenario B", then a "causal" relationship is supplemented. During the supplementation process, the time factor is fully considered to ensure that the supplemented attributes and relationships are reasonable in the time dimension. Through these processes, a complete set of virtualization features is generated, improving the integrity and availability of resources.

[0173] Step S430: Perform temporal correlation reconstruction on the resource units after virtualization feature generation and processing. Establish standardized temporal correlations between different resource units through the relation edge set, calculate the temporal change rate of correlation strength, and filter out weak temporal correlations below the strength change threshold.

[0174] Temporal association reconstruction is the process of rebuilding the associations between resource units after virtualization feature generation. It establishes standardized associations between different resource units in the time dimension. Establishing standardized temporal associations between different resource units through a set of relational edges involves using the previously generated set of relational edges to add standardized associations between resource units. Calculating the temporal change rate of association strength measures how the strength of the association changes over different times. Filtering out weak temporal associations below the strength change threshold removes those with weak and unstable association strength, thus improving the quality of the associations.

[0175] When reconstructing temporal relationships, associations are added between different resource units based on information in the relationship edge set. For example, if an edge in the relationship edge set represents a "causal" relationship connecting resource unit A and resource unit B, then a "causal" association is added to these two resource units in the virtualization features. For each association, the temporal rate of change of its association strength is calculated. This rate of change can be calculated by comparing the association strength values ​​across different time windows. A strength change threshold is set; if the temporal rate of change of the association strength is lower than this threshold, it is considered a weak temporal association and is filtered out. Through these processes, the temporal relationships between resource units are reconstructed, making the relationships clearer and more stable, facilitating knowledge analysis and mining.

[0176] Step S440: Encapsulate the reconstructed resource units in a temporal structure, organize the data according to a three-layer temporal structure of concept nodes-relationship edges-attribute labels, and generate virtualized labeled data units containing unique resource temporal identifiers.

[0177] Temporal structured encapsulation is the process of encapsulating reconstructed resource units according to a certain structure. It organizes data according to a three-layer temporal structure of concept nodes, relation edges, and attribute labels, giving the data a clearer hierarchy and structure. Generating virtualized labeled data units containing unique resource temporal identifiers involves assigning a unique identifier to each encapsulated data unit, facilitating data management and querying.

[0178] When performing time-series structured encapsulation, concepts within a resource unit are treated as concept nodes, relationships as relationship edges, and attribute information as attribute labels. This information is organized into a three-layer structure according to chronological order. For example, for a resource unit containing the concepts "artificial intelligence technology" and "intelligent medical application," the relationship "applied to," and the attribute "technology maturity (high)," "artificial intelligence technology" and "intelligent medical application" are treated as concept nodes, "applied to" as relationship edges, and "technology maturity (high)" as attribute labels, organized chronologically.

[0179] Step S450: Perform time-series consistency verification on multiple virtualized annotation data units, identify conflicting annotation items by comparing the attribute values ​​across resource units, correct conflicts based on rule-based time-series priority, and combine the corrected annotation data units according to time windows to generate a virtualized annotation data set.

[0180] Temporal consistency verification is the process of checking whether multiple virtualized annotation data units are consistent in the time dimension. It identifies conflicting annotations by comparing the temporal values ​​of attribute values ​​across resource units, i.e., comparing the changes in attribute values ​​of different resource units at different times to find conflicting annotations. Conflict correction based on rule temporal priority involves correcting conflicting annotations according to the temporal priority of the rules to ensure data consistency. The corrected annotation data units are then combined according to time windows to generate a virtualized annotation data set. This involves combining the verified and corrected data units in chronological order to form the final virtualized annotation data set. During temporal consistency verification, for multiple virtualized annotation data units, the changes in their attribute values ​​across different time windows are compared. For example, for the attribute "technical parameter X" in two resource units, if one resource unit has a value of 10 and the other has a value of 20 in time window T1, this may constitute a conflict. By comparing the temporal values ​​of attribute values ​​across resource units, all conflicting annotations are identified. Then, the conflicting annotations are corrected according to the rule temporal priority. If a rule specifies a higher priority, the correction is performed according to that rule. For example, if the rule stipulates that "when attribute values ​​conflict, the more recent data shall prevail," then the value of the more recent data shall be used. After correction, the labeled data units are combined according to time windows to generate a virtualized labeled data set.

[0181] Step S500: Construct an evolutionary heterogeneous resource knowledge base based on the virtualized labeled data set. The evolutionary heterogeneous resource knowledge base includes concept units, relation units, and attribute units.

[0182] An evolutionary heterogeneous resource knowledge base is a knowledge base that reflects the evolution of heterogeneous resources over time. It includes conceptual units (such as various concepts in science and technology projects), relational units (the relationships between concepts), and attribute units (the attribute information of concepts). Building this knowledge base based on a virtualized labeled dataset involves integrating and organizing the data after virtualization labeling to form a knowledge base with a clear structure and rich information.

[0183] When constructing an evolutionary heterogeneous resource knowledge base, the virtualized labeled data set is first classified and organized. Conceptual information is extracted from the data to form conceptual units; relational information is extracted to form relational units; and attribute information is extracted to form attribute units. Then, these units are organized according to a predetermined structure. For example, a graph structure can be used, with conceptual units as nodes, relational units as edges, and attribute units as attributes of the nodes. During the organization process, the time factor is fully considered, recording the state and changes of each unit at different times.

[0184] As one implementation method, step S500 can be specifically implemented as the following steps S510-S550:

[0185] Step S510: Perform knowledge unit temporal division processing on the virtualized labeled data set, and decompose the labeled data of different time windows into basic temporal knowledge units according to the three types of features: concept, relationship and attribute.

[0186] Knowledge unit temporal partitioning is the process of dividing the virtualized labeled data set along the time dimension. Labeled data from different time windows is decomposed according to three categories of features: concepts, relationships, and attributes, forming basic temporal knowledge units. Each basic temporal knowledge unit contains concept, relationship, or attribute information within its corresponding time window.

[0187] When performing temporal partitioning of knowledge units, each virtualized labeled data unit in the virtualized labeled dataset is decomposed into three categories based on the information it contains: concepts, relationships, and attributes. For example, for a data unit containing the concept "artificial intelligence algorithm," the relationship "applied to," and the attribute "algorithm complexity (high)," "artificial intelligence algorithm" is treated as a conceptual knowledge unit, "applied to" as a relational knowledge unit, and "algorithm complexity (high)" as an attribute knowledge unit. Simultaneously, the time window information for each knowledge unit is recorded. This partitioning method decomposes the virtualized labeled dataset into basic temporal knowledge units, providing foundational data for subsequent standardized encapsulation.

[0188] Step S510: Perform knowledge unit temporal division processing on the virtualized labeled data set, and decompose the labeled data of different time windows into basic temporal knowledge units according to the three types of features: concept, relationship and attribute.

[0189] The temporal partitioning of knowledge units aims to meticulously divide the virtualized labeled data set based on the time dimension. This process decomposes the labeled data under different time windows according to three types of features: concepts, relationships, and attributes, thereby forming basic temporal knowledge units. Each basic temporal knowledge unit accurately covers the concept, relationship, or attribute information within the corresponding time window, which helps to organize and manage the knowledge more systematically in the future.

[0190] When performing time-series partitioning of knowledge units, each virtualized labeled data unit in the virtualized labeled dataset is decomposed into three categories—concept, relation, and attribute—based on its contained information. For example, for a data unit containing the concept "blockchain technology," the relation "promotion," and the attribute "transaction efficiency (high)," "blockchain technology" is treated as a conceptual knowledge unit, "promotion" as a relational knowledge unit, and "transaction efficiency (high)" as an attribute knowledge unit. Simultaneously, the time window information corresponding to each knowledge unit is recorded in detail. This partitioning method decomposes the originally complex virtualized labeled dataset into basic time-series knowledge units, laying a solid data foundation for subsequent standardization and encapsulation work. Through this partitioning, the specific composition of knowledge within each time window can be clearly seen, providing a clear framework for further knowledge base construction, enabling the knowledge base to more accurately reflect the state and changes of heterogeneous resources at different times.

[0191] Step S520: Standardize the basic temporal knowledge units by encapsulating them into a time sequence, assigning a unique knowledge ID and time window marker to each knowledge unit, adding source resource identifiers and generation timestamps, and generating structured temporal knowledge units.

[0192] Standardized temporal encapsulation further processes basic temporal knowledge units, giving them a unified format and structure. Each knowledge unit is assigned a unique knowledge ID and a time window marker, facilitating identification and management. The knowledge ID allows for quick location of specific knowledge units, while the time window marker clearly defines the time range within which the knowledge unit exists. Source resource identifiers and generation timestamps are added. The source resource identifier traces the original source of the knowledge unit, and the generation timestamp records the specific time the knowledge unit was generated, helping to understand the timeliness and evolution of the knowledge. Through these operations, structured temporal knowledge units are generated, making the knowledge units more standardized and orderly.

[0193] When performing standardized temporal encapsulation, a unique knowledge ID is first generated for each basic temporal knowledge unit. A UUID (Universally Unique Identifier) ​​algorithm can be used to generate the knowledge ID, ensuring its uniqueness. For time window marking, the corresponding time information is directly obtained from the division of the basic temporal knowledge unit. The source resource identifier can be extracted from the original information of the virtualized labeled data unit, recording which specific resource the knowledge unit comes from. The generated timestamp can use the current system time to record the encapsulation time of the knowledge unit.

[0194] Step S530: Perform association indexing and time-series construction processing on the structured time-series knowledge units, establish bidirectional time-series references between concept knowledge units and relational knowledge units, and establish association time-series pointers between relational knowledge units and attribute knowledge units.

[0195] The temporal construction process of the association index is the process of establishing associations and indexes between structured temporal knowledge units, taking into account the time factor. Establishing bidirectional temporal references between conceptual knowledge units and relational knowledge units means that conceptual knowledge units can reference related relational knowledge units, and vice versa, and this referencing is time-dependent. Establishing temporal pointers between relational knowledge units and attribute knowledge units allows relational knowledge units to point to related attribute knowledge units, and clarifies the nature of this association at different times.

[0196] When constructing time-series association indexes, hash tables can be used to implement bidirectional time-series references between conceptual knowledge units and relational knowledge units. For example, for a conceptual knowledge unit "cloud computing technology," which is associated with the relational knowledge unit "applied to" in a certain time window T5-T6, a reference to the relational knowledge unit is added to the storage structure of the conceptual knowledge unit, and a reference to the conceptual knowledge unit "cloud computing technology" is added to the storage structure of the relational knowledge unit, recording the time window information. For time-series pointers between relational knowledge units and attribute knowledge units, linked lists can be used. For example, if the relational knowledge unit "depends on" is associated with the attribute knowledge unit "dependency (high)," a pointer to the attribute knowledge unit is added to the relational knowledge unit, recording the association time.

[0197] Step S540: The knowledge units after association indexing are processed by storage structure temporal organization. Concept nodes and relation edges of different time windows are stored in a graph database, attribute tags are stored in a temporal relation database, and cross-database temporal association is achieved through knowledge ID and time window marker.

[0198] Time-series storage structure organization is the process of designing and organizing the storage structure of knowledge units after association indexing, fully considering the time factor. Graph databases are suitable for storing concept nodes and relation edges, and can intuitively display the relationship structure between knowledge units. Time-series relational databases are suitable for storing attribute tags because attribute tags usually have structured data and can easily handle time-related information. Cross-database time-series associations are achieved through knowledge IDs and time window markers, enabling knowledge units in different databases to be associated and queried along the time dimension.

[0199] When organizing storage structures in a time-series manner, graph databases such as Neo4j can be chosen. Conceptual knowledge units are treated as nodes in the graph, and relational knowledge units as edges. Each node and edge carries a knowledge ID and a time window marker. For example, the concept node "IoT technology" and the relation edge "connected to" both have corresponding knowledge IDs and time window information. For time-series relational databases, products such as TimescaleDB can be used. Attribute knowledge units are stored in the time-series relational database, with each attribute record containing a knowledge ID and a time window marker. For example, the attribute "data transmission rate (100Mbps)" is stored in the time-series relational database and associated with a corresponding knowledge ID and time window. A connection is established between the graph database and the time-series relational database using the knowledge ID and time window marker. For example, when querying the attribute of a concept node within a target time window, a related query can be performed in both databases using the knowledge ID and time window marker. This storage structure effectively utilizes the advantages of different databases, achieving efficient storage and association of knowledge units in the time dimension.

[0200] Step S550: Perform temporal integrity verification on the organized knowledge units, detect isolated nodes and missing relationships in different time windows through knowledge graph reasoning algorithms, perform automatic temporal completion based on cross-modal virtualization mapping rules, and generate an evolutionary heterogeneous resource knowledge base containing concept units, relation units and attribute units.

[0201] Temporal integrity verification is the process of checking the integrity of organized knowledge units along the time dimension. Knowledge graph reasoning algorithms can be used to detect isolated nodes (i.e., nodes that have not established relationships with other nodes) and missing relationships (relationships that should exist but have not been recorded) in different time windows. Automatic temporal completion based on cross-modal virtualization mapping rules supplements the detected isolated nodes and missing relationships using previously generated cross-modal virtualization mapping rules, taking into account the time factor to ensure that the supplemented knowledge is reasonable in the time dimension. Finally, an evolutionary heterogeneous resource knowledge base containing concept units, relationship units, and attribute units is generated, which can accurately reflect the evolution of heterogeneous resources over time.

[0202] When performing temporal integrity verification, rule-based reasoning algorithms, such as Datalog rule reasoning, can be used for knowledge graph reasoning. This involves traversing the nodes and edges in the graph database to check for isolated nodes and missing relationships. For example, if a concept node "quantum computing technology" has no associated edges within a certain time window, it is marked as an isolated node. For detecting missing relationships, reasoning can be performed based on known relationship patterns and cross-modal virtualization mapping rules. For instance, if "artificial intelligence technology" is known to be "applied" to multiple fields, and no relationship with a specific application field is recorded within a certain time window, a missing relationship is inferred. When performing automatic temporal completion based on cross-modal virtualization mapping rules, appropriate associations are found according to the rules.

[0203] Please see Figure 2 , Figure 2 This is a schematic diagram of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or Central Processing Unit, CPU) is the computing and control core of the computer system, capable of parsing various instructions and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system used to store programs and data. It is understood that the memory 103 here can include the computer system's built-in memory, or it can include extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system; this invention does not limit this storage space.

[0204] In one embodiment, the processor 101 executes the heterogeneous resource virtualization annotation and knowledge base creation method provided above in the embodiments of the present invention by running a computer program in the memory 103.

Claims

1. A method for heterogeneous resource virtualization annotation and knowledge base creation, characterized in that, The method comprises: Step S100: acquiring a plurality of types of heterogeneous technology project resource sets, performing cross-modal semantic feature extraction operations on the heterogeneous technology project resource sets, and generating a cross-modal resource semantic feature set; Step S200: constructing a semantic association network based on the cross-modal resource semantic feature set, the semantic association network comprising semantic nodes, association edges, and evolution weights; Step S300: performing cross-modal virtualization mapping rule mining operations through the semantic association network, generating a cross-modal virtualization mapping rule set, and the cross-modal virtualization mapping rule set comprising text mapping rules, structure mapping rules, and interface mapping rules; Step S400: performing virtualization annotation processing on the heterogeneous technology project resource sets through the cross-modal virtualization mapping rule set, and generating a virtualization annotation data set; Step S500: constructing an evolution-type heterogeneous resource knowledge base based on the virtualization annotation data set, and the evolution-type heterogeneous resource knowledge base comprising concept units, relationship units, and attribute units; Step S300 comprises: performing time sequence feature vector extraction processing on concept nodes in the semantic association network, mapping concept nodes in different time windows into low-dimensional dense time sequence vectors, calculating cosine similarity between time sequence vectors to generate a concept similarity matrix; Performing evolution similar concept group identification processing based on the concept similarity matrix, filtering concept node clusters with similarity higher than a threshold in different time windows, performing time sequence hierarchical relationship analysis on nodes in the clusters to determine core evolution concepts, and generating text mapping rules comprising the core evolution concepts and subordinate evolution concepts; Performing time sequence type matching processing on association edges in the semantic association network, extracting time sequence type identifiers and association node pair sequences of the association edges, identifying isomorphic relationship patterns in different time windows, and combining node time sequence attribute labels to generate structure mapping rules comprising relationship time sequence conversion directions; Performing time sequence feature matching processing on attribute labels in the semantic association network, identifying an attribute set that can be cross-resource-type reused through attribute name time sequence similarity calculation, and analyzing attribute value time sequence distribution rules to generate interface mapping rules comprising data type time sequence conversion rules; Integrating the text mapping rules, the structure mapping rules, and the interface mapping rules, identifying mutually contradictory rule items in different time windows through a rule conflict detection algorithm, performing conflict resolution processing based on time sequence priorities of rule application scenarios, and generating a cross-modal virtualization mapping rule set; The time sequence type matching processing on the association edges in the semantic association network comprises: Performing time sequence feature extraction on the association edges in the semantic association network to generate relationship feature vectors; Performing time sequence type clustering processing on the relationship feature vectors, clustering relationship edges of the same type in different time windows into a class, calculating time sequence variance values of relationship strengths in the class, and deleting abnormal relationship edges with time sequence variance values higher than a threshold; Performing time sequence isomorphic pattern identification processing on the clustered relationship edge classes, identifying relationship paths with the same node type sequences in different time windows through a graph pattern matching algorithm, and generating relationship pattern templates; The relationship pattern template is subjected to a time direction consistency check, so that the starting node type and the ending node type of all relationship edges in the template are consistent in different time window directions, and a directed relationship pattern is generated; A structure mapping rule is generated based on the directed relationship pattern, a premise of the structure mapping rule contains a time constraint of an original relationship type and a node type, a conclusion of the structure mapping rule is a standardized relationship type, and a rule application range is set as a time scene matching the relationship pattern template.

2. The method of claim 1, wherein, The semantic association network is constructed based on the cross-modal resource semantic feature set, and includes the following steps: Concept extraction processing is performed on text semantic features in the cross-modal resource semantic feature set, concept units evolving over time are extracted, semantic drift detection processing is performed on the concept units, and a concept node set with time labels is generated; Relationship identification processing is performed on structure semantic features in the cross-modal resource semantic feature set, a type and a direction attribute of an association between concept units are determined, a relationship strength value is calculated based on a time sequence occurrence frequency of the type, and an association edge set is generated based on the type, the direction attribute, and the relationship strength value; Attribute extraction processing is performed on interface semantic features in the cross-modal resource semantic feature set, attribute names and value range sequences corresponding to the concept units are extracted, and an attribute label set containing time constraints is generated; The concept node set, the association edge set, and the attribute label set are input into a network construction model, a time-varying mapping relationship between the concept nodes and the association edges is established, time binding is performed on the attribute labels and the concept nodes, and an initial semantic association network is generated; Network structure optimization operations are performed on the initial semantic association network, closely connected concept node groups in different time windows are identified, redundant association edges with a time-varying connection strength lower than a threshold value in the groups are deleted, core association edges with a time-varying connection strength higher than the threshold value across the groups are retained, and an optimized semantic association network is obtained.

3. The method of claim 2, wherein, The relationship identification processing on the structure semantic features in the cross-modal resource semantic feature set includes the following steps: Time sequence syntax dependency analysis processing is performed on the structure semantic features, a relationship phrase structure is identified through a time sequence dependency tree, a time sequence change sequence of a relationship core verb is extracted, and a relationship verb set is generated; Time sequence semantic extension processing is performed on each relationship verb in the relationship verb set, a synonym set in different time windows is obtained through time sequence word vector similarity calculation, a standardized relationship type set is generated based on the synonym set and similar relationship types, and a relationship directed graph containing time sequence direction labels is generated through time sequence direction attribute determination processing on each relationship type in the standardized relationship type set; Time sequence strength calculation is performed on each directed edge in the relationship directed graph, a relationship strength value is generated by combining a time sequence importance weight of an associated concept and a change rate of an occurrence frequency of the same relationship type in different time window resource sets, and a relationship feature tuple is constructed based on the standardized relationship type, the time sequence direction label, and the relationship strength value. Multiple relationship feature tuples are combined to generate an association edge set according to time windows. ​ 4. The method of claim 1, wherein, The evolution similar concept group recognition processing based on the concept similarity matrix includes: The concept similarity matrix is subjected to time sequence threshold screening processing, a similarity lower limit threshold changing with a time window is set, concept node pairs with similarity values higher than the corresponding threshold in different time windows in the matrix are retained, and an initial similar node pair set is generated; The initial similar node pair set is subjected to time sequence clustering processing, node pairs in different time windows are merged into concept node clusters, and average similarity values of all nodes in the clusters in each time window are calculated as cluster time sequence consistency indexes; The concept node clusters are subjected to time sequence consistency verification processing, node clusters with cluster time sequence consistency indexes higher than a verification threshold in each time window are screened as candidate evolution similar concept groups, and time sequence noise nodes in the candidate evolution similar concept groups are detected, and abnormal nodes with similarity values lower than cluster average values in continuous time windows are deleted; The evolution similar concept groups after purification are subjected to time sequence hierarchical relationship analysis processing, a core evolution concept is identified as a concept node with the highest connection degree in different time windows in the group through time sequence centrality calculation, the remaining nodes are subordinate evolution concepts, and a core-subordinate time sequence relationship structure is generated; A text mapping rule is generated based on the core-subordinate time sequence relationship structure, a premise of the text mapping rule includes subordinate evolution concepts and a time sequence similarity threshold condition, a consequent of the text mapping rule is the core evolution concept, and a rule confidence is set as a weighted average value of the cluster time sequence consistency indexes.

5. The method of claim 1, wherein, The virtualization annotation processing of the heterogeneous scientific and technological project resource set through the cross-modal virtualization mapping rule set includes: Time sequence rule matching processing is performed on each resource unit in the heterogeneous scientific and technological project resource set, text mapping rules, structure mapping rules and interface mapping rules in different time windows are matched in turn, and a rule identifier and a mapping result of successful matching are recorded; Virtualization feature generation processing is performed based on the mapping result, an original concept field of the resource unit is replaced with a core evolution concept specified in a consequent of a rule, an original relationship type is converted into a standardized relationship type, and an original attribute is mapped into an attribute label containing a time sequence constraint; Time sequence association relationship reconstruction processing is performed on the resource unit after the virtualization feature generation processing, a relationship edge set is used to establish standardized time sequence associations between different resource units, a time sequence change rate of an association strength is calculated, and weak time sequence associations with a strength change threshold lower than the change rate are filtered; The reconstructed resource unit is subjected to time sequence structured packaging, data is organized in a three-layer time sequence structure of a concept node-relationship edge-attribute label, and a virtualization annotation data unit containing a unique resource time sequence identifier is generated; Time sequence consistency verification is performed on a plurality of virtualization annotation data units, conflict annotation items are identified through time sequence comparison of attribute values across resource units, conflict correction is performed based on rule time sequence priorities, and the corrected annotation data units are combined into a virtualization annotation data set according to time windows.

6. The method of claim 2, wherein, The attribute extraction processing of the interface semantic feature in the cross-modal resource semantic feature set includes: The interface semantic features are subjected to time sequence key-value pair identification, and the combination mode of attribute name and value in different time windows is matched through a time sequence regular expression to extract an original attribute key-value pair set; The original attribute key-value pair set is subjected to attribute name time sequence standardization, and the same attribute names in different time windows are unified into standard attribute names through time sequence matching of an attribute dictionary to generate a standardized attribute key-value pair; The attribute values in the standardized attribute key-value pair are subjected to time sequence data type identification processing, and the time sequence change of the data type is determined through time sequence analysis of the value format to generate an attribute description tuple containing a time sequence marker of the data type; The attribute description tuple is subjected to value range time sequence inference, and the time sequence change of the extreme value and mean value of a numerical attribute in different time windows is calculated through statistical analysis to identify the time sequence evolution of a value set of a text attribute to generate an attribute constraint condition; The standard attribute name, data type time sequence marker and attribute constraint condition are combined to generate an attribute label, and a plurality of attribute labels are grouped according to the concept node and time window to generate an attribute label set.

7. The method of claim 5, wherein, The virtualization feature generation processing based on the mapping result comprises: The text mapping rule matching record in the mapping result is parsed, the original concept field of the resource unit is replaced with the core evolution concept specified in the rule consequent, and the value of the original concept in different time windows is retained as a mapping source time sequence annotation; The structure mapping rule matching record in the mapping result is parsed, the original relationship type of the resource unit is converted into the standardized relationship type specified in the rule consequent, and the relationship direction is adjusted to meet the time sequence direction attribute requirement; The interface mapping rule matching record in the mapping result is parsed, the original attribute key-value pair is converted into an attribute label containing a time sequence constraint, the unit of a numerical attribute is subjected to time sequence uniform processing, and the text attribute is subjected to time sequence evolution processing of standardized expression; The converted concept, relationship and attribute are subjected to cross-rule time sequence consistency checking, so that there is no logical contradiction between different mapping results of the same resource unit in different time windows, and an intermediate virtualization feature is generated; The intermediate virtualization feature is subjected to time sequence integrity supplement processing, default value time sequence rules in the rule set are used to assign values to missing attributes, and missing relationships are supplemented through time sequence inference of the association rule to generate a complete virtualization feature set.

8. The method of claim 1, wherein, The evolution-type heterogeneous resource knowledge base is constructed based on the virtualization labeled data set, comprising: The virtualization labeled data set is subjected to knowledge unit time sequence division processing, and the labeled data in different time windows is decomposed into basic time sequence knowledge units according to the concept, relationship and attribute three types of features; The basic time sequence knowledge units are subjected to standardized time sequence packaging processing, a unique knowledge ID and a time window marker are assigned to each knowledge unit, a source resource identifier and a generation timestamp are added, and a structured time sequence knowledge unit is generated; The structured time sequence knowledge unit is subjected to association index time sequence construction processing, bidirectional time sequence references of the concept knowledge unit and the relationship knowledge unit are established, and an association time sequence pointer of the relationship knowledge unit and the attribute knowledge unit is established. The knowledge units after association indexing are stored in a time sequence organization process, the concept nodes and relationship edges of different time windows are stored in a graph database, the attribute labels are stored in a time sequence database, and cross-database time sequence association is realized through knowledge ID and time window marking; The organized knowledge units are subjected to time sequence integrity checking, isolated nodes and missing relationships of different time windows are detected through knowledge graph reasoning algorithm, automatic time sequence completion is performed based on cross-modal virtualization mapping rules, and an evolved heterogeneous resource knowledge base containing concept units, relationship units and attribute units is generated.

9. A computer system, characterized by It comprises: a memory in which a computer program is stored; a processor for loading the computer program to realize the heterogeneous resource virtualization labeling and knowledge base creation method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Multi-source knowledge processing and querying method and device, equipment and medium

    CN120197681A

  • Intelligent agent interpretable retrieval path generation system and verification method

    CN120654839A