Cross-disciplinary Scientific Research Potential Evaluation Method Based on Dynamic Multimodal Knowledge Graph

Through dynamic multimodal knowledge graphs and online learning technology, the hot spots of emerging disciplines are identified in real time, and the problem of difficulty in capturing interdisciplinary changes in traditional evaluation methods is solved, and dynamic optimization and accurate evaluation of scientific research resources in colleges and universities is achieved.

CN120181409BActive Publication Date: 2025-07-22GUANGDONG UNIVERSITY OF FOREIGN STUDIES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510665367.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-22
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

Traditional static assessment methods are difficult to capture the dynamic changes in emerging discipline hotspots and interdisciplinary cooperation in time, resulting in lag or mismatch in resource allocation and discipline layout in colleges and universities, affecting the accuracy and competitiveness of decision-making.

Method used

The interdisciplinary scientific research potential evaluation method based on dynamic multimodal knowledge graph is adopted, and the graph structure is updated in real time through technical means such as concept drift detection, online learning and transfer learning, and situational simulation, and the map structure is updated in real time, emerging hot spots are identified and evaluation models are optimized to form a closed-loop evaluation system.

Benefits of technology

Real-time capture of new academic concepts and sudden hot spots has been achieved, and the sensitivity and accuracy of interdisciplinary research has been improved, so as to help universities adjust resource allocation in a timely manner, and maintain the adaptability of the evaluation model and forward-looking decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181409B_ABST
    Figure CN120181409B_ABST
Patent Text Reader

Abstract

The present invention discloses an interdisciplinary scientific research potential evaluation method based on a dynamic multi-modal knowledge graph, which relates to the technical field of scientific research management and intelligence analysis. In response to the dynamic needs of the university scientific research ecosystem, relying on an incremental evaluation framework, it includes multi-source data collection and incremental knowledge graph construction, time series trend recognition, online and transfer learning optimization, scenario simulation and multi-scenario strategy verification, and multi-modal data fusion. The graph structure is updated in real time through concept drift detection and ontology extension, and emerging hotspots are identified by combining time series differential indicators such as citation growth rate and cooperation network. And with the help of online learning and cross-domain transfer strategies, the accuracy and continuous iteration ability of the evaluation model are maintained in multiple environments. Finally, the potential evolution is quantified through scenario simulation and strategy suggestions are output, supplemented by multi-modal embedding to further enhance the sensitivity to cross-disciplinary fields, forming a closed-loop evaluation system from data to decision-making to assist the interdisciplinary layout and resource allocation of universities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of scientific research management and intelligence analysis, and specifically to a method for cross-disciplinary scientific research potential evaluation based on a dynamic multi-modal knowledge graph. Background Art

[0002] In today's university scientific research environment, with the continuous advancement of digitalization and informatization, a large number of academic achievements emerge in various forms such as papers, patents, reports, and datasets. Researchers also conduct cross-field exchanges and knowledge sharing through diverse channels such as online academic conferences, social media, and video forums. Especially during the rapid rise of emerging disciplines and interdisciplinary fields, changes in funding support, orientation preferences, and social needs can lead to frequent evolution of the cooperation networks and topic layouts among different research institutions. For example, certain laboratories may achieve breakthroughs in the combination of biology and information engineering, or suddenly appear interdisciplinary research projects targeting public health emergencies; these dynamic factors make the scientific research ecosystem highly non-linear and decentralized. If relying solely on traditional text mining and annual aggregated literature databases to track the development of disciplines, it is often impossible to timely capture emerging hotspots and difficult to reflect the true context of scientific research activities in a timely manner; this not only affects the strategic deployment of universities in frontier disciplines but may also lead to missed funding opportunities or technology cooperation opportunities.

[0003] Therefore, the industry urgently needs a technical solution that can perform incremental acquisition and comprehensive processing on different sources (text, video, patent descriptions, etc.), realize real-time discovery and forward-looking evaluation of interdisciplinary potential, and accurately depict the new dynamics in the scientific research field under the drive of a complex social environment and multiple guiding factors.

[0004] However, when faced with these multi-source heterogeneous data, how to ensure the update speed while taking into account the accurate identification of new academic concepts and mutation hotspots has become a core technical problem. Usually, information systems need to transform a large amount of academic data in different modalities into computable knowledge structures and track the evolution of disciplinary themes or collaboration relationships in the time dimension. However, due to the irregular update cycle of scientific research information and the possible explosive growth or rapid decline of cross-domain cooperation in the short term, traditional knowledge graphs or static evaluation models constructed once are difficult to warn in a timely manner about these rapidly changing research trends; when the managers of research institutions or universities have not yet perceived the potential of interdisciplinary research, the rapidly changing disciplinary landscape has already been reorganized, resulting in lags or mismatches in resource allocation, disciplinary support, talent introduction, etc. Without an effective mechanism for integrating dynamic incremental data and a sensitive detection of sudden trends, the constructed evaluation system will be disconnected from the real scientific research ecosystem, ultimately leading to problems such as distorted decision-making information, uneven funding distribution, and damaged competitiveness. Therefore, how to conduct multi-modal and incremental data collection and analysis of cross-domain research activities in a short time and map this dynamic change to the evaluation of scientific research potential in real time has become a key technical problem that urgently needs to be solved. Summary of the Invention

[0005] (I) Technical problems to be solved

[0006] In view of the deficiencies of the prior art, the present invention provides a method for evaluating the interdisciplinary scientific research potential based on a dynamic multi-modal knowledge graph. By detecting concept drift and expanding the ontology, the graph structure is updated in real time. Combining time series difference metrics such as citation growth rate and cooperation network to identify emerging hotspots; and using online learning and cross-domain transfer strategies to maintain the accuracy and continuous iteration ability of the evaluation model in multiple environments. Finally, through scenario simulation, the potential evolution is quantified and strategic suggestions are output, supplemented by multi-modal embedding to further enhance the sensitivity to cross-disciplinary fields, forming a closed-loop evaluation system from data to decision-making, helping universities with interdisciplinary layout and resource allocation, and solving the technical problems described in the background art.

[0007] (II) Technical solutions

[0008] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for evaluating the interdisciplinary scientific research potential based on a dynamic multi-modal knowledge graph, including, if the amount of newly crawled academic data exceeds the threshold, text mining is used to extract entities and relationships, and after identifying emerging terms through concept drift detection, they are incrementally merged into the incremental knowledge graph, and then the ontology level is expanded to accommodate the new concepts;

[0009] After reaching a new time slice, calculate the multi-time window difference vector and aggregate it to highlight the theme growth or convergence trend, generate a cross-trend index and a list of hot topics and mark sudden rises and falls;

[0010] When the cross-trend indicator and the hot topic list show significant fluctuations within the domain, online parameter update is performed to optimize the model parameters, and the migration loss function is used to adaptively fine-tune the model parameters in different regions;

[0011] When the optimized model is obtained and the management end submits a scenario composed of external environment parameters and scientific research orientation variables, the scenario potential function is called to evaluate the potential trend of resource allocation and interdisciplinary collaboration, output comparable indicators, and when the difference between the actual and predicted values is greater than the expected value, a deviation signal is fed back to the online learning module;

[0012] Perform modal marking and semantic vector extraction on the generated multi-modal original dataset, align the cross-modal features by the aggregation function and output the multi-modal fusion library, incrementally integrate it into the multi-modal knowledge graph, and trigger the re-fine-tuning of the optimized model when anomalies occur.

[0013] Furthermore, after establishing a timed or real-time crawling strategy for multi-source data, set the crawling frequency and filtering conditions to obtain a stage-based original dataset; use text mining algorithms to perform entity annotation on the common structures in the academic fields in the stage-based original dataset, and preliminarily merge duplicate entities or name variants to obtain an entity set, and extract the association information between the documents to form a preliminary relationship set.

[0014] Furthermore, incrementally merge the entity set and the relationship set into the incremental knowledge graph; if overlap with existing nodes or attribute conflicts are detected, a merging strategy based on the confidence threshold is adopted for correction; perform concept drift calculation for nodes and edges in the incremental knowledge graph,

[0015] If the obtained concept drift measure exceeds the preset drift threshold, use it as a new concept node and trigger the corresponding ontology expansion or hot topic marking process; incorporate the uncollected interdisciplinary fields into the semantic structure of the graph and mark them as hot topic candidates, and output a hot topic candidate list.

[0016] Furthermore, add marks to the hot topic candidate list so that it includes an additional cross-potential dimension when generating dynamic feature vectors; aggregate and generate serialized graph slices based on the incremental knowledge graph, extract key information from each graph slice, and generate feature vectors of nodes through embedding learning; extract the difference or high-order change rate of the feature vectors of adjacent graph slices to form a time series difference vector.

[0017] Furthermore, for each topic node, calculate the aggregate function of the overall evolution trend across time periods; when the trend aggregate function exceeds a preset aggregation threshold, determine that the corresponding topic is a significantly potential intersection node and mark it as an intersection trend indicator, and summarize it into the overall graph dimension to form global trend information; combine the time series difference vector and the intersection trend indicator to screen the detected new concept nodes to obtain a list of hot topics.

[0018] Furthermore, after obtaining new academic information and combining it with the list of hot topics, select nodes and relationships with high relevance to the new hot topics in the incremental knowledge graph and construct an incremental training set.

[0019] Combine the incremental training set and the intersection trend indicator to construct an online update loss function, and perform iterative solution on the online update loss function to obtain an online update model after updating the model parameters.

[0020] Furthermore, by analyzing the node distribution and external environment labels in the incremental knowledge graph, distinguish the source domain and the target domain, and define a transfer loss function to characterize the distribution difference between the source environment and the target environment.

[0021] If the transfer loss function is greater than expected, based on the existing online update model in the source environment, perform transfer optimization on the incremental data in the target environment, and perform gradient solution on the defined comprehensive objective function to obtain an optimized model.

[0022] Furthermore, after collecting the external environment parameters and scientific research orientation variables, associate them with the topic nodes or cooperation networks in the incremental knowledge graph; set change nodes for the environment or orientation in segments within the simulation duration to form a scenario function that evolves over time for dynamic loading of scenario configurations; finally, output a scenario definition set, where each scenario corresponds to a set of environment-orientation combinations and their evolution details in the time dimension.

[0023] Furthermore, send the updated node embedding feature vector representation into the optimized model to output a new cross-potential prediction value; the simulation-driven function calculates the potential in each sub-period in combination with the scenario configuration and then accumulates it to the next period until the entire simulation duration is covered; calculate the scenario potential for all scenarios in the scenario definition set respectively, and then perform sorting or classification summary to obtain a comparison of multi-scenario potentials.

[0024] Furthermore, define a policy benefit function to select scenarios in combination with actual management requirements. When there are significant differences in scenario potential or obvious deviations from expectations, collect the deviation information and external feedback during the simulation process into a new batch of incremental data, perform model correction or weighted reinforcement, and output a policy plan and a risk list, including the priority policy combination and response plans under different environmental assumptions.

[0025] Furthermore, after introducing multi-modal information into the data collection pipeline, a multi-modal raw dataset is generated using their respective collection strategies; after preliminary modal differentiation and extraction of core metadata from the multi-modal raw dataset, semantic annotation is performed to generate multi-modal labeled data. If the same event appears in multiple modalities, alignment is performed through multiple factors and merged into the same entity or relationship node.

[0026] Furthermore, for data segments of each modality, a deep learning model is used to extract feature vectors to form a temporary multi-modal vector set; a cross-modal alignment transformation is introduced to map the features of different modalities and aggregate them in a unified representation space to output multi-modal semantic vectors; the multi-modal semantic vectors are cross-checked with keyword tags and hot candidate lists generated by text mining. If abnormal noise is found, a concept drift or inconsistency detection mechanism is enabled for correction.

[0027] Furthermore, corresponding nodes or relationships in the incremental knowledge graph are created or updated for each newly emerged or significantly changed entity, and online learning, transfer learning, scenario simulation, and strategy verification mechanisms are introduced for rapid verification. If multi-modal information indicates a sharp increase in the frequency of occurrence of a certain hot topic in videos or experimental logs, fine-tuning of the optimized model is triggered; if the collaborative research method in a certain field is denser, the corresponding potential value or investment ratio is increased in the simulation link; when multi-modal updates cause changes to the interdisciplinary trend or hot list, the corresponding changes and potential deviation information are fed back again.

[0028] (III) Beneficial Effects

[0029] The present invention provides an interdisciplinary scientific research potential evaluation method based on a dynamic multi-modal knowledge graph, having the following beneficial effects:

[0030] Through the incremental knowledge graph and the concept drift detection strategy, real-time capture of new academic concepts and sudden hotspots is achieved, avoiding the situation of missing important research directions due to the obsolescence of the knowledge base in traditional static evaluations;

[0031] Time series analysis focuses on differential indicators such as keywords, cooperation networks, and citation growth rates at different time slices, enabling rapidly emerging cross-topics to be identified and quantified to output a hot topic list for the first time , which not only helps university management track the pulse of disciplinary evolution but also lays a data foundation for subsequent online model fine-tuning;

[0032] Online learning and transfer learning endow the model with self-adaptive capabilities: when significant changes in the disciplinary cooperation pattern or external environment are detected, according to the incremental data Fine-tune the evaluation parameters or perform cross-domain migration based on distribution differences, taking into account both the original knowledge accumulation and the adaptation to new trends. This can respond more quickly to short-term academic fluctuations, retain the long-term experience of the algorithm, and avoid evaluation distortion.

[0033] Scenario simulation and strategy verification introduce multi-scenario configurations and scenario potential functions to evaluate the cross-disciplinary potential trends under different orientations and environmental combinations. This helps the management side intuitively compare multiple decision-making schemes and quickly select the best balance point. If deviations are found, the results can also be fed back to the online learning module for self-correction, forming a deep collaboration among data, model, and strategy.

[0034] Multi-modal data fusion combines multi-modal vectors such as images, audio, and conference videos with incremental knowledge graphs to capture subtle clues of cross-disciplinary interactions and incorporate them into real-time judgments of hot topics, significantly improving the perception accuracy of emerging research fields.

[0035] The above elements interact with each other on the core of the incremental knowledge graph: the data collection link provides fresh information, the time series trend analysis provides hot topic positioning, online and transfer learning maintain the model's agile adaptation to changing environments, scenario simulation transforms the above results into actionable decision support, and multi-modal fusion further broadens the information source. Through the collaborative effect of these multi-links, taking into account the continuity of data management and model training, it helps universities or research institutions to promptly grasp cross-disciplinary opportunities, optimize resource allocation, and maintain long-term competitiveness in a dynamic and multi-source academic environment. Brief Description of the Drawings

[0036] Figure 1 It is a schematic flowchart of the method for evaluating cross-disciplinary scientific research potential based on a dynamic multi-modal knowledge graph of the present invention. Detailed Embodiment

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Please refer to Figure 1 , the present invention provides a method for evaluating cross-disciplinary scientific research potential based on a dynamic multi-modal knowledge graph, including

[0039] Step 1: When it is monitored that the amount of new data such as academic literature, patents, and scientific research news exceeds the threshold, extract the entity set and the relationship set , then call the concept drift detection function to perform ontology expansion on the high-growth concepts, and perform incremental merging of the new nodes and associations into the incremental knowledge graph based on the semantic hierarchy structure , to achieve dynamic update of the discipline structure and real-time capture of cross-cutting hotspots;

[0040] The first step includes the following contents:

[0041] Step 101, Construction of multi-source data collection pipeline and preprocessing of entity associations

[0042] Configure a unified collection module, establish a timed or real-time crawling strategy for multi-source data such as academic papers, patent documents, scientific research news, and guiding documents, and obtain a phased raw data set after setting the crawling frequency and filtering conditions (such as specifying topics, keywords, or institutional sources) , where represents the collection time or time batch;

[0043] Use text mining algorithms to batch process the phased raw data set , perform entity annotation for common structures in the academic field (such as author names, institutional names, professional term topic concepts, etc.). To reduce noise, it is necessary to perform preliminary merging of duplicate entities or name variants to obtain an entity set ;

[0044] Extract the association information between documents in terms of citation, author cooperation, keyword co-occurrence, etc., to form a preliminary relationship set , and these entities and relationships are temporarily stored as temporary structured data ;

[0045] Assign unique identifier IDs to the newly collected data, and at the same time record the mapping relationship between the phased raw data set used as input and the temporary structured data used as output to ensure accurate backtracking and invocation in subsequent steps;

[0046] When in use, multi-source data is centrally managed through the same pipeline, ensuring that subsequent steps do not require repeated crawling or cleaning, reducing system resource waste, and assigning unique IDs to all newly identified entities and relationships and recording their time batches to achieve traceability and consistency in subsequent steps.

[0047] Step 102, Incremental knowledge graph generation and ontology expansion

[0048] After receiving the temporary structured data , incrementally merge the entity set and the relationship set into the incremental knowledge graph If overlap with existing nodes or attribute conflicts are detected, a merging strategy based on a confidence threshold is adopted for correction. Specifically, a comprehensive score is calculated based on name similarity and cosine similarity of embedding vectors. If the comprehensive score exceeds the confidence interval, it is determined as a duplicate node and merged according to the frequency of text attributes first and the weighted average of numerical attributes by credibility. If it is below the confidence interval, a new node is created. When it is between the two, it is marked as a conflict pending review and finally confirmed by manual or semi-automatic processes; the incremental knowledge graph generated in this process is the latest state of the time-oriented knowledge graph and serves as the benchmark for time series analysis in subsequent steps;

[0049] To promptly identify new or significantly changed academic concepts, concept drift calculations need to be performed on nodes (entities) and edges (relationships) in the incremental knowledge graph where:

[0050] To measure the dynamic changes in the embedding space of the knowledge graph over the time period the concept drift measure function is defined as:

[0051]

[0052] where: is the incremental knowledge graph at time and contains the node set and its associated relationship information; is the set of valid node (entity) indices at time ;

[0053] is the -dimensional vector embedding value of the node at consecutive time ; if only the embedding data at discrete time points is available, it can be constructed through linear or high-order interpolation ;

[0054] is the partial derivative or difference approximation of the above dimensional embedding value with respect to time and reflects the rate of change of the node embedding in the dimension over time; is the weight factor for the dimension, with a value between 0 and 1, used to strengthen or weaken certain feature dimensions; is a high-order distance exponent greater than 1; is the time interval;

[0055] If the obtained concept drift measure exceeds the preset drift threshold , it can be determined that significant semantic evolution has occurred, significant concept drift has been determined, and it is used as a new concept node, which can immediately trigger corresponding ontology expansion or hot spot marking processing, and automatically update the classification hierarchy of newly emerging or significantly changed concept nodes;

[0056] Incorporate cross-disciplinary fields that may not have been included into the semantic structure of the knowledge graph, and mark these newly added or updated nodes as hot spot candidates, and output a hot spot candidate list ;

[0057] When in use, ensure that the knowledge graph can evolve rapidly with data iteration through incremental merging, and present it in a unified structural form, providing reliable continuity for subsequent time series analysis. Based on the concept drift measure detection mechanism, automatically identify new hot spot concepts, reduce the cost of manually maintaining the ontology hierarchy, and avoid missing potential important cross-disciplinary fields; use the combination of entity embedding and change amplitude to not only capture new nodes, but also quantify the semantic drift of existing concepts to achieve early warning of subtle semantic changes.

[0058] Step 2: When the incremental knowledge graph is updated and reaches the next time slice , based on differential mapping and combined with indicators such as citation growth rate and cooperation frequency, perform time series slicing and trend aggregation on the discipline theme and cooperation network, and output cross-trend indicators and the hot topic list , and associate and verify them with the concept drift trajectory;

[0059] The said Step 2 includes the following contents:

[0060] Step 201: Time slicing and dynamic feature extraction

[0061] Based on the incremental knowledge graph , aggregate the data storage and marking of different time batches, and generate serialized graph slices according to the set time window (such as quarterly, annual or adaptive time window) , denoted as:

[0062]

[0063] where represents the time stamp or time interval index of each time period; these time-series graphs retain the node and relationship incremental update information after ontology expansion, and can accurately record the evolution status of discipline concepts and cooperation relationships in different time periods.

[0064] For each graph slice Extract key information such as subject themes, collaboration networks, citation relationships, etc., and generate feature vectors for nodes through embedding learning methods (such as embedding models based on graph representation learning or text representation learning). , where represents the node identifier existing at time , which contains multi-dimensional encodings such as topic labels, collaboration frequencies, citation relationships, etc.;

[0065] To identify interdisciplinary intersections, it is necessary to mark the hot candidate list so that it includes an additional cross-potential dimension when generating dynamic feature vectors; the higher the value of this dimension, the more neighbors of different disciplines the node is connected to in the existing collaboration network, and its semantic embedding has undergone a significant drift; this cross-potential dimension can be enhanced and encoded in combination with the results of the concept drift measure in step 102 (such as a large number of newly added concept nodes) for easier focus in subsequent trend detection;

[0066] Extract the difference or high-order rate of change of the feature vectors of adjacent graph slices to form a time-series difference vector , and this difference value can help judge the increase or decrease trend of a certain subject theme or collaboration relationship between different time periods, avoiding ignoring potential clues of sudden increase or decrease based only on absolute values.

[0067] Through time slicing and feature differences between slices, while maintaining the historical data of nodes, dynamic changes can be accurately captured. After marking the hot candidate list , the evolution of emerging interdisciplinary entities can be more directly focused on in subsequent analysis.

[0068] Step 202, Trend Aggregation and Cross-Hotspot Detection

[0069] To measure the overall evolution trend across time periods, define a multi-time-window trend aggregation function denoted as:

[0070]

[0071] Where: , represents the vectorized representation of a certain theme or node obtained at different time points (such as integrating information such as collaboration frequency, citation characteristics, interdisciplinary labels, etc.); is an ordered set of time slice or time interval endpoints;

[0072] is the vectorized representation of the theme or node at time with respect to Partial derivatives (if there are only discrete vectors, the difference can be approximated by interpolation followed by taking the difference). is a set of matrix weights that vary over time and are used to further correct or amplify the change in the nodal vector during integration. It can be regarded as a dynamic mapping for factors such as external research orientation, funding, or interdisciplinary degree;

[0073] is a user-defined higher-order or weighted vector norm, such as norm ; is the exponent used to amplify the result of this norm, ; is the weight coefficient of the time window and its value is greater than or equal to 0;

[0074] For each topic node and , the trend aggregation function is calculated over each time period , and the cumulative change value of the topic is obtained. When the trend aggregation function exceeds the preset aggregation threshold, it is determined that there is a significant anomaly in the research activity, citation frequency, or cooperation scale of this topic, and it is a significantly cross-potential node, marked as the cross-trend indicator . This indicator will be summarized into the global trend information in the entire map dimension and finally output as ;

[0075] Combining the higher-order difference data generated in step 201 and the cross-trend indicator , entities that are detected as newly emerging or whose embedded representation has a significant drift from the historical semantics, that is, new concept nodes, are screened with emphasis.

[0076] If some new concept nodes or significantly cross-potential nodes show a sudden increase in a short period of time and have a high degree of interdisciplinary association, they are inducted into the hot topic list ;

[0077] Finally, step 202 will produce two types of key results:

[0078] Cross-trend indicator : Measures the evolution amplitude of different topic nodes within a given time interval;

[0079] Hot topic list : Focuses on emerging, high-growth rate, or highly interdisciplinary disciplines and topics;

[0080] In use, it can take into account the importance of different time windows to ensure that both rapidly emerging hotspots and areas with sustained and stable development can be reasonably evaluated; highlight potential cross-outbreak points, magnify and identify those themes with extremely high change rates in a short period of time, and improve the sensitivity to capture explosive interdisciplinary points. It can not only output quantitative cross-trend indicators , but also additionally give a list of hot topics , providing an intuitive reference for subsequent model optimization and strategy formulation.

[0081] First, slice the incremental knowledge graph on the time axis and construct dynamic feature vectors, and then use the multi-time window trend aggregation formula and high-order norm measure to identify emerging interdisciplinary points and hotspots. The output cross-trend indicators and list of hot topics can not only provide an accurate basis for model update for the third-step online and transfer learning model optimization, but also lay a foundation for trend judgment and strategy evaluation in the fourth-step final scenario simulation stage.

[0082] Step 3. Obtain cross-trend indicators and the list of hotspots When a significant change in the scientific research environment is detected, perform weighted update on the incremental data , and use the transfer loss function to adaptively fine-tune the model parameters of different regions or universities, continuously maintaining the evaluation accuracy and environmental matching degree;

[0083] The said Step 3 includes the following contents:

[0084] Step 301. Online learning and model parameter fine-tuning

[0085] After receiving the cross-trend indicators , the list of hot topics and the incremental knowledge graph , retain the original evaluation model parameters obtained from the previous round of training for online learning initialization to avoid training from scratch;

[0086] When new academic information (papers, patents, scientific research news, etc.) arrives, combine the list of hot topics , and select nodes and relationships with high relevance to these new hot topics in the incremental knowledge graph to construct an incremental training set , and this incremental training set can focus on discipline topics that have appeared frequently or have high CTI values recently to strengthen the adaptation to new trends;

[0087] ​​​, define the following online update loss function :

[0088]

[0089] where: represents the prediction output by the current model for the input under the parameter (such as the feature vector of a paper or scholar node); is a user-defined high-order norm; is the exponent for amplifying the prediction deviation;

[0090] represents the cross-trend metric corresponding to the input and can be mapped by or hot spot markers. When is more highly associated with the hot topic, takes a larger value; , is the sample preference enhancement index, used to exacerbate or weaken the preference for high samples;

[0091] Through methods such as gradient descent or variational inference, the online update loss function is iteratively solved to update the model parameters, and finally the online update model is obtained, which carries the ability to adapt to the latest academic hotspots;

[0092]

[0093] where, is the step size coefficient of online learning, which determines the amplitude of each incremental update;

[0094] When in use, the hot topic list and cross-trend indicators are used to assign higher training weights to emerging topics, enabling the model to quickly learn the latest research trends and adapt to new trends under the condition of limited incremental data; The weighting strategy is different from traditional uniform loss. This formula dynamically adjusts the importance of each sample according to the cross-trend indicator and can better capture rapidly emerging research directions.

[0095] Step 302, Transfer Learning and Environment Adaptation

[0096] If it is necessary to deploy in different universities or regions, there are often differences in subject layout and data distribution, and hot topics and cross-trends are not exactly the same in different environments. Here, by analyzing the node distribution and external environment labels in the incremental knowledge graph , the source domain (such as the environment of a research-intensive university) and the target domain (such as the environment of a teaching-oriented university) are distinguished;

[0097] To characterize the distribution difference between the source environment and the target environment, a transfer loss function is defined :

[0098]

[0099] where: is the parameter vector or matrix of the evaluation model (possibly from online learning or initial training), which acts on the input to produce an output representation or prediction result;

[0100] , are the data sets of the source environment (such as a research-intensive university) and the target environment (such as a teaching-oriented university), respectively;

[0101] is the node or sample after being mapped by the model parameter to obtain the embedded vector or prediction output;

[0102] is the transformation matrix or operator for cross-domain alignment, which performs a weighted transformation on and can be learned simultaneously during transfer optimization or let it vary piecewise over time;

[0103] is a high-order or weighted vector norm (consistent with the previous one), which is used to magnify or calculate the weighted distance of . Commonly, it is norm ( is more sensitive when is the power exponent, and its value is greater than 1;

[0104] is the weighted coefficient of the cross-domain sample pair and can be assigned according to external priors (such as the similarity between the two domains, cross-trend indicators, etc.). , if this item is not used, it can be set to ;

[0105] If the transfer loss function is greater than expected, it means that greater efforts are needed for transfer to adapt to the target environment. Based on the existing online updated model in the source environment, the incremental data in the target environment is transferred and optimized, and the comprehensive objective function is defined as:

[0106]

[0107] where: Similar to the online learning loss form in step 301, except for the incremental data for the target environment here ; is the domain difference measurement term defined above. For specific reference, please refer to the above text; , is the balance coefficient, which controls the balance degree between the two in the transfer learning process;

[0108] By minimizing , on the one hand, it ensures the prediction accuracy for the target domain data, and on the other hand, it reduces the conflict with the feature distribution of the source domain model, thereby realizing environment-adaptive transfer learning;

[0109] After gradient solving or other optimization algorithms are used to solve, the optimized model is obtained, which has high adaptability to the emerging interdisciplinary trends in the target environment.

[0110] When in use, in different types of universities or research institutions, instead of blindly following the unified model, but by using the domain difference measurement term to measure the distribution difference and then making targeted transfers, the credibility of the evaluation results is significantly improved, and cross-domain adaptation is realized; on the basis, the general features learned in the source environment are retained, and the existing accumulations will not be completely abandoned due to domain change, so that the experience can be carried to the new field, and the source model experience can be retained; by balancing the differences between the target domain and the source domain, the cross-domain generality or local accuracy can be flexibly selected to adapt to the needs of various actual scientific research environments.

[0111] Integrate the cross-trend index with the hot topic list into the update mechanism of online learning and transfer learning, so as to realize the adaptive optimization of the model in the situation where new data emerges continuously and there are multiple environmental differences; step 301 focuses on quickly adapting to sudden academic hotspots in the same environment, and step 302 further enables the model to adapt to the local characteristics of different universities or research institutions; through the completion of this step, the optimized model can not only continuously maintain high-precision prediction in the dynamically evolving scientific research environment, but also provide real-time and reliable basis for the fourth-step scenario simulation and strategy verification, forming a closed loop from data collection to trend recognition and then to model iteration, and truly achieving the forward-looking and long-term evaluation of scientific research intersection potential.

[0112] Step Four: When the optimized model is obtained by solving and the management terminal submits the combination of external environment parameters and the scientific research orientation variable to obtain the scenario , the scenario simulation framework will call the scenario potential function , and after performing time-series iteration and high-order norm measurement on the subject evolution, output each potential result;

[0113] The fourth step includes the following contents:

[0114] Step 401, multi-scenario environment definition and parameter configuration

[0115] According to the optimized model obtained in the previous step and the incremental knowledge graph , define the external environment parameters and scientific research orientation variables to be simulated here, including:

[0116] External environment parameters : can cover government funding intensity, industrial demand changes, impacts of public emergencies, etc.;

[0117] Scientific research orientation variables : such as university fund allocation strategies, key discipline support plans, construction of interdisciplinary joint laboratories, etc.;

[0118] Assign unique identifiers to different combinations of and associate the external environment parameters and scientific research orientation variables with the theme nodes or cooperation networks in the incremental knowledge graph. For example, weighted coefficients can be set on graph feature dimensions such as cooperation frequency, citation speed, or new discipline introduction rate to reflect the intervention degree brought by a specific environment or orientation.

[0119] Select the simulation duration or critical time period , such as an interval of the next 3 years or 5 years, and set change nodes of the environment or orientation in segments within the simulation duration to form a scenario function that evolves over time , for dynamically loading scenario configurations; where:

[0120] Scenario function is defined in the form of a piecewise constant vector over the simulation interval :

[0121]

[0122] where: is the time segmentation point selected according to the orientation release time, fund appropriation cycle, or major event node, satisfying ;

[0123] is the th segment's environment-orientation vector, jointly composed of the external environment parameters (such as annual scientific research fund ratio, industrial demand heat, etc.) and orientation variables (such as discipline support weight, cooperation incentive degree) within this segment;

[0124] is an interval indicator function, which takes 1 when and 0 otherwise;

[0125] In practical applications, it can correspond to key moments such as the start date of major national scientific research projects, the date of reallocation of university budgets, or the outbreak date of public health emergencies; each and are then normalized proportionally or directly assigned according to indicators such as the number of current literature outputs, the number of project approvals, or the amount of special funds.

[0126] In this way, it can accurately reflect the multi-segmented guidance and environmental changes over time, providing clear and quantifiable inputs for calling the optimized model in sub-step 402 for scenario simulation;

[0127] Based on the above information, output the scenario definition set Each scenario corresponds to a set of environment-guidance combinations and the details of their evolution in the time dimension.

[0128] When used, by assigning unique identifiers to different guidance and environmental parameters and mapping them, subsequent simulations can flexibly call various scenario definitions to achieve rapid switching between multiple sets of assumptions under the same model framework; with the help of the optimized model and the atlas and trend information in the first and second steps, the scenario definition can accurately grasp the key intervention points of real scientific research activities.

[0129] Step 402, Multi-scenario Simulation and Calculation of Interdisciplinary Potential

[0130] After obtaining the scenario definition set , to uniformly analyze each scenario , introduce the simulation driving function to call the optimized model to evaluate the potential of the disciplinary nodes or cooperation networks at time . Among them, the simulation driving function refers to the mapping function that batch-calculates the potential of all nodes after inputting the scenario and time into the optimized model . It can be specifically defined as:

[0131]

[0132] where; is the concatenated feature vector of all nodes in the atlas at time ;

[0133] is the concatenation of the scenario vector (environmental parameters) defined in sub-step 401 with the guiding variable ;

[0134] Project the scenario vector into the node feature space; Denote vector concatenation;

[0135] and are the weights and biases of the first layer of the network, is the ReLU activation;

[0136] and are the weights and biases of the second layer of the network, is a smooth non-linear mapping such as Softplus or Sigmoid;

[0137] Output of the component, which is the predicted cross-potential value of node at this scenario and time;

[0138] This structure can not only fuse the time-point features and scenario intervention factors, but also ensure the capture of complex cross-influences through multi-layer non-linear mappings, providing a refined and differentiable driving signal for the time integration and policy comparison in sub-step 402.

[0139] Inside the simulation driving function, the external environmental parameters and the scientific research guiding variable will be comprehensively weighted and corrected for relevant features, and then the updated node embedding feature vector representation will be fed into the optimized model to output a new predicted cross-potential value;

[0140] For the time series evolution, piecewise iteration or interpolation methods can be adopted, allowing the simulation driving function to complete a potential calculation in each sub-period in combination with the scenario configuration, and then accumulate it to the next period until the entire simulation duration is covered ;

[0141] To quantitatively evaluate the disciplinary potential trends under different scenarios , define the scenario potential function :

[0142]

[0143] Where: are the parameters of the optimized model, which are obtained by the simulation driving function Call, used to calculate or predict the potential of each subject node in the scenario at any moment : Is an environment - oriented combined scenario;

[0144] Is a simulation - driven function, outputting the potential vector at the moment ; where represents the potential value of the node (such as a certain subject or cooperation group);

[0145] Is the partial derivative of the potential value of the node with respect to time (or can be approximated by differences under discrete - time conditions), characterizing the rate of change of the node potential over time; Is the set of nodes determined to be the key focus of interdisciplinary research in the scenario , at the moment ; Is a high - order or mixed vector norm, which is of type ( ) or other mixed - type norms; Is a magnification factor, or larger;

[0146] Is a time - weighting function, to is the simulation interval; where, ;

[0147] Meaning: Taking the center of the simulation interval as the highest weight point, symmetrically decaying forward and backward;

[0148] Parameters: , determines the center - of - gravity position (for example, is the mid - point of the interval);

[0149] , determines the width (the larger, the slower the decay);

[0150] Effect: When hoping to focus on the mid - term of the simulation or a certain critical moment, it can smoothly focus on that time period, giving lower weights to moments farther from the center of gravity;

[0151] The scenario potential function The larger it is, the more likely it is to achieve a breakthrough in interdisciplinary potential or output of results in this scenario. It can also be used to inversely evaluate the cost - benefit ratio of the implementation of the guiding direction according to requirements; for all in the scenario definition set Calculate the scenario potential functions respectively Then perform sorting or classification summary to obtain the multi-scenario potential comparison for strategy screening and risk assessment.

[0152] When in use, through the scenario definition set and the simulation driving function it is possible to test the impacts of multiple combinations of environments and orientations under a unified process, enabling the management side to intuitively compare different decision-making schemes, integrate the potential assessment over time, observe the differences between short-term and long-term effects, avoid only focusing on extreme values at a certain moment while ignoring cumulative gains or delayed effects, and use the optimized model obtained by optimizing the online and transfer learning model in the third step to predict the potential, and then weight the graph features in combination with the scenario intervention parameters to ensure that the simulation results have both accuracy and executability.

[0153] Step 403: Strategy screening and iterative feedback

[0154] According to the multi-scenario potential comparison define the strategy benefit function in combination with the actual management requirements For example, combine the multi-scenario potential comparison with factors such as the input resources and the difficulty of orientation execution, calculate the benefit-risk balance value for each scenario for sorting or clustering;

[0155] If a certain scenario has high potential but also high execution difficulty and risk, the management side can make a cautious choice in the final decision; if some other scenarios achieve a better balance between potential and risk, they can be preferentially entered into the actual deployment stage, thus realizing the selection of scenarios ;

[0156] When there are significant differences in the scenario potential function or obvious deviations from the expectations, the deviation information and external feedback during the simulation process can be collected into a new batch of incremental data and sent back to the optimization stage of the online and transfer learning model in the third step for model correction or weighted strengthening; this iterative feedback mechanism ensures that when the gap between the simulation and the reality is large, the model can update its parameters in a timely manner and be closer to the actual situation during subsequent simulations or real deployments; finally, step 403 outputs the strategy plan and risk list, including the priority strategy combination (such as increasing the investment in a certain cross-field at a specific time point) and the response plans under different environmental assumptions;

[0157] When in use, the deviation of the simulation results and the real data are fed back to the optimized model in the third step through the retraining mechanism , continuously improve the fitness of the optimized model to avoid rapid obsolescence with environmental changes after a one-time evaluation; comprehensively consider scenario potential, resource consumption, and execution risks, so that the management side can more rationally allocate limited resources, avoid blindly chasing a certain high-potential field while ignoring the implementation difficulty, and achieve multi-objective balance.

[0158] Integrate the optimized model The increment knowledge graph accumulated in the early stage and cross-trend indicators and other elements to build a customizable multi-scenario simulation framework, quantify the impacts of different orientations and external environments, and output targeted strategies and risk assessments based on this.

[0159] Combined with the online and transfer learning mechanisms in the third step, finally form a complete closed loop of data collection - trend recognition - model optimization - strategy simulation, ensuring that the assessment of the potential of scientific research intersections is both forward-looking and keeps pace with the times, providing strong technical support for universities and research institutions to make accurate decisions in a dynamic environment.

[0160] However, with the increasing diversification of the forms of scientific research activities and data sources, a single text or citation network has become difficult to fully capture the laws of interaction and knowledge diffusion in some emerging disciplines. Therefore, introduce more modal data (such as academic conference videos, experimental process data, personal social media interactions, etc.), and deeply integrate it with the existing increment knowledge graph to further enrich and improve the characterization of cross-research activities.

[0161] Step Five: When a significant increment in multi-modal sources is detected based on the graph text, perform modal tagging and semantic vector extraction on the generated multi-modal raw dataset and then use an aggregation function to align cross-modal features and output a multi-modal fusion library , and finally integrate its increment into the multi-modal knowledge graph to strengthen the three-dimensional recognition of emerging cross-disciplinary fields and the adaptability of the online model;

[0162] The above Step Five includes the following contents:

[0163] Step 501: Multi-modal data acquisition and preprocessing

[0164] Based on the text, patents, citation networks, etc. used in the first four steps, introduce multi-modal information such as videos, audios, images, and sensor logs, unify these new sources into the data collection pipeline, and use their corresponding collection strategies (such as regularly capturing live replays of academic conferences and obtaining output logs of scientific research equipment) to generate a multi-modal raw dataset ;

[0165] For the multi-modal raw dataset Perform preliminary modal discrimination and extract core metadata (such as video duration, audio transcription text, image scene description, etc.), and then perform semantic annotation on this core metadata information to generate multi-modal labeled data ; This step will be combined with text mining to discover disciplinary keywords, research institutions, scholar names, etc. mentioned in video or audio content. If the same event (such as a certain scientific research forum) appears in multiple modalities such as conference videos, social media posts, and paper citations, it is necessary to align through multiple factors such as timestamps, keywords, and geographical locations, and merge them into the same entity or relationship node to avoid duplicate insertion in the incremental knowledge graph.

[0166] At this time, refer to the concept drift detection mechanism in sub-step 102 to mark the sudden associations that may be brought by new modalities and output them to the temporary multi-modal structure ; Supplement with non-text modalities such as videos, audios, and images to achieve a more comprehensive capture of academic activities, thereby expanding the data coverage, and avoiding knowledge graph structure conflicts caused by duplicate mapping of different modalities through multi-angle matching of timestamps, metadata, etc.

[0167] Step 502, Multi-modal Embedding Representation and Semantic Fusion

[0168] Use a deep learning model (such as CNN, Transformer, or a hybrid neural structure) to extract feature vectors from data segments of each modality (video frames, audio segments, image regions, etc.) to form a temporary multi-modal vector set , if it involves text transcription or OCR (optical character recognition) content, the semantic labels can be fused with visual / audio features by combining the entity information of the aforementioned text mining;

[0169] Introduce cross-modal alignment transformation , and map the features of different modalities and aggregate them in a unified representation space, and define the following multi-modal aggregation function :

[0170]

[0171] Where: is the vector representation of the th multi-modal segment at time ; is the modal weight factor used to strengthen or suppress the importance of certain modalities; is a linear or non-linear mapping used to complete alignment in a unified space; is the higher-order or hybrid norm defined above; is the amplification exponent;

[0172] The aggregation function outputs a multi-modal semantic vector represents the comprehensive features of a certain entity (such as a scholar or a research team) in various modalities, and is used for the next step of knowledge graph merging; after alignment and aggregation, the multi-modal semantic vectors are cross-checked with the keyword tags and hot topic candidate lists generated by text mining.

[0173] If abnormal noise is found (for example, strong conflict between image recognition results and text entities), a concept drift or inconsistency detection mechanism similar to that in step 102 is enabled for correction. The finally retained multi-modal features will be uniformly stored in the multi-modal fusion library. ;

[0174] When in use, extract interdisciplinary traces from multiple aspects of video / audio / text to further enhance the sensitivity to emerging scientific research fields; it can adapt to the usage requirements of multi-modal information in different universities and different fields, coordinate with the keyword analysis in the second step, the online learning in the third step and other mechanisms, and add the new multi-modal semantic vectors to the input feature space of the subsequent evaluation model, which helps the optimized model to be further iterated in subsequent re-learning.

[0175] Step 503, Multi-modal incremental integration and dynamic update

[0176] Utilize the multi-modal fusion library to create / update the corresponding nodes or relationships in the incremental knowledge graph for each newly emerged or significantly changed entity (researcher, experimental project, video topic, etc.). The update rules are the same as the incremental merging strategies adopted in the first and second steps, but the attribute fields need to be expanded (such as adding visual description speech tags, etc.) to record multi-modal information.

[0177] Record the updated multi-modal knowledge graph as , and introduce the online learning and transfer learning in the third step and the scenario simulation and strategy verification mechanism in the fourth step for quick verification:

[0178] If the multi-modal information indicates that the frequency of a certain hot topic has increased sharply in the video or experimental log, it can trigger the fine-tuning of the optimized model ; if the new modality shows that the collaborative research methods in a certain field are more intensive, the corresponding potential value or the proportion of guiding investment can also be increased in the simulation session; when the multi-modal update causes changes to the interdisciplinary trend or hot topic list, these changes and potential deviation information need to be passed back to step 201, 202 or the online learning module in the third step for regular or real-time re-training, and new strategy simulations can also be triggered when necessary, so that the management side can understand the differences brought by multi-modal information to the evaluation results and resource allocation suggestions. ​

[0179] When in use, it can keep the incremental knowledge graph timely collect new modal entities, not only can merge text data, but also can be compatible with new interdisciplinary signals generated by information such as videos and audios; after multimodal fusion, some potential breakthrough points that are not easily detected in traditional text data (such as abnormal laboratory instrument data and the popularity of group discussions at academic conferences) can be quickly converted into monitorable nodes to strengthen the forward-looking perception of hot fields; it realizes the comprehensive collection and deep fusion of non-text information such as images, audios, videos, and sensor logs, and feeds the newly obtained multimodal features back to the optimized model. And the scenario simulation framework further improves the perception and decision-making support capabilities of this system for the rapidly iterative and cross-modal scientific research ecosystem, providing more comprehensive and three-dimensional interdisciplinary insights for the scientific research management side of universities.

[0180] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0181] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0182] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0183] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0184] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

Claims

1. A method for evaluating the interdisciplinary scientific research potential based on a dynamic multi-modal knowledge graph, characterized in that: including, if the newly crawled academic data volume exceeds the threshold, text mining is used to extract entities and relationships. After identifying emerging terms through concept drift detection, they are incrementally merged into the incremental knowledge graph, and then the ontology level is expanded to accommodate the new concepts; After reaching a new time slice, calculate the multi-time window difference vector, aggregate it to highlight the theme growth or convergence trend, generate cross-trend indicators and a list of hot topics, and mark sudden rises and falls; If the cross-trend indicators and the list of hot topics show significant fluctuations within the domain, perform online parameter updates to optimize the model parameters, and use the transfer loss function to adaptively fine-tune the model parameters in different regions; When the optimized model is obtained and the management end submits a scenario composed of external environmental parameters and scientific research orientation variables, call the scenario potential function to evaluate the potential trend of resource allocation and interdisciplinary collaboration, output comparable indicators, and when the actual and predicted differences are greater than expected, feedback a deviation signal to the online learning module; Perform modal marking and semantic vector extraction on the generated multi-modal raw data set. The aggregation function aligns the cross-modal features and outputs a multi-modal fusion library, which is incrementally incorporated into the multi-modal knowledge graph, and triggers re-fine-tuning of the optimized model when an anomaly occurs.

2. The interdisciplinary scientific research potential evaluation method according to claim 1, wherein: Establish a timed or real-time crawling strategy for multi-source data, set the crawling frequency and filtering conditions, and then obtain a phased raw data set; use text mining algorithms to perform entity annotation on the common structures in the academic fields in the phased raw data set, and perform preliminary merging on duplicate entities or name variants to obtain an entity set, and form a preliminary relationship set after extracting the association information between the documents.

3. The interdisciplinary scientific research potential evaluation method according to claim 2, wherein: Incrementally merge the entity set and the relationship set into the incremental knowledge graph. If overlap with existing nodes or attribute conflicts are detected, a merging strategy based on a confidence threshold is adopted for correction; Perform concept drift calculation for nodes and edges in the incremental knowledge graph; If the obtained concept drift measure exceeds the preset drift threshold, use it as a new concept node, trigger the corresponding ontology expansion or hot topic marking process, incorporate the uncollected interdisciplinary fields into the semantic structure of the graph, and mark them as hot topic candidates, and output a list of hot topic candidates.

4. The interdisciplinary scientific research potential evaluation method according to claim 1, wherein: Add marks to the list of hot topic candidates so that it includes an additional cross-potential dimension when generating dynamic feature vectors; Based on the incremental knowledge graph, aggregate and generate serialized graph slices, extract key information from each graph slice, generate feature vectors of nodes through embedding learning, and extract the difference or high-order change rate of the feature vectors of adjacent graph slices to form a time-series difference vector.

5. The interdisciplinary scientific research potential evaluation method according to claim 4, wherein: For each topic node, calculate an aggregation function that measures the overall evolution trend across time periods; When the trend aggregation function exceeds the preset aggregation threshold, it is determined that the corresponding theme is a node with significant cross-potential and marked as a cross-trend indicator. When summarized to the entire graph dimension to form global trend information, combined with the time series differential vector and the cross-trend indicator, the detected new concept nodes are screened to obtain a list of hot topics.

6. The interdisciplinary scientific research potential evaluation method according to claim 5, characterized in that: After obtaining new academic information and combining it with the list of hot topics, nodes and relationships with high relevance to the new hot topics are selected in the incremental knowledge graph to construct an incremental training set; Combined with the incremental training set and the cross-trend indicator, an online update loss function is constructed, and the online update loss function is iteratively solved. After updating the model parameters, an online update model is obtained.

7. The interdisciplinary scientific research potential evaluation method according to claim 6, characterized in that: By analyzing the node distribution and external environment labels in the incremental knowledge graph, after distinguishing the source domain and the target domain, a transfer loss function is defined to characterize the distribution difference between the source environment and the target environment; If the transfer loss function is greater than expected, based on the existing online update model in the source environment, the incremental data in the target environment is transferred and optimized, and the defined comprehensive objective function is solved by gradient to obtain an optimized model.

8. The interdisciplinary scientific research potential evaluation method according to claim 7, characterized in that: After collecting the external environment parameters and scientific research orientation variables, they are associated with the theme nodes or cooperation networks in the incremental knowledge graph; Change nodes of the environment or orientation are set in segments within the simulation duration to form a scenario function that evolves over time for dynamically loading scenario configurations, and finally a scenario definition set containing several scenarios is output.

9. The interdisciplinary scientific research potential evaluation method according to claim 8, characterized in that: The updated node embedding feature vector representation is sent into the optimized model to output a new cross-potential prediction value; The simulation-driven function calculates the potential in each sub-period in combination with the scenario configuration, and then accumulates it to the next period until the entire simulation duration is covered; the scenario potential is calculated for all scenarios in the scenario definition set respectively, and then sorted or classified and summarized to obtain a comparison of multi-scenario potentials.

10. The interdisciplinary scientific research potential evaluation method according to claim 9, characterized in that: A strategy benefit function is defined in combination with the actual management requirements to select scenarios. When there are significant differences in scenario potential or obvious deviations from expectations, the deviation information and external feedback during the simulation process are collected into a new batch of incremental data for model correction or weighted reinforcement, and a strategy plan and a risk list are output, including a priority strategy combination and response plans under different environmental assumptions.

11. The interdisciplinary scientific research potential evaluation method according to claim 10, characterized in that: After introducing multimodal information into the data acquisition pipeline, a multimodal raw dataset is generated using their respective acquisition strategies; after preliminary modal differentiation and extraction of core metadata from the multimodal raw dataset, semantic annotation is performed to generate multimodal labeled data. If the same event appears in multiple modalities, alignment is performed through multiple factors and merged into the same entity or relationship node.

12. The interdisciplinary scientific research potential evaluation method according to claim 11, wherein: Feature vectors are extracted from data segments of each modality using a deep learning model to form a temporary multimodal vector set; A cross-modal alignment transformation is introduced to map the features of different modalities and aggregate them in a unified representation space to output multimodal semantic vectors; the multimodal semantic vectors are cross-checked with keyword tags and hot candidate lists generated by text mining. If abnormal noise is found, a concept drift or inconsistency detection mechanism is enabled for correction.

13. The interdisciplinary scientific research potential evaluation method according to claim 12, wherein: Corresponding nodes or relationships in the incremental knowledge graph are created or updated for each newly emerged or significantly changed entity, and an online learning and transfer learning and scenario simulation and strategy verification mechanism is introduced for rapid verification. If the multimodal information indicates that the frequency of a certain hot topic surges in videos or experimental logs, fine-tuning of the optimized model is triggered; If the collaborative research method in a certain field is more intensive, the corresponding potential value or investment ratio is increased in the simulation link. When the multimodal update causes changes to the interdisciplinary trend or hot list, the corresponding changes and potential deviation information are fed back again.

Citation Information

Patent Citations

  • Power transmission and transformation project knowledge graph construction method based on machine learning

    CN114936641A

  • Intelligent doorbell visitor identification method and system based on deep learning

    CN118761034A