A Knowledge Graph Fusion Method and System for Academic Industry Based on Large Models

Academic and industrial knowledge graphs are constructed through large models and multi-label random forest algorithms, and the problem of unclear division of scientific and technological achievements is solved, accurate integration of industrial fields and maps is achieved, and effective utilization of academic achievements is supported.

CN119204191BActive Publication Date: 2025-07-22SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411693518.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-07-22
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

In the existing technology, the industrial fields of scientific and technological achievements are unclear, and there is a lack of standards for the integration of academic achievements and industrial knowledge graphs, which leads to vague classification and difficult to widely accept, affecting the accurate evaluation and utilization of scientific and technological achievements.

Method used

The semantic vector model of the large model and the multi-label random forest algorithm are used to construct academic and industrial knowledge graphs, transform and classify academic achievement attribute information, generate industrial field labels, and calculate the correlation degree to achieve graph fusion.

Benefits of technology

It improves the accuracy of the industrial field division of academic achievements, provides clear field division standards, supports talent selection and industrial development planning, and enriches knowledge graph information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204191B_ABST
    Figure CN119204191B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for fusing academic-industry knowledge graphs based on large models, which are applied to the field of artificial intelligence technology. The method includes: constructing an academic achievement knowledge graph based on academic achievement text data, and constructing an industry knowledge graph based on industry text data; using the semantic vector model of the large model to convert the processed academic achievement attribute information; adding labels to the vector-represented academic achievement attributes as a training set, constructing and training a multi-label random forest model based on the training set, classifying the academic achievements in the academic achievement knowledge graph to obtain the industrial field classification results of the academic achievements; dividing the industrial fields of scientific research talents through the industrial field classification results of the academic achievements to generate industrial labels; calculating the correlation degree between the industrial labels and the industry knowledge graph, and completing the fusion of the academic achievement knowledge graph and the industry knowledge graph based on the correlation degree results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method and system for fusing academic and industrial knowledge graphs based on large models. Background Art

[0002] Knowledge graph is a knowledge management technology with great application prospects. Currently, there are platforms that have launched technology services based on scientific and technological knowledge graphs to explore the combined application of knowledge graphs and technology services. However, in the technology of extracting entities to construct knowledge graphs using large models, and in the research on the application of realizing the division of scientific and technological achievement fields and talent fields based on the semantic vector model and multi-label random forest algorithm of large models, it is almost blank.

[0003] Text classification refers to the process of automatically determining the category of a text according to the content of the text under a certain classification system. Currently, the commonly used text classification technology classifies texts by using the similarity between the text to be classified and the existing texts in each category, including algorithms such as KNN algorithm, Bayesian algorithm, support vector machine algorithm, and concept inference network algorithm based on semantic networks. However, its classification accuracy is limited by the feature extraction of the text, and with the popularity of large models, the derived semantic vector model can better understand the semantic information of the text and convert the semantic information of the text into features of a fixed dimension, which is convenient for subsequent calculations. Currently, there is little research on classifying scientific and technological texts using the semantic vector model of large models combined with traditional classification methods, and there is a lack of clear classification criteria and classification methods.

[0004] Knowledge graph fusion refers to the process of merging the same entities, attributes, or relationships from different knowledge sources to form a complete, consistent, and high-quality knowledge graph. Methods such as entity alignment and representation learning can be used to achieve the fusion of knowledge graphs based on the semantic information or structural information of the knowledge graphs.

[0005] Currently, there are some problems in the division of the industrial fields to which scientific and technological achievements belong. First, with the rapid development and cross-integration of the scientific and technological fields, many scientific and technological achievements involve multiple industrial fields and are difficult to simply be classified into a specific category. This leads to the ambiguity and uncertainty of classification, affecting the accurate evaluation and effective utilization of scientific and technological achievements. Second, there may be differences in the division of industrial fields among different institutions or experts, and there is a lack of unified standards and norms, making the classification results difficult to be widely accepted and recognized. In the fusion of academic achievement knowledge graphs and industrial knowledge graphs, there is currently a lack of relevant technical exploration, and there is a gap between academic achievements, scientific research talents and specific technical fields and industrial fields, making it difficult to select talents or retrieve academic achievement information.

[0006] To overcome these defects, the present application proposes a method and system for fusing academic and industrial knowledge graphs based on large models. Summary of the Invention

[0007] The purpose of this application is to provide a method and system for fusing academic-industry knowledge graphs based on large models, aiming to solve the problems of unclear division of academic achievements in the industrial field and complex knowledge graph fusion process in the prior art.

[0008] To achieve the above purpose, this application provides the following technical solutions:

[0009] This application provides a method for fusing academic-industry knowledge graphs based on large models, including:

[0010] Obtain academic achievement text data and industrial text data, construct an academic achievement knowledge graph based on the academic achievement text data, and construct an industrial knowledge graph based on the industrial text data;

[0011] Perform text cleaning on the academic achievement attribute information in the academic achievement knowledge graph, and use the semantic vector model of the large model to transform the processed academic achievement attribute information to obtain the academic achievement attributes represented by vectors;

[0012] Add labels to the academic achievement attributes represented by vectors as a training set, construct and train a multi-label random forest model based on the training set, and classify the academic achievements in the academic achievement knowledge graph to obtain the industrial field classification results of the academic achievements;

[0013] Through the industrial field classification results of the academic achievements, divide the scientific research talents into industrial fields and generate industrial labels;

[0014] Calculate the correlation between the industrial labels and the industrial knowledge graph, and complete the fusion of the academic achievement knowledge graph and the industrial knowledge graph based on the correlation results.

[0015] Further, in the step of obtaining academic achievement text data and industrial text data, constructing an academic achievement knowledge graph based on the academic achievement text data, and constructing an industrial knowledge graph based on the industrial text data, the following steps are specifically included:

[0016] Preprocess and annotate the academic achievement text data, identify the preprocessed academic achievement text data by fine-tuning the large model, extract the preprocessed academic achievement text data using prompt engineering, and combine to obtain academic achievement entities; construct an academic achievement knowledge graph based on the academic achievement entities;

[0017] Preprocess and annotate the industrial text data, identify the preprocessed industrial text data through fine-tuning a large model, extract the preprocessed industrial text data using prompt engineering, and combine to obtain industrial domain entities; construct an industrial knowledge graph based on the industrial domain entities.

[0018] Further, in the step of text cleaning the academic achievement attribute information in the academic achievement knowledge graph, converting the processed academic achievement attribute information using the semantic vector model of the large model to obtain the vector-represented academic achievement attributes, the following specific steps are included:

[0019] Text clean the academic achievement attribute information, remove preset irrelevant characters using regular expressions; perform stemming on the academic achievement attribute information using the NLTK library to obtain the root words;

[0020] Through the semantic vector model of the large model, convert the preprocessed academic achievement attribute information into the corresponding vector representation to obtain the vector-represented academic achievement attributes.

[0021] Further, in the step of adding labels to the vector-represented academic achievement attributes as a training set, constructing and training a multi-label random forest model based on the training set, and classifying the academic achievements in the academic achievement knowledge graph to obtain the industrial domain classification results of the academic achievements, the following specific steps are included:

[0022] Construct a multi-label random forest model based on the training set, construct several decision trees during the training process, and use the multi-label learning strategy for classification;

[0023] After the training is completed, classify and predict the academic achievements in the academic achievement knowledge graph through the trained multi-label random forest model to obtain the industrial domain classification results of the academic achievements.

[0024] Further, the industrial knowledge graph divides the industrial domain into four layers, and the academic achievement knowledge graph includes papers published by scientific research talents and patents invented.

[0025] This application provides an academic-industrial knowledge graph fusion system based on a large model, including:

[0026] Graph construction module: Obtain academic achievement text data and industrial text data, construct an academic achievement knowledge graph based on the academic achievement text data, and construct an industrial knowledge graph based on the industrial text data;

[0027] Feature extraction module: Text clean the academic achievement attribute information in the academic achievement knowledge graph, convert the processed academic achievement attribute information using the semantic vector model of the large model to obtain the vector-represented academic achievement attributes;

[0028] Training module: Add labels to the academic achievement attributes represented by the vectors as the training set, construct and train a multi-label random forest model based on the training set, classify the academic achievements in the academic achievement knowledge graph, and obtain the industrial field classification results of the academic achievements;

[0029] Partition module: Through the industrial field classification results of the academic achievements, divide the scientific research talents into industrial fields and generate industrial labels;

[0030] Graph fusion module: Calculate the correlation between the industrial labels and the industrial knowledge graph, and complete the fusion of the academic achievement knowledge graph and the industrial knowledge graph based on the correlation results.

[0031] This application provides a device, which includes a processor and a memory coupled to the processor. Among them, the memory stores program instructions for implementing a method for fusing academic-industrial knowledge graphs based on a large model; the processor is used to execute the program instructions stored in the memory to implement the fusion of academic-industrial knowledge graphs based on a large model.

[0032] This application provides a storage medium storing program instructions that can be run by a processor, and the program instructions are used to execute a method for fusing academic-industrial knowledge graphs based on a large model.

[0033] This application provides a method and system for fusing academic-industrial knowledge graphs based on a large model, and has the following beneficial effects:

[0034] (1) The large model can more accurately capture the context information in the text, thereby improving the extraction accuracy; at the same time, the large model has stronger generalization ability, can effectively handle diverse text contents and structures, reduces the dependence on a large amount of labeled data, and enables good performance even in the case of scarce data; in addition, the large model can better handle unknown entities and improve the robustness of the system;

[0035] (2) Existing methods for dividing the industrial fields of academic achievements often rely on manual annotation or division based on the institutions or journals to which the achievements belong, lack clear criteria, and lack clear definitions and demarcations of the subordinate relationships in each industrial field; this application gives a clear definition for field division by fusing the academic achievement knowledge graph and the industrial knowledge graph, realizes the analysis of academic achievement text data through the semantic vector model of the large model, and combines the multi-label random forest algorithm to realize the automatic division of the industrial fields of academic achievements, making the classification more accurate and efficient;

[0036] (3) By classifying academic achievements and scientific research talents into industrial fields, it is conducive to the realization of subsequent talent selection, industrial development planning and other applications, and enriches the information of the knowledge graph. Description of the Drawings

[0037] Figure 1 It is a schematic flow chart of a method for fusing academic-industrial knowledge graphs based on a large model according to Embodiment 1 of the present application;

[0038] Figure 2 It is a schematic structural diagram of a method for fusing academic-industrial knowledge graphs based on a large model according to Embodiment 1 of the present application;

[0039] Figure 3 It is a schematic structural diagram of the academic achievement knowledge graph according to Embodiment 1 of the present application;

[0040] Figure 4 It is a schematic structural diagram of the industrial knowledge graph according to Embodiment 1 of the present application;

[0041] Figure 5 It is a schematic structural diagram of a system for fusing academic-industrial knowledge graphs based on a large model according to Embodiment 2 of the present application;

[0042] Figure 6 It is a schematic diagram of the device structure according to Embodiment 3 of the present application;

[0043] Figure 7 It is a schematic diagram of the storage medium structure according to Embodiment 4 of the present application. Detailed Embodiments

[0044] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0045] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0046] Embodiment 1

[0047] Please refer to Figure 1 , which is a schematic flow chart of a method for fusing academic-industrial knowledge graphs based on a large model according to Embodiment 1 of the present application; the steps include:

[0048] S1: Obtain academic achievement text data and industrial text data, construct an academic achievement knowledge graph based on the academic achievement text data, and construct an industrial knowledge graph based on the industrial text data.

[0049] In this embodiment, the academic achievement text data is preprocessed and labeled. The preprocessed academic achievement text data is recognized by fine-tuning a large model, and the preprocessed academic achievement text data is extracted using prompt engineering, and an academic achievement entity is obtained by combination; an academic achievement knowledge graph is constructed based on the academic achievement entity.

[0050] The industrial text data is preprocessed and labeled, and the preprocessed industrial text data is recognized by fine-tuning a large model. According to the characteristics of the industrial text data and the task requirements, a pre-trained model is selected, such as BERT, GPT, RoBERTa, etc. A training set, a validation set, and a test set are established based on the industrial text data. According to the task characteristics and model features, fine-tuning parameters are set, such as learning rate, batch size, number of training epochs, etc. The pre-trained model is further trained, and the pre-trained model is optimized by adjusting the model weights and parameters. The fine-tuned model is evaluated using the validation set, and the model structure and parameters are adjusted according to the evaluation results until satisfactory performance is achieved.

[0051] The preprocessed industrial text data is extracted using prompt engineering, and an industrial domain entity is obtained by combination. Specifically, it includes: selecting the types, ranges, and prompt words of the extracted text information. Variable elements are introduced into the prompt words to improve the reusability and flexibility of the prompt words. The designed prompt words and the preprocessed industrial text data are input into the fine-tuned model together. The model performs reasoning and extraction based on the prompt words and the text data to generate an industrial domain entity, and an industrial knowledge graph is constructed based on the industrial domain entity.

[0052] S2: The academic achievement attribute information in the academic achievement knowledge graph is text-cleaned, and the processed academic achievement attribute information is converted using the semantic vector model of the large model to obtain the academic achievement attributes in vector representation.

[0053] In this embodiment, the academic achievement attribute information is text-cleaned. Regular expressions are used to remove preset irrelevant characters, such as non-alphanumeric characters, such as punctuation marks, special symbols, etc., and redundant spaces, line breaks, and tab characters are deleted to ensure the unity and neatness of the text format. The NLTK library is used to perform stemming on the academic achievement attribute information to remove the prefixes and suffixes of words to obtain the root words. This helps to normalize words in different forms to the same form.

[0054] Through the semantic vector model of large models, such as the BGE model, convert the preprocessed academic achievement attribute information into corresponding vector representations to obtain the academic achievement attributes in vector representation. Specifically, it includes: preprocessing the academic achievement attribute information, including data cleaning, format unification, removing irrelevant information, etc., to ensure the quality and consistency of the input data. Set up the software and hardware environment and ensure the availability of computing resources; use corresponding tools and libraries, such as the Transformers library, to load the BGE model and its related tokenizers. Encode the preprocessed academic achievement attribute information through the tokenizer and convert it into the input format of the model; input the encoded text into the BGE model to calculate the corresponding semantic vector.

[0055] S3: Add labels to the academic achievement attributes in vector representation as the training set, build and train a multi-label random forest model based on the training set, classify the academic achievements in the academic achievement knowledge graph, and obtain the industrial field classification results of the academic achievements.

[0056] In this embodiment, build a multi-label random forest model based on the training set, set the parameters of the random forest, such as the number of trees, maximum depth, minimum number of samples, etc.; build several decision trees on the training set. Each tree is trained by randomly sampling samples and features from the training set. At each node of each decision tree, divide according to the selected feature until the preset stop condition is reached.

[0057] For multi-label classification tasks, adopt multiple strategies to handle. Convert the multi-label problem into multiple binary classification problems, with each label corresponding to a binary classifier; or adopt a method based on label correlation, considering the relationship between labels. In the random forest, each tree can adopt the same or different multi-label learning strategies.

[0058] On each decision tree, train based on the training set to obtain the division rules of each node and the class labels of the leaf nodes. Integrate the prediction results of all decision trees to obtain the final prediction result. Use the validation set or test set to evaluate the trained multi-label random forest model, and calculate indicators such as accuracy, recall rate, and F1 score. Adjust the model parameters according to the evaluation results to improve the model performance. After training, classify and predict the academic achievements in the academic achievement knowledge graph through the trained multi-label random forest model to obtain the industrial field classification results of the academic achievements.

[0059] S4: Through the industrial field classification results of the academic achievements, divide the scientific research talents into industrial fields and generate industrial labels.

[0060] In this embodiment, according to the classification results of the industrial fields of the academic achievements published by scientific research talents, assuming that there are N academic achievements in total for a talent, and these achievements are involved in m industrial fields. If the number of achievements in a certain field is greater than N / m, then the label of this industrial field is added to the talent.

[0061] S5: Calculate the correlation degree between the industrial label and the industrial knowledge graph, and complete the fusion of the academic achievement knowledge graph and the industrial knowledge graph based on the correlation degree result.

[0062] In this embodiment, a matching method is used to calculate the correlation degree between the industrial label and the industrial knowledge graph. That is, by matching the labels of the industrial knowledge graph with the industrial labels, and according to the positions and quantities of the labels in the industrial knowledge graph that are matched, the correlation degree between the industrial label and the labels of the industrial knowledge graph is calculated. The correlation degree here represents the degree of mastery of the technical field by the scientific research talent.

[0063] Assume that there are n labels at a certain third level in the industrial knowledge graph, and a scientific research talent has m labels related to it. If m / n is greater than 1 / 2, then the industrial label of the upper level is linked to the scientific research talent, otherwise only the industrial label of the same level is linked to it. At this time, the correlation degree between the industrial label of the upper level and the talent is m / n, and the correlation degree between the industrial label of the same level and the talent is 1. When m / n is equal to 1, or other correlation degrees equal to 1 indicate that there is a direct corresponding relationship between this technical field and the scientific research talent or the scientific research talent has fully mastered all related technical fields in this field. Through the correlation degree calculation, higher-level industrial labels are linked to the scientific research talent.

[0064] For example, the industrial labels of a certain scientific research talent include "sentiment analysis", "machine translation", "knowledge graph", and "image segmentation". In the industrial knowledge graph, there is a secondary classification label "natural language processing", and its sub-labels include "sentiment analysis", "machine translation", "knowledge graph", "named entity recognition", and "text classification"; there is a secondary classification label "computer vision", and it has 5 sub-labels. Through matching, if the industrial label occupies more than half of the sub-labels of "natural language processing", then the label "natural language processing" is linked to the scientific research talent to form an associated node, and the association degree between the scientific research talent and "natural language processing" is 3 / 5. On the other hand, "image segmentation" belongs to the third-level classification and is attributed to "computer vision", and its association degree with "computer vision" is 1 / 5, which does not exceed the specified 1 / 2. It is considered that the association degree between the scientific research talent and "computer vision" is insufficient, so only "image segmentation" is linked to the scientific research talent, and its association degree is 1. Through the calculation of the association degree, a higher-level industrial label is established for the scientific research talent, and it is considered that the higher-level industrial label can better highlight the contribution of the scientific research talent in the industrial field. By establishing the association between academic achievements, industrial labels, and scientific research talents, a link is established between the two knowledge graphs to complete the knowledge graph fusion.

[0065] Please refer to Figure 2 , which is a schematic structural diagram of a method for fusing academic-industrial knowledge graphs based on a large model in Embodiment 1 of this application. Specifically, it includes: first, preprocessing and annotating the obtained original documents, that is, the academic achievement text data and industrial text data; improving the accuracy of entity recognition and extraction through a large language model (LLM) to obtain industrial field entities and academic achievement entities, and constructing an industrial knowledge graph based on the industrial field entities; subsequently, using the LLM to convert the academic achievement attributes into vector representations, and using the multi-label random forest algorithm to divide the academic achievements into industrial fields, improving the accuracy of classification; finally, dividing the industrial labels according to the industrial field classification results of the academic achievements, and then calculating the association degree between the industrial labels and the industrial knowledge graph, realizing the fusion of the knowledge graphs.

[0066] In summary, the entity extraction technology based on the large model in this Embodiment 1 significantly improves the accuracy of knowledge graph construction. First, the large model is used to preprocess and annotate academic achievement text data and industrial text data. The accuracy of entity recognition is improved by fine-tuning the large model, and an academic achievement knowledge graph and an industrial knowledge graph are constructed in combination with prompt engineering. Subsequently, the semantic vector model of the large model is used to convert the academic achievement attributes into vector representations, and the multi-label random forest algorithm is used to classify the academic achievements into industrial fields, improving the accuracy of classification. Finally, industrial labels are divided according to the industrial field classification results of academic achievements, and the correlation between industrial labels and the industrial knowledge graph is calculated to achieve the fusion of knowledge graphs. It provides a clear industrial field division for academic achievements and scientific research talents, which is beneficial to subsequent applications such as talent selection and industrial development planning, and greatly enriches the information content of the knowledge graph.

[0067] Embodiment 2

[0068] Please refer to Figure 5 , which is a schematic structural diagram of an academic-industrial knowledge graph fusion system based on a large model according to Embodiment 2 of this application; the specific content includes:

[0069] Knowledge graph construction module: Obtain academic achievement text data and industrial text data, construct an academic achievement knowledge graph based on the academic achievement text data, and construct an industrial knowledge graph based on the industrial text data;

[0070] Feature extraction module: Perform text cleaning on the academic achievement attribute information in the academic achievement knowledge graph, and use the semantic vector model of the large model to convert the processed academic achievement attribute information to obtain the academic achievement attributes represented by vectors;

[0071] Training module: Add labels to the academic achievement attributes represented by vectors as a training set, construct and train a multi-label random forest model based on the training set, classify the academic achievements in the academic achievement knowledge graph, and obtain the industrial field classification results of the academic achievements;

[0072] Division module: Divide the scientific research talents into industrial fields through the industrial field classification results of the academic achievements to generate industrial labels;

[0073] Knowledge graph fusion module: Calculate the correlation between the industrial labels and the industrial knowledge graph, and complete the fusion of the academic achievement knowledge graph and the industrial knowledge graph based on the correlation calculation results.

[0074] Embodiment 3

[0075] Please refer to Figure 6 , which is a schematic structural diagram of the device according to Embodiment 3 of this application. The device 50 includes a processor 51 and a memory 52 coupled to the processor 51.

[0076] The memory 52 stores program instructions for implementing the above-mentioned method for fusing an academic-industrial knowledge graph based on a large model.

[0077] The processor 51 is configured to execute the program instructions stored in the memory 52 to implement a method for fusing an academic-industrial knowledge graph based on a large model.

[0078] Among them, the processor 51 can also be referred to as a CPU (Central Processing Unit).

[0079] The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0080] Embodiment 4

[0081] Please refer to Figure 7 , which is a schematic structural diagram of the storage medium according to Embodiment 4 of the present application. The storage medium of the embodiment of the present application stores a program file 61 capable of implementing all the above methods. Among them, the program file 61 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods according to various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes, or devices such as a computer, a server, a mobile phone, or a tablet.

[0082] It should be noted that in this article, the term "comprises", "comprising" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.

[0083] The foregoing are only preferred embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included in the patent protection scope of the present application.

[0084] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present application. The scope of the present application is defined by the appended claims and their equivalents.

[0085] Certainly, the present invention can also have other various implementation manners. Based on this implementation manner, other implementation manners obtained by those of ordinary skill in the art without any creative work belong to the scope protected by the present invention.

Claims

1. A method for fusing an academic industry knowledge graph based on a large model, characterized in that, Including: Obtain academic achievement text data and industrial text data, construct an academic achievement knowledge graph based on the academic achievement text data, and construct an industrial knowledge graph based on the industrial text data; Perform text cleaning on the academic achievement attribute information in the academic achievement knowledge graph, and use the semantic vector model of the large model to convert the processed academic achievement attribute information to obtain the academic achievement attributes represented by vectors; specifically including: performing text cleaning on the academic achievement attribute information, and using regular expressions to remove preset irrelevant characters; using the NLTK library to perform stemming on the academic achievement attribute information to obtain the root words; through the semantic vector model of the large model, convert the preprocessed academic achievement attribute information into the corresponding vector representation to obtain the academic achievement attributes represented by vectors; Add labels to the academic achievement attributes represented by vectors as the training set, construct and train a multi-label random forest model based on the training set, and classify the academic achievements in the academic achievement knowledge graph to obtain the industrial field classification results of the academic achievements; specifically including: constructing a multi-label random forest model based on the training set, constructing several decision trees during the training process, and using a multi-label learning strategy for classification; after the training is completed, classify and predict the academic achievements in the academic achievement knowledge graph through the trained multi-label random forest model to obtain the industrial field classification results of the academic achievements; Based on the industrial field classification results of the academic achievements, divide the scientific research talents into industrial fields and generate industrial labels; Calculate the correlation between the industrial labels and the industrial knowledge graph, and complete the fusion of the academic achievement knowledge graph and the industrial knowledge graph based on the correlation results; specifically including the following steps: Calculate the correlation between the industrial labels and the labels of the industrial knowledge graph by matching the labels of the industrial knowledge graph with the industrial labels and according to the positions and quantities of the labels in the industrial knowledge graph that are matched; In the step of obtaining academic achievement text data and industrial text data, constructing an academic achievement knowledge graph based on the academic achievement text data, and constructing an industrial knowledge graph based on the industrial text data, specifically including the following steps: Preprocess and annotate the academic achievement text data, identify the preprocessed academic achievement text data by fine-tuning the large model, extract the preprocessed academic achievement text data using prompt engineering, and combine to obtain academic achievement entities; construct an academic achievement knowledge graph based on the academic achievement entities; preprocess and annotate the industrial text data, identify the preprocessed industrial text data by fine-tuning the large model, extract the preprocessed industrial text data using prompt engineering, and combine to obtain industrial field entities; construct an industrial knowledge graph based on the industrial field entities.

2. The method for fusing an academic industry knowledge graph based on a large model according to claim 1, wherein The industrial knowledge graph divides the industrial fields into four layers, and the academic achievement knowledge graph includes papers published by scientific research talents and patents invented.

3. A system for a method of integrating an academic industry knowledge graph based on a large model according to claim 1, characterized in that, Including: Graph Construction Module: Obtain academic achievement text data and industrial text data, construct an academic achievement knowledge graph based on the academic achievement text data, and construct an industrial knowledge graph based on the industrial text data; Feature Extraction Module: Perform text cleaning on the academic achievement attribute information in the academic achievement knowledge graph, and use the semantic vector model of the large model to convert the processed academic achievement attribute information to obtain the academic achievement attributes represented by vectors; Training Module: Add labels to the academic achievement attributes represented by vectors as the training set, construct and train a multi-label random forest model based on the training set, classify the academic achievements in the academic achievement knowledge graph, and obtain the industrial field classification results of the academic achievements; Division Module: Divide scientific research talents into industrial fields through the industrial field classification results of the academic achievements to generate industrial labels; Graph Fusion Module: Calculate the correlation degree between the industrial labels and the industrial knowledge graph, and complete the fusion of the academic achievement knowledge graph and the industrial knowledge graph based on the correlation degree results.

4. A device, characterized in that, The device includes a processor and a memory coupled to the processor. Among them, the memory stores program instructions for implementing a method for fusing academic and industrial knowledge graphs based on a large model according to any one of claims 1-2; the processor is used to execute the program instructions stored in the memory to implement a method for fusing academic and industrial knowledge graphs based on a large model.

5. A storage medium, characterized in that, Store program instructions that can be run by a processor, and the program instructions are used to execute a method for fusing academic and industrial knowledge graphs based on a large model according to any one of claims 1-2.

Citation Information

Patent Citations

  • Policy recommendation method based on knowledge graph and related equipment thereof

    CN114398477A

  • Knowledge graph intelligent construction model based on knowledge transfer learning strategy

    CN118228810A

  • Intelligent pediatric disease diagnosis auxiliary system

    CN118538399A