Subject knowledge graph construction method and related device
By using the cosine similarity calculation method to screen the knowledge module data in the course of subject knowledge graph completion, the problem of excessive redundant information in the knowledge graph in the existing technology is solved, and the accuracy and relevance of the knowledge graph are improved.
Patent Information
- Application Number
- CN202510032750.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing discipline knowledge graph completion model lacks clear limits or standards, resulting in the completed map that may contain a large amount of information that is not directly related to the subject topic or redundant, reducing the accuracy of the knowledge graph.
The cosine similarity calculation method is used to filter the knowledge module data in the completed knowledge graph. By extracting keywords, converting them into semantic vectors and calculating cosine similarity, the data is retained or eliminated according to the preset threshold, and a more accurate knowledge graph is constructed.
Improve the accuracy of the generated knowledge graph, ensure that the information contained in the graph is closely related to the subject topic, reduce redundant information, and improve the quality and relevance of the knowledge graph.
Smart Images

Figure CN119940507A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and specifically to a subject knowledge graph construction method and related devices. Background Art
[0002] With the rapid development of informatization and digitalization, the field of education is undergoing unprecedented changes. As a structured knowledge representation method, subject knowledge graphs have shown great potential in the integration of educational resources, the recommendation of personalized learning paths, and the evaluation of teaching effectiveness. In order to enrich the content of knowledge graphs, existing subject knowledge graphs can be completed based on existing knowledge data through completion models. They can take advantage of the advantages of graph neural networks in capturing complex relationships between nodes and information dissemination, effectively add missing knowledge nodes and relationships to the graph, and thus enrich the content of knowledge graphs.
[0003] However, existing models often lack a clear limit or standard when completing knowledge graphs, resulting in the completed graph containing a large amount of information that is not directly related to the subject or is redundant, and the final generated knowledge graph has low accuracy. Summary of the invention
[0004] The embodiment of the present application provides a subject knowledge graph construction method and related devices. The cosine similarity calculation method can further eliminate knowledge module data that is not directly related to the subject topic or is redundant on the basis of enriching the initial subject knowledge graph, thereby improving the accuracy of the finally generated knowledge graph.
[0005] A first aspect of an embodiment of the present application provides a method for constructing a subject knowledge graph, the method comprising:
[0006] Acquire first knowledge module data;
[0007] Constructing an initial subject knowledge graph based on the first knowledge module data;
[0008] Using a graph neural network model to complete the initial subject knowledge graph to obtain a second subject knowledge graph;
[0009] Extracting second knowledge module data from the second subject knowledge graph;
[0010] The second knowledge module data is screened by using a cosine similarity calculation method to obtain third knowledge module data;
[0011] Based on the third knowledge module data, a third subject knowledge graph is constructed.
[0012] Furthermore, the method of using the cosine similarity calculation method to screen the second knowledge module data to obtain the third knowledge module data includes:
[0013] Extracting keywords of each knowledge module data from the second knowledge module data;
[0014] Convert the keywords of each knowledge module data into semantic vectors to obtain a semantic vector data set;
[0015] Performing cosine similarity calculation on the semantic vectors in the semantic vector data set in sequence to obtain a cosine similarity calculation result;
[0016] The second knowledge module data is screened according to the cosine similarity calculation result and a preset cosine similarity threshold to obtain third knowledge module data.
[0017] Furthermore, the second knowledge module data is screened according to the cosine similarity calculation result and a preset cosine similarity threshold to obtain the third knowledge module data, including:
[0018] If the cosine similarity calculation result is greater than the preset cosine similarity threshold, retaining the knowledge module data corresponding to the current cosine similarity calculation result;
[0019] If the cosine similarity calculation result is not greater than the preset cosine similarity threshold, the knowledge module data corresponding to the current cosine similarity calculation result is removed.
[0020] Furthermore, the graph neural network model is used to complete the initial subject knowledge graph to obtain a second subject knowledge graph, including:
[0021] Extracting the first knowledge module data from the initial subject knowledge graph;
[0022] Determine a training set and a semantic graph according to the first knowledge module data;
[0023] According to the training set and the semantic graph, the NeuralLP complex reasoning in the graph neural network is used to complete the initial subject knowledge graph to obtain a second subject knowledge graph.
[0024] Furthermore, determining a training set according to the first knowledge module data includes:
[0025] Constructing a connection tag according to the first knowledge module data;
[0026] Dividing the first knowledge module data into triple data;
[0027] A training set is constructed according to the connection labels and triplet data.
[0028] Further, determining a semantic graph according to the first knowledge module data includes:
[0029] Confirming the extended course data according to the first knowledge module data;
[0030] Extracting lecture data, video data and test data from the extended course data;
[0031] A semantic graph is determined based on the lecture data, video data and test data.
[0032] In this example, the second subject knowledge graph is first obtained by using a graph neural network to complete the constructed initial subject knowledge graph, and then the cosine similarity calculation method is used to screen the knowledge module data in the second subject knowledge graph to obtain the final third subject knowledge graph. On the basis of enriching the initial subject knowledge graph, the knowledge module data that is not directly related to the subject topic or is redundant can be further eliminated, thereby improving the accuracy of the final generated knowledge graph, so as to solve the technical problem that the existing models often lack a clear limit or standard when completing the knowledge graph, resulting in a large amount of information that is not directly related to the subject topic or is redundant in the completed graph, and the final generated knowledge graph has low accuracy.
[0033] A second aspect of an embodiment of the present application provides a subject knowledge graph construction device, the device comprising:
[0034] A first acquisition unit, used to acquire first knowledge module data;
[0035] A first processing unit, configured to construct an initial subject knowledge graph according to the first knowledge module data;
[0036] A second processing unit is used to complete the initial subject knowledge graph using a graph neural network model to obtain a second subject knowledge graph;
[0037] A second acquisition unit, used to extract second knowledge module data from the second subject knowledge graph;
[0038] A third processing unit is used to filter the second knowledge module data by using a cosine similarity calculation method to obtain third knowledge module data;
[0039] The fourth processing unit is used to construct a third subject knowledge graph based on the third knowledge module data. Further, in the aspect of using the graph neural network model to complete the initial subject knowledge graph to obtain the second subject knowledge graph, the second processing unit is used to:
[0040] Extracting the first knowledge module data from the initial subject knowledge graph;
[0041] Determine a training set and a semantic graph according to the first knowledge module data;
[0042] According to the training set and the semantic graph, the NeuralLP complex reasoning in the graph neural network is used to complete the initial subject knowledge graph to obtain a second subject knowledge graph.
[0043] Further, in the aspect of determining a training set according to the first knowledge module data, the second processing unit is further configured to:
[0044] Constructing a connection tag according to the first knowledge module data;
[0045] Dividing the first knowledge module data into triple data;
[0046] A training set is constructed according to the connection labels and triplet data.
[0047] Further, in the aspect of determining the semantic graph according to the first knowledge module data, the second processing unit is further used for:
[0048] Confirming the extended course data according to the first knowledge module data;
[0049] Extracting lecture data, video data and test data from the extended course data;
[0050] A semantic graph is determined based on the lecture data, video data and test data.
[0051] Further, in the aspect of using the cosine similarity calculation method to screen the second knowledge module data to obtain the third knowledge module data, the third processing unit is used to:
[0052] Extracting keywords of each knowledge module data from the second knowledge module data;
[0053] Convert the keywords of each knowledge module data into semantic vectors to obtain a semantic vector data set;
[0054] Performing cosine similarity calculation on the semantic vectors in the semantic vector data set in sequence to obtain a cosine similarity calculation result;
[0055] The second knowledge module data is screened according to the cosine similarity calculation result and a preset cosine similarity threshold to obtain third knowledge module data.
[0056] Further, in the aspect of filtering the second knowledge module data according to the cosine similarity calculation result and the preset cosine similarity threshold to obtain the third knowledge module data, the third processing unit is further used to:
[0057] If the cosine similarity calculation result is greater than the preset cosine similarity threshold, retaining the knowledge module data corresponding to the current cosine similarity calculation result;
[0058] If the cosine similarity calculation result is not greater than the preset cosine similarity threshold, the knowledge module data corresponding to the current cosine similarity calculation result is removed.
[0059] A third aspect of an embodiment of the present application provides a terminal, comprising a processor, an input device, an output device and a memory, wherein the processor, input device, output device and memory are interconnected, wherein the memory is used to store a computer program, and the computer program includes program instructions, and the processor is configured to call the program instructions to execute the step instructions in the subject knowledge graph construction method of the first aspect of the embodiment of the present application.
[0060] The fourth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the above-mentioned computer-readable storage medium stores a computer program for electronic data exchange, wherein the above-mentioned computer program enables a computer to execute part or all of the steps described in the subject knowledge graph construction method in the first aspect of the embodiment of the present application.
[0061] The fifth aspect of the embodiments of the present application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps described in the subject knowledge graph construction method in the first aspect of the embodiments of the present application. The computer program product can be a software installation package. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0063] Figure 1 A schematic diagram of the overall process of a subject knowledge graph construction method is provided for an embodiment of the present application;
[0064] Figure 2A schematic diagram of the completion process of an initial subject knowledge graph of a subject knowledge graph construction method is provided for an embodiment of the present application;
[0065] Figure 3 A schematic diagram of the specific steps of using a cosine similarity calculation method to screen the second knowledge module data in the completed second subject knowledge graph in a subject knowledge graph construction method provided in an embodiment of the present application;
[0066] Figure 4 A schematic diagram of the structure of a subject knowledge graph construction device is provided for an embodiment of the present application;
[0067] Figure 5 A schematic diagram of the structure of a terminal provided in an embodiment of the present application;
[0068] Reference numerals:
[0069] First acquisition unit-1, first processing unit-2, second processing unit-3, second acquisition unit-4, third processing unit-5, fourth processing unit-6. DETAILED DESCRIPTION
[0070] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0071] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.
[0072] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0073] In order to better understand a method for constructing a subject knowledge graph provided by an embodiment of the present application, the following first briefly introduces the scenario of applying the method for constructing a subject knowledge graph. In order to enrich the content of the knowledge graph, the existing subject knowledge graph can be supplemented by a completion model based on the existing knowledge data, and can effectively add missing knowledge nodes and relationships to the graph by taking advantage of the graph neural network in capturing complex relationships between nodes and information propagation, thereby enriching the content of the knowledge graph. However, the existing models often lack a clear limit or standard when completing the knowledge graph, resulting in the completed graph containing a large amount of information that is not directly related to the subject or redundant, and the final generated knowledge graph has low accuracy. Specifically, although the graph neural network model can infer new knowledge nodes and relationships based on the existing knowledge data, it does not have the ability to judge whether this information is closely related to the subject. Therefore, in the process of completion, the model may introduce some knowledge modules that are irrelevant to the subject or have low correlation, thereby reducing the accuracy and relevance of the knowledge graph. This will not only affect the learning effect of learners, but also may mislead their understanding and application of subject knowledge.
[0074] The subject knowledge graph construction method is applied to the subject knowledge graph construction device. Figure 1 A schematic diagram of the overall process of constructing a subject knowledge graph is shown. Figure 1 As shown, including:
[0075] S1. Obtain the first knowledge module data, and obtain the knowledge module data of the relevant subject from the subject curriculum standards, textbooks, and educational research literature.
[0076] S2. Based on the first knowledge module data, construct an initial subject knowledge graph, and use digital tools to convert the collected first knowledge module data into entities in the knowledge graph and the relationships between the entities.
[0077] S3. Use a graph neural network model to complete the initial subject knowledge graph to obtain a second subject knowledge graph. Specifically, this embodiment preferably includes:
[0078] S301. Extract the first knowledge module data from the initial subject knowledge graph.
[0079] S302: Determine a training set and a semantic graph according to the first knowledge module data. Specifically, first construct a training set according to the first knowledge module data, including:
[0080] S3021. Construct a connection tag based on the first knowledge module data.
[0081] S3022. Divide the first knowledge module data into triple data.
[0082] S3023. Construct a training set based on the connection labels and triplet data.
[0083] Then, according to the first knowledge module data, a semantic graph is determined, including:
[0084] S3024. Confirm the extended course data based on the first knowledge module data.
[0085] S3025. Extracting lecture data, video data and test data from the extended course data.
[0086] S3026. Determine a semantic graph based on the lecture data, video data and test data.
[0087] S303. Based on the training set and the semantic graph, NeuralLP complex reasoning in the graph neural network is used to complete the initial subject knowledge graph to obtain a second subject knowledge graph.
[0088] In this example, the NeuralLP model can identify the deficient nodes in the student's knowledge graph. Then, the model applies complex reasoning, combines the student's learning status and logical rules, finds out the knowledge points that the student has not yet mastered or has not mastered firmly, and selects the course resources associated with these deficient nodes from the knowledge base, such as video lectures, reading materials, exercises, and case studies, so as to fill the initial knowledge graph with these knowledge data.
[0089] S4. Extract the second knowledge module data from the second subject knowledge graph.
[0090] S5. Use a cosine similarity calculation method to screen the second knowledge module data to obtain third knowledge module data. Figure 3 A schematic diagram of the specific steps of using a cosine similarity calculation method to screen the second knowledge module data in the completed second subject knowledge graph in a subject knowledge graph construction method is shown, including:
[0091] S501. Extract keywords of each knowledge module data from the second knowledge module data.
[0092] S502, converting the keywords of each knowledge module data into semantic vectors to obtain a semantic vector data set, and using the WordVec model to convert the keywords of each knowledge module data into semantic vectors.
[0093] S503 , performing cosine similarity calculation on the semantic vectors in the semantic vector data set in sequence to obtain a cosine similarity calculation result.
[0094] S504: Filter the second knowledge module data according to the cosine similarity calculation result and the preset cosine similarity threshold to obtain third knowledge module data, specifically including:
[0095] S504A: If the cosine similarity calculation result is greater than the preset cosine similarity threshold, retain the knowledge module data corresponding to the current cosine similarity calculation result.
[0096] S504B: If the cosine similarity calculation result is not greater than the preset cosine similarity threshold, the knowledge module data corresponding to the current cosine similarity calculation result is removed.
[0097] In this example, the Word2Vec model plus the cosine correlation calculation method are used to compare and analyze the converted entities with the entity content formulated by the teacher based on the course content, and the entity texts with correlation greater than 0.7 are selected as the entity data of the knowledge graph. At the same time, the text content with correlation less than 0.7 is excluded to improve the accuracy of the knowledge graph.
[0098] S6. Construct a third subject knowledge graph based on the third knowledge module data.
[0099] In this example, the second subject knowledge graph is first obtained by using a graph neural network to complete the constructed initial subject knowledge graph, and then the cosine similarity calculation method is used to screen the knowledge module data in the second subject knowledge graph to obtain the final third subject knowledge graph. On the basis of enriching the initial subject knowledge graph, the knowledge module data that is not directly related to the subject topic or is redundant can be further eliminated, thereby improving the accuracy of the final generated knowledge graph, so as to solve the technical problem that the existing models often lack a clear limit or standard when completing the knowledge graph, resulting in a large amount of information that is not directly related to the subject topic or is redundant in the completed graph, and the final generated knowledge graph has low accuracy.
[0100] In line with the above, see Figure 4 , Figure 4 The present invention provides a schematic diagram of a device for constructing a subject knowledge graph. Figure 4 As shown, the device comprises:
[0101] The first acquisition unit 1 is used to acquire first knowledge module data.
[0102] The first processing unit 2 is used to construct an initial subject knowledge graph based on the first knowledge module data.
[0103] The second processing unit 3 is used to use a graph neural network model to complete the initial subject knowledge graph to obtain a second subject knowledge graph.
[0104] The second acquisition unit 4 is used to extract second knowledge module data from the second subject knowledge graph.
[0105] The third processing unit 5 is used to filter the second knowledge module data by using a cosine similarity calculation method to obtain third knowledge module data.
[0106] The fourth processing unit 6 is used to construct a third subject knowledge graph based on the third knowledge module data.
[0107] In a possible implementation, in the aspect of using a graph neural network model to complete the initial subject knowledge graph to obtain a second subject knowledge graph, the second processing unit 3 is used to:
[0108] The first knowledge module data is extracted from the initial subject knowledge graph.
[0109] According to the first knowledge module data, a training set and a semantic graph are determined.
[0110] According to the training set and the semantic graph, the NeuralLP complex reasoning in the graph neural network is used to complete the initial subject knowledge graph to obtain a second subject knowledge graph.
[0111] In a possible implementation, the second processing unit 3 is further configured to:
[0112] Construct a connection tag based on the first knowledge module data.
[0113] Dividing the first knowledge module data into triple data;
[0114] A training set is constructed according to the connection labels and triplet data.
[0115] In a possible implementation, the second processing unit 3 is further configured to:
[0116] Confirming the extended course data according to the first knowledge module data;
[0117] Extracting lecture data, video data and test data from the extended course data;
[0118] A semantic graph is determined based on the lecture data, video data and test data.
[0119] In a possible implementation, in the aspect of using the cosine similarity calculation method to screen the second knowledge module data to obtain the third knowledge module data, the third processing unit 5 is used to:
[0120] Keywords of each knowledge module data are extracted from the second knowledge module data.
[0121] The keywords of the data of each knowledge module are converted into semantic vectors to obtain a semantic vector data set.
[0122] The cosine similarity calculation is performed on the semantic vectors in the semantic vector data set in sequence to obtain a cosine similarity calculation result.
[0123] The second knowledge module data is screened according to the cosine similarity calculation result and a preset cosine similarity threshold to obtain third knowledge module data.
[0124] In a possible implementation, in the aspect of filtering the second knowledge module data according to the cosine similarity calculation result and the preset cosine similarity threshold to obtain the third knowledge module data, the third processing unit 5 is further used to:
[0125] If the cosine similarity calculation result is greater than the preset cosine similarity threshold, retaining the knowledge module data corresponding to the current cosine similarity calculation result;
[0126] If the cosine similarity calculation result is not greater than the preset cosine similarity threshold, the knowledge module data corresponding to the current cosine similarity calculation result is removed.
[0127] For the above embodiments, please refer to Figure 5 , Figure 5 A schematic diagram of the structure of a terminal provided in an embodiment of the present application, as shown in the figure, includes a processor, an input device, an output device and a memory, the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, the processor is configured to call the program instructions, and the program includes instructions for executing the following steps;
[0128] Acquire first knowledge module data;
[0129] Constructing an initial subject knowledge graph based on the first knowledge module data;
[0130] Using a graph neural network model to complete the initial subject knowledge graph to obtain a second subject knowledge graph;
[0131] Extracting second knowledge module data from the second subject knowledge graph;
[0132] The second knowledge module data is screened by using a cosine similarity calculation method to obtain third knowledge module data;
[0133] Based on the third knowledge module data, a third subject knowledge graph is constructed.
[0134] In this example, a second subject knowledge graph is obtained by using a graph neural network to complete the constructed initial subject knowledge graph, and then the cosine similarity calculation method is used to screen the knowledge module data in the second subject knowledge graph to obtain the final third subject knowledge graph. On the basis of enriching the initial subject knowledge graph, the knowledge module data that is not directly related to the subject topic or is redundant can be further eliminated, thereby improving the accuracy of the final generated knowledge graph, so as to solve the technical problem that the existing model often lacks a clear limit or standard when completing the knowledge graph, resulting in a large amount of information that is not directly related to the subject topic or is redundant in the completed graph, and the final generated knowledge graph has low accuracy.
[0135] The above mainly introduces the scheme of the embodiment of the present application from the perspective of the execution process on the method side. It is understandable that in order to realize the above functions, the terminal includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments provided herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0136] The embodiment of the present application can divide the terminal into functional units according to the above method example. For example, each functional unit can be divided according to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0137] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any subject knowledge graph construction method recorded in the above method embodiments.
[0138] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program enables a computer to execute part or all of the steps of any subject knowledge graph construction method recorded in the above method embodiments.
[0139] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0140] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0141] In the several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical or other forms.
[0142] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] In addition, the functional units in the various embodiments of the application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software program modules.
[0144] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, disk or optical disk and other media that can store program codes.
[0145] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, which can include: a flash drive, a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.
[0146] The embodiments of the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for constructing a subject knowledge graph, characterized in that: include: Acquiring first knowledge module data; Constructing an initial subject knowledge graph based on the first knowledge module data; Using a graph neural network model to complete the initial subject knowledge graph to obtain a second subject knowledge graph; Extracting second knowledge module data from the second subject knowledge graph; The second knowledge module data is screened by using a cosine similarity calculation method to obtain third knowledge module data; Based on the third knowledge module data, a third subject knowledge graph is constructed.
2. The subject knowledge graph construction method according to claim 1, characterized in that: The method of using the cosine similarity calculation method to screen the second knowledge module data to obtain the third knowledge module data includes: Extracting keywords of each knowledge module data from the second knowledge module data; Convert the keywords of each knowledge module data into semantic vectors to obtain a semantic vector data set; Performing cosine similarity calculation on the semantic vectors in the semantic vector data set in sequence to obtain a cosine similarity calculation result; The second knowledge module data is screened according to the cosine similarity calculation result and a preset cosine similarity threshold to obtain third knowledge module data.
3. The subject knowledge graph construction method according to claim 2, characterized in that: The step of filtering the second knowledge module data according to the cosine similarity calculation result and the preset cosine similarity threshold to obtain the third knowledge module data includes: If the cosine similarity calculation result is greater than the preset cosine similarity threshold, retaining the knowledge module data corresponding to the current cosine similarity calculation result; If the cosine similarity calculation result is not greater than the preset cosine similarity threshold, the knowledge module data corresponding to the current cosine similarity calculation result is removed.
4. The subject knowledge graph construction method according to claim 1, characterized in that: The method of using a graph neural network model to complete the initial subject knowledge graph to obtain a second subject knowledge graph includes: Extracting the first knowledge module data from the initial subject knowledge graph; Determine a training set and a semantic graph according to the first knowledge module data; According to the training set and the semantic graph, the NeuralLP complex reasoning in the graph neural network is used to complete the initial subject knowledge graph to obtain a second subject knowledge graph.
5. The subject knowledge graph construction method according to claim 4, characterized in that: The step of determining a training set according to the first knowledge module data includes: Constructing a connection tag according to the first knowledge module data; Dividing the first knowledge module data into triple data; A training set is constructed according to the connection labels and triplet data.
6. The subject knowledge graph construction method according to claim 4, characterized in that: Determining a semantic graph according to the first knowledge module data includes: Confirming the extended course data according to the first knowledge module data; Extracting lecture data, video data and test data from the extended course data; A semantic graph is determined based on the lecture data, video data and test data.
7. A device for constructing a subject knowledge graph, characterized in that: The device comprises: A first acquisition unit, used to acquire first knowledge module data; A first processing unit, configured to construct an initial subject knowledge graph according to the first knowledge module data; A second processing unit is used to complete the initial subject knowledge graph using a graph neural network model to obtain a second subject knowledge graph; A second acquisition unit, used to extract second knowledge module data from the second subject knowledge graph; A third processing unit is used to filter the second knowledge module data by using a cosine similarity calculation method to obtain third knowledge module data; The fourth processing unit is used to construct a third subject knowledge graph based on the third knowledge module data.
8. The subject knowledge graph construction device according to claim 7, characterized in that: In the aspect of using the cosine similarity calculation method to screen the second knowledge module data to obtain the third knowledge module data, the third processing unit is used to: Extracting keywords of each knowledge module data from the second knowledge module data; Convert the keywords of each knowledge module data into semantic vectors to obtain a semantic vector data set; Performing cosine similarity calculation on the semantic vectors in the semantic vector data set in sequence to obtain a cosine similarity calculation result; The second knowledge module data is screened according to the cosine similarity calculation result and a preset cosine similarity threshold to obtain third knowledge module data.
9. A terminal, characterized in that: The method comprises a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 11.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 11.