Multimodal Resource Association Generation Method for Subject Knowledge Points Based on Knowledge Graph

The method automates multi-modal resource association in knowledge graphs, addressing labor and efficiency issues in traditional methods, enhancing educational experiences through intelligent and efficient knowledge graph construction.

CN119622359BActive Publication Date: 2025-07-15GUOKAI ONLINE EDUCATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411750005.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-07-15
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

Traditional knowledge graph construction relies on manual annotation, which has problems such as high labor costs, lag in knowledge updates and poor scalability. Especially in the relationship between multimodal resource, there is data heterogeneity, labeling difficulties, information redundancy and architecture complexity, which limits its wide application.

Method used

The multi-modal resource association generation method of subject knowledge points based on knowledge graphs is adopted, and a variety of resources are automatically collected through the data acquisition module, and the knowledge points are automatically associated with natural language processing, multi-modal large model analysis and graph construction modules are used to provide personalized learning paths in combination with the user interface to realize automatic association and dynamic update of multimedia resources.

Benefits of technology

It improves the efficiency of knowledge graph construction, reduces labor costs, realizes the speed and efficiency of knowledge updates, provides personalized learning experience and rich multimedia resources, and promotes more efficient learning and teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622359B_ABST
    Figure CN119622359B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method for generating multi-modal resource associations for subject knowledge points based on a knowledge graph, including: obtaining knowledge data of different modalities, where the different modalities include at least one of text, video, audio, and image; associating the knowledge data of different modalities according to the belonging knowledge points; in the knowledge graph with knowledge points as nodes, displaying the mutually associated multi-modal knowledge data as multimedia resource links belonging to the knowledge points. This embodiment can realize the automatic association of multi-modal resources in the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph. Background Art

[0002] Knowledge points based on knowledge graphs can bring significant benefits to education, scientific research, and other fields. By graphically presenting the knowledge structure, it makes the learning context clear, promotes personalized learning and interdisciplinary connections. At the same time, knowledge graphs can integrate diverse resources such as videos, audios, and pictures, enriching the learning experience. Teachers can monitor students' mastery in real time, improve teaching efficiency, and adjust teaching strategies as needed. In addition, the vivid graph display can also stimulate students' learning interest, enhance their learning motivation, and overall improve teaching quality.

[0003] However, traditional knowledge graph construction usually relies on manual annotation and expert knowledge, with limitations such as high labor costs, lagging knowledge updates, and poor scalability. Especially in associating multi-modal resources (such as videos, audios, pictures), there are more technical difficulties such as data heterogeneity, annotation difficulties, information redundancy, architecture complexity, and computational resource consumption, resulting in a more complex and resource-intensive knowledge graph construction process and limiting its wide application.

[0004] Therefore, there is an urgent need for a method that can automatically associate multi-modal resources in a knowledge graph to reduce labor costs and improve the speed and efficiency of knowledge update and expansion. Patent CN112015955A discloses a multi-modal data association method and device, and patent CN117035081A discloses a method and device for constructing a multi-source multi-modal knowledge graph, neither of which can well solve the above problems. Summary of the Invention

[0005] The embodiments of the present invention provide a method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph to solve the above technical problems.

[0006] In a first aspect, the embodiments of the present invention provide a method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph, including:

[0007] Obtain knowledge data of different modalities, where the different modalities include at least one of text, video, audio, and image;

[0008] Associate the knowledge data of different modalities according to the knowledge points to which they belong;

[0009] In a knowledge graph with knowledge points as nodes, display the mutually associated knowledge data of each modality as a multimedia resource link of the corresponding knowledge point.

[0010] Second aspect, an embodiment of the present invention provides an electronic device, which includes:

[0011] One or more processors;

[0012] A memory for storing one or more programs,

[0013] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph according to any embodiment.

[0014] Third aspect, an embodiment of the present invention further provides a system for generating multi-modal resource associations of subject knowledge points based on a knowledge graph, including:

[0015] A data acquisition module for acquiring and collecting multi-modal data;

[0016] A knowledge extraction module for extracting knowledge points and knowledge data of different modalities from the collected multi-modal data;

[0017] A knowledge graph construction module for constructing a knowledge graph with knowledge points as nodes according to the knowledge points and knowledge data of different modalities, and associating the knowledge data of different modalities according to the knowledge points;

[0018] A user interface module for displaying the knowledge graph in a graphical manner and displaying the mutually associated knowledge data of each modality as a multimedia resource link of the corresponding knowledge point.

[0019] In summary, an embodiment of the present invention provides a method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph, which can automatically associate multi-modal knowledge data belonging to the same knowledge point and display them in the knowledge graph, and provide multimedia resource links for each knowledge point. Description of the Drawings

[0020] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 is a collaboration flowchart between the modules of a system for generating multi-modal resource associations of subject knowledge points based on a knowledge graph provided by an embodiment of the present invention;

[0022] Figure 2It is a flowchart of a method for generating multimodal resource associations of subject knowledge points based on a knowledge graph provided by an embodiment of the present invention;

[0023] Figure 3 It is a schematic structural diagram of a knowledge association model provided by an embodiment of the present invention;

[0024] Figure 4 It is a schematic structural diagram of an encoder and decoder model in the training stage provided by an embodiment of the present invention;

[0025] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0027] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0028] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0029] A method for generating multimodal resource associations of subject knowledge points based on a knowledge graph provided by an embodiment of the present invention. To illustrate this method, the knowledge graph-based multimodal resource association generation system that supports the implementation of this method is described first. The system adopts a modular structure and includes a data collection module, a knowledge extraction module, a knowledge graph construction module, and a user interface module. Each module has independent functions, which is convenient for maintenance and upgrade. At the same time, the system uses a distributed database to store the knowledge graph and related resources to ensure high availability and scalability of the data.

[0030] Among them, the data collection module is mainly used to implement the following functions:

[0031] Multi-source data acquisition: Automatically collect various forms of educational resources such as textbooks, literature, videos, audios, and pictures through API interfaces, web crawler technology, and online platforms.

[0032] Automated data scraping and processing: Use automated scripts and scheduling tools to regularly scrape and update data, and convert it into a standard format for subsequent processing.

[0033] The knowledge extraction module is mainly used to implement the following functions:

[0034] Natural language processing: Use large language models to analyze text data, automatically extract key knowledge points, definitions, and concepts, and identify context relationships.

[0035] Video analysis: Use multimodal large models to process video content, automatically extract visual information and speech content related to knowledge points, and identify important scenes and key points of explanation.

[0036] Audio analysis: Also adopt large model technology for audio processing, extract key knowledge information in explanations and discussions, and ensure comprehensive integration of knowledge.

[0037] Image processing: Use large models for image classification and feature extraction, identify relevant knowledge points in pictures, and achieve effective integration of multimodal information.

[0038] The knowledge graph construction module is mainly used to implement the following functions:

[0039] Automated construction: Combine the results of knowledge extraction and use algorithms to automatically generate a knowledge graph, forming a structure of nodes and edges.

[0040] Dynamic update mechanism: Set periodic update and real-time update functions to ensure that new knowledge and relationships are promptly incorporated into the knowledge graph, automatically detect data changes, and trigger updates.

[0041] The user interface module is mainly used to implement the following functions:

[0042] Visual display: Design an intuitive user interface to display the knowledge graph in a graphical way, enabling users to conveniently browse and interact.

[0043] Personalized learning path recommendation: Based on the user's learning history and preferences, provide personalized recommended knowledge points and learning resources to enhance the learning effect.

[0044] Multimedia resource links: Provide relevant video, audio, and picture links for each knowledge point, allowing users to directly access and enhancing the learning experience.

[0045] At the same time, an evaluation and feedback mechanism is also established in the system, mainly used to achieve the following functions:

[0046] Automated feedback collection: Design a feedback mechanism that allows users to evaluate the accuracy and practicality of the knowledge graph, and the system automatically records and analyzes the feedback data.

[0047] Effect evaluation: Regularly analyze the system usage data, evaluate the effect of knowledge extraction and user satisfaction, and use automated analysis tools to continuously optimize the system performance.

[0048] In summary, the system adopts a modular design to ensure efficient collaboration between functional modules. Figure 1 The collaboration process between the system modules is shown. This system combines automated technology and advanced large model technology to achieve the visual expression of subject knowledge points, and at the same time can automatically identify associated knowledge points and resources from multiple resources (such as videos, audio, and pictures). This integration not only improves the construction efficiency of the knowledge graph but also solves the problems of manual annotation and lagging knowledge update in traditional methods. Through this system, users can obtain a dynamic and intelligent knowledge graph platform that provides personalized learning paths and rich multimedia resources, promoting a more efficient learning and teaching experience.

[0049] After understanding the above system, Figure 2 is a flowchart of a method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph provided by an embodiment of the present invention. This method can be executed in cooperation by each module in the above system or by other electronic devices independent of the above system. As Figure 2 shown, the method specifically includes:

[0050] S110. Obtain knowledge data of different modalities, where the different modalities include at least one of text, video, audio, and image.

[0051] The knowledge data here can be understood as each piece of knowledge data associated with a certain knowledge point.

[0052] Optionally, multiple pieces of knowledge data can be collected through API interfaces, web crawler technologies, and online platforms, and content segmentation can be performed if necessary. Among them, the knowledge points associated with the text data can be determined in advance through natural language processing technologies, while the knowledge points associated with the knowledge data in other modalities are unknown.

[0053] S120. Associate the knowledge data in different modalities according to the affiliated knowledge points.

[0054] In this step, a knowledge point association model based on deep learning is used to associate the knowledge data belonging to the same knowledge point. Figure 3 It is a schematic structural diagram of a knowledge association model provided by an embodiment of the present invention. As Figure 3 shown, the knowledge point association model includes encoders of different modalities and association calculation branches under different reference modalities. Among them, the encoders of each modality are respectively used to encode the knowledge data of each modality; the association calculation branches under each reference modality all include a transformation matrix, a bias matrix, and a similarity calculation layer. The transformation matrix and the bias matrix are jointly used to uniformly transform the encoded features of the knowledge data in different modalities into the encoding space of the reference modality, and the similarity calculation layer is used to calculate the similarity between the features after the transformation of the knowledge data in different modalities.

[0055] Specifically, in this embodiment, according to the amount of information of different modality data, a higher feature dimension is set for the modality with a larger amount of information to represent the effective information in the high-information modality with as rich a dimension as possible. Optionally, the encoding feature dimension of the video modality > the encoding feature dimension of the audio modality > the encoding feature dimension of the image modality > the encoding feature dimension of the text modality.

[0056] Based on the above model, S120 may include the following steps:

[0057] Step 1. Encode the knowledge data in different modalities respectively through the encoders of different modalities to obtain the encoded features of the knowledge data in each modality. Combining Figure 3 , multiple pieces of knowledge data in the four modalities of text, video, audio, and image can be respectively input into the corresponding encoders to obtain the encoded features corresponding to the four modalities , , and .

[0058] Step 2: Select the reference modality for the current learning topic according to the maximized representation form of the learning topic. Since the encoding feature dimensions corresponding to different modalities are different, in this embodiment, one modality is selected as the reference modality according to the learning topic. This modality corresponds to the best knowledge form of the current learning topic and can accurately present the knowledge content under the current learning topic. In this embodiment, this best knowledge form is also referred to as the maximized representation form, where "maximized" means maximizing the information accuracy.

[0059] Optionally, if the current learning topic is descriptive knowledge, the maximized representation form "text" of the descriptive knowledge can be used as the reference modality for the current learning topic. Exemplarily, if the current learning topic is literature and the maximized representation form of literature is text, then text is used as the reference modality for literature.

[0060] If the current learning topic is static visual knowledge, select the maximized representation form "image" of the static visual knowledge as the reference modality for the current learning topic. Exemplarily, if the current learning topic is art and the maximized representation form of art is image, then image is used as the reference modality for art.

[0061] If the current learning topic is dynamic visual knowledge, select the maximized representation form "video" of the dynamic visual knowledge as the reference modality for the current learning topic. Exemplarily, if the current learning topic is animation design and the maximized representation form of animation design is video, then video is used as the reference modality for animation design.

[0062] If the current learning topic is auditory knowledge, select the maximized representation form "audio" of the auditory knowledge as the reference modality for the current learning topic. Exemplarily, if the current learning topic is music and the maximized representation form of music is audio, then audio is used as the reference modality for music.

[0063] Step 3: Uniformly transform the encoding features of the knowledge data of each modality under the current learning topic into the encoding space of the reference modality. Combining Figure 3 , after selecting the reference modality, the encoding features of the knowledge data of any modality under the current learning topic can be transformed into the encoding space of the reference modality through the following formula:

[0064]

[0065] where n represents the index of each modality type, n = 1, 2, 3, 4…, represents the encoding feature of the knowledge data of the modality type with index n; S represents the index of the reference modality, and S is one of all the modality type indexes, represents the feature after transferring to the encoding space of the reference modality, and respectively represent the transformation matrix and the bias matrix from the encoding space of the modality type with index n to the encoding space of the reference modality with index S. Exemplarily, if the feature dimension of the encoding space with index n is a and the feature dimension of the encoding space of the reference modality is b, then is a column vector with dimension = a, is a matrix with dimension = b×a, and are both column vectors with dimension = b, and The elements in are all obtained through pre-training.

[0066] Specifically, taking Figure 3 as an example, and respectively represent the transformation matrix and the bias matrix from the encoding space of the video modality to the encoding space of the reference modality when the text modality is the reference modality; and respectively represent the transformation matrix and the bias matrix from the encoding space of the audio modality to the encoding space of the reference modality when the text modality is the reference modality; and respectively represent the transformation matrix and the bias matrix from the encoding space of the image modality to the encoding space of the reference modality when the text modality is the reference modality; , and respectively represent the features after the encoding features of the video modality, audio modality, and image modality are transformed into the encoding space of the reference modality when the text modality is the reference modality. Similarly, and respectively represent the transformation matrix and the bias matrix from the encoding space of the text modality to the encoding space of the reference modality when the video modality is the reference modality; and respectively represent the transformation matrix and the bias matrix from the encoding space of the audio modality to the encoding space of the reference modality when the video modality is the reference modality; and respectively represent the transformation matrix and the bias matrix from the encoding space of the image modality to the encoding space of the reference modality when the video modality is the reference modality; , and respectively represent the features after the encoding features of the text modality, audio modality, and image modality are transformed into the encoding space of the reference modality when the video modality is the reference modality. The meanings of the remaining variable symbols are similar and will not be elaborated one by one.

[0067] Step 4. Determine the multi-modal data belonging to the same knowledge point according to the encoded features after the conversion of the knowledge data of each modality. Specifically, each piece of knowledge data of each modality after conversion corresponds to a , calculate every two features of different modalities (that is and , ), and the features with a similarity higher than the set threshold correspond to the multi-modal data belonging to the same knowledge point.

[0068] In practical applications, the structures of the encoders of each modality can be flexibly selected according to needs. When necessary, the knowledge data input to the encoder can be preprocessed to meet the input requirements of the encoder. Exemplarily, the text encoder can be a Glove structure or other text vectorization models; the audio encoder can adopt a VGGish structure. The audio knowledge data can be first subjected to a frequency domain conversion to generate a log-Mel spectrogram, and then the log-Mel spectrogram is converted into an encoded vector; the video encoder can adopt a ResNet structure. For the video knowledge data, several video images of frames can be extracted first, each frame image is adjusted to the specified size of the model input, and the BGR format of each frame is converted to the RGB format through cv2.cvtColor, and then input into the video encoder; the image encoder can also adopt an image processing structure such as ResNet.

[0069] Further, in a specific embodiment, in order to enable the Figure 3 model shown to implement the function of the above-mentioned knowledge point association, the model can be trained in the following manner:

[0070] First, based on the Figure 4 network structure shown, corresponding decoders are respectively connected after the encoders of each modality, and the encoder and the decoder are trained as a whole. Specifically, for any encoder-decoder model of a modality, after the original knowledge data is input into the encoder, the encoder extracts the deep features (i.e., encoded features) of the original knowledge data, and the decoder restores the original knowledge data according to the deep features. The encoder and the decoder as a whole can be trained by minimizing the difference between the restored knowledge data and the original knowledge data. Optionally, the following loss function can be set:

[0071]

[0072] where i represents the sample index, and respectively represent the original knowledge data and the restored knowledge data in the sample. It should be noted that if the data input to the encoder is the preprocessed knowledge data, then and They also respectively correspond to the preprocessed data. By minimizing L1, it can be ensured that the encoder extracts the most important information dimensions in the knowledge data, and maximally reduces the information loss during the encoding process.

[0073] Then, remove each decoder, and connect the correlation calculation branches under each benchmark modality after each modality encoder respectively, to obtain the model structure as Figure 3 shown. At the same time, construct a sample set for this model. Use the multi-modal knowledge data belonging to the same knowledge point under various learning topics with each modality as the benchmark modality as positive samples, and use the multi-modal knowledge data belonging to different knowledge points as negative samples to train the knowledge point association model. During the training process, the parameters of each encoder remain unchanged, and only the parameters in the transformation matrix and bias matrix under each modality are updated. Optionally, the model training can be completed by constraining the maximization of the final similarity of the multi-modal knowledge data belonging to the same knowledge point and the minimization of the final similarity of the multi-modal knowledge data belonging to different knowledge points. For example, set the following loss function:

[0074]

[0075] where p and q represent the knowledge point indices, ; m and n represent the modality indices, ; represents the encoded feature after transformation of the knowledge data of modality n belonging to knowledge point p, and have similar meanings; represents the feature and similarity of, has a similar meaning; and respectively represent the weights of the loss terms. By minimizing L2, it can be ensured that after the encoded features of each modality are transformed into the vector space of the benchmark modality, the mutually related knowledge data can be recognized in this space.

[0076] S130. In the knowledge graph with knowledge points as nodes, display the mutually related multi-modal knowledge data as multimedia resource links belonging to the knowledge points.

[0077] As shown in the user interface module in the above system, in the knowledge graph, text data can be displayed as the name or attributes of the associated knowledge points, and other modality data associated with the text data can be displayed as video, audio, and picture links of the knowledge points. Users can directly access them to enhance the learning experience.

[0078] Further, if a piece of text data is associated with more than one knowledge point, one knowledge point can be selected as the main knowledge point, and the multi-modal data can be associated according to the main knowledge point; or each knowledge point can be encoded by an encoder of the text modality, and the encoded features can be transformed into the benchmark modality space, and then the projection of the encoded features of each modality knowledge data after transformation on the encoded features of the specific knowledge point after transformation can be calculated, and the similarity of each modality knowledge data can be calculated according to the projection vector. At this time, the similarity is the similarity of the multi-modal knowledge data with respect to the specific knowledge point. Correspondingly, in the knowledge graph, the multi-modal knowledge data with high similarity should be displayed as the relevant attributes or links of the specific knowledge point. In practical applications, the above processing methods can be flexibly selected according to the knowledge granularity and scale in the knowledge graph, and this embodiment does not make specific limitations.

[0079] In summary, this embodiment provides a method for generating multi-modal resource associations for subject knowledge points based on a knowledge graph, which can automatically associate multi-modal knowledge data belonging to the same knowledge point and display them in the knowledge graph, providing multimedia resource links for each knowledge point. Specifically, in this embodiment, different coding dimensions are set for each modality according to the amount of information included in different modalities, and through the training of a model with an encoder-decoder structure, it is ensured that each encoder can extract the key features in each modality and minimize information loss; at the same time, in this embodiment, according to the maximized representation form of the learning theme, the information modality that expresses the learning content most accurately is selected as the benchmark modality, and each modality data is respectively transformed into the vector space of the benchmark modality, and the similarity of each modality data is calculated through the transformed encoded vectors, so as to determine the interrelated multi-modal data. The benchmark modality reflects the dimensions emphasized by different learning themes. Transforming other modality data into the encoding space emphasized by the learning theme can avoid losing too much information in the dimensions emphasized by the learning theme and ensure the accuracy of knowledge association as much as possible.

[0080] Figure 5 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, as Figure 5 shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more, Figure 5 taking one processor 60 as an example; the processor 60, the memory 61, the input device 62, and the output device 63 in the device can be connected through a bus or other means, Figure 5 taking the connection through the bus as an example.

[0081] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph in the embodiments of the present invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, that is, to implement the above-mentioned method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph.

[0082] The memory 61 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 61 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 61 may further include a memory remotely provided with respect to the processor 60, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0083] The input device 62 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 63 may include display devices such as a display screen.

[0084] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph in any embodiment.

[0085] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0086] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0087] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0088] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating multi-modal resource associations of subject knowledge points based on a knowledge graph, characterized in that Comprising: Obtaining knowledge data of different modalities, where the different modalities include at least one of text, video, audio, and image; Associating knowledge data of different modalities according to the knowledge points to which they belong; In a knowledge graph with knowledge points as nodes, displaying the mutually associated knowledge data of each modality as multimedia resource links of the knowledge points to which they belong; Among them, the associating the knowledge data of different modalities according to the knowledge points to which they belong includes: Encoding the knowledge data of different modalities respectively through encoders of different modalities to obtain the encoded features of the knowledge data of each modality, where the modality with greater information volume corresponds to a larger feature dimension; Selecting a reference modality for the current learning topic according to the maximized representation form of the learning topic, where the maximized representation form corresponds to the best knowledge form of the current learning topic and can achieve the maximization of information accuracy; specifically, if the current learning topic is descriptive knowledge, selecting the text, which is the maximized representation form of descriptive knowledge, as the reference modality for the current learning topic; if the current learning topic is static visual knowledge, selecting the image, which is the maximized representation form of static visual knowledge, as the reference modality for the current learning topic; if the current learning topic is dynamic visual knowledge, selecting the video, which is the maximized representation form of dynamic visual knowledge, as the reference modality for the current learning topic; if the current learning topic is auditory knowledge, selecting the audio, which is the maximized representation form of auditory knowledge, as the reference modality for the current learning topic; Converting the encoded features of the knowledge data of each modality under the current learning topic into the encoded space of the reference modality through the following formula: I n,S = T n,S × I n + B n,S Among them, I n represents the knowledge data coding feature of the modality type with index n, T n,S and B n,S respectively represent the transformation matrix and the bias matrix from the coding space of the modality type with index n to the coding space of the reference modality, I n,S represents I n the feature after being transformed to the coding space of the reference modality; Calculating the similarity of the encoded features after conversion of every two knowledge data of different modalities respectively; determining the multi-modal knowledge data with a similarity greater than the set threshold as the multi-modal knowledge data belonging to the same knowledge point; Among them, before encoding the knowledge data of different modalities respectively through encoders of different modalities to obtain the encoded features of the knowledge data of each modality, it further includes: Connecting corresponding decoders respectively behind the encoders of each modality, where each encoder is used to encode the original knowledge data, and each decoder is used to restore each knowledge data according to each encoded feature; Training each encoder and decoder with the knowledge data of each modality, and updating the parameters by constraining the difference between the restored knowledge data and the original knowledge data to be minimized during the training; Removing each decoder, and connecting an association calculation branch under each modality respectively behind the encoders of each modality, where each association calculation branch includes a transformation matrix, a bias matrix, and a similarity calculation layer; Fixing the parameters of each encoder, and training each transformation matrix and bias matrix with the multi-modal knowledge data under various learning topics, and updating the parameters by constraining the similarity of the multi-modal knowledge data belonging to the same knowledge point to be maximized finally, and the similarity of the multi-modal knowledge data belonging to different knowledge points to be minimized finally during the training.

2. The method according to claim 1, wherein The size relationship of the encoded feature dimensions of different modalities is: video > audio > image > text.

3. An electronic device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating multimodal resource associations of subject knowledge points based on a knowledge graph according to any one of claims 1-2.

4. A multi-modal resource association generation system for subject knowledge points based on a knowledge graph, characterized in that, It includes: A data acquisition module for acquiring and collecting multimodal data. A knowledge extraction module for extracting knowledge points and knowledge data of different modalities from the collected multimodal data. A knowledge graph construction module for constructing a knowledge graph with knowledge points as nodes according to the knowledge points and knowledge data of different modalities, and associating the knowledge data of different modalities according to the knowledge points. A user interface module for displaying the knowledge graph in a graphical manner and displaying the mutually associated knowledge data of each modality as a multimedia resource link of the corresponding knowledge point. Among them, the knowledge graph construction module associates the knowledge data of different modalities according to the following method based on the corresponding knowledge points: Encoding the knowledge data of different modalities respectively through encoders of different modalities to obtain the encoded features of the knowledge data of each modality, where the modality with greater information volume corresponds to a larger feature dimension. Selecting the benchmark modality of the current learning topic according to the maximized representation form of the learning topic. Specifically, if the current learning topic is descriptive knowledge, select the text of the maximized representation form of descriptive knowledge as the benchmark modality of the current learning topic; if the current learning topic is static visual knowledge, select the image of the maximized representation form of static visual knowledge as the benchmark modality of the current learning topic; if the current learning topic is dynamic visual knowledge, select the video of the maximized representation form of dynamic visual knowledge as the benchmark modality of the current learning topic; if the current learning topic is auditory knowledge, select the audio of the maximized representation form of auditory knowledge as the benchmark modality of the current learning topic. Converting the encoded features of the knowledge data of each modality under the current learning topic into the encoded space of the benchmark modality through the following formula: I n,S = T n,S × I n + B n,S Among them, I n represents the knowledge data coding feature of the modality type with index n, T n,S and B n,S respectively represent the transformation matrix and the bias matrix from the coding space of the modality type with index n to the coding space of the reference modality, I n,S represents I n the feature after being transformed to the coding space of the reference modality; Calculating the similarity of the encoded features after conversion of every two items of knowledge data of different modalities respectively; determining the multimodal knowledge data with a similarity greater than a set threshold as the multimodal knowledge data belonging to the same knowledge point. Among them, before encoding the knowledge data of different modalities respectively through encoders of different modalities to obtain the encoded features of the knowledge data of each modality, it further includes: Connecting corresponding decoders respectively behind the encoders of each modality. Each encoder is used to encode the original knowledge data, and each decoder is used to restore each knowledge data according to each encoded feature. Training each encoder and decoder with the knowledge data of each modality, and updating the parameters by minimizing the difference between the restored knowledge data and the original knowledge data during the training. Removing each decoder, and connecting an association calculation branch under each modality respectively behind the encoders of each modality. Each association calculation branch includes a transformation matrix, a bias matrix, and a similarity calculation layer. Fix the parameters of each encoder, and use the multi-modal knowledge data under various learning topics to train each transformation matrix and bias matrix. During the training, update the parameters by constraining the maximization of the final similarity of the multi-modal knowledge data belonging to the same knowledge point and the minimization of the final similarity of the multi-modal knowledge data belonging to different knowledge points.

Citation Information

Patent Citations

  • Multi-modal data association method and device

    CN112015955A

  • Construction method and device of multi-element and multi-mode knowledge graph

    CN117035081A

  • Multi-modal knowledge graph construction method

    CN112200317A

  • Multi-modal data knowledge information extraction method based on deep-width joint neural network

    CN113361559A