Marine field-oriented multi-modal knowledge graph construction and completion method

By employing retrieval enhancement, visual element generation, and image technology, the problems of insufficient utilization of triple semantics and data resources in the construction of multimodal knowledge graphs have been solved, achieving complete visualization and high-quality representation of marine multimodal knowledge graphs.

CN120874996APending Publication Date: 2025-10-31SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511044250.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing methods fail to fully utilize the global semantics of triples when constructing multimodal knowledge graphs. The lack of high-quality multimodal data resources on the Internet results in incomplete knowledge graphs that cannot effectively represent abstract relationships and provide high-quality visualizations.

Method used

By employing retrieval-enhanced triplet correction, visual element generation based on large language models, image retrieval and filtering, and image generation and editing techniques, we construct and complete a multimodal marine knowledge graph to ensure the accuracy and completeness of knowledge.

Benefits of technology

It significantly improves the completeness and coverage of the marine multimodal knowledge graph, solves the problem of visualizing abstract relationships, overcomes the limitation of scarce data resources, and ensures the accuracy and high-quality visualization of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874996A_ABST
    Figure CN120874996A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal knowledge graph construction and completion method oriented to the ocean field. The method comprises the steps of construction of a multi-modal knowledge graph oriented to the ocean field and completion of the multi-modal knowledge graph oriented to the ocean field. Wherein the multi-modal knowledge graph construction oriented to the ocean field comprises triple correction based on retrieval enhancement generation, visual element generation based on a large language model and image retrieval and filtering; the multi-modal knowledge graph completion facing the ocean field comprises multi-modal knowledge graph completion based on image generation and multi-modal knowledge graph completion based on image editing. According to the technical scheme, the integrity and quality of the marine multi-modal knowledge graph can be improved, and a solid foundation is provided for deep mining and application of marine knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of marine data processing, specifically relating to a method for constructing and completing a multimodal knowledge graph for the marine field. Background Technology

[0002] With the introduction of the strategy of "developing the marine economy, protecting the marine ecological environment, and accelerating the construction of a maritime power," technologies based on the ocean are constantly innovating, resulting in massive, heterogeneous, and multi-source marine data. Traditional marine data processing methods struggle to effectively integrate and utilize this heterogeneous data, thus hindering the in-depth mining and intelligent application of marine knowledge.

[0003] Multimodal knowledge graphs, as a form of structured knowledge representation, effectively overcome the limitations of traditional single-modal representations by integrating information from multiple modalities such as text and vision, making the presentation of marine knowledge more comprehensive and intuitive. Therefore, how to apply multimodal knowledge graphs to the marine field has become a research hotspot in marine data processing. However, the construction and completion of multimodal knowledge graphs for the marine field still face many challenges:

[0004] First, existing methods fail to fully utilize the global semantics of triples when constructing multimodal knowledge graphs. Specifically, current methods only focus on associating entities in triples with corresponding multimodal data (such as images and videos), but do not effectively utilize the relationships within the triples. This results in multimodal knowledge graphs being unable to represent the deep semantics established through relationships. Second, there is a lack of high-quality multimodal data resources on the internet, leading to incomplete marine multimodal knowledge graphs. Knowledge graphs constructed based on multimodal data retrieval suffer from the limitation of scarce high-quality multimodal data resources, resulting in incomplete multimodal knowledge graphs that affect their completeness and universality, thus limiting their application in downstream tasks in the marine field. Furthermore, even when multimodal data capable of expressing triple semantics is retrieved, the data itself suffers from low resolution, semantic non-focus, or distortion of local details, making it unsuitable as a high-quality visual representation of triples. Summary of the Invention

[0005] The purpose of this invention is to provide a method for constructing and completing multimodal knowledge graphs in the marine field, so as to solve the problems existing in the prior art.

[0006] To achieve the above objectives, this invention provides a method for constructing and completing a multimodal knowledge graph for the marine field, comprising:

[0007] Obtain the processed marine knowledge graph;

[0008] Construction of multimodal knowledge graphs for the marine domain;

[0009] Multimodal knowledge graph completion for the marine domain.

[0010] The specific method for constructing the multimodal knowledge graph for the marine field is as follows:

[0011] Retrieval-enhanced triplet correction;

[0012] Visual element generation based on a large language model;

[0013] Image retrieval and filtering.

[0014] The specific method for completing the multimodal knowledge graph for the marine field is as follows:

[0015] Image-based multimodal knowledge graph completion;

[0016] Multimodal knowledge graph completion based on image editing.

[0017] Furthermore, the correction based on retrieval enhancement-generated triples specifically includes:

[0018] Step 1: Use the head and tail entities of the original triples in the ocean knowledge graph as search keywords in the search engine to search for the text descriptions of the head and tail entities;

[0019] Step 2: Set fixed template questions for the triplet correction task;

[0020] Step 3: Combine the triples with the corresponding text descriptions of the head and tail entities into a fixed template sample and input it into GPT-3.5 to obtain the final result.

[0021] Repeat steps one through three until all triples have been checked and corrected to obtain a corrected ocean knowledge graph, ensuring the reliability of the knowledge.

[0022] Furthermore, visual element generation based on a large language model specifically includes:

[0023] Step 1: Manually define 10 visual elements corresponding to the relationships;

[0024] Step 2: Set a fixed question template for the visual element generation task;

[0025] Step 3: Utilize the contextual learning capabilities of the large language model to learn from the defined relation-visual element pairs, extend them to new relations, and quickly obtain visual elements of unseen relations.

[0026] Furthermore, image retrieval and filtering specifically include:

[0027] Step 1: Develop a web crawler to retrieve images related to the triplet in a search engine using (head entity, visual element, tail entity) as the query string;

[0028] Step 2: Manually evaluate the retrieved images and exclude those that cannot accurately express the semantics of the triples.

[0029] Furthermore, image-based multimodal knowledge graph completion specifically includes:

[0030] Step 1: Generate detailed image text descriptions using triples and their associated visual elements in the ocean knowledge graph;

[0031] Step 2: Input the image text description into Stable Diffusion to generate an image that can express the semantics of triples.

[0032] Furthermore, multimodal knowledge graph completion based on image editing specifically includes:

[0033] Step 1: Input the original image corresponding to the triple, go through the vision-language joint fine-tuning stage, and output the optimized source text embedding and the fine-tuned UNet;

[0034] Step 2: In the text-guided image editing stage, visual elements that express the semantics of triples are edited onto the original image.

[0035] The technical effects of this invention are as follows:

[0036] This invention ensures the reliability of fundamental knowledge by semantically verifying and correcting triples in a marine knowledge graph. By generating visual elements of relationships using a large language model, it effectively solves the visualization challenge of abstract marine concepts and invisible relationships, significantly increasing the types of relationships and the number of triples that the knowledge graph can represent, thereby greatly improving the completeness and coverage of the marine multimodal knowledge graph. Utilizing image generation technology, this invention overcomes the limitation of scarce high-quality data resources in the marine field, creating new visual representations for triples lacking multimodal content. Simultaneously, through image editing technology, it solves problems such as low data resolution, semantic non-focus, or distortion of local details, ensuring the accuracy and high quality of the final visualization results. In summary, this invention achieves the construction and completion of a multimodal knowledge graph through triple correction, visual element generation, image retrieval, image generation, and image editing steps. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0039] Figure 1 This is a flowchart illustrating the implementation of an embodiment of the present invention. Detailed Implementation

[0040] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0041] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Every smaller range between any stated value or intermediate value within a stated range, and any other stated value or intermediate value within said range, is also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.

[0042] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. While only preferred methods have been described herein, any methods similar or equivalent to those described herein may be used in the implementation or testing of this invention. All references to this specification are incorporated by way of citation to disclose and describe the methods associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.

[0043] Various modifications and variations can be made to the specific embodiments described in this specification without departing from the scope or spirit of the invention, as will be apparent to those skilled in the art. Other embodiments derived from this specification will also be obvious to those skilled in the art. This application specification and embodiments are merely exemplary.

[0044] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.

[0045] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0046] Example:

[0047] like Figure 1 As shown, this embodiment provides a method for constructing and completing a multimodal knowledge graph for the marine field, specifically including the following steps:

[0048] S1: Construction of a multimodal knowledge graph for the marine domain;

[0049] S2: Multimodal knowledge graph completion for the marine domain.

[0050] Step S1 aims to process the single-modal ocean knowledge graph and associate it with high-quality image modalities. The proposed method specifically includes:

[0051] S101: Triple correction based on retrieval enhancement generation;

[0052] S102: Visual element generation based on a large language model;

[0053] S103: Image retrieval and filtering.

[0054] Traditional knowledge graph construction processes may introduce factual errors, and when associating images, it is difficult to directly find corresponding visual representations for abstract concepts and invisible relationships prevalent in the marine domain. This results in inherent deficiencies in the accuracy and modal completeness of the constructed knowledge graph. Therefore, this invention ensures the accuracy of marine knowledge through triple correction; broadens the visual representation range of triples through the visual elements of relationships; and matches high-quality images to the knowledge graph through image retrieval and filtering.

[0055] The triplet correction based on retrieval enhancement includes the following:

[0056] This embodiment utilizes a large language model combined with external knowledge retrieved in real time from a search engine to automatically perform semantic verification and correction on triples. Specifically, for a triple (h, r, t) to be corrected in the ocean knowledge graph, its head entity h and tail entity t are extracted as search keywords. Textual descriptions d about these two entities are then retrieved through a search engine. h and d t .

[0057] After obtaining the text description information, a few-shot cue is constructed for the triplet correction task. This cue is derived from the retrieved entity text description d. h d tThis process, along with a pre-defined templated question q (e.g., "Based on the provided knowledge, determine whether the triple... contains a semantic error; if so, provide a correct modification"), constitutes the entire process. The constructed complete prompt is then input into a large language model (in this embodiment, GPT-3.5) for reasoning. The model outputs the judgment result and correction suggestions. This process is repeated until all triples in the knowledge graph have been verified and corrected.

[0058] The generation of visual elements based on a large language model includes the following:

[0059] To address the challenge of directly visually associating abstract relationships in the marine domain, this embodiment generates a set of concrete visual elements for the relationships, transforming abstract semantics into searchable, generateable, and editable visual features.

[0060] To achieve this, this embodiment first manually defines a set of visual elements corresponding to a small subset of representative marine domain relationships to construct contextual examples. Based on this, the contextual learning capabilities of a large language model are utilized to construct a complete inference prompt that includes a task guide and the aforementioned examples. For a new relationship not appearing in the examples, a new set of visual elements corresponding to that new relationship can be generalized and generated by posing a specific question to the model.

[0061] The image retrieval and filtering includes the following:

[0062] This embodiment utilizes the visual elements generated in the previous step to guide the web crawler in precise retrieval. Specifically, for a triple and a set of visual elements relating it to other triples, multiple query strings combining triples and visual elements are constructed for retrieval. After the retrieval is complete, to ensure image quality, marine experts will manually evaluate the retrieved images to determine whether they clearly and accurately represent the semantics of the corresponding triples. Statistical metrics such as Fleiss's Kappa are used to measure consistency among evaluators, ensuring the objectivity of the evaluation results and ultimately selecting high-quality images.

[0063] At this point, after stage S1, some triples in the ocean knowledge graph have been successfully associated with high-quality images.

[0064] Step S2 aims to create or improve the visual representation of triples for which suitable images could not be found in step S1 through retrieval, thereby achieving complete visualization of the knowledge graph. The proposed methods include:

[0065] S201: Image-based multimodal knowledge graph completion;

[0066] S202: Multimodal knowledge graph completion based on image editing.

[0067] Due to the limited image resources on the internet, especially in highly specialized fields like marine science, many triples cannot be matched with images through retrieval. Furthermore, even if relevant images are retrieved, issues such as image distortion and missing key semantics may exist. Therefore, this invention utilizes image generation and editing technologies to create or finely modify images from scratch, supplementing knowledge graphs with high-quality visual modalities.

[0068] The image-based multimodal knowledge graph completion includes the following:

[0069] For triples that have no corresponding images on the internet, a text-to-image generative model is used to create entirely new visual content for them. Specifically, for a triple without an image, its relationships and visual elements are combined to generate a detailed and vivid image description text using a large language model. t Subsequently, the text description I t As input, it is fed into Stable Diffusion to generate a high-fidelity image.

[0070] Stable Diffusion consists of a forward diffusion process and a reverse denoising process. The forward diffusion process starts with a given image x0 and gradually adds Gaussian noise at each time step t∈{0,…,T}. To obtain x t The mathematical definition of this process is as follows:

[0071]

[0072] Where α t These are diffusion scheduling parameters that satisfy 0 = α T <α T-1 …<α1<α0=1.

[0073] In the inverse denoising process, the model reconstructs the image from the noise. Given x t And text embedding e, time bars UNets∈ in the diffusion model θ (x t ,t,e) Predict noise ∈ t Using the denoising diffusion implicit model, the inverse process can be defined as:

[0074]

[0075] Using a latent diffusion model, the original image x0 is transformed into a latent representation z0 via a variational autoencoder ∈(x0). The overall training loss is defined as follows:

[0076]

[0077] Finally, the denoised z0 is restored to a pixel-level image by the variational autodecoder.

[0078] The image-edit-based multimodal knowledge graph completion includes the following:

[0079] For existing original images that lack complete or flawed semantic expression, text-guided image editing techniques are used for refined modification. This embodiment employs a two-stage approach.

[0080] In the vision-language joint fine-tuning stage, firstly, the original image corresponding to the triple is input, and BLIP is used to generate a text title describing the original image, called the source cue. Then, this source cue is input into the Stable Diffusion CLIP text encoder to generate the embedding of the source cue. src Subsequently, the original image is encoded using a VAE to obtain the latent representation z. t Noise is added to the text and then fed into UNet. Finally, the source text embedding is jointly optimized. src The optimization process focuses only on the encoding layers 0, 1, 2 and decoding layers 1, 2, 3 of the UNet, while freezing the parameters of its deeper layers. During optimization, different learning rates are used for the source text embedding and UNet due to the significant difference between text and images. The loss function for the optimization process is defined as:

[0081]

[0082] This stage aims to help the model "remember" the core content of the original image. Through this optimization process, the output includes the optimized source text embedding. src And the finely tuned UNet.

[0083] After fine-tuning, we move on to the editing stage. This stage requires a target text prompt, namely the adjusted triplet image I. t The text description. This target text cue is generated using triples and visual elements and output to the StableDiffusion text encoder to generate the target text embedding. tgt This embodiment uses vector subtraction to merge the source text embeddings. src and target text embedding e tgt This yields the final bootstrap embedding e.

[0084] e = γe tgt +(1-γ)e src =e src +γ(e tgt -esrc )

[0085] The interpolation coefficient γ ranges from 0 to 1 and is used to balance the target text embedding e. tgt and source text embedding e src The proportion in the final embedding e.

[0086] Finally, DDIM sampling and classifier-free guidance are used for editing. Guided by the fine-tuned UNet and the final guided embedding e, the latent representation of the original image is edited to generate an edited image that retains the original features while adding new semantics, thus mitigating the distortion problem caused by image generation.

[0087] Thus far, this embodiment has completed the construction and completion of the entire multimodal knowledge graph for the ocean domain after the S2 stage, which created high-quality visual representations for all triples for which no suitable image could be found in S1.

[0088] Compared with the prior art, the advantages and positive effects of this embodiment are as follows:

[0089] First, it ensures the accuracy and reliability of the underlying data of the knowledge graph. At the outset, this invention employs a retrieval-enhanced triplet correction step (S101) to verify and correct facts in the basic marine knowledge graph. This step avoids subsequent multimodal associations based on erroneous or inaccurate knowledge, fundamentally improving the quality of knowledge and laying a solid data foundation for constructing a high-quality, highly reliable marine multimodal knowledge graph. Second, it overcomes the technical bottleneck of expressing abstract and invisible relationships in a multimodal manner, significantly improving the completeness and coverage of the knowledge graph. Existing technologies often ignore abstract or invisible relationships, leading to incomplete knowledge graphs. This invention uses a step (S102) to generate visual elements using a large language model, transforming abstract semantics into searchable and generateable concrete visual features. Third, it overcomes the dependence of existing technologies on high-quality online multimodal data resources, solving the problem of data scarcity. Given the scarcity of images in specific marine domains on the internet, existing retrieval methods struggle to construct complete knowledge graphs. This invention, through a multimodal knowledge graph completion step (S201) based on image generation, can create entirely new, high-fidelity visual content "from scratch" for triples for which the corresponding image cannot be found through retrieval.

[0090] refer to Figure 1 This embodiment designs a method for constructing and completing a multimodal knowledge graph for the marine domain, which can be divided into the following two steps: constructing a multimodal knowledge graph for the marine domain (S1) and completing a multimodal knowledge graph for the marine domain (S2).

[0091] S1 comprises three core steps: triplet correction based on retrieval enhancement (S101), visual element generation based on a large language model (S102), and image retrieval and filtering (S103). S2 comprises two key steps: multimodal knowledge graph completion based on image generation (S201) and multimodal knowledge graph completion based on image editing (S202). The output of S101 is the input of S102, and the output of S102 is the input of S103; the output of S201 is the input of S202.

[0092] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for constructing and completing a multimodal knowledge graph for the marine domain, characterized in that, include: Obtain the processed marine knowledge graph; Construction of multimodal knowledge graphs for the marine domain; Multimodal knowledge graph completion for the marine domain.

2. The method for constructing and completing a multimodal knowledge graph for the marine field according to claim 1, characterized in that, The construction of the multimodal knowledge graph for the marine domain specifically includes: Retrieval-enhanced triplet correction; Visual element generation based on a large language model; Image retrieval and filtering.

3. The method for constructing and completing a multimodal knowledge graph for the marine field according to claim 1, characterized in that, The multimodal knowledge graph completion for the marine domain specifically includes: Image-based multimodal knowledge graph completion; Multimodal knowledge graph completion based on image editing.

4. The method for constructing and completing a multimodal knowledge graph for the marine field according to claim 2, characterized in that, The triplet correction based on retrieval enhancement specifically includes the following steps: 1) Use the head and tail entities of the original triples in the ocean knowledge graph as search keywords in the search engine to search for the text descriptions of the head and tail entities; 2) The issue of setting fixed templates for triplet correction tasks; 3) Combine the triples with the corresponding text descriptions of the head and tail entities into a fixed template sample and input it into GPT-3.5 to obtain the final result; Repeat steps 1) through 3) until the verification and correction of each triple are completed, to obtain the corrected marine knowledge graph and ensure the reliability of the knowledge.

5. The method for constructing and completing a multimodal knowledge graph for the marine field according to claim 2, characterized in that, The generation of visual elements based on a large language model specifically includes the following steps: 1) Manually define 10 visual elements corresponding to each relationship; 2) Set fixed question templates for visual element generation tasks; 3) Leveraging the contextual learning capabilities of large language models, learn from defined relation-visual element pairs and extend them to new relations to quickly acquire visual elements of unseen relations.

6. The method for constructing and completing a multimodal knowledge graph for the marine field according to claim 2, characterized in that, The image retrieval and filtering specifically includes the following steps: 1) Develop a web crawler to retrieve images related to the triplet in a search engine using (head entity, visual element, tail entity) as the query string; 2) The retrieved images are manually evaluated to exclude those that cannot accurately express the semantics of the triples.

7. The method for constructing and completing a multimodal knowledge graph for the marine field according to claim 3, characterized in that, The image-based multimodal knowledge graph completion specifically includes the following steps: 1) Generate detailed image and text descriptions using triples and their associated visual elements in the ocean knowledge graph; 2) Input the image text description into Stable Diffusion to generate an image that can express the semantics of triples.

8. The method for constructing and completing a multimodal knowledge graph for the marine field according to claim 3, characterized in that, The image-edit-based multimodal knowledge graph completion specifically includes the following steps: 1) Input the original image corresponding to the triple, go through the vision-language joint fine-tuning stage, and output the optimized source text embedding and the fine-tuned UNet; 2) Through the text-guided image editing stage, visual elements that express the semantics of triples are edited on the original image.