A method, device and medium for constructing a scientific information graphic and text knowledge base
By building a scientific information picture and text knowledge base and using deep learning neural network training models and the scientific information picture and text knowledge base ontology, the problem of insufficient contextual knowledge mining in the field of science and technology by large multimodal models has been solved, and in-depth analysis of scientific and technological pictures and texts and improved data utilization efficiency have been achieved.
Patent Information
- Application Number
- CN202411855157.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-12-16
AI Technical Summary
When processing scientific and technological images, existing large multimodal models lack the ability to deeply mine the contextual knowledge contained in them, making it difficult to meet the needs of the scientific and technological intelligence field.
Build a scientific information image and text knowledge base, mine the correlation between images and text from scientific papers, patents and news, use the preset deep learning neural network to train the scientific research image-text matching model, and combine it with the scientific information image and text knowledge base ontology to achieve in-depth analysis and matching of scientific research images and texts.
It improves the application value and data utilization efficiency of scientific and technological images, supports the fine-tuning of multimodal large models in the field of science and technology, and enhances the depth and breadth of scientific and technological intelligence analysis.
Smart Images

Figure CN119782505B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of knowledge base construction, and in particular to a method, device and medium for constructing a scientific information graphic knowledge base. Background Art
[0002] Large-scale model technology utilizes vast amounts of data and computing power to train deep learning models with a large number of parameters. Compared to other AI technologies, these models possess superior generalization and learning capabilities. Currently, major multimodal large-scale models include the Flamingo large-scale model released by Google DeepMind, OpenAI's Chat GPT4, and the Wudao 3.0 Emu large-scale model released by China Zhiyuan. Datasets used for multimodal large-scale models include: 1) Flickr30k and its Chinese translation, Flickr30k-CNA, which primarily consist of images and simple one-sentence descriptions of the scenes within them; 2) China Zhiyuan's WUKONG dataset, which consists of 100 million image pairs and primarily images and text descriptions; and 3) AIC-ICC (a Chinese image description dataset released at the inaugural AI Challenger global AI competition, jointly sponsored by Sinovation Ventures, Sogou, and Toutiao). These datasets consist of images and multiple text descriptions detailing the scenes within them. These datasets focus on task-oriented or everyday images, and their descriptions are limited to the image itself, without any contextual information. The large models trained using the above datasets have good application effects in image retrieval, target detection, and image recognition, but they are not deep enough in mining scientific and technological images containing contextual knowledge, and are not suitable for training multimodal large models for scientific and technological intelligence. Summary of the Invention
[0003] The purpose of this application is to provide a method, device and medium for constructing a scientific information graphic and text knowledge base, thereby constructing a scientific information graphic and text knowledge base containing graphic and text information, thereby improving the value and utilization efficiency of data.
[0004] To achieve the above objectives, this application provides the following solutions:
[0005] In a first aspect, the present application provides a method for constructing a scientific information graphic and text knowledge base, comprising:
[0006] Acquire multiple scientific research documents containing images and perform preprocessing to obtain scientific research text and scientific research images corresponding to each of the scientific research documents;
[0007] For each scientific research document, the corresponding scientific research text and the scientific research image are input into a preset scientific research image-text matching model to obtain a scientific research image-text pair; the preset scientific research image-text matching model is obtained by training a preset deep learning neural network using a training sample set;
[0008] Obtaining a preset scientific information graphic and text knowledge base ontology; the preset scientific information graphic and text knowledge base ontology includes a plurality of triples, each triple including two scientific research entities, a relationship between the two scientific research entities, and attributes of each scientific research entity;
[0009] Based on the scientific research image-text pair, the scientific research entity and the corresponding attributes, and the relationship between the two scientific research entities are extracted and mapped to the preset scientific information image-text knowledge base ontology to obtain the scientific information image-text knowledge base.
[0010] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a method for constructing a scientific information graphic and text knowledge base.
[0011] In a third aspect, the present application provides a computer-readable storage medium on which a computer program is stored, which, when executed by a processor, implements a method for constructing a scientific information graphic and text knowledge base.
[0012] According to the specific embodiments provided by the present application, the present application has the following technical effects: the present application provides a method, device and medium for constructing a scientific information picture and text knowledge base, which mines scientific research texts and scientific research pictures from scientific research documents, and then obtains a preset scientific research picture-text matching model based on a preset deep learning neural network training, which can capture the correlation and correspondence between pictures and texts. In the training process of the model, not only the language features of the scientific research pictures are utilized, but also the language features of the scientific research texts are utilized, so that the mining of scientific research pictures is more extensive. Finally, knowledge extraction is performed based on the obtained scientific research picture-text pairs and mapped to the preset scientific information picture and text knowledge base ontology to obtain a scientific information picture and text knowledge base. The scientific information picture and text knowledge base includes picture and text information, which realizes the deep mining of the information contained in the picture and improves the value of the data; subsequent picture and text retrieval based on the scientific information picture and text database can improve the efficiency of data utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0014] Figure 1 A flowchart of a method for constructing a scientific information graphic and text knowledge base provided in one embodiment of the present application;
[0015] Figure 2A schematic diagram of a preset scientific research image-text matching model provided in one embodiment of the present application. DETAILED DESCRIPTION
[0016] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0017] Building a high-quality multimodal image and text dataset is a key factor in promoting the development of multimodal large model technology, and the scientific information image and text knowledge base in the field of scientific and technological intelligence is the key to fine-tuning the general multimodal large model technology into a multimodal large model in the field of scientific and technological intelligence. Based on this, this application explores the use of artificial intelligence technology to mine the correlation between images and texts from scientific papers, patent articles and scientific news, and establish a scientific information image and text knowledge base for multimodal large models. This application scheme not only provides an image and text dataset for the improvement of the general multimodal large model, but also provides support for the establishment of a multimodal scientific and technological intelligence large model.
[0018] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] In an exemplary embodiment, Figure 1 As shown, a method for constructing a scientific information graphic knowledge base is provided. The method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or can be executed by a terminal and a server together. In an embodiment of the present application, the method includes the following steps 101 to 104.
[0020] Step 101, obtain multiple scientific research documents containing pictures and pre-process them to obtain scientific research text and scientific research pictures corresponding to each of the scientific research documents; wherein, the scientific research documents containing pictures include scientific research news documents, scientific research paper documents and scientific research patent documents. The pictures in scientific research news may be the appearance of equipment and products in new product releases, high-tech equipment and process flows of technological breakthroughs, actual scenes of scientific and technological applications, leaders in the science and technology industry or executives of science and technology companies, etc. From the pictures in scientific and technological articles, the structure and relationship of the research objects and the logic of the model algorithm can be mined. The pictures in scientific research papers and patents complement the content of the articles, concisely and clearly display the research ideas and research results, and can also enhance the breadth and depth of academic exchanges, such as architecture diagrams, flow charts, research roadmaps, experimental results diagrams, system operation diagrams, etc.
[0021] In a practical application, crawler technology can be used to collect and adopt multiple scientific research documents. That is, crawler technology is used to automatically collect the content of all accessible pages according to certain rules to obtain or update the content of these websites, and then the obtained data is saved locally.
[0022] In another practical application, the step of preprocessing the scientific research file containing the image includes:
[0023] (11) For the scientific research news files, web page content recognition technology is used to obtain web page text and web page images. Specifically, web page content recognition technology can identify the structure, content, style and other information of the web page, and can also extract the required information therefrom to obtain web page text and web page images, wherein the web page text includes text content and content structure, the text content includes text information such as title, content, and image description, and the content structure is a basic format, such as the order of title, abstract, and headline, and also includes the position information of the image.
[0024] (12) The scientific research paper and patent documents are generally in PDF format and are parsed using OCR (Optical Character Recognition, which converts image information into text) technology to obtain document text and document images; the web page text and document text both include text content and content structure. In addition, a third-party PDF file parsing library can be used for parsing, and information such as document format, images, and their locations can be retained.
[0025] (13) Marking the web page text and the document text as scientific research text, and marking the web page image and the document image as scientific research images.
[0026] Step 102: For each scientific research document, the corresponding scientific research text and the scientific research picture are input into a preset scientific research picture-text matching model to obtain a scientific research picture-text pair; the preset scientific research picture-text matching model is obtained by training a preset deep learning neural network using a training sample set.
[0027] In an application example, the preset scientific research image-text matching model includes a visual encoder, a text encoder, an external knowledge combination module, a cross-modal matching module, and a loss function module, such as Figure 2 As shown. The input end of the visual encoder is used to receive the scientific research image, and the output end of the visual encoder is connected to the first input end of the cross-modal matching module; the input end of the text encoder is used to receive the scientific research text, and the output end of the text encoder is connected to the second input end of the cross-modal matching module; and the output end of the cross-modal matching module is connected to the loss function module.
[0028] The external knowledge combination module is arranged between the visual encoder and the text encoder, and the external knowledge combination module is used to: measure the distance between the scientific research picture and the corresponding text paragraph position in the scientific research text, and send it to the visual encoder and the text encoder respectively as a picture-text pair distance feature.
[0029] Specifically, the step of measuring the distance between the scientific research image and the corresponding text paragraph position in the scientific research text in the external knowledge combination module includes: if there is a caption below or above the scientific research image, the distance between the scientific research image and the corresponding text paragraph position in the scientific research text is 0; if there is no caption below or above the scientific research image, the distance is calculated using the following formula:
[0030]
[0031] The visual encoder is used to encode the visual semantic information contained in scientific research images to obtain a visual feature representation. Specifically, a convolutional neural network is used to model the scientific research images as a whole to obtain a global visual representation of the image. Because global encoding loses fine-grained semantic information, a self-attention mechanism or graph convolutional network is used to strengthen the relationship between entities within the visual modality to further enhance visual semantic understanding. In addition, the image-text distance feature is combined to provide more accurate matching information for subsequent cross-modal matching.
[0032] The text encoder is used to capture the semantic information contained in scientific research text to obtain a text modality feature representation. Specifically, a multi-head attention mechanism is introduced into the text encoder using the Transformer model to model the text sentences and obtain corresponding representations. In addition, the image-text distance feature is combined to provide more accurate matching information for subsequent cross-modal matching.
[0033] The cross-modal matching module is used to bridge the semantic gap between modalities, better aligning the semantic information of the two modalities and thus more accurately estimating similarity. Specifically, it measures cross-modal semantic similarity by jointly embedding visual and textual features into a common space.
[0034] The loss function module is used to constrain the learning of the image-matching model using different objective functions. Specifically, the same ranking function is used for learning, such as the hinge ranking loss (also known as the triple ranking loss) and the bidirectional hinge ranking loss. In this module, the introduced loss function is used to improve learning performance. During the training process, the training ends when the loss function value reaches the preset loss condition.
[0035] The image-text matching implemented in step 102 measures the semantic similarity between the image and the sentence, determines the degree of association between the image and the text, and ultimately outputs the scientific image and the text segment that matches the scientific image. In scientific papers and patents, image annotation and the location of text related to the image are more rigorous, allowing scientific image-text matching tasks based on this to achieve higher matching accuracy.
[0036] Step 103, obtaining a preset scientific information graphic and text knowledge base ontology; the preset scientific information graphic and text knowledge base ontology includes multiple triples, each triple includes two scientific research entities, the relationship between the two scientific research entities and the attributes of each scientific research entity.
[0037] The goal of an ontology is to capture knowledge in a relevant domain, provide a shared understanding of that knowledge, identify commonly recognized vocabularies within the domain, and provide clear definitions of these vocabularies (terms) and their interrelationships at various levels of formalization to overcome terminological differences and achieve semantic interoperability. Ontologies represent domain knowledge by capturing relevant concepts, instances, their attributes, semantic relationships, and associated semantic constraints. They serve as the backbone of a scientific and technological image and text knowledge base, describing image and text datasets in the field and establishing corresponding knowledge specifications. Constructing an ontology can achieve a certain degree of knowledge sharing and reuse, as well as improve system communication, interoperability, and reliability.
[0038] In one application example, this application provides a preset scientific information and graphic knowledge base ontology. Specifically, interviews are conducted with experts in the field of science and technology to investigate the main scientific research concepts and processes in different scientific research fields, and to collect opinions and suggestions for the design of the meta-ontology in the field of scientific information and graphic. Relevant technical personnel then establish a meta-model that describes the basic entities, attributes, and relationships of scientific information and graphic data. On this basis, the entities, attributes, and relationships can be automatically expanded through deep learning-based semantic analysis technology to complete the automatic expansion of the ontology.
[0039] In another application example, the scientific research entity includes: author, publishing institution, affiliated school department, paper ID and picture ID.
[0040] The attributes of the scientific research entity include: author attributes, institution attributes, department attributes, paper attributes and image attributes; the author attributes include name, date of birth, age, gender, professional title, research field and academic qualifications; the institution attributes include institution ID, institution name and institution introduction; the department attributes include college ID, college name, college introduction and college research field; the paper attributes include: paper title, paper abstract, paper keywords, paper field, paper publication journal and paper publication time; the image attributes include image name, image type, image keywords, paper to which the image belongs and image field.
[0041] The relationship between the two scientific research entities includes: is-a relationship, subordinate relationship, causal relationship, time relationship, homonymous relationship and hierarchical relationship.
[0042] In this application, the image types include flow charts, architecture diagrams, design diagrams, experimental result diagrams, system interface diagrams, character diagrams and product schematics. In addition, scientific research images have strong domain attributes, category attributes and association relationships. The scientific research field has a relatively complete system architecture, but there are many interdisciplinary studies in existing scientific and technological articles, so the domain attributes of scientific research images belong to a one-to-one or one-to-many relationship. There are also different associations between different types of scientific research images. For example, a design diagram and an experimental result diagram may be a description of different stages of the same scientific research topic; the same design diagram may be a description of different aspects of the design of the same scientific research topic. Therefore, there are co-local relationships, hierarchical relationships, and affiliation relationships between different scientific research images.
[0043] Each type of scientific research image holds distinct value. For example, experimental results images can reveal performance indicators for a particular research area, person images can identify experts in that area, and product diagrams can provide competitive intelligence. The types of scientific research images are manually defined. Later, through analysis and mining of large numbers of images, deep learning techniques or machine learning cluster analysis are used to learn new entity types, attributes, and relationship types from the data. This new information is then used to expand and update the ontology.
[0044] Step 104: Based on the scientific image-text pairs, the scientific research entities, their corresponding attributes, and the relationship between the two scientific research entities are extracted and mapped into the pre-set scientific information image-text knowledge base ontology to obtain a scientific information image-text knowledge base. This scientific information image-text knowledge base supports dual search functions for scientific research images and scientific research texts to support scientific and technological intelligence analysis tasks. It also supports batch export of scientific research image-text pairs to generate cross-modal datasets, enabling the pre-training of large cross-modal scientific and technological intelligence models.
[0045] In an application example, for images in the same scientific research document, after completing the matching of scientific research image-text pairs, knowledge extraction is required, mainly including the extraction of the field to which the image belongs, the image type, the relationship between multiple images, etc. In the specific extraction process, if the scientific research entity is an image ID, that is, the extraction of scientific research entities, relationships or attributes related to the image, then knowledge extraction is performed based on the content of the text paragraph corresponding to the scientific research image (that is, the text content in the scientific research image-text pair); in addition, relevant content assistance can be retrieved from the entire scientific research text as needed, such as the title, abstract, etc.; if it involves the extraction of other scientific research entities, relationships or attributes related to the author or publishing institution, then knowledge extraction is performed based on the entire scientific research text.
[0046] In an application example, the step of extracting the attributes corresponding to the scientific research entity includes:
[0047] When the attribute corresponding to the scientific research entity is the field to which the image belongs, the extraction process includes: utilizing deep learning classification technology to classify the title, abstract, or research content of the scientific research document to which the scientific research image belongs, thereby determining the field to which the image belongs. After completing this classification, preliminary proofreading and verification of the field to which the image belongs may also be performed according to the classification code of the publicly available scientific research document, as needed.
[0048] When the attribute corresponding to the scientific research entity is a picture type, the extraction process includes: classifying the pictures based on the scientific research pictures and the corresponding texts in the scientific research picture-text pair to obtain a first picture type and a second picture type; when the first picture type is consistent with the second picture type, the first picture type is used as the final picture type and stored; when the first picture type is inconsistent with the second picture type, the scientific research pictures are stored in a waiting area; when the number of scientific research pictures in the waiting area reaches a preset number, all the scientific research pictures are clustered to obtain a new picture type; and the new picture type is added to the preset scientific information picture and text knowledge base ontology to realize ontology update.
[0049] Specifically, a deep learning convolutional network is used to classify scientific research images based on visual features to obtain a first image type. Simultaneously, a deep learning recurrent network is used to classify the scientific research images based on the text matching them to obtain a second image type. Furthermore, when clustering the pending areas, a threshold is set. When the value of a new image category exceeds the set threshold, it is manually reviewed and, upon approval, added as a new image type.
[0050] The step of extracting the relationship between the two scientific research entities includes: matching the image types of the scientific research images according to preset rules to obtain the relationship between the scientific research images. For relationships outside the preset rules, a deep learning algorithm is used to extract the text corresponding to the scientific research images to obtain the relationship between the images.
[0051] Finally, the text that matches the scientific research picture is used to extract the knowledge of named entities, and the extracted entities and attributes are mapped to the ontology library. If the entity does not exist in the ontology library, the ontology library is updated as a new entity. After the knowledge extraction, the scientific research pictures and texts as well as the extracted knowledge are mapped to an instance of the ontology to realize the construction of the scientific information picture and text knowledge base. In addition, in this application, the picture classification, picture clustering, picture-text matching, text knowledge extraction and other places that require automatic processing and analysis are all realized by neural networks of different structures in deep learning technology. Graph database technology is also used to effectively store, manage and query complex relational data, thereby improving the value and utilization efficiency of the data. That is, the construction of the scientific information picture and text knowledge base in this application is based on this graph database technology.
[0052] In summary, this application has the following advantages over the prior art:
[0053] (1) Construct a scientific information graphic text body based on scientific research pictures in scientific and technological literature, deeply analyze the scientific research information contained in scientific and technological pictures, form a knowledge system based on scientific research pictures, and improve the application value of scientific and technological pictures.
[0054] (2) Match scientific research images with corresponding texts in scientific and technological literature, complete cross-modal data pre-labeling, and use the rigorous writing style of scientific and technological literature to match images and texts in the literature, and then further mine relevant information about the images from the text describing the images. This application not only utilizes the visual features of scientific and technological images, but also utilizes the language features of related texts, making the mining of images more extensive.
[0055] (3) This application constructs a scientific information picture and text knowledge base, which has the function of exporting data sets for multimodal large model training, and also has the function of semantic analysis of scientific and technological pictures, making the use of scientific and technological pictures more extensive.
[0056] From an application perspective, this application can accelerate the development of large multimodal models in the vertical field of scientific and technological intelligence. The development of large multimodal models relies on the support of high-quality multimodal datasets. This application's multimodal scientific information image and text knowledge base can provide dataset support for fine-tuning large multimodal models in the field of science and technology, accelerating their development.
[0057] This application can drive scientific and technological innovation. Scientific research images contain design information on scientific research plans, methods, process flows, experimental results, research transformation achievements, and information on experts, scholars, and research institutions in the field. The scientific research image and text knowledge base can deeply interpret this information, providing strong support for scientific and technological innovation.
[0058] This application can improve think tank research efficiency. Semantic retrieval of scientific and technological images allows think tank researchers to quickly and accurately utilize the knowledge provided by images, maximizing research efficiency and helping them better grasp scientific and technological trends, providing timely and accurate intelligence support for decision-making.
[0059] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a method for constructing a scientific information graphic and text knowledge base.
[0060] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0061] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0062] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0063] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0064] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0065] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0066] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0067] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for constructing a scientific information graphic and text knowledge base, characterized in that: The method for constructing a scientific information graphic and text knowledge base comprises: Acquire multiple scientific research documents containing images and perform preprocessing to obtain scientific research text and scientific research images corresponding to each of the scientific research documents; For each scientific research document, the corresponding scientific research text and the scientific research image are input into a preset scientific research image-text matching model to obtain a scientific research image-text pair; the preset scientific research image-text matching model is obtained by training a preset deep learning neural network using a training sample set; Obtaining a preset scientific information graphic and text knowledge base ontology; the preset scientific information graphic and text knowledge base ontology includes a plurality of triples, each triple including two scientific research entities, a relationship between the two scientific research entities, and attributes of each scientific research entity; Based on the scientific research image-text pair, the scientific research entity and the corresponding attributes, and the relationship between the two scientific research entities are extracted, and mapped into the preset scientific information image-text knowledge base ontology to obtain the scientific information image-text knowledge base; The preset scientific research image-text matching model includes a visual encoder, a text encoder, an external knowledge combination module, a cross-modal matching module and a loss function module; The input end of the visual encoder is used to receive the scientific research image, and the output end of the visual encoder is connected to the first input end of the cross-modal matching module; the input end of the text encoder is used to receive the scientific research text, and the output end of the text encoder is connected to the second input end of the cross-modal matching module; the output end of the cross-modal matching module is connected to the loss function module; The visual encoder uses a convolutional neural network to model the scientific research images as a whole to obtain a global visual representation of the images, and then uses a self-attention mechanism or a graph convolutional network to strengthen the relationship between entities within the visual modality; the text encoder uses a Transformer model to introduce a multi-head attention mechanism into the text encoder to model text sentences to obtain corresponding representations; the cross-modal matching module measures cross-modal semantic similarity by jointly embedding visual features and text features into a common space to estimate cross-modal semantic similarity; the loss function module uses the same ranking function for learning, such as hinge ranking loss and bidirectional hinge ranking loss; during the training process, the training ends when the value of the loss function reaches a preset loss condition; The external knowledge combination module is arranged between the visual encoder and the text encoder, and the external knowledge combination module is used to: measure the distance between the scientific research picture and the corresponding text paragraph position in the scientific research text, and send it to the visual encoder and the text encoder respectively as a picture-text pair distance feature.
2. The method for constructing a scientific information graphic and text knowledge base according to claim 1, characterized in that: The scientific research documents containing images include scientific research news documents, scientific research paper documents and scientific research patent documents; The steps of preprocessing the scientific research file containing the image include: For the scientific research news files, web page text and web page images are obtained using web page content recognition technology; For the scientific research paper documents and the scientific research patent documents, OCR technology is used to parse them to obtain document text and document images; the web page text and the document text both include text content and content structure; The web page text and the document text are marked as scientific research texts, and the web page pictures and the document pictures are marked as scientific research pictures.
3. The method for constructing a scientific information graphic and text knowledge base according to claim 1, characterized in that: The scientific research entities include: author, publishing institution, affiliated school department, paper ID and image ID; The attributes of the scientific research entity include: author attributes, institution attributes, department attributes, paper attributes and image attributes; the author attributes include name, date of birth, age, gender, professional title, research field and academic qualifications; the institution attributes include institution ID, institution name and institution introduction; the department attributes include college ID, college name, college introduction and college research field; the paper attributes include: paper title, paper abstract, paper keywords, paper field, paper publication journal and paper publication time; the image attributes include image name, image type, image keywords, image paper and image field; The relationship between the two scientific research entities includes: is-a relationship, subordinate relationship, causal relationship, time relationship, homonymous relationship and hierarchical relationship.
4. The method for constructing a scientific information graphic and text knowledge base according to claim 3, characterized in that: The image types include flow charts, architecture diagrams, design diagrams, experimental result diagrams, system interface diagrams, character diagrams and product schematics.
5. The method for constructing a scientific information graphic and text knowledge base according to claim 1, characterized in that: The step of measuring the distance between the scientific research image and the corresponding text paragraph position in the scientific research text in the external knowledge integration module includes: If there is a caption below or above the scientific image, the distance between the scientific image and the corresponding text paragraph in the scientific text is 0; If there is no legend below or above the scientific image, the distance is calculated using the following formula:
6. The method for constructing a scientific information graphic and text knowledge base according to claim 1, characterized in that: The step of extracting the attributes corresponding to the scientific research entity includes: When the attribute corresponding to the scientific research entity is the field to which the image belongs, the extraction process includes: using deep learning classification technology to classify the title, abstract or scientific research content of the scientific research document to which the scientific research image belongs, so as to obtain the field to which the image belongs; When the attribute corresponding to the scientific research entity is a picture type, the extraction process includes: classifying the pictures based on the scientific research pictures and the corresponding texts in the scientific research picture-text pair to obtain a first picture type and a second picture type; when the first picture type is consistent with the second picture type, the first picture type is used as the final picture type; when the first picture type is inconsistent with the second picture type, the scientific research pictures are stored in a waiting area; when the number of scientific research pictures in the waiting area reaches a preset number, all the scientific research pictures are clustered to obtain a new picture type; and the new picture type is added to the preset scientific information picture and text knowledge base ontology to realize ontology update.
7. The method for constructing a scientific information graphic and text knowledge base according to claim 1, characterized in that: The step of extracting the relationship between the two scientific research entities includes: matching the image types of the scientific research images according to preset rules to obtain the relationship between the scientific research images.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for constructing a scientific information graphic and text knowledge base according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a scientific information graphic and text knowledge base according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and system for constructing domain multi-modal knowledge graph
CN118568271A
Scientific and technical literature flow chart entity and relation extraction method based on retrieval enhancement
CN119003788A