Map image-text data set construction method based on multi-technology fusion

Through the multi-technology fusion method, the Web of Science API, Sci-hub database and llama model are used to automatically collect and annotate map literature, solving the problems of inefficiency and inconsistency in traditional methods, and achieving efficient and accurate construction of map graphic data sets.

CN120337882APending Publication Date: 2025-07-18SOUTH CHINA NORMAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510410787.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional mapping literature data collection and labeling are inefficient, manual screening and labeling are time-consuming and labor-intensive, and the labeling quality is inconsistent.

Method used

The mapping graphic data set construction method based on multi-technology fusion is adopted, and the Web of Science API, Sci-hub database and Python scripts are used to automatically collect and convert data formats, and intelligent annotation is combined with the llama model to achieve accurate identification of pictures and texts and automatically generate parsed content.

Benefits of technology

It greatly improves data collection efficiency and coverage, ensures the consistency and accuracy of labeling, reduces labor costs, and generates high-quality image analysis content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337882A_ABST
    Figure CN120337882A_ABST
Patent Text Reader

Abstract

The invention provides a method for constructing a cartographic image-text data set based on multi-technology fusion, which comprises the following steps of: acquiring literature information from a source website based on keyword information, and generating a literature table based on the literature information; reading each piece of literature information in the literature table in sequence, and downloading a PDF document corresponding to the literature information from a paper website; using a document format conversion tool to convert the PDF document into a Markdown format document; identifying the content in the Markdown format document so as to find out the picture and the picture number; searching and extracting mention paragraphs menting the picture numbers in the full text, and sorting the pictures, the picture numbers and text information of the mention paragraphs; and processing the picture, the picture number and the text information of the mention paragraph based on a plama model to generate picture analysis content. According to the method and the device, concise and accurate general description and comprehensive detailed description can be generated, the manual annotation cost is greatly reduced, and the annotation consistency and accuracy are also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent annotation of cartographic literature data, and specifically relates to a method for constructing a cartographic text and image dataset based on multi-technology integration. Background Art

[0002] Traditional data collection mainly relies on manual retrieval of academic databases and professional websites by researchers. They need to screen through a large number of documents one by one, search on the database retrieval page by entering keywords, and then manually view information such as the paper title and abstract to determine whether it meets the requirements. When preprocessing data, general PDF conversion tools such as Adobe Acrobat are usually used for PDF to text conversion. However, the text format after conversion by these tools is chaotic, and the separation and recognition effects of charts, pictures and text are not good. In terms of data annotation, it mainly relies on manual annotation, and professional personnel write annotations according to the picture content. This method is extremely inefficient. Annotators need to spend a lot of time carefully observing the pictures and organizing language descriptions. Moreover, due to differences in professional levels and understandings among different annotators, the consistency and accuracy of annotations are difficult to guarantee. Summary of the Invention

[0003] The purpose of the present invention is to overcome the above technical deficiencies, and provide a method for constructing a cartographic text and image dataset based on multi-technology integration, so as to solve the technical problems of low data annotation efficiency and uneven annotation quality in the prior art.

[0004] To achieve the above technical purpose, in a first aspect, the technical solution of the present invention provides a method for constructing a cartographic text and image dataset based on multi-technology integration, including the steps of:

[0005] Obtain literature information from the source website based on keyword information, and generate a literature table based on the literature information;

[0006] Read each piece of literature information in the literature table in sequence, and download the PDF document corresponding to the literature information from the paper website based on the DOI number of the literature information;

[0007] Use a document format conversion tool to convert the PDF document into a Markdown format document;

[0008] Use a Python script to identify the content in the Markdown format document based on a preset text and image rule, so as to find pictures and captions, and extract picture numbers using regular expressions;

[0009] Based on the extracted picture numbers, search and extract the mentioned paragraphs in the full text that mention the picture numbers, and organize the text information of the pictures, the picture numbers and the mentioned paragraphs;

[0010] Process the text information of the said picture, the picture number and the mentioned paragraph based on the Llama model to generate picture parsing content.

[0011] Compared with the prior art, the beneficial effects of the present invention include:

[0012] The present invention designs an intelligent crawler program specifically for the field of cartography. Combining the characteristics of the Web of Science and Sci-hub databases, and using complex screening rules and advanced web parsing technologies, it can quickly and comprehensively collect cartography papers. The crawler program can accurately locate relevant papers according to the keywords in the sub-fields of cartography, and at the same time effectively process dynamic web pages and complex page structures, greatly improving the data collection efficiency and coverage, and providing sufficient data guarantee for building a large-scale graphic dataset. And it uses the open-source tool Marker combined with specific text and image processing algorithms to perform efficient format conversion and information extraction on cartography literature. It can not only accurately separate pictures and text, but also use technologies such as semantic analysis and regular expressions to accurately identify picture numbers, captions and related detailed descriptions, ensuring the accuracy and integrity of data preprocessing and laying a solid foundation for subsequent data processing and analysis. Finally, with the help of the advanced vision-language large model Llama 3.2 - vision 11b, automatic generation of picture annotations is realized. By learning the association patterns between pictures and text, the model can generate concise and accurate general descriptions and comprehensive detailed descriptions. This not only greatly reduces the manual annotation cost, but also improves the consistency and accuracy of annotations.

[0013] According to some embodiments of the present invention, obtaining literature information from a source website based on keyword information and generating a literature table based on the literature information, including the steps of:

[0014] Obtain the API interface of the Web of Science Starter website;

[0015] Input retrieval keywords related to cartography, send an HTTP request to the Web of Science Starter website through the API interface, and carry the set keyword information, so that the website database retrieves according to the keyword information;

[0016] Receive the literature table presented in the xlsx format returned by the website. Each literature table contains the titles, authors, DOIs, and publication journal information of multiple literatures.

[0017] According to some embodiments of the present invention, read each piece of literature information in the literature table in sequence, and download the PDF document corresponding to the literature information from the paper website based on the DOI number of the literature information, including the steps

[0018] Read each piece of the literature information in the literature table in sequence, and determine whether the literature contains a DOI number;

[0019] If the literature information contains a DOI number, send a get request to the Sci-hub website based on the DOI number; if the URL in the response can be successfully obtained, obtain the PDF document of the literature information through the URL.

[0020] According to some embodiments of the present invention, after determining whether the literature contains a DOI number, the following steps are included:

[0021] If the literature does not contain a DOI number or the URL cannot be obtained, continue to process the next piece of literature information.

[0022] According to some embodiments of the present invention, the llama model is used to process the text information of the picture, the picture number, and the mentioned paragraph to generate picture parsing content, including the following steps:

[0023] Build and deploy the running environment of the llama model;

[0024] After the running environment is built, initialize the dataset path and output path parameters of the picture, the picture number, and the text information of the mentioned paragraph;

[0025] Extract the text information of the picture, the caption, and the mentioned paragraph from the preprocessed dataset;

[0026] Initialize the prompt instructions, and combine the extracted text information of the picture, the caption, and the mentioned paragraph to generate short annotation instructions and long description instructions respectively;

[0027] Input the short annotation instructions and the long description instructions into the Llama model. The Llama model executes the prompt instructions according to the learned association pattern between the picture and the text and generates picture parsing content, obtaining a short annotation (module-caption) and a long description (module-description).

[0028] According to some embodiments of the present invention, the document format conversion tool is the Python-based document format conversion tool Marker.

[0029] In a second aspect, the technical solution of the present invention provides a cartographic graphic and text dataset construction system based on multi-technology integration, including:

[0030] A literature table generation module, which obtains literature information from a source website based on keyword information and generates a literature table based on the literature information;

[0031] The PDF document download module reads each piece of literature information in the literature table in sequence, and downloads the PDF document corresponding to the literature information from the paper website based on the DOI number of the literature information;

[0032] The document format conversion tool is communicatively connected to the PDF document download module and is used to convert the PDF document into a Markdown format document;

[0033] The content extraction module uses a Python script to identify the content in the Markdown format document based on preset graphic rules, so as to find pictures and captions, and extracts the picture numbers using regular expressions;

[0034] The text search module searches and extracts the mentioned paragraphs that mention the picture numbers in the full text based on the extracted picture numbers, and organizes the text information of the pictures, the picture numbers, and the mentioned paragraphs;

[0035] The parsing content generation module processes the text information of the pictures, the picture numbers, and the mentioned paragraphs based on the llama model to generate picture parsing content.

[0036] In the field of cartography, the effective collection and processing of paper data are crucial. This invention has made significant improvements in multiple key aspects. In terms of data collection, the traditional method of manually screening papers one by one is inefficient and prone to omissions. With the help of the applied Web of Science Starter API and combined with the keyword retrieval strategy, the retrieval efficiency of this invention has been greatly improved. In terms of obtaining the full text of literature, the literature crawling program developed based on the Sci-hub website can automatically obtain the PDF full text according to information such as the DOI number of the literature, broadening the coverage of data collection to build a large-scale and complete cartography graphic dataset.

[0037] Traditional PDF literature conversion tools often have problems such as format chaos and information loss. This invention uses the open-source tool Marker to efficiently convert PDF literature into Markdown format, effectively avoiding these drawbacks. At the same time, the information extraction technology of Python scripts based on graphic rules and regular expressions can accurately separate and match information when identifying pictures, captions, and related paragraphs, greatly improving the accuracy and integrity of data processing.

[0038] In the past, manually annotating picture notes not only consumed a large amount of manpower and time, but also due to differences in understanding among annotators, it was difficult to guarantee the quality of the notes. This invention uses the llama3.2-vision 11B model to achieve automatic generation of annotations, which not only greatly improves the efficiency but also ensures the consistency and accuracy of the annotations.

[0039] In a third aspect, the technical solution of the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the method for constructing a cartographic graphic dataset based on multi-technology fusion as described in any one of the first aspects.

[0040] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which the abstract drawing should be exactly the same as one of the specification drawings:

[0042] Figure 1 It is a flowchart of the method for constructing a cartographic graphic dataset based on multi-technology fusion provided by an embodiment of the present invention;

[0043] Figure 2 It is a schematic diagram of the process for obtaining literature information of the method for constructing a cartographic graphic dataset based on multi-technology fusion provided by an embodiment of the present invention;

[0044] Figure 3 It is a schematic diagram of the process for crawling PDF literature based on Sci-hub of the method for constructing a cartographic graphic dataset based on multi-technology fusion provided by an embodiment of the present invention;

[0045] Figure 4 It is a schematic diagram of the Markdown format conversion process of the method for constructing a cartographic graphic dataset based on multi-technology fusion provided by an embodiment of the present invention;

[0046] Figure 5 It is a schematic diagram of the technical process for extracting graphic-text pairs from Markdown-format literature of the method for constructing a cartographic graphic dataset based on multi-technology fusion provided by an embodiment of the present invention;

[0047] Figure 6 It is a schematic diagram of the process for generating graphic-text parsing based on the llama3.2-vision model of the method for constructing a cartographic graphic dataset based on multi-technology fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] It should be noted that although the functional modules are divided in the system schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the system or a different order from that in the flowchart. Terms such as "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence.

[0050] Llama 3.2-vision 11B is a medium-sized model in the Llama 3.2 series released by Meta that supports visual tasks and has the following characteristics:

[0051] Model architecture: To support image input, Meta trained a set of adapter weights to integrate a pre-trained image encoder into the pre-trained language model. The adapter consists of a series of cross-attention layers that can feed the image encoder representation to the language model and train the adapter on "text-image pair" data to align the image representation with the language representation, thus keeping all pure text capabilities unchanged and facilitating developers to replace Llama 3.1 with Llama 3.2.

[0052] Training process: Starting from the pre-trained Llama 3.1 text model, first add the image adapter and encoder, and then pre-train on a large-scale noisy paired (image, text) data. Also use synthetic data generation, filter and enhance the questions and answers on in-domain images with the Llama 3.1 model, and use the reward model to rank all candidate answers to provide high-quality fine-tuning data. In addition, to obtain a highly secure and useful model, safety mitigation data is also added.

[0053] Functional features: Support image reasoning, including document-level diagram understanding, image captioning, and visual localization tasks. For example, it can directly locate things in an image according to a natural language description, can handle visual and text reasoning tasks, has the ability of chain-of-thought reasoning, which enhances its problem-solving ability, especially when dealing with complex visual reasoning tasks. Support text input in various languages, such as English, German, French, Hindi, etc. Its 128k context length allows for extended multi-turn conversations. When processing images, focusing on only one image at a time can maintain quality and optimize memory usage.

[0054] Performance: It is comparable to the industry-leading basic models Claude 3Haiku and GPT4o-mini in a series of visual understanding tasks such as image recognition, and performs well in the evaluation on more than 150 benchmark datasets involving multiple languages, providing strong support for research and applications in related fields.

[0055] Refer toFigures 1 to 6 , Figure 1 is a flowchart of a method for constructing a cartographic text and image dataset based on multi-technology fusion provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the literature information acquisition process of a method for constructing a cartographic text and image dataset based on multi-technology fusion provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the Sci-hub-based PDF literature crawling process of a method for constructing a cartographic text and image dataset based on multi-technology fusion provided by an embodiment of the present invention; Figure 4 is a schematic diagram of the Markdown format conversion process of a method for constructing a cartographic text and image dataset based on multi-technology fusion provided by an embodiment of the present invention; Figure 5 is a schematic diagram of the text and image pair extraction technology process of a method for constructing a cartographic text and image dataset based on multi-technology fusion provided by an embodiment of the present invention;

[0056] Figure 6 is a schematic diagram of the text and image analysis generation process based on the llama3.2-vision model of a method for constructing a cartographic text and image dataset based on multi-technology fusion. The method for constructing a cartographic text and image dataset based on multi-technology fusion includes but is not limited to the following steps:

[0057] Step S110, obtain literature information from the source website based on keyword information, and generate a literature table based on the literature information;

[0058] Step S120, read each piece of literature information in the literature table in sequence, and download the PDF document corresponding to the literature information from the paper website based on the DOI number of the literature information;

[0059] Step S130, use a document format conversion tool to convert the PDF document into a Markdown format document;

[0060] Step S140, use a Python script to identify the content in the Markdown format document based on a preset text and image rule, so as to find pictures and captions, and extract picture numbers using regular expressions;

[0061] Step S150, based on the extracted picture numbers, search and extract the mentioned paragraphs that mention the picture numbers in the full text, and organize the text information of the pictures, picture numbers, and mentioned paragraphs;

[0062] Step S160, process the text information of the pictures, picture numbers, and mentioned paragraphs based on the llama model to generate picture analysis content.

[0063] In one embodiment, the method for constructing a cartographic graphic dataset based on multi-technology integration includes the steps of: obtaining literature information from a source website based on keyword information and generating a literature table based on the literature information; sequentially reading each piece of literature information in the literature table and downloading the PDF document corresponding to the literature information from a paper website based on the DOI number of the literature information; using a document format conversion tool to convert the PDF document into a Markdown format document; using a Python script to identify the content in the Markdown format document based on a preset graphic rule, so as to find pictures and captions, and extracting picture numbers using regular expressions; searching and extracting the mentioned paragraphs that mention the picture numbers in the full text based on the extracted picture numbers, and organizing the pictures, picture numbers, and text information of the mentioned paragraphs; processing the text information of the pictures, picture numbers, and mentioned paragraphs based on the llama model to generate picture parsing content.

[0064] By obtaining literature information from a source website based on keyword information, the present invention can widely collect various types of literature materials related to cartography, avoiding the limitations brought by a single data source, making the constructed dataset cover richer and more representative content, and can include cartography-related content from different perspectives and research directions to the greatest extent. The present invention uses the DOI number to download the corresponding PDF document from a paper website, ensuring the accuracy and authority of literature acquisition. As the unique identifier of the literature, the DOI number helps to accurately locate the required academic papers, ensuring that the research results included in the dataset are all professionally reviewed and of high quality, improving the overall quality of the dataset.

[0065] The present invention uses a document format conversion tool to convert the PDF document into a Markdown format document, which is convenient for subsequent unified processing and analysis of the document content. The Markdown format has the characteristics of being concise and easy to read, convenient for editing, and being able to better compatible with various text processing scripts, making subsequent operations such as content identification and extraction can be carried out on the basis of a relatively standardized text format, reducing the processing difficulty caused by format differences.

[0066] The present invention uses Python scripts to perform content recognition on Markdown format documents based on pre-set graphic rules. This rule-based approach makes the search for images and image numbers more logical and accurate. Through clear rule constraints, the target content can be accurately located, avoiding omissions or errors that may occur in manual searches, and laying a good foundation for subsequent operations such as extracting related paragraphs. The present invention extracts image numbers with the help of regular expressions. Regular expressions have powerful functions in processing text pattern matching, and can efficiently and accurately filter out image number information that conforms to a specific format from complex texts, further improving the accuracy and efficiency of data extraction, and ensuring the reliability of subsequent associations with other content based on numbers.

[0067] Based on the extracted image number, the full text is searched and the paragraphs mentioning the image number are extracted, which realizes the effective association between the image and the relevant text description. The constructed data set is not just an isolated image and a simple number, but an organic whole that forms a logical and corresponding combination of images and texts, which helps to better understand the meaning expressed by the image and its role in the entire study when using the data later.

[0068] The present invention uses the llama model to process the sorted pictures, picture numbers and text information of the mentioned paragraphs to generate picture analysis content. The llama model can rely on its powerful natural language processing capabilities to deeply understand, summarize and expand text information, and automatically generate high-quality picture analysis content, which greatly improves the efficiency of interpreting picture-related content during the data set construction process. At the same time, it can also make the generated analysis more in line with professional logic and language expression habits, and enhance the practicality and readability of the data set. In general, the construction method comprehensively uses a variety of technical means, from data acquisition, format conversion, content recognition to information integration and intelligent processing and other links to work together, and can efficiently and accurately construct a high-quality and rich cartographic graphic data set.

[0069] In one embodiment, the method for constructing a cartographic graphic dataset based on multi-technology integration includes the steps of: obtaining literature information from a source website based on keyword information and generating a literature table based on the literature information; sequentially reading each piece of literature information in the literature table, and downloading the PDF document corresponding to the literature information from a paper website based on the DOI number of the literature information; using a document format conversion tool to convert the PDF document into a Markdown format document; using a Python script to identify the content in the Markdown format document based on a preset graphic rule, so as to find pictures and captions, and extracting picture numbers using regular expressions; based on the extracted picture numbers, searching and extracting the mentioned paragraphs that mention the picture numbers in the full text, and organizing the pictures, picture numbers, and text information of the mentioned paragraphs; processing the text information of the pictures, picture numbers, and mentioned paragraphs based on the llama model to generate picture parsing content.

[0070] Obtaining literature information from a source website based on keyword information and generating a literature table includes the steps of: obtaining the API interface of the Web of Science Starter website; inputting retrieval keywords related to cartography, sending an HTTP request to the Web of Science Starter website through the API interface, and carrying the set keyword information, so that the website database retrieves according to the keyword information; receiving the literature table presented in the xlsx format returned by the website, and each literature table contains the titles, authors, DOIs, and publication journal information of multiple literatures.

[0071] After obtaining the xlsx document of the literature information, it is necessary to further obtain the full text PDF of the literature. First, initialize relevant parameters, including the xlsx document of the literature information, input path, output path, etc. Then sequentially read each piece of literature information in the xlsx document and judge whether the literature contains a DOI number. The DOI number is a digital object unique identifier that can uniquely identify a piece of literature. If the literature contains a DOI number, a get request is sent to the Sci-hub website based on the DOI number. If the URL in the response can be successfully obtained, the PDF document of the literature can be obtained through this URL. If it does not contain a DOI number or the URL cannot be obtained, continue to process the next piece of literature information. Through such a process, the full text materials of cartography papers are comprehensively and targeted collected, greatly improving the efficiency and coverage of data collection.

[0072] By obtaining the API interface of the Web of Science Starter website, the present invention can automate data acquisition. Instead of manually entering keywords one by one on the website for retrieval and flipping through pages to find literature, the system can automatically send requests to the website according to the set program, greatly saving time and labor costs. The API interface allows direct communication between systems and can obtain the required information from the website database more quickly compared to manual operations. Especially when a large amount of literature information is needed, this efficient data acquisition method can significantly improve work efficiency.

[0073] The API interface can accurately transfer the retrieval keywords entered by the user to the website database, avoiding problems such as spelling mistakes or inaccurate retrieval condition settings that may occur in manual input. This makes the retrieval results more accurate and can obtain literature information related to cartography more precisely. The data obtained by the present invention using the API interface usually has a unified format, which is convenient for subsequent processing and analysis. In this method, the literature information returned by the website is presented in the xlsx format, and this standardized format allows the data to be easily imported into various data analysis tools for further processing.

[0074] Entering retrieval keywords related to cartography can focus the retrieval results on the field of cartography. Users can flexibly set keywords according to their research needs to obtain literature information closely related to specific topics. This helps to reduce the interference of irrelevant information and improve the efficiency of literature screening. Users can adjust the retrieval keywords as needed to expand or narrow the retrieval scope. For example, if more extensive cartography-related literature is needed, broader keywords can be used; if literature in a specific direction is needed, more specific keyword combinations can be used. This flexibility enables this method to adapt to different research needs.

[0075] The xlsx format is a common spreadsheet format that is widely supported by various data analysis software, such as Microsoft Excel, the Pandas library in Python, etc. Researchers can conveniently use these tools to perform operations such as data cleaning, screening, and statistical analysis on the literature table to mine valuable information in the literature. The xlsx table clearly shows information such as the title, author, DOI, and publication journal of each literature in the form of rows and columns, facilitating researchers to quickly browse and view the basic situation of the literature. At the same time, the table form is also convenient for sorting, screening, and comparing literature information, helping researchers better understand the distribution and characteristics of the literature.

[0076] Further, read each piece of literature information in the literature table in sequence, and download the PDF document corresponding to the literature information from the paper website based on the DOI number of the literature information, including the steps of: reading each piece of literature information in the literature table in sequence, and determining whether the literature contains a DOI number; if the literature information contains a DOI number, send a get request to the Sci-hub website; if the URL in the response can be successfully obtained, obtain the PDF document of the literature information through the URL.

[0077] Reading each piece of literature information in the literature table in sequence ensures the orderliness of the processing process. This can avoid missing some literature information and prevent duplicate processing of the same literature, ensuring that the entire literature download process proceeds in an orderly manner. Determining whether the literature contains a DOI number can quickly screen out the literature that can be used for further download operations. The DOI number is the unique identifier of academic literature, with global uniqueness and permanence. Literature containing a DOI number usually has higher authority and traceability. Through this screening step, high-quality and accurately retrievable literature resources can be focused on.

[0078] Sci-hub is a platform that provides free downloads of a large number of academic literatures and integrates the resources of many academic databases. Sending a request to Sci-hub through the DOI number allows trying to obtain literatures from multiple sources on one platform, avoiding the cumbersome process of separately searching for and downloading literatures in multiple different academic databases, and greatly improving the efficiency of literature acquisition. Many academic literatures are difficult for ordinary users to directly access due to copyright or paywall restrictions. Sci-hub uses technical means to break through these restrictions, enabling researchers to obtain the required literatures in a relatively simple way. Sending a request to it based on the DOI number can take advantage of the platform's advantages and has a high probability of successfully obtaining the PDF document of the literature, especially for those high-quality academic resources restricted by access.

[0079] When the URL is successfully obtained from the response, the PDF document of the literature can be directly downloaded through this URL. This direct acquisition method reduces intermediate links, avoids time waste and possible errors caused by multiple jumps or complex operations, and can quickly and accurately save the required literature to the local. By implementing the process of downloading the PDF document from the URL through code, batch operations can be easily performed. For a table containing a large amount of literature information, each piece of literature information can be processed in sequence, realizing automated literature downloads, further improving work efficiency, and being particularly suitable for large-scale literature data collection work.

[0080] In the data preprocessing stage, open-source tools and specific algorithms are mainly used to process the collected cartographic literature. First, install the open-source tool Marker, which is a key tool for document format conversion. After installation, determine the storage location of the PDF document and the output location of the Markdown document. According to actual needs, you can choose to convert a single PDF document or convert them in batches, and then input the conversion instructions of the Marker tool. After processing, you can obtain the Markdown format file of the literature. The Markdown format is more convenient for subsequent extraction and processing of pictures and picture-text description information than the PDF format.

[0081] After determining whether the literature contains a DOI number, the steps include: if it does not contain a DOI number or the URL cannot be obtained, continue to process the next piece of literature information.

[0082] Process the pictures, picture numbers, and text information of the mentioned paragraphs based on the llama model to generate picture parsing content, including the steps of: building and deploying the running environment of the llama model; after the running environment is built, initialize the dataset paths and output path parameters of the pictures, picture numbers, and text information of the mentioned paragraphs; extract the pictures, captions, and text information of the mentioned paragraphs from the preprocessed dataset; initialize the prompt instructions, and combine the extracted pictures, captions, and text information of the mentioned paragraphs to generate short annotation instructions and long description instructions respectively; input the short annotation instructions and long description instructions into the llama model, and the llama model executes the prompt instructions and generates picture parsing content according to the learned association pattern between pictures and text, obtaining short annotations (module-caption) and long descriptions (module-description).

[0083] Extracting the required key information from the preprocessed dataset can remove the interference of irrelevant data, enabling the model to focus on processing the pictures, captions, and text information of the mentioned paragraphs. This can improve the processing efficiency and accuracy of the model and avoid performance degradation caused by processing too much useless information.

[0084] The present invention performs preprocessing before information extraction, which can perform operations such as data cleaning and conversion to ensure the quality and consistency of the data. This helps the model better understand and process the data, thereby generating more accurate and valuable picture parsing content.

[0085] By initializing the prompt instructions, short comments and long description instructions can be generated respectively in combination with the extracted information, and the output content of the model can be customized according to specific requirements. The short comment can concisely summarize the key information of the picture, while the long description can provide more detailed and in-depth interpretations to meet the needs of picture parsing in different scenarios. The prompt instructions provide a clear task orientation for the Llama model, guiding the model to reason and generate according to the learned association patterns between pictures and texts. This way can fully utilize the language generation ability of the model, making the generated content more in line with expectations and professional requirements.

[0086] The Llama model has powerful semantic understanding and language generation capabilities. It can deeply analyze and understand picture and related text information based on the input instructions and learned association patterns, and generate high-quality picture parsing content. The generated content can not only accurately reflect the features and meanings of the pictures, but also be expanded and interpreted to a certain extent, providing valuable references for researchers. The Llama model has been pre-trained on a large amount of data and has good learning ability and generalization. It can learn general association patterns from different pictures and text information and apply them to new input data to generate reasonable and effective parsing content. This makes the model have strong adaptability and flexibility when processing diverse cartographic pictures.

[0087] In one embodiment, a cartographic graphic and text data set construction system based on multi-technology integration includes: a table generation module, which obtains literature information from a source website based on keyword information and generates a literature table based on the literature information; a PDF document download module, which sequentially reads each piece of literature information in the literature table and downloads the PDF document corresponding to the literature information from a paper website based on the DOI number of the literature information; a document format conversion tool, which is communicatively connected to the PDF document download module and is used to convert the PDF document into a Markdown format document; a content extraction module, which uses a Python script to identify the content in the Markdown format document based on pre-set graphic and text rules, so as to find pictures and picture numbers, and extracts the picture numbers using regular expressions; a text search module, which searches and extracts the mentioned paragraphs that mention the picture numbers in the full text based on the extracted picture numbers, and organizes the pictures, picture numbers, and text information of the mentioned paragraphs; a parsing content generation module, which processes the pictures, picture numbers, and text information of the mentioned paragraphs based on the Llama model to generate picture parsing content.

[0088] A memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0090] In addition, an embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions, which are executed by a processor or a controller, for example, executed by a processor in the above terminal embodiment, so that the processor can execute the method for constructing a cartographic graphic dataset based on multi-technology integration in the above embodiment.

[0091] Those of ordinary skill in the art can understand that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0092] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present invention.

[0093] The specific embodiments of the present invention described above do not constitute a limitation on the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A method for constructing a cartographic graphic dataset based on the integration of multiple technologies, characterized in that, Including steps: Obtain literature information from the source website based on keyword information, and generate a literature table based on the literature information; Read each piece of literature information in the literature table in sequence, and download the PDF document corresponding to the literature information from the paper website based on the DOI number of the literature information; Use a document format conversion tool to convert the PDF document into a Markdown format document; Use a Python script to identify the content in the Markdown format document based on pre-set graphic rules, so as to find pictures and captions, and extract picture numbers using regular expressions; Based on the extracted picture numbers, search and extract the mentioned paragraphs that mention the picture numbers in the full text, and organize the pictures, the picture numbers, and the text information of the mentioned paragraphs; Process the text information of the pictures, the picture numbers, and the mentioned paragraphs based on the llama model to generate picture parsing content.

2. The method for constructing a cartographic graphic and text data set based on multi-technology integration according to claim 1, wherein Obtain literature information from the source website based on keyword information, and generate a literature table based on the literature information, including steps: Obtain the API interface of the Web of Science Starter website; Input retrieval keywords related to cartography, send an HTTP request to the Web of Science Starter website through the API interface, and carry the set keyword information, so that the website database retrieves according to the keyword information; Receive the literature table presented in xlsx format returned by the website, and each literature table contains the title, author, DOI, and publication journal information of multiple literatures.

3. The method for constructing a cartographic graphic and text data set based on multi-technology integration according to claim 1, wherein, Read each piece of literature information in the literature table in sequence, and download the PDF document corresponding to the literature information from the paper website based on the DOI number of the literature information, including steps Read each piece of the literature information in the literature table in sequence, and judge whether the literature contains a DOI number; If the literature information contains a DOI number, send a get request to the Sci-hub website based on the DOI number; if the URL in the response can be successfully obtained, obtain the PDF document of the literature information through the URL.

4. The method for constructing a cartographic graphic dataset based on multi-technology integration according to claim 3, wherein, After judging whether the literature contains a DOI number, including steps: If it does not contain a DOI number or the URL cannot be obtained, continue to process the next piece of literature information.

5. The method for constructing a cartographic graphic dataset based on multi-technology integration according to claim 1, wherein, Process the text information of the pictures, the picture numbers, and the mentioned paragraphs based on the llama model to generate picture parsing content, including steps: Build and deploy the operating environment of the llama model; After the operating environment is built, initialize the dataset paths and output path parameters of the text information of the pictures, the picture numbers, and the mentioned paragraphs; Extract the text information of pictures, captions, and mentioned paragraphs from the pre-processed dataset; Initialize the prompt instructions, and combine the extracted text information of pictures, captions, and mentioned paragraphs to generate short annotation instructions and long description instructions respectively; Input short comment instructions and long descriptions into the Llama model. The Llama model executes the prompt instructions and generates picture parsing content according to the learned association pattern between pictures and texts, obtaining short comments (module-caption) and long descriptions (module-description).

6. The method for constructing a cartographic graphic dataset based on multi-technology integration according to claim 1, wherein The document format conversion tool is the Python-based document format conversion tool Marker.

7. A cartographic graphic and text data set construction system based on multi-technology integration, characterized in that, It includes: A literature table generation module that obtains literature information from a source website based on keyword information and generates a literature table based on the literature information; A PDF document download module that sequentially reads each piece of literature information in the literature table and downloads the PDF document corresponding to the literature information from a paper website based on the DOI number of the literature information; A document format conversion tool that is communicatively connected to the PDF document download module and is used to convert the PDF document into a Markdown format document; A content extraction module that uses a Python script to identify the content in the Markdown format document based on preset picture and text rules, thereby finding pictures and captions, and extracting picture numbers using regular expressions; A text search module that searches and extracts the mentioned paragraphs referring to the picture numbers in the full text based on the extracted picture numbers, and organizes the text information of the pictures, the picture numbers, and the mentioned paragraphs; A parsing content generation module that processes the text information of the pictures, the picture numbers, and the mentioned paragraphs based on the llama model to generate picture parsing content.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the method for constructing a cartographic picture and text data set based on multi-technology integration according to any one of claims 1 to 6.