Dinosaur footprint image data set construction method and system based on large language model
Through a method based on a large language model, image data is mined from research results and image data sets are constructed, which solves the problem of lack of effective image data, and realizes automatic identification of dinosaur footprints and efficient data acquisition.
Patent Information
- Application Number
- PCT/CN2024/098859
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-06-13
- Publication Date
- 2025-05-08
AI Technical Summary
Due to the lack of sufficiently effective image data, artificial intelligence algorithms cannot be effectively trained to assist in the automatic recognition of dinosaur footprints.
Using a method based on a large language model, data sets that can be used to train image recognition algorithms are constructed by mining image data in research results, collecting, processing, annotating and verifying dinosaur footprint fossil images.
It realizes the construction of dinosaur footprint image datasets that can be used to train image recognition algorithms, supports automatic recognition of dinosaur footprints, improves data acquisition efficiency and accuracy, and reduces artificial errors.
Smart Images

Figure CN2024098859_08052025_PF_FP_ABST
Abstract
Description
Method and system for constructing dinosaur footprint image dataset based on large language model Technical Field
[0001] The present invention belongs to the technical field of image data processing, and in particular relates to a method and system for constructing a dinosaur footprint image dataset based on a large language model. Background Art
[0002] In paleontological research, the identification of dinosaur footprints is primarily performed manually by paleontologists. With the advancement of computer image recognition technology, computer algorithms may be able to automatically identify dinosaur footprints. However, due to a lack of sufficient image data, artificial intelligence algorithms cannot be effectively trained to assist in the identification of dinosaur footprints. Therefore, a method for constructing a dinosaur footprint image dataset would address the problem of the inability to apply automatic identification technology. Technical issues
[0003] Through the above analysis, the problems and defects of the existing technology are: due to the lack of sufficient and effective image data, the artificial intelligence algorithm cannot be effectively trained to assist in the identification of dinosaur footprints. Technical Solutions
[0004] In response to the problems existing in the prior art, the present invention provides a method and system for constructing a dinosaur footprint image dataset based on a large language model.
[0005] The present invention is implemented as follows: a method for constructing a dinosaur footprint image dataset based on a large language model, comprising:
[0006] S1, mining data source: the data comes from the image data contained in the research results, which are mainly published papers;
[0007] S2, collects images of dinosaur footprint fossils;
[0008] S3, processing dinosaur footprint fossil images;
[0009] S4, annotated dinosaur footprint fossil images;
[0010] S5, verify the validity of image data, verify the validity of image data through deep neural network.
[0011] Furthermore, the specific steps of S2 are as follows:
[0012] S201, write a Python crawler program to first crawl the download link of the PDF format paper from the web page and download the paper to the local computer;
[0013] S202: Use tools based on the Large Language Model (LLM), including chatPDF (a tool based on the ChatGPT API), to analyze the paper content and preliminarily determine whether there are footprint images in the paper. If so, proceed to the next step; otherwise, find a new paper.
[0014] S203, write a Python program to use the LLM analysis results to extract relevant information of the paper, including its DOI, title, author, and references; check whether the paper has been processed and registered in the historical information through the DOI; if not, proceed to the next step; otherwise, ignore the paper;
[0015] S204, write a Python program to read the papers one by one that are confirmed to contain new image data after analysis by the large model LLM, create a folder with the same name as the paper, and a series of subfolders under the folder with the same name, including a Label folder, a Processed folder, a Non-label folder, and an Original folder, which are used to store images with different attributes;
[0016] S205, write a Python program to automatically extract images from the paper and save the images to the corresponding Original folder. The image name is "paper title_image extraction serial number";
[0017] S206, using a tool based on a general large language model (LLM) to analyze the content of the paper, and based on the arguments and viewpoints of the paper on the image, determine whether the extracted image is a footprint image. If so, retain it, otherwise delete it.
[0018] Furthermore, S3 specifically includes: For the dinosaur footprint images in the paper, the authors will make some marks in the images in order to more clearly express the new discoveries, including using white powder to indicate the outline; at this time, these artificial marks need to be erased using photo editing software or web pages, and images that do not need to be processed are placed in the Non-label folder, and images that need to be processed are placed in the Label folder, and after processing, they are placed in the Processed folder.
[0019] Furthermore, in S4, the annotation steps are as follows:
[0020] S401, Image Binary Classification: Use the Large Language Model (LLM) to read the paper. Based on the paper description, read the footprint images from the corresponding Label and Non-label folders, determine whether the footprint images contain dinosaur footprints, and copy them to the Positive folder (if dinosaur footprint images exist) and the Negative folder (if no dinosaur footprint images exist).
[0021] S402, marking position information: marking the position information of the dinosaur footprints in the Positive folder, using a polygonal box or a rectangular box of the labeling software Labelme to mark the information, and the position information of each picture is saved in the "picture name.json" file.
[0022] Furthermore, S5 specifically includes:
[0023] S51, using a deep neural network-based classifier, such as ResNet50, a prediction score greater than 0.5 is used to determine that the image contains dinosaur footprint fossils, and an ACC (accuracy) index greater than 0.6 is a valid dataset for classification;
[0024] S52, using a target detector based on a deep neural network, including Yolox, the predicted area with a prediction score greater than 0.5 and an IoU (between the predicted area and the labeled area) greater than 0.3 is the range of dinosaur footprint fossils correctly detected in the image, and the ACC index greater than 0.6 is a valid data set that can be used for target detection, where,
[0025] ;
[0026] ;
[0027] .
[0028] Another object of the present invention is to provide a system for constructing a dinosaur footprint image dataset based on a large language model using the method for constructing a dinosaur footprint image dataset based on a large language model, comprising:
[0029] Data source mining module, used to mine data sources;
[0030] Image acquisition module, used to collect images of dinosaur footprint fossils;
[0031] Image processing module, used to process dinosaur footprint fossil images;
[0032] Image annotation module, used to annotate images of dinosaur footprint fossils;
[0033] The validity verification module is used to verify the validity of image data and verify whether the image data is valid through a deep neural network.
[0034] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for constructing a dinosaur footprint image dataset based on a large language model.
[0035] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the method for constructing a dinosaur footprint image dataset based on a large language model.
[0036] Another object of the present invention is to provide an information data processing terminal, which is used to implement the dinosaur footprint image data set construction system based on the large language model. Beneficial effects
[0037] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0038] First, the present invention proposes a method for constructing a dinosaur footprint image dataset. Through image data processing technology, an image dataset that can be used to train image recognition algorithms is constructed, providing an image data foundation for subsequent automated recognition of dinosaur footprints.
[0039] The purpose of this invention is to use image data processing technology and the semantic analysis capabilities of large language models to mine the image data contained in existing dinosaur footprint theoretical research results, and form a method that can be used to automatically construct a dinosaur footprint image dataset, so as to better support the design and development of automatic recognition artificial intelligence algorithms in this field.
[0040] Second, the method of constructing a dinosaur footprint image dataset based on a large language model has indeed brought significant technological progress compared to traditional methods:
[0041] 1. Automated data collection: By combining Python crawlers with large language models, we can automatically filter out relevant papers on dinosaur footprints from a large amount of literature, greatly improving the efficiency and completeness of data collection.
[0042] 2. Intelligent image recognition: Through the analysis of large language models, it can intelligently determine whether there are footprint images in the paper, further screen and extract images, avoiding a lot of manual screening work.
[0043] 3. Automatic metadata extraction: Using large language models, relevant information of the paper, such as DOI, title, author, and references, can be automatically extracted, which facilitates subsequent data management and retrieval.
[0044] 4. Reduce human error: Traditional methods often require manual image screening and data entry, which are prone to errors. However, methods based on large language models can significantly reduce these errors and improve data accuracy.
[0045] 5. Scalability and adaptability: This method is not limited to the construction of dinosaur footprint image datasets. Its core ideas and processes can be applied to the construction of image datasets in other fields, and it has strong scalability and adaptability.
[0046] 6. Rapid response and update: When a new dinosaur footprint research paper is published, the dataset can be quickly updated through this method to ensure the timeliness and cutting-edge nature of the dataset.
[0047] 7. Promote interdisciplinary research: The implementation of this method will help connect multiple disciplines such as paleontology, computer science, and data science, and promote interdisciplinary research and innovation.
[0048] In summary, the method of constructing a dinosaur footprint image dataset based on a large language model has brought significant technological progress to related research fields, and helps to collect, process and analyze data more efficiently and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0050] FIG1 is a flow chart of a method for constructing a dinosaur footprint image dataset based on a large language model according to an embodiment of the present invention;
[0051] FIG2 is a structural diagram of a system for constructing a dinosaur footprint image dataset based on a large language model according to an embodiment of the present invention;
[0052] FIG3 is a flowchart of a technical solution provided by an embodiment of the present invention; wherein (a) is the overall process, and (b) is the image acquisition process;
[0053] FIG4 is a schematic diagram of removing artificial markings of dinosaur footprints according to an embodiment of the present invention; wherein (a) is a marked image and (b) is a cleaned image;
[0054] FIG5 is an original image and outline image of a dinosaur footprint provided by an embodiment of the present invention; wherein (a) an expert takes a photo, and (b) an expert identifies the footprint outline;
[0055] FIG6 is a schematic diagram of labeling dinosaur footprints using the labeling software Labelme provided in an embodiment of the present invention;
[0056] FIG7 is a diagram of the ResNet50 network structure provided by an embodiment of the present invention;
[0057] FIG8 is a diagram of a Yolox network structure provided in an embodiment of the present invention;
[0058] FIG9 is a flowchart of collecting dinosaur footprint images provided by an embodiment of the present invention. Modes for Carrying Out the Invention
[0059] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0060] The present invention provides a method for constructing a dinosaur footprint image dataset. The following are two specific embodiments:
[0061] Example 1: Construction of a dinosaur footprint image dataset for difficult documents
[0062] 1. Data source mining: For those dinosaur footprints with controversial or multiple interpretations, we specifically searched for papers containing but not limited to keywords such as “controversial dinosaur footprints” and “dinosaur footprint interpretation”.
[0063] 2. Collect footprint fossil images:
[0064] Using Python crawlers, we download documents from a few well-known paleontology and geology paper databases.
[0065] Use LLM tools to analyze the content of the paper and find footprint images related to controversy or multiple solutions.
[0066] 3. Processing footprint fossil images: Pay special attention to preserving the original markings, as these suggestions may be key to the dispute. At the same time, create a "disputed points" folder and save the parts related to the dispute as separate images.
[0067] 4. Annotate footprint fossil images: In addition to basic dinosaur footprint information, add annotations about controversies, such as different interpretations and which researchers support which interpretation.
[0068] Example 2: Construction of a dinosaur footprint image dataset for a specific geographical area
[0069] 1. Data source mining: Conduct specialized literature searches for specific geographic areas, such as “dinosaur footprints in Northeast China.”
[0070] 2. Collect footprint fossil images: Prioritize obtaining information from publicly published literature of paleontological and geological institutes or universities in specific areas.
[0071] The paper was analyzed using LLM tools to obtain images of dinosaur footprints unique to the region.
[0072] 3. Processing footprint fossil images: Special processing is performed to account for the region's unique geological features or coloring methods. For example, the color of the stone in certain areas may differ from the common footprint fossils, so the contrast or color balance may need to be adjusted to better display the footprints.
[0073] 4. Label footprint fossil images: In addition to basic dinosaur footprint information, the footprint's geographical origin, specific location of discovery, and related geological era should also be labeled.
[0074] The implementation solutions of these two embodiments are appropriately modified and adjusted based on the provided method to meet specific needs and situations.
[0075] In response to the problems existing in the prior art, the present invention provides a method and system for constructing a dinosaur footprint image dataset based on a large language model. The present invention is described in detail below with reference to the accompanying drawings.
[0076] As shown in FIG1 , the method for constructing a dinosaur footprint image dataset based on a large language model provided by an embodiment of the present invention includes:
[0077] S1. Data mining. The data source comes from image data included in research findings in this field, primarily published papers. This can be done by looking at papers published by experts in the field of dinosaur footprints, the references cited in these papers, and even the papers cited in these references, layer by layer. Alternatively, search for papers using keyword-based search engines (such as Baidu Scholar and DBLP).
[0078] S2. Collect images of dinosaur footprint fossils. The steps are as follows:
[0079] S201. Write a Python crawler program to first grab the download link of the PDF format paper from the web page and download the paper to the local computer.
[0080] S202. Use tools based on the Large Language Model (LLM), such as chatPDF (a tool based on the ChatGPT API), to analyze the paper content and preliminarily determine whether there are footprint images in the paper. If so, proceed to the next step; otherwise, find a new paper.
[0081] S203. Write a Python program to use the LLM analysis results to extract relevant information about the paper, including its DOI, title, authors, and references. Use the DOI to check whether the paper has been processed and registered in the historical records. If not, proceed to the next step; otherwise, ignore the paper.
[0082] S204. Write a Python program to read the papers one by one that have been confirmed to contain new image data after analysis by the large model LLM, create a folder with the same name as the paper, and a series of subfolders under the folder with the same name, including a Label folder, a Processed folder, a Non-label folder, and an Original folder, which are used to store images with different attributes.
[0083] S205. Write a Python program to automatically extract images from the paper and save them to the corresponding Original folder. The image name is "Paper Title_Image Extraction Number".
[0084] S206. Use a tool based on the general large language model (LLM) to analyze the content of the paper and determine whether the extracted image is a footprint image based on the paper's arguments and views on the image. If so, retain it; otherwise, delete it.
[0085] S3. Process dinosaur footprint fossil images. Typically, authors of dinosaur footprint images in papers will add some markings, such as white powder, to clearly indicate the new discovery. These artificial markings should be removed using image editing software or a webpage. Images that do not require processing should be placed in the Non-label folder. Images that require processing should be placed in the Label folder, and after processing, placed in the Processed folder.
[0086] S4. Label the dinosaur footprint fossil image. To meet more technical requirements, the labeling steps are as follows:
[0087] S401. Image binary classification. Use the Large Language Model (LLM) to read the paper. Based on the paper description, read the footprint images from the corresponding Label and Non-label folders. Determine whether the footprint images contain dinosaur footprints and copy them to the Positive folder (if dinosaur footprint images are present) and the Negative folder (if no dinosaur footprint images are present), respectively.
[0088] S402. Label the location information. Label the location information of the dinosaur footprints in the Positive folder using a polygonal or rectangular box in the Labelme labeling software. The location information of each image is saved in the "imagename.json" file.
[0089] S5. Verify the validity of the image data. Verify the validity of the image data through a deep neural network, specifically including:
[0090] S51. Using a deep neural network-based classifier, such as ResNet50, a prediction score greater than 0.5 indicates that the image contains dinosaur footprint fossils. An ACC (accuracy) index greater than 0.6 indicates that the dataset is valid for classification.
[0091] S52. Using a deep neural network-based object detector, such as Yolox, the predicted area with a prediction score greater than 0.5 and an IoU (between the predicted area and the annotated area) greater than 0.3 is the range of dinosaur footprint fossils correctly detected in the image. An ACC index greater than 0.6 is a valid dataset that can be used for object detection.
[0092] ;
[0093] ;
[0094] ;
[0095] True prediction Positive (positive sample) Negative (negative sample) PositiveTP (True Positive) FN (False Negative) NegativeFP (False Positive) TN (Ture Negative)
[0096] As shown in FIG2 , the system for constructing a dinosaur footprint image dataset based on a large language model provided by an embodiment of the present invention includes:
[0097] Data source mining module, used to mine data sources;
[0098] Image acquisition module, used to collect images of dinosaur footprint fossils;
[0099] Image processing module, used to process dinosaur footprint fossil images;
[0100] Image annotation module, used to annotate dinosaur footprint fossil images.
[0101] The validity verification module is used to verify the validity of image data and verify whether the image data is valid through a deep neural network.
[0102] Figure 3 is a flowchart of the technical solution, which records the steps of the construction method. Figure 3 (a) shows the overall process, and Figure 3 (b) shows the detailed process of acquiring images.
[0103] Figure 4. Schematic diagram of removing artificial markings from dinosaur footprints. Figure 4(a) shows the image marked by the authors, and Figure 4(b) shows the cleaned image.
[0104] Figure 5 is the original image and outline image of the dinosaur footprints. Figure 5 (a) records the photos taken by the experts, and Figure 5 (b) records the outline of the footprints identified by the experts, indicating that there are indeed dinosaur tracks here.
[0105] Figure 6 shows the dinosaur footprints annotated using Labelme software based on Figure 5. The figure records the location information of the dinosaur footprints. Compared with Figure 5, it can be seen that the location information is annotated based on the outline identified by experts.
[0106] Figure 7 is the ResNet50 network structure diagram.
[0107] Figure 8 is a diagram of the Yolox network structure.
[0108] An application embodiment of the present invention provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of a method for constructing a dinosaur footprint image dataset based on a large language model.
[0109] An application embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of a method for constructing a dinosaur footprint image dataset based on a large language model.
[0110] An application embodiment of the present invention provides an information data processing terminal, which is used to implement a dinosaur footprint image data set construction system based on a large language model.
[0111] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0112] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A method for constructing a dinosaur footprint image dataset based on a large language model, characterized in that: include: S1, mining data source: the data comes from the image data contained in the research results, which are mainly published papers; S2, collects images of dinosaur footprint fossils; S3, processing dinosaur footprint fossil images; S4, annotated dinosaur footprint fossil images; S5, verify the validity of the image data, and verify the validity of the image data through a deep neural network.
2. The method for constructing a dinosaur footprint image dataset based on a large language model as claimed in claim 1, characterized in that: The specific steps of S2 are as follows: S201, write a Python crawler program to first grab the download link of the PDF paper from the web page and download the paper to the local computer; S202: Use tools based on the general large language model (LLM), including chatPDF, to analyze the content of the paper and make a preliminary judgment on whether there is a footprint image in the paper. If so, proceed to the next step; otherwise, find a new paper. S203, write a Python program to use the LLM analysis results to extract relevant information of the paper, including its DOI, title, author and references; check whether the paper has been processed and registered in the historically saved information through DOI, if not, proceed to the next step, otherwise ignore the document; S204, write a Python program to read the papers that are confirmed to have new image data after analysis by the large model LLM one by one, create a folder with the same name as the paper, and a series of subfolders under the folder with the same name, including a Label folder, a Processed folder, a Non-label folder, and an Original folder, which are used to store images with different attributes; S205, write a Python program to automatically extract images from the paper and save the images to the corresponding Original folder. The image name is "paper title_image extraction sequence number"; S206, using a tool based on the general large language model (LLM) to analyze the content of the paper, and judging whether the extracted image is a footprint image based on the arguments and opinions of the paper on the image, if so, retain it, otherwise delete it.
3. The method for constructing a dinosaur footprint image dataset based on a large language model as claimed in claim 1, characterized in that: S3 specifically includes: For the dinosaur footprint images in the paper, the authors will make some marks in the images, including using white powder to indicate the outline in order to more clearly show the new discovery; at this time, you need to use photo editing software or web pages to erase these artificial marks, put the images that do not need to be processed in the Non-label folder, and put the images that need to be processed in the Label folder, and put them in the Processed folder after processing.
4. The method for constructing a dinosaur footprint image dataset based on a large language model as claimed in claim 1, characterized in that: In S4, the annotation steps are as follows: S401, image classification: Use the large language model LLM to read the paper, read the footprint image from the corresponding Label and Non-label folders according to the paper description, determine whether there are dinosaur footprints in the footprint image, and copy them to the Positive folder and Negative folder respectively; S402, marking position information: marking the position information of the dinosaur footprints in the Positive folder, using a polygonal box or a rectangular box of the labeling software Labelme to mark the information, and the position information of each picture is saved in the "picture name.json" file.
5. The method for constructing a dinosaur footprint image dataset based on a large language model as claimed in claim 1, characterized in that: S5 specifically includes: S51, using a deep neural network-based classifier, such as ResNet50, a prediction score greater than 0.5 is considered to indicate the presence of dinosaur footprints in the image, and an ACC index greater than 0.6 is considered to be a valid data set that can be used for classification; S52, using a target detector based on a deep neural network, including Yolox, the prediction area with a prediction score greater than 0.5 and an IoU greater than 0.3 is the range of dinosaur footprint fossils that are correctly detected in the image, and the ACC index greater than 0.6 is a valid data set that can be used for target detection, where ; ; 。 6. A system for constructing a dinosaur footprint image dataset based on a large language model using the method for constructing a dinosaur footprint image dataset based on a large language model as described in any one of claims 1 to 5, characterized in that: include: Data source mining module, used to mine data sources; Image acquisition module, used to collect images of dinosaur footprint fossils; Image processing module, used to process dinosaur footprint fossil images; Image annotation module, used to annotate images of dinosaur footprint fossils; The validity verification module is used to verify the validity of image data and verify whether the image data is valid through a deep neural network.
7. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for constructing a dinosaur footprint image dataset based on a large language model as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to execute the steps of the method for constructing a dinosaur footprint image dataset based on a large language model as described in any one of claims 1 to 5.
9. An information data processing terminal, used for implementing the dinosaur footprint image data set construction system based on a large language model as described in claim 6.
Citation Information
Patent Citations
Micro paleontology fossil image detection, classification and discovery method, system and application
CN113128335A
Single-sample and small-sample microbody paleontology fossil image identification method and system
CN114399763A
Mesomorphic fossil image processing system based on hierarchical structure matrix
CN114926678A
Dinosaur footprint image data set construction method and system based on large language model
CN117496307A
Method of Detecting at Least One Geological Constituent of a Rock Sample
US20230154208A1
Cited By
Reading identification method and device, network training method and device, equipment and medium
CN120708206A