Ecological parameter of large language model and automatic extraction method of driving factor thereof

By designing a multi-level prompt word system and image interaction tools, and utilizing a large language model to automatically extract ecological parameters and their driving factors from ecological literature, the problem of low efficiency in ecological literature data processing is solved. This achieves efficient and accurate data extraction and standardized storage, and supports cross-research and cross-regional data comparison and analysis.

CN120910562BActive Publication Date: 2026-05-01RES INST OF FOREST RESOURCE INFORMATION TECHN CHINESE ACADEMY OF FORESTRY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RES INST OF FOREST RESOURCE INFORMATION TECHN CHINESE ACADEMY OF FORESTRY
Filing Date
2025-07-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and flexibly extract complex ecological parameters and their driving factors from ecological literature, resulting in low data processing efficiency and high costs.

Method used

We designed a multi-level prompting word system based on ecological laws using a large language model. Combining text and table parsing techniques, and through block prompting words and image interaction tools, we achieved automatic extraction and standardized processing of ecological parameters, forming a structured ecological database.

Benefits of technology

It improves the efficiency and quality of ecological literature data analysis, ensures the accuracy and consistency of extraction results, breaks down the barriers of scattered data storage and inconsistent formats, and promotes big data mining and data sharing in the field of ecology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910562B_ABST
    Figure CN120910562B_ABST
Patent Text Reader

Abstract

An ecological parameter and driving factor automatic extraction method of a large language model belongs to the technical field of artificial intelligence. The ecological literature PDF is obtained, the text blocking prompt word is constructed, the large language model is used to analyze the ecological literature, and the segmented and reorganized design based on the multi-level prompt word system of the ecological law is used to drive the large language model to extract the ecological parameters and driving factors in the ecological parameter structure, the ecological parameter fusion and correction are carried out based on the table data, the time, space and unit standards are unified, the semi-automatic extraction and manual calibration of the literature picture observation data are realized through the picture interactive extraction, the field alignment prompt word is designed to guide the large language model to automatically standardize the extraction result field, the database is imported, and the ecological standardized database is formed.
Need to check novelty before this filing date? Find Prior Art

Description

An Automatic Extraction Method for Ecological Parameters and Driving Factors of a Large Language Model Technical Field

[0001] This invention belongs to the field of artificial intelligence technology and relates to an automatic extraction method for ecological parameters and driving factors of a large language model. Background Technology

[0002] A vast body of ecological literature provides crucial empirical support for meta-analysis, ecological big data mining, and large-scale ecological research. However, this data is scattered across the text and figures of millions of academic research papers, severely hindering its effective integration. Although existing studies have compiled experimental datasets through manual extraction, the exponential growth in the number of ecological documents makes manual extraction inefficient and difficult to scale, resulting in high costs for data reuse.

[0003] Existing text mining technologies based on rules, traditional natural language processing, machine learning, and deep learning have made progress in areas such as automatically labeling taxonomic texts and identifying species names. However, due to their reliance on predefined rules or training models for specific tasks, they lack flexibility and adaptability, making it difficult to accurately understand the complex syntax and specialized semantics in ecological literature, and also difficult to efficiently handle data extraction tasks unique to the ecological field.

[0004] In recent years, Large Language Models (LLMs) have demonstrated outstanding performance in literature mining in fields such as materials science and medicine due to their powerful semantic understanding and logical reasoning capabilities. They can quickly extract key information and significantly improve data acquisition efficiency. In ecology, LLMs have also been initially used to extract simple information such as species names and geographical coordinates. However, existing methods still lack autonomy and flexibility in the efficient extraction and standardization of fine-grained quantitative data, such as experimental design, environmental parameters, and dynamic processes. Therefore, how to fully utilize the capabilities of LLMs to achieve efficient, structured, and standardized extraction of complex ecological parameters and their driving factors from ecological literature is a pressing issue that needs to be addressed. Summary of the Invention

[0005] This invention addresses the problems of existing technologies by providing an automatic extraction method for ecological parameters and driving factors of large language models.

[0006] An automatic extraction method for ecological parameters and their driving factors from a large language model includes the following steps: acquiring ecological literature PDFs, constructing text segmentation prompts, parsing the ecological literature using a large language model, segmenting and reorganizing the texts, designing a multi-level prompt system based on ecological principles, driving the large language model to extract ecological parameters and their driving factors from the ecological parameter structure, performing ecological parameter fusion and correction based on tabular data, unifying time, space, and unit standards, achieving semi-automatic extraction and manual calibration of literature image observation data through interactive image extraction, designing field alignment prompts to guide the large language model to automatically standardize the extracted result fields, and importing the data into a database to form a standardized ecological database.

[0007] The advantages of this invention are: it achieves accurate extraction and standardized processing of ecological literature data, improving data processing efficiency and quality, and has at least the following technical effects or advantages:

[0008] By designing segmented prompts that align with the characteristics of ecological literature, the large language model is guided to quickly identify core content such as "description of study area, sample plot survey methods, results analysis, and discussion conclusions" in the literature. It automatically completes segmentation and recombination, replacing the inefficient process of relying on manual page-by-page reading, paragraph-by-paragraph marking, and manual extraction. It intelligently filters out effective ecological information and redundant content, ensuring the accuracy of subsequent extraction from the data source and significantly improving the efficiency of literature analysis.

[0009] We designed a multi-level cue word system based on ecological principles to drive a large language model to deeply understand the intrinsic relationship between ecological parameters and driving factors in the literature, accurately extract the specific values ​​of ecological parameters and their driving factors, and automatically complete time alignment, spatial matching, and unit standardization conversion, reducing the error of manual post-processing and laying the foundation for cross-study and cross-regional ecological data comparison.

[0010] For unstructured image data such as line graphs and bar charts in literature, we developed image interaction tools and automated extraction technology to achieve a semi-automatic extraction and manual calibration closed loop of observation data from literature images, balancing data extraction efficiency and verifiability, and solving the problems of error-prone and time-consuming traditional manual image reading.

[0011] This study constructs field alignment prompts to guide a large language model in standardizing the extracted results and automatically imports them into a pre-defined ecological database, forming a standardized database with a reasonable structure, unified fields, and direct accessibility. This database supports multi-dimensional retrieval and cross-document comparative analysis, breaking down the barriers of "dispersed data storage, inconsistent formats, and difficulty in data reuse" in the field of ecology. It promotes big data mining and data sharing in the ecological field, providing high-quality data support for ecological theory innovation and ecological protection policy formulation. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. As shown in the figures:

[0013] Figure 1 is a flowchart of one aspect of the present invention.

[0014] Figure 2 is a flowchart of the second part of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Example 1: As shown in Figures 1 and 2, an automatic extraction method for ecological parameters and driving factors of a large language model is presented, which relates to the application of large language models in the construction of ecological literature datasets.

[0017] Large language models are artificial intelligence models trained on large-scale text data, possessing deep semantic understanding and knowledge reasoning capabilities. They can parse instructions and execute tasks based on prompts. Addressing the low efficiency of manual extraction from ecological literature and the insufficient flexibility of large language models, this invention proposes an automatic extraction method for ecological parameters and their driving factors from large language models. By designing an intelligent workflow tailored to the ecological field, combined with a multi-level prompt system and a human interaction mechanism, this method achieves efficient and accurate extraction of ecological parameters and their driving factors from ecological literature, providing powerful tools and data support for large-scale ecological research.

[0018] An automatic extraction method for ecological parameters and driving factors of a large language model, comprising the following steps:

[0019] Step 1: Obtaining Ecological Literature PDFs. Based on the pre-set ecological research goals and parameter requirements, design search keywords and search and download relevant PDF literature from literature databases such as CNKI as the basis for subsequent analysis and extraction.

[0020] Step 2, the step of parsing ecological literature, uses text parsing technology to convert PDFs into Markdown format text with heading symbols. The text is segmented according to preset heading symbols, and a large language model is constructed with the first prompt word input to guide it to reorganize the text according to the paper's chapter titles, outputting the titles of paragraphs such as research area overview, research methods, results and analysis, and discussion. The corresponding content is automatically retrieved through association to generate the main text Markdown text. Tables and images are extracted by recognizing keywords and regular expressions, and saved as table Markdown and image Markdown respectively. Table information includes table title and table content, and image information includes image title and image path. The reorganized literature content can help the large language model accurately understand the logic of each part and improve the targeting of data extraction.

[0021] Step 3: Extracting the ecological parameter structure. Based on the research objectives and ecological parameter characteristics, a second prompt word following ecological principles is designed. The main text Markdown and the second prompt word are input into the large language model together to enable it to deeply understand the main text content and accurately extract the specific observation values ​​of the target ecological parameters and their corresponding vegetation factors (such as vegetation type, tree height, diameter at breast height, and forest age), climate factors (such as temperature, precipitation, and wind speed), soil factors (such as soil pH, soil bulk density, and soil nitrogen content), and topographic factors (such as altitude, slope, and aspect). Time, space, and unit standards are unified to ensure consistent matching. The extraction results are formed into Markdown tables with spatiotemporal matching, such as metadata tables, full data tables, and ecological parameter time series observation tables, and are automatically saved as Excel files.

[0022] Step 4, Ecological Parameter Fusion and Correction: Table titles are obtained through rule-based filtering. A third prompt word is designed to guide the large language model to output a list of table numbers containing the target data. The corresponding table content is associated with the table Markdown text, forming a new paragraph composed of multiple tables. Based on the text extraction results obtained in Step 3 and the newly formed table information, a fourth prompt word is designed to drive the large language model to fuse table information, fill in missing data, enrich ecological field dimensions, and correct the extraction results again to obtain more accurate and comprehensive structured data. After verification, each driving factor is accurately classified into major categories such as vegetation, climate, soil, and topography. Finally, new metadata tables, vegetation / climate / soil / topography factor tables, and ecological parameter time-series observation tables with spatiotemporal matching are generated and automatically saved as Excel spreadsheets.

[0023] Step 5: Interactive image extraction. Image titles are obtained through rule-based filtering. A fifth prompt word is designed to guide the large language model to output a list of image IDs containing the target data. Image paths are associated with the corresponding image Markdown text. An interactive extraction tool is invoked according to user needs. The horizontal and vertical axis scale values ​​are set in the automatically pop-up image interface. Users click on the locations where data needs to be extracted, and the actual values ​​are calculated based on the image scale and pixel ratio. After data extraction, a storage tool is invoked. Users set the titles for the horizontal and vertical axes, and the data is automatically converted into column names and row values ​​upon saving, forming an Excel spreadsheet.

[0024] Step 6: Importing the database. After manually verifying the various Excel spreadsheets obtained in Steps 3, 4, and 5, create the corresponding MySQL database tables. Input the designed sixth prompt word and the Excel header information into the large language model, driving it to compare the fields of the Excel to be stored with the existing fields in the database. Set synonymous and similar fields as the same standard fields, add fields that do not exist in the database, automatically associate the corresponding data and import them into the database, forming a standardized ecological database that can be directly compared and analyzed.

[0025] Example 2: As shown in Figures 1 and 2, an automatic extraction method for ecological parameters and driving factors of a large language model includes the following steps:

[0026] Step 1: Based on the ecological research objectives (such as extracting litter volume and its driving factors), use thematic search to retrieve relevant literature (such as litter volume), conduct literature searches on authoritative database platforms such as CNKI, select literature that meets the research needs based on the literature title and abstract, and download them as PDF format.

[0027] Step 2: Use MinerU to parse the PDF text, converting it into Markdown formatted text marked with "#". Using "#" as the dividing line, extract first-level headings such as "Overview of Study Area", "Research Methods", "Results and Analysis", and "Discussion", along with their subheadings (e.g., "2.1 Sample Plot Setup"), forming a heading list {headers}. Design a first prompt word to drive the DeepSeek-V3 large language model to output body text titles containing key information, automatically associating the titles with their corresponding text and reorganizing them into body text Markdown. Extract tables using regular expressions and save them as table Markdown (including headings and cell content). Identify image titles and paths and save them as image Markdown. The following is a detailed explanation of the first prompt word in a specific scenario:

[0028] The following is a list of headings for a Markdown document: {headers}. Please select headings that relate to the overview of the study area (e.g., latitude and longitude, location name), the observation methods and results analysis of total litter volume, and the vegetation (type, age), climate (temperature, precipitation), soil (pH, total nitrogen), and topographic (altitude, slope) factors affecting litter volume. Requirements: 1. Only retain headings directly related to the above content, excluding irrelevant content such as litter decomposition, nutrient content, and current stock; 2. The returned headings must be completely identical to the original headings, and name modifications are prohibited; 3. The output format should be a list, such as: ["1.1 Overview of the Study Area", "3.2 Litter Volume Dynamics"].

[0029] Step 3: For the extraction target of "litter quantity and its driving factors", design a second prompt word, input it along with the main text Markdown into the large language model, output the corresponding table and text according to the extraction requirements, and automatically save it as an Excel file. The following is a detailed explanation of the second prompt word in a specific scenario:

[0030] As an ecological data extraction expert, please extract the total litter volume and its driving factors data from the following text and output three tables (Markdown format): I. Metadata Table (Fixed Background of Study Area): Fields include: Location, Survey Time (Start and End Years), Longitude, Latitude, Center Point Coordinates (Decimal), Climate Type, Overall Vegetation Type of Study Area (listing all types of litter volume measured), Average Annual Temperature (°C), Annual Precipitation (mm), Soil Type, and Geomorphological Features. Fields can be added or deleted, but values ​​must be consistent and fixed information throughout the entire literature. II. AL_DATA Table (Dynamic Data at Plot Scale): Table headers include: Year, Plot Number, Plot Description, Elevation (m), Slope (°), Aspect, and Litter Yield (kg / hm²). 2 ), vegetation type, tree height (m), diameter at breast height (cm), forest age (years), soil pH, total soil nitrogen (%), average annual temperature (°C), and annual precipitation (mm). Requirements: 1. Only retain rows containing total litter volume, each row corresponding to a unique time and space (year + plot); 2. Driver factors should be listed separately, missing values ​​should be left blank; 3. Units should be strictly consistent (e.g., "m" "°C"). III. Monthly Data Validation Table (if monthly data is available): Table header: Year, Month, Plot Number, Vegetation Type, Monthly Litter Volume (kg / hm) 2 If no data is available, no output will be provided. Note: Fabricating data is prohibited; time and sample plots must strictly correspond; leave missing values ​​blank.

[0031] Step 4: Extract table titles from the Markdown table to form a list {table_titles}. Design a third prompt word input large language model to obtain the table numbers containing the target data. The following is a detailed explanation of the third prompt word in a specific scenario:

[0032] The following are table titles from the literature: {table_titles}. Please filter out table titles that contain data on soil, vegetation, topography, and total litter volume, and output a list of matching table titles (e.g., ["Table 1", "Table 3", etc.]), excluding tables that only contain litter components.

[0033] Step 4.1: After obtaining the table numbers related to the research topic, filter out the table information corresponding to the title from the Markdown containing table information (e.g., "Table 3 Soil Physicochemical Properties"). Input this table information, along with the full data table generated in Step 3 and the designed fourth prompt word, into the large language model. Output the supplemented and improved table data and automatically save it as a corresponding Excel file. The following is a detailed explanation of the fourth prompt word in a specific scenario:

[0034] There are two sets of data: ① the all_data table (information extracted from the main text); ② tabular data (from tables in the literature). Please merge them according to the following rules: 1. Use the all_data table as the basis and fill in the missing values ​​with the tabular data (e.g., soil pH in plot A is missing in the main text, fill it with data from Table 3); 2. If there are data conflicts (e.g., the amount of litter in the main text and the table are inconsistent), the tabular data shall prevail;

[0035] 3. Split the driving factors into four tables: Vegetation Factors (plot number, vegetation type, forest age, tree height); Topographic Factors (plot number, altitude, slope); Soil Factors (plot number, soil pH, total nitrogen content); and Climate Factors (year, annual precipitation, annual average temperature). 4. Output the merged all_data table, metadata table, and the four driving factor tables (Markdown format).

[0036] Step 5: Extract image titles from the image Markdown to form a list {img_titles}. Input this list along with the designed fifth prompt word into the large language model to obtain the image numbers from which data needs to be extracted. The following is a detailed explanation of the fifth prompt word in a specific scenario:

[0037] The following are image titles: {img_titles}. Please filter all titles related to the study area, plot overview, and litter volume (such as yearly and monthly dynamics), as well as environmental variables affecting litter volume such as temperature, precipitation, vegetation, topography, and soil. Output a list of numbers that meet the criteria (such as ["Figure No. 1", "Figure No. 4", etc.]), excluding tables that only contain litter components.

[0038] Step 5.1: Based on the Markdown text of the image information, associate the image path corresponding to the image number, and the image will automatically pop up. The user can decide whether to extract the image and whether to extract the same image multiple times. If the image is extracted, an interactive extraction interface will pop up. The user can manually set the horizontal and vertical coordinates and scale values, click on the location where data needs to be extracted, and calculate the actual value based on the image scale and pixel ratio. By setting the titles of the horizontal and vertical coordinates, the structured data is formed and saved as an Excel file.

[0039] Step 6: After manually verifying all Excel spreadsheets, create a target table in MySQL, construct the sixth prompt word, and input it along with existing fields in the database and fields from the Excel spreadsheets into the large language model. Drive the model to perform comparisons, setting synonymous and similar fields to the same field, and aligning literature from different sources for unification and standardization (e.g., storing all instances of "altitude" in the "altitude_m" field). For fields not present in the database, convert them to appropriate names using the large language model and add them as new fields. Finally, automatically associate fields with corresponding data and store them in the database, forming a standardized database that can be continuously updated and directly compared and analyzed. The following is a detailed explanation of the sixth prompt word in a specific scenario:

[0040] Task: Map Excel headers to database fields in the table {table_name}, following these rules: 1. Prioritize predefined mappings: {mapping_text}; 2. If a field with the same meaning already exists in the database (refer to the existing field list), use the exact same field name; 3. If there is no predefined mapping and no matching field, translate the English directly, retaining the units, and connect them with underscores, e.g., convert "soil organic matter (0-20cm)" to "soil_organic_matter_0-20cm"; 4. Record all Excel headers, merge identical fields, and standardize them.

[0041] Excel header: {excel_headers}

[0042] The {table_name} table contains the following field: {existing_columns_text}

[0043] The output must be a dictionary: {{"database field name":"original Excel header"}}.

[0044] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for automatically extracting ecological parameters and driving factors of a large language model, characterized in that, The process includes the following steps: obtaining ecological literature PDFs; constructing text segmentation prompts; parsing ecological literature using a large language model, segmenting and reorganizing the text; designing a multi-level prompt system based on ecological principles to drive the large language model to extract ecological parameters and their driving factors from the ecological parameter structure; performing ecological parameter fusion and correction based on tabular data; standardizing time, space, and units; achieving semi-automatic extraction and manual calibration of literature image observation data through interactive image extraction; designing field alignment prompts to guide the large language model to automatically standardize the extracted result fields; importing the data into a database to form a standardized ecological database. The process of obtaining ecological literature PDFs includes the following steps: obtaining... The ecological literature analysis process, based on pre-defined ecological research objectives and parameter requirements, involves designing search keywords and downloading relevant PDF documents from the CNKI (China National Knowledge Infrastructure) database as the foundation for subsequent analysis and extraction. The process includes the following steps: converting the PDF into Markdown format text with heading symbols using text parsing technology; segmenting the text according to pre-defined heading symbols; constructing a large language model with initial prompts to guide the text reorganization according to the paper's chapter titles; outputting titles related to the study area overview, research methods, results and analysis, and discussion paragraphs; and automatically retrieving corresponding content to generate the main Markdown text. The tables and images are extracted using keyword recognition and regular expressions, and saved as table Markdown and image Markdown respectively. Table information includes table title and content, while image information includes image title and path. The ecological parameter extraction process involves the following steps: Based on the research objectives and ecological parameter characteristics, a second cue word following ecological principles is designed. The main text Markdown and the second cue word are input into a large language model to enable it to deeply understand the main text content and accurately extract the specific observation values ​​of the target ecological parameters and their corresponding driving factor data. The driving factor data includes vegetation factors, climate factors, soil factors, and topographic factors. The data includes vegetation type, tree height, diameter at breast height (DBH), and forest age; climate factors include temperature, precipitation, and wind speed; soil factors include soil pH, soil bulk density, and soil nitrogen content; and topographic factors include altitude, slope, and aspect. Time, space, and unit standards are standardized to ensure consistent matching. The extracted results form a metadata table, a full data table, and a time-series observation table of ecological parameters, which are automatically saved as Excel files. The ecological parameter fusion and correction steps include the following: obtaining table titles through rule filtering; designing third-party prompts; guiding the large language model to output a list of table numbers containing the target data; associating the corresponding table content with the table Markdown text; and forming a new paragraph composed of multiple tables.Based on the text extraction results and newly formed table information obtained in the ecological parameter structure extraction step, a fourth prompt word is designed to drive the large language model to integrate table information, fill in missing data, enrich ecological field dimensions, and further correct the extraction results to obtain more accurate and comprehensive structured data. After verification, each driving factor is accurately classified into vegetation, climate, soil, and topography categories, ultimately generating a new metadata table, a vegetation / climate / soil / topography factor table, and an ecological parameter time series observation table, which are automatically saved as an Excel spreadsheet. The interactive image extraction step includes the following steps: obtaining image titles through rule filtering; designing a fifth prompt word to guide the large language model to output a list of image numbers containing target data; associating the corresponding image paths based on the image Markdown text; and calling the interactive extraction tool according to user needs, in an automatically popping-up window. In the image interface, the scale values ​​of the horizontal and vertical axes are set. Clicking on the location where data needs to be extracted, and converting it to the actual value based on the image scale and pixel ratio, allows for data extraction. After data extraction, a storage tool is called. The user sets the titles for the horizontal and vertical axes, which are automatically converted into column names and row values ​​upon saving, forming an Excel spreadsheet. The database import process includes the following steps: After manually verifying the various Excel spreadsheets obtained from the ecological parameter structure extraction, ecological parameter fusion and correction, and interactive image extraction steps, corresponding MySQL database tables are created. The designed sixth prompt word and Excel header information are input into the large language model, driving it to compare the fields of the Excel to be stored with existing fields in the database. Synonymous and similar fields are set to the same standard field. Fields not existing in the database are added, and corresponding data is automatically associated and imported into the database, forming a standardized ecological database that can be directly compared and analyzed.

Citation Information

Patent Citations

  • Long text generation method, device and equipment based on large language model

    CN118536502A

  • Document question and answer method and system based on multi-modal primitives, terminal and medium

    CN119917686A