Roof greening plant data set construction method based on large model and knowledge graph

By applying large models and knowledge graph technology in the field of roof greening in China, a diversified roof greening plant information data set was constructed, which solved the problem of China's lack of roof greening plant information data sets and improved research efficiency and application breadth.

CN120144774APending Publication Date: 2025-06-13BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510027708.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

China lacks a data set of roof greening plants information, which leads to difficulties in selecting plant planting combinations, affecting the implementation and research of roof greening projects.

Method used

Using a method based on big model and knowledge graph, the academic paper literature library is used as a data source, and the information of roof greening plants is extracted and processed using big model API calls and knowledge graph technology to build a diversified Chinese roof greening plant information data set.

Benefits of technology

It significantly improved the efficiency of researchers in obtaining plant information, filled the gap in the selection data set of roof greening plant planting combinations in China, and promoted in-depth research and wide application in the field of roof greening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144774A_ABST
    Figure CN120144774A_ABST
Patent Text Reader

Abstract

The invention discloses a roof greening plant data set construction method based on a large model and a knowledge graph, and the method comprises the steps: taking a paper in an academic paper literature library as a data source, taking a roof greening plant as a keyword, and carrying out the retrieval, so as to construct an initial paper data set; the method comprises the following steps of: extracting plant information and constructing a preliminary data set by using large model API (Application Program Interface) calling and a knowledge graph technology, proofreading and expanding the plant information by using a plant classification name online proofreading tool, and constructing an accurate diversified Chinese roof greening plant information data set by writing a verification algorithm; the method specifically comprises construction methods of a plant detailed information statistical table, a plant type combination table, a plant combination and city corresponding table, plant data in an original paper, a knowledge graph and other data types. According to the invention, the efficiency of acquiring plant information by researchers can be remarkably improved, and deep research and wide application in the roof greening field are powerfully promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to long text processing, large models, and knowledge graph construction, and specifically relates to a method for constructing a rooftop greening plant dataset based on large models and knowledge graphs. Background Art

[0002] With the acceleration of global climate change and urbanization, the pressure on urban ecosystems is increasing day by day. As an effective way to alleviate the urban heat island effect, improve air quality, and increase the urban green area, rooftop greening has received extensive attention in recent years. However, the implementation effect of rooftop greening depends to a large extent on the selection of plant species, and the adaptability of plants varies due to differences in regional climate, environmental conditions, and maintenance requirements. At the same time, there is a lack of rooftop greening plant information datasets at the national level in China. Therefore, developing suitable rooftop greening plant schemes for different cities has become a key issue.

[0003] For this reason, this patent proposes a method for constructing a rooftop greening plant dataset based on large model calls and knowledge graphs. This method uses the academic paper literature library as the data source, applies large model API calls and knowledge graph technology, and constructs an accurate and diversified rooftop greening plant information dataset for China through a verification algorithm. The dataset will cover plant data records in multiple cities in China, including the construction methods of data types such as plant detailed information statistical tables, plant species combination tables, plant combination and city correspondence tables, as well as plant data and knowledge graphs in the original papers. The dataset constructed by this method aims to provide a comprehensive and accurate reference for plant information for the implementation of rooftop greening projects in China, and help urban planning and ecological design. Summary of the Invention

[0004] The technical problem solved by the present invention is: to propose a method for constructing a rooftop greening plant dataset based on large models and knowledge graphs, use the method of large model API calls to extract long text data from papers, screen and process the obtained information to construct a diversified rooftop greening plant information dataset for multiple cities in China, make up for the gap in the selection dataset of rooftop greening plant planting combinations at the national level in China, significantly improve the efficiency of researchers in obtaining plant information, and strongly promote the in-depth research and wide application in the field of rooftop greening.

[0005] The technical solution of the present invention is as follows: A method for constructing a rooftop greening plant dataset based on a large model and a knowledge graph. First, take the papers in the academic paper literature database as the data source, and retrieve with the keyword "rooftop greening plants" to construct an initial paper dataset. Then, use the large model API call and knowledge graph technology to extract plant information and construct a preliminary dataset, and use an online proofreading tool for plant classification names to proofread and expand the plant information, so as to construct a diversified Chinese rooftop greening plant information dataset. The specific steps are as follows:

[0006] (1) Obtain the original paper data, store and preprocess the data.

[0007] In step (1), for the obtained original paper data, data reduction is performed in a unified format. Among them, the acquisition of paper data is to retrieve papers with the keyword "rooftop greening plants" in the CNKI literature database and save them in pdf format.

[0008] Furthermore, preprocess the original paper data to construct an original paper dataset. The specific steps include:

[0009] (a) Data import. Import the target paper data into the dedicated literature analysis software CiteSpace to deeply mine and analyze the academic literature, providing a data basis for subsequent multi-dimensional information parsing;

[0010] (b) Time dimension analysis. Through the time series analysis algorithm built into the CiteSpace software, accurately sort out the distribution of papers on the time axis and screen papers within a specified time span;

[0011] (c) Keyword dimension analysis step. Analyze the papers in the CiteSpace software with "Keyword" as the node. With the help of the text mining and semantic analysis functions of the software, extract and count the keyword information in the papers, and deeply analyze the appearance frequency, co-occurrence relationship and clustering situation of the keywords. Screen the papers that meet "rooftop greening plants";

[0012] (d) Author and institution dimension analysis step. Analyze the papers in the CiteSpace software with "Author" and "Institution" as the nodes. Use the author cooperation network construction and institution data processing and analysis modules of the software to count indicators such as the number of papers published and the number of citations by each author and institution in the papers in this field, and identify individual authors and institutions with important influence in this field to ensure the quality of the paper data.

[0013] (2) Identify and extract key information from the original paper dataset obtained in step (1) to construct an initial rooftop greening plant information dataset. The specific steps include:

[0014] (a) Kimi large model API call. Use OpenAI's API to interact with the large model. First, configure the corresponding API key and base URL to establish a connection with the specified large model "moonshot-v1-32k". During the call, construct appropriate input prompt information and send a request to the large model, which includes setting the role of the large model (set as a botany expert) and the text content to be processed (text extracted from the PDF). At the same time, set parameters such as temperature and maximum number of tokens to control the characteristics of the generated results. And handle the retry mechanism for possible rate limit errors. By setting the initial waiting time, maximum number of retries, and doubling the waiting time after each retry, ensure that the API call is completed as much as possible within a reasonable range, making the program have a certain degree of stability and fault tolerance;

[0015] (b) Specified information extraction. Use the pdfplumber library to open the PDF file and extract the text content page by page, and integrate it into a complete text string. Then, use the extracted text as input and pass it to the large model. The large model returns the processed text result based on the set requirements (extracting plant names, classifications, and research cities). Then, parse this result and split each line of text into three parts: plant name, classification, and research city according to specific delimiter rules. Finally, organize the extracted data into a dictionary form and store it in a list, and further convert it into the DataFrame format of pandas and save it as an Excel file for convenient subsequent data viewing and analysis;

[0016] (c) Initial dataset construction. Summarize and organize the extracted information in the format of paper title, paper URL, plant name, species, and city to construct an initial dataset of rooftop greening plant information and save it in the xlsx format.

[0017] (3) Perform diversified processing on the initial rooftop greening plant information dataset by integrating factors such as rooftop greening plant combinations, research cities, and detailed plant information. The specific steps are as follows:

[0018] (a) Construction of Python deduplication algorithm. Use the pandas library to implement deduplication of data related to plant names in Excel files. First, read the Excel file at the specified path, determine the column containing plant names, then add a new column to store the deleted duplicate data. Subsequently, loop through each row of data, split the plant name strings in the cells according to the set delimiter, collect unique plant names and duplicate plant names respectively by using the method of judging whether elements in the list are repeated, then recombine the processed unique plant names and duplicate plant names into strings to update the data in the corresponding columns, and finally save the processed data as a new Excel file to achieve the purpose of deduplicating plant name data and recording duplicate data.

[0019] (b) Use the Excel filtering function to filter the initial rooftop greening plant information dataset for three types of plant information: trees, shrubs, and ground cover plants. Use the deduplication algorithm in (a) above to deduplicate the plant names of the three types and save the unique data to construct a rooftop greening plant combination classification data table;

[0020] (c) Use the Excel filtering function to extract the research cities involved in the initial rooftop greening plant information dataset and classify their corresponding plant information into three types. Use the deduplication algorithm in (a) above to remove the plant names in different species of each city for deduplication operations to construct a city rooftop greening plant combination data table;

[0021] (d) Use the Excel merging function to summarize the plant names in the initial rooftop greening plant information dataset, and use the deduplication algorithm in (a) to remove all duplicate names;

[0022] (e) Import the deduplicated plant names in (d) into the online plant classification name proofreading tool to proofread the plant names and match their corresponding detailed information such as Latin names, families, genera, naming persons, and PPBC species IDs to construct a general overview table of Chinese rooftop greening plants.

[0023] (f) Summarize the data tables constructed in (b), (c), and (e) to construct a diversified rooftop greening plant dataset at the national level in China.

[0024] (4) Construct a knowledge graph of rooftop greening plant information by combining plant information and its relationships. The specific steps are as follows:

[0025] (a) Use the Excel filtering function to extract the plant names, family names, and genus names from the general overview table of Chinese rooftop greening plants and save them as three csv files respectively as entity files to be imported into the neo4j graph database;

[0026] (b) Construction of the entity file import function. First, use the LOAD CSV WITH HEADERS statement to load the data of the CSV file with headers from the specified local path, and refer to each row of data by the alias 'row' for convenience in subsequent operations. Then, use the MERGE statement to create (if it does not exist) or match (if it already exists) nodes with the specified names in the graph database according to the content corresponding to the 'name' field in each row of data in the CSV file, so as to import the plant-related data in the CSV file into the graph database and build the corresponding node structure;

[0027] (c) Use the genus names and family names corresponding to the plant names in the General List of Greening Plants on Chinese Roofs to extract the genus names corresponding to the plant names and the family names corresponding to the genus names, etc., and save them as two csv files to be imported into the neo4j graph database as relationship files;

[0028] (d) Construction of the relationship file import function. First, use the LOAD CSV WITH HEADERS statement to load the data of the CSV file with headers from the specified path, and refer to each row of data by 'row'. Then, use the MATCH statement to search for nodes with the specified names and the 'name' attribute values corresponding to those in the CSV file in the graph database respectively. Finally, use the MERGE statement to create (if it does not exist) or confirm (if it already exists) a relationship named BELONG_TO between the two corresponding nodes found, so as to realize the function of building an association relationship in the graph database based on the data in the CSV file.

[0029] (5) Verify the accuracy and usability of the constructed Chinese roof greening plant information dataset and knowledge graph. The specific steps are as follows:

[0030] (a) Construction of the python keyword matching algorithm. First, define the paths of two table files and read them into DataFrame objects respectively. Then, specify the column for matching in the first table and the de-duplicated keyword list extracted from the keyword column of the second table. Subsequently, use the apply function combined with conditional judgment to check whether each value in the specified matching column of the first table is in the keyword list. If the condition is met (non-empty and in the keyword list), fill in the word'matched' in the newly added marked column as a mark. Finally, save the processed data of the first table as a new Excel file to construct an algorithm for matching and marking results based on the data in specific columns of the two tables;

[0031] (b) Verify the algorithm in (a) above for the extracted plant name column and the 'List of Chinese Plant Species (2024 Edition)' downloaded from the Plant Science Data Center to ensure the accuracy of the plant names;

[0032] (c) Method for configuring the Neo4j database in PyCharm. The operations related to configuring the connection with the Neo4j database are implemented through the custom KnowledgeGraph class. In the initialization method of the class, the connection address, username, and password of the Neo4j database are received as parameters, and the GraphDatabase.driver method is used to establish a connection with the database based on this information. At the same time, a close method is defined to close the established database connection to ensure the reasonable release of database resources;

[0033] (d) Construction of the Cypher query function. Rely on the QuestionParser class to implement the information query function based on natural language questions. In the initialization stage of the class, multiple question templates and their corresponding processing functions are defined, and different types of natural language questions are matched in the form of regular expressions. The parse_question method traverses these question templates and tries to match the natural language question input by the user with regular expressions. If the match is successful, the corresponding processing function and the key entity information extracted from the question are returned. Subsequently, each processing function will convert it into a corresponding Cypher query statement based on the extracted entity, execute the query by means of the previously configured Neo4j database connection, and obtain the desired information. This module effectively combines the parsing of natural language questions with database queries to achieve a convenient information query mechanism;

[0034] (e) Compile the code in (c) and (d) through PyCharm, input the corresponding information according to the prompt statement, and compare and verify the output result with the corresponding data in the Chinese rooftop greening plant information dataset to verify the accuracy and usability of the knowledge graph;

[0035] (f) Summarize the data verified in (b) and (e) to construct the Chinese rooftop greening plant information dataset.

[0036] The advantages of the present invention compared with the prior art are as follows:

[0037] 1. Taking the academic paper literature database as the data source, combined with the CiteSpace software, deeply mining and analyzing the paper data from multiple dimensions such as time, keywords, authors, and institutions, constructing the original paper dataset, enriching the data foundation, and improving the data quality and comprehensiveness of the analysis. On this basis, various technical means such as Kimi large model API calls, knowledge graph technology, python deduplication algorithms, and Excel function operations are used to extract, sort, deduplicate, classify, and summarize the data, constructing a diversified rooftop greening plant dataset, which is faster, more accurate, and more comprehensive than the prior art data processing methods.

[0038] 2. Build functions for importing entity files and relationship files, and accurately import data related to roof greening plants into the Neo4j graph database to construct a knowledge graph. This not only realizes the effective construction of plant information and its relationships in the knowledge graph, but also achieves a convenient interactive query function with the knowledge graph through custom classes and question parsing classes. It can parse natural language questions into Cypher query statements to obtain information, and has a significant improvement in the integrity of knowledge graph construction and the convenience of query applications compared with the existing technology.

[0039] 3. By constructing a Python keyword matching algorithm, match the extracted plant names with the professional "List of Chinese Plant Species (2024 Edition)" to ensure the accuracy of plant names. At the same time, configure the Neo4j database and Cypher query function in PyCharm and combine them with the knowledge graph, which can verify the accuracy and usability of the constructed data set and knowledge graph based on natural language questions, and effectively improve the accuracy and reliability of data information through multiple verification processes. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is the overall flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0041] In order to enable those skilled in the art to better understand the solutions of the embodiments of the present invention, the embodiments of the present invention will be further described in detail below in conjunction with the drawings and embodiments.

[0042] As Figure 1 shown, the present invention includes the following steps:

[0043] 1. Obtain the original paper data, store and preprocess the data. For the obtained original paper data, perform data reduction in a unified format. Among them, the acquisition of paper data is to retrieve papers with the keyword "roof greening plants" in the CNKI literature database and save them in pdf format. Preprocess the original paper data to construct an original paper data set. The specific steps include: data import, time dimension analysis, keyword dimension analysis, author and institution dimension analysis.

[0044] Data import: Import the target paper data into the specialized literature analysis software CiteSpace for in-depth mining and analysis of academic literature;

[0045] Time dimension analysis: Through the time series analysis algorithm built into the software, accurately sort out the distribution of papers on the time axis and screen papers within a specified time span;

[0046] Keyword dimension analysis: Extract and count the keyword information in the papers through the text mining and semantic analysis functions of the software, and screen the papers that meet the criteria of "roof greening plants".

[0047] Author and institution dimension analysis: Through the author cooperation network construction and institution data processing and analysis modules of the software, count the relevant indicators of each author and institution in the papers in this field, identify the individual authors and institutions with important influence in this field, and ensure the quality of the paper data.

[0048] 2. Identify and extract key information from the original paper dataset obtained in step 1 above to construct an initial roof greening plant information dataset. The specific steps include: Kimi large model API call, specified information extraction, and initial dataset construction.

[0049] Kimi large model API call: First, configure the corresponding API key and base URL to establish a connection with the specified large model "moonshot-v1-32k". During the call, construct appropriate input prompt information, which includes setting the role of the large model and the text content to be processed. At the same time, set parameters such as temperature and maximum number of tokens to control the characteristics of the generated results. And handle the retry mechanism for possible rate limit errors by setting the initial waiting time, maximum number of retries, and doubling the waiting time after each retry to ensure that the API call is completed as much as possible within a reasonable range.

[0050] Specified information extraction: Use the pdfplumber library to open the PDF file and extract the text content page by page, and integrate it into a complete text string. Then, use the extracted text as input and pass it to the large model. The large model returns the processed text result based on the set requirements. Then, parse this result and split each line of text into three parts: plant name, classification, and research city according to specific delimiter rules. Finally, organize the extracted data into a dictionary form and store it in a list, and further convert it into the DataFrame format of pandas and save it as an Excel file.

[0051] Initial dataset construction: Summarize and organize the extracted information in the format of paper title, paper URL, plant name, species, and city to construct an initial roof greening plant information dataset, and save it in the xlsx format.

[0052] 3. Perform diversified processing on the initial roof greening plant information dataset by integrating factors such as roof greening plant combinations, research cities, and detailed plant information. The specific steps include: constructing a python deduplication algorithm, constructing a classification data table for cities and plant combinations, and constructing a general overview table of Chinese roof greening plants.

[0053] Construction of Python deduplication algorithm: Use the pandas library to achieve deduplication. First, read the Excel file, determine the column containing the plant names, split the plant name strings in the cells by the set delimiter through looping through each row of data, collect the unique plant names and duplicate plant names respectively by using the method of judging whether elements in the list are repeated, then recombine the processed unique plant names and duplicate plant names into strings to update the data in the corresponding columns, and finally save the processed data as a new Excel file.

[0054] Construction of the city and plant combination classification data table: Use the Excel filtering function to filter the initial rooftop greening plant information dataset for the plant information of three types: trees, shrubs, and ground cover plants respectively, extract the research cities involved and classify their corresponding plant information into three types, use the deduplication algorithm to deduplicate the two data tables respectively, and construct the rooftop greening plant combination classification data table and the city rooftop greening plant combination data table;

[0055] Construction of the general overview table of Chinese rooftop greening plants: Use the Excel merging function to summarize the plant names in the initial rooftop greening plant information dataset, and use the deduplication algorithm to remove all duplicate names. Import the deduplicated plant names into the online proofreading tool for plant classification names, proofread the plant names and match their corresponding detailed information such as Latin names, families, genera, naming persons, and PPBC species IDs to construct the general overview table of Chinese rooftop greening plants.

[0056] 4. Construct a knowledge graph of rooftop greening plant information by combining plant information and its relationships. The specific steps are as follows: Import entity files, import relationship files.

[0057] Import of entity files: Extract and save the plant names, family names, and genus names as three CSV files respectively, load them through the LOADCSV WITH HEADERS statement, and refer to each row of data by the alias row. Then use the MERGE statement to create or match the specified nodes in the graph database according to the content corresponding to the name field in each row of the CSV file, and construct the corresponding node structure.

[0058] Import of relationship files: Save the corresponding genus names of plant names and the corresponding family names of genus names as two CSV files and import them as relationship files into the neo4j graph database using the LOAD CSV WITH HEADERS statement, and refer to each row of data by row. Search for the specified nodes with the specified names and the name attribute values corresponding to those in the CSV file in the graph database respectively through the MATCH statement, and finally use the MERGE statement to create or confirm the relationship named BELONG_TO, so as to construct the function of the association relationship in the graph database.

[0059] 5. Verify the accuracy and usability of the diversified rooftop greening plant information dataset and knowledge graph to construct the Chinese rooftop greening plant information dataset. The specific steps are as follows: plant name verification, knowledge graph verification, and construction of the Chinese rooftop greening plant information dataset.

[0060] Plant name verification: First, define a DataFrame object. Through the apply function combined with conditional judgment, if the condition is met, fill in the word "matched" as a mark in the newly added marked column. Finally, save the processed data as a new Excel file to construct an algorithm for matching and marking results based on specific column data in two tables. Verify the extracted plant name column with the "List of Chinese Plant Species (2024 Edition)" downloaded from the Plant Science Data Center using the above algorithm to ensure the accuracy of plant names.

[0061] Knowledge graph verification: Receive the connection address, username, and password of the Neo4j database as parameters, and use the GraphDatabase.driver method to establish a connection with the database. Rely on the QuestionParser class to implement the information query function based on natural language questions. The parse_question method uses regular expressions to match the natural language questions input by users, converts them into corresponding Cypher query statements based on the extracted entities, and executes the query through the Neo4j database connection to obtain the desired information; compile through PyCharm, input the corresponding information according to the prompt statements, and compare and check the output results with the corresponding data in the Chinese rooftop greening plant information dataset to achieve the verification of the knowledge graph.

[0062] Construction of the Chinese rooftop greening plant information dataset: Summarize the above-verified plant information and knowledge graph to form a diversified Chinese rooftop greening plant information dataset.

Claims

1. A method for constructing a roof greening plant dataset based on a large model and knowledge graph, characterized in that: The implementation steps of this method are as follows: Step (1) Using the academic paper library as the data source, searching with the keyword "roof green plants", selecting all papers that meet the requirements and saving them in PDF format, and constructing the initial paper data set through literature analysis software; Step (2) Write API call code based on the big model, use its ability to understand long texts, accurately understand and extract the plant information data required by each article and organize it according to unified standards; Step (3) using Python data analysis to write standard deduplication algorithms for specific types of data tables to remove duplicate plant names, using an online proofreading tool for plant classification names to verify the accuracy of the extracted plant information and expand plant detailed information, thus constructing a diversified plant information dataset; Step (4) The relationships between plant species and their corresponding families and genera are sorted into entity files and relationship files, and imported into the neo4j graph database to construct a knowledge graph; Step (5) verifies the extracted data through a keyword matching algorithm to ensure the accuracy of the plant names; uses the Cypher query function to verify the availability of the constructed knowledge graph, and finally constructs a roof greening plant information dataset.

2. The method for constructing a roof greening plant dataset based on a large model and a knowledge graph according to claim 1 is characterized in that: In step (1), the creation of the original paper dataset includes: (1) Search the CNKI document library with the keyword "roof greening plants" and select the papers that meet the requirements and save them in PDF format; (2) The papers were imported into the literature analysis software CiteSpace, and the obtained papers were analyzed and screened using "Keyword", "Author", and "Instution" as nodes to ensure the timeliness and reliability of the construction of the original paper dataset.

3. The method for constructing a roof greening plant dataset based on a large model and a knowledge graph according to claim 1 is characterized in that: In the step (2), extracting information about roof greening plants in the paper includes: (1) Apply for the API key of Kimi's large model and build an API call algorithm model; (2) Automatically extract the required plant information from the original paper data set based on the algorithm model; (3) The extracted information is saved according to the plant name and type and the corresponding paper source and research city to construct the initial roof greening plant information dataset; (4) Save the initial roof greening plant information dataset in xlsx format and set a link address for each paper to facilitate access to the original paper.

4. The method for constructing a roof greening plant dataset based on a large model and a knowledge graph according to claim 1 is characterized in that: In step (3), constructing a diversified green roof plant dataset includes: (1) Based on the initial green roof plant information dataset, the plant names were classified into three types: trees, shrubs, and ground cover plants. The Python deduplication algorithm was used to remove duplicate plant names to retain unique data and construct a green roof plant combination classification data table; (2) Based on the initial green roof plant information dataset, the plant names corresponding to the research city were selected and corresponded to the three plant types. The Python deduplication algorithm was used to perform deduplication operations to retain unique data to construct the urban green roof plant combination data table; (3) Based on the initial green roof plant information dataset, only all plant names were screened and all duplicate data were removed using the Python deduplication algorithm. The accuracy of the plant information was verified using the online proofreading tool for plant classification names and the corresponding Latin name, family, genus, namer, and PPBC species ID detailed information of the plant were expanded to construct an overview table of green roof plants. (4) After condition screening, the statistical tables are summarized to construct a diversified roof greening plant data set.

5. The method for constructing a roof greening plant dataset based on a large model and a knowledge graph according to claim 1 is characterized in that: In the step (4), the construction of the knowledge graph of roof greening plant information includes: (1) Save the plant name, family name, and genus name as three entity files in csv format and use the "LOAD CSV WITH HEADERS" function of neo4j to import the three entity files; (2) The relationship between the plant name and the genus name, and the relationship between the genus name and the family name are saved as two relationship files in csv files, and the two relationship files are imported using the "LOAD CSV WITH HEADERS" function of neo4j; (3) Use the "MATCH" function of neo4j to match entities and relationships and build a knowledge graph of roof greening plant information.

6. The method for constructing a roof greening plant dataset based on a large model and a knowledge graph according to claim 1 is characterized in that: In the step (5), the verification of the roof greening plant information dataset includes: (1) The extracted plant names were matched with the Plant Science Data Center through the constructed Python keyword matching algorithm to ensure the accuracy of the plant names; (2) Define multiple cypher query functions through pycharm, connect to the neo4j database and enter the corresponding account and password, run the program and enter the data you want to query according to the program prompts, and verify the accuracy of the output data to ensure the availability of the knowledge graph; (3) The above verified data are aggregated to construct a roof greening plant information dataset.