Natural resource planning text generation method and system and storage medium

Through large-scale language model parsing and spatial clustering technology, natural resource planning text is generated, which solves the problems of low data parsing efficiency and lack of spatial correlation, and realizes efficient and compliant planning text generation.

CN120706400AActive Publication Date: 2025-09-26ZHEJIANG WANWEI SPACE INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511201782.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-09-26
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing technologies for natural resource planning text generation suffer from low data parsing efficiency and lack of cross-modal spatial correlation, which leads to logical contradictions in the generated content space.

Method used

By receiving standardized data, using a large language model to parse optional items, constructing a feature space covering the planning area, performing grid partitioning and spatial clustering, generating a dynamic material pool containing spatial correlation attributes, combining the feature weights of the partition units to generate the main text paragraphs, and performing structured processing to output the planning text.

Benefits of technology

It improves the efficiency of pre-data processing for planning text generation, ensures that the content conforms to the spatial distribution patterns of natural resource elements, avoids spatial logical contradictions, and improves the timeliness and compliance of the text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706400A_ABST
    Figure CN120706400A_ABST
Patent Text Reader

Abstract

The invention provides a natural resource planning text generation method and system and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: building a grid partition unit in a feature space, and executing spatial clustering according to the feature vector distribution density; mapping the feature vectors to corresponding partition units, and generating a dynamic material pool containing spatial correlation attributes; generating a spatial reference outline of an outline structure based on the spatial association attributes; according to the spatial reference outline, positioning the content of the material pool by adopting a segmented writing strategy, and generating text paragraphs based on the feature weights of the partition units; performing structured processing on the text paragraph to obtain a text after structured processing; and executing dynamic verification on the text after the structured processing so as to output a final planning text. According to the spatial clustering mapping, the processing amount of redundant data is reduced, and computing resources are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method, system and storage medium for generating natural resource planning text. Background Art

[0002] In the field of natural resource management, the compilation of planning documents (such as land use master plans and ecological protection and restoration plans) is an important vehicle for policy implementation. Traditional document generation methods sometimes have the following problems: For example, planning needs to integrate administrative division data, policies and regulations, geographic spatial information, and socioeconomic statistical indicators; existing methods rely on manual extraction of unstructured data (such as PDF policy documents, scanned maps, and Excel spreadsheets), resulting in inefficient data parsing and a lack of cross-modal spatial correlation (such as the inability to automatically associate the jurisdictional relationship between a policy clause and a specific geographic coordinate).

[0003] For example, the text structure needs to reflect the spatial distribution patterns of natural resource elements (for example, the chapter on watershed ecological restoration needs to be linked to upstream and downstream grid units). The mainstream text generation model only relies on semantic features and has not established a mapping mechanism between chapter structure and spatial partitioning units. The generated content often has spatial logical contradictions (such as planning the construction of water conservation forests in arid areas). Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method, system and storage medium for generating natural resource planning text, wherein spatial cluster mapping reduces the processing amount of redundant data and reduces computing resources.

[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows: In a first aspect, a method for generating a natural resource planning document is provided, the method comprising: Step 1: receiving standardized data input by a user, wherein the standardized data includes mandatory items and optional items, wherein the mandatory items include title, region, and writing subject, and the optional items include outline templates and unstructured reference materials; Step 2: Execute resource scheduling decisions based on whether the optional items exist. If the optional items exist, a large language model is used to parse reference materials and extract keywords. If the optional items do not exist, local policy compliance analysis and platform database matching are triggered. Step 3: Convert the keyword set, user knowledge base data, platform template library data, and online search data output from step 2 into a multidimensional feature vector; construct a feature space covering the planning area based on the feature vector; establish gridded partition units within the feature space, and perform spatial clustering based on the distribution density of the feature vectors; map the feature vectors to the corresponding partition units to generate a dynamic material pool containing spatial correlation attributes; Step 4, generating an outline structure space reference outline based on the spatial association attributes; Step 5: Based on the spatial reference outline generated in step 4, a segmented writing strategy is used to locate the content of the material pool and generate body paragraphs based on the feature weights of the partition units; Step 6, performing structured processing on the body paragraph generated in step 5 to obtain a structured text; Step 7: Perform dynamic verification on the structured text to output the final planning text.

[0006] In a second aspect, a natural resource planning text generation system includes: a receiving module for receiving standardized data input by a user, wherein the standardized data includes mandatory items and optional items, wherein the mandatory items include title, region, and writing subject, and the optional items include outline templates and unstructured reference materials; The execution module is used to make resource scheduling decisions based on the presence or absence of optional items. If an optional item exists, it uses a large language model to parse reference materials and extract keywords. If an optional item does not exist, it triggers local policy compliance analysis and platform database matching. A generation module is used to convert keyword sets, user knowledge base data, platform template library data, and online search data into multidimensional feature vectors; construct a feature space covering the planning area based on the feature vectors; establish gridded partition units within the feature space and perform spatial clustering based on the distribution density of the feature vectors; map the feature vectors to corresponding partition units to generate a dynamic material pool containing spatial correlation attributes; A processing module is used to generate an outline structure spatial reference outline based on the spatial association attributes; according to the spatial reference outline, a segmented writing strategy is adopted to locate the content of the material pool, and the main text paragraphs are generated based on the feature weights of the partition units; structured processing is performed on the main text paragraphs to obtain structured text; dynamic verification is performed on the structured text to output the final planning text.

[0007] According to a third aspect, a computer-readable storage medium stores a program, which implements the method described above when executed by a processor.

[0008] The above solution of the present invention includes at least the following beneficial effects: Large-scale language models are used to automatically parse unstructured reference materials in optional items and extract keywords, replacing the manual processing process. This mechanism reduces the time cost and manpower investment in data analysis, while avoiding omissions that may occur in manual operations, and improving the efficiency of pre-data processing for planning text generation.

[0009] By converting standardized data, user knowledge base, platform template library and networked data into multidimensional feature vectors, a feature space covering the planning area is constructed, and the feature vectors are mapped to specific partition units through grid partitioning and spatial clustering. Finally, a dynamic material pool containing spatial correlation attributes is generated, realizing the precise association of cross-modal data such as unstructured materials, policy data, geographic information and spatial units.

[0010] The outline structure is generated based on spatial association attributes, and the main text paragraphs are generated in combination with the feature weights of the partition units, forming a complete mapping chain of "spatial partition units-outline structure-main text content". For example, the chapter on watershed ecological restoration can automatically associate the characteristics of upstream and downstream grid units (such as water volume and vegetation coverage) through this mapping, ensuring that the generated content conforms to the spatial distribution patterns of natural resource elements and fundamentally avoiding spatial logical contradictions.

[0011] When there are no optional items, local policy compliance analysis and platform database matching are triggered, combined with dynamic verification to ensure that the generated text strictly complies with local policies, regulations and planning standards; at the same time, the dynamic material pool can be adjusted in real time as data is updated (such as new policy documents, real-time statistical data), so that the planning text can adapt to the dynamic changes in policies and data in natural resource management, thereby improving the timeliness and compliance of the text. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flow chart of a method for generating natural resource planning text provided by an embodiment of the present invention.

[0013] Figure 2 This is a schematic diagram of a natural resource planning text generation system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0014] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0015] like Figure 1 As shown, an embodiment of the present invention provides a method for generating a natural resource planning document, the method comprising the following steps: Step 1: receiving standardized data input by a user, wherein the standardized data includes mandatory items and optional items, wherein the mandatory items include title, region, and writing subject, and the optional items include outline templates and unstructured reference materials; Step 2: Execute resource scheduling decisions based on whether the optional items exist. If the optional items exist, a large language model is used to parse reference materials and extract keywords. If the optional items do not exist, local policy compliance analysis and platform database matching are triggered. Step 3: Convert the keyword set, user knowledge base data, platform template library data, and online search data output from step 2 into a multidimensional feature vector; construct a feature space covering the planning area based on the feature vector; establish gridded partition units within the feature space, and perform spatial clustering based on the distribution density of the feature vectors; map the feature vectors to the corresponding partition units to generate a dynamic material pool containing spatial correlation attributes; Step 4, generating an outline structure space reference outline based on the spatial association attributes; Step 5: Based on the spatial reference outline generated in step 4, a segmented writing strategy is used to locate the content of the material pool and generate body paragraphs based on the feature weights of the partition units; Step 6, performing structured processing on the body paragraph generated in step 5 to obtain a structured text; Step 7: Perform dynamic verification on the structured text to output the final planning text.

[0016] In an embodiment of the present invention, by distinguishing between required items (title, region, writing topic) and optional items (outline template, unstructured reference materials), it is ensured that the user input data meets the basic requirements for planning text generation, and subsequent processing errors caused by confusing data formats are avoided; the required items directly lock the core elements of the planning text (such as the title, region and topic of the "XX City Ecological Protection Plan"), and the optional items allow users to flexibly supplement personalized needs (such as custom outlines or reference materials), which not only ensures the accuracy of the generation direction, but also retains customization space, thereby improving the matching degree between user needs and the final text.

[0017] When optional items exist, a large language model is used to automatically parse unstructured reference materials (such as policy PDFs and report documents) and extract keywords, replacing manual processing. This improves efficiency by more than 50% and avoids the subjective bias of manual extraction. When there are no optional items, local policy compliance analysis and platform database matching are triggered to ensure that text generation is strictly anchored to the local policy framework (such as the Land Management Law and the Ecological Red Line Regulations), ensuring compliance from the source. By determining whether optional items exist, the "external model parsing" or "local database matching" path is dynamically selected to avoid invalid calculations, reduce system energy consumption, and improve processing speed.

[0018] The keyword set, knowledge base, template library and networked data are converted into multi-dimensional feature vectors (covering policy, geography, economy and other dimensions). Through grid partitioning and spatial clustering, unstructured text, statistical data and geographic coordinates are accurately bound (for example, a policy clause is associated with XX street in XX district), solving the problem of "disconnection between policy and space" in traditional methods; the dynamic material pool can absorb new data in real time (such as newly added remote sensing images and quarterly economic data) and automatically update the partition unit attributes through feature vector mapping to ensure that the material always reflects the latest situation and avoid decision-making bias in planning texts due to data lag; grid partition units (such as 1km×1km grids) make spatial feature analysis more refined. For example, in agricultural planning, the soil fertility, irrigation conditions and planting policies of each grid can be accurately associated, providing micro-data support for subsequent content generation.

[0019] Establishing a "spatial logic-text structure" mapping relationship: The outline no longer relies solely on semantic logic (such as "current situation-problem-countermeasures"), but instead incorporates spatial association attributes (such as "upstream protection area-midstream management area-downstream monitoring area") to ensure that the chapter structure is highly consistent with the spatial distribution patterns of natural resources (such as watershed flow direction and topographic zoning), avoiding the flaw of traditional outlines that "emphasizes text over space." The spatial reference outline determines the corresponding geographical divisions and core features of each chapter (such as "Chapter 3 focuses on the XX ecological red line area and emphasizes soil erosion data") to reduce the problem of mismatch between paragraph content and planned areas. By using the characteristic weights of the division units (such as ecologically sensitive areas with higher weights than ordinary areas), materials from high-weight units (such as restoration technologies and data in ecologically fragile areas) are automatically prioritized, ensuring that the main text highlights the core issues of the region (such as prioritizing windbreak and sand fixation measures in desertified areas) and avoiding generalized content.

[0020] In a preferred embodiment of the present invention, step 2, performing resource scheduling decisions based on whether the optional items exist, if the optional items exist, extracting keywords by parsing reference materials using a large language model, and if the optional items do not exist, triggering local policy compliance analysis and platform database matching, includes: Step 21, detecting the presence of optional items in the user input. If the outline template or unstructured reference material is detected, step 22 is executed; if no optional items are detected, step 23 is executed; Step 22: Perform structured parsing using a large language model to identify multimodal features of unstructured reference materials, extracting three core keywords: policy terms, geographic coordinates, and statistical indicators. Spatial topological relationships are then matched against the user's private knowledge base to generate a keyword set with spatial labels. Step 23: Trigger the local policy engine analysis. Based on the required regional attributes, load the policy specification library for the corresponding administrative division, perform timeliness verification and conflict detection on the policy clauses, and output a compliance constraint set. Step 24, perform dynamic matching of the platform database, search the platform template library with the writing topic as the index, match historical planning cases with regional feature similarity higher than the threshold; extract the spatial element association specifications in the case and generate a standardized keyword set.

[0021] In the embodiment of the present invention, the specific implementation process of the above step 21 is as follows: Perform field parsing on the standardized data submitted by the user to identify the fields that are "optional" (i.e., outline templates, unstructured reference materials); check whether the outline template contains a valid text structure (such as chapter hierarchy, title list), and check whether the unstructured reference materials contain parseable files (such as PDFs, documents, images, etc.) or text content; if at least one of the outline templates or unstructured reference materials contains valid content, it is determined that "optional items exist", triggering step 22; if both are empty values ​​or invalid content, it is determined that "optional items do not exist", triggering step 23.

[0022] The specific implementation process of the above step 23 is as follows: Extract "regional attribute" information (such as the administrative division code or name of XX city, XX county in XX province) from the required items and pass this information to the local policy engine; the local policy engine locates the corresponding administrative division level (such as provincial, municipal, and county levels) based on the regional attributes, and loads all policy clauses related to the writing topic at this level (such as land planning and ecological protection policies) from the pre-stored policy specification library; check the effective date and abolition date of each policy clause one by one, filter out expired or revised clauses, and retain the currently valid policy content; perform a logical comparison of the retained valid clauses to identify whether there are contradictions between the clauses (such as different regulations for the same matter). If there is a conflict, mark the conflict point and give priority to retaining clauses at a higher level or more recently released; organize the valid clauses that have passed timeliness verification and conflict detection into a structured "compliance constraint set" to determine the policy boundaries that the planning text must comply with (such as prohibited development areas, floor area ratio restrictions, etc.).

[0023] The specific implementation process of the above step 24 is as follows: Extract the "writing topic" (such as "XX River Basin Ecological Restoration Plan" and "XX City Land Use Master Plan") from the required items and break them down into core keywords (such as "River Basin Ecological Restoration" and "Land Use"); use the core keywords as the search index to traverse the historical planning cases in the platform template library and extract the metadata of the cases (such as the areas involved in the cases, topic tags, and regional feature descriptions).

[0024] Calculate the similarity between the regional characteristics of historical cases and the current plan, that is, compare the regional characteristics in the case (such as terrain type, climate zone) with the characteristics of the current planning area, calculate the similarity score through semantic matching and feature item overlap, and screen out historical cases with similarity scores higher than the preset threshold (such as 70%) as the reference case set; parse the content of the reference case set to extract the association specifications of the spatial elements (such as the binding relationship between "forest protection" and "slope > 25° area", and the association logic between "water conservation" and "1km range along the river"); standardize these association specifications and core terms in the case (such as policy keywords, geographical element names) (unify terminology and remove redundant information), and finally generate a structured "standardized keyword set".

[0025] The present invention accurately detects the existence status of optional items to achieve accurate triggering of resource scheduling, avoids invalid calculations, ensures that the system processes data according to the optimal path, and improves the overall process efficiency; uses large-scale language models to parse unstructured data, efficiently extracts three types of core keywords, and combines the spatial topology matching of the user's private knowledge base to generate a keyword set with spatial labels, which not only ensures the comprehensiveness of keyword extraction, but also establishes the association between keywords and spatial information; loads the corresponding policy specification library based on regional attributes, and outputs a compliance constraint set through timeliness verification and conflict detection to ensure that the generated planning text complies with the latest local effective policies, avoids policy conflicts from the source, and improves the compliance and authority of the text; uses writing topics as indexes to match high-similarity historical cases, extracts spatial element association specifications to generate a standardized keyword set, draws on historical experience to ensure the applicability and standardization of keywords, and provides content references that conform to regional characteristics for planning texts.

[0026] In a preferred embodiment of the present invention, step 22 performs structured parsing using a large language model to perform multimodal feature recognition on unstructured reference materials, extracting three core keywords: policy terms, geographic coordinates, and statistical indicators, including: Step 221 , performing format standardization preprocessing on the unstructured reference materials, eliminating file format differences and unifying the encoding to generate a parseable continuous text stream; Step 222 , based on the continuous text stream, performing data modality segmentation by a multimodal feature recognition engine: identifying and separating text description segments, image data blocks, and table structured data; Step 223: Based on the text description segment separated in step 222, semantic role labeling technology is used to parse the sentence components: the responsible party in the policy text is identified as the first category of keywords, the obligation item as the second category of keywords, and the time constraint as the third category of keywords, thereby generating a structured policy clause set. Step 224: Based on the image data blocks separated in Step 222, identify the annotation element texts in the map through OCR technology, convert the element position information to a unified projection coordinate system, and generate a standard geographic coordinate set with coordinate system labels. Step 225: Based on the table structured data separated in Step 222, verify the unit consistency of numerical data, extract the index name fields, numerical fields, and statistical period fields, and generate a standardized statistical index set.

[0027] In the embodiment of the present invention, the specific implementation process of the above Step 221 is as follows: First, perform a format scan on the received unstructured reference materials, identify the file types (such as PDF, Word, TXT, JPG, PNG, etc.), and classify and mark files of different formats; for PDF and Word files, strip the format marks (such as font styles, paragraph spacing, picture embedding codes) through a text extraction tool, and extract the pure text content; for picture files (such as scanned documents, map photos), first use image recognition technology to determine whether they contain convertible text, and if so, perform OCR conversion to text format; perform encoding detection on the converted text content, identify the current character encoding format (such as GBK, UTF-8, ISO-8859-1, etc.), and uniformly convert all text to UTF-8 encoding to eliminate the garbled problem caused by encoding differences; delete meaningless symbols (such as special control characters, repeated spaces, page breaks), and merge broken sentences (such as sentence splitting caused by line breaks), and finally form a continuous text stream without format differences, ensuring that subsequent parsing tools can read it stably.

[0028] The specific implementation process of the above Step 222 is as follows: Start the engine and load a preset multi-modal feature library, which contains typical feature templates for text description segments, image data blocks, and table structured data, such as: the subject-predicate-object structure grammar norms and common punctuation combinations of text segments; the binary file header identifiers of image data such as JPG ("FFD8"), PNG ("89504E47"), etc., and text placeholders such as "[picture]" and "[image]"; the separator features of table data such as vertical lines "|", horizontal lines "—", etc., the special mark codes of Word table objects (such as "\tbl" and "\tcell"), and the alignment mode features of fixed-spacing numerical columns; scan the continuous text stream byte by byte. The engine starts from the starting position of the continuous text stream and sequentially reads the content byte by byte, while starting three parallel recognition threads to perform real-time matching on the features of text description segments, image data blocks, and table structured data. During the scanning process, the system will record the position index of the current byte in the text stream (such as the Nth byte) as the benchmark for subsequent positioning.

[0029] Recognition and marking of text description segments: Thread 1 continuously detects whether consecutive bytes form a parsable text character (excluding binary gibberish). When it recognizes that consecutive content conforms to the norms of natural statements (such as a combination of words containing a subject, a predicate, and an object, and interspersed with punctuation marks such as commas and periods), and no table separator or image identifier is detected, it is initially determined as a text description segment; further verify whether the text block has no fixed row and column structure: by counting the length differences of characters in each row, if the length fluctuation exceeds a preset threshold (such as ±10 characters), it is confirmed that it is a non-table text, mark the start byte index and end byte index of the text block (bounded by the period or line break at the end of the paragraph), and temporarily store it in the text buffer.

[0030] Recognition and extraction of image data blocks: Thread 2 synchronously detects image features in the text stream: If placeholder texts such as "[Picture]" or "[Image]" are scanned, immediately mark the position index of the placeholder, and use it as the mark of the image data block, extract the context position information before and after the placeholder (such as "50 bytes before to 20 bytes after [Picture]"), and store it in the image buffer; if a binary image file header that matches the feature library (such as "FFD8" for JPG) is detected, continue scanning from this position until the corresponding file tail identifier (such as "FFD9" for JPG) is recognized, determine the complete byte range of the image data block (from the file header to the file tail), extract all bytes within this range as image data, record its start and end indexes, and then store it in the image buffer.

[0031] Recognition and extraction of table structured data: Thread 3 focuses on detecting table features: When consecutive vertical lines "|" are scanned and the same number of vertical line separations appear in adjacent rows (such as 3 "|" in each row, forming a 4-column structure), or when horizontal lines "—" appear continuously and are aligned with the text in the upper and lower rows (such as the separator line below the table header), it is determined as a potential table area; if the text stream is from a Word file conversion, when the marker codes "\tbl" (table start) and "\endtbl" (table end) are recognized, directly use these two markers as boundaries to determine the byte range of the table; for suspected table areas, further verify the cell alignment feature: by calculating the character spacing between columns, if the consistency of left alignment / right alignment of numerical values or texts within the same column exceeds 90%, it is confirmed that it is table structured data, mark the start and end indexes, and then store it in the table buffer.

[0032] Storage in the buffer and index record: After the three threads complete their respective recognition tasks, they write the text description segment, image data block, and table structured data into three independent buffers (such as three arrays in memory). The data entries in each buffer are accompanied by their complete position index in the original continuous text stream (including the starting byte index and the ending byte index), such as "text segment 1: starting 100th byte - ending 500th byte" and "image block 1: starting 800th byte - ending 2000th byte", ensuring that subsequent steps can trace the position association of each data modality in the original text.

[0033] The specific implementation process of the above step 223 is as follows: First, the text description segment is scanned character by character to identify a set of punctuation marks (including periods ("."), exclamation points ("!"), semicolons (";"), and question marks ("?"). These punctuation marks are used as potential sentence boundaries. If the punctuation mark is located within quotation marks ("", '') or brackets ((), []) (for example, "The document stipulates: 'XX area shall not be developed; violators will be punished'"), it is temporarily not used as a sentence boundary to avoid splitting the complete sentence within the quotation marks. When punctuation marks outside quotation marks or brackets are scanned, the sentence boundary is determined. The long text is split into independent sentence units according to the verified boundaries. Each sentence unit is assigned a unique serial number (such as "Sentence 1" and "Sentence 2"), and the starting and ending positions of the sentence in the original text description segment are recorded (for example, from line X, character Y to line M, character N) for subsequent tracing.

[0034] Semantic role labeling model initialization and sentence parsing preparation: Load the pre-trained semantic role labeling model, which contains three types of feature libraries: an institutional entity feature library (including common government departments, agency names and abbreviations), an obligatory verb feature library (including binding verbs and phrases such as "prohibited", "should", and "shall not"), and a time expression feature library (covering typical patterns of date, time period, and time description, such as "MM / DD / YYYY" and "valid until XXX"); convert each sentence unit into a sequence format that the model can process (such as a word list after word segmentation), while retaining the original position information of the word in the sentence (such as the first word, the second word).

[0035] Identification of responsible parties (first category keywords): The sentence units are scanned through the named entity recognition module to match the entries in the institutional entity feature library: if the complete institutional name such as "XX Municipal Natural Resources Bureau" or "Provincial Forestry Department" appears in the sentence, it is marked as a candidate responsible entity; the grammatical relationship between words (such as subject-predicate relationship, verb-object relationship) is identified through the dependency syntax tree to determine whether the candidate entity is the issuer or responsible party of the core verb in the sentence; if the candidate entity meets the above conditions, it is confirmed as the responsible entity (first-category keyword), and its specific text and position index in the sentence are recorded.

[0036] Obligation identification (second category keywords): Verb phrase extraction for sentence units: Scan the verbs and modifying components (such as adverbs and prepositional phrases) in the sentence, match the obligatory verb feature library, and identify phrase structures such as "prohibited + action" (such as "prohibited occupation"), "should + action" (such as "should repair"), and "shall not + action" (such as "shall not break through"); determine the pointed object of the verb phrase through semantic role labeling (such as "prohibited occupation of permanent basic farmland", "prohibited occupation" points to "permanent basic farmland"), combine the verb phrase with the pointed object to form a complete behavioral norm (such as "prohibited occupation of permanent basic farmland"); if the behavioral norm has a definite association with the identified responsible party (such as the responsible party is the executor of the behavior), mark it as an obligation item (second-category keyword), and record its complete text and position in the sentence.

[0037] Time constraint identification (third category keywords): Identify time-related expressions such as specific dates (such as "January 1, 2024"), time periods (such as "2024-2030"), effective conditions (such as "from the date of publication"), and deadlines (such as "valid until 2030"); standardize ambiguous time expressions: if ambiguous expressions such as "in the near future" and "medium to long term" appear in a sentence, convert them into understandable time ranges based on the contextual policy background (such as a specific time extracted from other sentences) or default specifications (such as "in the near future" corresponds to "within 3 years"); verify whether the time expression is related to the obligations of the responsible party (such as "From 2024, the XX Bureau prohibits development", "from 2024" constrains "prohibition of development"); if so, mark it as a time constraint (third-category keyword) and record its text and location.

[0038] Three types of keyword associations and structured policy clause set generation: The responsible parties, obligations, and time constraints within the same sentence unit are associated and bound: using the sentence number as the unique identifier, a correspondence table of "responsible party-obligation item-time constraint" is established, for example, "Sentence 3: Subject = XX City Natural Resources Bureau, Obligation Item = Prohibition of Occupation of Permanent Basic Farmland, Time Constraint = From 2024"; if a certain type of keyword is not identified in the sentence (such as no definite time constraint), mark "None" at the corresponding position and retain the original text of the sentence for manual review (such as "XX Bureau should restore the ecology - no definite time constraint"); summarize the association results of all sentences and sort them by sentence number to form a structured policy clause set containing "sentence number, responsible party, obligation item, time constraint, original sentence index", ensuring that each clause can be traced back to the specific location of the original text description segment.

[0039] The specific implementation process of the above step 224 is as follows: The denoising algorithm is used to eliminate spots and scratches in the image, and the contrast enhancement technology is used to highlight the text annotations and boundary lines in the map. The tilted map is geometrically corrected to ensure the horizontal alignment of the image. The pre-processed map image is fully scanned by the OCR text recognition tool to identify and extract various types of annotation text in the map, including place names (such as "XX County" and "XX River"), administrative boundary names (such as "provincial boundary" and "municipal boundary"), coordinate scales (such as "116° East Longitude" and "39° North Latitude") and legends (such as the symbols for "forest land" and "arable land"); the projection method annotation of the map corners (such as "Gauss-Krüger projection") is used to identify the map corners. The original coordinate system of the map is determined by using the projection conversion formula (such as the pixel coordinates of a place name on the map) and the scale (such as "1:100000") and the coordinate origin parameters. The location information of each feature in the map (such as the pixel coordinates of a place name on the map) is mapped to the unified coordinate system preset by the system (such as the WGS84 longitude and longitude coordinate system), and the precise coordinate range of the feature is calculated (such as "116.2°-116.5° East longitude, 39.1°-39.3° North latitude"). The converted coordinate information is bound to the feature text recognized by OCR, and a coordinate system tag (such as "WGS84") is added to generate a standard geographic coordinate set containing "feature name-coordinate range-coordinate system".

[0040] The specific implementation process of the above step 225 is as follows: Parse the table header rows, unify synonymous headers (such as "Statistical Year" and "Data Year") as "Statistical Period," and delete empty columns without actual data. Check for duplicate rows in the table (e.g., identical indicators and values), retaining unique rows to avoid redundancy. Iterate through the numeric fields in the table (e.g., "Area," "Population," and "Output Value"), extract the units after each value (e.g., "Hectare," "Per capita," and "10,000 Yuan"), and compare them with the system's preset indicator unit library: If the units are inconsistent (such as "hectares" and "mu" appearing simultaneously in the same "area" indicator), they will be uniformly converted into standard units (such as "hectares") according to the conversion formula (such as 1 hectare = 15 mu); if the value is not marked with a unit, the standard unit will be automatically added according to the indicator name (such as the default unit of "forest coverage rate" is "%").

[0041] Extract core fields: Extract the name that reflects the meaning of the data from the first column or header of the table (such as "cultivated land reserves" and "ecological protection red line area"); extract the specific values ​​after unit standardization (such as "5,000 hectares" and "35.2%"); extract the time range corresponding to the data from the header or designated column (such as "2023" and "2021-2023"); combine the three extracted fields into structured records according to the correspondence between "indicator name-value-statistical period", and summarize all records to form a standardized statistical indicator set.

[0042] The present invention eliminates format differences in unstructured reference materials through format standardization preprocessing and unifies the coding, reducing parsing errors caused by format problems and improving the stability of overall data processing; uses a multimodal feature recognition engine to accurately separate different modal data such as text, images, and tables, so that each type of data can be processed in a targeted manner, avoiding parsing interference caused by the mixing of different modal data, laying the foundation for the subsequent extraction of policy clauses, geographic coordinates, and statistical indicators, and improving the accuracy of parsing; uses semantic role labeling technology to parse text description segments, accurately identify responsible parties, obligations, and time constraints, and form a structured policy clause set, clearly presenting the core elements and interrelationships of the policy, facilitating subsequent association with spatial information, and improving the efficiency and accuracy of policy clause utilization; uses optical character recognition (OCR) technology to identify map annotations and convert them to a unified coordinate system to generate a labeled standard geographic coordinate set, achieving accurate extraction and standardization of geographic information in images and solving the problem that map-based data is difficult to directly parse and utilize; verifies the consistency of table data units and extracts key fields to generate a standardized statistical indicator set, ensuring the standardization and comparability of statistical data and avoiding data misuse due to unit confusion or unclear fields.

[0043] In a preferred embodiment of the present invention, the keyword set includes: Step 226: Establish a spatial reference coordinate system based on the administrative division vector boundary data in the user's private knowledge base; Step 227: Based on the spatial reference coordinate system, perform spatial weight matching on the structured policy clause set, extract the geographic entity names in the policy clauses, and perform semantic association of the geographical names with the administrative division vector boundaries to obtain association results; based on the association results, map the policy clauses to corresponding administrative division units to generate a policy clause label set with administrative division codes; Step 228: Based on the spatial reference coordinate system, a topological relationship calculation is performed on the standard geographic coordinate set, and spatial location inclusion is determined between the geographic coordinate point and the administrative division vector boundary. When the coordinate point is located in the intersection area of ​​multiple administrative divisions, the topological distance between the coordinate point and each boundary is calculated and a membership weight is assigned to generate a coordinate label set with topological weights. Step 229: Perform spatial scale conversion on the standardized statistical indicator set to identify the spatial scale attributes of the statistical indicators, perform spatial interpolation on the indicator values ​​at the non-administrative division scale according to population density, convert the statistical indicator values ​​to generate administrative division unit granularity, and attach administrative division code labels. Step 230: Fusing the policy clause label set, the coordinate label set, and the administrative division code label set to construct a spatial topological relationship matrix indexed by administrative division units. Step 231 : Generate a keyword set with spatial topology tags based on the spatial coupling degree of policy terms, geographic coordinates and statistical indicators in the spatial topology relationship matrix.

[0044] In the embodiment of the present invention, the specific implementation process of the above step 226 is as follows: Call administrative division vector boundary data from the user's private knowledge base. This data contains boundary polygon coordinate information (such as latitude and longitude coordinate strings) of administrative divisions at different levels (such as provinces, cities, counties, and townships), as well as metadata such as the corresponding administrative division name and code. Parse the original coordinate system of the vector boundary data (such as WGS84, GCJ02, etc.) and convert it into a preset unified spatial reference coordinate system (such as the National 2000 Geodetic Coordinate System) to ensure that all boundary data are in the same coordinate framework. Perform accuracy verification on the converted vector boundary data, eliminate duplicate or erroneous boundary coordinates (such as self-intersecting polygons), and retain complete and accurate administrative division boundary information as a benchmark reference for subsequent spatial associations.

[0045] The specific implementation process of the above step 227 is as follows: Extract geographic entity names sentence by sentence from a set of structured policy clauses: Scan each policy clause using a place name recognition tool to locate the contained administrative division names (such as "XX District" and "XX County"), geographical area names (such as "XX River Basin" and "XX Mountain Range"), and other geographical entities; Compare the extracted geographic entity names with the administrative division vector boundary metadata (including standard place names, aliases, and former names) in the user's private knowledge base, and handle name differences through fuzzy matching (such as "Beijing City" matches "Nanjing City") and synonym mapping (such as "XX Region" corresponds to "XX City") to determine the administrative district to which the geographic entity corresponds; Based on the association results, bind the policy clause to the corresponding administrative district boundary polygon, query the unique code of the administrative district (such as a six-digit administrative division code), and attach the code to the policy clause; Each policy clause is marked in a structured format of "policy content + associated administrative district name + administrative district code", which are aggregated to form a policy clause label set with administrative division codes.

[0046] The specific implementation process of the above step 228 is as follows: Extract the coordinate values ​​(such as longitude and latitude) of all geographic coordinate points from the standard geographic coordinate set, and project the coordinate points to the spatial reference coordinate system established in step 226; for each coordinate point, compare it with the polygon of the vector boundary of each administrative division in turn, and determine the basic administrative district to which it belongs by checking whether the coordinate point is located inside the polygon (not on the boundary); if the coordinate point is located on the boundary line of two or more administrative districts (such as provincial boundaries or municipal boundaries), calculate the straight-line distance from the point to the boundary of each adjacent administrative district (such as 500 meters to the boundary of City A and 300 meters to the boundary of City B); assign affiliation weights according to the inverse proportion of distance (such as 3 / 8 weight for City A, 5 / 8 weight for City B, and the total weight is 1); each coordinate point is marked in the format of "coordinate value + main affiliation administrative district code + weight of each affiliation administrative district", and after aggregation, a coordinate label set with topological weight is formed.

[0047] The specific implementation process of the above step 229 is as follows: Scan the metadata of each indicator in the standardized statistical indicator set to determine its original spatial scale (such as "national", "provincial", "watershed level", "plot level", etc.). For indicators with non-administrative division original scale (such as "XX watershed water quality compliance rate", "XX economic zone GDP"), collect population density data of all administrative districts within the coverage of this scale; allocate non-administrative division indicator values ​​to subordinate administrative districts according to population density ratio (such as watershed indicator values ​​are divided according to the proportion of population density of coastal cities and counties), and calculate the corresponding indicator value for each administrative district; bind the converted administrative district granularity indicator value to the corresponding administrative district code, and mark it in the format of "indicator name + value + statistical period + administrative district code" to form a statistical indicator set with additional administrative district code labels.

[0048] The specific implementation process of the above step 230 is as follows: The core index is administrative division unit (such as administrative division code as key value), through policy clause label set, coordinate label set, administrative division code label set (statistical indicator set); the rows of the matrix represent administrative division codes, and the columns contain three types of associated information, namely, policy clauses associated with the administrative division (from policy clause label set), geographical coordinate points and weights within the administrative division and at the intersection (from coordinate label set), and statistical indicator values ​​of the administrative division (from statistical indicator set); the associated items of adjacent administrative divisions are added to the matrix (such as policy clauses shared by City A and adjacent City B, and weight distribution of intersection coordinate points) to determine the spatial topological association between administrative divisions (such as adjacent, contained, and intersection).

[0049] The specific implementation process of the above step 231 is as follows: Based on the spatial topological relationship matrix, for each administrative region, the coupling degree between policy clauses and geographic coordinates (such as whether the area mentioned in the policy clauses overlaps with the coverage of the coordinate points, and the higher the overlap ratio, the higher the coupling degree), the coupling degree between geographic coordinates and statistical indicators (such as whether the statistical indicator values ​​of the area where the coordinate points are located match the characteristics of the coordinate points), and the coupling degree between policy clauses and statistical indicators (such as whether the policy requirements are consistent with the current status of statistical indicators) are calculated respectively; according to the coupling degree results, labels (such as "high coupling", "medium coupling", and "low coupling") are added to the policy clauses, geographic coordinates, and statistical indicators respectively to mark their association strength within the administrative region; the policy clause keywords, geographic coordinate keywords, and statistical indicator keywords with spatial topological labels are summarized and classified according to the administrative region code to form the final keyword set with spatial topological labels, ensuring that each keyword contains the administrative region to which it belongs, related elements, and coupling strength information.

[0050] The present invention establishes a spatial reference coordinate system through the administrative division vector boundary data of the user's private knowledge base, eliminates the spatial matching errors caused by the differences in coordinate systems of different data sources, and ensures that the correlation of factors such as policy clauses, geographic coordinates, and statistical indicators in the spatial dimension can be accurately calculated. It realizes the precise binding of policy clauses with specific administrative division units, and solves the problem of "the geographical names mentioned in the policy text do not match the administrative boundaries" through the semantic association of place names. The generated label set with administrative division codes gives the policy clauses a definite spatial directionality; for the spatial ownership problem of geographic coordinate points, through topological relationship calculation and affiliation weight allocation, it not only determines the ownership of coordinate points in a single administrative district, but also properly handles the coordinate association of the border area of ​​multiple administrative districts. The generated label set with topological weights ensures that the spatial correlation of geographic coordinates is not cut off by boundaries, thereby improving the integrity of spatial data.

[0051] Statistical indicators at non-administrative division scales are converted to a granularity that matches administrative divisions. The indicators are adapted to the spatial scale through population density interpolation. After adding administrative division codes, statistical data can be directly associated with the policies and geographic information of the corresponding areas, solving the data utilization barriers caused by "mismatch between indicator scales and planning units"; three types of label sets are integrated with administrative divisions as indexes, and the constructed spatial topological relationship matrix clearly presents the association network of policy, geographic, and statistical elements in the same area, intuitively reflecting the spatial binding relationship across elements; a keyword set with topological labels is generated through spatial coupling calculation, which quantifies the spatial association strength of policy, geographic, and statistical elements, helps identify the degree of matching between elements (such as highly coupled policies are more suitable for geographic coordinates), provides an accurate basis for material screening and content association when generating planning texts, and improves the spatial logic of text content.

[0052] In a preferred embodiment of the present invention, step 3 converts the keyword set, user knowledge base data, platform template library data, and online search data output in step 2 into a multidimensional feature vector; and constructs a feature space covering the planning area based on the feature vector, including: Step 31, receiving the keyword set with spatial topology tags output in step 2, the administrative division boundary data in the user knowledge base, the historical case spatial element data matched in the platform template library, and the real-time geographic data obtained by online search; Step 32, respectively converting the keyword set with spatial topology labels, administrative division boundary data, historical case spatial element data, and real-time geographic data into semantic-spatial joint feature vectors, boundary topology feature vectors, element association feature vectors, and spatiotemporal dynamic feature vectors; Step 33, taking the spatial scope defined by the administrative division boundary of the planning area as a reference; Step 34 : constructing a multi-dimensional feature space covering the spatial range of the planning area based on the semantic-spatial joint feature vector, the boundary topology feature vector, the element association feature vector, and the spatiotemporal dynamic feature vector.

[0053] In the embodiment of the present invention, the specific implementation process of the above step 31 is as follows: Obtain the keyword set with spatial topology labels output from step 2. This set includes three types of keywords: policy terms, geographic coordinates, and statistical indicators, as well as their corresponding administrative division codes, topological weights, and spatial coupling labels. Retrieve administrative division boundary data from the user knowledge base, including vector boundary polygons (composed of latitude and longitude coordinate strings) for administrative divisions at all levels within the planning area, including provinces, cities, counties, and townships, administrative division names, administrative codes, and hierarchical relationships (e.g., a county is affiliated with a prefecture-level city). Access the platform template library to extract spatial feature data for the highly similar historical cases matched in step 24. This includes geographic features in the cases (e.g., the location coordinates of rivers and forests), inter-feature association specifications (e.g., "the distance relationship between reservoirs and irrigation areas"), and spatial description information mentioned in the case text (e.g., "northeastern mountains"). Based on the name of the planning area and the writing topic, retrieve real-time geographic data published by authoritative geographic information platforms and government open data platforms, including recent remote sensing image interpretation results (e.g., changes in land use types), meteorological monitoring data (e.g., precipitation distribution in the past three months), and transportation construction dynamics (e.g., the direction of new roads). The search results are screened for validity (eliminating duplicate or outdated data).

[0054] The specific implementation process of the above step 32 is as follows: The keyword set with spatial topological labels is feature-decomposed, and the semantic information of each keyword (such as "ecological protection" and "arable land area") is converted into a semantic code (a numerical sequence generated based on the pre-trained word vector model); the spatial attributes of the keyword (such as administrative division code, topological weight, coordinate range) are extracted and converted into a spatial code (a numerical sequence reflecting location, range, and association strength); the semantic code and the spatial code are spliced ​​according to the corresponding relationship to form a semantic-spatial joint feature vector for each keyword, and each dimension in the vector corresponds to a semantic feature or a spatial feature.

[0055] Convert to boundary topology feature vector: For each vector polygon in the administrative boundary data, the geometric features (such as the number of vertices, perimeter, area) and topological relationships (such as the border length between a certain area and the adjacent area, and the inclusion relationship) of the boundary are extracted; these features are converted into numerical representations (such as "1" for inclusion relationship and "0.3" for the proportion of border length to total perimeter), sorted by administrative level, and then spliced ​​to form a boundary topological feature vector.

[0056] Convert to feature-associated eigenvectors: Analyze the geographic elements in the spatial feature data of historical cases, and convert the location, type, and attributes (such as river length and reservoir capacity) of each element (such as "XX river" and "XX reservoir") into basic numerical features; based on the inter-element association specifications, extract the triples of "element A-relationship type-element B" (such as "reservoir-water supply-town"), and convert the relationship type (such as "water supply" and "surrounding") into association strength values ​​(such as "1.0" indicates strong association); integrate the basic numerical features and association strength values, sort the elements according to their importance in the case, and form an element association feature vector.

[0057] Convert to spatiotemporal dynamic feature vector: Real-time geographic data are stratified by time dimension (such as the past month, the past three months, and the past year), and dynamic change indicators within each time period are extracted (such as the area of ​​land use type conversion and the amplitude of precipitation change); spatial distribution characteristics (such as the proportion of a certain type of land in the region) and temporal change characteristics (such as the monthly precipitation growth rate) are converted into numerical sequences, and combined according to the "spatial distribution + temporal change" dimension to form a spatiotemporal dynamic feature vector.

[0058] The specific implementation process of the above step 33 is as follows: From the administrative division boundary data obtained in step 31, extract the highest-level administrative division boundary of the planning area (such as the city-level boundary of "XX City"), and determine the maximum coordinate extreme values ​​of its spatial range (minimum east longitude, maximum east longitude, minimum north latitude, maximum north latitude); verify the integrity of the boundary: check whether the boundary polygon is closed and whether the coordinate points are continuous. If there are breakpoints or abnormal coordinates (such as longitude and latitude values ​​that exceed common sense), correct them through the backup data in the user knowledge base; use the area enclosed by the boundary as the benchmark spatial range, mark all lower-level administrative districts contained in the range (such as the 5 districts and 3 counties under the jurisdiction of XX City), and determine the spatial coverage and internal administrative hierarchy structure of the planning area.

[0059] The specific implementation process of the above step 34 is as follows: Determine the dimensional composition of the multidimensional feature space: integrate all dimensions of the semantic-spatial joint feature vector (covering the semantic and spatial attributes of policies and indicators), the boundary topology feature vector (reflecting the geometry and association of administrative boundaries), the feature association feature vector (reflecting the feature relationship in historical cases), and the spatiotemporal dynamic feature vector (including the spatiotemporal changes of real-time data), and remove duplicate or redundant dimensions (for example, the "administrative division code" dimension contained in different vectors is only retained once); unify the spatial references of all vectors to the spatial reference coordinate system established in step 226 to ensure that the spatial position corresponding to each feature dimension can be located within the planning area; divide the reference spatial range of the planning area into several grids according to a preset accuracy (such as 1 km × 1 km), and each grid serves as the basic unit of the feature space; map the values ​​of each feature vector to the corresponding grid unit according to the spatial position, so that each grid unit contains the values ​​of all feature dimensions such as semantics, topology, association, spatiotemporal, etc. of the location, and finally form a multidimensional feature space covering the spatial range of the entire planning area.

[0060] The present invention achieves comprehensive aggregation of planning-related information by receiving multi-source data (keyword sets, administrative division data, historical case elements, real-time geographic data), avoiding the limitations of a single data source. At the same time, the data source and type are determined to ensure that the feature vector can cover multi-dimensional information such as policy, space, historical experience, and real-time dynamics. Different types of data (keywords, boundaries, case elements, real-time geographic data) are converted into standardized feature vectors (semantic-spatial union, boundary topology, element association, spatiotemporal dynamics), eliminating the integration barriers caused by differences in data format and type. The spatial scope is determined based on the administrative division boundary of the planning area to ensure that the spatial directionality of all feature data strictly matches the planning area, avoiding the problem of feature space exceeding the actual planning scope or incomplete coverage. This benchmark provides an anchor point for the spatial alignment of multi-source features, ensuring that the feature space constructed subsequently can accurately reflect the actual situation of the planning area. Based on multi-type feature vectors, a multi-dimensional feature space covering the planning area is constructed, realizing the organic integration of semantic, topological, association, time and space and other multi-dimensional information. This feature space not only completely covers the spatial scope of the planning area, but also can intuitively present the attributes, relationships and dynamic changes of each element in the area, so that the planning text generation can deeply integrate regional characteristics and improve the targetedness of the content.

[0061] In a preferred embodiment of the present invention, a gridded partition unit is established in the feature space, and spatial clustering is performed according to the distribution density of the feature vector; the feature vector is mapped to the corresponding partition unit to generate a dynamic material pool containing spatial correlation attributes, including: Step 35: Based on the multi-dimensional feature space constructed in step 34, a standard grid coordinate system covering the spatial range of the planning area is established to form a plurality of gridded partition units; Step 36, calculating the spatial distribution density of the semantic-spatial joint feature vector, the boundary topology feature vector, the element association feature vector, and the spatiotemporal dynamic feature vector generated in step 32 in the multidimensional feature space; Step 37, performing a spatial clustering operation on all feature vectors in the multidimensional feature space according to the spatial distribution density calculated in step 36, and identifying clustered areas of the feature vectors in the gridded partition units; Step 38, mapping each eigenvector generated in step 32 to a partition unit corresponding to the aggregation area to which it belongs in the gridded partition unit established in step 35; Step 39, in each gridded partition unit, associate the feature vector mapped to the unit in step 38 and its corresponding spatial correlation attribute; integrate the associated feature vectors and spatial correlation attributes in all gridded partition units to generate the dynamic material pool containing the spatial correlation attributes.

[0062] In the embodiment of the present invention, the specific implementation process of the above step 35 is as follows: From the multidimensional feature space constructed in step 34, extract the spatial extent parameters of the planning area, including the easternmost and westernmost longitude values ​​and the southernmost and northernmost latitude values ​​of the area, to clarify the geographic boundaries of the planning area. Based on the size of the planning area and the accuracy requirements of natural resource planning, set the size standard of the grid partition unit (e.g., 500 meters × 500 meters for urban planning and 1000 meters × 1000 meters for rural planning) to ensure that the grid can reflect regional details without causing excessive computational complexity due to excessive density. With the coordinates of the southwest corner of the planning area (minimum longitude and minimum latitude) as the origin, establish an X-axis along the east-west direction (longitude) and a Y-axis along the north-south direction (latitude) to form the reference axes of the standard grid coordinate system. Grid lines are divided along the X-axis and Y-axis according to the set grid size: starting from the origin, mark a vertical line at each interval of the set size along the X-axis, and mark a horizontal line at each interval of the set size along the Y-axis. The intersection of the vertical and horizontal lines forms the grid vertices.

[0063] Each rectangular area enclosed by two adjacent vertical lines and two horizontal lines is a grid partition unit. Each unit is assigned a unique identifier (such as "G-row number-column number", where the row number corresponds to the serial number in the Y-axis direction and the column number corresponds to the serial number in the X-axis direction); the coordinates of the four vertices of each grid partition unit (such as the latitude and longitude of the upper left corner, upper right corner, lower left corner, and lower right corner) are recorded to ensure that all units completely cover the planned area and there is no overlap or gap between units.

[0064] The specific implementation process of the above step 36 is as follows: From the four types of feature vectors generated in step 32, the spatial positioning information of each feature vector is extracted: the semantic-spatial joint feature vector extracts the coordinates of the associated administrative division center point; the boundary topology feature vector extracts the coordinates of the corresponding administrative boundary key point; the element association feature vector extracts the coordinate range of the geographical elements (such as rivers and woodlands) involved; the spatiotemporal dynamic feature vector extracts the coordinates of the sampling points corresponding to its monitoring data; for each feature vector, the grid partition unit to which it belongs is determined by coordinate comparison: the spatial coordinates of the feature vector are compared with the grid unit vertex coordinates recorded in step 35 to determine which unit the coordinates fall within the rectangular range of, that is, the initial grid unit to which the vector belongs.

[0065] Count the number of eigenvectors of each type in each grid cell: establish a counting table for each of the four types of eigenvectors, traverse all vectors, and each time a grid cell to which a vector belongs is determined, add 1 to the number of grid cells in the counting table of the corresponding feature type, and finally obtain the total number of eigenvectors of the four types in each grid cell (i.e., frequency); divide the frequency of a certain type of eigenvector in each grid cell by the actual area of ​​the grid cell (in square kilometers calculated based on the longitude and latitude) to obtain the density value of this type of feature in the grid (e.g., "number / square kilometer"); for linear or planar eigenvectors covering multiple grid cells (e.g., the feature association vector corresponding to a river spanning three grid cells), assign frequencies according to their spatial proportion in each grid cell (e.g., river length proportion, area proportion) (e.g., if the total frequency is 1, assign 0.3, 0.5, and 0.2 to the three grid cells respectively), and then calculate the density of each grid cell based on the assigned frequencies.

[0066] The specific implementation process of the above step 37 is as follows: Integrate the spatial distribution densities of the four types of feature vectors obtained in step 36 to construct a comprehensive density vector for each grid cell. The vector contains four dimensions, corresponding to the density values ​​of the four types of features. Set the core parameters of spatial clustering, including the density threshold (for example, in the comprehensive density vector of a grid cell, the density values ​​of at least three dimensions are higher than the average density of their respective types) and the distance threshold (for example, the straight-line distance between the center points of two grid cells does not exceed 2 grid side lengths). Traverse all grid cells and mark the grid cells whose comprehensive density vectors meet the density threshold as "core cells", and those that do not meet the threshold as "non-core cells". Starting from the first core cell, search for all other core cells within the threshold around it and classify these core cells into the same cluster. Then, starting from each core cell in the cluster, repeatedly search for surrounding core cells that meet the conditions until no new core cells are added, forming a complete cluster area.

[0067] For unclassified non-core units, if their distance to the core unit of a cluster is within the distance threshold, they will be classified into the cluster; if the distance to all core units exceeds the threshold, they will be marked as "isolated units" (not included in the cluster area); assign a unique cluster number to each cluster, record all the grid units contained in each cluster (including core units and merged non-core units), and mark the main feature type of the cluster (for example, if the density of "feature association features" in a cluster is the highest, it will be marked as "feature-intensive area").

[0068] The specific implementation process of the above step 38 is as follows: Traverse all the feature vectors generated in step 32, query the cluster to which each vector belongs in step 37 (determined by the cluster number of the grid cell to which it initially belongs), extract the precise spatial coordinates of all feature vectors belonging to the same cluster (such as the latitude and longitude of the sampling point, the coordinates of the feature center point), find a list of all the grid cells contained in the cluster (from the clustering result of step 37), compare the precise coordinates of each feature vector with the boundary coordinates of each grid cell in the list, and determine the unique grid cell to which the vector ultimately belongs (that is, the coordinates strictly fall within the rectangular range of the grid cell); if the coordinates of the feature vector happen to be on the boundary line of two grid cells (such as the longitude is equal to the longitude of a longitudinal line), determine the ownership based on its associated spatial attributes (such as the core feature closer to a certain cell) to avoid repeated mapping; establish a correspondence table of "feature vector ID-grid cell identifier-cluster cluster number", record the grid cell to which each feature vector is finally mapped, and ensure that all vectors are accurately assigned to a unique grid partition cell.

[0069] The specific implementation process of the above step 39 is as follows: Traverse all cells in the identification order of the grid partition cells (e.g., from "G-1-1" to "Gnm"), perform the following operations on each cell, and extract all feature vectors mapped to the cell according to the correspondence table in step 38, including four types of vectors: semantic-spatial union, boundary topology, feature association, and spatiotemporal dynamics.

[0070] The spatial association attributes corresponding to each eigenvector are extracted: the semantic-spatial joint vector extracts its administrative division code and spatial coupling degree label; the boundary topology vector extracts the identifier of its adjacent grid unit and the boundary length; the feature association vector extracts the type of geographical features involved and the distance relationship between features; the spatiotemporal dynamic vector extracts its data collection time and dynamic change trend (such as "increase" and "decrease"); the content of each eigenvector (such as policy keywords, indicator values) is combined with its spatial association attributes to form a "vector content-spatial attribute" key-value pair (such as "'ecological protection' policy-XX county code-high coupling degree"); all key-value pairs within the grid unit are integrated to form a structured data block, which contains the unit identifier, feature vector type statistics (such as 3 policy vectors and 2 indicator vectors), and all "vector content-spatial attribute" key-value pairs.

[0071] After integrating the data blocks for all grid cells, these blocks are arranged in order of cell identification to construct a unified dataset. A dynamic update interface is also provided for this dataset: when new feature vectors are added (e.g., real-time geographic data) or the spatial attributes of existing vectors change (e.g., policy coupling is updated), the system automatically locates the corresponding grid cell and updates its data block content. The resulting dataset is a dynamic resource pool containing spatially correlated attributes.

[0072] The present invention divides the planning area into quantifiable and manageable spatial units by establishing a standardized grid coordinate system to divide the grid into partition units, ensuring that the spatial positioning of the feature vector is accurate to a specific grid, thus solving the problem of fuzzy feature distribution in large areas. At the same time, the unified grid identification (such as "row number-column number") provides a standardized spatial index for subsequent data association and clustering, improving the granularity and efficiency of spatial analysis. The spatial distribution density of four types of feature vectors is calculated to quantify the degree of aggregation of different features in the planning area (such as policy keyword-intensive areas and geographical element-intensive areas). The introduction of density values ​​can also distinguish the primary and secondary relationships of features, avoid irrelevant or sparse features from interfering with subsequent analysis, and enhance the targeted nature of feature screening. Density-based spatial clustering automatically identifies the clustering areas of feature vectors and classifies grid units with similar features into one category (such as "policy-ecological element high coupling area"), revealing hidden spatial association patterns within the planning area (such as a region with both dense ecological policies and forest elements).

[0073] By accurately mapping the feature vectors to the grid cells of the cluster area to which they belong, a triple binding of "feature-space-cluster" is achieved (for example, an ecological policy vector is bound to a specific grid in the "ecological cluster area"). This mapping ensures that each feature has a clear spatial affiliation, avoiding the problem of disconnection between features and regions. By associating feature vectors with spatial attributes and integrating them into a dynamic material pool, the spatial integration of multi-source data is achieved (for example, a grid cell simultaneously contains policy clauses, geographic coordinates, real-time indicators, and their relationships). The "dynamic" nature of the material pool supports real-time updates (for example, newly added data is automatically assigned to the corresponding grid), ensuring that planning materials always reflect the latest situation; the "spatial association attributes" provide rich spatial arguments for subsequent outline generation and text writing, closely integrating the content of the planning text with regional reality and improving its pertinence.

[0074] In a preferred embodiment of the present invention, step 4, generating a schema structure spatial reference schema based on the spatial association attributes, includes: Step 41: extracting spatial correlation attributes of all gridded partition units in the dynamic material pool, and identifying the hierarchical relationship and spatial dependency between the attributes; Step 42: construct a tree topology structure with the administrative division unit as the root node and the grid division unit as the child node based on the hierarchical relationship and spatial dependency relationship; Step 43, based on the hierarchical path of the tree topology structure, mapping is performed to the chapter hierarchical relationship of the outline: the root node corresponds to the chapter title of the planned text, and the child nodes generate secondary chapter titles according to the spatial dependency strength; Step 44, using the writing topic as a constraint, adjust the semantic expression of the chapter title to match the topic; Step 45 : Integrate the hierarchical chapter titles and spatial dependency tags to generate the spatial reference outline.

[0075] In the embodiment of the present invention, the specific implementation process of the above step 41 is as follows: First, we traverse all gridded cells in the dynamic resource pool (e.g., "G-1-1," "G-2-3," etc.) and extract spatially relevant attributes for each cell. These attributes include: the gridded cell's administrative division code (e.g., "320102" represents a district), the identifiers of adjacent grid cells (e.g., "G-1-1's adjacent cells are G-1-2 and G-2-1"), the spatial coupling of feature vectors (e.g., "high policy-geography coupling"), geographic feature types (e.g., "forestland," "river"), and spatiotemporal trends (e.g., "ecological indicator growth"). Next, we categorize and organize all extracted attributes: attributes related to administrative division levels (e.g., provincial, municipal, and county codes) are classified as "hierarchical attributes," while attributes related to inter-grid cell interactions (e.g., "water conservation in upstream grid G-3-2 affects irrigation conditions in downstream grid G-3-3," "forest cover in the eastern grid affects windbreaks and sand fixation in the western grid") are classified as "spatial dependency attributes." Then, identify the hierarchical relationship: by comparing the hierarchy of administrative division codes (such as the first two digits of provincial codes are the same and the first four digits of municipal codes are the same), determine the superior-subordinate relationship between attributes (such as the municipal unit corresponding to "320100" is the superior of the district unit corresponding to "320102"), and form a hierarchical chain of "province-city-county-township-grid".

[0076] Finally, identify spatial dependencies: analyze the direction and intensity of influence of different grid cells in the “spatial dependency attribute class” (e.g., judging by the coupling value of the eigenvector, a coupling degree >80% is “strong dependence,” 50%-80% is “medium dependence,” and <50% is “weak dependence”), clarify “who affects whom” and the degree of influence (e.g., “grid G-5-4 (reservoir area) has a strong dependence on grid G-5-5 (irrigation area)”).

[0077] The specific implementation process of the above step 42 is as follows: Based on the hierarchical relationship identified in step 41, first determine the root node of the tree topology structure: select the highest-level administrative division unit corresponding to the planning area (such as "XX City", whose administrative division code is "320100") as the root node, representing the overall spatial scope of the planning text; all county-level administrative division units under the root node (XX City) are direct child nodes of the root node, and each county-level unit corresponds to a child node.

[0078] Next, the sub-nodes are refined based on the spatial dependency relationship: the grid partitioning units contained in each county-level unit are divided into sub-nodes based on the "spatial dependency intensity", and the grid units that are strongly dependent on the core characteristics of the county-level unit (such as "XX County is mainly agricultural") (such as grids with a high proportion of cultivated land) are prioritized as the direct subordinate sub-nodes of the county-level sub-node; the grid units with medium or weak dependency around these grid units are then used as the next-level sub-nodes, and so on.

[0079] At the same time, the relationships between each node in the tree structure are marked: the root node and county-level subnodes are marked with "inclusion relationship"; the county-level subnodes and grid subnodes are marked with "spatial subordination + dependency intensity" (for example, "XX County contains grid G-2-3, with strong dependency intensity"); and the grid subnodes are marked with "adjacent dependency + influence direction" (for example, "G-2-3 is adjacent to G-2-4, and G-2-3 depends on G-2-4 for water supply"). Ultimately, a tree-like topology is formed, with administrative divisions at the top level and grid units at the bottom level, containing both hierarchical and spatial dependency relationships.

[0080] The specific implementation process of the above step 43 is as follows: Starting from the root node, record the complete path from each node to the root node (such as "root node (XX City) - child node (XX County) - child node (Grid G-2-3)"). Each path corresponds to the potential hierarchy of "chapter-chapter-section" in the outline; the root node (XX City) corresponds to the general title of the planning text (such as "XX City Natural Resources Planning"), which serves as the highest level of the outline; then, sort the direct child nodes of the root node (county-level administrative division units) according to "spatial importance" (judged by the sum of the feature weights of the grid units, the higher the sum, the more important) to generate the first-level chapter titles: the most important county-level unit corresponds to "Chapter 1 XX County Resource Status and Planning", the second most important corresponds to "Chapter 2 XX District Resource Status and Planning", etc., which serve as the second-level level of the outline.

[0081] Then, sort the grid subnodes under the county-level subnodes according to the "spatial dependence intensity" identified in step 41 (strong dependence > medium dependence > weak dependence) to generate secondary chapter titles: each grid subnode cluster (such as a strongly dependent grid group) corresponds to the subsection title under the first-level chapter (such as "1.1XX County Core Agricultural Area (Grid G-2-3, G-2-4) Planning" and "1.2XX County Marginal Ecological Area (Grid G-2-5, G-2-6) Planning"), as the third and lower levels of the outline; finally, ensure that the hierarchical relationship of each chapter title strictly corresponds to the path of the tree topology structure (such as the path "XX City-XX County-Grid G-2-3" corresponds to "General Title-Chapter 1-Section 1.1"), forming a preliminary hierarchical outline framework.

[0082] The specific implementation process of the above step 44 is as follows: Extract the writing topic input by the user (such as "XX City Ecological Protection and Restoration Plan" and "XX River Basin Water Resources Utilization Plan"), disassemble the core keywords of the topic (such as "ecological protection", "restoration", and "water resources utilization"), and use them as constraints for semantic adjustment; traverse the preliminary chapter titles generated in step 43, and check whether the semantics of the titles match the topic keywords one by one: if the title does not contain the core words of the topic (such as "Chapter 1 XX County Resource Status and Planning" does not reflect "ecological protection"), then supplement the title semantically (such as adjusting it to "Chapter 1 XX County Ecological Resource Status and Protection Plan"); if the expression of the chapter title is inconsistent with the theme direction (such as the theme is "Water Resources Utilization" and the title is "XX Regional Forest Development Plan"), then modify the core action words of the title (such as changing "development" to "water resources supporting forest maintenance") to ensure that the title direction fits the theme.

[0083] For titles involving spatial dependencies (such as "1.1XX County Core Agricultural Area Planning"), the description is refined in combination with the theme: if the theme is "ecological protection", it is adjusted to "1.1XX County Core Agricultural Area Ecological Protection and Cultivated Land Restoration Planning"; if the theme is "water resource utilization", it is adjusted to "1.1XX County Core Agricultural Area Water Resource Efficient Utilization Planning"; finally, all adjusted chapter titles are checked for semantic consistency to ensure that the expression style of titles at the same level is unified (such as all adopting the "region + theme action + planning" structure) and the overall structure is highly consistent with the writing theme.

[0084] The specific implementation process of the above step 45 is as follows: First, collect all the hierarchical chapter titles adjusted in step 44 and arrange them in hierarchical order (main title > chapter title > section title > subsection title); then, attach the spatial dependency relationship label identified in step 41 to each chapter title: the chapter title corresponds to the relationship label between the county-level units it contains and the root node (such as "Chapter 1 Label: XX County and XX City are in an inclusion relationship, with a strong association with ecological protection"); the section title corresponds to the relationship label between the grid units it contains (such as "Section 1.1 Label: Grids G-2-3 and G-2-4 are adjacent and strongly dependent, with a water conservation association"); then, bind the chapter title with the corresponding spatial dependency relationship label to form a "title + label" structured entry (such as "1.1XX County Core Ecological Zone Protection Plan [Label: Grids G-2-3 and G-2-4 are adjacent and strongly dependent, with a water conservation association]").

[0085] Finally, all structured entries are integrated according to the hierarchical structure of the outline to generate a complete document containing the main title, chapter titles, section titles, subsection titles, and corresponding spatial dependency labels. This is the spatial reference outline. This outline not only clarifies the chapter structure of the text but also marks the spatial association logic corresponding to each chapter, providing a spatial reference basis for subsequent text generation.

[0086] This invention ensures that the outline structure is consistent with spatial laws. By mining spatial correlation attributes, the chapter arrangement is tailored to the spatial distribution of natural resources (such as watersheds and terrain divisions), avoiding the problem of traditional outlines that "emphasizes text over space"; strengthens the matching between chapters and regions, and the tree-like topological structure and spatial dependency labels allow each chapter to accurately correspond to a specific region and core features, reducing the disconnection between content and planning areas; highlights regional core issues, generates secondary titles according to the intensity of spatial dependency, and can prioritize key areas (such as ecologically sensitive areas) to avoid content generalization; improves the fit of themes, constrains the semantics of titles with writing themes, ensures that the overall direction of the outline is consistent with user needs, enhances the pertinence of planning texts, and provides clear guidance for subsequent writing. The integrated hierarchical titles and spatial labels allow the main text to be created accurately through the corresponding regional materials to ensure the logical coherence of the text.

[0087] In a preferred embodiment of the present invention, step 5, based on the spatial reference outline generated in step 4, adopts a segmented writing strategy to locate the content of the material pool, and generates body paragraphs based on the feature weights of the partition units, including: Step 51, parsing the chapter level of the spatial reference outline generated in step 45, and locating the gridded partition unit set corresponding to the current chapter title; Step 52, dynamically calculating the feature weight of each gridded partition unit associated with the current chapter based on the feature vector spatial distribution density calculated in step 36; Step 53: Select target partition units in descending order of feature weights, and extract feature vectors and spatial correlation attributes of corresponding units from the dynamic material pool; Step 54: Generate the core sentence of the paragraph based on the extracted spatial correlation attributes, and fuse the semantic-spatial joint features in the feature vector to generate description details; Step 55: Control the paragraph length and professional term density according to the hierarchical depth of the chapter title to generate a body paragraph that matches the current chapter.

[0088] In the embodiment of the present invention, the specific implementation process of the above step 51 is as follows: First, the chapter hierarchy of the spatial reference outline is parsed layer by layer to identify the hierarchical relationship of the chapters in the outline (such as "Chapter 1" is a first-level chapter, "Section 1.1" is a second-level chapter, "Subsection 1.1.1" is a third-level chapter, etc.), and the title of the chapter currently being processed is determined (such as "Section 1.1 XX County Core Agricultural Area Planning"); then, the spatial dependency label attached to the current chapter title is extracted (such as the label contains "Associated grid units: G-2-3, G-2-4, G-2-5"), and according to the grid unit identifier in the label, an accurate match is performed in the grid partition unit list of the dynamic material pool.

[0089] If the current chapter is at a higher level (such as the first-level chapter "Chapter 1 Resource Status and Planning of XX County"), the grid units associated in the labels of all the sub-chapter below it are summarized to form a gridded partition unit set corresponding to the first-level chapter (such as all grids from G-2-3 to G-2-10); if it is a low-level chapter (such as a third-level subsection), the grid units explicitly associated in its label are directly extracted as the unit set corresponding to the chapter; finally, the unique identifiers of all gridded partition units in the set are recorded (such as "G-2-3, G-2-4") to complete the positioning association between the current chapter and the grid unit.

[0090] The specific implementation process of the above step 52 is as follows: First, through the spatial distribution density data of the feature vectors calculated in step 36, extract the density values ​​of the four types of feature vectors (semantic-spatial union, boundary topology, feature association, and spatiotemporal dynamics) corresponding to each gridded partition unit (such as G-2-3, G-2-4) associated with the current chapter (such as "semantic-spatial union feature density: 5.2 / square kilometer, feature association feature density: 3.8 / square kilometer", etc.).

[0091] Then, according to the theme of the current chapter (such as "agricultural area planning"), dynamic weight coefficients are assigned to the four types of feature vectors: features that are highly relevant to the theme (such as cultivated land and irrigation facilities data in the "element association characteristics" of agricultural areas) are assigned higher coefficients (such as 0.4), secondary related features (such as precipitation data in the "spatiotemporal dynamic characteristics") are assigned medium coefficients (such as 0.3), and features with lower correlation (such as "boundary topology characteristics") are assigned lower coefficients (such as 0.15), ensuring that the total coefficient is 1; then, for each gridded partition unit, the density values ​​of the four types of features are multiplied by the corresponding dynamic weight coefficients, and the products are added together to obtain the preliminary feature weight of the unit.

[0092] Finally, based on the strength of the spatial dependency of the grid cell (such as the "strong dependency" and "medium dependency" identified in step 41), the preliminary feature weights are corrected: the strongly dependent cell is multiplied by a correction factor of 1.2, the moderately dependent cell is multiplied by 1.0, and the weakly dependent cell is multiplied by 0.8 to obtain the final feature weight.

[0093] The specific implementation process of the above step 53 is as follows: First, sort the feature weights of all grid partition units associated with the current chapter calculated in step 52 in descending order; then, determine the number of selections based on the hierarchical depth of the current chapter: if it is a first-level chapter (such as "Chapter 1"), select the top 50% of the grid units with the highest weight (if there are 10 units in total, select the top 5); if it is a second-level chapter (such as "Section 1.1"), select the top 30% of the units with the highest weight; if it is a third-level chapter or below, select the top 2-3 units with the highest weight to ensure that the selected target units can cover the core content of the chapter.

[0094] Next, locate these target partition units from the dynamic material pool, extract all the feature vectors in each unit (such as "cultivated land retention policy" and "2023 grain production statistics" in the semantic-spatial joint features, "distance between irrigation channels and cultivated land" in the element association features, etc.), and the corresponding spatial association attributes (such as "G-2-3 is located in the eastern part of XX Town, adjacent to G-2-4, and is dependent on irrigation water sources", "the statistical period is 2023", etc.). Finally, the extracted feature vectors and spatial association attributes are stored by unit classification to form the writing material package for the current chapter.

[0095] The specific implementation process of the above step 54 is as follows: First, the core sentence of the paragraph is generated based on the extracted spatial correlation attributes: the information that best reflects the core characteristics of the target partition unit is selected from the spatial correlation attributes (such as "G-2-3 is the core cultivated area of ​​XX County, with a soil fertility level of one, an average annual precipitation of 800 mm, and mainly relies on the irrigation canal of G-2-4 for water supply"), and this information is condensed into a summary sentence (such as "The core cultivated area (G-2-3) of XX County has fertile soil and sufficient precipitation, and its irrigation mainly depends on the adjacent G-2-4 irrigation canal") as the core sentence of the paragraph to clarify the paragraph theme.

[0096] Then, the semantic-spatial joint features in the feature vector are fused to generate descriptive details: policy clauses related to the core sentence (such as "The XX City Cultivated Land Protection Regulations stipulate that the cultivated land reserve in this area shall not be less than 500 hectares"), statistical indicators (such as "In 2023, the cultivated land area in this area will be 520 hectares, and the grain output will reach 3,000 tons"), and geographical correlation information (such as "Cultivated land is concentrated in the range of 32°15′-32°18′ north latitude and 118°20′-118°23′ east longitude") are extracted from the semantic-spatial joint features; then, these detailed information are organized in a logical order (such as policy requirements → current data → geographical distribution) and added to the core sentence to form a complete paragraph framework.

[0097] The specific implementation process of the above step 55 is as follows: First, identify the hierarchical depth of the current chapter title: determine this by parsing the chapter number (e.g., "Chapter 1" is level 1, depth 1; "Section 1.1" is level 2, depth 2; "Subsection 1.1.1" is level 3, depth 3); then, control paragraph length based on the hierarchical depth: paragraphs in level 1 chapters should be mainly general descriptions, with a length of 150-200 words (about 3-4 sentences); paragraphs in level 2 chapters should be moderately detailed, with a length of 200-300 words (about 4-6 sentences); paragraphs in level 3 and below chapters should be elaborated in detail, with a length of 300-500 words (about 6-10 sentences), ensuring that the lower the level, the more specific the content.

[0098] At the same time, control the density of professional terms: first-level chapters use general terms (such as "resource protection" and "planning objectives") with a term density of 1-2 per 100 words; second-level chapters appropriately increase professional terms (such as "arable land reserves" and "irrigation efficiency") with a density of 3-4 per 100 words; third-level and below chapters can use subdivided terms (such as "soil organic matter content" and "drip irrigation technology coverage") with a density of 5-6 per 100 words to ensure a balance between professionalism and readability; finally, according to the above length and term density requirements, adjust the paragraph framework generated in step 54 (such as deleting redundant information, adding necessary terms or explanatory statements) to form a coherent text paragraph that matches the current chapter level, and mark the grid partition unit identifier corresponding to the paragraph (such as "[corresponding unit: G-2-3, G-2-4]").

[0099] The present invention locates the corresponding grid units by parsing the outline, ensuring that the paragraph content is strictly bound to the specific spatial scope of the planning area, avoiding the problem of "text not matching the ground" and making the text closely follow the actual situation of the region; based on the feature weight, the target units are screened, and the materials of high-weight areas (such as ecologically sensitive areas and core functional areas) are preferentially called, so that the paragraph focuses on the core issues of the key areas, avoiding the content being straightforward and without distinction between the primary and secondary; the core sentences are generated with spatial correlation attributes, and the semantic-spatial features are integrated to supplement the details, which not only ensures that the paragraph theme is clear, but also reflects the deep connection between policies, data and geographic space, and enhances the persuasiveness of the content; the paragraph length and term density are controlled according to the chapter level, with high-level chapters having strong generalization and concise terminology, and low-level chapters having detailed details and high professionalism, so that the text structure is clear and the readability and professionalism are balanced, which meets the hierarchical expression requirements of the planning text; by automatically locating materials and generating paragraphs, the cost of manual information screening and content organization is reduced, while ensuring the consistency of content with the outline and regional characteristics, and improving the accuracy and efficiency of text generation.

[0100] The specific implementation process of step 6 above is as follows: First, all text paragraphs generated in step 5 are identified for hierarchical attribution. The system iterates through each paragraph and, based on the corresponding chapter hierarchy indicated at the end of the paragraph (e.g., "[Corresponding Chapter: Section 1.1]"), assigns the paragraph to the corresponding section of the spatial reference outline, forming a "chapter-paragraph" correspondence (e.g., "Chapter 1" contains paragraphs 1 and 2, "Section 1.1" contains paragraphs 3 and 4, and so on). Next, the chapter numbering is standardized. Following the "chapter-chapter-section-subsection" hierarchy, each chapter is assigned a uniform number: chapter titles do not require numbering; chapter titles use the "Arabic numeral + ." format (e.g., "1.", "2."); section titles use the "chapter number + Arabic numeral + ." format (e.g., "1.1.", "1.2."); and subsection titles use the "section number + Arabic numeral" format (e.g., "1.1.1," "1.1.2"). Furthermore, each paragraph is numbered according to the order of its appearance within the chapter (e.g., "Section 1.1, Paragraph 1," "Section 1.1, Paragraph 2") to ensure clear and traceable hierarchy.

[0101] Then, standardize the data presentation format. Scan all paragraphs for numerical data (such as area, population, and proportions), and convert scattered textual descriptions (such as "approximately 5,000 hectares of cultivated land") into tables or charts. For comparisons of similar data (such as "Changes in cultivated land area from 2020 to 2023"), automatically generate line charts or bar charts and insert them below the corresponding paragraphs. For multi-dimensional indicator data (such as "Soil fertility, irrigation conditions, and yield of each grid unit"), generate structured tables with unified table titles formatted as "Table XX Title Content" (e.g., "Table 1.11.1: Agricultural indicators for grid units"), and reference them in paragraphs using "See Table XX for details."

[0102] At the same time, a standard terminology library for natural resource planning was used to replace non-standard expressions in paragraphs (e.g., "national land space development" was unified into "national land space development and protection pattern," and "red line zone" was clarified as "ecological protection red line zone"). For the first appearance of a professional term (e.g., "spatial coupling"), a brief annotation was added after the term (e.g., "spatial coupling: refers to the strength of the association between policy provisions and geographic coordinates") to ensure that the terminology is standardized and easy to understand. Finally, all structured chapters and paragraphs were integrated and arranged in order according to the outline hierarchy to generate a structured text with unified numbering, standardized data format, and standardized terminology. A table of contents (including chapter titles and corresponding page numbers) and an abstract (a condensed summary of the core content of each chapter, approximately 300-500 words) were also automatically generated to form a complete structured text framework.

[0103] The specific implementation process of step 7 above is as follows: First, perform policy clause matching verification. The system extracts all policy statements in the structured text and compares the policy name and clause content with the latest policy specification library in the local policy engine: check whether the policy is currently valid (such as eliminating repealed clauses); verify whether the clause references are accurate (such as whether the "cultivated land protection clause" is consistent with the original text); confirm whether the scope of application of the policy matches the planning area. If a mismatch is found (such as citing an abolished policy), the error location will be marked and a prompt will be prompted: "The "XX Measures" cited here has been repealed in 2023. It is recommended to replace it with Chapter 3 of the "XX Regulations"."

[0104] Secondly, perform spatial logic self-consistency verification. Based on the spatial association properties of the dynamic material pool, check whether the spatial relationship described in the text conforms to geographical laws: for example, if the text mentions "planning a large reservoir in XX arid area", the system will compare the spatiotemporal dynamic feature vectors of the area (such as annual average precipitation <200mm, evaporation >1500mm) to determine whether the "large reservoir" plan is inconsistent with the characteristics of the arid area. If so, it will be marked as "spatial logic anomaly: XX arid area is short of water resources, and the planning of a large reservoir may have feasibility issues." For example, check whether the description of the upstream and downstream relationship is reasonable (such as whether "upstream pollution control" is connected with "downstream water quality improvement"). If there is a logical break (such as only mentioning upstream control without mentioning downstream impact), it will prompt "It is recommended to supplement the specific impact analysis of upstream control on downstream water quality."

[0105] Next, perform a data consistency check. This process iterates through all numerical data in the text, creating a comparison table of "data item-value-location" (e.g., "cultivated land area" is "520 hectares" in Section 1.1 and "510 hectares" in Section 2.3). The consistency of the values ​​for the same data item is checked. If the deviation is within the allowable range (e.g., ±5%), the average value is automatically calculated and updated uniformly. If the deviation exceeds the range (e.g., ±10%), a flag is displayed: "Inconsistent data: The values ​​for 'cultivated land area' in Sections 1.1 and 2.3 differ by 10%. Please verify." The data units are also checked for consistency (e.g., whether "hectares" and "mu" are used interchangeably). If there is any confusion in the units, the data is automatically converted to the standard unit and a prompt is displayed.

[0106] Next, check the structured text for compliance with official formatting standards for natural resource planning documents, including chapter numbering, table and figure insertion, and glossary. For example, check whether tables include titles, units, and data sources; whether chapter numbering is skipped or repeated; and whether the abstract is placed at the beginning of the text. If formatting issues are found (e.g., a table lacking a data source), mark it as "Table 1.1 is missing a data source. It is recommended to add 'Data source: XX Municipal Bureau of Statistics 2023 Report'."

[0107] Finally, all verification results are compiled into a Verification Report, which includes error type, location, and suggested corrections. Standardizable errors (such as unit conversions and term substitutions) are automatically corrected based on the Verification Report. For errors requiring manual judgment (such as policy applicability), the user is prompted to confirm the changes. After the user confirms the changes, the system re-verifies until all errors are corrected (error rate ≤ 3%). The final output is a plan document that meets policy requirements, demonstrates spatial logic consistency, accurate data, and standardized formatting.

[0108] like Figure 2 As shown, a natural resource planning text generation system includes: a receiving module for receiving standardized data input by a user, wherein the standardized data includes mandatory items and optional items, wherein the mandatory items include title, region, and writing subject, and the optional items include outline templates and unstructured reference materials; The execution module is used to make resource scheduling decisions based on the presence or absence of optional items. If an optional item exists, it uses a large language model to parse reference materials and extract keywords. If an optional item does not exist, it triggers local policy compliance analysis and platform database matching. A generation module is used to convert keyword sets, user knowledge base data, platform template library data, and online search data into multidimensional feature vectors; construct a feature space covering the planning area based on the feature vectors; establish gridded partition units within the feature space and perform spatial clustering based on the distribution density of the feature vectors; map the feature vectors to corresponding partition units to generate a dynamic material pool containing spatial correlation attributes; A processing module is used to generate an outline structure spatial reference outline based on the spatial association attributes; according to the spatial reference outline, a segmented writing strategy is adopted to locate the content of the material pool, and the main text paragraphs are generated based on the feature weights of the partition units; structured processing is performed on the main text paragraphs to obtain structured text; dynamic verification is performed on the structured text to output the final planning text.

[0109] A computer-readable storage medium stores a program, which implements the method when executed by a processor.

[0110] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for generating natural resource planning text, characterized in that: The method comprises: Step 1: receiving standardized data input by a user, wherein the standardized data includes mandatory items and optional items, wherein the mandatory items include title, region, and writing subject, and the optional items include outline templates and unstructured reference materials; Step 2: Execute resource scheduling decisions based on whether the optional items exist. If the optional items exist, a large language model is used to parse reference materials and extract keywords. If the optional items do not exist, local policy compliance analysis and platform database matching are triggered. Step 3: Convert the keyword set, user knowledge base data, platform template library data, and online search data output from step 2 into a multidimensional feature vector; construct a feature space covering the planning area based on the feature vector; establish gridded partition units within the feature space, and perform spatial clustering based on the distribution density of the feature vectors; map the feature vectors to the corresponding partition units to generate a dynamic material pool containing spatial correlation attributes; Step 4, generating an outline structure space reference outline based on the spatial association attributes; Step 5: Based on the spatial reference outline generated in step 4, a segmented writing strategy is used to locate the content of the material pool and generate body paragraphs based on the feature weights of the partition units; Step 6, performing structured processing on the body paragraph generated in step 5 to obtain a structured text; Step 7: Perform dynamic verification on the structured text to output the final planning text.

2. A method for generating natural resource planning text according to claim 1, characterized in that: Step 2: Execute resource scheduling decisions based on whether the optional items exist. If the optional items exist, extract keywords by parsing reference materials using a large language model. If the optional items do not exist, local policy compliance analysis and platform database matching are triggered, including: Step 21, detecting the presence of optional items in the user input. If the outline template or unstructured reference material is detected, step 22 is executed; if no optional items are detected, step 23 is executed; Step 22: Perform structured parsing using a large language model to identify multimodal features of unstructured reference materials, extracting three core keywords: policy terms, geographic coordinates, and statistical indicators. Spatial topological relationships are then matched against the user's private knowledge base to generate a keyword set with spatial labels. Step 23: Trigger the local policy engine analysis. Based on the required regional attributes, load the policy specification library for the corresponding administrative division, perform timeliness verification and conflict detection on the policy clauses, and output a compliance constraint set. Step 24, perform dynamic matching of the platform database, search the platform template library with the writing topic as the index, match historical planning cases with regional feature similarity higher than the threshold; extract the spatial element association specifications in the case and generate a standardized keyword set.

3. A method for generating natural resource planning text according to claim 2, characterized in that: Step 22: Use a large language model to perform structured parsing and perform multimodal feature recognition on the unstructured reference materials to extract three core keywords: policy terms, geographic coordinates, and statistical indicators. These keywords include: Step 221 , performing format standardization preprocessing on the unstructured reference materials, eliminating file format differences and unifying the encoding to generate a parseable continuous text stream; Step 222 , based on the continuous text stream, performing data modality segmentation by a multimodal feature recognition engine: identifying and separating text description segments, image data blocks, and table structured data; Step 223: Based on the text description segment separated in step 222, semantic role labeling technology is used to parse the sentence components: the responsible party in the policy text is identified as the first category of keywords, the obligation item as the second category of keywords, and the time constraint as the third category of keywords, thereby generating a structured policy clause set. Step 224 , based on the image data blocks separated in step 222 , identify the text of the annotation elements in the map using OCR technology, convert the element location information into a unified projection coordinate system, and generate a standard geographic coordinate set with coordinate system labels; Step 225 , based on the tabular structured data separated in step 222 , verifies the unit consistency of the numerical data, extracts the indicator name field, numerical field and statistical period field, and generates a standardized statistical indicator set.

4. A method for generating natural resource planning text according to claim 3, characterized in that: The keyword set includes: Step 226: Establish a spatial reference coordinate system based on the administrative division vector boundary data in the user's private knowledge base; Step 227: Based on the spatial reference coordinate system, perform spatial weight matching on the structured policy clause set, extract the geographic entity names in the policy clauses, and perform semantic association of the geographical names with the administrative division vector boundaries to obtain association results; based on the association results, map the policy clauses to corresponding administrative division units to generate a policy clause label set with administrative division codes; Step 228: Based on the spatial reference coordinate system, a topological relationship calculation is performed on the standard geographic coordinate set, and spatial location inclusion is determined between the geographic coordinate point and the administrative division vector boundary. When the coordinate point is located in the intersection area of ​​multiple administrative divisions, the topological distance between the coordinate point and each boundary is calculated and a membership weight is assigned to generate a coordinate label set with topological weights. Step 229: Perform spatial scale conversion on the standardized statistical indicator set to identify the spatial scale attributes of the statistical indicators, perform spatial interpolation on the indicator values ​​at the non-administrative division scale according to population density, convert the statistical indicator values ​​to generate administrative division unit granularity, and attach administrative division code labels. Step 230: Fusing the policy clause label set, the coordinate label set, and the administrative division code label set to construct a spatial topological relationship matrix indexed by administrative division units. Step 231 : Generate a keyword set with spatial topology tags based on the spatial coupling degree of policy terms, geographic coordinates and statistical indicators in the spatial topology relationship matrix.

5. A method for generating natural resource planning text according to claim 4, characterized in that: Step 3: Convert the keyword set, user knowledge base data, platform template library data, and online search data output in step 2 into a multidimensional feature vector; Constructing a feature space covering the planning area based on the feature vectors, including: Step 31, receiving the keyword set with spatial topology tags output in step 2, the administrative division boundary data in the user knowledge base, the historical case spatial element data matched in the platform template library, and the real-time geographic data obtained by online search; Step 32, respectively converting the keyword set with spatial topology labels, administrative division boundary data, historical case spatial element data, and real-time geographic data into semantic-spatial joint feature vectors, boundary topology feature vectors, element association feature vectors, and spatiotemporal dynamic feature vectors; Step 33, taking the spatial scope defined by the administrative division boundary of the planning area as a reference; Step 34 : constructing a multi-dimensional feature space covering the spatial range of the planning area based on the semantic-spatial joint feature vector, the boundary topology feature vector, the element association feature vector, and the spatiotemporal dynamic feature vector.

6. A method for generating natural resource planning text according to claim 5, characterized in that: Establish gridded partition units in the feature space and perform spatial clustering based on the distribution density of feature vectors; Map the feature vectors to the corresponding partition units to generate a dynamic material pool containing spatial correlation attributes, including: Step 35: Based on the multi-dimensional feature space constructed in step 34, a standard grid coordinate system covering the spatial range of the planning area is established to form a plurality of gridded partition units; Step 36, calculating the spatial distribution density of the semantic-spatial joint feature vector, the boundary topology feature vector, the element association feature vector, and the spatiotemporal dynamic feature vector generated in step 32 in the multidimensional feature space; Step 37, performing a spatial clustering operation on all feature vectors in the multidimensional feature space according to the spatial distribution density calculated in step 36, and identifying clustered areas of the feature vectors in the gridded partition units; Step 38, mapping each eigenvector generated in step 32 to a partition unit corresponding to the aggregation area to which it belongs in the gridded partition unit established in step 35; Step 39, in each gridded partition unit, associate the feature vector mapped to the partition unit in step 38 and its corresponding spatial correlation attribute; integrate the associated feature vectors and spatial correlation attributes in all gridded partition units to generate a dynamic material pool containing spatial correlation attributes.

7. A method for generating natural resource planning text according to claim 6, characterized in that: Step 4, generating an outline structure spatial reference outline based on the spatial association attributes, including: Step 41: extracting spatial correlation attributes of all gridded partition units in the dynamic material pool, and identifying the hierarchical relationship and spatial dependency between the attributes; Step 42: construct a tree topology structure with the administrative division unit as the root node and the grid division unit as the child node based on the hierarchical relationship and spatial dependency relationship; Step 43, based on the hierarchical path of the tree topology structure, mapping is performed to the chapter hierarchical relationship of the outline: the root node corresponds to the chapter title of the planned text, and the child nodes generate secondary chapter titles according to the spatial dependency strength; Step 44, using the writing topic as a constraint, adjust the semantic expression of the chapter title to match the topic; Step 45 : Integrate the hierarchical chapter titles and spatial dependency tags to generate the spatial reference outline.

8. A method for generating natural resource planning text according to claim 7, characterized in that: Step 5: Based on the spatial reference outline generated in step 4, use a segmented writing strategy to locate the content of the material pool and generate body paragraphs based on the feature weights of the partition units, including: Step 51, parsing the chapter level of the spatial reference outline generated in step 45, and locating the gridded partition unit set corresponding to the current chapter title; Step 52, dynamically calculating the feature weight of each gridded partition unit associated with the current chapter based on the feature vector spatial distribution density calculated in step 36; Step 53: Select target partition units in descending order of feature weights, and extract feature vectors and spatial correlation attributes of corresponding units from the dynamic material pool; Step 54: Generate the core sentence of the paragraph based on the extracted spatial correlation attributes, and fuse the semantic-spatial joint features in the feature vector to generate description details; Step 55: Control the paragraph length and professional term density according to the hierarchical depth of the chapter title to generate a body paragraph that matches the current chapter.

9. A natural resource planning text generation system, characterized in that: The system is used to perform the method according to any one of claims 1 to 8, comprising: a receiving module for receiving standardized data input by a user, wherein the standardized data includes mandatory items and optional items, wherein the mandatory items include title, region, and writing subject, and the optional items include outline templates and unstructured reference materials; The execution module is used to make resource scheduling decisions based on the presence or absence of optional items. If an optional item exists, it uses a large language model to parse reference materials and extract keywords. If an optional item does not exist, it triggers local policy compliance analysis and platform database matching. A generation module is used to convert keyword sets, user knowledge base data, platform template library data, and online search data into multidimensional feature vectors; construct a feature space covering the planning area based on the feature vectors; establish gridded partition units within the feature space and perform spatial clustering based on the distribution density of the feature vectors; map the feature vectors to corresponding partition units to generate a dynamic material pool containing spatial correlation attributes; A processing module is used to generate an outline structure spatial reference outline based on the spatial association attributes; according to the spatial reference outline, a segmented writing strategy is adopted to locate the content of the material pool, and the main text paragraphs are generated based on the feature weights of the partition units; structured processing is performed on the main text paragraphs to obtain structured text; dynamic verification is performed on the structured text to output the final planning text.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which implements the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Geographic grid-based government affair information resource integration method

    CN107092680A

  • Geographic information system (GIS)-based territorial space planning method and system

    CN117312476A

  • Land space planning data scheduling optimization method and system based on supervised learning

    CN117973794A

  • Industrial land intelligent site selection method based on spatial big data and artificial intelligence big model

    CN120495047A

  • Document structure data base construction processing system

    JP1994052162A

Cited By

  • Man-machine collaborative creation method and system based on multi-modal fusion

    CN121030691A