A natural resource planning text generation method and system, and a storage medium

By parsing unstructured data and constructing a feature space using a large-scale language model, natural resource planning texts are generated. This solves the problems of low data parsing efficiency and lack of spatial correlation, and achieves efficient and compliant planning text generation.

CN120706400BActive Publication Date: 2025-11-28ZHEJIANG WANWEI SPACE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511201782.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-28
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing methods for generating natural resource planning texts suffer from low data parsing efficiency and a lack of cross-modal spatial correlation, leading to spatial logical contradictions in the generated content.

Method used

By receiving standardized data, using a large language model to parse unstructured reference materials, constructing a feature space covering the planning area and dividing it into grids, generating a dynamic material pool containing spatial association attributes, generating an outline structure based on spatial association attributes, generating text paragraphs by combining the feature weights of the partition units, and performing structured processing and dynamic verification.

Benefits of technology

It improves data processing efficiency, ensures that the generated planning text conforms to the spatial distribution patterns of natural resource elements, avoids spatial logical contradictions, and enhances the timeliness and compliance of the text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706400B_ABST
    Figure CN120706400B_ABST
Patent Text Reader

Abstract

The application provides a natural resource planning text generation method and system and a storage medium, relates to the technical field of data processing, and comprises the following steps: establishing a grid partition unit in a feature space, performing spatial clustering according to feature vector distribution density; mapping the feature vector to a corresponding partition unit to generate a dynamic material pool containing spatial correlation attributes; generating an outline structure space reference outline based on the spatial correlation attributes; positioning the material pool content by adopting a segmented writing strategy according to the spatial reference outline, generating a text paragraph based on the feature weight of the partition unit; performing structured processing on the text paragraph to obtain a structured processed text; and performing dynamic checking on the structured processed text to output a final planning text. The spatial clustering mapping reduces the processing amount of redundant data and reduces the computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a natural resource planning text generation method and system and a storage medium. BACKGROUND

[0002] In the field of natural resource management, the compilation of planning texts (such as land use overall planning and ecological protection and restoration schemes) is an important carrier for policy implementation. Traditional text generation methods have the following problems:

[0003] For example, planning needs to integrate administrative division data, policies and regulations, geographic spatial information and social and economic statistical indicators. Existing methods rely on manual extraction of unstructured data (such as PDF policy documents, scanned maps and Excel tables), resulting in low data analysis efficiency and lack of cross-modal spatial correlation (such as the inability to automatically associate a specific policy provision with the jurisdiction of a specific geographic coordinate).

[0004] For example, the text structure needs to reflect the spatial distribution of natural resource elements (such as the need to associate upstream and downstream grid cells in the chapter on watershed ecological restoration). Mainstream text generation models only rely on semantic features and do not establish a mapping mechanism between chapter structure and spatial zoning units, resulting in spatial logic contradictions in the generated content (such as planning water conservation forest construction in arid regions). SUMMARY

[0005] The technical problem to be solved by the present application is to provide a natural resource planning text generation method, system and storage medium. Spatial clustering mapping reduces the amount of redundant data processing and reduces the use of computing resources.

[0006] To solve the above technical problems, the technical solution of the present application is as follows:

[0007] In a first aspect, a natural resource planning text generation method is provided, the method comprising:

[0008] Step 1: receiving user input of standardized data, the standardized data including mandatory items and optional items, wherein the mandatory items include title, region and writing theme, and the optional items include outline template and unstructured reference materials;

[0009] Step 2: performing resource scheduling decisions according to the presence or absence of optional items. If optional items are present, key words are extracted from reference materials by a large language model. If optional items are not present, local policy compliance analysis and platform database matching are triggered;

[0010] Step 3, converting the keyword set, user knowledge base data, platform template library data, and network search data output by step 2 into a multi-dimensional feature vector; constructing a feature space covering the planning area based on the feature vector; establishing a grid partition unit in the feature space, and performing spatial clustering according to the feature vector distribution density; mapping the feature vector to the corresponding partition unit to generate a dynamic material pool containing spatial correlation attributes;

[0011] Step 4, generating an outline structure space reference outline based on the spatial correlation attributes;

[0012] Step 5, according to the spatial reference outline generated in step 4, using a segmented writing strategy to locate the content of the material pool, and generating a text paragraph based on the feature weight of the partition unit;

[0013] Step 6, performing structured processing on the text paragraph generated in step 5 to obtain a structured processed text;

[0014] Step 7, performing dynamic verification on the structured processed text to output the final planning text.

[0015] The second aspect is a natural resource planning text generation system, comprising:

[0016] A receiving module for receiving user input standardization data, the standardization data including mandatory items and optional items, wherein the mandatory items include title, region, and writing theme, and the optional items include outline template and unstructured reference materials;

[0017] An execution module for executing resource scheduling decisions according to the presence or absence of optional items, and if the optional items are present, extracting keywords by analyzing reference materials through a large language model, and if the optional items are not present, triggering local policy compliance analysis and platform database matching;

[0018] A generation module for converting the keyword set, user knowledge base data, platform template library data, and network search data into a multi-dimensional feature vector; constructing a feature space covering the planning area based on the feature vector; establishing a grid partition unit in the feature space, and performing spatial clustering according to the feature vector distribution density; mapping the feature vector to the corresponding partition unit to generate a dynamic material pool containing spatial correlation attributes;

[0019] A processing module for generating an outline structure space reference outline based on the spatial correlation attributes; according to the spatial reference outline, using a segmented writing strategy to locate the content of the material pool, and generating a text paragraph based on the feature weight of the partition unit; performing structured processing on the text paragraph to obtain a structured processed text; and performing dynamic verification on the structured processed text to output the final planning text.

[0020] In a third aspect, a computer readable storage medium having stored therein a program, which, when executed by a processor, implements the method.

[0021] The above scheme of the present application at least includes the following beneficial effects:

[0022] The unstructured reference materials in the optional items are automatically parsed and keywords are extracted by the large language model, replacing the manual processing process. This mechanism reduces the time cost and labor input of data parsing, while avoiding the omissions that may be caused by manual operation, thereby improving the efficiency of pre-processing of planning text generation.

[0023] By converting standardized data, user knowledge base, platform template library and network data into multi-dimensional feature vectors, a feature space covering the planning area is constructed, and the feature vectors are mapped to specific partition units through gridding partition and spatial clustering, and finally a dynamic material pool containing spatial correlation attributes is generated, realizing the accurate correlation of cross-modal data such as unstructured materials, policy data and geographic information and spatial units.

[0024] Based on the spatial correlation attributes, an outline structure is generated, and the feature weights of the partition units are combined to generate text paragraphs, forming a complete mapping chain of "spatial partition unit-outline structure-text content". For example, the chapter of watershed ecological restoration can automatically associate the features (such as water quantity and vegetation coverage) of the upstream and downstream grid units through this mapping, ensuring that the generated content conforms to the spatial distribution law of natural resource elements, and fundamentally avoiding spatial logical contradictions.

[0025] When there are no optional items, local policy compliance analysis and platform database matching are triggered, combined with dynamic verification to ensure that the generated text strictly follows local policies and regulations and planning standards. At the same time, the dynamic material pool can be adjusted in real time with data updates (such as new policy documents and real-time statistical data), so that the planning text can adapt to the dynamic changes of policies and data in natural resource management, improving the timeliness and compliance of the text. BRIEF DESCRIPTION OF DRAWINGS

[0026] Figure 1 is a flowchart of a natural resource planning text generation method provided by an embodiment of the present application.

[0027] Figure 2 is a schematic diagram of a natural resource planning text generation system provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] Exemplary embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is to be understood that the present disclosure can be embodied in various forms without being limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0029] As Figure 1 shown, an embodiment of the present disclosure proposes a natural resource planning text generation method, the method comprising the following steps:

[0030] Step 1, receiving user input standardized data, the standardized data including mandatory items and optional items, wherein the mandatory items include title, region and writing theme, and the optional items include outline template and unstructured reference materials;

[0031] Step 2, performing resource scheduling decision according to the presence or absence of optional items, if the optional items exist, extracting keywords by analyzing reference materials through large language model, if the optional items do not exist, triggering local policy compliance analysis and platform database matching;

[0032] Step 3, converting the keyword set output by step 2, user knowledge base data, platform template library data and network search data into a multi-dimensional feature vector; constructing a feature space covering the planning area based on the feature vector; establishing a grid partition unit in the feature space, and performing spatial clustering according to the feature vector distribution density; mapping the feature vector to the corresponding partition unit to generate a dynamic material pool containing spatial correlation attributes;

[0033] Step 4, generating an outline structure based on the spatial correlation attributes to generate a spatial reference outline;

[0034] Step 5, according to the spatial reference outline generated in step 4, adopting a segmented writing strategy to locate the content of the material pool, and generating text paragraphs based on the feature weight of the partition unit;

[0035] Step 6, performing structured processing on the text paragraphs generated in step 5 to obtain the text after structured processing;

[0036] Step 7, performing dynamic verification on the text after structured processing to output the final planning text.

[0037] In the embodiments of the present application, by determining the distinction between mandatory items (title, region, writing theme) and optional items (outline template, unstructured reference materials), it is ensured that the user input data meets the basic requirements of the planning text generation, avoiding subsequent processing errors caused by data format confusion; the mandatory items directly lock the core elements of the planning text (such as the title, region and theme of “XX City Ecological Protection Planning”), and the optional items allow users to flexibly supplement personalized needs (such as customizing the outline or reference materials), which not only ensures the accuracy of the generation direction, but also retains the customization space, improving the matching degree of user needs and the final text.

[0038] When there are optional items, the unstructured reference materials (such as policy PDFs, report documents) are automatically parsed by large language models and keywords are extracted, replacing manual processing, with efficiency improved by more than 50%, and avoiding the subjective bias of manual extraction; when there are no optional items, local policy compliance analysis and platform database matching are triggered to ensure that the text generation is strictly anchored to the local policy framework (such as the Land Management Law, Ecological Red Line Regulations), ensuring compliance from the source; by judging whether the optional items exist, dynamically selecting the “external model analysis” or “local database matching” path, avoiding invalid calculations, reducing system energy consumption, and improving processing speed.

[0039] The keyword set, knowledge base, template library and network data are converted into multi-dimensional feature vectors (covering policy, geography, economy and other dimensions), through grid partitioning and spatial clustering, the precise binding of unstructured text, statistical data and geographic coordinates (such as a certain policy clause associated with XX District XX Street) is achieved, solving the problem of “policy and space disconnection” in traditional methods; the dynamic material pool can absorb new data in real time (such as newly added remote sensing images, quarterly economic data), and automatically update the partition unit attributes through feature vector mapping, ensuring that the materials always reflect the latest situation, avoiding the decision-making bias caused by data lag in the planning text; the grid partition unit (such as a 1km x 1km grid) makes spatial feature analysis more precise, for example, in agricultural planning, the soil fertility, irrigation conditions and planting policies of each grid can be accurately associated, providing micro-data support for subsequent content generation.

[0040] The mapping relationship of "spatial logic-text structure" is established: the outline no longer depends on semantic logic (such as "current situation-problem-countermeasure"), but integrates spatial correlation attributes (such as "upstream protected area-middle reach management area-downstream monitoring area"), ensures that the chapter structure is highly consistent with the spatial distribution law of natural resources (such as the flow direction of the river basin and the terrain partition), and avoids the defects of the traditional outline "paying more attention to text than space". The spatial reference outline determines the geographical partition and core features of each chapter (such as "Chapter 3 focuses on XX ecological red line area, and highlights soil erosion data"), reducing the problem of mismatch between paragraph content and planning area. Through the feature weight of the partition unit (such as the high weight of the ecological sensitive area), the material of the high weight unit (such as the restoration technology and data of the ecological fragile area) is automatically prioritized, ensuring that the content highlights the core problems of the area (such as prioritizing the description of wind and sand prevention measures in desertification areas), and avoiding generalization.

[0041] In a preferred embodiment of the present application, step 2, according to the existence of the optional item, if the optional item exists, the key words are extracted by analyzing the reference materials through the large language model, and if the optional item does not exist, the local policy compliance analysis and platform database matching are triggered, including:

[0042] Step 21, detecting the existence of the optional item in the user input, if the outline template or unstructured reference material exists, executing step 22; if the optional item is not detected, executing step 23;

[0043] Step 22, performing structured analysis through a large language model, identifying multi-modal features of unstructured reference materials, extracting three types of core keywords: policy provisions, geographic coordinates, and statistical indicators; matching the keywords with the user's private knowledge base to generate a set of keywords with spatial tags;

[0044] Step 23, triggering local policy engine analysis, loading the policy specification library of the corresponding administrative division according to the regional attributes in the mandatory item, performing time effectiveness verification and conflict detection of policy provisions, and outputting a compliance constraint set;

[0045] Step 24, performing platform database dynamic matching, retrieving the platform template library with the writing theme as the index, matching historical planning cases with a regional feature similarity higher than a threshold; extracting spatial element correlation specifications in the cases to generate a standardized keyword set.

[0046] In the embodiment of the present application, the specific implementation process of step 21 is as follows:

[0047] The standardized data submitted by the user is field parsed to identify the fields that are "optional items" (i.e., outline templates, unstructured reference materials); it is checked whether the outline template contains a valid text structure (such as chapter hierarchy, title list), and whether the unstructured reference materials contain a resolvable file (such as PDF, document, picture, etc.) or text content; if at least one of the outline template or the unstructured reference materials contains valid content, it is determined that "optional items exist", triggering step 22; if both are null or invalid content, it is determined that "optional items do not exist", triggering step 23.

[0048] The specific implementation process of the above step 23 is as follows:

[0049] The "region attribute" information (such as XX province XX city, XX county administrative division code or name) is extracted from the mandatory items and transmitted to the local policy engine; the local policy engine locates the corresponding administrative division level (such as provincial, municipal, and county) according to the region attribute, loads all policy provisions related to the writing topic (such as land planning and ecological protection policies) at this level from the pre-stored policy specification library; the effective date and the abolition date of each policy provision are checked one by one, and the provisions that have expired or have been replaced are filtered out, and the current effective policy content is retained; the retained effective provisions are logically compared to identify whether there are contradictions between the provisions (such as different provisions for the same matter), and if there are conflicts, the conflict points are marked and the provisions with higher level or updated publication are retained in priority; the effective provisions that have passed the time effectiveness verification and conflict detection are sorted into a structured "compliance constraint set" to determine the policy boundaries (such as prohibited development areas, volume rate limits, etc.) that the planning text must comply with.

[0050] The specific implementation process of the above step 24 is as follows:

[0051] The "writing topic" (such as "XX watershed ecological restoration planning" "XX city land use overall planning") is extracted from the mandatory items and disassembled into core theme words (such as "watershed ecological restoration" "land use"); the core theme words are used as search indexes to traverse the historical planning cases in the platform template library, and the metadata of the cases (such as the regions involved, theme tags, and regional feature descriptions) are extracted.

[0052] The similarity of the historical cases to the regional features of the current planning is calculated, that is, the regional features in the cases (such as terrain types and climate zones) are compared with the features of the current planning area, the similarity scores are calculated through semantic matching and feature item overlap, and the historical cases with a similarity score higher than a preset threshold (such as 70%) are selected as a reference case set; the reference case set is analyzed, and the associated specifications of the spatial elements (such as the binding relationship between "forest protection" and "slope > 25° area" and the associated logic between "water conservation" and "1km range along the river") are extracted; and the associated specifications and core terms (such as policy keywords and geographic element names) in the cases are standardized (unified term expression and redundant information removal), and finally a structured "standardized keyword set" is generated.

[0053] The application realizes accurate triggering of resource scheduling by accurately detecting the existence state of the selected items, avoids invalid calculation, ensures that the system processes data according to the optimal path, and improves the overall process efficiency; with the help of a large language model to analyze unstructured data, three types of core keywords are efficiently extracted, combined with the spatial topological matching of the user's private knowledge base, a keyword set with spatial tags is generated, which not only ensures the comprehensiveness of keyword extraction, but also establishes the association between keywords and spatial information; based on the loading of the corresponding policy specification library according to the regional attributes, through timeliness verification and conflict detection, a compliance constraint set is output, ensuring that the generated planning text conforms to the latest effective policies in the local area, avoiding policy conflicts from the source, and improving the compliance and authority of the text; taking the writing theme as an index to match high-similarity historical cases, extracting spatial element association specifications to generate a standardized keyword set, learning from historical experience to ensure the applicability and normativity of keywords, and providing content references that meet regional characteristics for planning texts.

[0054] In a preferred embodiment of the application, step 22, structured analysis is performed by a large language model, multi-modal feature recognition is performed on unstructured reference materials, and three types of core keywords, policy provisions, geographic coordinates and statistical indicators, are extracted, including:

[0055] Step 221, performing format standardization preprocessing on unstructured reference materials to eliminate file format differences and unify encoding, generating a parseable continuous text stream;

[0056] Step 222, based on the continuous text stream, performing data modal segmentation by a multi-modal feature recognition engine: identifying and separating text description segments, image data blocks and table structured data;

[0057] Step 223, based on the text description segments separated in step 222, using semantic role labeling technology to analyze sentence components: identifying the responsible subject in the policy text as the first type of keyword, the obligation item as the second type of keyword, and the time constraint as the third type of keyword, and generating a structured policy provision set;

[0058] Step 224, based on the image data block separated out in step 222, the text of the marked elements in the map is recognized by the OCR technology, the element position information is converted to the unified projection coordinate system, and the standard geographic coordinate set with the coordinate system label is generated;

[0059] Step 225, based on the table structured data separated out in step 222, the unit consistency of the numerical data is verified, the index name field, the numerical field and the statistical period field are extracted, and the standardized statistical index set is generated.

[0060] In the embodiment of the application, the specific implementation process of the above step 221 is as follows:

[0061] First, the received unstructured reference materials are scanned for format, the file types (such as PDF, Word, TXT, JPG, PNG, etc.) are identified, and the files of different formats are classified and marked; for PDF and Word files, the format marks (such as font style, paragraph spacing, picture embedding code) are stripped through a text extraction tool, and the pure text content is extracted; for picture files (such as scanned copies, map photos), it is first determined whether the text can be converted by image recognition technology, if yes, the OCR is executed to convert the text into a text format; the encoding of the converted text content is detected, the current character encoding format (such as GBK, UTF-8, ISO-8859-1, etc.) is identified, all the texts are uniformly converted into UTF-8 encoding, and the garbled code problem caused by encoding difference is eliminated; meaningless symbols (such as special control characters, repeated spaces, page breaks) are deleted, sentences are merged (such as sentence splitting caused by line breaks), and finally a coherent, format-difference-free continuous text stream is formed, ensuring that the subsequent parsing tools can be stably read.

[0062] The specific implementation process of the above step 222 is as follows:

[0063] The engine is started and a pre-defined multimodal feature library is loaded. This library contains typical feature templates for text descriptions, image data blocks, and tabular structured data, such as: subject-verb-object structure syntax rules and common punctuation combinations in text fields; binary file header identifiers for image data such as JPG ("FFD8") and PNG ("89504E47"), as well as text placeholders such as "[image]" and "[picture]"; separator features for tabular data such as vertical lines "|" and horizontal lines "—", special marker codes for Word table objects (such as "\tbl" and "\tcell"), and alignment mode features for fixed-spaced numerical columns. The continuous text stream is scanned byte by byte. Starting from the beginning of the continuous text stream, the engine reads the content byte by byte in order, while simultaneously starting three parallel recognition threads to perform real-time matching for the features of text descriptions, image data blocks, and tabular structured data. During the scanning process, the system records the position index of the current byte in the text stream (e.g., byte N) as a reference for subsequent positioning.

[0064] Text description segment recognition and tagging:

[0065] Thread 1 continuously detects whether consecutive bytes constitute parsable text characters (excluding binary garbled characters). When it identifies consecutive occurrences of content conforming to natural sentence rules (such as word combinations containing subject, predicate, and object, interspersed with punctuation marks such as commas and periods), and no table separators or image identifiers are detected, it is initially determined to be a text description segment. Further verification is made to determine whether the text block has no fixed row and column structure: by statistically analyzing the length difference of each line of characters, if the length fluctuation exceeds a preset threshold (such as ±10 characters), it is confirmed to be non-table text, the start byte index and end byte index of the text block are marked (with the period or newline character at the end of the paragraph as the boundary), and it is temporarily stored in the text buffer area.

[0066] Image data block recognition and extraction:

[0067] Thread 2 synchronously detects image features in the text stream: If placeholder text such as "[image]" or "[image]" is detected, the position index of the placeholder is immediately marked and used as a marker for the image data block. The context position information before and after the placeholder (such as "50 bytes before and 20 bytes after [image]") is extracted and stored in the image buffer. If a binary image file header matching the feature library is detected (such as "FFD8" for JPG), the scanning continues from that position until the corresponding file end identifier (such as "FFD9" for JPG) is identified. The complete byte range of the image data block (from the file header to the file end) is determined, all bytes within this range are extracted as image data, and their start and end indices are recorded and stored in the image buffer.

[0068] Tabular structured data identification and extraction:

[0069] Thread 3 focuses on detecting table features: when consecutive vertical lines “|” are scanned and the same number of vertical line separations appear in adjacent rows (such as 3 “|” in each row, forming a 4-column structure), or when consecutive horizontal lines “—” are detected and form an alignment relationship with the text above and below (such as the separator line below the table header), it is determined as a potential table area; if the text stream is derived from Word file conversion, when the marker codes “\tbl” (table start) and “\endtbl” (table end) are identified, the byte range of the table is directly determined as the boundary between the two markers; for suspected table areas, further verify the cell alignment features: by calculating the character spacing between columns, if the consistency of values or text left / right alignment within the same column exceeds 90%, it is confirmed as a table structured data, and the starting and ending indexes are marked and stored in the table buffer area.

[0070] Buffer area storage and index record:

[0071] When the three threads complete their respective identification tasks, the text description segment, image data block, and table structured data are written into three independent buffer areas (such as three arrays in memory); each data entry in each buffer area is accompanied by its complete position index in the original continuous text stream (including the starting byte index and the ending byte index), such as “text segment 1: starting at byte 100 - ending at byte 500” “image block 1: starting at byte 800 - ending at byte 2000”, ensuring that subsequent steps can trace the location association of each data modality in the original text.

[0072] The specific implementation process of the above step 223 is as follows:

[0073] First, scan the text description segment character by character, identify the set of punctuation marks (including period “.”, exclamation mark “!”, semicolon “;”, question mark “?”), and use these punctuation marks as potential sentence boundaries; if the punctuation mark is inside the quotation marks (“”, ‘’), or parentheses (( ), []), such as “File regulations: ‘XX area cannot be developed; violators will be punished’”, do not use it as a sentence boundary to avoid splitting the complete sentence inside the quotation marks; when a punctuation mark outside the quotation marks or parentheses is scanned, determine the sentence boundary; split the long text into independent sentence units according to the verified boundaries, assign a unique serial number to each sentence unit (such as “sentence 1” “sentence 2”), and record the starting and ending positions of the sentence in the original text description segment (such as X line Y word to M line N word), which facilitates subsequent tracing.

[0074] Initialization of semantic role labeling model and sentence parsing preparation:

[0075] Load a pre-trained semantic role labeling model, which contains three types of feature libraries: institutional entity feature library (recording common government departments, institutional names and abbreviations), obligatory verb feature library (containing "prohibit", "should", "must not" and other verbs and phrases with constraints), and time expression feature library (covering typical patterns of dates, time periods, and time limit descriptions, such as "YYYY year MM month DD", "valid period to XXX"). Convert each sentence unit into a sequence format that the model can process (such as a list of tokenized words), while preserving the original position information of the words in the sentence (such as the 1st word, the 2nd word).

[0076] Subject identification (first type of keyword):

[0077] Scan the sentence unit through the named entity recognition module and match the entries in the institutional entity feature library: if the sentence contains "XX City Natural Resources Bureau" or "provincial forestry department" and other complete institutional names, mark it as a candidate subject of responsibility; identify the grammatical relationship between words (such as subject-predicate relationship, verb-object relationship) through dependency syntax tree to determine whether the candidate subject is the issuer or responsibility bearer of the core verb in the sentence; if the candidate subject meets the above conditions, it is confirmed as the subject of responsibility (the first type of keyword), and its specific text and position index in the sentence are recorded.

[0078] Obligation item identification (second type of keyword):

[0079] Extract verb phrases from the sentence unit: scan the verbs and modifiers (such as adverbs, prepositional phrases) in the sentence, match the obligatory verb feature library, and identify phrases such as "prohibit + action" (such as "prohibit occupation"), "should + action" (such as "should repair"), and "must not + action" (such as "must not break through"); determine the pointing object of the verb phrase (such as "prohibit occupation permanent basic farmland" in "prohibit occupation permanent basic farmland") through semantic role labeling, combine the verb phrase with the pointing object to form a complete behavior specification (such as "prohibit occupation permanent basic farmland"); if the behavior specification has a definite association with the identified subject of responsibility (such as the subject of responsibility being the performer of the behavior), mark it as an obligation item (the second type of keyword), and record its complete text and position in the sentence.

[0080] Time limit constraint identification (third type of keyword):

[0081] Identify specific dates (e.g., "January 1, 2024"), time periods (e.g., "2024-2030"), effective conditions (e.g., "from the date of publication"), expiration dates (e.g., "valid until 2030"), and other time-related expressions; Standardize ambiguous time expressions: If the sentence contains ambiguous expressions such as "short-term" or "medium and long-term", combine the context policy background (e.g., the determined time extracted from other sentences) or the default specification (e.g., "short-term" corresponds to "within 3 years") to convert it into an understandable time range; Verify whether the time expression is related to the obligation item of the responsible subject (e.g., "XX Bureau prohibits development from 2024" in "XX Bureau prohibits development from 2024"), if related, mark it as time constraint (third type of keyword), record its text and location.

[0082] Three types of keyword association and structured policy clause set generation:

[0083] Correlation and binding of responsible subjects, obligation items, and time constraints within the same sentence unit: Use sentence number as unique identifier, establish "responsible subject-obligation item-time constraint" correspondence table, for example "sentence 3: subject = XX City Natural Resources Bureau, obligation item = prohibit occupation of permanent basic farmland, time constraint = from 2024"; If no keywords of a certain type are identified in the sentence (e.g., no definite time constraint), mark "none" in the corresponding position and keep the original sentence text for manual review (e.g., "XX Bureau should repair the ecology - no definite time constraint"); Aggregate the association results of all sentences, sort by sentence number, form a structured policy clause set containing "sentence number, responsible subject, obligation item, time constraint, original sentence index", ensure that each clause can be traced back to the specific location of the original text description segment.

[0084] The specific implementation process of the above step 224 is as follows:

[0085] The speckles and scratches in the image are removed by a denoising algorithm, the contrast enhancement technique is used to highlight the text annotations and boundary lines in the map, the geometric correction is performed on the tilted map to ensure the image is horizontally aligned; the OCR text recognition tool is used to scan the pre-processed map image, identify and extract various annotations in the map, including place names (such as "XX County" and "XX River"), administrative boundary names (such as "provincial boundary" and "city boundary"), coordinate scales (such as "East 116°" and "North 39°"), and legend explanations (such as "forest" and "farmland" corresponding symbols); the projection mode annotations (such as "Gauss-Kruger projection"), scale (such as "1:100000"), and coordinate origin parameters in the map corners are identified to determine the original coordinate system of the map; the position information of each element in the map (such as the pixel coordinates of a place name in the map) is mapped to the system's preset unified coordinate system (such as the WGS84 latitude and longitude coordinate system) through the projection conversion formula, and the accurate coordinate range of the element (such as "East 116.2°-116.5°, North 39.1°-39.3°") is calculated; the converted coordinate information and the element text recognized by OCR are bound, a coordinate system label (such as "WGS84") is added, and a standard geographic coordinate set containing "element name-coordinate range-coordinate system" is generated.

[0086] The specific implementation process of the above step 225 is as follows:

[0087] The table header row is parsed, synonymous headers (such as "statistical year" and "data year") are uniformly named as "statistical period", and empty columns without actual data are deleted; duplicate rows (such as identical indicators and values) in the table are checked, and only unique rows are retained to avoid redundancy; the value type fields (such as "area", "population", and "output value") in the table are traversed, the units after each value (such as "hectare", "person", and "ten thousand yuan") are extracted, and the system's preset index unit library is compared:

[0088] If the units are inconsistent (such as "hectare" and "mu" appearing in the same "area" indicator), they are uniformly converted to the standard unit (such as "hectare") according to the conversion formula (such as 1 hectare = 15 mu); if the value is not labeled with a unit, the standard unit is automatically supplemented according to the indicator name (such as "forest coverage" with a default unit of "%").

[0089] Extract the core fields:

[0090] Extract the names reflecting the meaning of the data (such as "cultivated land holding amount" "ecological protection red line area") from the first column of the table or the table header; extract the specific numerical values after unit standardization (such as "5000 hectares" "35.2%"); extract the time range corresponding to the data from the table header or the specified column (such as "2023" "2021-2023"); combine the three extracted fields according to the correspondence of "index name-value-statistical period" to form a structured record, and all records are aggregated to form a standardized statistical index set.

[0091] The application eliminates the format difference of unstructured reference materials and unifies the coding through format standardization preprocessing, reduces the parsing error caused by format problems, and improves the stability of overall data processing; through the multi-modal feature recognition engine, the different modal data such as text, image and table are accurately separated, so that each type of data can be processed specifically, the parsing interference caused by the mixing of different modal data is avoided, the foundation is laid for subsequent extraction of policy clauses, geographic coordinates and statistical indicators, and the accuracy of parsing is improved; the semantic role labeling technology is used to analyze the text description section, accurately identifies the responsibility subject, obligation item and time limit and forms a structured policy clause set, clearly presents the core elements and mutual relationship of the policy, and is convenient for subsequent association with spatial information, improves the utilization efficiency and accuracy of the policy clauses; the map label is recognized through the OCR technology and converted to a unified coordinate system, a standard geographic coordinate set with label is generated, the accurate extraction and standardization of geographic information in the image are realized, and the problem that the map data is difficult to be directly parsed and utilized is solved; the unit consistency of the table data is verified, the key fields are extracted to generate a standardized statistical index set, the standardization and comparability of the statistical data are ensured, and the misuse of data caused by unit confusion or unclear fields is avoided.

[0092] In a preferred embodiment of the application, the keyword set comprises:

[0093] Step 226, establishing a spatial reference coordinate system according to the administrative division vector boundary data in the user private knowledge base;

[0094] Step 227, based on the spatial reference coordinate system, performing spatial weight matching on the structured policy clause set, extracting the geographic entity name in the policy clause, and performing name semantic association with the administrative division vector boundary to obtain an association result; mapping the policy clause to the corresponding administrative division unit according to the association result to generate a policy clause label set with administrative division code;

[0095] Step 228, based on the spatial reference coordinate system, performing topological relationship calculation on the standard geographic coordinate set, and performing spatial position inclusion judgment on the geographic coordinate point and the administrative division vector boundary; when the coordinate point is located in the intersection area of multiple administrative divisions, the topological distance of the coordinate point from each boundary is calculated and the membership weight is allocated, and a coordinate label set with topological weight is generated;

[0096] Step 229, performing spatial scale conversion on the standardized statistical indicator set, identifying the spatial scale attribute of the statistical indicators, and spatially interpolating the indicator values at non-administrative division scales according to population density; converting to generate statistical indicator values at administrative division unit granularity, and attaching administrative division code labels;

[0097] Step 230, fusing the policy clause label set, the coordinate label set, and the administrative division code label set to construct a spatial topological relationship matrix indexed by administrative division units;

[0098] Step 231, generating a keyword set with spatial topological labels based on the spatial coupling degree of policy clauses, geographic coordinates, and statistical indicators in the spatial topological relationship matrix.

[0099] In the embodiments of the present application, the specific implementation process of the above step 226 is as follows:

[0100] The administrative division vector boundary data is called from the user's private knowledge base, which contains boundary polygon coordinate information (such as longitude and latitude coordinate strings) of different levels of administrative divisions (such as provinces, cities, counties, and towns), as well as corresponding administrative division names, codes, and other metadata; the original coordinate system (such as WGS84, GCJ02, etc.) of the vector boundary data is parsed and converted into a preset unified spatial reference coordinate system (such as the national 2000 geodetic coordinate system), ensuring that all boundary data are under the same coordinate framework; the converted vector boundary data is subjected to precision verification to eliminate duplicate or incorrect boundary coordinates (such as self-intersecting polygons), and the complete and accurate administrative division boundary information is retained as a reference for subsequent spatial correlation.

[0101] The specific implementation process of the above step 227 is as follows:

[0102] The geographic entity names are extracted from the structured policy clause set sentence by sentence: the administrative division names (such as "XX district" and "XX county") and geographic region names (such as "XX watershed" and "XX mountain range") contained in each policy clause are located by scanning each policy clause with a place name recognition tool; the extracted geographic entity names are compared with the administrative division vector boundary metadata (including standard place names, aliases, and former names) in the user's private knowledge base, and the name differences are processed through fuzzy matching (such as "Jing City" matching "Nanjing City") and synonym mapping (such as "XX area" corresponding to "XX city") to determine the corresponding administrative division of the geographic entity; according to the correlation results, the policy clause is bound with the corresponding administrative division boundary polygon, the unique code (such as a six-digit administrative division code) of the administrative division is queried, and the code is attached to the policy clause; each policy clause is marked in a structured format of "policy content + associated administrative division name + administrative division code", and the policy clause label set with administrative division codes is formed after being aggregated.

[0103] The specific implementation process of the above step 228 is as follows:

[0104] The coordinate values (such as longitude and latitude) of all geographic coordinate points are extracted from the standard geographic coordinate set, and the coordinate points are projected into the coordinate system established in step 226. For each coordinate point, it is compared with the polygons of the administrative division vector boundary in turn, and the basic administrative division to which it belongs is determined by checking whether the coordinate point is inside the polygon (not on the boundary). If the coordinate point is located on the boundary line of two or more administrative divisions (such as provincial boundary, municipal boundary), the straight line distance from the point to the boundary of each adjacent administrative division is calculated (such as 500 meters to A city boundary, 300 meters to B city boundary). According to the inverse proportion of the distance, the membership weight is allocated (such as A city weight 3 / 8, B city weight 5 / 8, and the total weight is 1). Each coordinate point is marked as "coordinate value + main membership administrative division code + membership administrative division weight" format, and after being collected, the coordinate tag set with topological weight is formed.

[0105] The specific implementation process of the above step 229 is as follows:

[0106] The metadata of each index in the standardized statistical index set is scanned to determine its original spatial scale (such as "national", "provincial", "basin level", "plot level", etc.). For indexes with original scale of non-administrative division (such as "XX basin water quality compliance rate", "XX economic zone GDP"), collect the population density data of all administrative divisions within the scale coverage. According to the proportion of population density, the non-administrative division index value is allocated to the subordinate administrative divisions (such as the basin index value is split according to the proportion of population density of each city and county along the coast), and the corresponding index value of each administrative division is calculated. The converted administrative division granularity index value is bound with the corresponding administrative division code and marked as "index name + value + statistical period + administrative division code" format to form the statistical index set with additional administrative division code tags.

[0107] The specific implementation process of the above step 230 is as follows:

[0108] Taking the administrative division unit as the core index (such as taking the administrative division code as the key value), the policy clause tag set, the coordinate tag set, and the administrative division code tag set (statistical index set) are respectively used. The row of the matrix represents the administrative division code, and the column contains three types of associated information, including the policy clauses associated with the administrative division (from the policy clause tag set), the geographic coordinate points and weights within and at the boundary of the administrative division (from the coordinate tag set), and the statistical index values of the administrative division (from the statistical index set). The associated items of adjacent administrative divisions (such as the policy clauses shared by A city and adjacent B city, the weight allocation of the coordinate points at the boundary) are added in the matrix to determine the spatial topological association (such as adjacent, containing, and boundary) between administrative divisions.

[0109] The specific implementation process of the above step 231 is as follows:

[0110] Based on the spatial topological relationship matrix, the coupling degree of each administrative region, policy provisions and geographic coordinates (such as whether the area mentioned in the policy provisions overlaps with the coordinate point coverage range, the higher the overlap ratio, the higher the coupling degree), the coupling degree of geographic coordinates and statistical indicators (such as whether the statistical indicator value of the area where the coordinate point is located matches the coordinate point characteristics), and the coupling degree of policy provisions and statistical indicators (such as whether the policy requirements are consistent with the current situation of statistical indicators) are calculated respectively; according to the coupling degree result, labels (such as "high coupling", "medium coupling", "low coupling") are added to the policy provisions, geographic coordinates and statistical indicators respectively to mark the association strength in the administrative region; the spatial topological labeled policy provision keywords, geographic coordinate keywords and statistical indicator keywords are summarized, classified according to the administrative region code, and the final keyword set with spatial topological labels is formed to ensure that each keyword contains the information of the administrative region to which it belongs, the associated elements and the coupling strength.

[0111] The application establishes a spatial reference coordinate system through the administrative division vector boundary data of the user's private knowledge base, eliminates the spatial matching errors caused by the difference of coordinate systems of different data sources, and ensures that the association of policy provisions, geographic coordinates, statistical indicators and other elements in the spatial dimension can be accurately calculated. The accurate binding of policy provisions and specific administrative division units is realized, the problem of "mismatch between geographical names mentioned in policy text and administrative boundaries" is solved through geographical name semantic association, and the generated label set with administrative division code makes the policy provisions have a definite spatial orientation; for the spatial attribution of geographic coordinate points, through topological relationship calculation and membership weight distribution, the attribution of coordinate points in a single administrative region is determined, and the coordinate association of multi-administrative region boundary areas is properly handled, and the generated topological weight labeled set ensures that the spatial association of geographic coordinates is not fragmented by the boundary, and the integrity of spatial data is improved.

[0112] The statistical indicators of non-administrative division scale are converted to the granularity matching with the administrative region, the index is adapted in the spatial scale through population density interpolation, and after adding the administrative division code, the statistical data can be directly associated with the corresponding policy and geographical information, solving the data utilization obstacle caused by the mismatch between the index scale and the planning unit; the three types of label sets are integrated with the administrative division as the index, and the spatial topological relationship matrix constructed clearly presents the association network of policy, geography and statistical elements in the same region, and intuitively reflects the spatial binding relationship between cross-elements; the keyword set with topological labels is generated through spatial coupling degree calculation, which quantifies the association strength of policy, geography and statistical elements in space, helps to identify the matching degree between elements (such as high-coupling policy and geographic coordinates are more suitable), provides accurate basis for material selection and content association during planning text generation, and improves the spatial logic of text content.

[0113] In a preferred embodiment of the present application, step 3, the keyword set output by step 2, user knowledge base data, platform template library data and network search data are converted into a multi-dimensional feature vector; based on the feature vector, a feature space covering the planning area is constructed, including:

[0114] Step 31, receiving the keyword set with spatial topology label output by step 2, administrative division boundary data in the user knowledge base, spatial element data of the matched historical cases in the platform template library, and real-time geographic data obtained by network search;

[0115] Step 32, respectively converting the keyword set with spatial topology label, administrative division boundary data, historical case spatial element data and real-time geographic data into semantic-spatial joint feature vector, boundary topology feature vector, element correlation feature vector and space-time dynamic feature vector;

[0116] Step 33, taking the spatial range defined by the administrative division boundary of the planning area as the reference;

[0117] Step 34, based on the semantic-spatial joint feature vector, boundary topology feature vector, element correlation feature vector and space-time dynamic feature vector, a multi-dimensional feature space covering the spatial range of the planning area is constructed.

[0118] In the embodiment of the present application, the specific implementation process of step 31 is as follows:

[0119] The keyword set with spatial topology label output by step 2 is obtained, which contains policy provisions, geographic coordinates, statistical indicators and corresponding administrative division codes, topology weights and spatial coupling degree labels; administrative division boundary data is called from the user knowledge base, including vector boundary polygons (composed of longitude and latitude coordinate strings) of administrative divisions at various levels such as provinces, cities, counties and towns within the planning area, administrative division names, administrative codes and hierarchical relationships (such as a county belonging to a prefecture-level city); the platform template library is accessed to extract spatial element data of high-similarity historical cases matched in step 24, including geographic elements (such as the location coordinates of rivers and forest land) in the cases, correlation specifications between elements (such as the distance relationship between reservoirs and irrigation areas) and spatial description information mentioned in the case text (such as "northeastern mountainous area"); according to the name and writing theme of the planning area, real-time geographic data published by authoritative geographic information platforms and government open data platforms is retrieved, including recent remote sensing image interpretation results (such as land use type changes), meteorological monitoring data (such as rainfall distribution in the past 3 months), traffic construction dynamics (such as the orientation of newly built roads), etc., and the search results are effectively screened (repeated or outdated data is removed).

[0120] The specific implementation process of step 32 is as follows:

[0121] The keyword set with spatial topological labels is decomposed into features. The semantic information of each keyword (such as "ecological protection" and "arable land area") is converted into semantic codes (numerical sequences generated based on pre-trained word vector models). The spatial attributes of the keywords (such as administrative division codes, topological weights, and coordinate ranges) are extracted and converted into spatial codes (numerical sequences reflecting location, range, and correlation strength). The semantic codes and spatial codes are concatenated according to the correspondence to form a semantic-spatial joint feature vector for each keyword. Each dimension of the vector corresponds to a semantic feature or a spatial feature, respectively.

[0122] Convert to boundary topological feature vector:

[0123] For each vector polygon in the administrative division boundary data, extract the geometric features (such as the number of vertices, perimeter, and area) and topological relationships (such as the border length and inclusion relationship between a region and its adjacent regions). Convert these features into numerical representations (such as using "1" to represent inclusion relationship and "0.3" to represent the proportion of border length to total perimeter), sort them according to administrative division level, and then concatenate them to form the boundary topological feature vector.

[0124] Convert to feature vectors related to features:

[0125] Analyze the geographic elements in the spatial element data of historical cases, and convert the location, type, and attributes (such as river length and reservoir capacity) of each element (e.g., "XX River" and "XX Reservoir") into basic numerical features; based on the relationship norms between elements, extract the triples of "Element A - Relationship Type - Element B" (e.g., "Reservoir - Water Supply - Town"), and convert the relationship types (e.g., "Water Supply" and "Enclosure") into relationship strength values ​​(e.g., "1.0" indicates a strong relationship); integrate the basic numerical features and relationship strength values, sort the elements according to their importance in the case, and form an element relationship feature vector.

[0126] Convert to spatiotemporal dynamic feature vectors:

[0127] Real-time geographic data is layered by time dimension (e.g., the past month, the past three months, the past year), and dynamic change indicators (e.g., land use type conversion area, precipitation change amplitude) are extracted for each time period. Spatial distribution characteristics (e.g., the proportion of a certain type of land use in the region) and temporal change characteristics (e.g., monthly precipitation growth rate) are converted into numerical sequences and combined according to the dimensions of "spatial distribution + temporal change" to form a spatiotemporal dynamic feature vector.

[0128] The specific implementation process of step 33 above is as follows:

[0129] From the administrative division boundary data obtained in step 31, the highest-level administrative district boundary of the planning area (such as the municipal boundary of "XX City") is extracted, and the maximum coordinate extreme value (east longitude minimum value, east longitude maximum value, north latitude minimum value, north latitude maximum value) of its spatial range is determined. The integrity of the boundary is checked: check whether the boundary polygon is closed and the coordinate points are continuous, if there are breakpoints or abnormal coordinates (such as longitude and latitude values beyond the normal range), the backup data in the user knowledge base is used for correction; take the area enclosed by the boundary as the reference spatial range, mark all the lower-level administrative districts contained in the range (such as 5 districts and 3 counties under the jurisdiction of XX City), and determine the spatial coverage range and internal administrative hierarchy of the planning area.

[0130] The specific implementation process of the above step 34 is as follows:

[0131] Determine the dimension composition of the multi-dimensional feature space: integrate all dimensions of semantic-spatial joint feature vectors (covering the semantics and spatial attributes of policies and indicators), boundary topology feature vectors (reflecting the geometry and correlation of administrative boundaries), element correlation feature vectors (reflecting the relationship between elements in historical cases), and spatio-temporal dynamic feature vectors (containing the spatio-temporal changes of real-time data), and remove redundant dimensions (such as the "administrative division code" dimension contained in different vectors is only kept once); unify the spatial reference of all vectors to the spatial reference coordinate system established in step 226, to ensure that the spatial position corresponding to each feature dimension is locatable within the planning area; divide the reference spatial range of the planning area into a number of grids according to the preset accuracy (such as 1km x 1km), and each grid is taken as the basic unit of the feature space; map the numerical values of each feature vector to the corresponding grid unit according to the spatial position, so that each grid unit contains the numerical values of all feature dimensions such as semantics, topology, correlation, and spatio-temporal at this position, and finally form a multi-dimensional feature space covering the entire spatial range of the planning area.

[0132] The application realizes comprehensive collection of planning related information by receiving multi-source data (keyword set, administrative division data, historical case elements, real-time geographic data), avoiding the limitation of single data source. At the same time, the data source and type are determined to ensure that the feature vector can cover multi-dimensional information such as policy, space, historical experience and real-time dynamics. Different types of data (keywords, boundaries, case elements, real-time geographic data) are converted into standardized feature vectors (semantic-space joint, boundary topology, element association, time-space dynamics), eliminating the integration obstacles caused by data format and type differences. The spatial range is determined based on the administrative division boundary of the planning area to ensure that the spatial orientation of all feature data strictly matches the planning area, avoiding the problem that the feature space exceeds or covers the actual planning range. This benchmark provides an anchor point for the spatial alignment of multi-source features, ensuring that the feature space constructed later can accurately reflect the actual situation of the planning area. Based on the multi-type feature vector, a multi-dimensional feature space covering the planning area is constructed, realizing the organic integration of multi-dimensional information such as semantics, topology, association and time-space. The feature space not only completely covers the spatial range of the planning area, but also can intuitively present the attributes, relationships and dynamic changes of various elements in the region, so that the planning text generation can deeply combine the regional characteristics and improve the pertinence of the content.

[0133] In a preferred embodiment of the application, a grid partition unit is established in the feature space, and spatial clustering is performed according to the feature vector distribution density; the feature vector is mapped to the corresponding partition unit to generate a dynamic material pool containing spatial association attributes, including:

[0134] Step 35, based on the multi-dimensional feature space constructed in step 34, a standard grid coordinate system covering the spatial range of the planning area is established, forming a plurality of grid partition units;

[0135] Step 36, the spatial distribution density of the semantic-space joint feature vector, the boundary topology feature vector, the element association feature vector and the time-space dynamic feature vector generated in step 32 in the multi-dimensional feature space is calculated;

[0136] Step 37, according to the spatial distribution density calculated in step 36, spatial clustering operation is performed on all feature vectors in the multi-dimensional feature space to identify the aggregation area of the feature vector in the grid partition unit;

[0137] Step 38, each feature vector generated in step 32 is mapped to the partition unit corresponding to its belonging aggregation area in the grid partition unit established in step 35;

[0138] Step 39, in each grid partition unit, the feature vector and its corresponding spatial correlation attribute associated with step 38 are mapped to the unit; all the feature vectors and spatial correlation attributes associated in all the grid partition units are integrated to generate the dynamic material pool containing spatial correlation attributes.

[0139] In the embodiment of the present application, the specific implementation process of the above step 35 is as follows:

[0140] From the multi-dimensional feature space constructed in step 34, the spatial range parameters of the planning area are extracted, including the easternmost and westernmost longitude values and the southernmost and northernmost latitude values of the area, to clearly define the geographical boundaries of the planning area; according to the area size of the planning area and the accuracy requirement of natural resource planning, the size standard of the grid partition unit is set (such as 500m x 500m for city planning and 1000m x 1000m for rural planning), to ensure that the grid can reflect the details of the area and will not cause excessive calculation due to excessive density; taking the coordinates (minimum longitude and minimum latitude) of the southwest corner of the planning area as the origin, the X axis is established along the east-west direction (longitude direction) and the Y axis is established along the south-north direction (latitude direction) to form the reference axes of the standard grid coordinate system; according to the set grid size, the grid lines are divided along the X axis and the Y axis: starting from the origin, a vertical line is marked every interval along the X axis and a horizontal line is marked every interval along the Y axis, and the intersection of the vertical line and the horizontal line forms a grid vertex.

[0141] Each rectangular area surrounded by two adjacent vertical lines and two horizontal lines is a grid partition unit, and each unit is assigned a unique identifier (such as "G-row number-column number", where the row number corresponds to the sequence number in the Y axis direction and the column number corresponds to the sequence number in the X axis direction); the coordinates of the four vertices (such as the longitude and latitude of the upper left corner, the upper right corner, the lower left corner and the lower right corner) of each grid partition unit are recorded to ensure that all units completely cover the planning area without overlapping or gaps between units.

[0142] The specific implementation process of the above step 36 is as follows:

[0143] From the four types of feature vectors generated in step 32, the spatial positioning information of each feature vector is extracted: the semantic-spatial joint feature vector extracts the associated administrative center point coordinates; the boundary topological feature vector extracts the corresponding administrative boundary key point coordinates; the element association feature vector extracts the coordinate range of the geographical elements (such as rivers and forest land) involved; the spatio-temporal dynamic feature vector extracts the sampling point coordinates corresponding to the monitoring data; for each feature vector, the grid partition unit to which it belongs is determined through coordinate comparison: the spatial coordinates of the feature vector are compared with the vertex coordinates of the grid unit recorded in step 35 to determine which unit the coordinates fall within, which is the initial grid unit to which the vector belongs.

[0144] Count the number of each type of feature vector in each grid cell: four types of feature vectors are established respectively, and all vectors are traversed. When a vector is determined to belong to a grid cell, the number of the grid cell in the corresponding feature type count table is increased by 1. Finally, the total number (i.e. frequency) of each type of feature vector in each grid cell is obtained. The frequency of a certain type of feature vector in each grid cell is divided by the actual area of the grid cell (calculated in square kilometers according to latitude and longitude) to obtain the density value of the feature in the grid (such as "individuals per square kilometer"). For linear or planar feature vectors that cover multiple grid cells (such as a river that crosses three grids), the frequency is distributed according to the spatial proportion (such as river length proportion, area proportion) in each grid cell (such as a total frequency of 1, distributed in three grids as 0.3, 0.5, 0.2), and then the density of each grid is calculated according to the distributed frequency.

[0145] The specific implementation process of the above step 37 is as follows:

[0146] Integrate the spatial distribution density of the four types of feature vectors obtained in step 36 to construct a comprehensive density vector for each grid cell. The vector contains four dimensions, corresponding to the density values of the four types of features. Set the core parameters of spatial clustering, including the density threshold (such as the density value of at least three dimensions in the comprehensive density vector of a grid cell being higher than the average density of the respective type) and the distance threshold (such as the straight-line distance between the center points of two grid cells not exceeding 2 grid edge lengths). Traverse all grid cells and mark the grid cells whose comprehensive density vectors meet the density threshold as "core cells" and those that do not meet the threshold as "non-core cells". Starting from the first core cell, find all other core cells within the threshold distance in its periphery and group these core cells into the same cluster. Then, take each core cell in the cluster as the starting point and repeatedly find the core cells in the periphery that meet the conditions until no new core cell is added, forming a complete aggregation area.

[0147] For non-core cells that are not classified, if they are within the distance threshold from a core cell of a cluster, they are grouped into the cluster. If they are all beyond the threshold distance from all core cells, they are marked as "isolated cells" (not included in the aggregation area). Assign a unique cluster number to each cluster, record all grid cells included in each cluster (including core cells and merged non-core cells), and mark the main feature type of the cluster (such as a cluster with the highest "element association feature" density, marked as "element dense area").

[0148] The specific implementation process of the above step 38 is as follows:

[0149] Traverse all feature vectors generated in step 32, query the cluster to which each vector belongs in step 37 (determined by the cluster number of its initial belonging grid cell), for all feature vectors belonging to the same cluster, extract its precise spatial coordinates (such as sampling point latitude and longitude, element center point coordinates), find the list of all grid cells contained in the cluster (from the clustering results in step 37), compare the precise coordinates of each feature vector with the boundary coordinates of each grid cell in the list, determine the unique grid cell to which the vector finally belongs (i.e. the coordinates strictly fall within the rectangular range of the grid cell); if the coordinates of the feature vector are exactly on the boundary line of two grid cells (such as longitude equal to the longitude of a certain vertical line), determine the attribution according to its associated spatial attributes (such as closer to the core element of a certain cell) to avoid repeated mapping; establish a correspondence table of "feature vector ID-grid cell identification-cluster number" to record the final mapping grid cell of each feature vector, ensuring that all vectors are accurately allocated to a unique grid partition cell.

[0150] The specific implementation process of the above step 39 is as follows:

[0151] Traverse all cells in the order of identification of the grid partition cells (such as from "G-1-1" to "G-n-m"), for each cell, perform the following operations, according to the correspondence table in step 38, extract all feature vectors mapped to the cell, including semantic-spatial joint, boundary topology, element association, and spatio-temporal dynamics.

[0152] Extract the spatial association attributes corresponding to each feature vector: extract the administrative division code and spatial coupling degree label of the semantic-spatial joint vector; extract the identification of adjacent grid cells and the length of the boundary of the boundary topology vector; extract the type of geographical elements involved and the distance relationship between elements of the element association vector; extract the data collection time and dynamic change trend (such as "growth" "reduction") of the spatio-temporal dynamics vector; combine the content of each feature vector (such as policy keywords, index values) with its spatial association attributes to form a "vector content-spatial attribute" key-value pair (such as "ecological protection" policy-XX county code-high coupling degree"); integrate all key-value pairs in the grid cell to form a structured data block, the data block contains cell identification, feature vector type statistics (such as containing 3 policy vectors, 2 index vectors), and all "vector content-spatial attribute" key-value pairs.

[0153] After the data block integration of all grid cells is completed, the data blocks are arranged in order of cell identification to construct a unified data set. At the same time, a dynamic update interface is set for the data set: when new feature vectors are added (such as new real-time geographic data) or the spatial attributes of existing vectors change (such as policy coupling degree update), the system automatically locates the corresponding grid cell and updates its data block content. The final data set is a dynamic material pool containing spatially related attributes.

[0154] The present application divides the grid partition unit by establishing a standard grid coordinate system, converts the planning area into a quantifiable and manageable spatial unit, ensures the spatial positioning of feature vectors to a specific grid, and solves the problem of feature distribution ambiguity in a large area. At the same time, the unified grid identification (such as "row number-column number") provides a standardized spatial index for subsequent data correlation and clustering, improving the granularity and efficiency of spatial analysis. The spatial distribution density of four types of feature vectors is calculated, quantifying the aggregation degree of different features in the planning area (such as policy keyword dense area, geographic feature dense area). The introduction of density value can also distinguish the primary and secondary relationship of features, avoid irrelevant or sparse features interfering with subsequent analysis, and enhance the pertinence of feature selection. Based on the density-based spatial clustering, the aggregation area of feature vectors is automatically identified, and grid cells with similar features are classified into a category (such as "policy-ecological element high coupling area"), revealing the hidden spatial correlation pattern in the planning area (such as a region with dense ecological policy and forest land elements).

[0155] The accurate mapping of feature vectors to the grid cells of the corresponding aggregation area realizes the triple binding of "feature-space-cluster" (such as a certain ecological policy vector binding to a specific grid in the "ecological aggregation area"), which ensures that each feature has a clear spatial attribution and avoids the problem of feature disconnection from the region. By correlating feature vectors with spatial attributes and integrating them into a dynamic material pool, the spatialization integration of multi-source data is realized (such as a grid cell containing policy clauses, geographic coordinates, real-time indicators, and their correlation). The "dynamic" of the material pool supports real-time updating (such as automatically assigning new data to the corresponding grid), ensuring that the planning material always reflects the latest situation; the "spatially related attributes" provide rich spatialized arguments for subsequent outline generation and text writing, making the content of the planning text closely integrated with the actual situation and improving the pertinence.

[0156] In a preferred embodiment of the present application, step 4, generating an outline structure based on the spatially related attributes, includes:

[0157] Step 41, extract the spatially related attributes of all grid partition units in the dynamic material pool, identify the hierarchical relationship and spatial dependency relationship between the attributes;

[0158] Step 42, according to the hierarchical relationship and spatial dependence relationship, construct a tree topology structure with administrative unit as root node and grid partition unit as child node;

[0159] Step 43, based on the hierarchical path of tree topology structure, map to the outline chapter hierarchical relationship: root node corresponds to the chapter title of planning text, child node generates secondary chapter title according to spatial dependence intensity;

[0160] Step 44, taking writing theme as constraint condition, adjusting the semantic expression of chapter title to match the theme;

[0161] Step 45, integrating hierarchical chapter title and spatial dependence relationship label, generating the spatial reference outline.

[0162] In the embodiment of the present application, the specific implementation process of the above step 41 is as follows:

[0163] Firstly, all grid partition units (such as "G-1-1", "G-2-3", etc.) in the dynamic material pool are traversed, and the spatial correlation attributes of each unit are extracted one by one. These attributes include: the administrative division code of the grid partition unit (such as "320102" representing a certain district), the identification of adjacent grid units (such as "G-1-1 adjacent units are G-1-2, G-2-1"), the spatial coupling degree of feature vector (such as "high coupling of policy-geography"), the type of geographic element (such as "woodland" and "river"), the spatio-temporal dynamic trend (such as "ecological index growth"), etc. Then, all the extracted attributes are classified and sorted: the attributes related to the hierarchical level of administrative division (such as provincial, municipal and county codes) are classified as "hierarchical attribute class"; the attributes related to the influence relationship between different grid units (such as "the water conservation of upstream grid G-3-2 affects the irrigation conditions of downstream grid G-3-3" and "the woodland coverage of eastern grid affects the wind prevention and sand fixation capacity of western grid") are classified as "spatial dependence attribute class". Then, the hierarchical relationship is identified: by comparing the hierarchical level of administrative division code (such as the first two digits of provincial code are the same, the first four digits of municipal code are the same), the superior-inferior relationship between attributes is determined (such as "320100" corresponding to the municipal unit is the superior of "320102" corresponding to the district-level unit), forming the hierarchical chain of "province-city-county-town-grid".

[0164] Finally, the spatial dependence relationship is identified: the influence direction and intensity of different grid units in "spatial dependence attribute class" are analyzed (such as through the coupling degree value of feature vector, coupling degree > 80% is "strong dependence", 50%-80% is "moderate dependence", and < 50% is "weak dependence"), and the "who influences whom" and the influence degree are determined (such as "grid G-5-4 (reservoir area) has strong dependence on grid G-5-5 (irrigation area)").

[0165] The specific implementation process of the above step 42 is as follows:

[0166] Based on the hierarchical relationship identified in step 41, the root node of the tree topology is first determined: the highest-level administrative division unit corresponding to the planning area (such as "XX City", whose administrative division code is "320100") is selected as the root node, representing the overall spatial range of the planning text; all county-level administrative division units under the root node (XX City) are taken as the direct child nodes of the root node, with each county-level unit corresponding to a child node.

[0167] Next, the child nodes are refined in combination with the spatial dependency relationship: the grid partition units contained in each county-level unit are divided into child nodes based on "spatial dependency strength", with grid units (such as grid units with high farmland proportion) that have strong dependency on the core characteristics of the county-level unit (such as "XX County is mainly agricultural") being given priority as the direct lower-level child nodes of the county-level child node; grid units around these grid units that have medium or weak dependency are taken as the next-level child nodes, and the process is repeated.

[0168] At the same time, the association relationship of each node in the tree structure is marked: the "containment relationship" is marked between the root node and the county-level child nodes; the "spatial subordination + dependency strength" is marked between the county-level child nodes and the grid child nodes (such as "XX County contains grid G-2-3, with strong dependency strength"); the "adjacent dependency + influence direction" is marked between the grid child nodes (such as "G-2-3 and G-2-4 are adjacent, and G-2-3 has water supply dependency on G-2-4"). Finally, a tree topology is formed, with administrative divisions as the top layer and grid units as the bottom layer, containing hierarchical and spatial dependency relationships.

[0169] The specific implementation process of the above step 43 is as follows:

[0170] Starting from the root node, the complete path of each node to the root node is recorded (such as "root node (XX City) - child node (XX County) - child node (grid G-2-3)"), with each path corresponding to the potential hierarchy of "chapter-section" in the outline; the root node (XX City) corresponds to the general title of the planning text (such as "XX City Natural Resources Planning"), serving as the highest level of the outline; then, the direct child nodes (county-level administrative division units) of the root node are sorted according to "spatial importance" (judged according to the total sum of the feature weights of the grid units, the higher the total sum, the more important), generating a first chapter title: the most important county-level unit corresponds to "Chapter 1 XX County Resource Status and Planning", the less important one corresponds to "Chapter 2 XX District Resource Status and Planning", etc., serving as the second level of the outline.

[0171] Then, the grid child nodes under the county-level child node are sorted according to the "spatial dependence strength" identified in step 41 (strong dependence > moderate dependence > weak dependence), and a secondary chapter title is generated: each grid child node cluster (such as a strong dependence grid group) corresponds to a sub-section title under the first chapter (such as "1.1 XX County Core Agricultural Area (Grid G-2-3, G-2-4) Planning" and "1.2 XX County Ecological Edge Area (Grid G-2-5, G-2-6) Planning"), which is the third level and below of the outline; finally, ensure that the hierarchical relationship of each chapter title strictly corresponds to the path of the tree topology (such as the path "XX City-XX County-Grid G-2-3" corresponds to "General Title-First Chapter-1.1 Section"), forming a preliminary hierarchical outline framework.

[0172] The specific implementation process of the above step 44 is as follows:

[0173] Extract the user input writing topic (such as "XX City Ecological Protection and Restoration Planning" and "XX River Basin Water Resources Utilization Planning"), and disassemble the core keywords of the topic (such as "ecological protection", "restoration", and "water resources utilization") as constraint conditions for semantic adjustment; traverse the preliminary chapter titles generated in step 43, and check the semantics of each title to see if it matches the topic keywords: if the title does not contain the core keywords of the topic (such as "Chapter 1 XX County Resource Status and Planning" does not reflect "ecological protection"), then supplement the semantics of the title (such as adjusting it to "Chapter 1 XX County Ecological Resource Status and Protection Planning"); if the expression of the chapter title is inconsistent with the topic direction (such as the topic is "water resources utilization", but the title is "XX Regional Forest Land Development Planning"), then modify the core action word of the title (such as changing "development" to "water resources supporting forest land maintenance"), to ensure that the title direction matches the topic.

[0174] For titles involving spatial dependence relationships (such as "1.1 XX County Core Agricultural Area Planning"), refine the expression in combination with the topic: if the topic is "ecological protection", then adjust it to "1.1 XX County Core Agricultural Area Ecological Protection and Cultivated Land Restoration Planning"; if the topic is "water resources utilization", then adjust it to "1.1 XX County Core Agricultural Area Water Resources Efficient Utilization Planning"; finally, perform semantic consistency verification on all adjusted chapter titles to ensure that the expression style of titles at the same level is uniform (such as using the structure "region + topic action + planning") and that the overall topic is highly consistent with the writing topic.

[0175] The specific implementation process of the above step 45 is as follows:

[0176] First, collect all the hierarchical chapter titles adjusted in step 44 and arrange them in hierarchical order (main title > chapter title > section title > subsection title). Then, attach the spatial dependency relationship labels identified in step 41 to each chapter title: the chapter title corresponds to the relationship label between its contained county-level units and the root node (e.g., "Chapter 1 label: XX County and XX City are contained, strong correlation in ecological protection"); the section title corresponds to the relationship label between its contained grid units (e.g., "Section 1.1 label: Grids G-2-3 and G-2-4 are adjacent and strongly dependent, related to water conservation"). Next, bind the chapter title with the corresponding spatial dependency relationship labels to form a structured entry of "title + label" (e.g., "1.1 XX County Core Ecological Zone Protection Plan [label: Grids G-2-3 and G-2-4 are adjacent and strongly dependent, related to water conservation]").

[0177] Finally, all structured entries are integrated according to the outline's hierarchical structure to generate a complete document containing the main title, chapter titles, section titles, subsection titles, and corresponding spatial dependency tags—this is the spatial reference outline. This outline clarifies the text's chapter structure and annotates the spatial relationships between each chapter, providing a spatial reference for the subsequent generation of the main text.

[0178] This invention ensures that the outline structure aligns with spatial patterns. By mining spatial correlation attributes, it arranges chapters to fit the spatial distribution of natural resources (such as watersheds and topographical zones), avoiding the problem of traditional outlines that "emphasize text over space." It strengthens the matching degree between chapters and regions. The tree-like topology and spatial dependency tags allow each chapter to accurately correspond to specific regions and core features, reducing the disconnect between content and the planned area. It highlights core regional issues by generating secondary headings based on spatial dependency strength, prioritizing key areas (such as ecologically sensitive areas) and avoiding content generalization. It improves thematic relevance by constraining the semantics of headings with the writing theme, ensuring that the overall direction of the outline is consistent with user needs, enhancing the relevance of the planned text, providing clear guidance for subsequent writing, and integrating hierarchical headings and spatial tags to ensure that the main text creation accurately uses materials from corresponding regions, guaranteeing the logical coherence of the text.

[0179] In a preferred embodiment of the present invention, step 5, based on the spatial reference outline generated in step 4, employs a segmented writing strategy to locate the content in the material pool, and generates main text paragraphs based on the feature weights of the partition units, including:

[0180] Step 51: Analyze the chapter hierarchy of the spatial reference outline generated in step 45, and locate the set of gridded partition units corresponding to the current chapter title;

[0181] Step 52: Based on the feature vector spatial distribution density calculated in step 36, dynamically calculate the feature weight of each gridded partition unit associated with the current chapter;

[0182] Step 53, select target partition units in descending order of feature weight, extract feature vectors and spatial correlation attributes of corresponding units from the dynamic material pool;

[0183] Step 54, generate paragraph core sentences based on the extracted spatial correlation attributes, and generate description details by fusing semantic-spatial joint features in the feature vector;

[0184] Step 55, according to the hierarchical depth of the chapter title, control the paragraph length and the professional term density, and generate the text paragraph matched with the current chapter.

[0185] In the embodiment of the present application, the specific implementation process of the above step 51 is as follows:

[0186] First, the chapter hierarchy of the spatial reference outline is parsed layer by layer, the hierarchical relationship of the chapters in the outline is identified (such as "Chapter 1" is a first-level chapter, "Section 1.1" is a second-level chapter, "Section 1.1.1" is a third-level chapter, etc.), and the chapter title currently being processed (such as "Section 1.1 XX County Core Agricultural Area Planning") is determined; then, the spatial dependence relationship label attached to the current chapter title (such as the label contains "related grid units: G-2-3, G-2-4, G-2-5") is extracted, and the grid unit identifier in the label is used to accurately match in the grid partition unit list of the dynamic material pool.

[0187] If the current chapter is a higher level (such as the first-level chapter "Chapter 1 XX County Resource Status and Planning"), the grid partition unit set corresponding to the first-level chapter is formed by collecting all the grid units associated in the labels of all the secondary chapters under it (such as containing all grids G-2-3 to G-2-10); if it is a low-level chapter (such as a third-level section), the grid units directly associated in its label are extracted as the unit set corresponding to the chapter; finally, the unique identifiers of all grid partition units in the set (such as "G-2-3, G-2-4") are recorded, and the positioning and association of the current chapter and the grid unit are completed.

[0188] The specific implementation process of the above step 52 is as follows:

[0189] First, the density values of the four types of feature vectors (semantic-spatial joint, boundary topology, element correlation, and spatio-temporal dynamics) corresponding to each grid partition unit (such as G-2-3, G-2-4) associated with the current chapter (such as "semantic-spatial joint feature density: 5.2 per square kilometer, element correlation feature density: 3.8 per square kilometer") are extracted through the feature vector spatial distribution density data calculated in step 36.

[0190] Then, according to the theme of the current chapter (such as "agricultural area planning"), dynamic weight coefficients are assigned to the four types of feature vectors: features highly relevant to the theme (such as cultivated land and irrigation facility data in the "element-related features" of the agricultural area) are given higher coefficients (such as 0.4), secondary relevant features (such as precipitation data in the "spatio-temporal dynamic features") are given medium coefficients (such as 0.3), and features with lower relevance (such as "boundary topological features") are given lower coefficients (such as 0.15), ensuring that the sum of the coefficients is 1; then, for each grid partition unit, the density values of its four types of features are multiplied by the corresponding dynamic weight coefficients, and the product results are added to obtain the preliminary feature weight of the unit.

[0191] Finally, the preliminary feature weight is corrected in combination with the strength of the grid unit in the spatial dependence relationship (such as "strong dependence" and "moderate dependence" identified in step 41): the strong dependence unit is multiplied by a correction coefficient of 1.2, the moderate dependence unit is multiplied by 1.0, and the weak dependence unit is multiplied by 0.8 to obtain the final feature weight.

[0192] The specific implementation process of the above step 53 is as follows:

[0193] First, the feature weights of all grid partition units associated with the current chapter calculated in step 52 are sorted in descending order; then, the number of selected units is determined according to the hierarchical depth of the current chapter: if it is a first-level chapter (such as "Chapter 1"), the top 50% of grid units in the weight ranking are selected (such as a total of 10 units, the top 5 units are selected); if it is a second-level chapter (such as "Section 1.1"), the top 30% of units in the weight ranking are selected; if it is a third-level or lower chapter, the top 2-3 units with the highest weight are selected, ensuring that the selected target units cover the core content of the chapter.

[0194] Next, the target partition units are located from the dynamic material pool, and all feature vectors (such as "cultivated land retention policy" and "2023 grain yield statistics" in the "semantic-spatial joint features") and corresponding spatial correlation attributes (such as "G-2-3 is located in the east of XX town, adjacent to G-2-4, and has irrigation water dependence" and "the statistical period is 2023") within each unit are extracted. Finally, the extracted feature vectors and spatial correlation attributes are stored by unit classification to form the writing material package of the current chapter.

[0195] The specific implementation process of the above step 54 is as follows:

[0196] First, generate the core sentence of the paragraph based on the extracted spatial correlation attributes: filter the information that best reflects the core characteristics of the target partition unit from the spatial correlation attributes (such as "G-2-3 is the core arable land in XX County, with soil fertility level of first class, annual precipitation of 800 mm, and mainly relying on the irrigation canal of G-2-4 for water supply"), condense these information into a summary sentence (such as "XX County core arable land (G-2-3) has fertile soil and sufficient precipitation, and its irrigation mainly depends on the adjacent G-2-4 irrigation canal"), as the core sentence of the paragraph, to clarify the theme of the paragraph.

[0197] Then, generate the description details by fusing the semantic-spatial joint features in the feature vector: extract the policy provisions related to the core sentence (such as "The XX City Arable Land Protection Ordinance stipulates that the arable land holding capacity of this area shall not be less than 500 hectares") from the semantic-spatial joint features, statistical indicators (such as "In 2023, the arable land area of this region was 520 hectares, and the grain output reached 3000 tons"), and geographic correlation information (such as "The arable land is concentrated in the range of north latitude 32°15'-32°18', east longitude 118°20'-118°23'"); then, organize these detail information in logical order (such as policy requirements → current status data → geographic distribution) and supplement to the core sentence to form a complete paragraph framework.

[0198] The specific implementation process of the above step 55 is as follows:

[0199] First, identify the hierarchical depth of the current chapter title: determine the hierarchical depth by analyzing the chapter number (such as "Chapter 1" is one level, depth 1; "Section 1.1" is two levels, depth 2; "Subsection 1.1.1" is three levels, depth 3); then, control the paragraph length according to the hierarchical depth: the paragraph of one-level chapter is mainly summarized, with length controlled in 150-200 words (about 3-4 sentences); the paragraph of two-level chapter needs moderate detail, with length controlled in 200-300 words (about 4-6 sentences); the paragraph of three-level and below chapters needs detailed elaboration, with length controlled in 300-500 words (about 6-10 sentences), to ensure that the lower the level, the more specific the content.

[0200] Meanwhile, control the density of professional terms: use general terms (such as "resource protection" and "planning target") in the first-level chapter, with a term density of 1-2 per 100 words; appropriately increase professional terms (such as "arable land retention" and "irrigation efficiency") in the second-level chapter, with a density of 3-4 per 100 words; use subdivided terms (such as "soil organic matter content" and "drip irrigation technology coverage") in the third-level and lower chapters, with a density of 5-6 per 100 words, to ensure the balance between professionalism and readability; finally, according to the length and term density requirements described above, adjust the paragraph framework generated in step 54 (such as deleting redundant information, supplementing necessary terms or explanatory sentences), form coherent text paragraphs matching the current chapter level, and mark the corresponding grid partition unit identifier of the paragraph (such as "[corresponding unit: G-2-3, G-2-4]").

[0201] The present application ensures that the paragraph content is strictly bound to the specific spatial range of the planning area by analyzing the outline to locate the corresponding grid unit, avoiding the problem of "text not matching the area", and keeping the text closely related to the actual area; based on the feature weight screening target unit, the material of high weight area (such as ecological sensitive area and core functional area) is preferentially called, so that the paragraph focuses on the core problems of key areas, avoiding the problem of content being narrated in a flat and unstructured manner; the core sentence is generated based on the spatial correlation attribute, the details are supplemented by integrating semantic-spatial features, which ensures the clear theme of the paragraph and reflects the deep correlation between policy, data and geographical space, enhancing the persuasiveness of the content; according to the chapter level, the paragraph length and term density are controlled, the high-level chapter is strong in generalization and the term is simple, and the low-level chapter is detailed and professional, so that the text structure is clear, the readability and professionalism are balanced, and the hierarchical expression requirements of the planning text are met; through automatic positioning of materials and generation of paragraphs, the cost of manual screening of information and organization of content is reduced, while ensuring the consistency of content with the outline and regional features, improving the accuracy and efficiency of text generation.

[0202] The specific implementation process of the above step 6 is as follows:

[0203] First, identify the hierarchical attribution of all text paragraphs generated in Step 5. The system traverses each paragraph and, according to the corresponding chapter level marked at the end of the paragraph (such as "[Corresponding Chapter: Section 1.1]"), classifies the paragraph under the corresponding chapter of the spatial reference outline, forming a "chapter-paragraph" correspondence (such as "Chapter 1" containing paragraphs 1, 2, "Section 1.1" containing paragraphs 3, 4, etc.); then, perform chapter number standardization processing. According to the hierarchical logic of "chapter-section-subsection", assign a uniform number to each chapter: the chapter title does not need to be numbered; the chapter title uses "Arabic number +." format (such as "1." "2."); the section title uses "chapter number + Arabic number +." format (such as "1.1." "1.2."); the subsection title uses "section number + Arabic number" format (such as "1.1.1" "1.1.2"). At the same time, add the paragraph number to each paragraph according to its order of appearance in the chapter (such as "1.1 Section 1" "1.1 Section 2"), ensuring that the hierarchy is clear and traceable.

[0204] Then, unify the data presentation format. Scan the numerical data in all paragraphs (such as area, population, proportion, etc.), and convert the scattered textual description data (such as "about 5,000 hectares of arable land") into tables or charts: for comparison of similar data (such as "2020-2023 arable land area changes"), automatically generate line charts or bar charts and insert them below the corresponding paragraph; for multi-dimensional index data (such as "soil fertility, irrigation conditions, yield of each grid unit"), generate structured tables, with table title format unified as "Table X.X Title Content" (such as "Table 1.11.1 Section Involved Grid Unit Agricultural Index Table"), and referenced in the paragraph as "See Table X.X for details".

[0205] At the same time, call the standard terminology library in the field of natural resource planning to replace non-standard expressions in the paragraph (such as "national space development" unified as "national space development and protection pattern", "red line area" clearly as "ecological protection red line area"). For the first time, add a brief note to the term (such as "spatial coupling degree: refers to the correlation strength of policy provisions and geographic coordinates") to ensure that the term is used in a standardized and easy-to-understand manner; finally, integrate all structured chapters and paragraphs in the order of the outline hierarchy, generate structured text containing uniform numbering, standardized data format, and standardized terminology, and automatically generate a table of contents (containing chapter titles and corresponding page numbers) and an abstract (a condensed version of the core content of each chapter, about 300-500 words), forming a complete structured text framework.

[0206] The specific implementation process of the above Step 7 is as follows:

[0207] First, the system performs a policy clause matching check. It extracts all policy-related expressions from the structured text and compares them with the latest policy specification library in the local policy engine. It checks if the policy is currently effective (e.g., excludes clauses that have been abolished); verifies the accuracy of clause references (e.g., whether the "farmland protection clause" matches the original text); and confirms that the policy's scope matches the planning area. If any mismatches are found (e.g., a reference to an abolished policy), the system marks the error location and suggests replacing it with a different policy (e.g., "This reference to the 'XX Measures' has been abolished since 2023. Please replace it with 'XX Regulations, Chapter 3'").

[0208] Second, the system performs a spatial logic consistency check. Based on the spatial correlation attributes of the dynamic material pool, it checks if the spatial relationships described in the text conform to geographical rules. For example, if the text mentions "large reservoirs in the XX arid region planning," the system compares the region's time-space dynamic feature vector (e.g., annual precipitation <200mm, evaporation >1500mm) to determine if the "large reservoir" planning contradicts the characteristics of the arid region. If there is a contradiction, the system marks "spatial logic anomaly: water resources are scarce in the XX arid region, and large reservoir planning may have feasibility issues." Another example is checking if the upstream and downstream relationship descriptions are reasonable (e.g., "upstream pollution control" is consistent with "downstream water quality improvement"). If there is a logical break (e.g., only mentioning upstream control without mentioning downstream impact), the system suggests "please add a specific impact analysis of upstream control on downstream water quality."

[0209] Then, the system performs a data consistency check. It traverses all numerical data in the text and creates a "data item-value-occurrence position" table (e.g., "farmland area" is "520 hectares" in section 1.1 and "510 hectares" in section 2.3). It checks if the values of the same data item are consistent. If the deviation is within the allowed range (e.g., ±5%), it automatically takes the average value and updates it uniformly. If the deviation exceeds the range (e.g., ±10%), it marks "data inconsistency: 'farmland area' differs by 10% in sections 1.1 and 2.3, please check." At the same time, it checks if the data units are uniform (e.g., whether "hectares" and "mu" are mixed). If there is a unit confusion, it automatically converts to the standard unit and prompts.

[0210] Next, the system checks the structured text against the official format standards of natural resource planning texts. It checks if the chapter numbers, chart insertion, and term annotations conform to the specifications: for example, whether the table contains a title, unit, and data source; whether the chapter numbers have any jumps or repetitions; whether the abstract is at the beginning of the text, etc. If any format issues are found (e.g., a table is missing a data source), it marks "Table 1.1 is missing a data source, please add 'Data source: XX City Statistical Bureau 2023 Report'".

[0211] Finally, all the check results are summarized into a "check report" containing error types, locations, and modification suggestions. According to the "check report", automatic correction can be made for standardizable errors (such as unit conversion, term replacement). For errors that require manual judgment (such as policy clause applicability), the user is prompted to confirm and modify. After the user confirms and modifies, the system performs the check again until all errors are processed (error rate ≤ 3%), and finally outputs a planning text that meets the policy requirements, is spatially logically consistent, has accurate data, and is in a standard format.

[0212] As shown in Figure 2 , a natural resource planning text generation system includes:

[0213] A receiving module is configured to receive user-input standardized data, which includes mandatory items and optional items. The mandatory items include a title, a region, and a writing topic, and the optional items include an outline template and unstructured reference materials.

[0214] An execution module is configured to make resource scheduling decisions based on the presence or absence of optional items. If optional items are present, the reference materials are parsed to extract keywords using a large language model. If optional items are not present, local policy compliance analysis and platform database matching are triggered.

[0215] A generation module is configured to convert a set of keywords, user knowledge base data, platform template library data, and network search data into a multi-dimensional feature vector. A feature space covering a planning area is constructed based on the feature vector. A grid partition unit is established within the feature space, and spatial clustering is performed based on the feature vector distribution density. The feature vector is mapped to the corresponding partition unit to generate a dynamic material pool containing spatial correlation attributes.

[0216] A processing module is configured to generate an outline structure based on the spatial correlation attributes. According to the spatial reference outline, a segmented writing strategy is used to locate the content of the material pool, and based on the feature weights of the partition unit, text paragraphs are generated. Structured processing is performed on the text paragraphs to obtain a structured processed text. Dynamic checking is performed on the structured processed text to output a final planning text.

[0217] A computer-readable storage medium stores a program that, when executed by a processor, implements the method.

[0218] The above describes preferred embodiments of the present application. It should be noted that those skilled in the art can make several improvements and refinements without departing from the principles of the present application. These improvements and refinements should also be considered within the scope of the present application.

Claims

1. A method for natural resource planning text generation, the method comprising: The method comprises: Step 1, receiving user input standardized data, the standardized data including mandatory items and optional items, wherein the mandatory items include title, region and writing theme, and the optional items include outline template and unstructured reference materials; Step 2, performing resource scheduling decision according to the presence or absence of optional items, if the optional items exist, extracting keywords by analyzing the reference materials through a large language model, if the optional items do not exist, triggering local policy compliance analysis and platform database matching; Step 3, converting the keyword set output by step 2, user knowledge base data, platform template library data and network search data into a multi-dimensional feature vector; constructing a feature space covering the planning area based on the feature vector; establishing a grid partition unit in the feature space, and performing spatial clustering according to the feature vector distribution density; mapping the feature vector to the corresponding partition unit to generate a dynamic material pool containing spatial correlation attributes; Step 4, generating an outline structure space based on the spatial correlation attributes; Step 5, positioning the material pool content using a segmented writing strategy based on the spatial reference outline generated in step 4, and generating text paragraphs based on the feature weight of the partition unit; Step 6, performing structured processing on the text paragraphs generated in step 5 to obtain the structured processed text; Step 7, performing dynamic verification on the structured processed text to output the final planning text; in step 2, the keywords are matched with the user's private knowledge base based on the spatial topological relationship to generate a keyword set with spatial tags; the keyword set includes: step 226, establishing a spatial reference coordinate system based on the administrative division vector boundary data in the user's private knowledge base; step 227, based on the spatial reference coordinate system, performing spatial weight matching on the structured policy clause set, extracting the geographical entity name in the policy clause, and performing name semantic association with the administrative division vector boundary to obtain the association result; mapping the policy clause to the corresponding administrative division unit based on the association result to generate a policy clause tag set with administrative division code; step 228, based on the spatial reference coordinate system, performing topological relationship calculation on the standard geographical coordinate set, and performing spatial position inclusion judgment on the geographical coordinate point and the administrative division vector boundary; when the coordinate point is located in the intersection area of multiple administrative divisions, the topological distance of each boundary is calculated and the membership weight is assigned to generate a coordinate tag set with topological weight; step 229, performing spatial scale conversion on the standardized statistical index set, identifying the spatial scale attribute of the statistical index, and performing spatial interpolation on the index value of non-administrative division scale according to population density; converting to generate statistical index values of administrative division unit granularity, and adding administrative division code tags; step 230, fuse the policy clause tag set, coordinate tag set and administrative division code tag set to construct a spatial topological relationship matrix indexed by administrative division unit; step 231, generating a keyword set with spatial topological tags based on the spatial coupling degree of policy clauses, geographical coordinates and statistical indicators in the spatial topological relationship matrix.

2. The method of claim 1, wherein, Step 2, resource scheduling decision is made based on the presence of optional items. If optional items are present, key words are extracted from reference materials through large language model analysis. If no optional items are present, local policy compliance analysis and platform database matching are triggered, including: Step 21, detect the presence of optional items in user input. If a outline template or unstructured reference material is detected, proceed to step 22. If no optional items are detected, proceed to step 23. Step 22, perform structured analysis using a large language model to identify multi-modal features in unstructured reference materials, extract three types of core keywords: policy provisions, geographic coordinates, and statistical indicators. Step 23, trigger local policy engine analysis. Based on the region attribute in the mandatory items, load the policy specification library corresponding to the administrative division, perform time effectiveness verification and conflict detection of policy provisions, and output the compliance constraint set. Step 24, perform platform database dynamic matching. Use the writing topic as an index to search the platform template library and match historical planning cases with a regional feature similarity higher than the threshold. Extract spatial element association specifications from the cases to generate a standardized keyword set.

3. The method of claim 2, wherein, Step 22, perform structured analysis using a large language model to identify multi-modal features in unstructured reference materials, extract three types of core keywords: policy provisions, geographic coordinates, and statistical indicators, including: Step 221, perform format standardization preprocessing on unstructured reference materials to eliminate file format differences and unify encoding, generating a continuous text stream that can be parsed. Step 222, based on the continuous text stream, perform data modality segmentation using a multi-modal feature recognition engine: identify and separate text description segments, image data blocks, and table structured data. Step 223, based on the text description segments separated in step 222, use semantic role labeling technology to analyze sentence components: identify the responsible subject in the policy text as the first type of keyword, the obligation item as the second type of keyword, and the time effectiveness constraint as the third type of keyword, generating a structured policy provision set. Step 224, based on the image data blocks separated in step 222, use OCR technology to recognize the text of the labeled elements in the map, convert the element position information to a unified projection coordinate system, and generate a standard geographic coordinate set with coordinate system labels. Step 225, based on the table structured data separated in step 222, verify the consistency of numerical data units, extract index name fields, numerical fields, and statistical period fields, and generate a standardized statistical indicator set.

4. The method of claim 3, wherein, Step 3, convert the keyword set output in step 2, user knowledge base data, platform template library data, and network search data into a multi-dimensional feature vector. Based on the feature vector, construct a feature space covering the planning area, including: Step 31, receive the keyword set with spatial topology labels output in step 2, administrative division boundary data in the user knowledge base, spatial element data of the matched historical cases in the platform template library, and real-time geographic data obtained through network search. Step 32, respectively convert the keyword set with spatial topology label, administrative boundary data, historical case spatial feature data and real-time geographic data into semantic-spatial joint feature vector, boundary topology feature vector, element correlation feature vector and time-space dynamic feature vector; Step 33, take the spatial range defined by the administrative boundary of the planning area as the reference; Step 34, based on the semantic-spatial joint feature vector, the boundary topology feature vector, the element correlation feature vector and the time-space dynamic feature vector, construct a multi-dimensional feature space covering the spatial range of the planning area.

5. The method of claim 4, wherein, Establish a grid partition unit in the feature space, and perform spatial clustering according to the feature vector distribution density; Map the feature vector to the corresponding partition unit to generate a dynamic material pool containing spatial correlation attributes, including: Step 35, based on the multi-dimensional feature space constructed in step 34, establish a standard grid coordinate system covering the spatial range of the planning area, forming a number of grid partition units; Step 36, calculate the spatial distribution density of the semantic-spatial joint feature vector, the boundary topology feature vector, the element correlation feature vector and the time-space dynamic feature vector generated in step 32 in the multi-dimensional feature space; Step 37, according to the spatial distribution density calculated in step 36, perform spatial clustering operation on all feature vectors in the multi-dimensional feature space, and identify the aggregation area of the feature vector in the grid partition unit; Step 38, map each feature vector generated in step 32 to the partition unit corresponding to its own aggregation area in the grid partition unit established in step 35; Step 39, in each grid partition unit, associate the feature vector and its corresponding spatial correlation attribute mapped to the partition unit in step 38; integrate all associated feature vectors and spatial correlation attributes in all grid partition units to generate a dynamic material pool containing spatial correlation attributes.

6. The method of claim 5, wherein, Step 4, generate an outline structure space reference outline based on the spatial correlation attributes, including: Step 41, extract the spatial correlation attributes of all grid partition units in the dynamic material pool, and identify the hierarchical relationship and spatial dependency relationship between the attributes; Step 42, according to the hierarchical relationship and spatial dependency relationship, construct a tree topology structure taking the administrative unit as the root node and the grid partition unit as the child node; Step 43, based on the hierarchical path of the tree topology structure, map to the chapter hierarchical relationship of the outline: the root node corresponds to the chapter title of the planning text, and the child node generates the secondary chapter title according to the spatial dependency strength; Step 44, take the writing theme as a constraint condition, adjust the semantic expression of the chapter title to match the theme; Step 45, integrate the hierarchical chapter title and spatial dependency relationship label to generate the space reference outline.

7. The method of claim 6, wherein, Step 5, according to the space reference outline generated in step 4, adopt a segmented writing strategy to locate the content of the material pool, and generate text paragraphs based on the feature weight of the partition unit, including: Step 51, analyze the chapter hierarchy of the space reference outline generated in step 45, and locate the grid partition unit set corresponding to the current chapter title; Step 52, dynamically calculate the feature weight of each grid partition unit associated with the current chapter according to the feature vector space distribution density calculated in step 36; Step 53, select the target partition unit in descending order of feature weight, extract the feature vector and spatial association attribute of the corresponding unit from the dynamic material pool; Step 54, generate a paragraph core sentence based on the extracted spatial association attribute, and generate a description detail by fusing the semantic-spatial joint features in the feature vector; Step 55, according to the hierarchical depth of the chapter title, control the paragraph length and the professional term density, and generate a text paragraph matched with the current chapter.

8. A natural resource planning text generation system characterized by, The system is used to perform the method of any one of claims 1 to 7, comprising: a receiving module for receiving user input standardized data, the standardized data including mandatory items and optional items, wherein the mandatory items include title, region and writing topic, and the optional items include outline template and unstructured reference materials; an execution module for executing resource scheduling decisions according to the presence or absence of optional items, and if the optional items exist, extracting keywords by analyzing reference materials through a large language model, and if the optional items do not exist, triggering local policy compliance analysis and platform database matching; a generation module for converting the keyword set, user knowledge base data, platform template library data and network search data into a multi-dimensional feature vector; constructing a feature space covering the planning area based on the feature vector; establishing a grid partition unit in the feature space, and performing spatial clustering according to the feature vector distribution density; mapping the feature vector to the corresponding partition unit to generate a dynamic material pool containing spatial association attributes; a processing module for generating an outline structure based on the spatial association attributes; positioning the material pool content using a segmented writing strategy according to the spatial reference outline; generating a text paragraph based on the feature weight of the partition unit; performing structured processing on the text paragraph to obtain a structured processed text; and performing dynamic verification on the structured processed text to output the final planning text.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a program which is executed by the processor to implement the method of any one of claims 1 to 7. The computer readable storage medium stores a program which is executed by the processor to implement the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Geographic information system (GIS)-based territorial space planning method and system

    CN117312476A

  • Industrial land intelligent site selection method based on spatial big data and artificial intelligence big model

    CN120495047A