A method for summarizing key information of a normative text, an electronic device, and a medium

By acquiring the key information types of normative texts and using a large language model for text extraction and index configuration, the problems of high data acquisition costs, coarse text parsing, and strong subjectivity in analysis in existing technologies are solved, enabling efficient and economical extraction of key information from normative texts and integration of valuable data.

CN121480485BActive Publication Date: 2026-06-02URBAN PLANNING & DESIGN INST OF SHENZHEN UPDIS

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
URBAN PLANNING & DESIGN INST OF SHENZHEN UPDIS
Filing Date
2025-12-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies suffer from high data acquisition costs, coarse text parsing, fragmented knowledge, and strong subjectivity in processing normative texts, making it difficult to efficiently and cost-effectively analyze and process normative texts, extract key information, and integrate valuable data.

Method used

By acquiring normative documents and key information types from multiple target domains, text extraction is performed using a large language model, key information indexes are configured, and key information is divided and archived based on the indexes to form a normative text key information database.

Benefits of technology

It enables efficient and cost-effective analysis and processing of normative texts, extracting key information and integrating valuable data, overcoming the systemic limitations of existing technologies, and improving the automation level of data acquisition and the accuracy of analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121480485B_ABST
    Figure CN121480485B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of text analysis, in particular to a key information collating method for normative text, an electronic device and a medium. According to the key information collating method for normative text, a plurality of normative files corresponding to target fields and key information types corresponding to each target field are first obtained; the plurality of normative files are input into a large language model for text extraction to obtain field text information corresponding to each target field; key information indexes are configured for each target field according to the key information types corresponding to each target field; and based on the key information indexes configured for each target field, key information division and archiving are performed on the corresponding field text information to form a key information library of normative text. In this way, normative text can be efficiently and cost-effectively analyzed and processed, key information can be extracted therefrom, and valuable data can be integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text analysis technology, and in particular to a method, electronic device, and medium for organizing key information in normative texts. Background Technology

[0002] In the era of big data, the number of normative texts issued across various industries has exploded, forming a vast ocean of information characterized by its sheer volume, complexity, and diverse formats. Related technologies in the field of normative text processing suffer from systemic limitations. Specifically, data acquisition relies excessively on manual collection or third-party commercial libraries, resulting in high costs, poor timeliness, and weak autonomy. Text parsing remains at a macro-level, holistic analysis, lacking detailed segmentation studies of internal structure. The few attempts at segmentation often employ mechanical methods using fixed lengths or delimiters, disrupting semantic coherence and failing to establish index connections between paragraphs, leading to knowledge fragmentation. Analysis tools remain limited to traditional software, struggling to handle unstructured text rich in semantic information. While some research has introduced large language models, they rely excessively on subjective expert interpretation, lacking universality and objectivity.

[0003] Therefore, how to efficiently and cost-effectively analyze and process a large volume and variety of normative texts, extract key information, and integrate valuable data has become a major challenge that urgently needs to be addressed in the industry. Summary of the Invention

[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method for organizing key information in normative texts, which enables efficient and cost-effective analysis and processing of normative texts, extracting key information and integrating valuable data.

[0005] The method for organizing key information of normative text according to the first aspect of this application includes:

[0006] Obtain the normative documents corresponding to multiple target domains, as well as the key information types corresponding to each target domain;

[0007] Multiple normative documents are input into a large language model for text extraction to obtain domain text information corresponding to each target domain;

[0008] Configure a key information index for each target domain according to the key information type corresponding to each target domain;

[0009] Based on the key information index configured for each target domain, the key information of the corresponding domain text information is divided and archived to form a standardized text key information database.

[0010] According to some embodiments of this application, obtaining normative documents corresponding to multiple target domains includes:

[0011] Obtain the key information organization requirements for each target domain and the domain key fields corresponding to the key information organization requirements;

[0012] Determine the target search engine that meets the requirements for organizing the key information, and the corresponding retrieval syntax rules of the target search engine;

[0013] Based on the requirements for organizing key information and the search syntax rules, the key fields of the domain are integrated to form key search statements for the domain.

[0014] The key search terms for the target domain are input into the target search engine to search for normative documents, thereby obtaining the normative documents corresponding to each target domain.

[0015] According to some embodiments of this application, the step of integrating the key domain fields into a key domain retrieval statement based on the key information organization requirements and the retrieval syntax rules includes:

[0016] Based on the aforementioned information organization requirements, the key fields in the aforementioned fields are divided into full fields, selected fields, and irrelevant fields;

[0017] Based on the retrieval syntax rules, construct logical retrieval statements for the full range of fields;

[0018] Based on the retrieval syntax rules, construct or logical retrieval statements for the selected fields;

[0019] Based on the retrieval syntax rules, construct non-logical retrieval statements for the irrelevant fields;

[0020] Based on the retrieval syntax rules, the AND logical retrieval statement, the OR logical retrieval statement, and the non-logical retrieval statement are integrated to obtain the domain key retrieval statement.

[0021] According to some embodiments of this application, before inputting the domain key search statement into the target search engine for normative document search, the method further includes:

[0022] Obtain the authoritative domain name field corresponding to each of the target domains;

[0023] Construct authoritative domain constraint statements based on the authoritative domain field corresponding to each of the target domains;

[0024] The step of inputting the key search terms of the domain into the target search engine to perform a normative document search, and obtaining the normative documents corresponding to each target domain, includes:

[0025] The key search terms for the domain and the authoritative domain constraint terms are input into the target search engine to search for normative documents, thereby obtaining the normative documents corresponding to each target domain.

[0026] According to some embodiments of this application, the step of inputting multiple normative documents into a large language model for text extraction to obtain domain text information corresponding to each target domain includes:

[0027] Determine the domain text structure pattern corresponding to each of the aforementioned normative documents;

[0028] For each of the aforementioned normative documents, text paragraphs are segmented based on the corresponding domain text structure pattern to obtain multiple corresponding normative text segments;

[0029] Multiple normative text segments are input into a large language model for text extraction to obtain domain text information corresponding to each target domain.

[0030] According to some embodiments of this application, determining the domain text structure pattern corresponding to each of the normative documents includes:

[0031] Title hierarchy identification is performed on the aforementioned normative documents;

[0032] If, in the title hierarchy identification, it is determined that the normative document contains at least two levels of text titles, and the text titles satisfy the hierarchical progression condition, then the hierarchical text category is determined as the domain text structure pattern.

[0033] If, during the title level identification process, it is determined that the normative document does not contain at least two levels of the aforementioned document titles, or that the aforementioned document titles do not meet the preset hierarchical progression conditions, then title normative identification is performed on the normative document.

[0034] If, in the title standardization identification, it is determined that the title of the standard document meets the preset writing paradigm conditions, the hierarchical framework writing category is determined as the domain text structure pattern;

[0035] If, during the title standardization identification, it is determined that the standardization document does not meet the writing paradigm conditions, the linear list of writing categories is identified as the domain text structure pattern.

[0036] According to some embodiments of this application, for each of the normative documents, text segmentation is performed based on the corresponding domain text structure pattern to obtain multiple corresponding normative text segments, including:

[0037] When the text structure pattern of the domain belongs to the hierarchical structure text category, the normative text is segmented into context headings to obtain multiple corresponding normative text segments;

[0038] When the domain text structure pattern belongs to the hierarchical framework writing category, hierarchical indexing and segmentation are performed on the normative text to obtain multiple corresponding normative text segments.

[0039] When the text structure pattern of the domain belongs to the linear list of text categories, semantic segmentation is performed on the normative text to obtain multiple corresponding normative text segments.

[0040] According to some embodiments of this application, after the key information is divided and archived based on the key information index configured for each target domain to form a standardized text key information database, the method further includes:

[0041] In the aforementioned normative text key information database, the text information for each of the aforementioned fields is vectorized.

[0042] After vectorizing the text information in each of the aforementioned domains, a standardized text question-and-answer program is generated based on the standardized text key information database.

[0043] According to some embodiments of this application, determining the domain text structure pattern corresponding to each of the normative documents includes:

[0044] Obtain multiple typical document structure categories, and construct corresponding document structure identification instructions based on each of the typical document structure categories;

[0045] The text structure identification instruction and the plurality of normative documents are input into the large language model so that the large language model can determine the domain text structure pattern corresponding to each normative document from the plurality of normative documents.

[0046] According to some embodiments of this application, after obtaining the normative documents corresponding to multiple target fields, the method further includes:

[0047] Each of the aforementioned normative documents is subjected to an authority verification, and the authority verification result is obtained;

[0048] If the authority verification result meets the preset authority judgment conditions, the timeliness verification is performed on each of the normative documents to obtain the timeliness verification result.

[0049] If the authority verification result does not meet the preset authority judgment condition, or if the timeliness verification result does not meet the preset invalidation judgment condition, the corresponding normative document is filtered out.

[0050] According to some embodiments of this application, the step of inputting multiple normative documents into a large language model for text extraction to obtain domain text information corresponding to each target domain includes:

[0051] Based on the key information type corresponding to each target domain, construct a key information extraction instruction;

[0052] The key information extraction instructions and the plurality of normative documents are input into the large language model so that the large language model can extract the domain text information corresponding to each target domain from the plurality of normative documents.

[0053] Secondly, embodiments of this application provide an electronic device, including: a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method for organizing key information of normative text as described in any one of the embodiments of the first aspect of this application.

[0054] Thirdly, embodiments of this application provide a computer-readable storage medium storing a program that is executed by a processor to implement the method for organizing key information of normative text as described in any one of the embodiments of the first aspect of this application.

[0055] The method for organizing key information in normative text according to the embodiments of this application has at least the following beneficial effects:

[0056] According to the method for organizing key information in normative texts presented in this application, it is necessary to first obtain normative documents corresponding to multiple target domains, as well as the key information types corresponding to each target domain; input the multiple normative documents into a large language model for text extraction to obtain domain text information corresponding to each target domain; configure a key information index for each target domain according to the key information types corresponding to each target domain; and based on the key information index configured for each target domain, classify and archive the corresponding domain text information to form a normative text key information database. In this way, it is possible to efficiently and cost-effectively analyze and process normative texts, extract key information, and integrate valuable data.

[0057] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0058] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0059] Figure 1 A flowchart illustrating a method for organizing key information in normative texts provided in this application embodiment;

[0060] Figure 2 Another flowchart illustrating the method for organizing key information in normative texts provided in this application embodiment;

[0061] Figure 3 Another flowchart illustrating the method for organizing key information in normative texts provided in this application embodiment;

[0062] Figure 4 Another flowchart illustrating the method for organizing key information in normative texts provided in this application embodiment;

[0063] Figure 5 Another flowchart illustrating the method for organizing key information in normative texts provided in this application embodiment;

[0064] Figure 6 Another flowchart illustrating the method for organizing key information in normative texts provided in this application embodiment;

[0065] Figure 7 Another flowchart illustrating the method for organizing key information in normative texts provided in this application embodiment;

[0066] Figure 8 Another flowchart illustrating the method for organizing key information in normative texts provided in this application embodiment;

[0067] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0068] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0069] In the description of this application, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0070] In the description of this application, it should be understood that the orientation descriptions, such as up, down, left, right, front, and back, are based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0071] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0072] In the description of this application, it should be noted that, unless otherwise explicitly defined, terms such as "setting," "installation," and "connection" should be interpreted broadly. Those skilled in the art can reasonably determine the specific meaning of the above terms in this application based on the specific content of the technical solution. Furthermore, the identification of specific steps in the following text does not imply a limitation on the order of steps or execution logic. The execution order and logic between each step should be understood and inferred from the content described in the embodiments.

[0073] Normative texts refer to a general term for formal documents formulated and issued by functional departments, industry authorities, and other organizations with public management functions in accordance with their authority and procedures. These documents aim to adjust social behavior, regulate the exercise of power, clarify the relationship between rights and obligations, guide practical activities in specific fields, and have universal binding force and repeated applicability. They include, but are not limited to, industry guidelines, standards and procedures, implementation opinions, and work guidelines. Their format usually follows the norms of official document writing.

[0074] In the era of big data, the number of normative documents issued for various industries has exploded, forming a vast ocean of information characterized by its sheer volume, complexity, and diverse formats, exhibiting significant features of "abundance, complexity, and disorder." Faced with such a large volume and variety of normative documents, how to achieve efficient, systematic, and cost-effective analysis and processing, extract key information, and integrate valuable data has become a long-standing and pressing problem in the field of information management and normative document research.

[0075] Current practices for obtaining normative texts mainly rely on two methods: manual collection and downloading from third-party libraries. However, both have significant limitations, specifically:

[0076] Manual collection is not only time-consuming and labor-intensive, but also easily affected by human factors, leading to omissions or errors in information;

[0077] While third-party libraries can reduce manpower burden, they require high subscription fees or face access restrictions. The comprehensiveness and integrity of their data are difficult to guarantee, and the update frequency, coverage, and file accessibility are all subject to the platform's operational strategy.

[0078] It is worth noting that existing technologies for analyzing normative texts generally remain at a macro level in terms of granularity, lacking detailed research on the internal structure of normative texts. This "one-size-fits-all" approach ignores the internal logical connections and key differences of the text, which greatly limits the practicality of the analysis results.

[0079] At the level of analytical tools, existing technologies are mostly limited to traditional software with single functions, making it difficult to meet complex needs. Although artificial intelligence, especially large-scale language model technology, has developed rapidly, its application in the field of intelligent analysis of normative texts is still insufficient. It has failed to fully leverage its potential to automatically extract key information, identify trends and problems, and provide accurate and forward-looking decision support for the formulation of normative texts.

[0080] Existing technical solutions have some limitations in the three aspects of data acquisition, text processing and intelligent analysis.

[0081] Regarding data acquisition mechanisms, both manual collection and automatic downloading using third-party libraries are susceptible to human error and incur high operational costs, making efficient acquisition difficult. Some related technologies propose manually importing standardized text files collected from multiple channels, but manual import requires a huge human investment in massive data scenarios, while manual update strategies lead to significant information lag issues, and the lack of autonomy in encapsulation tools makes it difficult to meet diverse needs.

[0082] In terms of the accuracy of text analysis, current research relies too much on macroscopic analysis of the whole text, and lacks detailed segmentation analysis of the internal structure of normative texts. Although a few studies have attempted to segment normative text content, they generally adopt mechanical methods with fixed lengths or fixed delimiters, which destroy the logical structure and contextual coherence of the text. Moreover, the segmentation results are isolated from the original text and lack an effective correlation mechanism.

[0083] At the level of analytical tools, existing practices are generally limited by traditional software frameworks. These tools mainly focus on numerical data and have weak capabilities for processing unstructured normative text rich in semantic information.

[0084] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method for organizing key information in normative texts, which enables efficient and cost-effective analysis and processing of normative texts, extracting key information and integrating valuable data.

[0085] The following explanation is based on the accompanying drawings.

[0086] Reference Figure 1 The method for organizing key information in normative text according to embodiments of this application may include:

[0087] Step S101: Obtain the normative documents corresponding to multiple target domains, as well as the key information types corresponding to each target domain;

[0088] Step S102: Input multiple normative documents into the large language model for text extraction to obtain the domain text information corresponding to each target domain;

[0089] Step S103: Configure a key information index for each target domain according to the key information type corresponding to each target domain;

[0090] Step S104: Based on the key information index configured for each target domain, the key information of the corresponding domain text information is divided and archived to form a standardized text key information database.

[0091] In response to the systemic limitations of existing technologies in normative text processing, such as high data acquisition costs, coarse text parsing, fragmented knowledge, and strong subjectivity in analysis, this application proposes a complete method from data collection to knowledge archiving, which solves the above problems one by one through four closely related steps.

[0092] In some embodiments, step S101 involves obtaining normative documents corresponding to multiple target domains and key information types corresponding to each target domain.

[0093] It should be noted that the core task of step S101 is to complete data collection and target definition, establishing the direction and scope for the entire normative text processing flow. Specifically, this step requires completing two basic tasks: first, obtaining the corresponding normative documents from multiple target domains; and second, clarifying the types of key information to be extracted from each domain.

[0094] Obtaining normative documents from multiple target domains signifies that the method for organizing key information from normative texts in this application has multi-domain adaptability. In practice, it is necessary to construct separate collection strategies for different industry characteristics. For example, the urban planning field may focus on documents related to land management and construction standards, while the environmental protection field needs to cover documents such as pollutant emission standards and ecological red line delineation. By simultaneously covering multiple domains, economies of scale can be achieved, thereby reducing overall development and maintenance costs. The collection process can be automated, such as by writing customized collection programs to obtain texts from authoritative channels such as public platforms based on constraints such as keywords for each domain, domain names of publishing institutions, and file formats, ensuring the reliability and timeliness of data sources.

[0095] Clearly defining the key information types for each target domain is crucial for establishing a clear information extraction framework before data collection. While normative documents across different domains may vary in format, each type of text has its core components. For example, land management normative documents typically include key points such as scope of application, approval processes, compensation standards, and liability delineation, while industry regulatory normative documents may focus on entry requirements, operational standards, and penalties. By pre-defining these key information types, text extraction avoids blindly capturing irrelevant information. This configuration process often requires the participation of domain experts who, based on actual business needs and text characteristics, identify valuable information dimensions and create standardized key information index templates, providing a basis for automated processing in subsequent steps.

[0096] Reference Figure 2 According to some embodiments of this application, step S101, obtaining normative documents corresponding to multiple target fields, may include:

[0097] Step S201: Obtain the key information organization requirements for each target domain and the domain key fields corresponding to the key information organization requirements;

[0098] Step S202: Determine the target search engine that meets the requirements for organizing key information, and the corresponding retrieval syntax rules of the target search engine;

[0099] Step S203: Based on the key information, organize the requirements and retrieval syntax rules, and integrate the key fields of the domain to form the key retrieval statement of the domain;

[0100] Step S204: Input the domain key search statement into the target search engine to search for normative documents and obtain the normative documents corresponding to each target domain.

[0101] In some embodiments, step S201 involves obtaining the key information organization requirements for each target domain and the domain key fields corresponding to the key information organization requirements.

[0102] It's important to note that the initial step is requirements analysis and field extraction, which forms the starting point of the entire data collection process. By gathering key information from each target domain, the requirements are organized to clarify which types of normative documents need to be collected and what core information elements each type should contain. Based on this, key fields closely related to the domain are extracted. For example, in the field of urban renewal, this might include terms such as "inefficient land use," "compensation standards," and "approval processes." These fields serve as the foundational vocabulary for constructing subsequent search queries and are also the core basis for judging document relevance. This step, through communication with domain experts or business stakeholders, transforms vague information requirements into actionable technical parameters, laying the foundation for accurate data collection.

[0103] In some embodiments, step S202 involves determining the target search engine that meets the requirements for organizing key information and the retrieval syntax rules corresponding to the target search engine.

[0104] It's important to note that identifying the target search engine and its retrieval syntax rules is a crucial choice in the technical implementation path. Different search engines vary in their indexing scope, update frequency, and syntax support, requiring a comprehensive evaluation based on the data collection needs. For example, when targeting domestic normative documents, choosing a search engine with better support for Chinese and more comprehensive official website coverage is more suitable. After determining the search engine, it's essential to thoroughly study its supported retrieval syntax rules, including field restrictions (such as title:), logical operations (AND / OR / NOT), and exact matching methods such as wildcards and quotation marks. Understanding these rules is a prerequisite for constructing efficient search statements and directly relates to the ability to quickly locate target documents within massive amounts of information.

[0105] At the implementation level, the retrieval syntax rules of different target search engines differ objectively, which affects the construction and execution effect of key domain search statements. Each search engine has different designs in terms of operator definitions, field qualifier support, and syntax precedence, requiring targeted adaptation in the technical implementation.

[0106] In some embodiments, step S203 involves integrating key domain fields into domain key search statements based on the requirements for organizing key information and search syntax rules.

[0107] It should be noted that constructing key domain search statements based on the aforementioned needs and rules is a transformation between needs and technology. This step combines and arranges the key domain fields extracted in step S201 according to the syntax rules determined in step S202 to form a structured search expression. For example, fields such as "inefficient land use" and "redevelopment" are connected using logical operators, and the `title:` keyword is used to ensure that the title contains the core words, excluding distracting words such as "meeting" and "notification". This statement construction is not a simple listing of keywords, but rather combines the ranking mechanism and weight rules of the search engine to maximize the recall of relevant documents and minimize the introduction of noise through syntax combination, reflecting a refined approach to information retrieval technology.

[0108] Reference Figure 3 According to some embodiments of this application, step S203, based on the requirements for organizing key information and retrieval syntax rules, integrates key domain fields to form key domain retrieval statements, which may include:

[0109] Step S301: Based on the requirement to organize key information, divide the key fields in multiple fields into full fields, selected fields, and irrelevant fields;

[0110] Step S302: Construct logical search statements for all fields according to the search syntax rules;

[0111] Step S303: Construct or logical search statements for the selected fields according to the search syntax rules;

[0112] Step S304: Construct non-logical search statements for irrelevant fields according to the search syntax rules;

[0113] Step S305: Based on the retrieval syntax rules, integrate logical retrieval statements, or logical retrieval statements and non-logical retrieval statements to obtain domain key retrieval statements.

[0114] In step S301 of some embodiments, based on the requirement of organizing key information, the key fields in multiple fields are divided into full fields, selected fields and irrelevant fields;

[0115] It should be noted that, firstly, based on the need to organize key information, key fields from multiple fields are categorized into three types: full fields, selective fields, and irrelevant fields. Full fields refer to core terms that must appear simultaneously in the target document, such as keywords like "inefficient land use" and "redevelopment" in the field of urban renewal; the absence of any one of them makes it impossible to accurately define the document's topic. Selective fields refer to extended terms that only require one of them to appear, such as synonyms or related concepts, used to broaden the search coverage. Irrelevant fields are distracting words that need to be explicitly excluded, such as "meeting," "notice," and "draft for comments," to avoid recalling non-standard documents. This classification method transforms vague search intentions into clear logical rules, providing a clear structural foundation for subsequently constructing compound search statements.

[0116] In some embodiments, steps S302 to S304 involve constructing logical search statements for all fields according to search syntax rules; constructing logical search statements for selected fields according to search syntax rules; and constructing non-logical search statements for irrelevant fields according to search syntax rules.

[0117] It should be noted that steps S302 to S304 construct different logical retrieval statements for the three types of fields. For all fields, an AND logical retrieval statement is constructed according to the retrieval syntax rules, meaning that the AND operator or plus sign (+) is used, requiring all all fields to exist simultaneously. For example, the retrieval expression "inefficient land use + redevelopment + compensation standards" ensures highly relevant results and improves precision. For selection fields, an OR logical retrieval statement is constructed, using the OR operator or vertical bar (|) to connect them, allowing any one of them to appear. For example, the expression "agreement transfer | targeted land supply | bidding and auction" can cover different expressions of the same concept, avoiding the omission of valid documents due to terminology differences and improving recall. For irrelevant fields, a non-logical retrieval statement is constructed, using the NOT operator or minus sign (-) for exclusion. For example, the setting "-meeting-notification-draft" can filter out informal documents and reduce noise interference.

[0118] In some embodiments, step S305 involves integrating logical search statements, or logical search statements and non-logical search statements, according to search syntax rules to obtain domain key search statements.

[0119] It should be noted that step S305 integrates the above three logical retrieval statements to form the final domain key retrieval statement. The integration process follows the syntax priority and execution order of the search engine, typically placing non-logical statements at the end, or grouping logical statements with parentheses, using these logical statements as the main framework. For example, the final statement might present a structure like "title:(Inefficient Land Use AND Redevelopment) AND (Agreement Transfer OR Targeted Land Supply) - Meeting - Notice". This composite retrieval statement combines the precision of logical statements with the inclusiveness of logical statements and the exclusivity of non-logical statements, enabling precise positioning within massive amounts of information. It ensures the relevance of the recalled documents while controlling the number of results within a reasonable range, providing an efficient and reliable technical tool for the search execution in step S204, significantly improving the automation level and result quality of normative document collection.

[0120] In step S204 of some embodiments, the domain key search statement is input into the target search engine to search for normative documents and obtain the normative documents corresponding to each target domain.

[0121] It should be noted that performing the search and retrieving files is the final stage of the data collection process. After inputting the key search query constructed in step S203 into the target search engine, this application automatically executes the search task and returns a list of files that meet the criteria. To ensure the quality of the results, it is usually necessary to deduplicate, filter, and verify the search results, removing duplicates, entries from non-authoritative sources, or those with incorrect formats, ultimately obtaining a set of normative documents corresponding to each target domain. This step realizes the implementation from the search query to the actual files, completing the data collection loop. Through automated script batch execution, a large number of files from multiple domains can be obtained at once, significantly improving collection efficiency and coverage, and solving the problems of low efficiency and incomplete coverage of traditional manual collection methods.

[0122] Reference Figure 4 According to some embodiments of this application, before step S204, which involves inputting the domain key search statement into the target search engine for normative document searching, the following may also be included:

[0123] Step S401: Obtain the authoritative domain name field corresponding to each target domain;

[0124] Step S402: Construct authoritative domain constraint statements based on the authoritative domain field corresponding to each target domain;

[0125] In step S204, the domain-specific search query is input into the target search engine to search for normative documents, obtaining the normative documents corresponding to each target domain, which may include:

[0126] Step S403: Input the domain key search statement and authoritative domain constraint statement into the target search engine to search for normative documents and obtain the normative documents corresponding to each target domain.

[0127] Adding authoritative domain constraints before inputting key search terms into the target search engine is an important technical measure to improve the quality of normative document collection. This application's embodiment effectively solves the problems of insufficient search result authority and excessive noise by introducing a domain filtering mechanism, ensuring that the obtained documents are reliable in origin and compliant in content.

[0128] In some embodiments, step S401 involves obtaining the authoritative domain name field corresponding to each target domain;

[0129] It's important to note that obtaining the authoritative domain name field corresponding to each target field is a prerequisite for constructing domain name constraints. Authoritative domain names refer to the website domains of institutions that publish regulatory documents; different regulatory departments in specific industries have different dedicated domain names. For the urban renewal field, the authoritative domain name field can cover the official website domains of local natural resources, housing and urban-rural development departments at all levels. This process requires establishing a domain name field database and maintaining and updating it regularly to ensure the accuracy and timeliness of domain name information and prevent constraints from becoming ineffective due to institutional adjustments or website redesigns.

[0130] In step S402 of some embodiments, an authoritative domain constraint statement is constructed based on the authoritative domain field corresponding to each target domain;

[0131] It's important to note that constraint statements are constructed based on the acquired authoritative domain field, converting the domain information into a syntax expression recognizable by the search engine. Technically, the search engine-supported `site:` syntax can be used. If multiple authoritative domains are involved, OR logic must be used to combine them. When constructing constraint statements, attention must be paid to the search engine's limitations on the number and length of `site:` syntax. When there are many authoritative domains, it may be necessary to group them into multiple constraint statements or use other equivalent filtering methods. Furthermore, the domain hierarchy must be considered. Top-level domains offer a wider range of constraints but may include non-target organizations, while second-level domains offer more precise constraints but may miss lower-level departments. A trade-off must be struck based on the required data collection accuracy.

[0132] In step S403 of some embodiments, the domain key search statement and authoritative domain name constraint statement are input into the target search engine to search for normative documents and obtain the normative documents corresponding to each target domain.

[0133] It should be noted that the search is performed by combining key domain search statements with authoritative domain constraint statements to achieve dual filtering. Specifically, in this embodiment, the two types of statements are combined using AND logic and submitted to the search engine, requiring that the returned results simultaneously meet the conditions of content relevance and source authority. This joint search mechanism effectively filters out non-authoritative sources such as commercial document platforms, personal blogs, and reprint websites, significantly reducing the noise reduction cost of the result set. For example, without domain constraints, searching for "redevelopment of inefficient land use" may return a large amount of information interpretation, news reports, and even advertising information. After adding the corresponding authoritative domain constraint statements, the result set only retains official documents published on the official websites of relevant departments, significantly improving collection efficiency and data quality.

[0134] This application's embodiment uses pre-domain filtering to move authority verification from the post-collection cleaning stage to the search stage, reducing invalid downloads and subsequent processing workload. Simultaneously, the synergistic effect of domain constraints and keyword retrieval enables this embodiment to accurately locate target files within massive amounts of internet information, solving the data pollution problem caused by mixed sources in traditional methods. This design not only ensures the official attributes and validity of normative documents but also provides a clean and reliable data source for subsequent text extraction and knowledge base construction, serving as a key control point for ensuring data quality throughout the entire method chain.

[0135] This application's embodiment transforms document collection from a vague, fragmented, and manually dependent traditional model into a clear, systematic, and automated one through a four-step process: requirements analysis, tool selection, statement construction, and execution acquisition. Each step provides precise input for subsequent stages, ensuring the accuracy, comprehensiveness, and timeliness of the collection results, and providing a high-quality raw data foundation for subsequent text extraction and knowledge base construction.

[0136] According to some embodiments of this application, after obtaining the normative documents corresponding to multiple target fields in step S101, the method may further include:

[0137] Each normative document undergoes an authority verification process to obtain the authority verification results.

[0138] If the authority verification result meets the preset authority judgment conditions, the timeliness verification is performed on each normative document to obtain the timeliness verification result.

[0139] If the authority verification result does not meet the preset authority judgment condition, or the timeliness verification result does not meet the preset invalidation judgment condition, the corresponding normative document will be filtered out.

[0140] It should be noted that after obtaining the normative documents, a dual verification process of authority and timeliness is added, forming a key barrier for data quality control to ensure that the documents entering subsequent processes are reliable in origin and valid in content. Authority verification, as the primary criterion, aims to filter out unofficial or low-credibility documents. The verification process can be carried out from multiple dimensions, including the issuing organization, file format, and website domain: the issuing organization must be a functional department or an authorized official organization; the file format must be a formal official document format such as PDF or Word, not a web snapshot or scanned copy; and the download source should be limited to an authoritative domain or a certified official website.

[0141] This application embodiment constructs an authoritative feature library, comparing the metadata of each document with the standards in the library to generate an authority verification result containing compliant and non-compliant items. Only documents that meet preset authority judgment conditions (such as matching all authoritative features) can proceed to the next stage. This design effectively filters out informal documents such as those reposted on commercial platforms, personally interpreted versions, and draft proposals for comments, avoiding data pollution caused by mixed sources and providing a clean and reliable data source for subsequent processing.

[0142] Timeliness verification, conducted after ensuring the document meets authority standards, is used to identify and remove expired or revised documents. This step involves extracting key information such as the document's publication date, effective date, and repeal clauses, and comparing it with the current timestamp and a document update database. For example, if a document explicitly states "effective from [date]," and the current date is earlier than that, the document has not yet taken effect and should be removed. Similarly, if a document states "this document is repealed from [date]," or if a subsequent revised version exists, the document is expired and must also be removed.

[0143] This application embodiment can also incorporate an information dynamic tracking mechanism to automatically capture newly released replacement documents or revision notices, establish version relationships between documents, and ensure the accuracy of timeliness verification. The introduction of timeliness verification solves the problems of delayed updates and difficulty in identifying expired documents in traditional methods, ensuring the timeliness and accuracy of knowledge base content.

[0144] The screening mechanism, as the final execution step in the verification process, is logically clear and strictly enforced. When a document's authority verification result fails to meet preset conditions—for example, if the source domain is not on the whitelist or the publishing organization is not certified—this embodiment directly removes it from the process and prevents it from proceeding further. Similarly, if a document passes authority verification but fails timeliness verification (e.g., it is obsolete or not yet effective), it will also be screened out. This dual-judgment mechanism ensures that only documents that simultaneously meet both the authority and validity prerequisites are retained. The screening operation is not a simple deletion; instead, detailed reasons are recorded and archived in the pending review area for easy manual review or traceability, avoiding data loss due to accidental deletion. This mechanism fundamentally improves the purity and usability of the normative document library, providing a high-quality data foundation for subsequent text extraction, knowledge archiving, and intelligent question answering.

[0145] In step S102 of some embodiments, multiple normative documents are input into a large language model for text extraction to obtain domain text information corresponding to each target domain.

[0146] It should be noted that the core task of step S102 is to hand over the normative documents collected in the previous steps to a large language model for in-depth processing, completing the initial transformation from raw text to structured information. This step changes the traditional information extraction method that relies on manual reading or simple rule matching. By leveraging the semantic understanding capabilities of the large language model, it automatically identifies and extracts the text content required for each target domain.

[0147] In practice, this step uses multiple normative documents as input sources, feeding them into a large language model for processing. Unlike traditional keyword matching or fixed-delimiter segmentation tools, the large language model, through pre-trained language knowledge, can understand the contextual relationships, logical hierarchy, and semantic emphasis of text. When the model receives a normative document, it locates relevant content within the entire text based on the key information types identified in step S101. For example, for a land management normative document, the large language model can identify that "scope of application" corresponds to the descriptions of geographical area and subject validity, and "approval process" corresponds to the procedures and requirements, rather than simply mechanically extracting information based on paragraph order or keyword frequency.

[0148] This semantic understanding-based extraction method effectively overcomes the shortcomings of traditional methods that disrupt text coherence. Traditional mechanical segmentation often forcibly breaks down logically complete paragraphs, resulting in incomplete information or loss of context. Large language models, on the other hand, can determine the natural boundaries of text and extract information units while maintaining semantic integrity, ensuring that the extracted domain-specific text information is logically self-consistent. Furthermore, large language models can handle documents with different formats and expression styles, accurately identifying the location of key information regardless of whether it is a clause-based, chapter-based, or hybrid structure, demonstrating strong adaptability.

[0149] After extraction, this embodiment of the application categorizes and organizes the results according to the target domain, forming a domain-specific text information set. This means that information from different documents but belonging to the same domain will be grouped together; for example, the content regarding "compensation standards" in all normative documents involving "redevelopment of inefficient land" will be extracted centrally. This domain-based aggregation method provides a clear structural framework for subsequent indexing and archiving, enabling massive amounts of text information to begin to appear organized and manageable, and laying the foundation for building a high-quality normative text key information database.

[0150] Reference Figure 5 According to some embodiments of this application, step S102, which involves inputting multiple normative documents into a large language model for text extraction to obtain domain text information corresponding to each target domain, may include:

[0151] Step S501: Determine the domain text structure pattern corresponding to each normative document;

[0152] Step S502: For each normative document, perform text segmentation based on the corresponding domain text structure pattern to obtain multiple normative text segments.

[0153] Step S503: Input multiple normative text segments into the large language model for text extraction to obtain the domain text information corresponding to each target domain.

[0154] In some embodiments, step S501 involves determining the domain text structure pattern corresponding to each normative document.

[0155] It should be noted that the core of step S501 is to determine the domain text structure pattern of each normative document. In practice, this embodiment requires pre-analysis of the input documents to identify their inherent textual structure features. Normative documents in different domains often follow specific organizational conventions. For example, urban renewal documents often adopt a framework of "general provisions - specific provisions - responsibility attribution - supplementary provisions," while industry regulatory documents may present a structure of "access conditions - operational standards - supervision mechanisms - penalties." This step automatically determines whether a document belongs to a structured (clear hierarchy, complete headings), semi-structured (hierarchical structure exists but headings are incomplete), or unstructured (only linear clause numbers) mode by combining sample learning and rule matching. This identification process provides a basis for subsequent segmentation, avoiding the drawbacks of traditional methods that ignore textual differences and adopt uniform mechanical segmentation, and is a prerequisite for refined processing.

[0156] Reference Figure 6 According to some embodiments of this application, step S501, determining the domain text structure pattern corresponding to each normative document, may include:

[0157] Step S601: Identify the title level of the normative document;

[0158] Step S602: If, in the title hierarchy identification, it is determined that the normative document contains at least two levels of text titles, and the text titles satisfy the hierarchical progression condition, the hierarchical structure text category is determined as the domain text structure pattern.

[0159] Step S603: If, in the title level identification, it is determined that the normative document does not contain at least two levels of document titles, or the document titles do not meet the preset hierarchical progression conditions, title normative identification is performed on the normative document.

[0160] Step S604: If, in the title standardization identification, it is determined that the title of the standard document meets the preset writing paradigm conditions, the hierarchical framework writing category is determined as the domain text structure pattern.

[0161] Step S605: If it is determined in the title normativity identification that the normative document does not meet the writing paradigm conditions, the linear list of writing categories is determined as the domain text structure pattern.

[0162] In some embodiments, steps S601 to S602 involve identifying the title hierarchy of a normative document. If, in the title hierarchy identification, it is determined that the normative document contains at least two levels of headings, and the headings satisfy the hierarchical progression condition, the hierarchical structure heading category is determined as the domain text structure pattern.

[0163] It should be noted that steps S601 to S602 constitute the first judgment criterion, the core of which is to identify whether the document has a complete and progressive multi-level heading system. Specifically, the embodiments of this application first scan the entire text to detect whether there are at least two levels of headings, such as "Chapter 1" with "Section 1" under it. More importantly, these headings must meet the condition of hierarchical progression, that is, the upper-level headings and lower-level headings must present an inclusive and extended relationship in terms of numbering rules and semantic logic, rather than simply being parallel. If the document meets both of these requirements, it is determined to be a hierarchical document type. Such documents are usually rigorous in system and clear in structure, and are commonly found in formal texts such as industry documents and departmental regulations. Their hierarchical completeness provides a natural semantic boundary for subsequent accurate segmentation.

[0164] In some embodiments, steps S603 to S604 involve: if, in the title hierarchy identification, it is determined that the normative document does not contain at least two levels of text titles, or that the text titles do not meet the preset hierarchical progression conditions, then title normative identification is performed on the normative document; if, in the title normative identification, it is determined that the text titles of the normative document meet the preset text paradigm conditions, then the hierarchical framework text category is determined as the domain text structure pattern.

[0165] It should be noted that when a normative document fails the title level check, a second judgment path is initiated, shifting to title standardization identification. This shift reflects the flexibility of the methodology, recognizing that not all normative documents possess complete multi-level headings, but may still retain certain structural characteristics. At this point, the system no longer examines the number of heading levels, but focuses on whether the heading itself conforms to preset writing paradigm conditions. Paradigm conditions include, but are not limited to, whether the heading naming follows a general framework such as "General Provisions—Specific Provisions—Attribution of Responsibility—Supplementary Provisions," or whether it adopts a unified format such as "I. General Requirements" or "II. Basic Principles." If the document title exhibits standardized naming characteristics, even if the hierarchy is incomplete, it is still classified as a hierarchical framework document. These documents are often official notices, implementation opinions, etc., and although the hierarchy is brief, the structural framework is still discernible.

[0166] In step S605 of some embodiments, if it is determined in the title normativity identification that the normative document does not meet the writing paradigm conditions, the linear listing of writing categories is determined as the domain text structure pattern.

[0167] It should be noted that, as a final catch-all classification, documents that lack clear multi-level headings and do not follow a unified naming convention are categorized as linearly listed documents. These documents are typically arranged sequentially with simple serial numbers such as "Article 1, Article 2," lacking chapter divisions and structural identifiers, and their content is loosely organized. They are commonly found in drafts, provisional documents, or local guidance documents.

[0168] According to some other embodiments of this application, step S501, determining the domain text structure pattern corresponding to each normative document, may include:

[0169] Obtain multiple typical document structure categories, and construct corresponding document structure identification instructions based on each typical document structure category;

[0170] Input the text structure identification instructions and multiple normative documents into the large language model so that the large language model can determine the domain text structure pattern corresponding to each normative document from the multiple normative documents.

[0171] It should be noted that acquiring multiple typical document structure categories and constructing corresponding identification instructions is a fundamental step in the methodology. These typical document structure categories are derived from a systematic analysis and summary of massive amounts of normative documents and can include hierarchical structure, hierarchical framework, and linear listing patterns. Each category has distinct formal characteristics: hierarchical structure documents possess a complete multi-level heading system and progressive numbering rules; hierarchical framework documents have headings but incomplete hierarchical levels; and linear listing documents are simply arranged by article number. Based on these category characteristics, structured identification instructions need to be constructed. These instructions not only include category definitions but also key identification markers, hierarchical progression conditions, heading naming paradigms, and other judgment criteria. These instructions are described in natural language, clearly defining the boundaries between categories and providing a clear basis for judgment for the large language model, enabling it to examine input documents according to a unified standard.

[0172] The core execution step of the method is to simultaneously input the text structure identification instructions and multiple normative documents into the large language model. After receiving the instruction set and the documents to be analyzed, the model reads the document content one by one and performs matching and judgment based on the structural features defined in the instructions. For example, the model scans the numbering format of document titles, checking for hierarchical relationships such as "Chapter 1, Section 1, Article 1," and verifies whether the title naming conforms to the paradigm of "General Principles—Specific Principles—Supplementary Principles." If a document simultaneously has multiple levels of headings with hierarchical progression, it is determined to be a hierarchical structured document; if it only has headings but the hierarchy is incomplete, it is classified as a hierarchical framework document; if it has neither hierarchy nor normative naming, it is identified as a linear list of documents. The model output is usually a structural pattern label, clearly indicating the category of each document, providing direct input for subsequent differential processing.

[0173] The technical advantage of this method lies in its automation and generalization capabilities. Traditional rule-matching methods require writing a large number of regular expressions or parsing rules for each structure, resulting in high maintenance costs and difficulty in adapting to format variations. Large language models, leveraging their pre-trained language understanding capabilities, can flexibly identify various variations of structural features. Even in edge cases involving non-standard numbering or inconsistent title naming, they can make reasonable judgments through semantic reasoning. Furthermore, the model can process multiple documents in parallel, significantly improving recognition efficiency and solving the problem of time-consuming and labor-intensive manual review. Through an instruction-driven approach, domain experts only need to define clear category standards without intervening in the specific recognition process, lowering the technical application threshold.

[0174] Overall, this application's embodiments transform structural pattern recognition, a process reliant on human experience, into an automated model task, ensuring consistent classification standards and traceable judgment processes. The recognition results directly impact the selection of subsequent text segmentation strategies, making their accuracy crucial. Utilizing a large language model for structural identification not only improves processing efficiency but also guarantees classification quality through semantic-level analysis, providing key support for the refinement and intelligentization of the entire process of organizing key information in normative texts.

[0175] Thus, through a progressive process of elimination, all normative documents have been assigned clear structural pattern labels. This classification not only provides a direct basis for the differentiated segmentation strategy in step S502, avoiding semantic damage caused by a "one-size-fits-all" approach, but more importantly, through structural standardization, it enables subsequent information extraction and archiving to be more efficient, significantly improving the relevance and accuracy of the entire processing flow.

[0176] In step S502 of some embodiments, for each normative document, text paragraph segmentation is performed based on the corresponding domain text structure pattern to obtain multiple corresponding normative text segments.

[0177] It should be noted that each document is segmented into multiple standardized text segments based on the identified structural patterns. Specifically, for structured documents, this embodiment segments according to heading levels to ensure each segment corresponds to a complete semantic unit; for semi-structured documents, semantic continuity is added on top of the hierarchical level to prevent excessive segmentation from causing information breaks; for unstructured documents, natural language processing techniques are used to identify topic boundaries for segmentation. This differentiated processing method preserves the original logical relationships and contextual coherence of the text, making the segmented segments both independent and complete yet interconnected. It solves the problem of fixed-length or delimiter-based segmentation destroying semantics and causing knowledge fragmentation, providing high-quality input for model extraction.

[0178] Reference Figure 7According to some embodiments of this application, step S502, for each normative document, performs text segmentation based on the corresponding domain text structure pattern to obtain multiple corresponding normative text segments, which may include:

[0179] Step S701: When the domain text structure pattern belongs to the hierarchical structure text category, the context headings of the normative text are segmented to obtain multiple corresponding normative text segments.

[0180] Step S702: When the domain text structure pattern belongs to the hierarchical framework writing category, hierarchical indexing and segmentation are performed on the normative text to obtain multiple corresponding normative text segments.

[0181] Step S703: When the domain text structure pattern belongs to the linear list of text categories, semantic segmentation is performed on the normative text to obtain multiple corresponding normative text segments.

[0182] In step S701 of some embodiments, when the domain text structure pattern belongs to the hierarchical structure text category, the context headings of the normative text are segmented to obtain multiple corresponding normative text segments.

[0183] It should be noted that for hierarchical document types, a contextual heading segmentation method is used. These documents possess a complete and progressive heading system, such as a hierarchical structure of "Chapter 1—Section 1—Article 1," with each heading having a clear name. This embodiment directly uses these headings as natural segmentation points to divide the text into several independent segments. Each segment begins with a heading and ends before the next heading, ensuring the integrity of the content and the clarity of its boundaries. The advantage of this segmentation method is that it fully respects the original organizational logic of the text, achieving high-precision segmentation without complex algorithms, preserving the document's inherent structure and contextual relationships to the greatest extent, and effectively avoiding semantic breaks caused by mechanical segmentation.

[0184] In step S702 of some embodiments, when the domain text structure pattern belongs to the hierarchical framework writing category, hierarchical indexing and segmentation are performed on the normative text to obtain multiple corresponding normative text segments.

[0185] It should be noted that a hierarchical indexing method is used for hierarchical framework document categories. These documents, while having titles, lack a complete structure and may have missing subheadings or inconsistent naming conventions. The hierarchical indexing not only relies on existing titles for initial segmentation but also identifies potential structural connections by analyzing the hierarchical relationships between titles. For parts with missing titles, this embodiment determines their appropriate level based on contextual semantics, and if necessary, backtracks and merges overly subdivided paragraphs to prevent information fragmentation due to insufficient structural information. This method overcomes the limitations of contextual title segmentation in scenarios with incomplete titles. Through dual semantic and structural verification, it achieves accurate parsing of semi-structured documents, ensuring that the segmentation results both conform to the actual text and meet the granularity requirements of subsequent processing.

[0186] In step S703 of some embodiments, when the domain text structure pattern belongs to the linear list of text categories, semantic segmentation is performed on the normative text to obtain multiple corresponding normative text segments.

[0187] It's important to note that semantic segmentation is used for linearly listed text categories. These documents are typically arranged linearly only by numbers like "Item 1," "Item 2," etc., lacking headings, making it difficult to determine segmentation boundaries using traditional methods. Semantic segmentation utilizes natural language processing technology to automatically identify implicit topic boundaries in the text by analyzing thematic relevance, semantic coherence, and logical transitions between sentences. Even without headings, it can group several articles discussing the same topic into a single paragraph, thus avoiding thematic fragmentation caused by mechanically segmenting by number. This method overcomes the limitations of heading dependence, empowering the system to process purely unstructured text, ensuring that each segmented paragraph remains relatively thematically focused, providing semantically complete input units for subsequent information extraction.

[0188] This application's embodiments employ a categorized approach to refine paragraph segmentation, respecting the original structural features of the text while incorporating semantic analysis to compensate for structural deficiencies. The resulting output of multiple standardized text segments achieves a balance between completeness, independence, and relevance. This processing strategy effectively addresses the problems of knowledge fragmentation and semantic coherence disruption, providing a reliable preprocessing guarantee for building a high-quality key information database.

[0189] In step S503 of some embodiments, multiple normative text segments are input into a large language model for text extraction to obtain domain text information corresponding to each target domain.

[0190] It should be noted that step S503 inputs the segmented normative text segments into a large language model for text extraction. Compared to traditional methods that directly feed the entire document into the model, segment-level input allows the model to focus on smaller-granular information units, significantly improving extraction accuracy. Based on the key information types preset in step S101, the model identifies and extracts corresponding content within each segment. For example, it extracts guiding principles and scope of application from the "General Principles" segment and specific implementation clauses from the "Specific Principles" segment. Since the input segments already have structural labels, the model can automatically annotate the information source hierarchy during output, achieving a binding between content and structure. This approach leverages the advantages of deep semantic understanding while reducing the uncertainty of model output through structural constraints, avoiding excessive reliance on expert subjective interpretation, and improving the objectivity and reproducibility of the extraction process.

[0191] Overall, the embodiments of this application achieve classification processing through structure recognition, ensure semantic integrity through intelligent segmentation, and improve accuracy and controllability through segment extraction. These three aspects work together to solve the technical bottlenecks of traditional text parsing, such as lack of specificity, crude segmentation, and fragmented knowledge. This enables the large language model to output structured, traceable, and high-precision domain text information in the processing of normative documents, laying a data foundation for subsequent index configuration and archiving.

[0192] According to some embodiments of this application, step S102, which involves inputting multiple normative documents into a large language model for text extraction to obtain domain text information corresponding to each target domain, may include:

[0193] Based on the key information type corresponding to each target domain, construct key information extraction instructions;

[0194] Input the key information extraction instructions and multiple normative documents into the large language model so that the large language model can extract the domain text information corresponding to each target domain from the multiple normative documents.

[0195] It should be noted that, in this embodiment, step S102 is explicitly implemented as an instruction-driven mode. By constructing key information extraction instructions and having them executed by a large language model, the goal of accurately extracting domain text information from normative documents is achieved. This method changes the traditional processing logic that relies on manually setting rules or the model's free understanding, instead using structured instructions to guide model behavior, achieving a balance between automation and controllability.

[0196] The construction of key information extraction instructions must be based on the key information types specified in step S101, transforming abstract information requirements into executable extraction specifications for the model. Specifically, the instructions should clearly define the connotation and denotation of each type of key information. For example, specifying "scope of application" refers to the geographical scope, subject type, and time validity of the document, while "approval process" covers specific stages such as application, acceptance, review, and decision. Instructions can also specify the output format, such as requiring the model to return results in JSON or Markdown table format, including field names, original text snippets, and location. Simultaneously, extraction principles must be set, requiring the model to prioritize extracting sentences that directly express key information while maintaining semantic integrity, avoiding over-inference or generalization. The instruction construction process is typically completed collaboratively by domain experts and technical engineers to ensure precise alignment between business needs and technical implementation, providing the model with clear and unambiguous operational guidelines.

[0197] After inputting the key information extraction instructions along with multiple normative documents into the large language model, the model enters the understanding and execution phase. The model first parses the instructions, identifies the objectives, scope, and constraints of the extraction task, and forms an internal execution plan. Then, the model reads each normative document one by one, scanning the entire text under the guidance of the instruction framework to locate text fragments matching the type of key information. During the extraction process, the model uses its semantic understanding capabilities to judge the relevance of fragments, distinguishing between direct explanations and background information, and avoiding the extraction of redundant information. For complex expressions, the model can appropriately merge them according to the instruction requirements, such as integrating approval time limits scattered in multiple places into a complete process description. Finally, the model outputs the extraction results in a preset format, generating a structured domain text information for each document, clearly marking the original text content and source corresponding to each key point, facilitating subsequent traceability and verification.

[0198] In step S103 of some embodiments, a key information index is configured for each target domain according to the key information type corresponding to each target domain;

[0199] It should be noted that the core task of step S103 is to configure a structured key information index for each target domain based on the pre-determined key information types, thereby establishing a clear mapping relationship between the extracted text information and the knowledge system. This step connects the previous and subsequent steps, transforming the domain text information extracted in the previous steps from a loose state into organized and searchable knowledge units, laying the foundation for subsequent archiving and efficient utilization.

[0200] In practical implementation, configuring the key information index requires a hierarchical and systematic index framework based on the key information types specified in step S101. Configuring the key information index is not simply about attaching tags; rather, it involves designing a multi-level index system according to the inherent logical structure of domain knowledge, ensuring that each piece of key information has a unique and fixed position in the index tree. The index configuration process can leverage the experience of domain experts, combining the general structural characteristics of standardized texts with the specific business needs of the domain, to ensure that the index framework is both scientifically sound and practically valuable.

[0201] From an organizational perspective, the configuration of key information indexes follows the principle of moving from abstract to concrete and from general to specific. First, common index dimensions for each target domain are determined, such as scope of application, basic principles, principal responsibilities, and procedural requirements. Then, individual index items are added based on the specific characteristics of each domain. The index itself contains clear semantic definitions and scope delineations, avoiding overlap or duplication of meaning between different index items. Through standardized configuration, textual information from different periods and publishing institutions can be categorized according to a unified index framework, resolving organizational barriers caused by differences in text formats and enabling the centralized association of similar information that was originally scattered across various files.

[0202] This step plays a pivotal role in the entire method, facilitating the transition from text extraction to knowledge management. Even when extracted, unindexed text information remains difficult to locate quickly and retrieve in batches. After configuring the index, each piece of domain-specific text is assigned clear coordinates, facilitating not only the automated archiving in step S104 but also providing a direct path for subsequent retrieval queries. The index system essentially constitutes the directory structure of the knowledge base, determining the final organizational form and efficiency of the information repository. A well-configured index can significantly improve the accuracy of information retrieval; users can quickly obtain the required content simply by locating it through the index, without needing to traverse all the text.

[0203] Overall, step S103 effectively addresses the issues of knowledge fragmentation and lack of correlation through systematic index configuration. It transforms the extracted text information from isolated fragments into interconnected knowledge network nodes. This structured organization overcomes the limitations of traditional analysis tools in handling unstructured text, provides clear domain constraints for the application of large language models, reduces uncertainty in the analysis process, and significantly improves the automation level and practical value of the entire method.

[0204] In step S104 of some embodiments, based on the key information index configured for each target domain, key information is divided and archived for the corresponding domain text information to form a standardized text key information database.

[0205] It should be noted that step S104 is the final step of the method. Its core task is to systematically divide and archive the domain text information extracted in the previous steps according to the configured key information index, ultimately forming a well-structured and easily accessible standardized text key information database. This step completes the final transformation from raw text to usable knowledge assets, making the entire processing flow a closed loop.

[0206] In the specific implementation process, this step, based on the index framework established in step S103, accurately categorizes the text information extracted from each target domain. This application traverses all domain texts, analyzes their content attributes segment by segment, and assigns them to the corresponding index nodes. For example, a paragraph in a document about the redevelopment of inefficient land involving compensation standards will be automatically categorized under the index path "Real Estate - Compensation and Resettlement"; content involving approval time limits will be categorized under the node "Transaction Arrangement - Approval Process". This archiving is not a simple accumulation of documents, but a precise matching based on semantic understanding, ensuring that each piece of information can find its accurate position in the knowledge system. The archiving process can be implemented using automated scripts, completing batch classification by comparing the semantic similarity between text vectors and index nodes, while retaining a manual review interface to correct any questionable classifications.

[0207] This step effectively solves the problems of knowledge fragmentation and lack of index association. Isolated text fragments generated by traditional methods are reorganized into interconnected knowledge networks through the linking of the index system. Information under the same index node comes from different files but has the same theme, forming horizontal aggregation; the indexes at different levels form vertical logical chains, restoring the inherent structural relationships of the normative text. This organizational method makes knowledge no longer scattered points, but a hierarchical and contextual system. Users can quickly locate relevant information clusters based on the index when querying, rather than blindly searching through numerous files.

[0208] Meanwhile, step S104 overcomes the limitations of analytical tools and the problem of over-reliance on expert subjective interpretation. The archiving process is algorithm-driven and automatically executed according to a preset index framework, reducing the uncertainty caused by human intervention. The resulting key information database is stored in structured data format, which can be directly connected to various analytical tools. Whether it is statistical software or visualization platforms, they can all be processed in batches based on standardized indexes, expanding the application boundaries of traditional tools. In addition, although the index framework requires expert participation in the initial design, once established, it becomes an objective standard. Subsequent archiving no longer depends on individual expert judgment, and the results of different personnel's operations are consistent, significantly improving the objectivity and universality of the method.

[0209] The resulting standardized text key information database is highly organized and scalable. It uses vector or relational databases as its carriers, with each record containing multi-dimensional information such as the original text, its domain, key point index, and semantic tags. It supports multi-condition combined queries and rapid retrieval. This standardized text key information database not only serves for immediate information retrieval but also provides a standardized data foundation for applications such as long-term trend analysis, comparative studies of standardized documents, and intelligent question-and-answer programs, achieving efficient integration and sustainable utilization of standardized text-related information resources.

[0210] Reference Figure 8 According to some embodiments of this application, after step S104, which involves dividing and archiving the key information based on the key information index configured for each target domain to form a standardized text key information database, the method may further include:

[0211] Step S801: In the normative text key information database, vectorize the text information for each field.

[0212] Step S802: After vectorizing the text information in each field, a standardized text question-and-answer program is generated based on the standardized text key information database.

[0213] After forming a standardized text key information database, the embodiment further adds two key steps: vectorization processing and question-and-answer program generation. This upgrades the static knowledge base into an interactive intelligent question-and-answer program, significantly improving the convenience and depth of information utilization.

[0214] In step S801 of some embodiments, the text information of various fields is vectorized in the normative text key information database;

[0215] It should be noted that the core of vectorizing text information from various domains in the knowledge base lies in converting unstructured text content into numerical vector representations. Specifically, this embodiment calls a pre-trained embedding model (such as Nomic-Embed-Text or similar large language model embedding interfaces) to map each key information segment into a dense vector in a high-dimensional space. This transformation preserves the semantic features of the text, enabling computers to understand its meaning through mathematical calculations rather than keyword matching. The vectorized text information is stored in a specialized vector database (such as Qdrant, Milvus, etc.), which is optimized for fast similarity retrieval of high-dimensional vectors and supports multiple metrics such as cosine similarity and Euclidean distance. The significance of this processing is that it endows the knowledge base with semantic retrieval capabilities. Even if a user's query is not entirely consistent with the description in the database, this embodiment can still find semantically similar content through vector similarity calculation, greatly improving the accuracy and flexibility of retrieval and solving the problems of incomplete and inaccurate information retrieval caused by traditional methods relying solely on literal matching.

[0216] In step S802 of some embodiments, after vectorizing the text information in various fields, a standardized text question-and-answer program is generated based on a standardized text key information database.

[0217] It should be noted that the standardized text question-answering program generated based on the vectorized knowledge base completes the final transformation from knowledge storage to intelligent service. This program can employ a Retrieval Augmentation (RAG) architecture. When a user asks a question, this embodiment first vectorizes the question, quickly retrieves the most relevant key information segments from the vector database, and then inputs these segments as context into a large language model, which generates an answer by synthesizing the context. This mechanism avoids the illusion problem that may arise from the large language model relying solely on its internal knowledge to answer, ensuring that each answer has a clear knowledge source. The generated question-answering program can take various forms, such as web services, desktop applications, or API interfaces, and supports natural language interaction. Users do not need to master complex search syntax; they can obtain accurate answers simply by asking questions in everyday language. Compared to the limitations of traditional analysis tools, which are limited in function and rely on expert interpretation, this program automates, simplifies, and democratizes knowledge retrieval, significantly lowers the information access threshold, and improves the real-time performance and accuracy of decision support.

[0218] The value of the embodiments in this application lies in their transformation of prior knowledge organization into intelligent applications that can directly serve practical needs. Vectorization processing endows the knowledge base with deep semantic understanding capabilities, while the generation of question-answering programs provides a user-friendly human-computer interaction interface. The combination of the two forms a complete technical closed loop of "collection-organization-storage-retrieval-generation." This not only solves the original problems of information dispersion, difficulty in querying, and reliance on manual labor in normative text processing, but also realizes the efficient reuse and value amplification of knowledge through technical means. It upgrades the normative text key information base from a static database to a dynamic intelligent question-answering program, providing industry users with a solution from information acquisition to decision support.

[0219] Reference Figure 9 , Figure 9 This illustration shows the hardware structure of an electronic device according to another embodiment. The electronic device may include:

[0220] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0221] The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 to execute the method for organizing key information of normative text according to the embodiments of this application.

[0222] The input / output interface 903 is used to implement information input and output;

[0223] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0224] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0225] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0226] This application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the aforementioned method for organizing key information in normative text.

[0227] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0228] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0229] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0230] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0231] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium may include: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code.

[0232] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.

[0233] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A method for organizing key information in normative texts, characterized in that, include: Obtain the normative documents corresponding to multiple target domains, as well as the key information types corresponding to each target domain; Determine the domain text structure pattern corresponding to each of the aforementioned normative documents; specifically including: Title hierarchy identification is performed on the aforementioned normative documents; If, in the title hierarchy identification, it is determined that the normative document contains at least two levels of text titles, and the text titles satisfy the hierarchical progression condition, then the hierarchical structure text category is determined as the domain text structure pattern. If, during the title level identification process, it is determined that the normative document does not contain at least two levels of the aforementioned document titles, or that the aforementioned document titles do not meet the preset hierarchical progression conditions, then title normative identification is performed on the normative document. If, in the title standardization identification, it is determined that the title of the standard document meets the preset writing paradigm conditions, the hierarchical framework writing category is determined as the domain text structure pattern; If, during the title standardization identification, it is determined that the standardization document does not meet the writing paradigm conditions, the linear listing of writing categories will be determined as the domain text structure pattern. When the text structure pattern of the domain belongs to the hierarchical structure text category, the normative text is segmented into context headings to obtain multiple corresponding normative text segments; When the domain text structure pattern belongs to the hierarchical framework writing category, hierarchical indexing and segmentation are performed on the normative text to obtain multiple corresponding normative text segments. When the text structure pattern of the domain belongs to the linear list of text categories, semantic segmentation is performed on the normative text to obtain multiple corresponding normative text segments; Multiple normative text segments are input into a large language model for text extraction to obtain domain text information corresponding to each target domain. Configure a key information index for each target domain according to the key information type corresponding to each target domain; Based on the key information index configured for each target domain, the key information of the corresponding domain text information is divided and archived to form a standardized text key information database.

2. The method according to claim 1, characterized in that, The acquisition of normative documents corresponding to multiple target domains includes: Obtain the key information organization requirements for each target domain and the domain key fields corresponding to the key information organization requirements; Determine the target search engine that meets the requirements for organizing the key information, and the corresponding retrieval syntax rules of the target search engine; Based on the requirements for organizing key information and the search syntax rules, the key fields of the domain are integrated to form key search statements for the domain. The key search terms for the target domain are input into the target search engine to search for normative documents, thereby obtaining the normative documents corresponding to each target domain.

3. The method according to claim 2, characterized in that, The process of integrating the key fields of the domain into domain key search statements based on the requirements for organizing the key information and the search syntax rules includes: Based on the aforementioned information organization requirements, the key fields in the aforementioned fields are divided into full fields, selected fields, and irrelevant fields; Based on the retrieval syntax rules, construct logical retrieval statements for the full range of fields; Based on the retrieval syntax rules, construct or logical retrieval statements for the selected fields; Based on the retrieval syntax rules, construct non-logical retrieval statements for the irrelevant fields; Based on the retrieval syntax rules, the AND logical retrieval statement, the OR logical retrieval statement, and the non-logical retrieval statement are integrated to obtain the domain key retrieval statement.

4. The method according to claim 2, characterized in that, Before inputting the domain-specific key search terms into the target search engine for normative document searching, the method further includes: Obtain the authoritative domain name field corresponding to each of the target domains; Construct authoritative domain constraint statements based on the authoritative domain field corresponding to each of the target domains; The step of inputting the key search terms of the domain into the target search engine to perform a normative document search, and obtaining the normative documents corresponding to each target domain, includes: The key search terms for the domain and the authoritative domain constraint terms are input into the target search engine to search for normative documents, thereby obtaining the normative documents corresponding to each target domain.

5. The method according to claim 1, characterized in that, After the key information index configured for each target domain is used to divide and archive the corresponding domain text information to form a standardized text key information database, the method further includes: In the aforementioned normative text key information database, the text information for each of the aforementioned fields is vectorized. After vectorizing the text information in each of the aforementioned domains, a standardized text question-and-answer program is generated based on the standardized text key information database.

6. An electronic device, characterized in that, include: The system includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method for organizing key information of normative text as described in any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that, The storage medium stores a program that is executed by a processor to implement the method for organizing key information of normative text as described in any one of claims 1 to 5.