Method and system for automatic generation and multi-dimensional review of standard documents based on large models
By using a large-model-based method for automatic generation and multi-dimensional review of standard documents, the problems of low efficiency and unstable quality in traditional standard document writing have been solved. This method achieves intelligent generation and automated review, ensuring the accuracy and standardization of standard documents and improving writing efficiency and security.
Patent Information
- Application Number
- CN202511093418.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Traditional standard document writing is inefficient and of inconsistent quality. Manual data collection is time-consuming and poses safety hazards, failing to effectively guarantee the safety of the charging process.
The method for automatic generation and multi-dimensional review of standard documents based on large models establishes a distributed database of multi-source documents, performs natural language processing and entity recognition, forms a standardized knowledge network, extracts a structured parameter library, automatically generates standard document outlines, and conducts multi-dimensional review to ensure the accuracy and standardization of the content.
It enables intelligent generation and automated review of standard documents, shortens the cycle from conception to draft, ensures data accuracy and standardization, reduces reliance on the professional experience of writers, and improves the quality and security of standard documents.
Smart Images

Figure CN120597846B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for automatic generation and multi-dimensional review of standard documents based on large models. Background Technology
[0002] Traditional methods have some drawbacks in the writing of standard documents. Traditional standard writing requires a lot of manual work to collect, organize and analyze data. Taking the writing of the standard for charging interfaces of new energy vehicles as an example, the writers have to select useful information from various documents, industry reports and past relevant standards. Just collecting the charging interface technical parameters of different car companies and the existing national electrical safety standards may take several weeks, which affects the progress of standard writing.
[0003] Furthermore, due to differences in the professional level and experience of the writers, the quality of the standards varies. Experienced electrical engineers may comprehensively consider key safety indicators such as the withstand voltage of the interface and leakage protection, while inexperienced writers may overlook some important parameters, resulting in safety hazards in the actual application of the standards and failing to effectively guarantee the safety of the charging process. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide a method and system for automatic generation and multi-dimensional review of standard documents based on large models, thereby improving the efficiency and quality of standard document generation and enhancing accuracy through multi-dimensional review.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, a method for automatic generation and multi-dimensional review of standard documents based on a large model, the method including:
[0007] Step 1: Establish a distributed database of multi-source documents and analyze heterogeneous texts through natural language processing to obtain a standardized knowledge network;
[0008] Step 2: Based on the standardized knowledge network, extract indicator elements and form a structured parameter library through verification and validation;
[0009] Step 3: Based on the structured parameter library, build a template library, analyze user needs by combining semantic matching, and automatically generate standard document outlines;
[0010] Step 4: Based on the standard document outline, retrieve similar document fragments from the knowledge base, and combine semantic generation to complete the writing of standard chapter content, generating a standard document draft;
[0011] Step 5: Based on the draft standard document, convert key data into point cloud data, determine benchmark points and delineate the detection area; set detection points inside and outside the area, construct an arc-shaped detection path and generate path correction parameters to obtain review conclusions and modification suggestions;
[0012] Step 6: Integrate the review conclusions and modification suggestions into the knowledge base, and adjust the parameter weights and generation rules through machine learning mechanisms.
[0013] Secondly, a standard document automatic generation and multi-dimensional review system based on a large model includes:
[0014] The knowledge network module is used to build a distributed database of multi-source documents and analyze heterogeneous texts through natural language processing to obtain a standardized knowledge network.
[0015] The parameter extraction module is used to extract indicator elements based on a standardized knowledge network and form a structured parameter library through verification and validation.
[0016] The document outline module is used to build a template library based on a structured parameter library, analyze user needs by combining semantic matching, and automatically generate standard document outlines.
[0017] The document content module is used to retrieve similar document fragments from the knowledge base based on the standard document outline, and combine semantic generation to complete the writing of standard chapter content and generate a standard document draft;
[0018] The audit and evaluation module is used to convert key data into point cloud data based on the standard document draft, determine the benchmark point and delineate the detection area; set detection points inside and outside the area, construct an arc-shaped detection path and generate path correction parameters to obtain audit conclusions and modification suggestions;
[0019] The knowledge feedback module is used to integrate the review conclusions and modification suggestions into the knowledge base, and adjust the parameter weights and generation rules through machine learning mechanisms.
[0020] Thirdly, a computing device, comprising:
[0021] One or more processors;
[0022] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0023] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0024] The above-described solution of the present invention has at least the following beneficial effects:
[0025] By constructing a standardized knowledge network, the system automates the integration of heterogeneous texts from multiple sources, such as policy documents, patent literature, and experimental reports, alleviating the tedious process of manually collecting and screening materials one by one. Leveraging a structured parameter library and semantic matching technology, the system achieves intelligent-driven processes throughout the entire process, from requirements analysis to outline generation and chapter writing, shortening the cycle from conception to draft completion of standard documents and alleviating the core pain points of inefficiency and time-consuming manual writing. Through format verification and cross-document cross-validation mechanisms, the system accurately identifies conflicting descriptions in multi-source data and corrects them to consistent core parameter benchmark values, ensuring the data accuracy of the standard content. An innovative review framework of detection points, sub-regions, and dynamic adjustment values is designed to conduct automated reviews from multiple dimensions, including the completeness of essential elements (scope definition, core clauses, etc.), the standardization of professional expressions (terminology, symbols, etc.), and the logical coherence of citations (document, chart numbering, etc.), thereby improving the structural rationality and professional rigor of the standard text.
[0026] By mapping structured parameters to a template library, the system automatically generates standardized terminology definitions, quantitative indicator clauses, and verification method clauses, transforming complex standardization rules into visual generation logic and reducing reliance on the professional experience and industry background of writers. The standardization knowledge network integrates cross-domain and cross-type standardization resources, achieving efficient reuse of the parameter system through case matching and fragment recombination technologies, avoiding repetitive work. Knowledge feedback continuously integrates review results, manual revision trajectories, and error distribution data through machine learning, dynamically adjusting conflict weight thresholds, semantic matching rules, and template mapping relationships, enabling it to adapt to different industry standards and new specification requirements. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating the method for automatic generation and multi-dimensional review of standard documents based on a large model, provided by an embodiment of the present invention.
[0028] Figure 2 This is a schematic diagram of a standard document automatic generation and multi-dimensional review system based on a large model provided by an embodiment of the present invention. Detailed Implementation
[0029] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0030] like Figure 1As shown, embodiments of the present invention propose a method for automatic generation and multi-dimensional review of standard documents based on a large model. The method includes the following steps:
[0031] Step 1: Establish a distributed database of multi-source documents and analyze heterogeneous texts through natural language processing to obtain a standardized knowledge network;
[0032] Step 2: Based on the standardized knowledge network, extract indicator elements and form a structured parameter library through verification and validation;
[0033] Step 3: Based on the structured parameter library, build a template library, analyze user needs by combining semantic matching, and automatically generate standard document outlines;
[0034] Step 4: Based on the standard document outline, retrieve similar document fragments from the knowledge base, and combine semantic generation to complete the writing of standard chapter content, generating a standard document draft;
[0035] Step 5: Based on the draft standard document, convert key data into point cloud data, determine benchmark points and delineate the detection area; set detection points inside and outside the area, construct an arc-shaped detection path and generate path correction parameters to obtain review conclusions and modification suggestions;
[0036] Step 6: Integrate the review conclusions and modification suggestions into the knowledge base, and adjust the parameter weights and generation rules through machine learning mechanisms.
[0037] In this embodiment of the invention, a distributed database is established by integrating various types of documents such as policies, patents, and experimental reports, solving the problems of traditional document fragmentation and difficulty in linkage, and realizing centralized management of cross-domain knowledge. Natural language processing is used to complete entity recognition, semantic association, and data fusion, allowing originally isolated texts to form an interconnected knowledge network, avoiding content bias caused by incomplete data. Based on the knowledge network, core indicator elements are extracted, and the data presentation format is standardized through format verification to reduce problems such as missing fields or chaotic formatting. Cross-document cross-validation identifies and corrects conflicting descriptions in multi-source data, forming unified benchmark parameters, providing structured and high-quality basic data for standard content generation, and ensuring the accuracy of standards from the source.
[0038] Based on a structured parameter library and standardized specifications, a template library is built. Through semantic matching, user needs are accurately interpreted, and standard outlines containing terminology definitions and core clause frameworks are automatically generated. This avoids the problems of arbitrary outline structure and non-compliance with industry standards in traditional manual drafting. It ensures that the outline both meets actual user needs and complies with basic standardization requirements, reducing the time cost of repeated outline adjustments. Based on the outline, similar document fragments are analyzed and retrieved, and chapter content is written using semantic analysis, replacing the tedious process of manually reviewing materials and writing chapter by chapter, shortening the generation cycle from outline to draft. Case matching enables the reuse of high-quality experience, avoiding repetitive creation and improving efficiency. This enhances the professionalism and efficiency of standard content writing; it focuses on key review points through fixed detection points and sub-region segmentation, and conducts automated review from three dimensions—element completeness, professional expression standardization, and logical coherence of citations—combined with adjustment values, accurately locating issues such as missing elements, incorrect symbols, and confused citations in drafts; compared to traditional manual review, it reduces subjective oversights, provides clear revision suggestions, and improves review efficiency and the standardization of standard documents; it integrates review results, human feedback, and other information, and adjusts parameter weights and generation rules through machine learning, continuously absorbing experience from practical applications and dynamically adapting to the writing needs of different industries and types of standards.
[0039] In a preferred embodiment of the present invention, step 1 above, establishing a distributed database of multi-source documents and analyzing heterogeneous texts through natural language processing to obtain a standardized knowledge network, may include:
[0040] In this embodiment of the invention, five major categories of text—policy documents, patent documents, experimental reports, standard documents, and internal enterprise documents—are collected through targeted acquisition (such as policy columns on government websites, patent databases, and internal enterprise document management systems) and batch import. The collected documents are categorized and labeled by type, and metadata information such as the source, publication time, and applicable field of each type is recorded to form an initial document set. A distributed storage architecture is designed based on document type, scale, and access requirements, allocating the initial document set to different storage nodes according to type or field. Each node independently stores document data of its corresponding category, while a centralized index node records the storage location, metadata, and access permissions of each document. A data synchronization mechanism between nodes ensures the consistency and accessibility of data across all storage nodes, forming a distributed database covering multiple text types, achieving centralized document management and efficient retrieval.
[0041] Heterogeneous text in a distributed database undergoes unified preprocessing. For documents in different formats (e.g., PDF, Word, TXT), plain text content is extracted using format conversion tools, removing redundant information (e.g., headers, footers, irrelevant symbols, and repeated paragraphs). The extracted plain text is then cleaned, including standardizing character encoding, correcting typos, and regulating punctuation to ensure readability and consistency, providing well-organized text material for subsequent natural language processing. Based on the preprocessed text material, and combining domain features and text semantics, entity recognition is established by manually defining entity types (e.g., regulation names and responsible entities in policy documents; technical terms and inventors in patent documents; indicator names and experimental equipment in experimental reports). Rule base: Based on the rule base, the text is scanned sentence by sentence to identify and extract key entities that conform to the type definition. At the same time, the occurrence position, context and document type of each entity in the text are recorded to form a preliminary entity list. For the identified entity list, the contextual logical relationship between entities in the text (such as "reference", "inclusion", "causation", "implementation object" etc.) is analyzed to establish association clues. For policy documents, the relationship between the regulatory name and the responsible entity is analyzed; for patent documents, the correspondence between technical terms and application scenarios is analyzed; for experimental reports, the relationship between indicator names and test methods is analyzed, etc. By summarizing the association clues between entities in different documents, multi-dimensional association relationships between entities are sorted out to form a preliminary entity association map.
[0042] For different expressions of the same entity in different documents (such as full name and abbreviation, professional term and common name), the system compares the semantic similarity of the entity's context and the consistency of its occurrence scenarios to determine whether they are the same entity and merges them for annotation. For conflicts in entity relationships in different documents (such as the same term being associated with different test methods in patents and experimental reports), the system combines the authority of the documents and the adaptability of the domain to make a judgment, retaining reasonable relationships. Finally, the integrated entities and their relationships are organized in a unified structure to form a standardized knowledge network that includes entity types, attributes, and multi-dimensional relationships.
[0043] In a preferred embodiment of the present invention, step 2 above, which involves extracting indicator elements based on a standardized knowledge network and forming a structured parameter library through verification and validation, may include:
[0044] Step 220: Based on the standardized knowledge network, semantic analysis is used to identify and extract three core elements from the knowledge network: quantitative indicators, performance parameters, and testing methods, to obtain descriptive text fragments.
[0045] Step 221: Match the descriptive text fragment with the preset standardized format, verify missing fields and trigger an alarm;
[0046] Step 222: For the same core element that has been verified, calculate the similarity of descriptive text fragments in documents from different sources, identify and record descriptive conflicts whose conflict weights exceed a preset threshold.
[0047] Step 223: For description conflicts, cross-validation is performed by combining contextual relevance and preset rules to determine the baseline value and make numerical corrections to generate a consistent description;
[0048] Step 224: Integrate the core elements and consistent descriptions, calculate the hierarchical affiliation and add classification tags to form a structured parameter library.
[0049] In this embodiment of the invention, all documents in the standardized knowledge network are first scanned in full, and the text content is split into paragraphs. In the semantic analysis stage, the characteristic words in the text are identified sentence by sentence. For quantitative indicators, the focus is on capturing expressions containing features such as "value + unit", "percentage", "multiple", and "range". For example, "standby current: 50mA±5mA" is extracted from "the standby current of the equipment is 50mA±5mA" as a descriptive text fragment of the quantitative indicator. For performance parameters, the focus is on performance keywords such as "speed", "precision", "efficiency", and "strength". For example, "rotational precision: 0.01mm" is extracted from "the bearing rotational precision reaches 0.01mm" as a descriptive text fragment of the performance parameter. For test methods, process-related expressions such as "test steps", "measurement conditions", and "operation procedures" are locked. For example, the complete sentence "when testing, first adjust the voltage to 220V, maintain it for 30 minutes, and then record the data" is extracted as a descriptive text fragment of the test method. After extraction, the name, chapter, and page number information of the source document are marked for each text fragment.
[0050] The pre-defined standardized format specifies field requirements for different core element types. For example, quantitative indicators need to include five fields: "indicator name, benchmark value, fluctuation range, measurement unit, and measurement environment." The extracted descriptive text fragments are compared with these fields one by one. First, it checks whether there is an explicit description of the corresponding field in the text. If a text fragment only mentions "pressure is 10MPa" without mentioning the measurement environment and fluctuation range, then the fields "fluctuation range" and "measurement environment" are missing. The missing field ratio is calculated as "number of missing fields ÷ total number of fields", i.e., 2 ÷ 5 = 40%. If the threshold is set to 20%, an alarm will be triggered if 40% exceeds the threshold, and a prompt will pop up and mark the missing fields. If a text fragment is missing 1 field, the missing ratio is 20%, which does not exceed the threshold, and only the name of the missing field is marked in red.
[0051] For the same core element (such as "engine power") that passes the format check, collect the descriptive text fragments in all source documents. When calculating the similarity, first remove the function words such as "of" and "in" in the text, and retain the core words such as "power", "120kW", and "rated speed". Count the number of overlapping core words in different text fragments. The proportion of the overlapping number to the total number of core words is the similarity. For example, if both texts contain "power", "120kW", and "rated speed 3000r / min", and the core word overlap rate is 80%, then the similarity is 80%. The conflict weight calculation starts from four aspects. The numerical difference is calculated as "|text 1 value - text 2 value|÷text 1 value", such as the difference rate between 120kW and 110kW is (10÷120)≈8.3%. The range difference is calculated as "|upper limit difference|+|lower limit difference|÷total length of the standard range", such as the difference value between 20 - 30°C and 22 - 28°C is (2 + 2)÷10 = 40%. If the unit difference is of the same type of unit (such as m and cm), it is calculated according to the numerical difference after conversion. If it is of different types of units (such as kg and m), it is directly recorded as a 100% difference. The conditional attribute difference is calculated as the proportion of the number of different conditions to the total number of conditions. For example, if one test environment is "room temperature" and the other is "high temperature", the difference rate is 50%. Multiply the difference rates of the four by the preset weights (numerical 0.3, range 0.2, unit 0.3, condition 0.2) and then sum them. If the result exceeds the 50% threshold, it is recorded as a description conflict.
[0052] For the conflict description, first calculate the context relevance, and count the number of times the core element is associated with other elements in the knowledge network. For example, the associations of "engine power" with "fuel consumption" and "speed" appear 10 times. The relevance is calculated as "actual association frequency÷maximum possible association frequency" and is 80%. Combining the conflict resolution rules (such as "authoritative manuals take precedence over ordinary reports" and "the latest version takes precedence over the old version"), if one conflict text comes from an authoritative manual in 2024 (relevance 80%) and the other comes from an ordinary report in 2022 (relevance 60%), then select the value in the authoritative manual as the reference value, correct the part in the conflict description that does not match the reference value. For example, correct "110kW" in the ordinary report to "120kW", and add the note "The original 110kW was corrected to the reference value 120kW due to the lower authority of the source" to generate a consistent description.
[0053] After integrating core elements and consistent descriptions, the hierarchical affiliation is calculated according to predefined classification rules (such as "subject area → technology category → specific parameters"). It checks whether the element contains keywords such as "mechanical" or "electronics", and matches it to the "mechanical engineering" field (90% matching degree). Then it determines whether it belongs to categories such as "power system" or "transmission system", and matches it to "power system" (85% matching degree). Finally, the specific level is determined and the label "mechanical engineering - power system - engine power" is added. At the same time, the missing field markers in the verification stage and the correction records in the conflict resolution stage are entered. All information is arranged in a fixed structure of "element name → consistent description → verification value → conflict record → classification label" to form a structured parameter library.
[0054] By precisely calculating the missing field ratio during format validation, the completeness of core element descriptions can be strictly controlled. When the missing field ratio exceeds a threshold, an alert is issued promptly to prevent incomplete information from entering the parameter library, ensuring that each parameter contains necessary key information. Conflict weight calculation analyzes differences from multiple dimensions—numerical, range, unit, and condition—and determines a baseline value based on contextual relevance and resolution rules, allowing for targeted correction of conflicting descriptions. This process eliminates description contradictions from documents from different sources, making parameter values closer to reality and reducing data errors. Similarity calculation identifies description differences for similar elements, and cross-validation generates consistent descriptions, ensuring that the same core element has only one standardized representation in the parameter library, avoiding parameter discrepancies caused by different document sources. The structured parameter library ensures data consistency across the entire knowledge network; hierarchical classification tags, based on clear rules, calculate the attribution of parameters, arranging them in a logical order. Users can quickly locate the required parameters through classification tags, reducing search time and improving the ease of use of the parameter library; the structured parameter library contains information such as check values and conflict resolution records, clearly recording the entire process of parameter extraction, verification, and correction. Users can trace the source and processing details of each parameter, enhancing data credibility and transparency. The automated calculation process replaces the tedious work of traditional manual sentence-by-sentence comparison and verification. Alarm mechanisms and automatic tagging functions reduce the workload of manual investigation, and the application of conflict resolution rules also reduces the difficulty of manual judgment, improving the efficiency of parameter library construction.
[0055] In a preferred embodiment of the present invention, step 3 above, which involves constructing a template library based on a structured parameter library, analyzing user needs through semantic matching, and automatically generating a standard document outline, may include:
[0056] Step 330: Based on the classification tags and parameter system, match the preset chapter framework rules and establish a mapping relationship table between parameter classification tags and document chapters;
[0057] Step 331: Based on the mapping relationship table, extract domain keywords, standard type codes and core constraint expressions, calculate the matching degree with the template and locate the target chapter structure;
[0058] Step 332: Based on the target chapter structure, index the parameter system of matching domain keywords and standard type codes from the structured parameter library to automatically generate the terminology definition set of the standard document;
[0059] Step 333: Based on the core constraint expression, combined with performance parameters and test method data, generate quantitative indicator threshold clauses and verification method clauses;
[0060] Step 334: Integrate the terminology definition set with the core clause framework to generate a standard document outline.
[0061] In this embodiment of the invention, firstly, all hierarchical classification tags are extracted from a structured parameter library. These tags are hierarchically divided according to fields, standard types, etc. For example, first, large technical fields are divided, and then specific standard categories are divided under each field. At the same time, predefined chapter framework rules are obtained from the standardization basic specifications to clarify the topic scope and content of each chapter. Next, each hierarchical classification tag is compared with the chapters in the chapter framework rules one by one. For each classification tag, its core meaning and the scope of its content are analyzed, and then it is seen whether the corresponding chapter topic is related to it. For example, if the classification tag is "electronic device security" and the chapter topic is "security requirements", then it is determined that the two are related. Then, the related classification tags and chapters are recorded to establish a preliminary correspondence. After that, the preliminary correspondence is verified to check whether there is a situation where one classification tag corresponds to multiple chapters or one chapter corresponds to multiple classification tags. If so, the tightness of the relationship is further analyzed, and finally a unique mapping relationship is determined to form a mapping relationship table.
[0062] Based on the mapping table generated in step 330, relevant information is extracted. First, user needs and related data are analyzed through semantic rules to extract domain keywords, standard type codes, and core constraint expressions. Domain keywords are words that reflect the domain to which the current standard document belongs, such as "mechanical manufacturing" and "information technology". Standard type codes are specific codes that identify the standard type, such as codes that represent national standards, industry standards, etc. Core constraint expressions are descriptions of the core requirements of the standard.
[0063] When calculating the matching score between domain keywords and chapter topics, first determine the keywords contained in the chapter topics, then compare the domain keywords with these keywords, count the number of domain keywords that are the same as or semantically similar to the chapter topic keywords, divide by the total number of domain keywords to obtain a basic ratio, and then adjust according to the importance of the words. If important words are matched successfully, a certain score is added, and finally the matching score is obtained.
[0064] When calculating the matching weight between standard type codes and template types, first, the standard type code range corresponding to different template types in the template library is determined. The extracted standard type codes are compared with the code ranges corresponding to each template type to see which template type's code range the code belongs to. If it completely belongs to the code range of a certain template type, the matching weight is 1. If there is partial overlap, the weight is determined according to the degree of overlap. The higher the proportion of overlap, the greater the weight, and vice versa. Then, a weighted calculation is performed. The matching score is multiplied by the corresponding weight, and the matching weight between the standard type code and the template type is multiplied by the corresponding coefficient. The two are added together to obtain a comprehensive score. The chapter structures in the template library are sorted from high to low according to the comprehensive score, and the one with the highest score is the target chapter structure.
[0065] Based on the target chapter structure determined in step 331, the domains and standard types involved in the chapter structure are clarified. Parameter systems related to the current domain keywords are selected from the structured parameter library. This requires comparing the descriptive information in the parameter systems with the domain keywords, retaining those parameter systems whose descriptions contain domain keywords or are semantically related. Simultaneously, based on the standard type encoding, the parameter systems are further filtered, retaining only the parameter systems under the standard type corresponding to the encoding. The filtered parameter systems are analyzed to extract key terms, which are words that appear repeatedly in the parameter systems and have specific meanings. Then, a definition is determined for each key term. The definition content comes from the explanation and description of the term in the parameter system. These terms and their corresponding definitions are compiled into a set to form a standardized term definition set. During the compilation process, the accuracy and consistency of the term definitions must be ensured to avoid ambiguity.
[0066] First, the core constraint expression is parsed to clarify the performance parameters, test methods, and constraint requirements involved. Performance parameter data related to the core constraint expression is obtained from a structured parameter library. This data includes the value range and common values of the performance parameter under different conditions. Based on the constraint requirements in the core constraint expression, the quantified threshold is determined. For example, if the constraint is "the equipment operating temperature shall not exceed 80 degrees Celsius," then the quantified threshold clause is "the equipment operating temperature threshold is 80 degrees Celsius." When determining the threshold, the performance parameter data in the parameter library should be referenced to ensure its rationality. The highest common value for this performance parameter in the library is 75 degrees Celsius, and the threshold may be set to 80 degrees Celsius based on the constraints. For the verification method clause, it is generated based on the core constraint expression and the test method data in the parameter library. The test method data is analyzed to include information such as test steps, test equipment, and test environment applicable to the current performance parameter. This information is then organized into the content of the standardized clause to determine how to verify the threshold of the quantitative indicator. For example, if the performance parameter is temperature, and the test method data includes a method for measuring it at a specific location using a thermometer, then the verification method clause will describe in detail the steps and requirements for measuring it at the specified location using the thermometer.
[0067] The terminology definition set generated in step 332 is organized and arranged according to the importance and logical order of the terms to ensure clarity in the outline. Simultaneously, the core clause framework generated in step 333 is compiled, including quantitative indicator threshold clauses and verification method clauses, clarifying the logical relationships between clauses. The scope definition section is determined based on the target chapter structure and domain keywords, clarifying the applicable scope of the standard document, such as applicable product types and application scenarios. This is derived by comprehensively analyzing the scope covered by the themes of the target chapter structure and domain keywords. Then, the scope definition, terminology definition set, and core clause framework are integrated in a logical order. First, the scope definition is listed to clarify the document's applicable boundaries. Next, the terminology definition set ensures readers have a unified understanding of the key terms in the document. Finally, the core clause framework details the content of each clause. During the integration process, the smoothness of the connections between each part and the completeness of the content are checked, ultimately forming the standard document outline.
[0068] Through detailed calculations, such as precise calculations of the matching degree between domain keywords and chapter topics, and the fit between standard type codes and template types, the target chapter structure can be accurately located. This ensures that the generated outline highly matches user needs and standard specifications, reducing deviations caused by human error. When generating a set of normative terminology definitions, the accuracy and consistency of terminology definitions are ensured by using a parameter system that has been screened and verified in a structured parameter library. This avoids the problem of different definitions for the same term in different documents, making it easier for readers to understand and use standard documents. When generating quantitative indicator threshold clauses and verification method clauses, performance parameters and test method data in the structured parameter library are referenced. By combining core constraint expressions, the quantitative indicator thresholds are made to conform to the actual situation, and the verification method is operable, which improves the scientificity and practicality of the standard document. The entire process reduces manual intervention through automated calculation and analysis steps. From establishing mapping relationships to integrating and generating the outline, each step has a clear calculation process, which can quickly process large amounts of data and information and shorten the generation time of the standard document outline. Through a step-by-step integration process, the scope definition, terminology definition set and core clause framework are integrated in an orderly manner, ensuring that the standard document outline contains all necessary components, and that the logic between each part is clear and the connection is smooth, thus improving the overall quality of the standard document.
[0069] In a preferred embodiment of the present invention, step 4 above, which involves retrieving similar document fragments from the knowledge base based on the standard document outline and combining semantic generation to complete the writing of standard chapter content and generate a standard document draft, may include:
[0070] Step 440: Decompose the standard document outline by chapter, generate chapter feature vectors and logical position identifiers, and calculate the similarity score between the fragment set and the current chapter based on the chapter feature vectors.
[0071] Step 441: For the set of fragments whose similarity scores exceed the threshold, adjust them in conjunction with logical position identifiers to obtain the reorganized clause units;
[0072] Step 442: Match the reorganized clause unit with the structured parameter library, retrieve missing parameter items and inject quantitative indicator thresholds and test conditions to obtain clause content with complete parameters.
[0073] Step 443: Perform industry terminology compliance verification on the complete parameter clauses and generate a complete standard document draft.
[0074] In this embodiment of the invention, firstly, the standard document outline is broken down into chapters, and the title, keywords, and core content description of each chapter are extracted. For each chapter, a feature vector is generated. This process requires analyzing the text content in the chapter to determine the key concepts, terms, and relationships between them. For example, the feature vector is constructed by analyzing high-frequency words, important phrases, and their semantic relationships in the chapter. At the same time, a logical position identifier is generated for each chapter. This identifier contains the chapter's hierarchical information in the entire outline, the relationship between preceding and following chapters, etc. For example, the identifier may contain information such as the first-level heading, second-level heading, etc., of the chapter, as well as the chapter's sequential position among its peer chapters.
[0075] Based on the generated chapter feature vectors, a search is performed in a standardized knowledge base to find matching clause fragments. For each clause fragment, its similarity score with the current chapter is calculated. This calculation process involves semantic analysis of the clause fragments and chapter feature vectors. First, the keywords and topics in the clause fragments are determined. Then, these keywords and topics are compared with the content in the chapter feature vectors. The number of identical or semantically similar keywords is counted, and the importance of these keywords is considered. Important keywords will receive higher scores if they are successfully matched. Finally, the similarity score between the clause fragments and the current chapter is obtained.
[0076] For fragment sets with similarity scores exceeding a threshold, contextual connection adjustments are made using logical position identifiers. First, the expected position and contextual relationship of these fragments in the standard document are determined based on the logical position identifiers. For example, it is determined whether the fragment belongs to the introduction, main body, or conclusion, and its logical coherence with the preceding and following content. Then, the fragments are adjusted to naturally integrate into the context of the current chapter. This may include modifying some expressions in the fragments to maintain consistency with the overall style and terminology of the chapter; or rearranging the order of the fragments to ensure logical coherence. For example, if some fragments in the fragment set have a semantic progression, but their order in the original knowledge base may not conform to this progression, their order needs to be adjusted.
[0077] The reorganized clause units are matched with a structured parameter library to retrieve missing parameter items. This process requires detailed analysis of the clause unit content to identify the parts that need parameter support. For example, a clause might mention "the operating temperature of the equipment should meet the requirements" but not specify a specific temperature range. In this case, the corresponding temperature parameter needs to be retrieved from the structured parameter library. Once the missing parameter item is identified, the corresponding quantitative indicator threshold and test conditions are extracted from the structured parameter library. For example, for the aforementioned temperature parameter, the specific temperature threshold, such as "the operating temperature range of the equipment is -20℃ to 60℃," and the corresponding test conditions, such as "measured under standard atmospheric pressure using a thermometer with an accuracy of ±0.5℃." These retrieved parameter items are then injected into the clause unit to ensure that the clause content contains all the necessary parameter information, resulting in a clause with complete parameters. During the injection process, the accuracy and applicability of the parameters must be ensured, and they must match the context of the clause.
[0078] Performing industry terminology compliance checks on clauses with complete parameters requires establishing an industry terminology standard library, which contains commonly used standard terms, abbreviations, and their correct usage in the target industry. The terms in the clauses are compared with this library to check for any non-compliant terminology. If non-compliant terms are found, appropriate corrections are made, such as replacing non-standard synonyms with industry-standard terms, or checking abbreviations to ensure they conform to industry conventions. Chapters and paragraphs conforming to the target industry's expression standards are then generated and integrated according to the standard outline's chapter order. During integration, it's crucial to ensure smooth transitions and logical coherence between chapters, checking for correct citations and duplicate or contradictory content, ultimately generating a complete draft standard document.
[0079] By automating the chapter breakdown, matching retrieval, and content generation process, the time and workload of manually writing standard documents are reduced. Relevant content can be quickly found and integrated from the knowledge base, improving document generation speed. Parameter injection based on a structured parameter library ensures the accuracy of quantitative indicators and test conditions included in the document. Simultaneously, industry terminology compliance verification ensures that the use of terminology in the document conforms to industry standards, improving the consistency and professionalism of the document content. Contextual connection adjustments ensure the logical coherence of the document content, making the document structure more reasonable. Parameter completeness processing ensures that the clauses in the document are complete, containing all necessary information, improving the overall quality of the document. Case matching retrieval of similar document fragments in the knowledge base fully utilizes previously accumulated knowledge and experience, avoiding repetitive work. It also ensures that the generated standard document draft has a certain degree of authority and reliability. The generated standard document draft is based on a standardized outline and knowledge base, with a clear structure and standardized content, making document maintenance and updates more convenient. Only modifications to the relevant content in the knowledge base are needed to quickly generate a new version of the standard document.
[0080] In a preferred embodiment of the present invention, step 5 above, based on the standard document draft, converts key data into point cloud data, determines benchmark points and delineates the detection area; sets detection points within and outside the area, constructs an arc-shaped detection path and generates path correction parameters to obtain review conclusions and modification suggestions, may include:
[0081] Step 550: Based on the standard document draft, extract technical parameters, test methods and performance indicators as key data, and map them to a three-dimensional coordinate system to form a point cloud set;
[0082] Step 551: Based on the point cloud set, construct the baseline and sector detection area using the point cloud of the industry's mandatory compliance indicators as the benchmark.
[0083] Step 552: Set up dual detection points within the sector detection area to verify the synergistic relationship between parameters and the matching degree between parameters and test methods;
[0084] Step 553: Based on the coordinates of the reference point, dual detection points and single detection points, connect them according to the original time sequence to generate an arc-shaped detection path, and analyze the path curvature change rate to generate path correction parameters.
[0085] Step 554: Integrate the parameter fluctuation information reflected by the path correction parameters with the content of the draft standard document, and perform a comprehensive verification to obtain multi-dimensional verification results;
[0086] Step 555: Convert the path correction parameters into parameter compliance risk levels, and combine them with the logical conflicts, terminological deviations and structural omissions identified in the verification results to obtain the audit conclusions and modification suggestions.
[0087] In this embodiment of the invention, a text recognition tool is used to comprehensively analyze the draft standard document chapter by chapter. During the analysis, technical parameters such as product specifications and material performance parameters, testing methods including experimental procedures and operating specifications, and performance indicators such as product pass rate and service life are accurately located and extracted from the document. These contents are identified as key data. Then, clear mapping rules are formulated for each attribute of these key data. For example, the numerical value of technical parameters, the complexity of testing methods, and the compliance requirements of performance indicators are all mapped to the X-axis, Y-axis, and Z-axis of a three-dimensional coordinate system. According to this rule, each attribute of each key data is converted into a specific coordinate value in the three-dimensional coordinate system, so that each key data corresponds to a point in three-dimensional space. When all key data has been transformed in this way, these points are gathered together to form a point cloud set, where each point clearly represents a key data in the document and all its attribute features.
[0088] From the existing point cloud set, point clouds corresponding to mandatory compliance standards within the industry are selected, and these point clouds are designated as benchmarks, serving as core references for document compliance. The Cuckoo Algorithm is introduced, treating each point in the point cloud set as a "nest," with the benchmarks representing the initial and final "nests." Simulating the cuckoo's search for a suitable nest, the correlation between each point cloud and the benchmarks is calculated. Specifically, this involves analyzing the correlation between the key data represented by each point cloud and the mandatory compliance indicators represented by the benchmarks, such as whether the parameters revolve around the mandatory indicators. The analysis includes factors such as whether the testing method is based on mandatory indicators. Through this analysis, adjacent points with a strong correlation to the benchmark are selected. Then, these adjacent points with a strong correlation are connected to the benchmark in sequence to construct a baseline. This baseline reflects the main correlation path between key data and core compliance indicators. Then, with this baseline as the central axis, a preset threshold range is expanded outward. This threshold range is determined based on the industry's conventional error tolerance range and the tightness of data correlation. After expansion, a fan-shaped detection area is formed, which concentrates all key data points that are closely related to the benchmark.
[0089] Two detection points are set up within the fan-shaped detection area. The workflow of the first detection point is to collect the technical parameters represented by all point cloud data within the area, and then conduct a comprehensive analysis of these parameters. For example, it analyzes whether the numerical combinations of different technical parameters are reasonable, and whether there are any inconsistencies caused by parameters being too high or too low, thereby verifying the synergistic relationship between parameters and ensuring that all parameters can cooperate and operate normally in practical applications. The second detection point also collects point cloud data within the area, focusing on the technical parameters and their corresponding test methods. It analyzes whether the test methods can accurately verify the technical parameters. For example, for a certain precision parameter, does the measurement accuracy of the test method match it, and does the test procedure effectively reflect the actual situation of the parameter, thereby verifying the matching degree between the parameters and the test methods. Outside the fan-shaped detection area, a single detection point is set up at a representative location based on the distribution of point clouds outside the area. This single detection point is mainly used to cover key data that are not in the core detection area but are equally important, ensuring that all key data in the document can be included in the detection scope and avoiding detection loopholes.
[0090] First, accurately record the specific coordinates of the reference point, dual detection points, and single detection points in the 3D coordinate system. Then, review the draft standard document to determine the order in which the key data represented by these points appears in the document, i.e., the original time sequence. According to this original time sequence, connect the reference point, dual detection points, and single detection points sequentially. Since these points are located in different positions in 3D space and appear in different orders, they naturally form an arc-shaped detection path after connection. This path completely reflects the presentation logic and sequential relationship of the key data in the document. Next, analyze the generated arc-shaped detection path, focusing on the curvature change of the path. The curvature change rate reflects the change in the degree of curvature of the path. When the curvature change rate fluctuates abnormally, it indicates that there may be a problem with the key data at the corresponding position. Borrowing the idea of optimizing inferior bird nests in the Cuckoo algorithm, analyze the reasons for these abnormal curvature change rates, such as inconsistencies in key data or logical incoherence. Based on the analysis results, generate path correction parameters. These parameters record in detail the position, direction, and degree of correction in the path that need to be corrected.
[0091] Key data fluctuations reflected by path correction parameters, such as abnormal changes in parameter values and inconsistencies in logical relationships, are mapped and integrated one-to-one with the specific content of the draft standard document. A comprehensive verification process is then conducted. First, rule verification is performed, comparing the document's content against relevant industry standards, specifications, and regulations to ensure compliance. For example, do technical parameters meet minimum industry standards? Do testing methods conform to industry operating procedures? Next, terminology is checked to verify that the professional terminology used in the document is consistent with industry-standard terminology, and to ensure the accuracy and standardization of terminology usage. Finally, structural analysis is performed to analyze the overall framework of the document, ensuring its logical arrangement, completeness, and the absence of structural gaps or disorganized chapters. The results of rule verification, terminology detection, and structural analysis are then integrated with the parameter fluctuation information to obtain a comprehensive, multi-dimensional verification result.
[0092] Based on the path correction parameters and industry-wide risk assessment standards, parameter fluctuations are converted into specific parameter compliance risk levels. If parameter fluctuations are small and within acceptable limits, the corresponding risk level is low. If parameter fluctuations are large, exceeding a certain range but still adjustable, the risk level is medium. If parameter fluctuations are extremely large and severely fail to meet requirements, the risk level is high. This is combined with logical conflicts identified in the multi-dimensional verification results, such as logical contradictions between technical parameters, logical mismatches between testing methods and performance indicators, terminology deviations (e.g., incorrect terminology usage, non-standard terminology expression), and structural deficiencies (e.g., missing important chapters, incomplete content). By comprehensively considering the parameter compliance risk levels and these specific issues, the overall compliance and completeness of the document are assessed, generating a tiered review conclusion, such as "passed," "requires minor modifications," or "requires major modifications." For each issue, based on the correction direction and extent indicated by the path correction parameters, specific and feasible modification suggestions are formulated, forming a corresponding modification suggestion set that clearly indicates the content to be modified, the modification methods, and the post-modification goals.
[0093] By transforming key data into point clouds and constructing 3D coordinates, combined with the Cuckoo algorithm for precise screening and correlation analysis of key data points, the inherent relationships between key data can be deeply explored. This approach breaks through the limitations of traditional manual review, which focuses on superficial textual checks. It can accurately identify deeper issues such as improper parameter coordination and mismatched testing methods, thus improving the accuracy of the review. By setting benchmark points, dual detection points, and single detection points, and constructing fan-shaped detection areas and arc-shaped detection paths, a comprehensive detection network is formed. This network not only focuses on key data in core compliance areas but also covers important data outside these areas. Combined with rule verification, terminology detection, and structural analysis, this ensures a comprehensive review of all aspects of the document, avoiding omissions in the review process. After introducing the Cuckoo algorithm, a certain degree of automation in point cloud clustering, detection point optimization, and path correction is achieved, reducing the need for manual step-by-step analysis. The tedious work of comparison and analysis, coupled with structured review steps and clear verification standards, enables reviewers to work more efficiently, shortening the review cycle and improving review efficiency. The generated set of modification suggestions is based on specific parameter fluctuation information, logical conflicts, terminology deviations, and structural deficiencies, combined with parameter compliance risk levels, making it highly targeted. Each suggestion clearly points out the problem and the direction of modification, allowing document compilers to clearly understand what needs to be modified and how to modify it, greatly enhancing the practicality of the modification suggestions. By combining point cloud technology, time series analysis, and the Cuckoo algorithm, an intelligent standard document review system has been constructed. This system changes the traditional review model that relies on human experience, realizing the digitalization and intelligence of the review process, providing new technical means and methods for the field of standard document review, and promoting the intelligent development of industry review work.
[0094] In a preferred embodiment of the present invention, step 6 above, which integrates the review conclusions and modification suggestions into the knowledge base and adjusts the parameter weights and generation rules through a machine learning mechanism, may include:
[0095] Step 660: Integrate the review results with the manual revision feedback to generate a fused feedback dataset, and associate the revision opinions with the conflict description nodes of the standardized knowledge network to generate associated node information;
[0096] Step 661: Based on the associated node information, extract the error type distribution and revision operation trajectory from the audit records to generate an error distribution dataset;
[0097] Step 662: Store the error distribution dataset in the knowledge base to form a set of audit cases with weighted labels, and analyze them through a machine learning mechanism to generate conflict weight threshold parameters.
[0098] Step 663: Based on the conflict weight threshold parameter, dynamically update the terminology compliance verification rules and adjust the mapping weights between the chapter framework and the parameter system in the template library.
[0099] In this embodiment of the invention, based on the previously generated correction parameters, a detailed weighted fusion calculation is performed on the review results and manual revision feedback. The correction parameters comprehensively consider various factors such as the semantic density and conflict records of each sub-region, acting like an "intelligent regulator." It assigns different weights to the review results and manual revision feedback according to the actual situation of different sub-regions. The weight allocation is based on strict criteria. For sub-regions with high semantic density and few conflict records, indicating relatively high content quality, the weight of manual revision feedback is appropriately reduced, as these regions may only require minimal human intervention to achieve good results. Conversely, for sub-regions with many conflict records and low semantic density, the weight of manual revision feedback is increased, as these regions require more human expertise for correction and improvement. During the weighted fusion calculation process, each piece of information in the review results and manual revision feedback is analyzed and compared in depth. For errors found in the review results and proposed revisions, the views and handling methods of the manual revision feedback on these issues are carefully studied.
[0100] If the manual revision feedback matches the review result, the credibility of this information is enhanced, and it is retained as an important reference. If the manual revision feedback does not match the review result, a judgment is not hastily made, but the reasons are further analyzed in depth. More data is consulted to see which approach is more effective in similar situations. Domain knowledge is also used to evaluate the rationality of the two opinions from a professional perspective. Through this rigorous analysis and comparison, the accuracy and reliability of the fused feedback data can be ensured. After weighted fusion calculation, a fused feedback dataset is generated. This dataset contains the integrated and verified review results and manual revision feedback information. The revision opinions in the fused feedback dataset are linked to the conflict description nodes of the standardized knowledge network. The standardized knowledge network is a vast and complex knowledge system, like a "knowledge network" that organically connects various domain knowledge and standards.
[0101] The conflict description nodes are like "key points" on this network, recording various possible conflict situations and corresponding solutions. Through semantic analysis and pattern matching, revision opinions are accurately associated with conflict description nodes. During the association process, associated node information is generated. This information records in detail the relationship between revision opinions and conflict description nodes, including the strength of the association and the basis for the association. The strength of the association reflects the closeness between the revision opinions and conflict description nodes, while the basis for the association explains why they are associated together. This associated node information provides an important basis for knowledge integration and rule adjustment, just like providing a detailed "map" for decision-making.
[0102] The association node information is deeply integrated with the fusion feedback dataset. During the integration process, advanced data matching and merging techniques are used to accurately match and organically merge each record in the association node information with the corresponding record in the fusion feedback dataset. For example, for a certain revision opinion, the conflict description node information corresponding to it in the association node information is carefully searched and added to the fusion feedback dataset to form a more complete and richer record. Through this integration, the revision opinions can be organically combined with the relevant knowledge in the standardized knowledge network. The error type distribution and revision operation trajectory are extracted from the review records. The error type distribution refers to the number and proportion of various error types found during the review process. Each error information in the review records is meticulously classified and statistically analyzed.
[0103] Errors are categorized into different types, such as terminology errors, formatting errors, logical errors, and data errors. The frequency of each type is then calculated to analyze which types are more common and which are less common. This analysis helps identify key and challenging issues in the review process. The revision operation trajectory is a detailed record of all operations performed during the revision process. Each revision operation in the review log is comprehensively and meticulously recorded, including the time, content, executor, and purpose of the operation. Analyzing the revision operation trajectory reveals behavioral habits and preferences during the revision process, such as preferred revision methods and different strategies employed when handling different types of issues. The extracted error type distribution and revision operation trajectory information are integrated to generate an error distribution dataset. This dataset contains the distribution of various error types and revision operation trajectory information during the review process.
[0104] The error distribution dataset is stored in a knowledge base, forming a weighted set of audit cases. The knowledge base is a vast knowledge storage system, like a "knowledge library," storing knowledge and audit experience from various domains. During storage, weight tags are added to each error record and revision operation. The size of the weight tags is not arbitrarily determined but allocated based on the severity of the error and the effectiveness of the revision operation. Errors that seriously affect document quality, may lead to misunderstandings, or misapplication are assigned higher weight tags because these errors require more attention and importance. Effective revision operations, such as revision methods that can quickly and accurately solve problems, are also assigned higher weight tags because these operations have high reference value. By adding weight tags, audit cases can be managed and analyzed more granularly, enabling a clearer understanding of the value and significance of each case. The stored audit case set is analyzed through machine learning mechanisms to generate conflict weight threshold parameters. The machine learning mechanism uses a variety of advanced algorithms and technologies to deeply mine and analyze the data in the audit case set. First, the machine learning algorithm conducts a comprehensive and in-depth analysis of the error type distribution and revision operation trajectory, attempting to find hidden patterns and rules.
[0105] For example, the algorithm might discover that certain types of errors frequently appear in specific chapters or areas, suggesting common problems in those chapters or areas that require special attention. Or, the algorithm might find that certain revisions are more effective in resolving specific types of errors. Based on these discovered patterns and rules, the machine learning algorithm performs complex calculations and analyses, ultimately generating weight threshold parameters for various conflict scenarios. The conflict weight threshold parameter is a crucial indicator used to determine the severity of a conflict. When the weight of a conflict exceeds this threshold, it is considered severe and requires focused attention, necessitating more resources and effort to resolve. Conversely, when the weight of a conflict is below this threshold, it is considered minor and can be handled with general consideration, requiring less attention. Through continuous learning, the machine learning mechanism can dynamically adjust the conflict weight threshold parameter based on changes in the audit case set, making it more accurately reflect the actual situation.
[0106] By applying a conflict weight threshold parameter to the semantic matching process, the terminology compliance verification rules are dynamically updated. During semantic matching, a detailed comparison and analysis of the similarity between terms in the document and industry standard terms is performed. The conflict weight threshold parameter has a significant impact on the calculation method of this similarity. For terms marked with high conflict weights, it indicates that these terms are prone to problems in use, thus increasing the strictness of their matching and requiring a higher similarity between these terms and industry standard terms. Only when a certain similarity threshold is reached will they be considered compliant. This can effectively reduce the problem of non-standard terminology usage and improve the professionalism and accuracy of the document. The terminology compliance verification rules are dynamically adjusted according to the conflict weight threshold parameter. If a term has a high conflict weight, it indicates that the term poses a significant risk in use, and the compliance verification of that term will be strengthened by adding more verification conditions and rules.
[0107] The process involves checking whether the context in which the term is used conforms to industry standards and whether there is any possibility of ambiguity or misunderstanding. It also checks whether the logical relationship between the term and other related terms is correct. Conversely, if a term has a low conflict weight, it indicates that the term is relatively safe to use, which simplifies the compliance verification of the term, reduces unnecessary verification steps, and improves verification efficiency. By dynamically updating the verification rules, compliance issues with terms in the document can be identified more accurately, improving the accuracy and efficiency of verification and ensuring that the use of terms in the document conforms to industry standards and specifications.
[0108] The conflict weight threshold parameter is applied to the template library to adjust the weight of the mapping relationship between the chapter framework and the parameter system. The template library contains various standard document templates, each of which defines the document's chapter framework and parameter system. The chapter framework specifies the document's structure and chapter divisions, while the parameter system contains various parameters and data required in the document. There is a certain mapping relationship between the chapter framework and the parameter system. This mapping relationship determines how to accurately map the content in the parameter system to the corresponding position in the chapter framework when generating the document. The conflict weight threshold parameter affects the weight of this mapping relationship. For some chapters and parameters that are found to have many conflicts during the review process, it indicates that there are certain matching problems between these chapters and parameters, and the weight of their mapping relationship is increased.
[0109] When generating documents, we pay more attention to the matching degree between these chapters and parameters, and devote more effort to ensuring that the parameter content can be accurately mapped to the appropriate chapter position to reduce the occurrence of conflicts. Conversely, for some chapters and parameters with fewer conflicts, we reduce the weight of their mapping relationship. By adjusting the weight of the mapping relationship, we can optimize the structure and function of the template library, improve the quality and efficiency of document generation, and when using the template library to generate documents, we can more reasonably allocate parameter content to each chapter according to the conflict weight threshold parameter, making the document structure more reasonable, the content more coherent, the logic stronger, reducing the occurrence of conflicts, and improving the quality of the document and the user experience.
[0110] By weighted fusion of audit results and manual revision feedback, and linking them to a standardized knowledge network, the knowledge base can be continuously enriched and improved. The knowledge in the knowledge base will be more accurate and comprehensive, better reflecting industry standards and practical applications. The machine learning mechanism's analysis and learning of audit case sets can continuously discover new patterns and rules, improving intelligence and decision-making capabilities. Based on data and current conditions, parameter weights and generation rules are dynamically adjusted to better adapt to different document audit needs. The application of conflict weight threshold parameters focuses more on high-risk areas and issues during the audit process, improving the relevance and efficiency of the audit. By dynamically updating terminology compliance verification rules and adjusting template library mapping weights, potential problems can be identified and resolved in advance, reducing manual intervention and audit costs. Through verification rules and template library mapping weights, the generated documents are more in line with industry standards and specifications, reducing errors and conflicts. The logic, coherence, and accuracy of the documents will be improved, thereby increasing their usability and influence. Parameter weights and generation rules are continuously adjusted based on the accumulated audit cases and feedback.
[0111] like Figure 2 As shown, embodiments of the present invention also provide a standard document automatic generation and multi-dimensional review system based on a large model, including:
[0112] The knowledge network module is used to build a distributed database of multi-source documents and analyze heterogeneous texts through natural language processing to obtain a standardized knowledge network.
[0113] The parameter extraction module is used to extract indicator elements based on a standardized knowledge network and form a structured parameter library through verification and validation.
[0114] The document outline module is used to build a template library based on a structured parameter library, analyze user needs by combining semantic matching, and automatically generate standard document outlines.
[0115] The document content module is used to retrieve similar document fragments from the knowledge base based on the standard document outline, and combine semantic generation to complete the writing of standard chapter content and generate a standard document draft;
[0116] The audit and evaluation module is used to convert key data into point cloud data based on the standard document draft, determine the benchmark point and delineate the detection area; set detection points inside and outside the area, construct an arc-shaped detection path and generate path correction parameters to obtain audit conclusions and modification suggestions;
[0117] The knowledge feedback module is used to integrate the review conclusions and modification suggestions into the knowledge base, and adjust the parameter weights and generation rules through machine learning mechanisms.
[0118] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0119] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0120] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0121] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for automatic generation and multi-dimensional review of standard documents based on a large model, characterized in that: The method includes: Step 1: Establish a distributed database of multi-source documents and analyze heterogeneous texts through natural language processing to obtain a standardized knowledge network; Step 2: Based on the standardized knowledge network, extract indicator elements and form a structured parameter library through verification and validation; Step 3: Based on the structured parameter library, build a template library, analyze user needs by combining semantic matching, and automatically generate standard document outlines; Step 4: Based on the standard document outline, retrieve similar document fragments from the knowledge base, and combine semantic generation to complete the writing of standard chapter content, generating a standard document draft; Step 5: Based on the draft standard document, key data is converted into point cloud data, benchmark points are determined, and the detection area is delineated. Detection points are set within and outside the area, an arc-shaped detection path is constructed, and path correction parameters are generated to obtain audit conclusions and modification suggestions. Specifically, this includes: extracting technical parameters, test methods, and performance indicators as key data based on the draft standard document and mapping them to a three-dimensional coordinate system to form a point cloud set; based on the point cloud set, using the point cloud of industry mandatory compliance indicators as benchmark points, constructing a baseline and a sector-shaped detection area; setting dual detection points within the sector-shaped detection area to verify the synergistic relationship between parameters and the matching degree between parameters and test methods; connecting the coordinates of the benchmark point, dual detection points, and single detection point according to the original time sequence to generate an arc-shaped detection path, and analyzing the path curvature change rate to generate path correction parameters; integrating the parameter fluctuation information reflected by the path correction parameters with the content of the draft standard document, performing comprehensive verification to obtain multi-dimensional verification results; converting the path correction parameters into parameter compliance risk levels, and combining them with logical conflicts, terminological deviations, and structural missing items identified in the verification results to obtain audit conclusions and modification suggestions. Step 6: Integrate the review conclusions and modification suggestions into the knowledge base, and adjust the parameter weights and generation rules through machine learning mechanisms.
2. The method for automatic generation and multi-dimensional review of standard documents based on a large model as described in claim 1, characterized in that, Based on a standardized knowledge network, indicator elements are extracted and, through verification and validation, a structured parameter library is formed, including: Based on a standardized knowledge network, semantic analysis is used to identify and extract three core elements from the knowledge network: quantitative indicators, performance parameters, and testing methods, to obtain descriptive text fragments. The descriptive text snippets are matched against a pre-defined standardized format to verify missing fields and trigger alarms. For the same core element that has been verified, calculate the similarity of descriptive text fragments in documents from different sources, identify and record descriptive conflicts whose conflict weights exceed a preset threshold; To address description conflicts, cross-validation is performed by combining contextual relevance and preset rules to determine baseline values and make numerical corrections, thereby generating consistent descriptions. The core elements and consistent descriptions are integrated, hierarchical affiliation is calculated and classification tags are added to form a structured parameter library.
3. The method for automatic generation and multi-dimensional review of standard documents based on a large model as described in claim 2, characterized in that, Based on a structured parameter library, a template library is built. Combined with semantic matching, user needs are analyzed to automatically generate standard document outlines, including: Based on the classification label and parameter system, and matching the preset chapter framework rules, a mapping relationship table between parameter classification labels and document chapters is established; Based on the mapping relationship table, domain keywords, standard type codes and core constraint expressions are extracted, the matching degree with the template is calculated and the target chapter structure is located. Based on the target chapter structure, the parameter system that matches domain keywords and standard type codes is indexed from the structured parameter library to automatically generate a set of terminology definitions for standard documents; Based on the core constraint expressions, combined with performance parameters and test method data, quantitative indicator threshold clauses and verification method clauses are generated. The terminology definition set is integrated with the core clause framework to generate a standard document outline.
4. The method for automatic generation and multi-dimensional review of standard documents based on a large model according to claim 3, characterized in that, Based on the standard document outline, similar document fragments are retrieved from the knowledge base, and semantic generation is used to complete the writing of standard chapter content, generating a standard document draft, including: The standard document outline is broken down into chapters, generating chapter feature vectors and logical position identifiers. Based on the chapter feature vectors, the similarity score between the fragment set and the current chapter is calculated. For the set of fragments with similarity scores exceeding the threshold, adjustments are made in conjunction with logical position identifiers to obtain the reorganized clause units; The reorganized clause units are matched with the structured parameter library, missing parameter items are retrieved and quantitative indicator thresholds and test conditions are injected to obtain clause content with complete parameters. Perform industry terminology compliance checks on clauses with complete parameters and generate a complete draft standard document.
5. The method for automatic generation and multi-dimensional review of standard documents based on a large model according to claim 4, characterized in that, The comprehensive verification includes rule verification, which calls a pre-set logical rule library to verify the causal rationality of technical clauses; terminology detection, which compares with a standard terminology library to verify the consistency of professional terminology expressions; and structure analysis, which verifies the integrity of core chapters through a chapter weighting mechanism.
6. The method for automatic generation and multi-dimensional review of standard documents based on a large model as described in claim 5, characterized in that, The review conclusions and modification suggestions are integrated into a knowledge base, and parameter weights and generation rules are adjusted through machine learning mechanisms, including: The review results are integrated with the feedback from manual revisions to generate a fused feedback dataset. The revision opinions are then linked to the conflict description nodes of the standardized knowledge network to generate associated node information. Based on the associated node information, the error type distribution and revision operation trajectory are extracted from the audit records to generate an error distribution dataset. The error distribution dataset is stored in a knowledge base to form a set of audit cases with weighted labels, and analyzed using a machine learning mechanism to generate conflict weight threshold parameters. Based on the conflict weight threshold parameter, the terminology compliance verification rules are dynamically updated, and the mapping relationship weight between the chapter framework and the parameter system in the template library is adjusted.
7. A standard document automatic generation and multi-dimensional review system based on a large model, wherein the system implements the method as described in any one of claims 1 to 6, characterized in that, include: The knowledge network module is used to build a distributed database of multi-source documents and analyze heterogeneous texts through natural language processing to obtain a standardized knowledge network. The parameter extraction module is used to extract indicator elements based on a standardized knowledge network and form a structured parameter library through verification and validation. The document outline module is used to build a template library based on a structured parameter library, analyze user needs by combining semantic matching, and automatically generate standard document outlines. The document content module is used to retrieve similar document fragments from the knowledge base based on the standard document outline, and combine semantic generation to complete the writing of standard chapter content and generate a standard document draft; The review and evaluation module is used to convert key data into point cloud data based on the standard document draft, determine the benchmark point and delineate the detection area; Detection points are set up inside and outside the area, an arc-shaped detection path is constructed and path correction parameters are generated to obtain audit conclusions and modification suggestions; The knowledge feedback module is used to integrate the review conclusions and modification suggestions into the knowledge base, and adjust the parameter weights and generation rules through machine learning mechanisms.
8. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Training method of controllable and credible official document generation model
CN119988648A
Standard analysis atlas construction method and system using large-scale language model
CN120012891A
Cited By
Multi-technology fusion intelligent document processing method and system
CN121682898A
A multi-technology integrated intelligent document processing method and system
CN121682898B
Large language model driven report generation agent method, system, equipment and medium
CN121859853A
Intelligent compilation method and system for scientific research achievement standard draft based on big data
CN121936431A
Intelligent compilation method and system for scientific research achievement standard draft based on big data
CN121936431B