Software development document automatic modification scheme generation method for ambiguous review comments

CN122816641APending Publication Date: 2026-09-25BEIJING INST OF COMP TECH & APPL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611030708.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-12
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0011]本发明要解决的技术问题是如何提供一种针对模糊评审意见的软件研发文档自动修改方案生成方法,以解决自动理解和执行模糊意见的问题

Benefits of technology

[0029]本发明提出一种针对模糊评审意见的软件研发文档自动修改方案生成方法,与现有技术相比,具有以下有益效果:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816641A_ABST
    Figure CN122816641A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of software development document automatic modification scheme generation method for fuzzy review opinion, belong to the cross technical field of document quality guarantee, natural language processing and automation planning in software engineering.This method of the present application includes: document structuring and reference relationship diagram construction;Fuzzy opinion semantic analysis;Problem positioning;Modification scheme generation based on HTN planning;Three-level simulation verification and sorting of candidate scheme;Modification execution and effect verification;Man-machine collaborative feedback.The present application realizes the full automatic generation from fuzzy opinion to executable modification scheme, and engineers only need one-key confirmation;HTN planning guarantees that operation sequence meets precondition and reference integrity, and three-level simulation verification and final verifier further guarantee, with the characteristics of strong explainability, can be landed, reproducible, and wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary technical field of document quality assurance, natural language processing and automated planning in software engineering, and specifically relates to a method for generating automatic modification schemes for software development documents in response to fuzzy review comments. Background Technology

[0002] In the software development process, expert review of documents such as requirements, design, and testing is a crucial step in ensuring product quality. The feedback generated after the review needs to be addressed point by point by developers or the authors, and then reviewed by the reviewers to verify the rectification. In recent years, the industry has attempted to use automated tools for regression testing of review feedback, such as comparing differences before and after document modifications and checking whether the modified locations are consistent with the feedback descriptions. However, these tools can only handle "precise instructions" (such as "change 'throughput' to 'response time' in Section 5.2"), and are powerless against the large number of ambiguous feedback comments.

[0003] Vague comments are common because reviewers often point out problems at the system level rather than providing specific solutions. For example, an expert might write, "The class diagram lacks a connection to the access control module," but won't specify which class, attribute, or method should be added. Currently, understanding and implementing such vague comments relies entirely on engineers' personal experience and domain knowledge, which is not only time-consuming but also prone to repeated revisions and multiple reviews due to misunderstandings.

[0004] The main limitations of existing technologies include:

[0005] 1. Only process explicit instructions: Existing document comparison tools and review comment tracking systems require comments to include precise location and modification content; otherwise, they cannot be processed automatically.

[0006] 2. Lack of semantic understanding: Existing methods cannot extract high-level intentions such as "what is missing" or "where is inconsistent" from natural language opinions.

[0007] 3. Lack of automatic planning capabilities: Even if the intention of the opinion is understood, no system can automatically generate a series of modification steps to meet that intention.

[0008] 4. Inability to verify the adequacy of modifications: The current tool can only check "whether modifications have been made", but cannot determine "whether the modifications have truly resolved the issues raised in the comments".

[0009] Therefore, a method is needed that can automatically understand ambiguous opinions and generate executable modification schemes. Summary of the Invention

[0010] (a) Technical problems to be solved

[0011] The technical problem to be solved by this invention is how to provide a method for automatically generating modification schemes for software development documents in response to ambiguous review comments, so as to solve the problem of automatically understanding and implementing ambiguous comments.

[0012] (II) Technical Solution

[0013] To address the aforementioned technical problems, this invention proposes a method for automatically generating modification schemes for software development documents in response to fuzzy review comments. This method includes:

[0014] Step S1: Document structuring and reference graph construction

[0015] Convert the original document into a machine-readable reference graph structure with reference relationships;

[0016] Step S2: Semantic analysis of fuzzy opinions

[0017] Defect types and key elements are extracted from the review comments text; a fine-tuned BERT model is used for multi-label classification to extract defect types and sequence labeling to extract key elements. At the same time, a rule backup mechanism based on dependency parsing is provided to handle marginal cases with low model confidence.

[0018] Step S3, Problem Identification

[0019] Based on the key elements output in step S2, find the set of seed nodes most likely to be affected in the document reference relationship graph, and determine the affected area through variable step graph propagation;

[0020] Step S4: Generation of Modification Scheme Based on HTN Planning

[0021] Based on the defect type and key elements output in step S2, and the seed node set and affected area output in step S3, a candidate modification scheme composed of operation primitives is automatically generated using hierarchical task network (HTN) planning.

[0022] Step S5: Three-level simulation verification and ranking of candidate solutions

[0023] A three-level verification system is used to quickly evaluate the structural feasibility, citation integrity, and semantic correctness of each candidate modification scheme, and output the confidence ranking.

[0024] Step S6: Modify execution and verify results

[0025] Apply the selected changes to the actual document and perform final verification;

[0026] Step S7, Human-Machine Collaborative Feedback

[0027] Record user feedback behavior to continuously optimize the model and task library.

[0028] (III) Beneficial Effects

[0029] This invention proposes a method for automatically generating modification schemes for software development documents in response to ambiguous review comments. Compared with existing technologies, it has the following advantages:

[0030] 1. High degree of automation: From vague opinions to executable modification plans, the entire process is generated automatically, and engineers only need to confirm with one click.

[0031] 2. Correctness of modifications is guaranteed: HTN planning ensures that the operation sequence meets the preconditions and reference integrity, and three-level simulation verification and the final verifier further ensure correctness.

[0032] 3. High interpretability: The decomposition process of HTN is traceable, and the reasons for each modification step are clear.

[0033] 4. Practical and reproducible: This invention discloses key technical contents such as task library definition, operation primitive mapping, and model fine-tuning details, which can be used by those skilled in the art to implement the system.

[0034] 5. Broad application prospects: It can be integrated into R&D management platforms such as Jira, GitLab, and ZenTao, and is suitable for industries with strict document quality requirements such as finance, communications, automotive electronics, and aerospace and military. Attached Figure Description

[0035] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0036] To make the objectives, contents, and advantages of the present invention clearer, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples.

[0037] This invention belongs to the interdisciplinary field of document quality assurance, natural language processing, and automated planning in software engineering. Specifically, it involves: after software development documents (such as requirements specifications, design specifications, test cases, etc.) have undergone expert review, performing semantic analysis and problem localization on ambiguous review comments that do not specify the location or method of modification, and automatically generating executable document modification plans using hierarchical task network (HTN) planning technology to assist engineers in quickly completing document rectification and ensuring the quality of rectification.

[0038] This invention not only generates modification schemes but also automatically executes the modifications and verifies the effects. This invention is applicable to editable Word documents, PDF documents containing text layers, and documents containing model elements such as UML class diagrams. For scanned PDFs, OCR processing is required first, but OCR is not within the scope of this invention. For non-text elements in the document (such as text embedded in images or embedded OLE objects), this invention does not process them and requires manual user intervention. It should be specifically noted that if model elements such as UML class diagrams and flowcharts exist in the document in a parsable structured data form (e.g., a format exported from a modeling tool containing explicit node and edge information), this method can treat them as special semantic units for parsing and processing; if they exist in an unparsable bitmap form, they are still considered non-text elements and are not processed.

[0039] (Problem scenario description: In software development, expert reviewers often give vague comments such as "class diagrams lack exception handling mechanisms"—neither specifying which chapter or page is missing, nor providing clear suggestions on how to fix it. Engineers have to repeatedly search and ponder the intent, then manually design modification schemes, often spending tens of minutes or even hours. This invention aims to automatically understand these vague comments, intelligently locate the problem in the document, and automatically generate directly executable modification schemes, freeing engineers from the heavy burden of document rework.)

[0040] The objective of this invention is:

[0041] 1. Enables fully automated conversion from vague review comments to executable modification plans, significantly reducing the understanding and design burden on engineers.

[0042] 2. Establish an opinion intent understanding model covering 8 types of defects, which can accurately extract deeper information such as missing elements and inconsistencies in opinions.

[0043] 3. Solve the unique challenge of automatically extracting natural language reference relationships in the document domain, and construct a usable document reference relationship graph.

[0044] 4. Use HTN planning technology to generate multi-step modification plans, and design a dedicated template library for the content generation requirements of document modifications to ensure the certainty and auditability of the generated content.

[0045] 5. Provides three-level simulation verification (structure, references, semantics) for document modification, quickly assessing the quality of the solution before modification, replacing the compilation verification in code repair.

[0046] Basic Concepts

[0047] 1. Vague review comments: These are review comments that do not specify any exact changes (such as chapter numbers, page numbers, or figure / table numbers) or the specific method of modification. For example, "The overall logic of the document is contradictory; please check and adjust it yourself." These comments rely entirely on human interpretation and identification.

[0048] 2. Semi-fuzzy comments (a special subclass of fuzzy comments): These comments only contain location identifiers such as chapter numbers and chart numbers, but do not specify the exact paragraph, sentence, or cell, nor do they provide precise modification text. For example: "The architecture diagram in Chapter 5 is inconsistent with the requirements; please adjust it." Although this type of comment gives a general location, the modification location is still imprecise (e.g., the entire chapter or the entire diagram), and no specific modification content is given. For semi-fuzzy comments, the system will skip keyword matching, directly locate the specified position, and limit the number of steps the diagram propagates.

[0049] 3. R&D Engineering Documents: Software R&D documents with a clear chapter structure (e.g., written according to GB / T 8567-2006 or similar standards), including but not limited to requirements specifications, software design specifications, test cases, user manuals, deployment documents, etc. Semantic unit types in the documents include sections, paragraphs, tables, and figures. For model elements such as UML class diagrams, subtypes can be defined, such as uml_class (class node) and uml_interface (interface node). Their diagram structure processing is consistent with ordinary nodes, but the node subtypes can be distinguished when generating content templates.

[0050] 4. Defect Types: Here are 8 predefined document problem categories, including: missing elements, naming conflicts, inconsistent structure, logical contradictions, non-standard formatting, incomplete descriptions, incorrect citations, and redundant content.

[0051] 5. Document Reference Graph: A directed graph where nodes are semantic units such as chapters, paragraphs, tables, and graphics in a document, and edges represent explicit references extracted from natural language text (e.g., "see Section 3.2" or "as shown in Table 2"). Unlike code dependency graphs, document references are not provided by the compiler and must be automatically extracted from natural language, which is a unique technical challenge in the document domain.

[0052] 6. Modification operation primitives: The smallest executable unit for modifying a document, consisting of 6 types: ADD_NODE, DELETE_NODE, UPDATE_TEXT, UPDATE_FORMAT, MOVE_NODE, and SPLIT_NODE. Each primitive has a defined mapping relationship with a specific document editing API.

[0053] 7. Hierarchical Task Network (HTN) Planning: An artificial intelligence planning method. High-level tasks are recursively decomposed into subtasks using predefined methods, down to atomic tasks (i.e., modification operation primitives). This invention uses HTN planning to generate candidate modification schemes.

[0054] 8. Three-level simulation validator: Includes structural checks, reference integrity checks, and semantic similarity checks, used to evaluate the quality of candidate solutions without actually modifying the document.

[0055] The core idea of ​​this invention is to decompose the cognitive chain of "understanding fuzzy opinions → locating problems → designing modifications → verifying effects" into computable sub-tasks, which are then implemented sequentially using natural language understanding, graph reasoning, HTN planning, and simulation verification.

[0056] The overall process is as follows:

[0057] 1. Document structuring and reference graph construction (including explicit reference extraction and semantic fallback based on chapter titles);

[0058] 2. Semantic parsing of fuzzy opinions: achieved through defect type identification and key element extraction;

[0059] 3. Problem localization: Locate the set of affected nodes on the referencing graph and determine the affected area through dynamic graph propagation;

[0060] 4. Generation of modification schemes based on HTN planning: Select the top-level task according to the defect type, recursively decompose it into a sequence of operation primitives, and use the content template library to generate natural language text for the new nodes;

[0061] 5. Three-level simulation verification (structural check, citation integrity check, semantic similarity check) and ranking of candidate solutions;

[0062] 6. Modify execution and verify results (rule layer, semantic layer, reference consistency layer);

[0063] 7. Human-machine collaborative feedback and model optimization.

[0064] The technical solution of the present invention specifically includes:

[0065] Step S1: Document structuring and reference graph construction

[0066] The original document is converted into a machine-readable, reference-based graph structure, providing a foundation for subsequent problem localization and modification planning. Unlike code repair, which directly obtains the symbol dependency graph, this step requires automatically identifying and extracting reference relationships from natural language text.

[0067] Step 1.1 Document Format Parsing: Use python-docx (Word) or PyMuPDF (editable PDF) to extract the ID, type, content, and location (chapter number, page number) of each semantic unit (chapter, paragraph, table cell, text block within a figure). Output a list of nodes in JSON format. This method is applicable to editable Word documents and PDF documents containing text layers. For scanned PDFs, OCR processing is required first, but OCR is not within the scope of this invention. For model elements such as UML class diagrams, subtypes such as uml_class can be defined, whose text content includes the class name, attributes, and method descriptions.

[0068] Step 1.2 Explicit Reference Extraction: Write regular expressions to match common reference patterns, such as "see \s([Chapter, Table, Figure]\s[\d.-]+)", "as shown in Table [\d.-]+", "as shown in Figure [\d.-]+" (those skilled in the art can expand the regular expression library according to document type). If the match is successful, add a directed edge, recording the reference type and the ID of the referenced node.

[0069] Step 1.3 Semantic Backtracking Based on Chapter Titles (Reference Relationship Supplementation). When the number of reference edges extracted by regular expressions is less than a preset threshold (e.g., 5, which can be adjusted according to document size) or the number of reference edges in any chapter of the document is 0, the system uses semantic matching based on chapter titles as a supplement. Specifically: For any two chapter nodes A and B, Sentence-BERT (paraphrase-multilingual-MiniLM-L12-v2) is used to encode their titles into vectors, and cosine similarity is calculated. If the similarity is ≥0.75, a bidirectional edge (of type "semantically related") is added and marked in the log. This step is only used to enhance the reference relationship graph and is not used for locating review comments.

[0070] Step 1.4 Graph Structure Storage: Nodes = Semantic Units; Edges = References (type 'See also' or 'Semantically Related'). Use a graph data structure for storage.

[0071] Step S2: Semantic analysis of fuzzy opinions

[0072] Extract defect types and key elements from review comments. This step uses a fine-tuned BERT model as the primary method for multi-label classification to extract defect types and sequence labeling to extract key elements. It also provides a rule-based backup mechanism based on dependency parsing to handle marginal cases with low model confidence.

[0073] Step 2.1 Data Labeling Standards: Define the labeling boundaries for the following eight defect types: missing elements, naming conflicts, structural inconsistencies, logical contradictions, non-compliant formatting, incomplete descriptions, incorrect citations, and redundant content. See Explanation A for detailed definitions. The labeling standards adopt the BIO tagging system (see Explanation A) to sequentially label the following key elements in the opinion text: Missing Object (MISS), Domain Name (DOMAIN), and Reference Location (REF, such as chapter number, figure / table number). The corresponding BIO tags are: B-MISS, I-MISS, B-DOMAIN, I-DOMAIN, B-REF, I-REF, and O (representing a non-feature character).

[0074] Step 2.2 Model Fine-tuning:

[0075] (1) Training Data Collection: 5,000 historical review comments were collected and independently annotated by six domain experts, with a Kappa coefficient of no less than 0.85. These 5,000 comments came from software projects in multiple fields, covering three types of documents: requirements, design, and testing. Data Augmentation: To prevent overfitting, the following data augmentation techniques were used: synonym replacement (based on Synonyms dictionary), random insertion or deletion of stop words (no more than 2 per sentence). After augmentation, the script automatically checked whether the key elements (words corresponding to BIO tags) had been modified; if the key elements were modified, the augmented sample was discarded or the original sample was reverted. After augmentation, the effective training sample size was expanded to 15,000.

[0076] (2) Training parameters: The learning rate is The batch size is 16, the epochs are set to 3, and BERT-base-chinese is used as the pre-trained weights.

[0077] (3) Model output: For a review comment, fine-tune the BERT model output defect type (multi-label classification result) and BIO label sequence (used to extract key elements).

[0078] Step 2.3 Rule Backup Mechanism (Optional): When the BERT model's classification confidence for a certain opinion falls below a preset threshold (e.g., 0.7) or key elements are missing in the sequence labeling results, the system can fall back to the rule extraction method based on dependency parsing. This method does not replace the main BERT process; it is only a fallback strategy. Rule Extraction Process:

[0079] (1) LTP was used to perform sentence segmentation and dependency parsing on the original comments. LTP (Language Technology Platform) is an open-source Chinese natural language processing system developed by Harbin Institute of Technology, which integrates a series of basic functions such as sentence segmentation, word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing. LTP has a high authority in the field of Chinese processing. Here, LTP is used to perform basic syntactic structure parsing on the comment text to provide support for subsequent rule extraction.

[0080] (2) Identify the following trigger words (grouped by defect type, example): missing elements: missing, lacking, omission; naming conflicts: conflict, duplicate name, ambiguity; inconsistent structure: incorrect hierarchy, inconsistent structure, disordered chapters; logical contradictions: contradiction, conflict, inconsistency; non-standard format: incorrect format, inconsistent style, incorrect font; incomplete description: incomplete, missing parameters, missing steps; citation errors: citation errors, reference errors, link errors; redundant content: redundant, repetitive, superfluous.

[0081] (3) By analyzing syntactic dependency relations (such as dobj direct object, nmod modification relation, etc.), the objects (i.e. key elements) affected by the trigger words are extracted, including: missing objects, domain words, and reference positions (such as chapter numbers). The rule extraction results will be converted into BIO tag sequences and structured dictionaries consistent with the BERT output format.

[0082] Step 2.4 Output Structured Dictionary: Convert the model output (or rule fallback output) into standard JSON format. Example:

[0083] (1) Output example (missing features):

[0084] {

[0085] "defect_type": "Function missing",

[0086] "elements": {

[0087] "missing_object": "Exception handling class",

[0088] "domain": "Network timeout, database connection failed",

[0089] "reference": "Section 4.3"

[0090] }

[0091] }

[0092] (2) Output example (logical contradiction, does not include node_id):

[0093] {

[0094] "defect_type": "Logical contradiction",

[0095] "elements": {

[0096] "contradiction_pairs": [

[0097] ["Chapter 2 'Interface Returns XML'", "Chapter 5 'Interface Returns JSON'"] ]

[0099] }

[0100] }

[0101] Step S3, Problem Identification

[0102] Based on the key elements output in step S2, find the set of seed nodes most likely to be affected in the document reference relationship graph, and determine the affected area through variable step graph propagation.

[0103] Step 3.1 Precise Matching Based on Position Identifiers (Handling Semi-Fuzzy Comments): If a review comment is semi-fuzzy (i.e., contains explicit chapter or table numbers, such as "Section 5.2" or "Table 3-1"), the position identifier is directly parsed into the corresponding document node ID, used as a seed node, and assigned a confidence level of 1.0. Subsequent keyword matching and semantic expansion steps are skipped. Parsing rules: Regular expressions are used to match chapter number patterns (e.g., Section (\d+(?:\.\d+)) or Table (\d+(?:\.\d+))), searching for nodes with matching position attributes in the document reference graph. If the parsed position identifier does not exist in the document (e.g., the chapter number is out of range), a fallback to manual user specification is triggered.

[0104] Step 3.2 Keyword Matching (for non-semi-fuzzy opinions): For non-semi-fuzzy opinions (those without explicit location identifiers), extract key elements from the output of Step 2 (such as missing_object, domain, conflict_name, and text phrases in contradiction_pairs), and perform fuzzy matching on the titles and texts of all nodes. The matching method is Jaro-Winkler distance, with a threshold of 0.85. Return a list of all successfully matched node IDs as candidate seed nodes. If no match is found, proceed to Step 3.3.

[0105] Step 3.3 Semantic Expansion (Supplement when Keyword Matching Fails): When keyword matching yields no results, Sentence-BERT is used to calculate the cosine similarity between key elements and the summary (title + first sentence) of each node. The specific model is paraphrase-multilingual-MiniLM-L12-v2, selecting the top-3 nodes with the highest similarity as candidate seed nodes. If the highest similarity is <0.4, it is considered a match failure, triggering a fallback to user-specified nodes. If there are multiple candidate nodes, all are used as seed nodes (because they may be semantically related). It is important to emphasize that this step differs from the semantic matching in Step 1.3: Step 1.3 constructs a citation relationship graph based on the similarity between chapter titles (offline, global); this step locates seed nodes based on the similarity between key elements in review comments and the node summary (online, specific to each comment). The two steps operate independently and do not interfere with each other.

[0106] Step 3.4 Variable-Step Graph Propagation: Starting from the seed node set, propagate outwards in the document reference graph to determine the affected area. The number of propagation steps is dynamically determined: For each seed node, calculate its out-degree + in-degree (i.e., reference degree). If the degree > threshold T, propagate 2 steps; otherwise, propagate 1 step. The maximum number of steps is capped at 2 to avoid space explosion. The threshold T is dynamically calculated: T = max(3, min(10, total number of nodes / 20)), where the total number of nodes is the total number of nodes in the document reference graph.

[0107] Special restrictions: For seed nodes from semi-fuzzy opinions, the propagation step limit is set to 1 step (propagating only to directly referenced nodes) to avoid overexpansion. If the number of nodes in the affected area after propagation exceeds 50, it will be truncated (only the first 50 nodes visited from the seed node in BFS order will be retained) and the user will be prompted.

[0108] Step 3.5 Output Location Results: Output a list of node IDs (affected area) and the confidence score for each node. The confidence score calculation rule is as follows:

[0109] (1) Direct localization of semi-fuzzy opinions: base confidence level 1.0 (seed node);

[0110] (2) Keyword match successful: Basic confidence level 1.0 (seed node);

[0111] (3) Semantic expansion match successful: base confidence 0.8 (seed node).

[0112] The number of nodes obtained from graph propagation is calculated as: base confidence × 0.6 (multiplied by 0.6 for each propagation step, up to a maximum of two steps).

[0113] Special case: If the location result is empty (no seed nodes), the entire document graph is considered the affected area, and the user is prompted for manual confirmation. If the affected area is too large (more than 50 nodes), it has been truncated in Step 3.4, and the confidence calculation only applies to the retained nodes.

[0114] Output example:

[0115] {

[0116] "seed_nodes": [

[0117] {"node_id": "sec_4.3", "confidence": 1.0, "source": "semi-fuzzy opinion"},

[0118] {"node_id": "para_4.3.1", "confidence": 0.6, "source": "Graph propagation (1 step)"}

[0119] ],

[0120] "affected_region": ["sec_4.3", "para_4.3.1", "fig_4-2"],

[0121] "truncated": false

[0122] }

[0123] Step S4: Generate modification scheme based on HTN planning (including content template generation)

[0124] Based on the defect type and key elements output in step S2, and the seed node set and affected area output in step S3, candidate modification schemes composed of operation primitives are automatically generated using hierarchical task network (HTN) planning.

[0125] Technical Approach: A Hierarchical Task Network (HTN) is employed for planning. A predefined task library (method set) is used, and an HTN planner (such as PyHop or a self-developed lightweight recursive decomposer) recursively decomposes the top-level tasks into atomic tasks. Atomic tasks include: document modification primitives (mapped to the document editing API) and internal graph operations (modifying only the reference relationship graph). For operations requiring the addition or updating of text content, natural language text is generated by filling in pre-defined content templates, ensuring the determinism and traceability of the generated content.

[0126] Planning Layer Description: This invention introduces a planning layer abstraction during the HTN planning process. The planning layer maintains a lightweight view of the document, containing only the structure, referencing relationships, and text summaries, without storing the complete long text, to ensure planning efficiency. A mapping relationship exists between the planning layer state and the real document state: nodes in the planning layer correspond to semantic units in the real document, and operation primitives in the planning layer are ultimately mapped to document editing API calls. During template population, the system reads the required context from the real document.

[0127] Step 4.1 Planning Layer State Representation: The planning layer maintains a lightweight document state for precondition judgment and effect simulation. The state consists of the following information:

[0128] (1) Node set: Each node contains: id (unique identifier, consistent with the node ID generated in step 1), type (semantic unit type, such as section / paragraph / table / figure / uml_class, etc.), name (node ​​name, chapter title, figure title, table title, or paragraph first sentence summary), text_summary (text summary, title + first sentence, or the first 200 characters, used for semantic matching and name checking), keywords (top-5 keywords extracted from text_summary, used for fast matching), parent_id (parent node ID, chapter hierarchy relationship), position (position in the document, chapter number, page number, etc.), format_props (format attributes, optional, used for UPDATE_FORMAT).

[0129] (2) Reference edge set: Directed edge (from, to, ref_type), representing the reference relationship from node from to node to (e.g., "see also").

[0130] Note: The planning layer does not store the complete long text content of nodes to keep the state lightweight. When the template is populated, the system temporarily reads the required context from the original document.

[0131] Step 4.2 Definition of Operation Primitives:

[0132] (1) Document modification primitives (mapped to document editing API)

[0133] Original language parameter Prerequisites Effect ADD_NODE parent_id, position, node_type, content_text If the parent node exists, the position is valid (e.g., "end" or a specified index). Add a new node under the parent node and set its full text content. DELETE_NODE node_id Node exists Delete the node and all its child nodes (reference integrity is checked during the verification phase). UPDATE_TEXT node_id, new_text Node exists Replace the complete text content of the node MOVE_NODE node_id, new_parent, new_position The node exists, the new parent node exists, and `new_position` is valid. Move the node to the specified position under the new parent node. UPDATE_FORMAT node_id, format_props Node exists Update the node's formatting attributes (such as font and style). SPLIT_NODE node_id, split_point The node type is paragraph, and split_point is the character offset within the paragraph or sentence boundary. Split a paragraph node into two consecutive paragraph nodes.

[0134] (2) Internal diagram operations (only modify the internal reference relationship diagram of the system, without changing the document text)

[0135] operate parameter Effect ADD_EDGE from_id, to_id, ref_type Add a directed edge to the reference graph. REMOVE_EDGE from_id, to_id Delete specified directed edge

[0136] These internal operations are used to maintain referential integrity (e.g., adding reference edges to other nodes after a new node is added), and do not directly generate document text. When the final modification plan is executed, the system will synchronously modify the document text as needed (e.g., inserting reference statements at appropriate positions via UPDATE_TEXT).

[0137] Step 4.3 HTN Task Library Definition: Define the top-level task and its decomposition method for each defect type. The design philosophy of the task library is compatible with classic HTN planning systems (such as SHOP2), but this invention does not limit the specific planner implementation. The decomposition logic is illustrated below using "missing elements" and "naming conflicts" as examples; other defect types can be deduced by analogy.

[0138] 1. Missing elements

[0139] (1) Top-level task: add-missing-element(type, name, parent_region), where type is the missing element type (e.g., uml_class, paragraph), name is the name or identifier of the element (e.g., “anomaly handling class”), and parent_region is the suggested parent node region (from the affected region in step 3).

[0140] (2) Decomposition method (default): Prerequisites include the existence of a parent node (parent) that is a chapter type (or a container that allows nodes of this type to be mounted), and the absence of a node with the same name in the current document (has-name(node, name)=false). Subtask sequence (executed in order):

[0141] Generate text: Call the template filling engine to generate natural language text for the new node based on type, defect type (missing feature), and key features. text=generate_template_text(type, "missing feature", "add",{name, domain, missing_object}).

[0142] Add a node: Creates a new node under the parent node (mapped to the ADD_NODE primitive). ADD_NODE(parent, position="end", node_type=type, content_text=text).

[0143] Maintaining Reference Relationships (Optional): Within the affected region, find existing nodes that semantically reference the new node. Calculate the Jaccard similarity between the existing node's text keywords and the new node's name / type keywords. If the similarity is ≥0.3, the node is considered a candidate. This threshold is independent of the 0.5 threshold used for text replacement discrimination in step 4.6, and is optimized separately for adding reference relationships and local replacement scenarios. For each candidate node, add an edge pointing to the new node in the internal reference relationship graph (ADD_EDGE(candidate, new_node, "See also")). Optionally, insert a reference statement (e.g., "See [new node name]") into the candidate node text using UPDATE_TEXT. Note: If there are no candidate nodes within the affected region, skip the reference relationship maintenance step.

[0144] 2. Naming conflict

[0145] (1) Top-level task: resolve-naming-conflict(conflict_name, node1, node2, parent_region), where conflict_name is the name of the conflict, node1 and node2 are the two nodes that are in conflict (may be located in different positions), and parent_region is the affected region (from step 3).

[0146] (2) The system provides two decomposition methods, which can be selected according to the preconditions.

[0147] Method 1: Rename one of the nodes (keeping both).

[0148] Prerequisites: Both nodes exist and both have the name conflict_name; the two nodes are not equivalent (or should not be merged).

[0149] Subtask sequence:

[0150] Generate a unique identifier for the new name: `generate_unique_name(conflict_name, node1)`. Specifically, if a numeric suffix can be added to the name of `node1`, try names like "original_name_2", "original_name_3", etc., until it does not conflict with any existing node name; if `node1` is a class name, a prefix can be added based on its section (e.g., "Sec4_" + original name). Those skilled in the art can implement this function according to the specific document type.

[0151] Update the text content of node1: replace all occurrences of conflict_name with new_name (mapped to the UPDATE_TEXT primitive, using a local replacement mechanism). UPDATE_TEXT(node1, replace_old_name_with_new(original_text, conflict_name, new_name)).

[0152] Update all reference edges pointing to node1: Traverse the edges pointing to node1 in the reference graph, optionally updating the old name in the text of the reference source node to the new name (via UPDATE_TEXT).

[0153] Method 2: Merge two nodes (keep one, delete the other)

[0154] Prerequisite: The two nodes are semantically substitutable (determined by the is-substitutable(node1, node2) function). This function is implemented by using Sentence-BERT to calculate the cosine similarity between node1.text_summary and node2.text_summary. If the similarity is ≥0.85, then they are determined to be substitutable.

[0155] Subtask sequence: Redirect all reference edges pointing to node2 to node1: For each (from, node2) edge, execute REMOVE_EDGE(from, node2) and ADD_EDGE(from, node1, ref_type). Optionally, update the reference target name in the text of the from node (from node2 to node1).

[0156] Delete node2 (mapped to the DELETE_NODE primitive). DELETE_NODE(node2).

[0157] Brief description of task definitions for other defect types:

[0158] Defect types Top-level task Main subtasks (primitives) Remark Structural inconsistency fix_structure(node) MOVE_NODE, UPDATE_FORMAT Adjust chapter hierarchy or table structure Logical contradiction resolve_contradiction(pair) UPDATE_TEXT Modify the description of one of the contradictory points. Reference error fix_reference(edge) UPDATE_TEXT Correct the target chapter / figure number in the cited text. Redundant content remove_redundancy(node1, node2) DELETE_NODE, merge operation Delete duplicate nodes or merge content The format does not conform to the specifications. fix_format(node) UPDATE_FORMAT Adjust font, style, indentation, etc. Incomplete description complete_description(node) UPDATE_TEXT Add missing parameters or steps using template filling.

[0159] For the above types, the system selects an appropriate decomposition method based on the specific defects identified in step 2 and the nodes located in step 3, and finally generates an atomic operation sequence.

[0160] Extensibility of the task library: This invention does not limit the complete enumeration of the task library. Those skilled in the art can add new defect types, top-level tasks, and decomposition methods according to actual document types and defect patterns. Updates to the task library can be performed offline through human-computer collaborative feedback in step 7.

[0161] Step 4.4 Execution: Using the HTN planner (this example uses the lightweight PyHop; in actual deployment, a high-performance planner such as Fast Downward can be used), set a time limit of 3 seconds. The planner outputs a sequence of atomic tasks, each of which is a primitive or internal operation from Step 4.2. The input is:

[0162] (1) Initial state: The document state defined in step 4.1;

[0163] (2) Top-level task: generated based on the defect type in step 2 and the area of ​​impact in step 3 (e.g., (add-missing-element "uml_class" "ExceptionManager" sec_4_2)).

[0164] (3) Planning parameters: Time limit of 3 seconds (if the time limit is exceeded, the longest legal prefix currently generated will be returned and the scheme will be marked as "partially generated"), maximum atomic task steps of 10 steps (if the number of steps exceeds the limit, it will be truncated to the first 10 steps and a warning log will be recorded).

[0165] Step 4.5 Candidate solution generation: Generate multiple candidate solutions (up to 5) through parent node relaxation and multi-method decomposition.

[0166] 1. Parent Node Relaxation (only applicable to scenarios requiring new nodes, such as missing elements): For the top-level task `add-missing-element`, the default parent node (or the chapter specified in the comment) of the seed node in step 3 is used. Additionally, the system attempts to place the new node under other candidate parent nodes, generating different solutions:

[0167] (1) Default parent node: the chapter where the seed node is located (or the chapter specified in the opinion).

[0168] (2) Distance priority: Chapter nodes that are 1 distance away from the seed node graph (i.e., chapters that directly reference the seed node or are directly referenced by the seed node).

[0169] (3) Semantic priority: Other sub-chapter nodes that share the same grandparent node (same chapter or section) as the seed node.

[0170] (4) Hierarchical priority: Chapters in the document outline that are at the same level or one level higher than the seed node (e.g., promoted from section to chapter).

[0171] The four strategies described above are all heuristic rules, covering the most likely insertion positions for new nodes in the document structure. The system will try all applicable strategies simultaneously, generating corresponding candidate solutions. The optimal solution will ultimately be determined by simulation verification and user selection.

[0172] Each choice corresponds to a candidate solution. If a choice results in the preconditions not being met (such as a mismatch in parent node types), it is skipped.

[0173] 2. Multi-method decomposition: For the same top-level task, multiple methods may be defined in the task library (such as "rename" and "merge" for naming conflicts). Each method generates an independent candidate solution.

[0174] Ultimately, the system collects all valid candidate solutions, retaining a maximum of 5 (by generation order or heuristic scoring).

[0175] Step 4.6 Content Template Population: For the ADD_NODE and UPDATE_TEXT primitives, specific natural language text needs to be generated. This invention uses a predefined template library plus keyword population to ensure the determinism and auditability of the generated content.

[0176] 1. Template Library Structure: The template library uses (node_type, defect_type, operation_type) as keys, where operation_type takes the value of add or update. Template strings contain placeholders, such as {class_name}, {domain}, {old_name}, {new_name}, etc. Example (JSON format):

[0177] {

[0178] "uml_class_missing_feature_add": "The class {class_name} is responsible for handling functions related to {domain}."

[0179] "paragraph_missing_add": "Added explanation regarding {domain}. Specifically includes: {missing_object}."

[0180] "uml_class_naming conflict_update": "The {new_name} class is used to handle the {domain} functionality, replacing the original {old_name} class to eliminate ambiguity."

[0181] "paragraph_naming conflict_update": "The original term "{old_name}" had different meanings in different chapters, so it has been uniformly changed to "{new_name}".

[0182] }

[0183] 2. Placeholder sources: (1) Key elements output in step 2 (such as missing_object, domain, conflict_name, etc.); (2) Seed node attributes output in step 3 (such as node name, ID); (3) Document structure information (such as parent node title, chapter number); (4) System generated (such as new_name generated by the generate_unique_name function).

[0184] 3. Partial Replacement Mechanism (for UPDATE_TEXT): The system defaults to full replacement based on regular expressions (replacing all matching strings in the node text with the new string). If you need to retain the original text of the non-replaced parts, you can enable sentence-level replacement through configuration (this relies on accurate sentence boundary recognition and may be suitable for highly structured documents). Full replacement is used by default to ensure reliability.

[0185] (1) Extracting old names: Extract from the original node text by named entity recognition or keyword matching; if extraction is not possible, the entire node text will be replaced by default.

[0186] (2) Template usage: The template is only used to generate new content for the replaced sentence. After replacement, the other sentences in the original paragraph remain unchanged.

[0187] 4. Template Population Timing: Template population occurs during the planning and execution phase. When `call get-template-text` is invoked, the system generates text in real time based on the current context (defect type, node type, key element). The generated text is then passed as a parameter to the `ADD_NODE` or `UPDATE_TEXT` primitive. See Explanation B for the complete template library (covering at least four common defect types).

[0188] Step 4.7 Output Candidate Solutions: The output of Step 4 is a list of candidate solutions. Each solution includes a solution ID, a top-level task description (e.g., "Add an ExceptionManager class under Section 4.3"), an atomic task sequence (including operation primitives and parameters, with the text content already filled in), and a generation method flag (e.g., "Parent node relaxation - default" or "Method - rename"). This list will be passed to Step 5 for three-level simulation verification and sorting.

[0189] Step S5: Three-level simulation verification and ranking of candidate solutions

[0190] The system quickly evaluates the structural feasibility, reference integrity, and semantic correctness of each candidate modification, and outputs a confidence ranking. Unlike compilation verification in code fixing, a three-level verification system is designed specifically for the characteristics of the documentation.

[0191] Step 5.1 First Level: Structure Check (Precondition Verification): Without simulation execution, first check whether each operation primitive in the operation sequence satisfies its precondition. If any operation violates the precondition, the scheme is discarded directly and will not proceed to subsequent verification.

[0192] 1. Inspection rules (based on the definition in Step 4.2):

[0193] ADD_NODE: The parent node exists, and the position is valid;

[0194] DELETE_NODE: The node exists;

[0195] UPDATE_TEXT: The node exists;

[0196] MOVE_NODE: The node exists, and the new parent node exists;

[0197] UPDATE_FORMAT: The node exists;

[0198] SPLIT_NODE: The node exists and its type is paragraph;

[0199] 2. Pass condition: All operations meet the prerequisite conditions.

[0200] Step 5.2 Simulate Execution (for schemes that pass Level 1): Create a lightweight copy of the current document graph (including the affected region and all external reference nodes pointing to that region). Apply primitives sequentially according to the operation sequence, updating node attributes and reference edges in the copy.

[0201] Copy strategy: All nodes within the affected region and all source nodes of reference edges pointing to the affected region from outside must be included. If the number of external reference nodes pointing to the affected region exceeds a preset threshold (e.g., 100), all nodes will still be copied (without truncation), but the verification report will mark it as "Many external reference nodes, verification time may be long." For UPDATE_TEXT operations, if they affect the external reference nodes themselves, the complete text attributes of these nodes must also be copied.

[0202] During the simulation: Each primitive is executed sequentially, and the replica state is updated. If a primitive fails to execute (e.g., due to an API exception), the scheme is marked as "execution exception" and will not proceed to subsequent verification.

[0203] Step 5.3 Level 2: Referential Integrity Check: On the document graph copy after the simulation execution, check the validity of all reference edges.

[0204] The algorithm checks the integrity of references by traversing all reference edges (from, to) in the original graph (before simulation). If the "to" node is deleted after simulation, and none of the candidate solutions contain any of the following operations, then the algorithm fails the integrity check.

[0205] (1) The operation to delete the edge (REMOVE_EDGE(from, to));

[0206] (2) An operation that updates the text of the from node to remove or modify the reference (UPDATE_TEXT and the new text no longer contains the reference);

[0207] (3) Redirection operation (first REMOVE_EDGE then ADD_EDGE to the new node).

[0208] For newly added nodes, check if they have at least one incoming edge (i.e., referenced by other nodes). If not, mark them as "orphaned nodes" and deduct their confidence score, but do not directly reject them (because some newly added content may not need to be referenced). Users can configure whether to enforce incoming edges (via the `require_incoming_edges` parameter, which defaults to false).

[0209] Pass condition: There are no unprocessed dangling references (i.e., no target deletion with a corresponding update operation). Orphaned nodes only affect the confidence level, not the pass / fail status.

[0210] Step 5.4 Level 3: Semantic Similarity Check (Modification Adequacy Assessment): For UPDATE_TEXT and ADD_NODE operations, assess whether the modified text maintains semantic coherence and adequately addresses the problem.

[0211] (1) Inspection rules

[0212] Operation type Inspection content Similarity calculation methods By threshold Consequences of failure UPDATE_TEXT Semantic similarity between the text before and after modification Sentence-BERT cosine similarity ≥0.7 The scheme is marked as "semantic risk" and will not pass Level 3. ADD_NODE Semantic similarity between the newly added node text and its parent node text Sentence-BERT cosine similarity ≥0.5 The scheme is marked as "semantically incoherent" and does not pass Level 3.

[0213] Note: The threshold can be adjusted experimentally (user-configurable), and the default value is obtained based on the internal validation set. If the scheme contains multiple UPDATE_TEXT or ADD_NODE operations, the lowest similarity score is taken as the semantic similarity score of the scheme. If the scheme does not contain these two types of operations, it will automatically pass the third level.

[0214] (2) Passing condition: If the similarity of all related operations is not lower than the corresponding threshold, then the third level passes; otherwise, it fails.

[0215] Step 5.5 Confidence Calculation: Calculate a confidence score (0~1) for each candidate solution for ranking. Solutions that pass the first level participate in subsequent validation and confidence calculation. Solutions that pass the second level have their confidence scores calculated normally and participate in the ranking; solutions that do not pass the second level also have their confidence scores calculated, but are placed after the passing group in the ranking and are only for manual reference.

[0216] 1. Base score: Assuming pass_structure=1 (first level passed, otherwise the solution has been discarded), pass_integrity=1 (second level passed) or 0 (failed), pass_semantic=1 (third level passed) or 0 (failed), then the base score = (pass_structure+pass_integrity+pass_semantic) / 3 (range 0~1).

[0217] 2. Step Penalty Factor: Let step_count be the number of atomic operation steps in the scheme (maximum 10, which is truncated in step 4.4). Then the step penalty factor = 1 - 0.05 × (step_count / 10). The fewer the steps, the closer the factor is to 1, and it is 0.95 when the number of steps is 10.

[0218] 3. Isolated node penalty: If there is at least one newly added isolated node (without incoming edges) in the solution, the penalty is multiplied by 0.9 (regardless of the number of isolated nodes, the penalty is only applied once, and this coefficient is configurable).

[0219] 4. Final confidence formula: Confidence = Base score × Step penalty factor × Isolated node penalty factor.

[0220] Example: Level 3 passed, 5 steps, no isolated nodes: 1×(1-0.05×0.5)=0.975; Level 2 failed, 2 steps, no isolated nodes: (1+0+1) / 3×(1-0.05×0.2)=0.6667×0.99≈0.66.

[0221] The confidence level is rounded to two decimal places (not rounded, but truncated).

[0222] Step 5.6 Sorting and Output: Arrange the candidate solutions according to the rules and output them to the front end for display. The sorting rules are as follows:

[0223] (1) Grouping by whether the second level is passed: All schemes that pass the second level (reference integrity) are listed first, and those that do not pass are listed later.

[0224] (2) Within the group, sort by confidence level in descending order (from high to low).

[0225] (3) If the confidence levels are the same, they are arranged in ascending order of the number of steps (those with fewer steps are given priority).

[0226] Output format: Each scheme includes a scheme ID, a list of atomic tasks, confidence level, and pass status flags for each level.

[0227] Step 5.7 User selects candidate solutions: The front end displays a sorted list of candidate solutions, and the user can select any solution (the solution with the highest confidence level is highlighted by default). After the user makes a selection, proceed to Step 6. The system can also be configured to automatically select the solution with the highest confidence level.

[0228] Step S6: Modify execution and verify results

[0229] Apply the selected modification scheme (which can be selected by the user or the system will automatically select the highest-scoring scheme) to the real document and perform final verification.

[0230] Step 6.1 Execute the modifications: Using the python-docx (Word) or PyMuPDF (editable PDF) API, execute each operation primitive sequentially according to the operation sequence. Log the operation type, parameters, and execution result after each step. Establish a transaction rollback mechanism:

[0231] 1. Use a document copy mechanism: Perform all operations sequentially on a copy of the original document.

[0232] 2. If all operations are executed successfully, replace the original document with a copy (while backing up the original document).

[0233] 3. If any step fails, the copy is discarded, the system is rolled back to the original document state, and an exception is thrown for the user to handle.

[0234] Step 6.2 Perform integrity verification: Check whether the changes have successfully taken effect in the actual document. Verification content:

[0235] 1. For ADD_NODE: Check if the target node exists in the document (by matching position and content).

[0236] 2. For DELETE_NODE: Check if the node no longer exists.

[0237] 3. For UPDATE_TEXT: Check if the node text has been updated to the new content.

[0238] 4. For MOVE_NODE: Check if the node has already appeared under the new parent node.

[0239] 5. For UPDATE_FORMAT: Check if the format attributes have been changed.

[0240] 6. For SPLIT_NODE: Check if the original node has been split and if two new nodes exist.

[0241] Pass condition: All operations are verified successfully. If any operation fails, the operation is deemed unqualified, and the reason for failure is output (e.g., "New node not found").

[0242] Step 6.3 Defect Fix Adequacy Verification: Extract the complete text of the relevant areas (affected areas + newly added nodes) in the modified document, and use an independent, rule-based validator (not the classification model in Step 2) to check whether the modifications have resolved the issues pointed out in the original review comments. The complete rule set can be found in Explanation C.

[0243] Optional Reference Signal (Step 2 Model): The model from Step 2 (defect classification + feature extraction) can be called again, and its output can be used as a reference signal, but not as a basis for judgment. The reference signal is only used to assist users in understanding, and the specific logic is as follows:

[0244] 1. If the rule verification passes → it is deemed qualified (the reference signal is only used for log recording).

[0245] 2. If the rule verification fails, but the defect type output by the model in step 2 is consistent with the original opinion and has a high confidence level (≥0.8), mark it as "requires manual review" (it may be due to incomplete verification rules).

[0246] 3. If the rule validation fails and the model output in step 2 also does not match, it is judged as unqualified.

[0247] Pass condition: Rule verification passes (reference signal not considered); if rule verification fails but is marked as "requires manual review", the system pauses and waits for user judgment.

[0248] Step 6.4 Reference Consistency Verification: Check whether all references to the modified nodes in the revised document are still valid. Verification content:

[0249] 1. Traverse the document reference graph (based on the modified documents, re-extract the reference relationships according to the method in step 1 to construct a temporary graph, and then verify it. If performance is critical, the original graph can be reused and incrementally updated, but it is recommended to re-extract to ensure accuracy).

[0250] 2. For each reference edge (from, to): If the to node has been deleted and the reference has not been cleaned up using UPDATE_TEXT or REMOVE_EDGE, an error is reported. If the to node has been moved, check if the position identifier in the reference text needs to be updated (optional; complex cases should be manually reviewed).

[0251] 3. For newly added nodes, if the user configures them to "must be referenced", check if there is at least one reference pointing to them.

[0252] Pass condition: No invalid references exist (i.e., the reference target exists and is valid). If an invalid reference exists, the application is deemed unqualified, and the specific reference location is output.

[0253] Step 6.5 Comprehensive Judgment and Report Output:

[0254] 1. Qualification criteria: Steps 6.2, 6.3 (rule verification passed), and 6.4 all pass. If Step 6.3 is marked as "requires manual review," the system will pause and wait for user confirmation before continuing.

[0255] 2. Output content: Judgment result ("qualified" or "unqualified"), modification basis report (Markdown or HTML format), including: original review comments text, selected candidate solutions (solution ID, atomic task list), comparison of document differences before and after modification (text changes in affected areas), verification results at all levels (execution integrity, defect repair adequacy, reference consistency) and specific reasons for failure (if any), and optional reference signal output for the model in step 2.

[0256] Example report summary:

[0257] 1. Judgment result: Qualified;

[0258] 2. Basis for modification: Omitted;

[0259] 3. Comment: The class diagram lacks an exception handling mechanism;

[0260] 4. Solution: Add an ExceptionManager class under section 4.3;

[0261] 5. Execution integrity: Passed (new node has been added);

[0262] 6. Defect Repair Verification: Passed (There are nodes of type "Class" with text containing "Exception Handling");

[0263] 7. Reference consistency: Passed (3 nodes referenced the newly added class).

[0264] Step S7, Human-Machine Collaborative Feedback

[0265] Record user feedback behavior to continuously optimize the model and task library.

[0266] Step 7.1 Log Collection: Record the following information and store it in JSON format: original review comments text, defect type identified in Step 2, candidate solution list generated by the system (including solution ID, confidence level, atomic task sequence), solution ID selected by the user, user modifications to the solution (if the parent node, text content, or operation order is modified, record the modification details; otherwise, leave it empty), final verification result (qualified / unqualified / manually reviewed), and timestamp.

[0267] Step 7.2 Offline Retraining: Retraining Trigger Conditions (any one of the following conditions must be met):

[0268] 1. Every 500 newly adopted and validated data points are accumulated (this number can be configured based on the actual data accumulation rate); "fully adopted" is defined as a user selecting a solution from the system-generated candidate list (rather than manually creating a solution), and the final validation result is "qualified". Regardless of whether the user adjusts the solution selection (e.g., from solution B instead of solution A), as long as the final adopted solution originates from system generation, it is included in the retraining data. Partially adopted data (qualified after user modification) is stored separately for manual analysis and is not automatically added to the training set.

[0269] 2. Triggered monthly (if the accumulated data is less than 500 records, it will be postponed to the next month).

[0270] Retraining process:

[0271] 1. Merge the newly accumulated fully adopted data with the historical training set, and fine-tune the BERT-base-chinese model (parameters are the same as in step 2.2).

[0272] 2. Calculate the failure rate of HTN decomposition for various defects. Failure criteria: planning timeout (no complete solution is generated within 3 seconds in step 4.4) or the user selects "no available solution" among the generated candidate solutions / the confidence level of all solutions is lower than 0.5.

[0273] 3. If the failure rate of a certain type of defect exceeds 30% for two consecutive months, a manual optimization prompt will be triggered, and domain experts will supplement new task decomposition methods (added to the task library in step 4.3).

[0274] Step 7.3 Model Version Management: After each retraining, a new version of the model is generated (version number increments), and the 5 most recent historical versions are retained. The new version of the model is evaluated using a fixed validation set (consistent with Step 2.2), and the performance metric is the macro-average F1 score for defect type classification. If the macro-average F1 score of the new version of the model drops by more than 5 percentage points in absolute terms compared to the current production version (e.g., from 0.85 to <0.80), it is automatically rolled back to the previous version, an alarm log is recorded, and automatic retraining is paused, awaiting manual analysis.

[0275] This invention also provides a system for automatically generating modification schemes for software development documents in response to ambiguous review comments. The system adopts a microservice architecture, and the module connection relationships are as follows:

[0276] 1. Document parsing and graph construction module: The output is connected to the location inference module.

[0277] 2. Semantic understanding module: The output is connected to the localization reasoning module and the HTN planning module.

[0278] 3. Location Inference Module: The output is connected to the HTN planning module.

[0279] 4. HTN Planning Module (including content template library): The output is connected to the three-level simulation verification and sorting module.

[0280] 5. Three-level simulation verification and sorting module: The output is connected to the front-end interactive interface.

[0281] 6. Front-end interactive interface: The output end is connected to the document execution module (to receive the scheme selected by the user).

[0282] 7. Document Execution Module: The output is connected to the complete verification module.

[0283] 8. Complete Validation Module: The validated samples (original opinions, comparison of documents before and after modification, defect types) are used as positive samples and sent to the offline retraining component of the semantic understanding module through the feedback interface.

[0284] This forms a closed-loop optimization system. The data flow between the modules forms a closed loop, supporting continuous optimization.

[0285] Explanation A: Defect Type Labeling Specification

[0286] Defect types definition Positive example Counterexamples (not belonging to this category) missing elements Content that should have been included in the document but was not "Missing error code definition" "Error code definition is inaccurate" (This is due to an incomplete description). Naming conflicts The same name points to different entities The variable name "timeout" has different meanings in two different places. "Spelling error" Structural inconsistency The chapter hierarchy and table structure do not meet expectations. "Section 3.1.1 appears directly in Chapter 3, skipping a level." Formatting issues Logical contradiction Conflict between descriptions Chapter 2 discusses interfaces returning XML, and Chapter 5 discusses returning JSON. Chapter 2 states that the interface returns XML, while Chapter 5 states that the interface returns XML, but in a different format. The format does not conform to the specifications. Font, indentation, and numbering format errors The heading is not using a standard style. - Incomplete description Missing necessary parameters and steps "Use case description lacks postconditions" - Reference error Incorrect chapter number and figure number reference See Chapter 5, which should actually be Chapter 6. - Redundant content Repeated description Section 3 and Section 5 contain overlapping content. -

[0287] Example of a key element BIO tag system:

[0288] B-MISS: Missing object start

[0289] I-MISS: Missing object internal

[0290] B-DOMAIN: Domain terminology begins

[0291] I-DOMAIN: Domain terminology

[0292] B-REF: Starts with the chapter number.

[0293] I-REF: Reference to the chapter number

[0294] O: Other irrelevant characters (optional)

[0295] Explanation B: Content Template Library

[0296] Placeholder explanation: Placeholders in the template, such as {class_name}, {domain}, {old_level}, {new_level}, {table_id}, {modification_description}, {incorrect_ref}, {correct_ref}, etc., are extracted not only from the key elements in step 2, but also include system predefined variables obtained from the document structure parsing results.

[0297] The template library is stored in JSON files (keys include the suffix _add or _update, and node types support extended types such as uml_class):

[0298] {

[0299] "uml_class_missing_feature_add": "The class {class_name} is responsible for handling functions related to {domain}."

[0300] "paragraph_missing_add": "Added explanation regarding {domain}. Specifically includes: {missing_object}."

[0301] "uml_class_naming conflict_update": "The {new_name} class is used to handle the {domain} functionality, replacing the original {old_name} class to eliminate ambiguity."

[0302] "paragraph_naming conflict_update": "The original term "{old_name}" had different meanings in different chapters, so it has been uniformly changed to "{new_name}".

[0303] "section_structure_inconsistency_update": "Adjusted section level: changed the {old_level} level heading "{title}" to the {new_level} level."

[0304] "table_structure_inconsistency_update": "Modifies the structure of table {table_id}: {modification_description}".

[0305] "paragraph_reference_update": "Corrected the reference "{incorrect_ref}" in the original text to "{correct_ref}".

[0306] "section_reference_update": "The reference "{incorrect_ref}" in section {section_num} should have been "{correct_ref}", and has been corrected.

[0307] }

[0308] Explanation C: Defect Repair Verification Rule Set (used in step 6.3)

[0309] Defect types Validation rules (pseudocode) missing elements exists(node ​​in modified_region such that node.type ==missing_object and (node.text contains domain ornode.keywords contains any keyword of domain)) Naming conflicts (not exists(node1, node2 in modified_region such thatnode1.name == node2.name and node1 != node2)) AND (foreach conflicting node originally, it is either deleted or its name changed) Structural inconsistency (for all node in modified_region: node.level_depth <=max_allowed_depth) AND (if a node was moved, then itsnew parent must satisfy hierarchical constraints) Logical contradiction (for each contradiction pair (text1, text2), not(positive_keywords in text1 and negative_keywords intext2)) AND (at least one node in the pair has beenmodified by UPDATE_TEXT) The format does not conform to the specifications. For all node in modified_region: node.format instandard_format_set Incomplete description (for all node in modified_region where node.type ==requirement: node.text contains at least one of {"shall", "must", "should", "will"}) OR (node ​​was updatedand new text length > old text length) Reference error For all ref_edge in modified_region: target_node existsand ref_edge.target_id == expected_target_id Redundant content (for all (node1, node2) in modified_region: similarity(node1.text, node2.text) < 0.85) AND (redundant nodesare either deleted or merged)

[0310] Note: This validator is independent of the classification model in step 2. It is based on rule and keyword matching, ensuring the determinism and interpretability of the judgment results. Furthermore, for "logical contradictions," only the contradictions are eliminated; the correctness of the business logic is not guaranteed.

[0311] Explanation D: Test dataset description and performance evaluation method

[0312] 1. Test environment:

[0313] Hardware: Intel Core i7-12700 CPU, 32GB RAM, 512GB SSD

[0314] Software: Windows 11, Python 3.9, python-docx 0.8.11, transformers 4.20.0

[0315] Test document: A "Software Requirements Specification for xx System", totaling 287 pages, which, after processing in step 1, yields 1560 semantic unit nodes and 389 reference edges.

[0316] 2. Evaluation Indicators:

[0317] Solution generation success rate: The proportion of solutions that generate at least one with a confidence level ≥ 0.8.

[0318] Manual modification rate: The percentage of users who need to modify the design (including adjusting the position and modifying the content).

[0319] Processing time: The time from when feedback is input to when the user confirms the solution (excluding user reading time).

[0320] First-pass rate: The percentage of cases that pass the modified full validator. First-pass rate = (Number of cases that pass the modified full validator) / (Number of cases that generate at least one candidate solution) × 100%. Due to the possibility of solution generation failures, the overall pass rate (the percentage of all test cases) is lower than this value.

[0321] 3. Compare to the baseline:

[0322] Baseline 1: Purely manual processing (handled by senior engineers).

[0323] Baseline 2: Template matching scheme (only supports simple template filling for "missing features").

[0324] 4. Evaluation Results

[0325] index This invention artificial Template matching Solution generation success rate 76% N / A 31% No manual adjustment of proportions required 38% N / A 12% Average processing time (minutes) 2.8 45 15 First pass rate 87% 55% 41%

[0326] Note: The above data is based on internal testing, and the actual results may vary depending on the document type and the complexity of the opinions.

[0327] Example 1: Missing Element (Main Example)

[0328] Reviewer's comment: "The class diagram lacks exception handling mechanisms, such as classes for handling network timeouts or database connection failures." This is a typical example of a vague comment.

[0329] System execution process:

[0330] (1) Document structuring: Parse the document and construct a reference relationship graph. The class diagram is located in Chapter 4 and contains 12 class nodes. Regular expression extraction successfully obtained the class diagram nodes.

[0331] (2) Semantic analysis: The defect type is determined to be "missing element". Key elements: {Missing object: "class", domain: "exception handling, network timeout, database connection failure"}.

[0332] (3) Problem localization: The keyword "abnormality" led to Section 4.2 "Abnormality Handling Instructions" and Section 4.3 Class Diagram. Graph propagation: Since the out-degree of the class diagram nodes is 3 (referenced by 3 other chapters), a 2-step propagation was adopted to include the 5 classes that depend on "network requests" and "database operations" in the affected area.

[0333] (4) HTN planning: Top-level task (add-missing-element class ExceptionManager sec_4.2). The task library decomposes and generates scheme A (adding a new class and adding dependencies). The parent node relaxes and generates scheme B (adding to the interpretation) and scheme C (not adding a new class, modifying existing classes). Scheme A has a confidence of 0.92, scheme B has a confidence of 0.78, and scheme C has a confidence of 0.61.

[0334] (5) User confirmation: The engineer selects option A and clicks “Execute”.

[0335] (6) Execution and Verification: The system automatically modifies the document and runs the validator. The rule layer check passes; the semantic layer verification passes (the independent rule validator detects that the newly added class contains the keyword "exception handling"); the reference consistency check passes. A "qualified" report is output.

[0336] Results: In a typical test environment (approximately 200 pages of document, standard hardware configuration), the time from input of feedback to solution generation (excluding user confirmation time) is approximately 10-30 seconds, and the time to execute modifications after user confirmation is approximately 5-10 seconds, with an overall processing time of approximately 2.8 minutes (including user reading and confirmation time). This represents a significant improvement in efficiency compared to manual processing (which typically takes tens of minutes to several hours).

[0337] Example 2: Naming Conflict

[0338] Reviewer's comment: "The 'User' class in Section 3.2 has a different meaning from the 'User' class in Section 5.1, resulting in a naming conflict."

[0339] System execution process:

[0340] (1) Document structuring: Parse the document and locate the two User class nodes.

[0341] (2) Semantic analysis: The defect type is determined to be "name conflict". Key elements: {conflict_name: "User",locations: "Section 3.2 and Section 5.1"}.

[0342] (3) Problem location: Due to the semi-fuzzy opinion, it is directly located to two nodes. The number of graph propagation steps is limited to 1 step. The affected area includes these two nodes and their direct references.

[0343] (4) HTN Planning: Top-level task resolve-naming-conflict. The system generates two candidate solutions: Solution A (rename User in Section 3.2 to User_Authentication) and Solution B (merge the two classes, retaining User in Section 5.1). After simulation verification, the confidence level of Solution A is 0.89, and the confidence level of Solution B is 0.72 (because multiple references need to be updated after merging). The user chooses Solution A.

[0344] (5) Execution and Verification: The system renames the User class in Section 3.2 to User_Authentication and updates all text references to this class. Verification passes, and a pass report is output.

[0345] Example 3: Logical Contradiction

[0346] Reviewer's comment: "Chapter 2 states that the interface returns XML format, but Chapter 5 states that the interface returns JSON format, which is contradictory."

[0347] System execution process:

[0348] (1) Document structuring: Parse the document and locate two related paragraphs.

[0349] (2) Semantic analysis: The defect type is determined to be "logical contradiction". Key elements: {contradiction_pairs:[["interface returns XML format", "interface returns JSON format"]]}.

[0350] (3) Problem location: Keyword matching locates two paragraph nodes, the graph propagates in 2 steps, and the affected area includes other chapters that reference these two paragraphs.

[0351] (4) HTN Planning: Top-level task resolve_contradiction. System generation scheme: Modify the description in Chapter 5 to "the interface returns XML format". Simulation verification shows that the semantic similarity meets the standard (0.82). User confirmation.

[0352] (5) Execution and verification: The system modifies the text of Chapter 5, re-verifies the contradictory pairs, and the rule verification is passed (positive and negative keywords no longer exist at the same time), and outputs a qualified report.

[0353] Novelty and inventiveness of the present invention

[0354] Novelty

[0355] 1. Apply HTN planning to the automatic modification of software document review comments, especially for ambiguous review comments.

[0356] 2. The document modification scheme generation integrates automatic extraction of natural language reference relationships (regular expressions + semantic fallback) and content template generation, solving the problem of the lack of a compiler to provide a dependency graph in the document domain.

[0357] 3. A three-level simulation verification (structure + reference + semantics) for evaluating document modification schemes is proposed to replace the compilation verification in code repair.

[0358] creativity

[0359] When addressing the problem of automatically modifying ambiguous comments in documents, those skilled in the art typically employ methods such as enhancing semantic understanding and template matching, or using end-to-end generation models. This invention creatively introduces HTN planning and designs modules for natural language citation extraction (regular expressions + semantic fallback), template content generation, and three-level verification, tailored to document characteristics. These modules are not simply a collection of known technologies, but rather non-obvious designs addressing the unique challenges of the document domain (no compiler, no precise syntax, and the need to generate natural language). In particular, the sub-problem of "extracting citation relationships from natural language" has long existed in document analysis and remains unsolved. The explicit citation extraction mechanism proposed in this invention, based on a combination of regular expressions and semantic matching of chapter titles, combined with HTN planning, forms a complete technical solution, demonstrating originality.

[0360] Non-obviousness: HTN planning in the field of code repair relies on precise information provided by the compiler. Those skilled in the art have no incentive to use HTN planning for document modification because the documents lack structured information. This invention solves this adaptation problem by designing regular expression extraction of references and semantic fallback of chapter titles, as well as template content generation. This adaptation is not obvious. Specifically:

[0361] (1) Adaptation barrier due to lack of structured information in documents: Software development documents lack dependency graphs provided by the compiler, and the reference relationships are hidden in the natural language text, resulting in inconsistencies, omissions, and errors. The conventional approach is to enhance natural language understanding rather than introduce HTN planning. This invention breaks through this mindset by constructing a usable reference relationship graph through "regular expression extraction + chapter title semantic backtracking" and combining this graph with HTN planning.

[0362] (2) Deterministic requirement of content generation: Document modification needs to generate natural language text that conforms to the context and is readable. Those skilled in the art may directly use large language models to generate it, but large models have problems with uncertain output and unauditable nature. This invention chooses to fill the predefined template library, which ensures the determinism and auditability of the modification scheme, which is not a conventional practice.

[0363] (3) Fundamental difference in verification methods: Document modification lacks automated verification methods such as compilers or unit tests. The three-level simulation verification (structure, reference, semantics) of this invention is specifically designed for document modification, and there is no similar verification system in the prior art. Decomposing verification into three progressive levels and designing specific inspection rules for each level is an innovative architecture in itself.

[0364] This invention solves a long-standing technical problem: the automated processing of fuzzy review comments is a recognized challenge in the field of software engineering document quality assurance, and provides an end-to-end solution.

[0365] Compared to end-to-end generation methods based on Large Language Models (LLM): In existing technologies, some researchers have attempted to directly generate modified document fragments using LLM. However, in certain scenarios, LLM suffers from output uncertainty, difficulty in guaranteeing citation integrity, and uninterpretable modification steps. The HTN planning + template generation method of this invention provides a deterministic modification scheme with better auditability and engineering feasibility.

[0366] Unexpected technical effects: In the aforementioned test environment, compared to purely manual processing, the average processing time for a single ambiguous opinion was reduced from 45 minutes to 2.8 minutes (a reduction of approximately 94%) after adopting this invention; compared to a baseline system based on rule and keyword matching, the first-pass yield increased from 41% to 87% (an increase of approximately 112%). This significant effect exceeded the reasonable expectations of those skilled in the art based on the simple superposition of the performance of each module (NLP, planning, and verification).

[0367] This invention provides a method for automatically generating modification schemes for software development documents in response to ambiguous review comments. Compared with existing technologies, it has the following advantages:

[0368] 1. High degree of automation: From vague opinions to executable modification plans, the entire process is generated automatically, and engineers only need to confirm with one click.

[0369] 2. Correctness of modifications is guaranteed: HTN planning ensures that the operation sequence meets the preconditions and reference integrity, and three-level simulation verification and the final verifier further ensure correctness.

[0370] 3. High interpretability: The decomposition process of HTN is traceable, and the reasons for each modification step are clear.

[0371] 4. Practical and reproducible: This invention discloses key technical contents such as task library definition, operation primitive mapping, and model fine-tuning details, which can be used by those skilled in the art to implement the system.

[0372] 5. Broad application prospects: It can be integrated into R&D management platforms such as Jira, GitLab, and ZenTao, and is suitable for industries with strict document quality requirements such as finance, communications, automotive electronics, and aerospace and military.

[0373] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for automatically generating modification plans for software development documents in response to ambiguous review comments, characterized in that, The method includes: Step S1: Document structuring and reference graph construction Convert the original document into a machine-readable reference graph structure with reference relationships; Step S2: Semantic analysis of fuzzy opinions Defect types and key elements are extracted from the review comments text; a fine-tuned BERT model is used for multi-label classification to extract defect types and sequence labeling to extract key elements. At the same time, a rule backup mechanism based on dependency parsing is provided to handle marginal cases with low model confidence. Step S3, Problem Identification Based on the key elements output in step S2, find the set of seed nodes most likely to be affected in the document reference relationship graph, and determine the affected area through variable step graph propagation; Step S4: Generation of Modification Scheme Based on HTN Planning Based on the defect type and key elements output in step S2, and the seed node set and affected area output in step S3, a candidate modification scheme composed of operation primitives is automatically generated using hierarchical task network (HTN) planning. Step S5: Three-level simulation verification and ranking of candidate solutions A three-level verification system is used to quickly evaluate the structural feasibility, citation integrity, and semantic correctness of each candidate modification scheme, and output the confidence ranking. Step S6: Modify execution and verify results Apply the selected changes to the actual document and perform final verification; Step S7, Human-Machine Collaborative Feedback Record user feedback behavior to continuously optimize the model and task library.

2. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 1, characterized in that, S1 includes: Step 1.1 Document Format Parsing: Use python-docx or PyMuPDF to extract the ID, type, content, and location of each semantic unit; output a list of nodes in JSON format; Step 1.2 Explicit Reference Extraction: Write regular expressions to match common reference patterns; if a match is successful, add directed edges and record the reference type and the ID of the referenced node; Step 1.3 Semantic backtracking based on chapter titles: When the number of reference edges extracted by regular expressions is less than a preset threshold or the number of reference edges for any chapter in the document is 0, semantic matching based on chapter titles is used as a supplement; the specific method is as follows: for any two chapter nodes A and B, Sentence-BERT is used to encode their titles into vectors respectively, and the cosine similarity is calculated; if the similarity is ≥0.75, a bidirectional edge is added and marked in the log; Step 1.4 Graph structure storage: Node = semantic unit; Edge = reference relationship, stored using graph data structure.

3. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 1, characterized in that, S2 includes: Step 2.1 Data Labeling Specifications: Define the labeling boundaries for the following 8 defect types: missing elements, naming conflicts, structural inconsistencies, logical contradictions, non-compliant formatting, incomplete descriptions, incorrect citations, and redundant content. The labeling specifications adopt the BIO tag system, and perform sequence labeling on the following key elements in the opinion text: missing object MISS, domain term DOMAIN, and citation location REF; the corresponding BIO tags are: B-MISS, I-MISS, B-DOMAIN, I-DOMAIN, B-REF, I-REF, and O; Step 2.2 Model Fine-tuning: (1) Training data collection: Collect historical review opinions, which are independently annotated by domain experts, with a Kappa coefficient of not less than 0.85; the opinions come from software projects in multiple domains, covering three types of documents: requirements, design, and testing; perform data augmentation: synonym replacement, random insertion or deletion of stop words, and after augmentation, the script automatically checks whether the key elements have been modified; if the key elements have been modified, the augmented sample is discarded or the original sample is reverted; (2) Training parameters: Set the learning rate, batch size, and epochs, and use BERT-base-chinese as the pre-training weights; (3) Model output: For a review comment, the BERT model outputs the defect type and BIO label sequence. The defect type is a multi-label classification result, and the BIO label sequence is used to extract key elements. Step 2.3 Rule Backup Mechanism: When the BERT model's classification confidence for a certain opinion falls below a preset threshold or key elements are missing in the sequence labeling results, it reverts to a rule extraction method based on dependency parsing; Rule extraction process: (1) Use LTP to perform clause segmentation and dependency parsing analysis on the original opinion; (2) Identify the following trigger words: missing elements: missing, lacking, omission; naming conflicts: conflict, duplicate names, ambiguity; inconsistent structure: incorrect hierarchy, inconsistent structure, disordered chapters; logical contradictions: contradiction, conflict, inconsistency; non-standard format: incorrect format, inconsistent style, incorrect font; incomplete description: incomplete, missing parameters, missing steps; citation errors: citation errors, reference errors, link errors; redundant content: redundant, repetitive, superfluous; (3) By analyzing syntactic dependency relations, the objects that the trigger words affect are extracted, including: missing objects, domain words, and reference positions; the rule extraction results will be converted into BIO tag sequences and structured dictionaries consistent with the BERT output format; Step 2.4 Output Structured Dictionary: Convert the model output or rule backup output into standard JSON format.

4. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 1, characterized in that, S3 includes: Step 3.1 Precise matching based on position identifiers: If the review comments are semi-fuzzy, i.e., contain explicit chapter or figure numbers, the position identifier is directly parsed into the corresponding document node ID, used as a seed node, and assigned a confidence level of 1.0; subsequent keyword matching and semantic expansion steps are skipped; parsing rules: use regular expressions to match chapter number patterns and search for nodes with matching position attributes in the document reference relationship graph; if the parsed position identifier does not exist in the document, a fallback to manual specification by the user is triggered; Step 3.2 Keyword Matching: For non-semi-fuzzy opinions, extract key elements and perform fuzzy matching on the titles and texts of all nodes; the matching method is Jaro-Winkler distance, and the threshold is set to 0.85; return a list of all successfully matched node IDs as candidate seed nodes; if there are no matches, proceed to Step 3.3; Step 3.3 Semantic Expansion: When no keyword matching results are found, Sentence-BERT is used to calculate the cosine similarity between the key element and the summary of each node; the top-3 nodes with the highest similarity are selected as candidate seed nodes; if the highest similarity is <0.4, it is considered a matching failure and a fallback to manual specification by the user is triggered. Step 3.4 Variable Step Graph Propagation: Starting from the set of seed nodes, propagate outwards in the document citation graph to determine the area of ​​influence. The number of propagation steps is dynamically determined: For each seed node, calculate its out-degree + in-degree, i.e., the citation degree; if the degree > threshold T, propagate 2 steps; otherwise, propagate 1 step; the maximum number of steps is 2 to avoid space explosion; the threshold T is dynamically calculated: T = max(3, min(10, total number of nodes / 20)), where the total number of nodes is the total number of nodes in the document citation graph; for seed nodes from semi-fuzzy opinions, the maximum number of propagation steps is set to 1 step; if the number of nodes in the area of ​​influence after propagation exceeds 50, the propagation is truncated and the user is notified. Step 3.5 Output Location Results: Output a list of node IDs in the affected area and the confidence score for each node; the confidence score calculation rule is as follows: (1) Direct localization of semi-fuzzy opinions: seed node base confidence 1.0; (2) Keyword matching successful: Seed node basic confidence level 1.0; (3) Semantic expansion match successful: seed node base confidence 0.8; The nodes obtained from graph propagation are calculated as follows: base confidence × 0.6, multiplied by 0.6 for each propagation step, up to a maximum of two steps.

5. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 1, characterized in that, In S4, a hierarchical task network (HTN) is used for planning; A predefined task library is used to recursively decompose top-level tasks into atomic tasks using the HTN planner. Atomic tasks include document modification primitives and internal graph operations. Document modification primitives are mapped to the document editing API, and internal graph operations only modify the reference relationship graph. For operations that require adding or updating text content, natural language text is generated by filling in the pre-built content template library. In the HTN planning process, a planning layer abstraction is introduced. The planning layer maintains a lightweight view of the document, containing only the structure, reference relationships, and text summary, without storing the complete long text, in order to ensure planning efficiency. There is a mapping relationship between the planning layer state and the real document state: the nodes in the planning layer correspond to the semantic units in the real document, and the operation primitives in the planning layer are ultimately mapped to document editing API calls. When the template is filled, the system reads the required context from the real document.

6. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 1, characterized in that, S4 specifically includes: Step 4.1 Planning Layer State Representation: The planning layer maintains a lightweight document state for precondition judgment and effect simulation; the state consists of the following information: (1) Node set: Each node contains: unique identifier id, semantic unit type type, node name name, text summary text_summary, keywords keywords, parent node ID parent_id, position in the document, and format attributes format_props; (2) Reference edge set: directed edges, representing the reference relationship from node from to node to; Step 4.2 Definition of Operation Primitives: (1) Define document modification primitives, including: ADD_NODE, DELETE_NODE, UPDATE_TEXT, MOVE_NODE, UPDATE_FORMAT, and SPLIT_NODE, which are used to map to the document editing API; (2) Define internal graph operations, including: ADD_EDGE, REMOVE_EDGE, which only modify the internal reference relationship graph of the system and do not change the document text; Step 4.3 HTN Task Library Definition: Define the top-level task and its decomposition method for each defect type; Missing elements: (1) Top-level task: add-missing-element(type, name, parent_region), where type is the type of missing element, name is the name or identifier of the element, and parent_region is the suggested parent node region; (2) Decomposition method: The prerequisites are that the parent node exists and is of chapter type, and there is no node with the same name in the current document; subtasks are executed in sequence: Generate text: Call the template filling engine to generate natural language text for new nodes based on type, defect type, and key features; text=generate_template_text(type, "Feature missing", "add", {name, domain,missing_object}); Add a node: Creates a new node under the parent node; ADD_NODE(parent, position="end", node_type=type, content_text=text); Maintaining reference relationships: Within the influence region, find existing nodes that should semantically reference the new node: Calculate the Jaccard similarity between the text keywords of the existing node and the name / type keywords of the new node. If it is ≥0.3, it is determined as a candidate node; for each candidate node, add an edge pointing to the new node in the internal reference relationship graph. Naming conflict: (1) Top-level task: resolve-naming-conflict(conflict_name, node1, node2, parent_region), where conflict_name is the name of the conflict, node1 and node2 are the two nodes that are in conflict, and parent_region is the affected region; (2) The system provides two decomposition methods, which can be selected according to the preconditions; Method 1: Rename one of the nodes Prerequisites: Two nodes exist, and both have the name `conflict_name`; the two nodes are not equivalent; Subtask sequence: Generate a unique identifier for the new name: generate_unique_name(conflict_name, node1); specifically, if a numeric suffix can be added to the name of node1, try adding a numeric suffix until it does not conflict with any existing node name; if node1 is a class name, add a prefix based on its chapter; Update the text content of node1: replace all occurrences of conflict_name with new_name, map them to the UPDATE_TEXT primitive, and use a local substitution mechanism; Update all reference edges pointing to node1: Traverse the reference graph and update the old name in the text of the reference source node to the new name; Method 2: Merge two nodes Precondition: The two nodes are semantically substitutable, which is determined by the is-substitutable(node1, node2) function. This function is implemented by using Sentence-BERT to calculate the cosine similarity between node1.text_summary and node2.text_summary. If the similarity is ≥0.85, then they are determined to be substitutable. Subtask sequence: Redirect all reference edges pointing to node2 to node1: For each (from, node2) edge, execute REMOVE_EDGE(from, node2) and ADD_EDGE(from, node1, ref_type); Update the reference target name in the text of the from node; Delete node2: Mapped to the DELETE_NODE primitive, DELETE_NODE(node2); Step 4.4 Execution: Using the HTN planner, the planner outputs a sequence of atomic tasks, each atomic task being a primitive or internal operation from Step 4.2; where the input is: (1) Initial state: The document state defined in Step 4.1; (2) Top-level task: generated based on the defect type in step S2 and the area of ​​influence in step S3; (3) Planning parameters: Time limit is 3 seconds. If the time limit is exceeded, the longest valid prefix generated so far will be returned and the scheme will be marked as "partially generated". The maximum number of atomic task steps is 10. If the number of steps is exceeded, the task will be truncated to the first 10 steps and a warning log will be recorded. Step 4.5 Candidate solution generation: Multiple candidate solutions are generated through parent node relaxation and multi-method decomposition; Parent node relaxation: For the top-level task `add-missing-element`, the parent node of the seed node in step S3 is used by default; in addition, attempts are made to place the new node under other candidate parent nodes to generate different solutions: (1) Default parent node: the chapter where the seed node is located; (2) Distance priority: Chapter nodes that are 1 distance away from the seed node graph; (3) Semantic priority: Other sub-sections that share the same grandparent node as the seed node; (4) Hierarchical priority: Chapters in the document outline that are at the same level as or one level higher than the seed node; The above four strategies are all heuristic rules, covering the most likely insertion positions of new nodes in the document structure; try all applicable strategies, generate corresponding candidate solutions, and finally determine the optimal solution by simulation verification and user selection; Each choice corresponds to a candidate solution; if a choice causes the preconditions to be unmet, it is skipped. Multi-method decomposition: For the same top-level task, multiple methods may be defined in the task library; each method generates an independent candidate solution; Ultimately, the system collects all valid candidate solutions, retaining a maximum of 5. Step 4.6 Content Template Population: For the ADD_NODE and UPDATE_TEXT primitives, specific natural language text needs to be generated; a predefined template library + keyword population method is used to ensure the determinism and auditability of the generated content; Template library structure: The template library uses (node_type, defect_type, operation_type) as the key, where operation_type can be either add or update; Placeholder sources: (1) Key elements output in step S2; (2) Seed node attributes output in step S3; (3) Document structure information; (4) System generation; Local replacement mechanism: Full replacement based on regular expressions is used; if it is necessary to retain the original text of the non-replaced parts, sentence-level replacement can be enabled through configuration; (1) Extracting old names: Extracting from the original node text through named entity recognition or keyword matching; if extraction is not possible, the entire node text will be replaced by default; (2) Template usage: The template is only used to generate new content for the replaced sentence, and the other sentences in the original paragraph remain unchanged after the replacement; Template population timing: Template population occurs during the planning execution phase. When `call get-template-text` is invoked, the system generates text in real time based on the current context. The generated text is then passed as a parameter to the `ADD_NODE` or `UPDATE_TEXT` primitive. Step 4.7 Output Candidate Solutions: The output of step S4 is a list of candidate solutions, each solution containing a solution ID, top-level task description, atomic task sequence, and generation method label; this list will be passed to step S5 for three-level simulation verification and sorting.

7. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 6, characterized in that, Other defect types in Step 4.3 include: structural inconsistencies, logical contradictions, incorrect citations, redundant content, non-standard formatting, and incomplete descriptions.

8. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 6, characterized in that, S5 includes: Step 5.1 First level: Structure check: Without simulation execution, first check whether each operation primitive in the operation sequence satisfies its precondition. If any operation violates the precondition, the scheme is directly discarded and will not proceed to subsequent verification. Step 5.2 Simulate execution: Copy a lightweight copy of the current document graph, including the affected region and all external reference nodes pointing to that region; apply primitives sequentially according to the operation sequence to update node attributes and reference edges in the copy; Copy strategy: It must include all nodes within the affected area and all source nodes of all reference edges pointing from the outside to the affected area; if the number of external reference nodes pointing to the affected area exceeds the preset threshold, all nodes will still be copied, but the verification report will mark "There are many external reference nodes, and the verification time may be long"; for UPDATE_TEXT operations, if they affect the external reference nodes themselves, the complete text attributes of these nodes must also be copied. During the simulation: each primitive is executed sequentially, and the replica state is updated; if a primitive fails to execute, the scheme is marked as "execution error" and will not proceed to subsequent verification; Step 5.3 Level 2: Referential Integrity Check: On the document graph copy after the simulation execution, check the validity of all reference edges; The algorithm checks the integrity of references by traversing all reference edges (from, to) in the original graph. If the "to" node is deleted after simulation, and none of the candidate solutions contain any of the following operations, then the algorithm fails the integrity check. (1) The operation of deleting this edge; (2) Update the text of the from node to remove or modify the reference; (3) Redirection operation; For newly added nodes, check if they have at least one incoming edge; if not, mark them as "isolated nodes" and deduct confidence points, but do not directly determine them as failing; users can configure whether to require incoming edges. Pass condition: There are no unprocessed dangling references; orphaned nodes only affect the confidence level, not the pass / fail status; Step 5.4 Level 3: Semantic Similarity Check: For UPDATE_TEXT and ADD_NODE operations, evaluate whether the modified text maintains semantic coherence and adequately addresses the problem; (1) Inspection rules (2) Passing condition: If the similarity of all related operations is not lower than the corresponding threshold, then the third level is passed; otherwise, it is not passed. Step 5.5 Confidence Calculation: Calculate a confidence score for each candidate solution for ranking; all solutions that pass the first level participate in subsequent validation and confidence calculation; solutions that pass the second level have their confidence scores calculated normally and participate in ranking; solutions that do not pass the second level also have their confidence scores calculated, but are placed after the passing group in the ranking and are only for manual reference. Base score: Assuming the first level is passed with pass_structure=1, the second level is passed with pass_integrity=1, the third level is passed with pass_semantic=1, and failing is 0, then the base score = (pass_structure + pass_integrity + pass_semantic) / 3; Step penalty factor: Let step_count be the number of atomic operation steps in the scheme; then the step penalty factor = 1 - 0.05 × (step_count / 10); the fewer the steps, the closer the factor is to 1, and it is 0.95 when the number of steps is 10; Orphan node penalty: If there is at least one newly added orphan node in the scheme, the penalty is multiplied by an additional 0.

9. Final confidence formula: Confidence = Base score × Step penalty factor × Isolated node penalty factor; Step 5.6 Sorting and Output: Arrange the candidate solutions according to the rules and output them to the front end for display; the sorting rules are as follows: (1) Grouping by whether the second level has been passed: all schemes that have passed the second level are listed first, and those that have not passed are listed last; (2) Within each group, groups are sorted in descending order of confidence level; (3) If the confidence levels are the same, sort them in ascending order by the number of steps; Output format: Each solution includes a solution ID, a list of atomic tasks, a confidence level, and pass / fail status flags for each level; Step 5.7 User selects candidate solution: The front end displays a sorted list of candidate solutions; after the user selects, proceed to step S6.

9. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 1, characterized in that, S6 includes: Step 6.1 Execute the modifications: Using the python-docx or PyMuPDF API, execute each operation primitive in sequence according to the operation sequence; log after each step; establish a transaction rollback mechanism; Step 6.2 Perform integrity verification: Check whether the changes have successfully taken effect in the actual document; Verification content: For ADD_NODE: Check if the target node exists in the document; For DELETE_NODE: Check if the node no longer exists; For UPDATE_TEXT: Check if the node text has been updated with the new content; For MOVE_NODE: Check if the node has already appeared under the new parent node; For UPDATE_FORMAT: Check if the format attributes have been changed; For SPLIT_NODE: Check if the original node has been split and if the two new nodes exist; Pass condition: All operations are verified successfully; if any operation fails to take effect, it is judged as unqualified, and the reason for failure is output. Step 6.3 Defect Fix Adequacy Verification: Extract the complete text of the relevant areas of the modified document and use an independent, rule-based validator to check whether the modifications have resolved the issues pointed out in the original review comments; If the rule verification passes → it is deemed qualified; If the rule verification fails, but the defect type output by the model in step S2 is consistent with the original opinion and the confidence level is ≥0.8, it is marked as "requires manual review". If the rule validation fails and the model output in step S2 does not match, it is judged as unqualified. Pass condition: Rule verification passes; if rule verification fails but is marked as "requires manual review", the system will pause and wait for user judgment. Step 6.4 Reference Consistency Verification: Check whether all references to the modified nodes in the revised document are still valid; Verification content: Traverse the document reference graph; For each reference edge (from, to): if the to node has been deleted and the reference has not been cleaned up using UPDATE_TEXT or REMOVE_EDGE, an error is reported; if the to node has been moved, check if the position marker in the reference text needs to be updated. For newly added nodes, if the user configures them to "must be referenced", then check if there is at least one reference pointing to them; Pass condition: No invalid references exist; if invalid references exist, the result is considered invalid, and the specific reference location is output. Step 6.5 Comprehensive Judgment and Report Output: Acceptance criteria: Steps 6.2, 6.3, and 6.4 must all pass; if Step 6.3 is marked as "Requires manual review", the system will pause and wait for user confirmation before continuing. Output content: Judgment results, modification basis report, including: original review comments text, selected candidate solutions, comparison of document differences before and after modification, verification results at each level and specific reasons for failure.

10. The method for automatically generating software development document modification schemes for fuzzy review comments as described in claim 9, characterized in that, S7 includes: Step 7.1 Log Collection: Record the following information and store it in JSON format: original review comments text, defect type identified in step S2, list of candidate solutions generated by the system, solution ID selected by the user, user modifications to the solution, final verification result, and timestamp; Step 7.2 Offline Retraining: Retraining Trigger Conditions: Every 500 pieces of new data that are fully adopted and verified by users, or a monthly scheduled trigger; Retraining process: The newly accumulated fully adopted data is combined with the historical training set, and the BERT-base-chinese model is fine-tuned. Statistical analysis of the HTN decomposition failure rate for various defects; If the failure rate of a certain type of defect exceeds 30% for two consecutive months, a manual optimization prompt will be triggered, and domain experts will supplement new task decomposition methods. Step 7.3 Model Version Management: After each retraining, a new version of the model is generated, and the 5 most recent historical versions are retained; the new version of the model is evaluated using a fixed validation set, and the performance metric is the macro average F1 score of the defect type classification; if the macro average F1 score of the new version of the model drops by more than 5 percentage points in absolute terms compared with the current production version, it is automatically rolled back to the previous version, an alarm log is recorded, and automatic retraining is paused for manual analysis.