A document automatic generation method based on code change difference

CN122816692APending Publication Date: 2026-09-25SHENZHEN ZHILIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610892036.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]随着软件迭代速度加快,代码版本频繁更新,研发人员在完成代码修改后,往往需要人工梳理代码变更背后的技术改进、创新逻辑,并撰写对应专利文档,流程繁琐、效率低下;

Benefits of technology

本发明提供一种基于代码变更差异的文档自动生成方法,通过代码变更完成后即可自动生成技术文档,消除开发与专利撰写之间的时间差,降低新颖性丧失风险;准确性:基于代码语义而非人工描述识别创新点,避免遗漏和误述;系统性:通过变更影响图识别跨文件、跨模块的系统性创新,而非仅关注单个函数的修改;保护范围合理:从代码抽象到技术方案时自动剥离实现细节,使权利要求覆盖合理范围。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816692A_ABST
    Figure CN122816692A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of information intelligent processing, and particularly relates to a document automatic generation method based on code change difference, which comprises the following steps: S1, obtaining the code change difference, and parsing the text-level change difference into a semantic-level change unit, wherein each change unit comprises a change type, a change position and a change context; technical documents can be automatically generated after the code change is completed, the time difference between development and patent writing is eliminated, and the risk of novelty loss is reduced; accuracy: based on code semantics rather than artificial description to identify innovation points, and omission and misstatement are avoided; systematicness: through a change influence graph, systematic innovation across files and modules is identified, rather than only focusing on the modification of a single function; reasonable protection scope: implementation details are automatically stripped when abstracting from the code to the technical scheme, so that the claims cover a reasonable range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information intelligent processing technology, specifically a method for automatically generating documents based on code change differences. Background Technology

[0002] In today's era of rapid software iteration, technical improvements such as functional optimization, performance upgrades, vulnerability fixes, and architectural iterations of software systems are all achieved through code changes. The differences in changes generated during code version iterations are the most core and direct carriers of software technological innovation. In the context of intellectual property application, R&D personnel need to sort out the innovation logic, refine the technical solutions, and write complete patent specifications, claims, and drawings based on the technical improvements made during code iterations to complete the patent application.

[0003] Firstly, for the R&D side, it can systematically sort out scattered R&D ideas, fragmented technical improvement points, and iterative technical solutions, transforming non-standardized R&D results such as verbal ideas, code optimization, and technical improvements into standardized and written technical data, so as to effectively preserve and retain technological innovation results and avoid the loss of core technical ideas with personnel movement and R&D iteration.

[0004] With the rapid pace of software iteration and frequent code version updates, developers often need to manually analyze the technical improvements and innovative logic behind code changes after completing code modifications, and write corresponding patent documents, which is a cumbersome and inefficient process. First, code changes are only presented in text form. Basic differences such as additions, deletions, and modifications cannot directly reflect the technical semantics. Manually interpreting the intent of the changes is time-consuming and easily overlooks core technical points. Secondly, there are complex dependencies such as calls and references between code modules. A single code change can have a chain reaction along the dependency chain, making it difficult for humans to fully understand the scope of the change's propagation and the overall technical logic. Third, patent documents have extremely high requirements for the structure and logical standardization of the technical solutions. Manual drafting is prone to problems such as non-standard format, inaccurate extraction of technical points, and unreasonable layout of claims. Therefore, a document automatic generation method based on code change differences is proposed to address the above problems. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies and solve at least one of the technical problems mentioned in the background, this invention proposes a method for automatically generating documents based on code change differences.

[0006] The technical solution adopted by this invention to solve its technical problem is: a document automatic generation method based on code change differences, comprising: S1. Obtain code change differences and parse text-level change differences into semantic-level change units. Each change unit contains change type, change location, and change context. S2. Using change units as nodes and symbolic dependencies as edges, construct a change impact graph to describe the propagation path of changes in the codebase; S3. Based on the change impact diagram, identify technological innovations using predefined innovation point identification rules; S4. Structure each technological innovation point into a triplet of problem, solution, and effect; S5. Automatically generate patent documents based on triples, including the specification, specification drawings, and claims.

[0007] Preferably, the change types of the semantic-level change unit include: adding functions, modifying call relationships, renaming symbols, adding data structures, adding inter-thread communication mechanisms, adding synchronization mechanisms, and adding conditional branching logic; the change context includes associated header files, build scripts, and cross-file dependencies.

[0008] Preferably, the innovation identification rules include: when the change introduces a new inter-thread communication mechanism, it is identified as a concurrent and asynchronous architecture innovation; when the change involves renaming multiple component symbols or splitting files, it is identified as a modular and isolated architecture innovation; when the change introduces a conditional branch containing degradation processing logic, it is identified as a fault-tolerant and adaptive innovation; and when the change introduces a new memory allocation path, it is identified as a resource management innovation.

[0009] Preferably, the problem, solution, and effect triple is generated as follows: the technical problem is inferred from the original defects in the replaced or deleted code; the technical solution is extracted from the core mechanism in the added or modified code; and the technical effect is derived from the inherent attributes of the solution.

[0010] Preferably, the automatic generation method of the patent document is as follows: the background technology of the specification is derived from the problem of the triple, the invention content is derived from the solution, the beneficial effects are derived from the effect, and the specific implementation method is extracted from the code implementation in the change impact diagram; the drawings of the specification are extracted from the architecture diagram and flowchart in the change impact diagram; the independent claims of the claims cover the core solution abstracted from the common features of the change cluster, and the dependent claims cover the specific implementation details of each change unit.

[0011] Preferably, the change impact graph is constructed through the following steps: analyzing the symbol definitions and symbol references involved in each change unit; establishing calling relationships, header file inclusion relationships, and system registration relationships between symbols; constructing a directed graph with change units as nodes and symbol dependencies as directed edges; performing connectivity component analysis on the directed graph, aggregating the interconnected change units into change clusters, with each change cluster corresponding to a potential technological innovation point.

[0012] Preferably, the semantic parsing module is used to parse text-level code change differences into semantic-level change units; The Change Impact Graph Construction Module is used to construct a directed graph of change propagation based on symbolic dependencies. The innovation point identification module is used to identify technological innovation points from the change impact diagram based on predefined rules; The triplet generation module is used to structure technological innovation points into triplets of problem, solution, and effect. The document generation module is used to automatically generate specifications, specification drawings, and claims based on triples and change impact diagrams.

[0013] Preferably, a computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the method described in any of the preceding claims.

[0014] The beneficial effects of this invention are: This invention provides an automatic document generation method based on code change differences. Technical documents can be automatically generated immediately after code changes are completed, eliminating the time lag between development and patent drafting and reducing the risk of loss of novelty. It boasts several advantages: accuracy (identifying innovation points based on code semantics rather than manual description, avoiding omissions and misstatements); systematicity (identifying systemic innovations across files and modules through change impact diagrams, rather than focusing solely on modifications to individual functions); and reasonable scope of protection (automatically stripping implementation details when abstracting from code to a technical solution, ensuring the claims cover a reasonable scope).

[0015] This invention provides a method for automatically generating documents based on code change differences. By parsing code change differences into semantic-level change units, constructing a change impact diagram, and automatically identifying innovation points according to rules, it achieves accurate identification of systematic innovations from code semantics without relying on human experience, thus avoiding omissions of innovation points and description deviations.

[0016] This invention provides a method for automatically generating documents based on code change differences. After code changes are completed, structured technical documents can be automatically generated, streamlining the entire process from code development to patent drafting. This enables development and patent drafting to proceed simultaneously, eliminating time lags and effectively reducing the risk of loss of novelty due to technology disclosure.

[0017] This invention provides an automatic document generation method based on code change differences. By abstracting the common features of the change cluster, automatically stripping away code implementation details, and generating a triplet of problem, solution, and effect, the scope of protection of the claims is reasonably defined, avoiding an overly narrow scope of protection while ensuring sufficient technical feature support, thereby enhancing the value of patent protection. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention.

[0019] In the attached diagram: Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram illustrating the working principle of the semantic parsing engine in this invention; Figure 3 This is a schematic diagram illustrating the impact of changes in this invention; Figure 4 This is a schematic diagram of the rule mapping for identifying the technical innovation points in this invention; Figure 5 This is a schematic diagram illustrating the mapping relationship between triples and patent documents in this invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Specific implementation examples are given below.

[0022] Please see Figure 1 , Figure 4 This invention provides a method for automatically generating documents based on code change differences, comprising the following steps: S1. Code Change Difference Acquisition and Semantic Analysis Obtain code change differences from a version control system (such as Git), for example, by obtaining text-level differences through gitdiff, and then parse them into semantic-level change units, each of which contains the change type, change location, and change context.

[0023] Change types include: adding functions, modifying call relationships, renaming symbols, adding data structures, adding inter-thread communication mechanisms, adding synchronization mechanisms, and adding conditional branching logic; change contexts include associated header files, build scripts, and cross-file dependencies.

[0024] This example uses the refactoring of an embedded audio system code: This change involves renaming functions in 9 source files, creating 2 new files, and updating 2 build scripts. An overview of the file-level changes is as follows: Delete: webrtc / webrtc / alg_sram_alloc.h, alg_sram_alloc.c New: webrtc / webrtc / webrtc_sram_alloc.h / .c, audio, alg, core / include / alg_sram_alloc.h / src / alg_sram_alloc.c Modify: a total of 9 files including webrtc / webrtc / agc / agc.c and webrtc / webrtc / aec / ring_buffer.cpp. Update: CMakeLists.txt build scripts for the relevant components. S2, semantic parsing is a change unit The above text-level diff is further parsed into structured, semantic change units: Change Unit 1: The symbol is renamed alg_malloc → webrtc_alg_malloc, and its scope covers multiple files such as webrtc / agc / agc.c and webrtc / aec / *.cpp; Change Unit 2: Added a new thread function ob_algo_task_main, located in sndcard_onboard.c; Change Unit 3: Added an inter-thread communication mechanism, the algo_queue, located in sndcard_onboard.c; Change Unit 4: Added a new synchronization mechanism, rtos_enter_critical, for critical section protection, located in sndcard_onboard.c; Change Unit 5: Added degradation and fault tolerance logic, including multi-level conditional branches such as if(algo_queue)...elseif..., located in sndcard_onboard.c.

[0025] S3, Construction of Change Impact Diagram Using change units as nodes and symbolic dependencies as directed edges, a change impact graph is constructed and connectivity component analysis is performed to obtain interconnected change clusters. Each cluster corresponds to a systemic improvement. Changes to Cluster A: Asynchronous Algorithm Pipeline A new independent algorithm thread, ob_algo_task_main, has been added. A new inter-thread communication queue, algo_queue, has been added. Add critical section protection to ensure mutual exclusion between access to algo_out and readyflag; Add multi-level conditional branches to implement degradation handling logic when the algorithm is overloaded.

[0026] Change Cluster B: Component-level SRAM isolation The memory interface alg_malloc was globally renamed to webrtc_alg_malloc, involving 9 source files; Create a new webrtc_sram_alloc.c / h file to independently implement SRAM allocation for the WebRTC component; Create a new alg_sram_alloc.c / h file to provide an independent memory copy for the audio algorithm side; Update the build scripts for both components to achieve compile-time symbol isolation and linking isolation.

[0027] S4. Identification of Technological Innovation Points and Generation of a Triplet of Problem, Solution, and Effect Based on the change impact diagram and predefined rules, two core technological innovations were identified, and a triplet of problem, solution, and effect was automatically generated: Asynchronous algorithm pipeline and adaptive degradation control The original synchronous algorithm and the acquisition thread were executed serially. Fluctuations in the algorithm's processing time would block the acquisition thread, and in the event of a timeout, it would directly cause audio frame drops and stuttering. Independent algorithm threads and frame queues are introduced for decoupling, along with output buffers, single-frame latency models, and conditional branching logic for automatic degradation when overloaded; The acquisition thread is zero-blocking, the algorithm latency is deterministic, and the system can adaptively degrade under high load to ensure continuous audio output.

[0028] Dual-heap architecture SRAM allocation and component-level memory isolation The system uses the PSRAM main heap by default. When the audio algorithm accesses it frequently, the access latency is 10-30 times higher than that of the on-chip SRAM, which affects real-time performance and stability. The allocation side uses on-chip SRAM in a targeted manner, and the release side reclaims it uniformly. Combined with global symbol renaming and independent memory for components, a dual-heap architecture is formed. It can significantly reduce memory access latency without intruding on business logic, ensure complete memory isolation between components and prevent failures from affecting each other, and has cross-platform portability compatibility.

[0029] S5. Automatic generation of patent documents Based on the above triplet and change impact diagram, a complete patent document is automatically generated: The background technology in the instruction manual is automatically derived from the "technical problems" corresponding to the two innovations; the invention content is abstracted and summarized from the "technical solutions"; the beneficial effects are summarized from the "technical effects"; and the specific implementation methods are filled in with the code change details corresponding to the change impact diagram. The accompanying diagrams in the manual are automatically extracted from the change impact diagram, including the asynchronous algorithm pipeline architecture diagram, the SRAM isolation module partitioning diagram, and the degradation control flowchart. The common features of the modified cluster are abstracted into independent claims, and dependent claims are drafted with details of each modified unit, forming a claim system with clear hierarchy and reasonable scope of protection.

[0030] Step 5: Automatically generate patent documents Automatically generate three patent documents based on triples: The background technology is derived from the "problem", the invention content is derived from the "solution", the beneficial effects are derived from the "effects", and the specific implementation methods are extracted from the code implementation in the change impact diagram; The accompanying diagrams in the manual include: an architecture diagram (node ​​→ module, edge → data flow) extracted from the change impact diagram, a flowchart extracted from the degradation branch, and a sequence diagram extracted from inter-thread communication. Claims: Independent claims cover the core scheme (abstracted from the common features of the change cluster), while dependent claims cover the specific implementation details of each change unit.

[0031] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A method for automatically generating documentation based on code change differences, comprising: S1. Obtain code change differences and parse text-level change differences into semantic-level change units. Each change unit contains change type, change location, and change context. S2. Using change units as nodes and symbolic dependencies as edges, construct a change impact graph to describe the propagation path of changes in the codebase; S3. Based on the change impact diagram, identify technological innovations using predefined innovation point identification rules; S4. Structure each technological innovation point into a triplet of problem, solution, and effect; S5. Automatically generate patent documents based on triples, including the specification, specification drawings, and claims.

2. The method according to claim 1, characterized in that, The change types of the semantic-level change unit include: adding functions, modifying call relationships, renaming symbols, adding data structures, adding inter-thread communication mechanisms, adding synchronization mechanisms, and adding conditional branching logic; the change context includes associated header files, build scripts, and cross-file dependencies.

3. The method according to claim 1, characterized in that, The rules for identifying innovations include: when a change introduces a new inter-thread communication mechanism, it is identified as a concurrent and asynchronous architecture innovation; when a change involves renaming multiple component symbols or splitting files, it is identified as a modular and isolated architecture innovation; when a change introduces a conditional branch containing degradation processing logic, it is identified as a fault-tolerant and adaptive innovation; and when a change introduces a new memory allocation path, it is identified as a resource management innovation.

4. The method according to claim 1, characterized in that, The problem, solution, and effect triple is generated as follows: the technical problem is inferred from the original defects in the replaced or deleted code; the technical solution is extracted from the core mechanism in the added or modified code; and the technical effect is derived from the inherent attributes of the solution.

5. The method according to claim 1, characterized in that, The automatic generation method of the patent document is as follows: the background technology of the specification is derived from the problem of the triple, the invention content is derived from the solution, the beneficial effects are derived from the effect, and the specific implementation method is extracted from the code implementation in the change impact diagram; the drawings of the specification are extracted from the architecture diagram and flowchart in the change impact diagram; the independent claims of the claims cover the core solution abstracted from the common features of the change cluster, and the dependent claims cover the specific implementation details of each change unit.

6. The method according to claim 1, characterized in that, The change impact graph is constructed through the following steps: analyzing the symbol definitions and symbol references involved in each change unit; establishing calling relationships, header file inclusion relationships, and building system registration relationships between symbols; and constructing a directed graph with change units as nodes and symbol dependencies as directed edges. Connectivity component analysis is performed on the directed graph to aggregate interconnected change units into change clusters, and each change cluster corresponds to a potential technological innovation point.

7. A document automatic generation system based on code change differences, comprising: The semantic parsing module is used to parse text-level code change differences into semantic-level change units; The Change Impact Graph Construction Module is used to construct a directed graph of change propagation based on symbolic dependencies. The innovation point identification module is used to identify technological innovation points from the change impact diagram based on predefined rules; The triplet generation module is used to structure technological innovation points into triplets of problem, solution, and effect. The document generation module is used to automatically generate specifications, specification drawings, and claims based on triples and change impact diagrams.

8. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1 to 6.