Generating a software documentation hierarchy

HGEN, an automated pipeline using large language models, addresses the challenge of inadequate software documentation by generating a well-organized hierarchy of artifacts, enhancing code comprehension and maintenance tasks through high-quality documentation.

WO2025184519A1PCT designated stage Publication Date: 2025-09-04UNIV OF NOTRE DAME DU LAC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017862
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Creating and maintaining high-quality, multi-level software documentation is incredibly time-consuming, leading to inadequate documentation in many code bases, which results in inefficient code structure, difficult debugging, and wasted resources.

Method used

An automated pipeline, HGEN, uses large language models to transform source code into a well-organized hierarchy of artifacts with trace links, addressing the challenges of inadequate documentation by generating a multi-layer hierarchy of software artifacts.

Benefits of technology

HGEN produces high-quality documentation hierarchies that accelerate code comprehension and maintenance tasks, improving code updates and impact analysis by providing a comprehensive and readable documentation structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017862_04092025_PF_FP_ABST
    Figure US2025017862_04092025_PF_FP_ABST
Patent Text Reader

Abstract

A system for generating a software documentation hierarchy. The system includes an electronic processor. The electronic processor is configured to generate clusters of lower-level artifacts based on natural language descriptions. Each lower-level artifact includes a natural language description. The electronic processor is also configured to, for each cluster, create a higher-level artifact. The higher-level artifact includes a natural language description of the cluster. The electronic processor is further configured to refine the higher-level artifacts containing overlapping information and generate trace links to connect the lower-level artifacts with the higher-level artifacts.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATING A SOFTWARE DOCUMENTATION HIERARCHYGOVERNMENT LICENSE RIGHTS

[0001] This invention was made with government support under Grant No. CCF 1909007 and IIP2122689 awarded by the National Science Foundation (NSF). The government has certain rights in the invention.RELATED APPLICATIONS

[0002] This application claim priority to U.S. Provisional Application No. 63 / 559,413, filed February 29, 2024, the entire content of which is hereby incorporated by reference.FIELD

[0003] Implementations described herein relate to generating software documentation.SUMMARY

[0004] Software documentation supports a broad set of software maintenance tasks. However, creating and maintaining high-quality, multi-level software documentation can be incredibly time-consuming. Therefore, many code bases suffer from a lack of adequate documentation, which results in inefficient code structure, difficult and inefficient debugging and code updates, and wasted resources associated with the same. Implementations described herein provide a (e.g., fully) automated pipeline that leverages large language models (LLMs) to transform source code through a series of six blocks into an organized hierarchy of formatted documents to address the above-noted problems of inadequate source code documentation. This automated pipeline is referred to herein as HGEN. HGEN produces artifact hierarchies with high quality and coverage of the core concepts as compared to existing approaches. Accordingly, HGEN is a useful tool for accelerating code comprehension and maintenance tasks. In some implementations, the artifacts hierarchy or documentation hierarchy produced by HGEN for source code included in a code base may be used to modify the source code to generate improved source code. For example, the documentation hierarchy may be used in impact analysis to find locations in source code to apply a change. Therefore, a documentation hierarchy generated by HGEN may act as aguide in making appropriate changes to source code (e.g., in an automated fashion) and, therefore, aid in the generation of improved source code.

[0005] For example, one implementation provides a system for generating a software documentation hierarchy. The system includes an electronic processor. The electronic processor is configured to generate clusters of lower-level artifacts based on natural language descriptions. Each lower-level artifact includes a natural language description. The electronic processor is also configured to, for each cluster, create a higher-level artifact. The higher-level artifact includes a natural language description of the cluster. The electronic processor is further configured to refine the higher-level artifacts containing overlapping information and generate trace links to connect the lower-level artifacts with the higher-level artifacts.

[0006] Another implementation provides a method for generating a software documentation hierarchy. The method includes generating clusters of lower-level artifacts based on natural language descriptions. Each lower-level artifact includes a natural language description. The method also includes, for each cluster, creating a higher-level artifact. The higher-level artifact includes a natural language description of the cluster. The method further includes refining the higher-level artifacts containing overlapping information and generating trace links to connect the lower-level artifacts with the higher-level artifacts.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] FIG. 1 is a flowchart of an example method for generating a software documentation hierarchy, in accordance with some implementations.

[0008] FIG. 2 is an example summary of HERO source code generated when the method of FIG. 1 is performed, in accordance with some implementations.

[0009] FIG. 3 is an example of clusters generated for the HERO source code when the method of FIG. 1 is performed, in accordance with some implementations.

[0010] FIG. 4 is an example of user stories generated for the HERO source code when the method of FIG. 1 is performed, in accordance with some implementations.

[0011] FIG. 5 is an example of user stories clustered during refinement when the method of FIG. 1 is performed, in accordance with some implementations.

[0012] FIG. 6 is an example of refined user stories generated for the HERO source code when the method of FIG. 1 is performed, in accordance with some implementations.

[0013] FIG. 7 is an example of trace links generated for a user story included in FIG. 5 when the method of FIG. 1 is performed, an accordance with some implementations.

[0014] FIG. 8 illustrates additional trace links generated for user stories when the method of FIG. 1 is performed, in accordance with some implementations.

[0015] FIG. 9 is an example system for generating a software documentation hierarchy, in accordance with some implementations.DETAILED DESCRIPTION

[0016] One or more implementations are described and illustrated in the following description and accompanying drawings. These implementations are not limited to the specific details provided herein and may be modified in various ways. Furthermore, other implementations may exist that are not described herein. Also, the functionality described herein as being performed by one component may be performed by multiple components in a distributed manner.Likewise, functionality performed by multiple components may be consolidated and performed by a single component. Similarly, a component described as performing particular functionality may also perform additional functionality not described herein. For example, a device or structure that is “configured” in a certain way is configured in at least that way but may also be configured in ways that are not listed. Furthermore, some implementations described herein may include one or more electronic processors configured to perform the described functionality by executing instructions stored in non-transitory, computer-readable medium. Similarly, implementations described herein may be implemented as non-transitory, computer-readable medium storing instructions executable by one or more electronic processors to perform the described functionality. As used in the present application, “non-transitory computer-readable medium” comprises all computer-readable media but does not consist of a transitory, propagating signal. Accordingly, non-transitory computer-readable medium may include, forexample, a hard disk, a CD-ROM, an optical storage device, a magnetic storage device, a ROM (Read Only Memory), a RAM (Random Access Memory), register memory, a processor cache, or any combination thereof.

[0017] In addition, the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. For example, the use of “including,” “containing,” “comprising,” “having,” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items. The terms “connected” and “coupled” are used broadly and encompass both direct and indirect connecting and coupling. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings and can include electrical connections or couplings, whether direct or indirect. In addition, electronic communications and notifications may be performed using wired connections, wireless connections, or a combination thereof and may be transmitted directly or through one or more intermediary devices over various types of networks, communication channels, and connections. Moreover, relational terms such as first and second, top and bottom, and the like may be used herein solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions.

[0018] Unless the context of their usage unambiguously indicates otherwise, the articles “a,” “an,” and “the” should not be interpreted as meaning “one” or “only one.” Rather these articles should be interpreted as meaning “at least one” or “one or more.” Likewise, when the terms “the” or “said” are used to refer to a noun previously introduced by the indefinite article “a” or “an,” “the” and “said” mean “at least one” or “one or more” unless the usage unambiguously indicates otherwise.

[0019] Also, it should be understood that the illustrated components, unless explicitly described to the contrary, may be combined or divided into separate software, firmware and / or hardware. For example, instead of being located within and performed by a single electronic processor, logic and processing described herein may be distributed among multiple electronic processors. Similarly, one or more memory modules and communication channels or networksmay be used even if implementations described or illustrated herein have a single such device or element. Also, regardless of how they are combined or divided, hardware and software components may be located on the same computing device or may be distributed among multiple different devices. Accordingly, in the claims, if an apparatus, method, or system is claimed, for example, as including a controller, control unit, electronic processor, computing device, logic element, module, memory module, communication channel or network, or other element configured in a certain manner, for example, to perform multiple functions, the claim or claim element should be interpreted as meaning one or more of such elements where any one of the one or more elements is configured as claimed, for example, to make any one or more of the recited multiple functions, such that the one or more elements, as a set, perform the multiple functions collectively.I. INTRODUCTION

[0020] Software documentation supports a broad set of software maintenance tasks, such as, for example, impact analysis, change analysis, requirements validation, safety assessment, and new developer onboarding. However, creating and maintaining consistent multi-level software documentation and its associated trace links is incredibly time-consuming. The process of documenting an implemented software system, defining the implemented software system, and maintaining documentation that describes the implemented system is often viewed as overly burdensome by developers and stakeholders. This perception leads to the documentation process being ignored, delayed, or inadequately sustained, especially in startups and small companies where speed is often prioritized over comprehensive requirements engineering processes. Consequently, despite the many benefits of a systematic software documentation process, many code bases suffer from a lack of adequate documentation.

[0021] While there have been advancements in automating certain types of software documentation, such as, for example, application programming interface (API) specifications or the continuous deployment of embedded software documentation, efforts to automate the generation of comprehensive, multi-layered artifacts describing system features remain underexplored. Similarly, it remains challenging to ensure that documentation correctly represents theunderlying code base, is readable, is understandable, is well formatted, and is properlyorganized so that the documentation is useful to practitioners maintaining software systems and avoids wasted resources (including human and computational).

[0022] To address these and other challenges, implementations described herein provide HGEN. HGEN is an automated pipeline that generates a multi-layer hierarchy of documentation including artifacts. Artifacts may be, for example, low-level design descriptions or sub-system and system-level requirements formatted according to the norms of the currently adopted lifecycle process. HGEN constructs these artifacts and generates trace links that connect the artifacts into a meaningful hierarchy. Therefore, HGEN provides well organized documentation, designed to effectively support diverse software maintenance activities.

[0023] The description provided below describes the HGEN process and provides a simple running example taken from the open-source gaming domain. The below description also includes results from a study evaluating HGEN in which the quality of the HGEN generated hierarchy is assessed for three different software development projects. In the study, a key software developer was recruited from each of the three projects to compare HGEN’s generated documentation against documentation manually created for the project and against documentation created using an off-the-shelf LLM baseline. For each software project included in the study, the quality of the documentation was systematically evaluated by assessing the individual quality of each generated artifact, the overall coverage of concepts included in the software project, and the relationships between layers in the generated hierarchy.

[0024] The description below is structured into the following sections. Section I provides an introduction. Section II presents an example process for generating a software artifact hierarchy using HGEN. Section III describes the design of an example experiment for evaluating HGEN. Section IV describes a quantitative evaluation of the HGEN process described in Section II using the experiment described in Section III. Section V summarizes the overall benefits of the HGEN approach to creating source code documentation and describes future work.II. HGEN PROCESS

[0025] HGEN is a pipeline or method for generating hierarchies of software artifacts. As depicted in the example flowchart illustrated in FIG. 1, in some implementations, a method 100 forperforming HGEN includes six blocks. In block 105, a set of lower-level artifacts are accepted as inputs. In blocks 115-125, internal processing is performed, and, in block 130, higher-level artifacts are produced that constitute a new layer of documentation. This new layer of documentation may be used to provide a set of artifacts to block 105 as input to restart the process to create the next layer in the hierarchy. The HGEN process utilizes a pipeline to produce each layer of the documentation hierarchy. In some implementations, the lowest layer accepts source code as input and generates a natural language summary (block 110). Blocks 105 and 115-130 may form a pipeline, which is used to generate each subsequent layer, thereby incrementally constructing a hierarchy of progressively higher-level artifacts formatted according to the norms of the current software development process. For each block in the pipeline, example underlying artificial intelligence (Al) models used to support the transformation of lower-level to higher-level documentation are shown. However, the HGEN process is designed to be model-agnostic and, therefore different LLM models may be used instead of those illustrated in FIG. 1. The blocks for generating a single layer of documentation, according to some implementations, are described below.

[0026] At block 105, a set of lower-level natural language artifacts may be accepted and clustering may be performed on the lower-level natural language artifacts to identify related features and / or functionalities. In the case of generating a lowest-level of a hierarchy, where the inputs are source code artifacts, an additional pre-processing block (block 110) may be performed prior to performing block 105 to generate a natural language summary of the source code included in the artifacts. The summary generated for source code included in an artifact may serve as a proxy for the source code included in the artifact throughout the remaining blocks. In some implementations, as a result of performing the functionality described in relation to block 105, a set of clusters of lower-level artifacts is generated as output.

[0027] In some implementations, at block 1 15, for each cluster identified in block 105, a natural language description is generated using a targeted artifact format (for example, user story, feature description, and the like). The natural language description may serve as a body of the new layer of documentation.

[0028] In some implementations, at block 120, the content of artifacts containing overlapping information is refined to improve clarity, improve conciseness, and ensure each artifact focuses distinctly on one specific feature or functionality.

[0029] In some implementations, at block 125, the refined artifacts generated at block 120 are connected to lower-level input artifacts (when lower-level input artifacts exist) by dynamically generating trace links.

[0030] In some implementations, at block 130, the relationships established at block 125 are used to detect and remedy redundant artifacts. The relationships established at block 125 may also be used to produce a final set of output artifacts for the current layer.

[0031] In some implementations, when a higher-layer of artifacts is desired, the output artifacts produced at block 130 are passed as inputs to generate the next layer. In other words, blocks 105 and 115-130 of the method 100 may be repeated until all targeted layers have been generated.

[0032] A result of performing the method 100 may be a hierarchy of software artifacts, referred to as an artifact tree.

[0033] Throughout the remainder of the description, a small, straightforward code repository, referred to as HERO, will be used as a running example to support the description of the HGEN process. HERO is openly accessible at https: / / github.com / gbaman / QUB-CSC1011 -Module- Hero-Game.

[0034] Each block included in the HGEN pipeline or the method 100 is described in further detail below.Block 110: Code Summarization

[0035] The lowest level of the documentation hierarchy starts with source-code, and, therefore, in some implementations, pre-processing is performed to summarize the code into natural language. The pre-processing may serve two purposes: 1) the preprocessing transforms source code into natural language comparable to input artifacts of all other documentation layers and 2) the summary generated by performing the pre-processing has higher information density and less redundancy than raw code. The summary enables a large amount of information to be conveyedwithin a single context window of the LLM, thereby enhancing the LLM’s capacity to comprehend a broader scope of the system. Summarization tasks may be performed using one or more generative models. For example, Anthropic’s Claude 2.0 model, which returns similar results to OpenAI’s GPT-4 and has a large context window of 100-k tokens, may be utilized at block 110 to perform summarization. In one example, Anthropic’s Claude 2.0 is prompted to summarize source code by (i) initially outlining the functionality provided to the user by the code, and (ii) then creating a polished summary that explains how the code supports the described user behavior. The summary created at block 110 may become the starting input for the first iteration of the blocks 105 and 115-130 and represent the initial tier in the documentation hierarchy. An example summary dynamically generated by HGEN for Hero.java in the HERO code-base is depicted in FIG. 2.Block 105: Form clusters from Lower-level Artifacts

[0036] In some implementations, performing the functionality described in relation to block 105 may involve performing a multi -technique clustering approach with the following internal steps, labeled C1-C8 and described in further detail below. In some implementations, steps C1-C8 may be performed in an alternative order to the one described herein.

[0037] Cl. Preprocessing: In some implementations, natural language artifacts are converted into embeddings using one or more transformers, such as, for example, a Sentence-BERT transformer. The Sentence-BERT transformer model has the capacity to encode entire sentences rather than rely on word-level encoding. The Sentence-BERT transformer model also has consistent performance across a diverse range of tasks. As a result, in some implementations, the Sentence-BERT transformer is used for the transformations to embeddings described throughout the remainder of this description.

[0038] C2. Multi-Technique Clustering: A consensus-based approach including five different unsupervised clustering algorithms may be used to achieve diversity of cluster size, outlier detection, and geometric considerations. These five different unsupervised clustering algorithms are, for example, OPTICS, Spectral, Agglomerative, Affinity Propagation, and K-means. Insome implementations, each technique is used to individually cluster the vector representations or embeddings of the input artifacts, producing a diverse set of candidate clusters.

[0039] C3. Filter by Size The set of candidate clusters generated by using a plurality of clustering techniques may be highly diverse but include overlapping and redundant clusters. A solution to this problem of increased redundancy, that is created when multiple filtering techniques are used to achieve a highly diverse set of candidate clusters, may be to filter clusters based on size, such as, for example, by removing singletons and large clusters. For example, singleton clusters containing only one artifact may be temporarily set aside. Large clusters containing five or more artifacts may be discarded, as these tend to inhibit a LLM’s ability to identify and extract finer details. While the decision to remove large clusters may limit the potential for constructing higher-level abstractions across larger artifact groups, this is at least partially addressed below in the discussion regarding block 120 when clusters are allowed to reform.

[0040] C4. Cluster Scoring: In some implementations, an importance value (importance score) is assigned to or generated for each cluster remaining after filtering by size is performed at C3. In some implementations, an importance value is calculated using Equation 1. importance = (a ■ log(.s) + h) v (1)

[0041] In Equation 1, h may be the cohesion score for the cluster, v may be the voting score from the five clustering techniques, 5 may represent the cluster size (for example, the number of input artifacts included in a cluster), a may be the weight applied to the cluster size (for example, a small constant value). The voting score (v) may be computed by counting the number of times the exact cluster, with the same input artifacts, appears in the candidate set of clusters. Cohesion Qi) may be computed by averaging the cosine similarity of each artifact’s embedding to all its neighbors within a cluster. In some implementations, Equation 2 is used to calculate cohesion.

[0042] The size metric (a • log( )), included in Equation 1, may be used to compensate for the tendency for smaller clusters to have higher cohesion. The use of logarithm moderates theimpact of larger clusters, ensuring that their contribution to the importance score grows at a decreasing rate and prevents them from disproportionately dominating the score due to their size alone. The cohesion score may measure how cohesive a cluster is. Cohesion may not be the same as importance. The cohesion score may represent whether artifacts in a cluster are closely related (highly cohesive) versus not well related (low cohesive).

[0043] C5. Cluster Ranking: In some implementations, clusters are ranked in descending order by importance score. FIG. 3, depicts several clusters generated from the HERO code base arranged by their respective importance scores. While clusters 300 and 305 have lower cohesion scores than cluster 310, their overall importance score places them higher in the ranking.

[0044] C6. Cluster Cleansing: In some implementations, artifact outliers that deviate by 1.5 standard deviations or more from the average similarity to their neighbors (other artifacts included in the cluster) are removed from their clusters to eliminate dissimilar artifacts from the clusters. For example, the artifact “Crime.java” is removed from cluster 320 in FIG. 3.

[0045] C7. Cluster Selection: In some implementations, the clusters are iterated through in order of their importance (or rank) to determine which clusters to select for the final set. At each iteration, the cluster possessing the next highest importance score (referred to herein as the FocusCluster), alongside the set of clusters already chosen for inclusion (referred to herein as the InclusionSet), is considered. Given the initial prevalence of overlapping or redundant clusters in the consensus-based approach to clustering, in some implementations, it is assessed whether the FocusCluster contains artifacts not already present in the InclusionSet. After removing any artifacts shared with clusters in the InclusionSet, the FocusCluster is included in the final cluster set only if it maintains a predetermined size (e.g., two or more artifacts) and has a cohesion score greater than or equal to a predetermined amount (for example, greater than or equal to a lowest cohesion score of the top 75% of cohesion scores of clusters included in the candidate cluster set or greater than or equal to a twenty fifth percentile of cohesion scores of clusters included in the candidate cluster set). In the example illustrated in FIG. 3, cluster 315 contains four artifacts already included in cluster 310, so each of these artifacts are removed, leaving the cluster 315 with only one unique artifact. As a result, the cluster 315 fails to meet thesize threshold for inclusion in the final cluster set and is therefore excluded from the final cluster set.

[0046] C8. Handle Orphans In some implementations, it is determined whether there exist artifacts not assigned to a cluster in the final cluster set. Artifacts not assigned (unassigned) to a cluster may be referred to as orphan artifacts. For each orphaned artifact, its most similar cluster may be identified by, for example, computing the average cosine similarity between the orphan and all members of each cluster. A similarity that is close to (within a predetermined threshold of) the cluster’s overall cohesion score (for example, within 0.1 of the cohesion score), may indicate that the orphan can be added to the cluster without reducing the overall cohesion of the cluster. As a result, when a similarity between an orphan and each member of a cluster is close to the cohesion score, the orphan is incorporated into or assigned to that cluster. Unplaced orphans (orphans not added to other clusters) may be retained (added to the final cluster set) as singleton clusters.Block 115: Generate Documentation Content

[0047] Prior studies have shown the importance of well-formatted documentation. Therefore, the functionality described in relation to block 115 focuses on formatting the generated artifacts according to user configurations. In some implementations, a user is able to choose the type of artifacts they would like to include in their hierarchy. For example, a user may configure the HGEN to generate an agile documentation hierarchy composed of source code (lowest layer), user stories (middle layer), and epics (top layer); or a user may configure the HGEN to generate a traditional hierarchy composed of source code, design specifications, and multiple layers of requirements. In block 115, an LLM may be prompted to format the output of the desired artifacts by specifying (i) the output artifact type, (ii) the desired format of the output artifact, and (iii) the targeted number of document artifacts to be generated from the current cluster. In some implementations, the LLM utilized at block 115 may be Anthropic’s Claude 2.0. The artifact type may be specified by the user as part of system configurations. In some implementations, the format of the output artifact can either be predefined by the user as part of system configurations or generated by the LLM in a separate context window, which increases the degree of automation of the HGEN. Given that most LLM’s pre-training data contains examples of diverse commonartifact types, the LLM utilized at block 115 may be prompted to generate a standardized format for the artifact type. For instance, the LLM may generate the following user story template: “As a [type of user], I want to [action or goal] so that [reason or benefit]”.

[0048] In some implementations, the number of high-level artifacts (n targets) to generate for each cluster is based on two factors. The first factor may be concept diversity. Cohesion may measure the extent to which an artifact focuses on a single topic. Clusters with low cohesion typically encapsulate more topics and require more higher-level elements. Therefore, concept diversity may be determined to be the inverse of the cohesion (for example, the cohesion calculated in Equation 2), normalized so that the maximum concept diversity equals 1. The second factor may be information density. The amount of information within a cluster’s artifacts may play a key role in determining the number of higher-level artifacts required. The information density of the cluster may be estimated by comparing the size of the cluster’s artifacts to the average size of all artifacts of the same type. The number of high-level artifacts (for example, n_targets) may be the product of concept diversity and information density. To promote the emergence of a tree-like documentation structure, a constraint that n targets must be greater than 50% and less than 100% of the current cluster’s artifact count may be imposed.

[0049] Returning to the HERO example, it may be determined that cluster 320 (see FIG. 3) requires three higher level artifacts. In this example, the four code files included in the cluster 320 have a total of 730 lines of code (LOC), while the average file in the current layer of the hierarchy has 109 LOC. Information density may be estimated to be approximately 6.7 (730 / 109), reflecting the complex game logic contained within these core character-related classes. Normalized concept diversity may be computed to be 0.56. Therefore, by computing and truncating 0.56 * 6.7 n targets is determined to be three artifacts. FIG. 4 shows the three example subsequent user stories (artifacts 400, 405, and 410) generated for cluster 320.Block 120: Refine Content & Clusters

[0050] Because automatic clustering may not always match human judgment, some level of conceptual overlap across clusters is inevitable, resulting in duplicated content in the generated artifacts. The functionality described in relation to block 120 addresses this technical problemcreated by automated clustering by reducing duplicated content and refining the artifact clusters through three steps, labeled D1-D3. In some implementations, steps D1-D3 may be performed in an alternative order to the one described herein.

[0051] DI. Duplicate Identification'. To identify potential duplicates, the generated artifacts may be clustered using the algorithm described in relation to block 105. This creates groups of similar artifacts that are currently spread across different clusters. The most cohesive clusters are those most likely to contain duplicated content and thus are identified as duplicate clusters. In other words, clusters of high-level artifacts with a cohesion score above a predetermined threshold are determined to be duplicate clusters. For example, the artifact 400 is a generated artifact that is detected as similar to other character-related user stories from different clusters (for example, artifact 500 and 505 included in FIG. 5). FIG. 5 is an example of generated user stories (artifacts) for HERO that were clustered together in block 120. Although each user story (artifact) in FIG.5 originated from a different cluster in block 115, the user stories were clustered together in block 120 due to their shared focus on character customizations.

[0052] D2. Duplicate Content Identification'. In this step, it is determined what source artifacts (lower-level artifacts) led to the overlapping content so these source artifacts may be re-clustered together. The source artifacts contributing to the overlap may be identified by selecting those artifacts with the highest semantic similarity to the parent artifact (high-level artifact). A new cluster may be formed containing the selected source artifacts for each generated artifact in the duplicate cluster.

[0053] D3. Re-generation. At step D2, for each duplicate cluster, the source artifacts containing the overlapping content have been identified. At step D3, an LLM may regenerate new artifacts based on this focused context. Given the source artifacts determined at step D2, block 115 is repeated to generate a fresh set of artifacts centered around the core theme. These newly generated artifacts replace those in the duplicate clusters. For example, the overlapping artifacts 500, 505, and 510 are re-generated as shown in FIG. 6. The artifacts 600, 605 are now focused on a more distinct topic.Block 125: Generate Intra-cluster Trace Links

[0054] After the new artifacts have been generated, the artifacts may be connected to the current layer of the hierarchy via trace links. Given that a single cluster can produce multiple higher- level artifacts, it cannot be assumed that every lower-level artifact in a cluster should link to each of the resulting higher-level artifacts. Consequently, in some implementations, trace links are only created between artifacts that demonstrate strong semantic similarity using standard automated tracing techniques.

[0055] To generate the connections via trace links, embeddings for the higher-level artifacts may be generated. In some implementations, the cosine similarity between the embeddings for the higher-level artifacts and each low-level artifact from their originating clusters are calculated. Each cluster’s scores may be scaled using min-max scaling so that the highest score is adjusted to 1. In some implementations, considering the variability in scores across different clusters, only links where the similarity score is within two standard deviations (a predetermined threshold) of the maximum normalized score are generated. Typically, this results in a cutoff of approximately 0.8. However, if no links are generated for a lower-level artifact, a link with the higher-level artifact that has the greatest similarity may be established.

[0056] For example, the trace links established for artifact 605 are depicted in FIG. 7. Following scaling, both Hero.java and V illain.java attain high similarity scores and are consequently linked to the user story. Although SuperheroGameController.java receives a considerably lower score, its similarity to artifact 605 exceeds its scores with all other user stories, allowing it to trace to artifact 605 as well.Block 130: Generate Inter-cluster Trace Links between Duplicates

[0057] At block 130, any remaining artifacts with over-lapping content may be identified. This identification may aid in detecting potential trace links between clusters. The cosine similarity between each pair of generated artifacts may be computed and pairs with similarity scores more than two standard deviations above the mean may be marked as potential duplicates. In some implementations, for each duplicate pair, designated as A and B, it is determined whether any of B’s trace links should also trace to A and vice versa. Trace links may be formed if the similarityscore between B and the child is of a similar strength between A and the child (for example, within a difference of 0.1).

[0058] When two pairs of highly similar artifacts result in trace links with identical child artifacts, it signifies that the pair are likely duplicates. In such cases, one of the duplicates may be removed since the risk of losing crucial information is significantly reduced.

[0059] FIG. 8 is an illustration of this re-tracing process. Artifacts that were originally traced to the artifacts 410, 605 are shown shaded. After re-tracing, each user story gains an additional trace link to an artifact, represented as the unshaded boxes. Notably, artifact 410 possesses one trace link (Crime.java) not linked to artifact 605; however, in this example, both artifact 410 and artifact 605 are retained after block 130 is performed, as artifact 410 is linked to Crime.java, while artifact 605 is not. Despite artifacts 410 and 605 being flagged as similar, neither are removed from the final output because they have different trace links and, therefore, it cannot be confidently concluded that artifacts 410 and 605 are duplicates.III. EXPERIMENT DESIGN

[0060] Evaluating the effectiveness of documentation hierarchies is complex because there is no single ground-truth solution. While it is tempting to use automated assessment techniques, such as, for example, BLEU, METEOR, ROUGE, CIDEr, and SPICE, to detect overlapping terms across documents, it has been shown that the metrics do not align with human judgment about the quality of documentation. Therefore, the evaluation described herein primarily leverages human judgment. In the evaluation study, for each of three different projects (code base developments), a knowledgeable project stakeholder (for example, a person involved in the development of a code base) was recruited to systematically compare HGEN’s performance against project documentation which was created by developers of the code base. In this section the study is described.

[0061] Projects

[0062] The three projects involved in the first study are summarized in Table I. The three projects represent diverse domains and cover both traditional and agile development processes. Each project included source code and at least two types of natural language artifacts, organizedinto layers, and connected by trace links. Dronology is an unmanned aerial vehicle (UAV) project that includes 423 Java files, 211 design definitions, and 99 requirements, respectively. Dronology is an open-source small Unmanned Aerial System (sUAS) written in Java. It provides a platform for controlling and coordinating multiple sUAS to support search-and-rescue, surveillance, and scientific data collection missions. Given 27 Java code files, for Dronology, HGEN produces 25 design definitions (artifacts) and 13 functional requirements (artifacts), the baseline approach generated 4 design definitions and 7 functional requirements, the manually created documentation (documentation created by a human) includes 12 design definitions and 4 functional requirements.

[0063] SAFA is a software safety tool and includes 242 Vue source-code files, 262 functional requirements, and 101 features. SAFA is a software documentation management platform that leverages live traceability to build a knowledge graph and support change impact analysis. Given 49 code files, for SAFA, HGEN produces 47 functional requirements (artifacts) and 25 features (artifacts), the baseline approach generated 9 functional requirements and 9 features, the manually created documentation (documentation created by a human) includes 35 functional requirements and 10 features.

[0064] Finally, Jack of Clubs is an open-source first-person shooter game and includes 36 C++ files, 7 user stories, and 3 epics. Jack of Clubs is a re-creation of Ace of Spades, a voxel-based first-person shooter game. Given 36 code files, for Jack of Clubs, HGEN produces 21 user stories (artifacts) and 5 epics (artifacts), the baseline approach generated 10 user stories and 4 epics, the manually created documentation (documentation created by a human) includes 7 user stories and 3 epics.

[0065] Each project had an available senior developer with prior experience in developing documentation, to serve as the project expert and review the output of the HGEN process when applied to the projects.Techniques under Comparison

[0066] The first study involves three different treatments including (i) the HGEN generated documentation, (ii) a baseline LLM approach, and (iii) the project documentation previouslydeveloped manually by project personnel. The HGEN documentation was generated for the three projects following the functionality outlined in Section II of this paper.

[0067] The baseline LLM approach (“BASELINE”) was created for comparison purposes. The baseline LLM approach starts with the identical set of summarized code as the HGEN process and uses the Claude 2.0 LLM to generate comprehensive documentation for each artifact type used in HGEN. As with HGEN, in the baseline LLM, the process is repeated for each layer and artifacts are connected across layers by generating embeddings (using, for example, Sentence-BERT) for both the lower and higher-level artifacts. In the baseline LLM, trace links are established between artifacts across the two layers if the normalized cosine similarity score for the artifacts is greater than 0.7. One difference between the BASELINE and HGEN is the use of clustering techniques in HGEN as well as the later refinement steps.

[0068] Due to the size of Dronology and SAFA, a subset of source-code files and their generated documentation was extracted so that the project experts could conduct a more in-depth analysis of each artifact. In each case, the project expert identified the most critical source files, and the study included those files and their linked documentation.IV. QUANTITATIVE COMPARISON OF THE DOCUMENTATION QUALITY

[0069] Table 1 includes evaluation guidelines for assessing documentation quality.TABLE I

[0070] To evaluate the quality of the generated documentation the following research questions (RQ) were addressed.

[0071] RQ1 Artifact Quality: How does the quality of individual machine-generated artifacts compare to that of expert-made artifacts? This RQ was addressed by asking the project experts to evaluate the language, content, and effectiveness of each artifact. Language includes readability and appropriateness, content includes conciseness and importance, and effectiveness includes usefulness and helpfulness. The definitions provided to the human assessors are summarized in Table I.

[0072] RQ2. Coverage: Towhat extent are concepts that appear in the original documentation covered by the generated documentation? To answer this RQ concept coverage was measured and the extent to which concepts appearing in the original documentation appeared in the generated documentation was evaluated. Examples of concepts in HERO might be the ability to select different characters or purchase specific items at a shop.

[0073] RQ3. Relationships: How effectively does the machine-generated documentation build appropriate parent-child relationships inherent in multi-layer data structures? Relationships are represented by trace links, and, therefore, the relationships were evaluated using standard traceability metrics of recall and precision for the generated artifact tree.

[0074] RQ1. Artifact Quality

[0075] Project experts comparatively evaluated artifacts in their own project for (a) the manually constructed documentation, (b) HGEN, and (c) the baseline approach using a qualitative rubric associated with each quality in Table III. Evaluations were performed against a Likert scale ranging from 1-5, where 1 signified low quality and 5 represented high quality. The project experts were allowed to move freely between the different types of artifacts during this process.

[0076] Analysis of results: Due to non-normal score distributions, the Mann-Whitney U test, which is non-parametric, was employed to test if there was a notable difference in score distributions between each of the three treatments. Further, to account for multiple comparisons in the first study, 18 tests across six metrics and three documentation groups (human-made, base, HGEN), the Holm-Sidak method was used to adjust p-values to control for error rates, ensuring areliable statistical analysis when comparing documentation quality across groups. Table II reports the mean scores for the six quality attributes assigned by the project experts across all three projects, as well as the adjusted p-values. Table II highlights instances where the null hypothesis is rejected (p < 0.05), suggesting that scores from one distribution tend to have higher scores than the other.TABLE IIRESULTS OF MANN-WHITNEY U TESTS. CASES IN WHICH THE NULL HYPOTHESIS IS REJECTED (I E., WHERE (P < 0.05)) ARE BOLDED AND DEPICT CASES WHERE ONE TREATMENT OUTPERFORMED THE OTHER WITH RESPECT TO THE QUALITY ATTRIBUTE.Human vs. BaselineMetric Human Mean Baseline Mean Corrected P-valueReadability 3.90 4.44 0.002445Appropriateness 3.70 4.00 0.214476Conciseness 4.30 4.44 0.596927Importance 4.10 4.35 0.191984Usefulness 3.51 3.67 0.794294Helpfulness 3.82 3.88 0.990378Human vs. HGENMetric Human Mean HGEN Mean Corrected P-valueReadability 3.90 4.38 0.000104Appropriateness 3.70 4.16 0.001493Conciseness 4.30 4.32 0.618787Importance 4.10 4.33 0.095525Usefulness 3.51 4.14 7.09e-06Helpfulness 3.82 4.21 0.002445Baseline vs. HGENMetric Baseline Mean HGEN Mean Corrected P-valueReadability 4.44 4.38 0.990378Appropriateness 4.00 4.16 0.462604Conciseness 4.44 4.32 0.990378Importance 4.35 4.33 0.990378Usefulness 3.67 4.14 0.020646Helpfulness 3.88 4.21 0.172422

[0077] Discussion of results: The comparison between the human constructed documentation and the baseline approach returned comparable scores on all metrics except Readability, whichwas higher for the baseline approach. The same comparison between human and HGEN documentation showed that HGEN returned higher quality scores across four of the six metrics: Readability, Appropriateness, Usefulness, and Helpfulness. On the other hand, in a direct comparison of HGEN versus the Baseline method, the only significant difference observed was for Usefulness, with HGEN’ s higher score indicating that project experts thought its documentation was more useful than the baselines. These results confirm the findings of previous studies that LLMs are able to produce software documentation of comparable quality to humans. Notably, the main difference between HGEN and the baseline method is in the way HGEN constructs the hierarchy and not in the way it generates individual documents.

[0078] RQ2: Coverage

[0079] Each expert was asked to evaluate concept coverage by identifying concepts in the manual documentation and checking whether the concepts were adequately reflected in the generated documentation. The proportion of concepts was computed from the original documentation that were addressed in each of the generated documentations.

[0080] Analysis of results: Results are reported in Table III (below) in the column labeled “% Covered”. The percentage of artifacts in the generated documentation that included these concepts is also reported in the column labeled “Covered by.” The “Covered by” percentages for both the Baseline and HGEN documentation suggest that many artifacts focus on concepts that were not highlighted in the manually generated documentation.TABLE IIIPERCENTAGE OF CONCEPTS FROM ORIGINAL DOCUMENTATION CAPTURED PER HGEN VERSION, AS DETERMINED BY THE PRODUCT EXPERTS (E1-E3)

[0081] Discussion of results: HGEN demonstrates a notable increase in concept coverage compared to the baseline approach across all three projects, capturing twice as many concepts inboth JOC and SAFA, and an impressive 80% increase in the case of Dronology. Additionally, a larger portion of the HGEN-generated documentation, as indicated by the “Covered by” metric, centers on core concepts from the manual documentation, particularly in the case of Dronology. Given that the experts did not identify duplicate artifacts, it appears that both the Baseline and HGEN uncovered project aspects not emphasized in the manual documentation. Matched with the increase in ‘helpfulness’ returned by experiments for RQ1, it appears that the additional information could be helpful to project developers.

[0082] RQ3: Relationships

[0083] To evaluate the quality of relationships within each generated hierarchy, a basic tracing tool which visualized the generated documentation tree was developed. Project experts were also allowed to approve or decline existing links, adding new links if needed. Each expert performed this task twice — once for HGEN and once for the Baseline approach, resulting in the expert’s version of a ground truth solution for each generative technique. The generated trees for HGEN and Baseline were then evaluated against the modified ground truth version for each one and computed mean average precision (mAP), precision, and recall for each project using standard formulas. To compute mAP, which assesses the extent to which correct links appear at the top of a ranked list, the links were ordered according to their original cosine similarity scores. In addition, the number of orphan artifacts generated by each approach was assessed, as this aspect emerged as a key distinguishing factor through discussions with the three experts. Table IV presents these results.

[0084] Discussion of results: The results of evaluating the quality of relationships within each hierarchy generated with HGEN are shown in Table IV (provided below). The relatively high mAP scores (80% to 95.5%) and recall scores (47.2% - 81.4%) indicate that Sentence-BERT was able to capture a range of semantic similarities between artifacts. This is likely attributable to the LLM’s use of lower-level artifacts for generating higher-level ones, thereby creating a shared vocabulary. However, a significantly higher number of orphans was also observed in the Baseline approach versus HGEN, which could have lowered recall whilst increasing precision. HGEN’s lower orphan count suggests that it’s enhanced clustering techniques (specifically its use of a plurality of clustering algorithms) enabled it to capture the concepts in the low-level artifactsat higher levels of abstraction, identifying concepts that might otherwise have been overlooked in the baseline approach.TABLE IV Traceability Accuracy Metrics for Generative ApproachesProject Approach mAP Precision Recall # OrphansDronology Baseline 84.5% 47.2% 89.5% 17HGEN 94.0% 56.3% 93.4% 0SAFA Baseline 91.9% 49.2% 100% 28HGEN 94.5% 54.3% 98.4% 9JOC Baseline 95.5% 67.3% 74.5% 11HGEN 96.7% 81.4% 80.2% 1

[0085] HGEN may be capable of generating high quality documentation hierarchies for an extremely wide range of software projects. For example, HGEN may be utilized in the automotive industry, information technology (IT) services industry, aerospace industry, and the like. Additionally, HGEN may be used on code bases in a variety of software languages (for example, C, C++, C#, Java, Type Script, and the like). HGEN may also be used to generate documentation for code bases of a variety of sizes. HGEN has the potential for aiding in the comprehension of complex systems lacking sufficient documentation (especially in scenarios where the original authors are unavailable), expediting the onboarding of new developer, and ensuring regulatory compliance. In some implementations HGEN may generate documentation that includes details from external libraries. In some implementations, HGEN may be used to generate an intermediate level of documentation for code bases for which a high-level of documentation already exists. In some implementations, HGEN may be used to provide incremental support for documentation alongside code development.V. CONCLUSION

[0086] Thus, implementations described herein provide HGEN. HGEN is an LLM-based approach to automatically generating hierarchies of requirements documentation from source code. HGEN engineers a process that addresses some of the deficiencies identified with unaided models. The evaluation of HGEN provided herein is designed to target three aspects of the documentation and supports existing literature that LLM-generated documentation can matchor exceed the quality of documentation written by humans. It is also shown that HGEN is able to capture meaningful relationships across varied artifact levels and can identify nearly all of the core concepts found in expert-produced documentation, showing a considerable enhancement over a baseline LLM. The results of the evaluation of HGEN indicate that HGEN could substantially reduce the time and effort needed for comprehensive documentation creation, thereby aiding in software maintenance tasks. Moreover, HGEN presents potential for automating additional aspects of requirements engineering, paving the way forward towards ubiquitous documentation and traceability.

[0087] It should be understood that the functionality described herein can be performed via one or more computing devices, communication networks, and the like. For example, FIG. 9 illustrates a system 900 for implementing HGEN as described herein according to some implementations. As illustrated in FIG. 9, the system 900 includes a computing device 905, such as a server. The computing device 905 may communicate with one or more other computing devices over one or more wired or wireless communication networks 920. Portions of the wireless communication networks 920 may be implemented using a wide area network, such as the Internet, a local area network, such as a Bluetooth™ network or Wi-Fi, and combinations or derivatives thereof. It should be understood that the system 900 may include more or fewer computing devices and the specific configuration of devices illustrated in FIG. 9 is purely for illustrative purposes. For example, in some implementations, the functionality described herein is performed via a plurality of computing devices, such as a plurality of devices in a distributed or cloud-computing environment. Also, in some implementations, the computing device 905 may communicate with one or more other computing devices through one or more intermediary devices (not shown).

[0088] As illustrated in FIG. 9, the computing device 905 includes an electronic processor 950, a memory 955, and a communication interface 960. The electronic processor 950, the memory 955, and the communication interface 960 communicate wirelessly, over wired communication channels or buses, or a combination thereof. The computing device 905 may include additional components than those illustrated in FIG. 9 in various configurations. For example, in some implementations, the computing device 905 includes multiple electronic processors, multiplememory modules, multiple communication interfaces, or a combination thereof. Also, as previously noted, it should be understood that the functionality described herein as being performed by the computing device 905 may be performed in a distributed nature by a plurality of computers located in one or more geographic locations.

[0089] The electronic processor 950 may be a microprocessor, an application-specific integrated circuit (ASIC), and the like. The electronic processor 950 is generally configured to execute software instructions to perform a set of functions, including the functions (HGEN) described herein. The memory 955 includes a non-transitory computer-readable medium and stores data, including instructions executable by the electronic processor 950. The communication interface 960 may be, for example, a wired or wireless transceiver or port, for communicating over the communication network 920 and, optionally, one or more additional communication networks or connections.

[0090] As illustrated in FIG. 9, the memory 955 includes the HGEN pipeline 965 as described herein. It should be understood that, in some implementations, the functionality described herein as being provided by the HGEN pipeline 965 may be distributed and combined in various configurations, such as through multiple separate software applications. In some implementations, output generated via the HGEN pipeline 965, such as the hierarchy of artifacts can be stored, such as in the memory 955 or one or more other storage locations within the computing device 905 or external to the computing device 905. In some implementations, the computing device 905 also includes or communicates with one or more input / output devices, such as a display device, a keyboard, a mouse, a printer, a speaker, a microphone, a touchscreen, etc. to receive user input, provide user output, or a combination thereof. For example, the documentation described herein may be provided to a display device included in the computing device 905 or to a display device in a separate computing device, such as a client device communicating with the computing device 905.

[0091] Various features and advantages of some implementations are set forth in the following claims.

Claims

CLAIMSWhat is claimed is:1 . A system for generating a software documentation hierarchy, the system comprising: an electronic processor, the electronic processor configured to generate clusters of lower-level artifacts based on natural language descriptions, wherein each lower-level artifact includes a natural language description; for each cluster, create a higher-level artifact, wherein the higher-level artifact includes a natural language description of the cluster; refine the higher-level artifacts containing overlapping information; and generate trace links to connect the lower-level artifacts with the higher-level artifacts.

2. The system according to claim 1, wherein the electronic processor is further configured to: when creating a lowest level of artifacts in the software documentation hierarchy, generate a natural language description for each artifact included in a code base; generate clusters of code based artifacts based on the generated natural language descriptions; for each cluster, create a higher-level artifact, wherein the artifact includes a natural language description of the cluster; and refine the higher-level artifacts containing overlapping information.

3. The system according to claim 1, wherein the electronic processor configured to generate clusters of lower-level artifacts based on natural language descriptions by:converting the natural language description included in each lower-level artifact to an embedding; generating a set of candidate clusters based on the embeddings using a plurality of unsupervised clustering algorithms; filtering clusters from the set of candidate clusters based on size; assigning an importance score to each cluster included in the filtered set of candidate clusters; ranking each cluster included in the filtered set of candidate clusters in descending order by importance score; eliminating artifact outliers from each cluster included in the filtered set of candidate clusters; iteratively selecting clusters from the filtered set of candidate clusters for inclusion in a final cluster set; and when orphan artifacts exist, for each orphan artifact, determining whether to assign the orphan artifact to a cluster included in final cluster set or add a cluster containing only the orphan artifact to the final cluster set, wherein an orphan artifact is a lower-level artifact unassigned to a cluster included in the final cluster set.

4. The system according to claim 3, wherein filtering clusters from the set of candidate clusters based on size includes: filtering clusters containing a single lower-level artifact from the set of candidate clusters; and filtering clusters containing five or more lower-level artifacts from the set of candidate clusters.

5. The system according to claim 3, wherein artifact outliers deviate by 1.5 or more standard deviations from an average similarity to their neighbors.

6. The system according to claim 3, wherein iteratively selecting clusters from the filtered set of candidate clusters for inclusion in a final cluster set includes for each cluster included in the filtered set of candidate cluster, in order of rank, removing, from the cluster, lower-level artifacts included in a cluster included in an inclusion set, wherein the inclusion set includes clusters chosen for inclusion in the final cluster set; and when, after removing lower-level artifacts included in a cluster included an inclusion set, the cluster includes at least two lower-level artifacts and has a cohesion score greater than or equal to a twenty fifth percentile of cohesion scores of clusters included in the candidate cluster set, adding the cluster to the inclusion set.

7. The system according to claim 3, wherein determining whether to assign the orphan artifact to a cluster included in final cluster set or add a cluster containing only the orphan artifact to the final cluster set includes computing a similarity score for each cluster by computing an average cosine similarity between the orphan artifact and each lower-level artifact included in the cluster; and when the similarity score for a cluster is within a predetermined threshold of a cohesion score of the cluster, adding the orphan artifact to the cluster.

8. The system according to claim 1, wherein the electronic processor is configured to, for each cluster, create a higher-level artifact by: determining a number of higher-level artifacts to create for a cluster based on concept diversity of the cluster and information density of the cluster.

9. The system according to claim 1, wherein the electronic processor is configured to refine the higher-level artifacts containing overlapping information by: clustering the higher-level artifacts based on natural language descriptions, wherein each higher-level artifact includes a natural language description; determining duplicate clusters, wherein duplicate clusters are clusters of higher-level artifacts with a cohesion score above a predetermined threshold; for each higher-level artifact included in a duplicate cluster, forming a new cluster including one or more lower-level artifacts with high semantic similarity to the higher-level artifact, wherein the higher-level artifact is based on the one or more lower-level artifacts; generating new higher-level artifacts based on the new clusters; and replacing the higher-level artifacts included in the duplicate clusters with the new higher- level artifacts.

10. The system according to claim 1, wherein the electronic processor is configured to generate trace links to connect the lower-level artifacts with the higher-level artifacts by: for each higher-level artifact, generating an embedding for the higher-level artifact; and for each lower-level artifact from a cluster the higher-level artifact originated from, calculating a cosine similarity between the embeddings for the higher-level artifact and the lower-level artifact; normalizing the cosine similarity; and when the cosine similarity is within a predetermined threshold of a maximum normalized score, generating a trace link; andfor each pair of higher-level artifacts, computing a cosine similarity between each pair of higher-level artifacts; determining potential duplicates, wherein potential duplicates are a pair of higher-level artifacts with a cosine similarity more than two standard deviations above a mean of cosine similarities determined for pairs of higher-level artifacts; and when potential duplicates have trace links to identical lower-level artifacts, removing one of the potential duplicates.

11. A method for generating a software documentation hierarchy, the method comprising: generating clusters of lower-level artifacts based on natural language descriptions, wherein each lower-level artifact includes a natural language description; for each cluster, creating a higher-level artifact, wherein the higher-level artifact includes a natural language description of the cluster; refining the higher-level artifacts containing overlapping information; and generating trace links to connect the lower-level artifacts with the higher-level artifacts.

12. The method according to claim 11, the method further comprising: when creating a lowest level of artifacts in the software documentation hierarchy, generating a natural language description for each artifact included in a code base; generating clusters of code based artifacts based on the generated natural language descriptions; for each cluster, creating a higher-level artifact, wherein the artifact includes a natural language description of the cluster; and refining the higher-level artifacts containing overlapping information.

13. The method according to claim 11, wherein generating clusters of lower-level artifacts based on natural language descriptions includes: converting the natural language description included in each lower-level artifact to an embedding; generating a set of candidate clusters based on the embeddings using a plurality of unsupervised clustering algorithms; filtering clusters from the set of candidate clusters based on size; assigning an importance score to each cluster included in the filtered set of candidate clusters; ranking each cluster included in the filtered set of candidate clusters in descending order by importance score; eliminating artifact outliers from each cluster included in the filtered set of candidate clusters; iteratively selecting clusters from the filtered set of candidate clusters for inclusion in a final cluster set; and when orphan artifacts exist, for each orphan artifact, determining whether to assign the orphan artifact to a cluster included in final cluster set or add a cluster containing only the orphan artifact to the final cluster set, wherein an orphan artifact is a lower-level artifact unassigned to a cluster included in the final cluster set.

14. The method according to claim 13, wherein filtering clusters from the set of candidate clusters based on size includes: filtering clusters containing a single lower-level artifact from the set of candidate clusters; andfiltering clusters containing five or more lower-level artifacts from the set of candidate clusters.

15. The method according to claim 13, wherein artifact outliers deviate by 1.5 or more standard deviations from an average similarity to their neighbors.

16. The method according to claim 13, wherein iteratively selecting clusters from the filtered set of candidate clusters for inclusion in a final cluster set includes for each cluster included in the filtered set of candidate cluster, in order of rank, removing, from the cluster, lower-level artifacts included in a cluster included in an inclusion set, wherein the inclusion set includes clusters chosen for inclusion in the final cluster set; and when, after removing lower-level artifacts included in a cluster included an inclusion set, the cluster includes at least two lower-level artifacts and has a cohesion score greater than or equal to a twenty fifth percentile of cohesion scores of clusters included in the candidate cluster set, adding the cluster to the inclusion set.

17. The method according to claim 13, wherein determining whether to assign the orphan artifact to a cluster included in final cluster set or add a cluster containing only the orphan artifact to the final cluster set includes computing a similarity score for each cluster by computing an average cosine similarity between the orphan artifact and each lower-level artifact included in the cluster; and when the similarity score for a cluster is within a predetermined threshold of a cohesion score of the cluster, adding the orphan artifact to the cluster.

18. The method according to claim 11, wherein, for each cluster, creating a higher-level artifact includes:determining a number of higher-level artifacts to create for a cluster based on concept diversity of the cluster and information density of the cluster.

19. The method according to claim 11, wherein refining the higher-level artifacts containing overlapping information includes: clustering the higher-level artifacts based on natural language descriptions, wherein each higher-level artifact includes a natural language description; determining duplicate clusters, wherein duplicate clusters are clusters of higher-level artifacts with a cohesion score above a predetermined threshold; for each higher-level artifact included in a duplicate cluster, forming a new cluster including one or more lower-level artifacts with high semantic similarity to the higher-level artifact, wherein the higher-level artifact is based on the one or more lower-level artifacts; generating new higher-level artifacts based on the new clusters; and replacing the higher-level artifacts included in the duplicate clusters with the new higher- level artifacts.

20. The method according to claim 1 1, wherein generating trace links to connect the lower-level artifacts with the higher-level artifacts incudes: for each higher-level artifact, generating an embedding for the higher-level artifact; and for each lower-level artifact from a cluster the higher-level artifact originated from, calculating a cosine similarity between the embeddings for the higher-level artifact and the lower-level artifact; normalizing the cosine similarity; andwhen the cosine similarity is within a predetermined threshold of a maximum normalized score, generating a trace link; and for each pair of higher-level artifacts, computing a cosine similarity between each pair of higher-level artifacts; determining potential duplicates, wherein potential duplicates are a pair of higher-level artifacts with a cosine similarity more than two standard deviations above a mean of cosine similarities determined for pairs of higher-level artifacts; and when potential duplicates have trace links to identical lower-level artifacts, removing one of the potential duplicates.

Citation Information

Patent Citations

  • Methods, systems, and computer program products for automating releases and deployment of a softawre application along the pipeline in continuous release and deployment of software application delivery models

    US20190129701A1

  • Automatic generation of code documentation

    US20210357210A1

  • Computer-implemented verification of a hardware design implementation against a natural language description of the hardware design or software code against a natural language description of a software application

    US20230148108A1