A context-aware-based self-admission technical debt identification method, device and electronic equipment

CN122526609BActive Publication Date: 2026-09-18HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611006928.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-18
Estimated Expiration
2046-07-08

AI Technical Summary

Technical Problem

[0005]第一,计算效率不足

Benefits of technology

[0022] First, an efficient pre-screening mechanism is introduced. Through the annotation-based rapid screening strategy, a considerable number of annotation samples can be directly identified without the need for large language model inference on these samples, which greatly reduces inference overhead and significantly improves recognition efficiency without compromising detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122526609B_ABST
    Figure CN122526609B_ABST
Patent Text Reader

Abstract

The application provides a context-aware self-admission technical debt identification method, device and electronic equipment, which comprises the following steps: classifying code annotations in a target code file to obtain the category of each code annotation in the target code file, wherein the category comprises SATD annotations, non-SATD annotations and to-be-reasoned annotations; for each to-be-reasoned annotation, determining a code segment related to the semantic of the to-be-reasoned annotation as the context of the to-be-reasoned annotation; inputting the to-be-reasoned annotation and the context thereof into a pre-trained large language model, guiding the large language model to perform semantic reasoning and attribution analysis through a structured prompt word, and generating a response result of the to-be-reasoned annotation; and standardizing the response result output by the large language model, and outputting a structured identification result containing a classification label and reasoning basis. The application can greatly reduce the model reasoning overhead, improve the identification efficiency and reduce false positives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of software engineering technology, and specifically to a context-aware, self-acknowledgment-based debt identification method, apparatus, and electronic device. Background Technology

[0002] Software development teams typically need to deliver high-quality software products within a given timeframe or budget. However, the real-world software development process often faces various challenges, such as changing customer requirements, earlier project deadlines, and market competition. To address these challenges, developers are often forced to adopt temporary workarounds, such as hardcoding certain parameters, implementing only partial functionality, or copying unreviewed code from elsewhere. Furthermore, due to inexperience or oversight, developers may unintentionally introduce suboptimal code implementations. These phenomena are collectively referred to as Technical Debt (TD).

[0003] Experienced developers often proactively mark technical debt in code comments; this type of technical debt is called self-admitted technical debt (SATD). Compared to unintentionally introduced technical debt, SATD is more observable and traceable because it is text-based. Research shows that SATD is highly prevalent in open-source software projects, sometimes covering up to 30% of the files. Therefore, accurately identifying and promptly cleaning up SATD is considered an important means of improving software quality.

[0004] To achieve efficient SATD recognition, relevant SATD recognition methods can be broadly categorized into supervised and unsupervised methods. Unsupervised methods match annotations using manually summarized SATD patterns (specific words and phrases); supervised methods train models on labeled datasets, enabling them to learn the semantic relationships between annotations and technical debt. Although supervised methods can model more complex semantic information and their recognition performance is generally superior to unsupervised methods, they still have the following two shortcomings in practical applications:

[0005] First, computational efficiency is insufficient. Supervised methods require complex semantic modeling of explicit SATD annotations, which can actually be efficiently identified through simple string matching. On the other hand, many normal annotations have simple semantics, and the benefits of supervised methods in performing complex modeling on them are very limited, while also incurring additional computational overhead.

[0006] Second, the use of context is insufficient. Supervised methods still do not fully utilize the contextual information of annotations, leading to a large number of false positives in the recognition results. In some scenarios, the model misclassifies annotations as requirements to be completed, when in fact the annotations describe functions that have already been implemented. Although a few studies have introduced contextual information to assist recognition, these studies crudely treat the next line of code after the annotation as context. In complex cases, such as when the context is not adjacent or contains multiple lines, it is easy to extract incorrect or semantically incomplete context, thus resulting in misrecognition.

[0007] Furthermore, both supervised and unsupervised methods suffer from insufficient interpretability. While unsupervised methods can output recognition results and matched patterns, for implicit SATDs, the development team still needs to manually review annotations and locate the pattern to ultimately determine whether it constitutes an SATD. On the other hand, supervised methods based on deep learning typically only output classification results and probability scores, lacking a basis for judgment and making it difficult to trace the recognition mechanism, thus leading to insufficient trust in the results by the development team.

[0008] In summary, the existing SATD identification methods suffer from excessive computational overhead when processing explicit samples, insufficient context awareness leading to a large number of false positives, and a lack of interpretable identification criteria, which restricts the application of SATD identification technology in practical software engineering scenarios. Summary of the Invention

[0009] In view of this, this application proposes a context-aware, self-acknowledging technical debt identification method, apparatus, and electronic device. Specifically, this application is implemented through the following technical solution:

[0010] According to a first aspect of the embodiments of this specification, a context-aware self-acknowledgment technology debt identification method is provided, comprising:

[0011] Step S1: Based on the preset self-acknowledgment technical debt (SATD) high probability pattern and preset SATD keywords, classify the code comments in the target code file to obtain the category of each code comment in the target code file. The category includes SATD comments, non-SATD comments, and comments to be reasoned.

[0012] Step S2: For each annotation to be reasoned, determine the code segment that is semantically related to the annotation to be reasoned and use it as the context of the annotation to be reasoned;

[0013] Step S3: Input the annotation to be reasoned and its context into the pre-trained large language model, and guide the large language model to perform semantic reasoning and attribution analysis through structured prompts to generate the response result of the annotation to be reasoned.

[0014] Step S4: Standardize the response results output by the large language model to output a structured recognition result containing classification labels and reasoning basis.

[0015] According to a second aspect of the embodiments of this specification, a context-aware self-acknowledgment technology debt identification device is provided, comprising:

[0016] The annotation classification unit is used to classify code annotations in the target code file based on a preset high-probability pattern of self-acknowledgment technical debt (SATD) and preset SATD keywords, and to obtain the category of each code annotation in the target code file. The category includes SATD annotations, non-SATD annotations, and annotations to be reasoned.

[0017] The context extraction unit is used to determine the code segment that is semantically related to each annotation to be inferred and use it as the context of the annotation to be inferred.

[0018] The annotation reasoning unit is used to input the annotation to be reasoned and its context into a pre-trained large language model, and guide the large language model to perform semantic reasoning and attribution analysis through structured prompt words, and generate the response result of the annotation to be reasoned.

[0019] The result parsing unit is used to standardize the response results output by the large language model and output a structured recognition result containing classification labels and reasoning basis.

[0020] According to a third aspect of the embodiments of this specification, an electronic device is provided, comprising: a processor; and a computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method as described in the first aspect.

[0021] The embodiments of this application have at least the following technical effects:

[0022] First, an efficient pre-screening mechanism is introduced. Through the annotation-based rapid screening strategy, a considerable number of annotation samples can be directly identified without the need for large language model inference on these samples, which greatly reduces inference overhead and significantly improves recognition efficiency without compromising detection performance.

[0023] Second, it can accurately perceive the context. The embodiments of this application can accurately locate code segments related to the annotation semantics as the context. Compared with the mechanical extraction of adjacent code lines, the embodiments of this application can effectively reduce false positives and significantly improve the accuracy of context extraction.

[0024] Third, it provides interpretable reasoning and attribution. By using structured prompts to guide the large language model to generate judgment results and reasoning basis, developers can understand the judgment logic of the recognition results, increase their trust in the recognition results, and provide good interpretability while having better recognition performance than the comparison methods.

[0025] Fourth, it requires no training and has strong generalization ability. The embodiments of this application do not require additional supervised training of large language models, do not depend on specific dataset distributions, have stronger adaptability to changes in data distribution, and can be generalized and applied across different software projects and technical fields. Attached Figure Description

[0026] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Some specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings indicate the same or similar parts or components. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0027] Figure 1 This is a schematic diagram illustrating the overall process of self-acknowledgment technology debt identification in an exemplary embodiment of this application;

[0028] Figure 2 This is a schematic flowchart illustrating an exemplary embodiment of the present application of a context-aware self-acknowledgment technology debt identification method;

[0029] Figure 3 This is a schematic diagram illustrating the prompt words used in a context-aware strategy according to an exemplary embodiment of this application;

[0030] Figure 4 This is a schematic diagram of structured cue words used in the reasoning and attribution stage, as illustrated in an exemplary embodiment of this application;

[0031] Figure 5 This is a structural block diagram of an electronic device illustrated in an exemplary embodiment of this application;

[0032] Figure 6 This is a block diagram illustrating a context-aware self-acknowledgment technology debt identification device according to an exemplary embodiment of this application. Detailed Implementation

[0033] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0034] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0035] First, a brief introduction to the terms used in the embodiments of this application:

[0036] Self-Acknowledged Technical Debt (SATD): This refers to technical debt that developers actively mark in code comments. In other words, developers acknowledge that there are technical issues in the code that need to be improved, fixed, or perfected, including temporary solutions, known defects, incomplete implementations, etc.

[0037] Large Language Model (LLM): A deep learning model based on the Transformer architecture and pre-trained on large-scale text corpora, possessing powerful natural language understanding, semantic reasoning, and instruction following capabilities. In the embodiments of this application, the LLM is used for two subtasks: context localization and extraction, and SATD inference decision-making.

[0038] SATD-Probable Patterns: These are specific word combinations or phrase patterns that, when appearing in annotated text, indicate a high probability that the annotation belongs to SATD. They are discovered by existing research through data mining methods.

[0039] Comment context: refers to the code snippets in a code file that are highly semantically related to the comments, used to help determine whether the technical state described by the comments has been implemented.

[0040] The embodiments described in this specification will now be described in detail.

[0041] This application provides a context-aware, self-acknowledgment-based debt identification method. Figure 1 This is a schematic diagram illustrating the overall process of self-acknowledgment technology debt identification in an exemplary embodiment of this application. Figure 2This is a schematic flowchart illustrating an exemplary embodiment of a context-aware self-acknowledgment technology debt identification method, as shown in the following example. Figure 1 and Figure 2 As shown, the self-acknowledgment technology debt identification method includes the following steps:

[0042] Step S1: Based on the preset self-acknowledgment technical debt (SATD) high probability pattern and preset SATD keywords, classify the code comments in the target code file to obtain the category of each code comment in the target code file. The category includes SATD comments, non-SATD comments, and comments to be reasoned.

[0043] Here, SATD keywords refer to predefined words or symbols that, when appearing in code comments, can be identified as SATD comments without needing to consider the context. For example, the SATD keywords include TODO, FIXME, HACK, XXX, and @deprecated.

[0044] SATD high-probability patterns refer to word combinations or phrases discovered through statistical analysis from known SATD annotations, which have a higher probability of belonging to SATD code when the pattern appears in code annotations than in non-SATD annotations. For example, SATD high-probability patterns include Abandon, Access, Actually, Add, Aggressive, Allocat, Ambiguous, Angry, Annoyed, Anxious, etc.

[0045] It should be noted that in this embodiment, there are no identical words or symbols between the SATD keywords and the SATD high-probability patterns.

[0046] Step S2: For each annotation to be inferred, determine the code segment that is semantically related to the annotation to be inferred and use it as the context of the annotation to be inferred.

[0047] Step S3: Input the annotation to be reasoned and its context into the pre-trained large language model, and guide the large language model to perform semantic reasoning and attribution analysis through structured prompts to generate the response result of the annotation to be reasoned.

[0048] Step S4: Standardize the response results output by the large language model to output a structured recognition result containing classification labels and reasoning basis.

[0049] Through the above steps S1 to S4, the method provided in this embodiment achieves the following technical effects:

[0050] First, in step S1, code comments are quickly filtered based on preset SATD high-probability patterns and SATD keywords. This allows a large number of comments containing explicit keywords or without any high-probability patterns to be directly identified as SATD comments or non-SATD comments without calling a large language model, thereby significantly reducing the overall computational overhead and significantly improving recognition efficiency.

[0051] Secondly, step S2 uses code segments related to the semantics of the annotation to be reasoned as context, avoiding the problem of missing context or incomplete semantics caused by mechanically extracting adjacent code lines in the prior art. This provides accurate semantic basis for subsequent reasoning and effectively reduces false positives.

[0052] Furthermore, step S3 utilizes a pre-trained large language model combined with structured cue words for semantic reasoning and attribution analysis. This not only captures implicit semantic cues in the annotations but also generates interpretable reasoning while outputting classification labels, enabling developers to understand and trust the recognition results. This overcomes the poor interpretability of traditional black-box models.

[0053] Finally, step S4 standardizes the response results and outputs a structured recognition result containing classification labels and reasoning basis, ensuring the consistency and readability of the recognition results and facilitating subsequent automated processing or manual review.

[0054] In summary, Figure 2 The self-acknowledgment technical debt identification method shown achieves efficient and interpretable self-acknowledgment technical debt identification while maintaining high identification accuracy, and does not require additional model training, and has good cross-project generalization ability.

[0055] In some embodiments, step S1, which categorizes the code comments in the target code file, includes:

[0056] If a code comment contains at least one of the SATD keywords, the code comment is determined to be an SATD comment; if the code comment does not contain any of the SATD keywords, it is determined whether the code comment matches the SATD high-probability pattern. If the code comment does not match any of the SATD high-probability patterns, the code comment is determined to be a non-SATD comment; otherwise, the code comment is determined to be a comment to be inferred.

[0057] This step uses a pre-defined set of SATD keywords and a set of SATD high-probability patterns to quickly filter code comments in the target code file, directly identifying some comments without calling a large language model.

[0058] Specifically, the SATD (Self-Deprecated and Targeted) comment determination is performed first: Each code comment is traversed, and if a code comment contains any one of the keywords TODO, FIXME, HACK, XXX, or @deprecated, then the comment is directly determined as SATD and output. These five keywords are highly indicative of SATD comments and can be directly used as sufficient conditions for SATD determination.

[0059] Non-SATD determination: For code comments that do not match the above keywords, further check whether they match any pattern in the SATD high-probability pattern set. If a comment does not match any pattern, it is highly likely that it is a non-SATD comment. In this embodiment, it is directly determined as a non-SATD comment and output without further semantic analysis.

[0060] It is worth noting that when filtering code comments based on the SATD keyword set and the SATD high probability pattern set in this embodiment, the characters should be uniformly converted in format, for example, all of them should be converted to lowercase letters before matching.

[0061] In this embodiment, the SATD high-probability pattern set is obtained by summarizing and filtering relevant patterns discovered in existing research. For example, it includes patterns such as Abandon, Access, Actually, and Add to cover typical word combinations and phrase expressions representing technical debt.

[0062] Determination of comments to be inferred: For code comments that are not covered by the above two determination conditions, it indicates that the code comments contain a high probability pattern of SATD but do not contain explicit keywords. Their SATD attributes have semantic uncertainty and need to be comprehensively inferred and determined in combination with the code context. That is, in this embodiment, the remaining part of the code comments is marked as comments to be inferred and proceeds to the subsequent step S2 processing.

[0063] This embodiment filters and classifies code comments based on preset filtering rules, which can directly determine that a considerable number of samples do not need to enter the large language model inference stage, including code comments containing SATD keywords and code comments that do not contain any SATD high probability patterns, thereby greatly reducing the inference time of the subsequent large language model and without introducing additional false negative or false positive risks.

[0064] In some embodiments, step S2 determines a code segment that is semantically related to the annotation to be inferred and uses it as the context of the annotation to be inferred, including:

[0065] The annotation to be reasoned and its corresponding method-level or statement-block-level code snippet are input into a pre-trained large language model. Context-aware prompts guide the large language model to locate the code snippet that is semantically related to the annotation to be reasoned.

[0066] This step provides the annotation to be inferred and its corresponding complete code snippet to the large language model. Through context-aware prompts, the large language model is guided to locate and extract the code snippet that is highly related to the semantics of the annotation as the context.

[0067] In practical applications, the method-level code snippets or statement-block-level code snippets are preprocessed to filter out non-English comments, license or copyright notice comments, semantically incomplete comments with a string length of less than 10, and non-natural language comments such as commented-out code.

[0068] In this embodiment, the context-aware prompt words are as follows: Figure 3 As shown, this includes the positional relationships between comments and code, which include adjacency, internal, and separation relationships; where:

[0069] The adjacency relationship refers to a comment being located on the line preceding, on the same line as, or following a code segment that is semantically related to it. This type of relationship is the most common and includes three specific scenarios. The context-aware strategy in this embodiment extracts the complete associated code block rather than just the next line of code to maintain semantic integrity.

[0070] The term "internal relationship" refers to a comment located within a code block, where the code segments preceding and following the comment together constitute its context. For this type of relationship, the context-aware strategy in this embodiment extracts the complete code block containing the comment, including the related code before and after it, rather than just extracting the single line of code following the comment.

[0071] The separated relationship refers to a situation where a comment and its semantically related code snippet are not adjacent in text position, and there is one or more lines of unrelated code between them. For this type of relationship, the context-aware strategy in this embodiment uses semantic understanding to navigate across unrelated code and accurately locate the target code snippet as the context.

[0072] Compared to mechanically using the next line of code after a code comment as context, this step can provide effective context for all samples, accurately capturing the real semantic relationship between comments and code in three scenarios: adjacency, internal, and separation, effectively avoiding problems such as incorrect context extraction or incomplete semantics.

[0073] In some embodiments, in step S3, an open-source large language model with 20B parameters is selected as the inference model. This model adopts a hybrid expert architecture and quantizes the weights, thereby significantly reducing the memory usage and enabling it to run efficiently in a single-card environment with 16GB of memory, balancing inference performance and deployment cost.

[0074] This embodiment does not require additional supervised training of the large language model and does not depend on a specific dataset distribution. Users can flexibly adjust the SATD decision boundary by adjusting the relevant constraints and decision criteria in the structured prompt words according to specific application needs, so as to adapt to different application scenarios.

[0075] like Figure 4 As shown, the structured prompts in step S3 include:

[0076] SATD definition constraints are used to provide a semantic benchmark for SATD in large language models. That is, the conceptual scope of SATD annotations is clarified through SATD definition constraints, including temporary solutions or expedient implementations (such as "TODO:fix later"), incomplete or delayed implementations (such as "FIXME: this is a workaround"), known defects or limitations, missing or delayed documents and tests, etc., so as to provide a unified semantic benchmark for large language models.

[0077] Feature indications are used to guide large language models to directly identify based on predefined explicit signals, and to reason based on predefined implicit semantic cues when explicit signals are lacking.

[0078] It is worth noting that the SATD keywords used in step S1 of this embodiment refer to a precise and fixed set of strings used for fast, context-free rule matching of code comments. For example, the keyword set includes TODO, FIXME, HACK, etc. When these strings appear in code comments, the system directly identifies them as SATD comments. This matching process is a simple normalization based on preset rules. Its advantage is that it has extremely high computational efficiency, but its disadvantage is that it cannot recognize SATD comments caused by spelling errors, such as TODO where the letter O is misspelled as the number 0, FIXME where an extra space is written as FIX ME, or minor variations.

[0079] In contrast, in the structured prompts of step S3, explicit signals are a broader and more robust semantic-level concept. Explicit signals include not only the aforementioned SATD keyword, but also common spelling variations, case-insensitive forms, abbreviations, separator variations, and approximate forms due to typos. For example, for the keyword TODO, its corresponding explicit signals could cover variations such as TOD0 (the number 0 replaces the letter O), TO-DO, and TO DO: (with a space in between). The purpose of explicit signals is to enable the large language model, with its semantic understanding capabilities, to generalize and recognize code comments that, due to human error or formatting differences, were not captured by the exact matching rules of step S1, but which actually have explicit indicative meaning.

[0080] In short, the SATD keyword in step S1 is a set of rules for exact matching, and the explicit signal in step S3 is a semantic category description that includes the set of rules and their common variations and typos.

[0081] In the structured prompts of step S3, implicit semantic cues include negative evaluations, such as "bad implementation" and "not ideal"; semantic uncertainty, such as "might cause issues"; and expressions of future obligation, such as "should be refactored". Implicit semantic cues help large language models to reason based on tone and semantic cues even in the absence of explicit signals.

[0082] The context consistency determination rule is used to constrain the large language model to determine whether an annotation belongs to SATD annotation based on the semantic consistency between the annotation and its context.

[0083] To avoid misjudgments based solely on annotation semantics, this embodiment introduces contextual consistency constraints. If the context has already implemented the behavior indicated by the annotation, the code annotation is determined to be a non-SATD annotation; if the context is missing, weakened, or contradicts the annotation description, the code annotation is determined to be a SATD annotation; if the evidence is insufficient or the semantics are ambiguous, it is determined to be a non-SATD annotation by default, in order to control the conservatism of the model and reduce the risk of false positives.

[0084] The decision boundary criterion limits the judgment result to two types: acceptable design choice or defect that should be improved. It instructs the large language model to classify an annotation as a SATD annotation only if the annotation belongs to the defect that should be improved type. This criterion abstracts the decision into an intentionally acceptable design choice or a known defect or temporary implementation that should be improved. If the annotation belongs to an intentionally acceptable design choice, it is judged as a non-SATD annotation; if the annotation belongs to a known defect or temporary implementation that should be improved, it is judged as a SATD annotation, to ensure overall accuracy.

[0085] Output format constraints are used to instruct the large language model to output classification labels and reasoning rationale. For example, the output format is: {"label":"SATD|NON-SATD", "reason":"brief reasoning explanation"}.

[0086] This embodiment, through the above-described structured design, enables the large language model to complete semantic-level attribution analysis under explicit rule constraints, thereby providing more stable, reliable, and interpretable judgment results.

[0087] In some embodiments, the standardization process for the response output of the large language model in step S4 includes:

[0088] The response results returned by the large language model are subjected to structural integrity verification; the response results that pass the structural integrity verification are deserialized and parsed to extract classification labels and reasoning basis; when the response result has structural abnormalities or deserialization fails, a retry mechanism is triggered, and after the number of retries reaches a preset threshold, a default classification result is returned and the abnormal information is recorded.

[0089] For example, response validity verification refers to performing a structural integrity check on the results returned by the large language model interface to confirm that the response contains valid candidate output fields, such as determining whether it contains label and reason fields; if the structure is abnormal, such as interface timeout or empty response, it is determined to be an invalid response and triggers the retry mechanism.

[0090] Structured result parsing refers to deserializing the JSON-formatted text generated by the large language model, extracting the classification result (label), brief reason, and reasoning from the JSON string, and forming a standardized structured output.

[0091] Exception handling and retries refer to the automatic triggering of a retry mechanism if issues such as response errors, structure mismatches, or JSON parsing failures occur during the parsing process. To improve stability, a maximum of three retries are performed. If parsing still fails, a default result (without the SATD comment) is returned, and the exception information is logged to ensure the robustness of the method. Response errors here include interface errors, structure mismatches include missing necessary fields, and JSON parsing failures include invalid formats.

[0092] This embodiment, through the above steps, can output structured recognition results for any code comment. The recognition results include classification labels, brief reasoning basis, and optional complete reasoning process, thereby achieving interpretable SATD recognition.

[0093] Next, the self-acknowledgment technology debt identification scheme of this application will be explained in detail through different scenarios.

[0094] (1) For the rapid identification scenario of displaying SATD annotations.

[0095] The annotation to be identified is: " / / TODO: refactor this method to improve performance".

[0096] Based on the SATD keyword "TODO" from step S1, the annotation to be identified can be directly determined as a SATD annotation and output without proceeding to subsequent steps, with inference time approaching zero.

[0097] (2) Context recognition scenario for implicit SATD annotation.

[0098] The comment to be identified is: " / / catch the exception to prevent executor from shutting down, log it instead". The code snippet corresponding to this comment is a Catch block, in which the code only logs the exception and does not handle the exception.

[0099] Based on the high-probability SATD pattern identified in step S1, the annotation to be identified contains "Exception". Therefore, this annotation is considered an annotation to be inferred. Furthermore, step S2 identifies the annotation as having an adjacency relationship, and the `Catch` code block is extracted as context. The annotation and its context are then input into the Big Prophet model. The Big Language model parses the annotation and determines that it indicates a bug in the code and avoids proper error handling by logging and continuing execution, which is a typical expedient measure. Therefore, it is classified as a SATD annotation, and inference criteria are generated. These criteria could be that "the annotation explicitly acknowledges a temporary error handling solution, and the context code confirms the existence of this solution."

[0100] (3) Context recognition scenarios for non-SATD annotations.

[0101] The comment to be identified is: " / / fix for issue: replace commas and colons with similar Unicode characters". The code snippet described by this comment contains character replacement logic.

[0102] Based on the SATD high-probability pattern in step S1, it can be seen that the annotation to be identified contains words such as "Fix". The annotation to be identified belongs to the annotation to be reasoned. Based on step S2, its context is extracted and the annotation to be identified and its context are input into the big oracle model. The big language model determines that the annotation to be identified is only a description of the repair solution. The context code fully implements the repair logic and does not imply that it is temporary, incomplete or needs to be improved in the future. It is determined to be a non-SATD annotation, effectively avoiding the wrong over-association of the word "fix".

[0103] Based on the above experimental verification, it can be seen that the embodiments of this application outperform the representative comparative methods in multiple evaluation metrics such as precision, recall, and F1 score. Furthermore, the embodiments of this application do not require supervised training of a large language model and exhibit strong generalization ability across different data distributions.

[0104] Figure 5This is a schematic diagram of an electronic device illustrated in this specification according to an exemplary embodiment. Please refer to... Figure 5 At the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, memory 508, a hardware acceleration device 510, and non-volatile memory 512, and may also include other hardware required for its functions. One or more embodiments of this application can be implemented in software, for example, the processor 502 reads the corresponding computer program from the non-volatile memory 512 into memory 508 and then runs it. Of course, in addition to software implementation, one or more embodiments of this application do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the above processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0105] Figure 6 This is a structural block diagram illustrating a context-aware self-acknowledgment technology debt identification device according to an exemplary embodiment of this application. The self-acknowledgment technology debt identification device can be applied to, for example... Figure 6 The electronic device shown implements the technical solution of this application. The self-acknowledging technical debt identification device includes: an annotation classification unit 610, a context extraction unit 620, an annotation reasoning unit 630, and a result parsing unit 640, wherein:

[0106] The annotation classification unit 610 is used to classify code annotations in the target code file based on a preset high-probability pattern of self-acknowledgment technical debt (SATD) and preset SATD keywords, so as to obtain the category of each code annotation in the target code file. The category includes SATD annotations, non-SATD annotations and annotations to be reasoned.

[0107] The context extraction unit 620 is used to determine, for each annotation to be inferred, a code segment that is semantically related to the annotation to be inferred and use it as the context of the annotation to be inferred;

[0108] The annotation reasoning unit 630 is used to input the annotation to be reasoned and its context into a pre-trained large language model, and guide the large language model to perform semantic reasoning and attribution analysis through structured prompt words to generate the response result of the annotation to be reasoned.

[0109] The result parsing unit 640 is used to standardize the response results output by the large language model and output a structured recognition result containing classification labels and reasoning basis.

[0110] In some embodiments, the annotation classification unit 610 is configured to: if a code annotation contains at least one of the SATD keywords, then determine that the code annotation is an SATD annotation; if the code annotation does not contain any of the SATD keywords, then determine whether the code annotation matches the SATD high-probability pattern; if the code annotation does not match any of the SATD high-probability patterns, then determine that the code annotation is a non-SATD annotation; otherwise, determine that the code annotation is an annotation to be inferred.

[0111] In some embodiments, the context extraction unit 620 is used to input the annotation to be reasoned and the method-level code segment or statement block-level code segment in which it is located into a pre-trained large language model, and guide the large language model to locate the code segment that is semantically related to the annotation to be reasoned from the code segment through context-aware prompts.

[0112] In some embodiments, the result parsing unit 640 is used to perform structural integrity verification on the response results returned by the large language model; and to perform deserialization parsing on the response results that pass the structural integrity verification to extract classification labels and reasoning basis.

[0113] In some embodiments, the result parsing unit 640 is further configured to trigger a retry mechanism when the response result structure is abnormal or deserialization parsing fails, and return the default classification result and record the abnormal information after the number of retries reaches a preset threshold.

[0114] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0115] Accordingly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the above embodiments.

[0116] Accordingly, embodiments of this application also provide a computer program product configured to perform the methods described in any of the above embodiments.

[0117] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0118] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0119] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0120] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0121] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0122] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0123] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0124] It should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0125] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A context-aware based self-admission technical debt identification method, characterized in that, Includes the following steps: Step S1: Based on the preset self-acknowledgment technical debt (SATD) high probability pattern and preset SATD keywords, classify the code comments in the target code file to obtain the category of each code comment in the target code file. The category includes SATD comments, non-SATD comments, and comments to be reasoned. Step S2: For each annotation to be reasoned, determine the code segment that is semantically related to the annotation to be reasoned and use it as the context of the annotation to be reasoned; Step S3: Input the annotation to be reasoned and its context into the pre-trained large language model, and guide the large language model to perform semantic reasoning and attribution analysis through structured prompts to generate the response result of the annotation to be reasoned. Step S4: Standardize the response results output by the large language model to output a structured recognition result containing classification labels and reasoning basis; Step S1 involves classifying the code comments in the target code file, including: If a code comment contains at least one of the aforementioned SATD keywords, then the code comment is determined to be a SATD comment; If the code comment does not contain any of the SATD keywords, then it is determined whether the code comment matches the SATD high-probability pattern. If the code comment does not match any of the SATD high-probability patterns, the code comment is determined to be a non-SATD comment; otherwise, the code comment is determined to be a comment to be inferred. Step S2, which determines the code segment related to the semantics of the annotation to be inferred and uses it as the context of the annotation, includes: The annotation to be reasoned and its corresponding method-level or statement-block-level code snippet are input into a pre-trained large language model. Context-aware prompts guide the large language model to locate the code snippet that is semantically related to the annotation to be reasoned from the code snippet. The context-aware prompts contain the positional relationship between the annotation and the code, which includes adjacency, internal, and separation relationships.

2. The method according to claim 1, characterized in that: The SATD keyword in step S1 refers to a predefined word or symbol in a code comment that can be determined as a SATD comment without considering the context when the keyword appears in the code comment. The high-probability SATD pattern in step S1 refers to a word combination or phrase that is discovered through statistical analysis from known SATD annotations, and whose probability of belonging to SATD code is higher than that of non-SATD annotations when the pattern appears in the code annotation. Furthermore, there are no identical words or symbols between the SATD keywords and the SATD high-probability patterns.

3. The method according to claim 1, characterized in that: The adjacency relationship refers to the fact that a comment is located on the line before, on the same line as, or on the line after a code segment that is semantically related to it. The internal relationship refers to the fact that a comment is located inside a code block, and the code segments before and after the comment together constitute its context; The separation relationship refers to the fact that the annotation and its semantically related code segment are not adjacent in text position, and there is one or more lines of unrelated code between them.

4. The method of claim 1, wherein, The structured prompts in step S3 include: SATD defines constraints to provide a semantic benchmark for large language models. Feature indications are used to guide large language models to directly identify based on predefined explicit signals, and to reason based on predefined implicit semantic cues when explicit signals are lacking. The context consistency determination rule is used to constrain the large language model to determine whether an annotation belongs to SATD annotation based on the semantic consistency between the annotation and its context. The decision boundary criterion is used to limit the decision result to two types: acceptable design choice or defect that should be improved, and instructs the large language model to determine the annotation as an SATD annotation only if the annotation belongs to the defect that should be improved type. Output format constraints are used to instruct large language models to output classification labels and reasoning bases.

5. The method of claim 1, wherein, The standardization process for the response output of the large language model in step S4 includes: Perform structural integrity verification on the response results returned by the large language model; The response results that pass the structural integrity check are deserialized and parsed to extract classification labels and reasoning basis.

6. The method according to claim 5, characterized in that, The standardization process for the response output of the large language model in step S4 further includes: When the response result is structurally abnormal or deserialization fails, a retry mechanism is triggered, and after the number of retries reaches a preset threshold, the default classification result is returned and the abnormal information is recorded.

7. A context-aware, self-acknowledgment-based debt identification device, characterized in that, include: The annotation classification unit is used to classify code annotations in a target code file based on a preset high-probability pattern of self-acknowledged technical debt (SATD) and preset SATD keywords, obtaining a category for each code annotation in the target code file. The category includes SATD annotations, non-SATD annotations, and annotations to be inferred. Specifically, if a code annotation contains at least one of the SATD keywords, the code annotation is determined to be a SATD annotation; if the code annotation does not contain any of the SATD keywords, it is determined whether the code annotation matches the high-probability pattern of SATD. When the code annotation does not match any of the high-probability patterns of SATD, the code annotation is determined to be a non-SATD annotation; otherwise, the code annotation is determined to be an annotation to be inferred. The context extraction unit is used to determine the code segment that is semantically related to each annotation to be inferred and use it as the context of the annotation to be inferred. Specifically, the annotation to be reasoned and its corresponding method-level or statement-block-level code segment are input into a pre-trained large language model. Context-aware prompts guide the large language model to locate the code segment semantically related to the annotation to be reasoned from the code segment. The context-aware prompts include the positional relationship between the annotation and the code, which includes adjacency, internal, and separation relationships. The annotation reasoning unit is used to input the annotation to be reasoned and its context into a pre-trained large language model, and guide the large language model to perform semantic reasoning and attribution analysis through structured prompt words, and generate the response result of the annotation to be reasoned. The result parsing unit is used to standardize the response results output by the large language model and output a structured recognition result containing classification labels and reasoning basis.

8. An electronic device, characterized in that, include: processor; as well as, A computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Self-acceptance technology debt detection method and device based on large and small model fusion

    CN117873558A

  • SATD identification method based on query strategy combining representativeness and uncertainty

    CN120180257A