Operable code smell fusion detection method and system based on large language model and multi-index call graph

CN121764768BActive Publication Date: 2026-09-25HANGZHOU DIANZI UNIVERSITY BINJIANG INSTITUTE CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511854251.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-09-25
Estimated Expiration
2045-12-10

AI Technical Summary

Technical Problem

[0007]为解决现有代码异味检测方法在处理可操作代码异味上存在的异味标签生成策略未考虑历史演化验证,代码语义表达不够充分,以及代码特征上下文建模单一问题,本发明提出了一种基于大语言模型和多指标调用图的可操作代码异味融合检测方法及系统

Benefits of technology

[0018]本发明的有益效果:本发明通过使用大语言模型的特征提取能力,在数据集上进行微调后充当代码语义提取器,设计了一个融合调用图将静态代码指标与代码调用图结合起来,通过图神经网络提取代码结构信息,通过代码结构和语义渐进式融合得到充分的特征建模,提高了可操作代码异味的检测能力,实现了向可操作异味检测范式的转变,提升了在真实开发场景中的检测效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121764768B_ABST
    Figure CN121764768B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on big language model and multi-index calling graph operable code smell fusion detection method and system.The application first collects potential code smell in open source project, judges its operability according to code smell evolution history, simultaneously constructs code calling graph, and constructs code smell data set;Second, fine-tune big language model, extract intermediate hidden layer state and generate code semantic vector;Then static code metric index is embedded in the node of code calling graph to form multi-index calling graph, and the graph neural network with attention mechanism is used to represent code structure vector;Finally, code semantic vector and structure vector are gradually fused, and whether the code contains operable code smell is detected by multi-layer classification network.The application uses fine-tuned big language model to obtain code semantic features, involves multi-index calling graph to obtain code structure features, and further improves the detection ability of operable code smell by representing code features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of code smell detection in software engineering, specifically to an operable code smell fusion detection method and system based on a large language model and a multi-index call graph. Background Technology

[0002] As modern software systems continue to grow in size and complexity, software developers, constrained by factors such as experience, timeliness, and awareness, are not always able to write code that meets software engineering requirements during the development process. This poorly designed code is known in the software engineering field as code smells. While code smells do not directly lead to software functional errors, they can cause maintenance difficulties, reduce code quality, and increase technical debt.

[0003] Code smells, as a concept reflecting potential design problems and maintainability risks, have received widespread attention in software engineering research and practice. In recent years, due to the popularization of continuous integration and rapid iteration concepts, effectively detecting, managing, and fixing code smells has become a crucial aspect of ensuring software quality.

[0004] Code smell detection, as a prerequisite for smell management, has been extensively studied. However, existing research generally treats all detected code smells as equally important, failing to adequately distinguish which smells will ultimately be fixed in subsequent version iterations. To address this issue, this application uses a new research concept—actionable code smells. They are defined as code smells that are explicitly removed or improved through manual or automated refactoring by developers in the subsequent evolution of software versions. Unlike traditional code smells that remain only at the level of "potential problems," actionable code smells have historical evidence of actual fixes, thus possessing higher reference value in prioritization, technical debt management, and refactoring suggestion generation.

[0005] Existing code smell detection methods can be broadly categorized into three types: rule-based and heuristic methods, metric and statistical methods, and machine learning and deep learning methods. However, regardless of the type, all methods suffer from the following shortcomings when dealing with such "operability" problems.

[0006] First, previous code smell detection methods lacked historical evolution verification during the dataset construction phase, failing to utilize version control evolution data to verify whether a particular smell would eventually be fixed. Second, existing code smell detection methods are insufficient in terms of code semantic representation, and research on code semantic representation through large models remains a vacuum in the field of code smell detection. Furthermore, current code smell detection methods model code features too simply, failing to deeply integrate semantic and structural information, resulting in insufficient information representation. Summary of the Invention

[0007] To address the shortcomings of existing code smell detection methods, such as the lack of historical evolution verification in smell tag generation strategies, insufficient code semantic expression, and limited code feature context modeling, this invention proposes a method and system for the fusion detection of operational code smells based on a large language model and a multi-index call graph.

[0008] In a first aspect, the present invention provides an operable code smell fusion detection method based on a large language model and a multi-index call graph, comprising the following steps:

[0009] The code smell detection tool scans the historical version information of open source projects to identify code smells in each historical version information. Based on the evolution history of the code smell, its operability is judged. At the same time, the source code of the project is compiled and built to generate a code call graph. Using the open source project and historical version information as the organizational unit, the smell instances are associated with multi-source information to form an operable code smell dataset.

[0010] Based on the code smell dataset, fine-tune the large language model, extract the intermediate hidden layer states, and generate code semantic vectors as the semantic features of the code to be tested by weighted averaging.

[0011] Traverse the code call graph and embed static code metrics into the nodes of the code call graph according to the method name or class name of the node in the graph to form a multi-metric call graph. At the same time, introduce weight information at the edge level to enhance the call strength and dependency relationship, and learn the multi-metric call graph to represent the code structure features through graph neural network.

[0012] A progressive fusion strategy is adopted to fuse the semantic features and structural features of the code. The fused features are processed through a multi-layer classification network, and a discriminant function is used to determine whether the code unit belongs to the operable code smell.

[0013] Secondly, this invention provides an operable code smell fusion detection system based on a large language model and multi-index call graph, comprising:

[0014] The code smell detection module is used to scan the historical version information of open source projects with code smell detection tools to identify code smells in the historical version information. Based on the evolution history of code smells, its operability is judged. At the same time, the project source code is compiled and built to generate a code call graph. Using open source projects and historical version information as organizational units, smell instances are associated with multi-source information to form an operable code smell dataset.

[0015] The code semantic feature construction module is used to fine-tune the large language model based on the code smell dataset, extract the intermediate hidden layer states, and generate a code semantic vector as the code semantic feature to be tested by weighted averaging.

[0016] The code structure feature construction module is used to traverse the code call graph, embed static code metrics into the nodes of the code call graph according to the method name or class name of the node in the graph, form a multi-metric call graph, introduce weight information at the edge level to enhance call strength and dependency relationship, and learn the multi-metric call graph to represent code structure features through graph neural network.

[0017] The feature fusion module is used to fuse the semantic features and structural features of the code using a progressive fusion strategy, process the fused features through a multi-layer classification network, and use a discriminant function to determine whether a code unit belongs to an operable code smell.

[0018] The beneficial effects of this invention are as follows: By using the feature extraction capabilities of a large language model and fine-tuning it on a dataset, this invention acts as a code semantic extractor. It designs a fusion call graph that combines static code metrics with the code call graph, extracts code structure information through a graph neural network, and obtains sufficient feature modeling through the progressive fusion of code structure and semantics. This improves the detection capability of operable code smells, realizes the transformation to an operable code smell detection paradigm, and enhances the detection efficiency in real development scenarios. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 The flowchart of an operable code odor fusion detection method based on a large language model and multi-index call graph provided by the present invention is shown. Detailed Implementation

[0021] To more clearly describe the technical solution proposed in this invention, the invention will be further described in detail below with reference to the accompanying drawings and specific examples.

[0022] Conversely, this invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the invention as defined in the claims. Furthermore, to provide a better understanding of the invention, certain specific details are described in detail below. However, those skilled in the art will fully understand the invention even without these detailed descriptions.

[0023] like Figure 1As shown, this application provides an operable code odor fusion detection method based on a large language model and multi-index call graph, which includes the following steps:

[0024] S1. A labeled, operable code odor detection dataset is obtained through data collection and processing, including the basic attributes of the odor (location, type, etc.), its operability labels (generated based on version evolution and auxiliary verification), static indicators carried by nodes and edges (from a multi-indicator call graph), and basic information such as project and version number:

[0025] Optionally, the operable code odor includes key-value pairs of relevant code and annotation results, defined as follows:

[0026]

[0027]

[0028] in Representative version All the code smells present in it, Represents a specific version of the project. The first one that exists in A code smell, Represents version number greater than Subsequent versions of the project The function is used to determine whether the currently selected code smell will appear in subsequent versions, and It is the collection of all operable code smells that exist in a specific version of a project.

[0029] Optionally, the collection and construction of the operable code odor dataset specifically includes the following steps:

[0030] S11: Select a batch of open-source software projects from the code hosting platform as research objects. To ensure the representativeness and validity of the data, the project selection followed these criteria: a. Written in Java to ensure compatibility with subsequent code analysis tools; b. Projects have high community attention and usage frequency; c. Projects have relatively complete historical version records to support long-term evolutionary analysis; d. Projects have maintained a certain level of updates in recent years. Based on the above criteria, the complete historical versions of the target projects were selected and downloaded. Subsequently, the data was organized according to project name, with each project containing dozens of different versions. At this point, the initial data required for the experiment was prepared.

[0031] S12: For each version of different projects, existing code smell detection tools are used to scan for smells, thereby obtaining the smell distribution of that version. Based on these detection results, the changes in smells during the project's version evolution are further analyzed. When a certain smell was detected in a previous version but no longer appears in subsequent versions, it is marked as an actionable code smell. This process not only relies on comparing detection results between versions but also comprehensively utilizes project commit messages and code diffs for auxiliary verification, thereby improving the accuracy of actionable label generation.

[0032] S13. Compile and build the Java project source code in the data, and use existing parsing tools to extract the program call relationships from the generated JAR package, parsing the dependency and call information between classes, between classes and methods, and between methods, thereby constructing a complete call graph. Based on this, further calculate various static code metrics for the classes and methods in the project, such as scale metrics (e.g., lines of code), complexity metrics (e.g., cyclomatic complexity), and object-oriented design metrics (e.g., inheritance depth, cohesion, and coupling).

[0033] S14: Take each project and its multiple versions as organizational units, and associate the odor instances involved with the corresponding multi-source information in a unified manner. Finally, unify the originally scattered multi-dimensional data into the same data framework to form a structured and semantically rich research dataset.

[0034] S2: Fine-tune the large language model on the constructed operational code smell dataset, and use the fine-tuned model for code semantic representation, specifically including the following steps:

[0035] S21: Adaptively fine-tune the CodeLlama model using LoRa parameter efficient fine-tuning technology and the operable code smell data collected in S1.

[0036] S22: Use a fine-tuned CodeLlama model to represent code semantics. First, code fragments are segmented into tokens (the smallest units or basic symbols with independent meaning in the source code) and encoded into sequences acceptable to the model. For long code fragments that exceed the maximum context length of the model, a sliding window and overlapping concatenation strategy is adopted to input long sequences in blocks and concatenate and average them in the feature space, thereby mitigating the information loss caused by long context truncation.

[0037] S23: Extract hidden states from several hidden layers in the model, and generate the final code semantic vector representation as the semantic features of the input code to be tested by layer-by-layer weighted averaging.

[0038]

[0039] in, Indicates the first The hidden state of the layer This represents the total number of layers in the model. The selected layer number.

[0040] S3: Learn the multi-indicator call graph based on graph neural network to obtain the structural features of the code.

[0041] In a preferred example, S3 includes the following steps:

[0042] S31: Integrate code call graphs and static code metrics. By traversing the call graph, code metrics are embedded into the nodes of the call graph based on the method name or class name of the nodes in the graph, thereby extending the traditional call relationship into a multi-metric call graph that combines structural and metric features.

[0043] Optionally, at the edge level, weight information such as call frequency and call direction can be introduced so that the edges can reflect the call intensity and the importance of the dependency.

[0044] S32: For static code metrics of nodes in a multi-metric call graph, standardize them so that metrics of different dimensions are on the same scale; for edges, introduce an edge pruning strategy to remove edges that are far from the target code (hop count > 4) to reduce the size of the graph.

[0045] S33: Use a graph neural network to learn a multi-metric call graph, capture cross-node call dependencies and metric interactions through a message passing mechanism, and use a graph attention network to weight and aggregate node features to highlight high-importance nodes related to operable code smells, ultimately obtaining a graph-level representation vector to represent code structure features.

[0046] S4: The two types of code features obtained in the above process are combined and then detected using an operable code odor detection model.

[0047] In a preferred example, S4 includes the following steps:

[0048] S41: For the code semantic features obtained in S2 and the code structural features obtained in S3, a progressive fusion strategy is adopted. By gradually merging the two types of features, their interaction and information sharing are gradually strengthened while maintaining their independence.

[0049] Optionally, the semantic representation vector is first... And structure representation vector projection To achieve the same dimension, dimension alignment is performed through two independent fully connected layers or a linear transformation. In subsequent multi-layer fusion networks, each layer uses an attention mechanism to dynamically compute fusion features, as shown in the following formula:

[0050]

[0051]

[0052] in It is a randomly initialized parameter matrix. It is a random bias offset vector. These are the semantic representation vector and the structural representation vector after linear transformation. It is the first The transformation results of the layers, These are learnable weights. Then it represents the first The fusion of layers allows for the retention of more original feature information in shallow layers and the enhancement of mixed features in deeper layers.

[0053] Furthermore, residual connections preserve some of the original features and directly connect them to the fusion output of the last layer. This ensures that no key information is lost during the fusion process. This approach better reflects the synergistic effect of semantic and structural information, thereby enhancing the expressive power of features.

[0054] S42: A multi-layer classification network is used to process the fused features. This network consists of several fully connected layers, with the number of neurons in each layer gradually decreasing, forming a structure that gradually compresses the feature dimension. Each layer undergoes a non-linear transformation using an activation function (ReLU) to enhance the model's expressive power. Finally, the output of the softmax function is used to determine whether a code unit is an operable code smell.

[0055] To clearly illustrate the present invention, the following detailed description is provided in conjunction with specific embodiments.

[0056] Optionally, this embodiment uses the LLM4ACSD operable odor dataset as an application example for detailed explanation, and the specific implementation steps are as follows:

[0057] S1: Collect actionable code smells data. This involves collecting historical versions from multiple open-source projects and processing the data to obtain an experimental dataset. Specifically, this includes the following sub-steps:

[0058] S11: Select ten open-source projects from GitHub, including cassandra, cayenne, dbeaver, dubbo, and easyexcel. The number of versions, application areas, code size, number of classes, and number of functions for each project are shown in Table 1.

[0059] Table 1. Version number, domain, code size, number of classes, and number of methods for each project.

[0060]

[0061] S12: All version codes for the different projects collected. Static analysis tools PMD, Designitejava, checkstyle, and Décor were used to cross-detect five code smells: God Class, Feature Envy, Long Method, Complex, and Long Parameter List, thereby obtaining the distribution of smells across different versions of the project. Then, following the version evolution timeline of each project, for All odors detected in this version are marked. If there are any odors in this version... If the odor is not detected in subsequent versions of this project, and the commit messages and diffs related to the odor show that a change was indeed made, then the odor will be considered an odor. Marked as an operable code smell and add to The collection of operable code odors.

[0062] S13: Compile and build the Java source code of the ten collected projects, use Dependecy Finder to parse the generated JAR packages to obtain the code call relationships, and then use a Python script to parse the dependency and call information between the codes and construct a complete call graph. .in These are the edges in the graph, encompassing the relationships between classes, between classes and methods, and between methods. The nodes in the diagram are of two types: functions and classes. Subsequently, the CK tool was used to calculate various static code metrics for classes and methods in the project. Examples include scale metrics (such as lines of code), complexity metrics (such as cyclic complexity), and object-oriented design metrics (such as inheritance depth, cohesion, and coupling).

[0063] S14: Organize the ten projects and their multiple versions as organizational units, and associate the odor instances involved with the corresponding call diagrams. Static indicators A unified association is performed to form a structured research dataset, LLM4ACSD.

[0064] S2: Fine-tune the large language model on the constructed operable code smell dataset LLM4ACSD, and use the fine-tuned model for code semantic representation. This includes the following steps:

[0065] S21: Construct a fine-tuning dataset on the LLM4ACSD dataset Subsequently, the CodeLlama model was fine-tuned for domain adaptation using LoRa parameter efficiency fine-tuning technology through the LLaMa-Factory fine-tuning framework.

[0066] S22: Using a fine-tuned CodeLlama model to represent code semantics, first, the code snippets... The code is segmented into tokens (the smallest unit or basic symbol with independent meaning in the source code) and encoded into a sequence acceptable to the model using CodeLlamaTokenizer. For long code snippets exceeding the maximum context length of the model, a sliding window and overlapping concatenation strategy is used to input the long sequence in blocks. And perform splicing and averaging in the feature space. This helps to mitigate information loss caused by long context truncation.

[0067] S23: Extract the hidden states from layers 20, 21, and 22 of the model, and generate the final code semantic vector representation as the semantic features of the input code to be tested by layer-by-layer weighted averaging.

[0068]

[0069] S3: Learn the multi-indicator call graph based on graph neural networks to obtain the structural features of the code, specifically including the following steps:

[0070] S31: Traversing the call graph of the code under test Nodes in Analyze the static code metrics corresponding to nodes by function name or class name. Adding these parameters to the node's attribute vector yields a multi-metric call graph. At the edge level, weight information such as call frequency and call direction is introduced so that the edges can reflect the intensity of the call and the importance of the dependency.

[0071] S32: For multi-indicator call charts Middle node For static metrics, Min-Max scaling is applied; for edges, an edge pruning strategy is introduced to remove edges that are far from the target code (hop count > 4) to reduce the size of the graph.

[0072] S33: Use a Graph Neural Network to learn a multi-metric call graph. Capture cross-node call dependencies and metric interactions in the graph through a message passing mechanism. Use a Graph Attention Network to weighted aggregate node features, and finally obtain a graph-level representation vector. .

[0073] S4: Combine the two types of code features obtained in the preceding process, and then detect them using an actionable code smell detection model, specifically including the following steps:

[0074] S41: Regarding the code semantic features obtained in S2 and the code structure features obtained in S3 A progressive fusion strategy is adopted, which gradually enhances the interaction and information sharing between the two while maintaining their independence through feature dimensionality reduction and residual mechanisms. First, the semantic representation vectors are... Structural representation of vector projection To the same dimension Dimension alignment is achieved through two independent fully connected layers or linear transformations. In subsequent multi-layer fusion networks, each layer uses an attention mechanism to dynamically compute fusion features, as shown in the following formula:

[0075]

[0076]

[0077]

[0078] Among them is Learnable weights allow for the retention of more original feature information in shallow layers and the enhancement of blended features in deeper layers. Furthermore, residual connections preserve a portion of the original features, directly connecting to the fused output of the final layer. This ensures that no critical information is lost during the integration process.

[0079] S42: A multi-layer classification network is used to process the fused features. This network consists of several fully connected layers, with the number of neurons in each layer gradually decreasing. This forms a structure that gradually compresses the feature dimensions. Each layer undergoes a non-linear transformation using an activation function (ReLU), and the output of the softmax function is used to determine whether a code unit is an operable code smell.

[0080] Optionally, to specifically illustrate the advantages of the present invention, this embodiment is compared with mainstream methods on the LLM4ACSD dataset, and the results are shown in Table 2:

[0081] Table 2 Comparison results of this embodiment with other methods

[0082]

[0083] As can be seen from the table, the method proposed in this embodiment achieved optimal or near-optimal results across all six evaluation metrics. (LLM4ACSD) outperforms all baseline methods in overall performance. In terms of Accuracy, Precision, F1-score, and AUC, this embodiment shows improvements of approximately 14.2%, 15.8%, 4.8%, and 11.7% respectively compared to the best-performing baseline (ACSDetector), significantly enhancing overall detection accuracy and the model's discriminative ability.

[0084] Based on the same concept as the above method embodiments, this application also proposes an operable code odor fusion detection system based on a large language model and a multi-index call graph, including:

[0085] The code smell detection module is used to scan the historical version information of open source projects with code smell detection tools to identify code smells in the historical version information. Based on the evolution history of code smells, its operability is judged. At the same time, the project source code is compiled and built to generate a code call graph. Using open source projects and historical version information as organizational units, smell instances are associated with multi-source information to form an operable code smell dataset.

[0086] The code semantic feature construction module is used to fine-tune the large language model based on the code smell dataset, extract the intermediate hidden layer states, and generate a code semantic vector as the code semantic feature to be tested by weighted averaging.

[0087] The code structure feature construction module is used to traverse the code call graph, embed static code metrics into the nodes of the code call graph according to the method name or class name of the node in the graph, form a multi-metric call graph, introduce weight information at the edge level to enhance call strength and dependency relationship, and learn the multi-metric call graph to represent code structure features through graph neural network.

[0088] The feature fusion module is used to fuse the semantic features and structural features of the code using a progressive fusion strategy, process the fused features through a multi-layer classification network, and use a discriminant function to determine whether a code unit belongs to an operable code smell.

[0089] The embodiments of the present invention have been described in detail above, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations can be made to these embodiments, including components, without departing from the principles and spirit of the present invention, and these variations still fall within the protection scope of the present invention.

Claims

1. An operable code odor fusion detection method based on a large language model and multi-index call graph, characterized in that, Includes the following steps: The code smell detection tool scans the historical version information of open source projects to identify code smells in each historical version information. Based on the evolution history of the code smell, its operability is judged. At the same time, the source code of the project is compiled and built to generate a code call graph. Using the open source project and historical version information as the organizational unit, the smell instances are associated with multi-source information to form an operable code smell dataset. Based on the code smell dataset, fine-tune the large language model, extract the intermediate hidden layer states, and generate code semantic vectors as the semantic features of the code to be tested by weighted averaging. Traverse the code call graph and embed static code metrics into the nodes of the code call graph according to the method name or class name of the node in the graph to form a multi-metric call graph. At the same time, introduce weight information at the edge level to enhance the call strength and dependency relationship, and learn the multi-metric call graph to represent the code structure features through graph neural network. A progressive fusion strategy is adopted to fuse the semantic features and structural features of the code. The fused features are processed through a multi-layer classification network, and a discriminant function is used to determine whether the code unit belongs to the operable code smell.

2. The method according to claim 1, characterized in that, The operable code odor dataset includes the basic attributes of the odor, operability labels, static indicators carried by nodes and edges, and project and version number information; wherein the operable code odor contains key-value pairs of related codes and annotation results.

3. The method according to claim 2, characterized in that, When a code smell exists in a previous version but disappears in a subsequent version, it is marked as an actionable code smell. This is then verified by combining commit information with code differences to improve the accuracy of actionable code smells.

4. The method according to claim 1, characterized in that, The CodeLlama model is fine-tuned using the LoRa parameter efficient fine-tuning method and an operable code smell dataset. The fine-tuned CodeLlama model is then used to represent the semantic features of the code, including the following steps: First, the code snippet is segmented into tokens and encoded into a sequence acceptable to a large language model; Secondly, for long sequences that exceed the maximum context length of a large language model, a sliding window and overlapping concatenation strategy is adopted. Then, by processing long sequences in blocks and concatenating and averaging them in the feature space, the information loss caused by long context truncation is mitigated. Finally, the hidden layer states are extracted and weighted to generate a code semantic vector representing the code semantic features.

5. The method according to claim 1 or 4, characterized in that, Weight information is introduced at the edge level to enhance the intensity of calls and dependencies. The weight information includes call frequency and direction information.

6. The method according to claim 1, characterized in that, The static code metrics of the nodes in the multi-metric call graph are standardized to unify the scale of metrics with different dimensions; at the same time, an edge pruning strategy is adopted to remove edges with more than 4 hops to simplify the graph structure.

7. The method according to claim 1 or 6, characterized in that, The graph neural network is used to learn the multi-indicator call graph. Node dependencies are captured through message passing and attention mechanisms, highlighting highly important nodes related to operable code smells.

8. The method according to claim 1, characterized in that, When the progressive fusion strategy fuses code semantic features and code structural features, it gradually strengthens their interaction and information sharing while maintaining their independence through feature dimensionality reduction and residual mechanisms. Specifically, it includes the following steps: First, the semantic representation vector and the structural representation vector are projected to the same dimension, and then dimension alignment is performed through two independent fully connected layers or linear transformations. Secondly, attention mechanisms are used to dynamically compute fusion features at each layer of the multi-layer fusion network; Finally, residual connections are used to preserve some of the original features to prevent the loss of key information.

9. The method according to claim 1 or 8, characterized in that, The multi-layer classification network progressively compresses the feature dimension through fully connected layers, uses the ReLU activation function for non-linear processing in each layer, and uses the softmax function to determine whether there is any code smell that is not operable.

10. An operable code odor fusion detection system based on a large language model and multi-index call graph, characterized in that, include: The code smell detection module is used to scan the historical version information of open source projects with code smell detection tools to identify code smells in the historical version information. Based on the evolution history of code smells, its operability is judged. At the same time, the project source code is compiled and built to generate a code call graph. Using open source projects and historical version information as organizational units, smell instances are associated with multi-source information to form an operable code smell dataset. The code semantic feature construction module is used to fine-tune the large language model based on the code smell dataset, extract the intermediate hidden layer states, and generate a code semantic vector as the code semantic feature to be tested by weighted averaging. The code structure feature construction module is used to traverse the code call graph, embed static code metrics into the nodes of the code call graph according to the method name or class name of the node in the graph, form a multi-metric call graph, introduce weight information at the edge level to enhance call strength and dependency relationship, and learn the multi-metric call graph to represent code structure features through graph neural network. The feature fusion module is used to fuse the semantic features and structural features of the code using a progressive fusion strategy, process the fused features through a multi-layer classification network, and use a discriminant function to determine whether a code unit belongs to an operable code smell.

Citation Information

Patent Citations

  • Code peculiar smell recognition method based on code semantics and measurement

    CN115952076A

  • Code peculiar smell detection method and system based on AST code measurement and code semantics

    CN119248277A