Software development system and software development method

By constructing a multi-dimensional technical debt quantization model and self-supervised learning technology, the problem of objective quantization and insufficient code semantic understanding of technical debt management is solved, and highly targeted reconstruction suggestions are generated, resource allocation is optimized, and software development efficiency and quality are improved.

CN120335775AInactive Publication Date: 2025-07-18BEIJING CAOMU TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510485914.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing technology, technical debt management lacks objective quantitative means, insufficient understanding of code semantics, and lack of scientific basis for refactoring decisions, resulting in a decline in code quality and inefficient development.

Method used

Build a multi-dimensional technical debt quantitative model, build a multi-level code semantic representation model through abstract syntax trees, data flow graphs and control flow graphs, and use a self-supervised learning model to train code representations, generate context-related reconstruction suggestions, and optimize reconstruction path planning.

Benefits of technology

It realizes objective quantification of technical debt, improves the depth of code semantics understanding, enhances the targeted nature of reconstruction suggestions, optimizes resource allocation, significantly reduces code maintenance costs and defect repair cycles, and improves developer productivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335775A_ABST
    Figure CN120335775A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of software development, and discloses a software development system and a software development method. The software development method comprises the following steps: constructing a multi-dimensional technical debt quantitative model, and calculating a technical debt score based on factors such as code complexity, change frequency and defect association degree; a code semantic multi-level representation model is constructed, and multi-level representation of codes is constructed through combined analysis of an abstract syntax tree, a data flow diagram and a control flow diagram; a self-supervised learning model is applied to train code representation, and development intentions and business concepts contained in codes are recognized; generating a context-dependent reconstruction suggestion; and optimizing the reconstruction path planning. Through objective quantification of the technical debt and deep understanding of code semantics, the technical problems that in the prior art, technical debt management is difficult to quantify and reconstruction decision-making lacks scientific basis are solved, code maintenance cost is remarkably reduced, and development efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software development, and more specifically, to a software development system and a software development method. Background Art

[0002] As software systems grow in size and complexity, maintaining and evolving existing codebases has become an extremely important part of software development. In the software development process, due to time pressure, changing requirements, or technical limitations, developers often adopt short-term, convenient solutions, leading to the accumulation of "technical debt". These technical debts manifest themselves in problems such as reduced code quality, complex structure, and poor maintainability. Long-term accumulation will seriously affect the stability of the software system and the productivity of the development team.

[0003] There are three main types of existing technical debt management methods: methods based on static code analysis, methods based on development history data analysis, and methods based on expert experience. Methods based on static code analysis detect potential problems in the code through a predefined set of rules, but are often limited to superficial grammatical features and have difficulty understanding the deep semantics and development intent of the code. Methods based on development history data analysis identify frequently changing or error-prone code areas by mining data from version control systems and defect tracking systems, but have limited use of code structure and business domain knowledge. Methods based on expert experience rely on the subjective judgment of developers and are difficult to form objective and systematic evaluation criteria.

[0004] In addition, existing code refactoring tools usually provide some mechanized code conversion operations based on predefined refactoring patterns, but lack a deep understanding of code semantics and business domain concepts, making it difficult to provide customized refactoring suggestions based on specific business scenarios and development intentions. At the same time, existing methods also lack systematic considerations in terms of priority sorting and resource allocation, making it difficult to maximize the value return of refactoring work under limited resource constraints.

[0005] Therefore, how to objectively quantify technical debt, deeply understand code semantics and development intent, automatically generate targeted refactoring suggestions, and optimize the refactoring execution path are technical problems that need to be solved urgently in the current software development field. Summary of the invention

[0006] The present invention provides a software development system and a software development method to solve technical problems in related technologies such as difficulty in quantifying technical debt, insufficient understanding of code semantics, and lack of scientific basis for refactoring decisions.

[0007] The present invention discloses a software development method, comprising the following steps: constructing a multi-dimensional technical debt quantification model, which calculates a technical debt score value based on factors such as code complexity, change frequency, defect correlation degree, and dependency complexity; constructing a multi-level code semantics representation model, which constructs a multi-level representation of the code syntax layer, semantics layer, intention layer, and domain layer through combined analysis of the abstract syntax tree, data flow graph, and control flow graph; applying a self-supervised learning model to train the code representation, where the self-supervised learning model learns the mapping relationship between levels by minimizing a loss function that includes reconstruction loss, contrastive loss, and domain consistency loss, and identifies the development intentions and business concepts contained in the code; generating context-related refactoring suggestions, where the refactoring suggestions are based on a historical refactoring case library and the learned semantic representation, and provide refactoring types, refactoring scopes, and specific implementation steps for specific code segments; optimizing the refactoring path planning, where the optimization calculates a refactoring benefit prediction value by combining business value assessment and technical debt quantification results, and generates an optimal refactoring execution path considering resource constraints.

[0008] Further, the step of constructing the multi-dimensional technical debt quantification model includes: collecting historical data of the code library, where the historical data includes code version history, defect records, and change frequency information; calculating code complexity metrics, where the code complexity metrics include cyclomatic complexity, inheritance depth, method length, and class cohesion; analyzing code change patterns to identify code areas with frequent changes but unstable structures; calculating the correlation degree between code areas and historical defects; constructing a technical debt scoring algorithm, where the scoring algorithm synthesizes multi-dimensional metrics into a single technical debt score value.

[0009] Further, the step of constructing the multi-level code semantics representation model includes: analyzing the source code to generate an abstract syntax tree representation, capturing the syntax structure information of the code; based on the abstract syntax tree, constructing a data flow graph to represent the flow and transformation relationship of data in the program; based on the abstract syntax tree, constructing a control flow graph to represent the execution path and conditional branches of the program; integrating the abstract syntax tree, data flow graph, and control flow graph to construct a unified code representation model; using code comments, variable naming, and document information to extract developer intentions and business domain concepts.

[0010] Further, the steps of training the code representation using the self-supervised learning model include: constructing self-supervised learning tasks, which include code completion tasks, context prediction tasks, and syntax structure reconstruction tasks; establishing a loss function for the self-supervised learning model, which includes reconstruction loss, contrast loss, and domain consistency loss; using code snippets in the code library as training data, and optimizing the model parameters by minimizing the loss function; generating deep semantic representations of the code, and improving the intent layer representation and domain layer representation; establishing a code-natural language mapping model to align the code representation with the natural language description.

[0011] Further, the steps of generating context-related refactoring suggestions include: constructing a historical refactoring case library, collecting previous successful refactoring instances and their context information; for the identified high-tech debt code regions, extracting their code representations and context information; applying a similarity matching algorithm to find cases semantically similar to the current code in the historical refactoring case library; generating customized refactoring suggestions based on the characteristics of the matched historical cases and the current code; validating and optimizing the generated refactoring suggestions to ensure the feasibility and effectiveness of the suggestions.

[0012] Further, the steps of optimizing the refactoring path planning include: evaluating the implementation cost of each refactoring suggestion, where the implementation cost includes the required time, human resources, and technical difficulty; evaluating the business value of each refactoring suggestion, where the business value includes the improvement of code maintainability, system stability, and development efficiency; calculating the refactoring benefit prediction value, which is jointly determined by the implementation cost and the business value; considering the dependency relationships between refactoring suggestions and constructing a refactoring dependency graph; applying an optimization algorithm to generate a refactoring execution path that maximizes the overall benefit under given resource constraints.

[0013] The present invention also provides a software development system, including:

[0014] A technical debt quantification module for constructing a multi-dimensional technical debt quantification model and calculating a technical debt score according to factors such as code complexity, change frequency, and defect correlation;

[0015] A code representation construction module for constructing a multi-level code semantic representation model, and constructing a multi-level representation of the code syntax layer, semantic layer, intent layer, and domain layer through the combined analysis of abstract syntax trees, data flow graphs, and control flow graphs;

[0016] A self-supervised learning module for training code representations using self-supervised learning algorithms, learning the mapping relationships between levels, and identifying the development intent and business concepts contained in the code by combining code-natural language contrast learning;

[0017] A refactoring suggestion generation module, which is used to automatically generate refactoring suggestions and implementation steps for specific code snippets based on a historical refactoring case library and learned semantic representations;

[0018] A path planning module, which is used to calculate the predicted refactoring benefit value by combining business value evaluation and technical debt quantification results, and generate an optimal refactoring execution path considering resource constraints

[0019] The software development method based on multi-dimensional analysis and self-supervised learning provided by the present invention solves the technical problems in the prior art such as the difficulty in quantifying technical debt, insufficient understanding of code semantics, and lack of scientific basis for refactoring decisions, and has achieved the following beneficial effects:

[0020] The objective quantification of technical debt is realized: through multi-dimensional analysis and comprehensive scoring algorithms, the present invention converts the subjective judgment of technical debt into an objective numerical evaluation, eliminates the subjective judgment bias in traditional methods, and provides a scientific basis for refactoring decisions.

[0021] The depth of code semantic understanding is improved: by constructing a multi-level code semantic representation model and applying self-supervised learning technology, the present invention can deeply understand the syntax structure, semantic relationship, developer intention, and business domain concepts of the code, greatly improving the ability to understand the essence of the code.

[0022] The pertinence of refactoring suggestions is enhanced: based on a historical refactoring case library and in-depth semantic understanding, the present invention can generate customized refactoring suggestions for specific code snippets and business scenarios, improving the accuracy and operability of refactoring suggestions.

[0023] The refactoring resource allocation is optimized: through business value evaluation and refactoring path optimization, the present invention ensures the maximization of technical debt reduction and business value improvement under limited resource constraints, making the refactoring work have a clear input-output ratio.

[0024] The technical threshold is reduced: the present invention automates the process of technical debt identification and refactoring suggestion generation, enabling junior developers to also participate in high-quality code refactoring work and reducing the professional threshold of code refactoring.

[0025] According to actual test data, using the software development method of the present invention can reduce the code maintenance cost by 41%, shorten the defect repair cycle by 56%, and improve the productivity of developers by 32%, significantly improving the efficiency and quality of software development and maintenance. Description of the Drawings

[0026] Figure 1 is the main step flowchart of the software development method of the present invention;

[0027] Figure 2 is the flowchart of the steps for constructing the technical debt quantification model of the present invention;

[0028] Figure 3 is the flowchart of the steps for constructing the multi-level representation model of the code semantics of the present invention;

[0029] Figure 4 is the flowchart of the steps for training the code representation of the self-supervised learning model of the present invention;

[0030] Figure 5 is the flowchart of the steps for generating context-related refactoring suggestions of the present invention;

[0031] Figure 6 is the flowchart of the steps for optimizing the refactoring path planning of the present invention. Detailed implementation manners

[0032] Before describing the present application in detail, in order to help understand the technical solution of the present application, the terms to be used in the present application are first explained below.

[0033] Technical debt: Refers to the additional maintenance costs and potential risks generated by sacrificing long-term code quality due to the adoption of short-term convenient solutions during the software development process.

[0034] Code refactoring: Refers to the process of adjusting and optimizing the internal structure without changing the external behavior of the software, aiming to improve the code quality and maintainability.

[0035] Abstract Syntax Tree (AST): A tree-like data structure representing the syntax structure of source code, where each node represents a syntax construct in the source code.

[0036] Data Flow Graph (DFG): A graphical representation describing how data flows and transforms in a program, where nodes represent data processing units and edges represent data flow paths.

[0037] Control Flow Graph (CFG): A graphical representation describing the execution paths of a program, where nodes represent basic blocks and edges represent possible execution paths.

[0038] Self-supervised learning: A machine learning paradigm that learns data representations by automatically generating labels from the data itself without manual annotation.

[0039] The problems existing in the prior art are mainly reflected in the following aspects:

[0040] Difficulty in quantifying technical debt: Existing software development methods lack objective means for quantifying technical debt and usually rely on the subjective judgment of developers, resulting in the lack of a scientific basis for technical debt management.

[0041] Limited recognition of code refactoring opportunities: The traditional code refactoring process highly depends on the personal experience and intuition of developers and it is difficult to systematically identify the code areas that most need refactoring.

[0042] Insufficient semantic understanding: Existing code analysis tools are mostly based on predefined syntax rules and pattern matching, lacking the ability to understand the deep semantics and domain concepts of code.

[0043] Difficulty in extracting development intentions: The original intentions of developers and business domain knowledge embedded in software code are usually difficult to automatically extract and utilize, and this information is crucial for high-quality refactoring.

[0044] Lack of business value association: Existing technical debt management methods are difficult to quantify the business value of refactoring work, resulting in unreasonable resource allocation and insufficient priority for refactoring work.

[0045] According to an embodiment of the present application, a software development method is provided. This method can objectively quantify technical debt, deeply understand code semantics and development intentions, automatically generate context-related refactoring suggestions, and optimize the refactoring execution path. It should be understood that through the technical solution of the present application, the code maintenance cost can be effectively reduced, the defect repair cycle can be shortened, the productivity of developers can be improved, and at the same time, the refactoring decision-making can be made more objective and scientific.

[0046] It should be noted that the technical solution of the present application is mainly applied to software development and maintenance scenarios, and is particularly applicable to the following situations:

[0047] Maintenance of large legacy systems: For large software systems that have been running for many years, a large amount of technical debt is usually accumulated. The method provided by the present application can help the maintenance team identify the most critical technical debt and provide priority ranking.

[0048] Agile development process: In a fast-paced iterative agile development environment, due to delivery pressure, the team often accumulates technical debt. The method provided by the present application can continuously monitor technical debt during the development process and provide refactoring suggestions at the appropriate time.

[0049] Team knowledge transfer: When new members join the development team or existing members leave, the development intentions and domain knowledge embedded in the code may be lost. The method provided by the present application can help extract and preserve this knowledge, reducing the cost of knowledge transfer.

[0050] Code quality management: For organizations that focus on code quality, the method provided by the present application can be used as part of the code quality management process to continuously monitor and improve code quality.

[0051] Training for development teams: For junior developers, the refactoring suggestions provided by the present application can be used as learning materials to help them understand the characteristics of high-quality code and refactoring techniques.

[0052] The present application will be described in detail below with reference to specific embodiments. It should be noted that the following embodiments are only used to illustrate the present application and do not limit the protection scope of the present application.

[0053] According to an embodiment of the present application, the method of this embodiment includes the following steps:

[0054] Step S1: Construct a multi-dimensional technical debt quantification model.

[0055] In this step, according to the embodiment of the present application, by analyzing code characteristics in multiple dimensions, a technical debt quantification model is constructed to convert subjective technical debt judgments into objective numerical evaluations. Specifically, this step includes the following sub-steps:

[0056] Step S1-1: Collect historical data of the code library, including information such as code version history, defect records, and change frequency.

[0057] Step S1-2: Calculate code complexity metrics, including static metrics such as cyclomatic complexity, inheritance depth, method length, and class cohesion.

[0058] Step S1-3: Analyze code change patterns to identify code areas with frequent changes but unstable structures.

[0059] Step S1-4: Calculate the correlation between code areas and historical defects based on defect records.

[0060] Step S1-5: Construct a technical debt scoring algorithm that combines the above-mentioned metrics in each dimension into a single technical debt score value. The technical debt scoring algorithm can be expressed as:

[0061] TD score = α·C complexity + β·C change + γ·C defect + δ·C dependency

[0062] Wherein, TD score represents the technical debt score, C complexity represents the code complexity metric, C change represents the change frequency metric, C defect represents the defect correlation metric, C dependency represents the dependency complexity metric, and α, β, γ, and δ are the corresponding weight parameters.

[0063] Optionally, the technical debt scoring algorithm can be adjusted according to the characteristics of different software projects. For example, for safety-critical software, the security vulnerability correlation metric C security can be added, and the scoring algorithm is modified to:

[0064] TDscore = α · C complexity + β · C change + γ · C defect + δ · C dependency + ∈ · C security

[0065] Where ∈ is the weight parameter of the security vulnerability correlation index.

[0066] In some embodiments, the technical debt scoring algorithm may also consider factors such as code documentation integrity and test coverage. Optionally, these weight parameters can be automatically adjusted by machine learning methods based on historical refactoring effects, so that the scoring algorithm can adapt to the characteristics of different teams and projects.

[0067] Through step S1, this application establishes an objective and multi-dimensional technical debt assessment model. Therefore, this model can accurately identify areas with high technical debt in the code library and provide a basis for subsequent refactoring decisions. As Figure 2 shown, step S1 specifically includes sub-steps such as collecting code library historical data, calculating code complexity metrics, analyzing code change patterns, calculating defect correlation, and constructing a technical debt scoring algorithm.

[0068] Step S2: Construct a multi-level representation model of code semantics.

[0069] In this step, this application performs multi-level semantic representation on the source code, from the syntax structure to the domain concept, to construct a complete code understanding model. Specifically, this step includes the following sub-steps:

[0070] Step S2-1: Analyze the source code to generate an Abstract Syntax Tree (AST) representation, capturing the syntax structure information of the code.

[0071] Step S2-2: Based on the Abstract Syntax Tree, construct a Data Flow Graph (DFG) to represent the flow and transformation relationship of data in the program.

[0072] Step S2-3: Based on the Abstract Syntax Tree, construct a Control Flow Graph (CFG) to represent the execution path and conditional branches of the program.

[0073] Step S2-4: Integrate the above three graph structures to construct a unified code representation model. This model can be formally represented as:

[0074] CR = {T syn , T sem , T int , T dom}

[0075] Where CR represents the complete code representation, and T synRepresents the syntactic layer representation (mainly based on AST), T sem Represents the semantic layer representation (mainly based on DFG and CFG), T int Represents the intent layer representation (developer intent), T dom Represents the domain layer representation (business domain concepts).

[0076] In some embodiments, the syntax and semantic analysis methods can be customized for different programming languages. For example, for object-oriented languages, features such as class inheritance relationships, interface implementations, and method overrides can be additionally analyzed; for functional languages, features such as higher-order functions, currying, and immutable data structures can be analyzed.

[0077] Optionally, the code representation model can also integrate a Program Dependence Graph (PDG) to capture control and data dependence relationships in the program. In this case, the code representation model can be extended to:

[0078] CR = {T syn , T sem , T int , T dom , T dep}

[0079] where T dep represents the dependence layer representation, mainly constructed based on PDG.

[0080] Step S2-5: Utilize information such as code comments, variable naming, and documentation to initially extract developer intent and business domain concepts, and construct initial intent layer and domain layer representations.

[0081] Through step S2, according to the embodiments of the present application, a multi-level code representation model is established. In addition, this model not only captures the syntax and semantic information of the code, but also initially extracts developer intent and business domain concepts, laying a foundation for subsequent in-depth understanding. As Figure 3 shown, step S2 includes sub-steps such as analyzing the source code, constructing data flow graphs and control flow graphs, integrating multiple graph structures, and extracting intent and domain concepts.

[0082] Step S3: Apply a self-supervised learning model to train the code representation.

[0083] In this step, according to an embodiment of the present application, self-supervised learning technology is used to directly learn the deep semantic representation of the code from the code library without manual annotation. It should be understood that this step includes the following sub-steps:

[0084] Step S3-1: Construct self-supervised learning tasks, including code completion tasks, context prediction tasks, and syntax structure reconstruction tasks.

[0085] Step S3-2: Establish the loss function of the self-supervised learning model. This loss function consists of three parts: reconstruction loss, contrastive loss, and domain consistency loss, and can be expressed as:

[0086] L(CR) = L rec + λ1L con + λ2L dom

[0087] where L rec represents the reconstruction loss, which is used to ensure that the model can accurately reconstruct the input code; L con represents the contrastive loss, which is used to learn the semantic mapping between the code and the natural language description; L dom represents the domain consistency loss, which is used to maintain the consistency between the code representation and the business domain concepts; λ1 and λ2 are weight parameters.

[0088] Optionally, in some embodiments, the loss function can add a structure preservation loss term L str , which is used to ensure that the representation learned by the model retains the structural information of the code. The modified loss function is:

[0089] L(CR) = L rec + λ1L con + λ2L dom + λ3L str

[0090] where λ3 is the weight parameter of the structure preservation loss.

[0091] For example, for a code library containing a large number of repetitive patterns, a pattern recognition task can be added to the self-supervised learning task, enabling the model to automatically identify and extract common patterns and design patterns in the code. This is particularly valuable for identifying refactorable code snippets.

[0092] Step S3-3: Train the self-supervised learning model. Use the code snippets in the code library as training data without manual annotation, and optimize the model parameters by minimizing the above loss function.

[0093] Step S3-4: Use the trained model to generate the deep semantic representation of the code. In particular, this step will further refine the intent layer representation T int and the domain layer representation T dom .

[0094] Step S3-5: Establish a code-natural language mapping model to align the code representation with the natural language description (such as comments, documents) to deepen the understanding of the code intent.

[0095] Through step S3, according to the technical solution of the present application, in-depth semantic understanding of the code is achieved. It can be seen that the present application can capture the original intentions of developers and business domain concepts, and these information are crucial for high-quality code refactoring. As Figure 4 shown, step S3 includes sub-steps such as constructing a self-supervised learning task, establishing a loss function, training a model, generating deep semantic representations, and establishing a code-natural language mapping model.

[0096] Step S4: Generate context-related refactoring suggestions.

[0097] In this step, according to the embodiments of the present application, based on the technical debt assessment and code semantic understanding obtained in the foregoing steps, refactoring suggestions for specific code snippets are automatically generated. It should be noted that this step includes the following sub-steps:

[0098] Step S4-1: Construct a historical refactoring case library, and collect previous successful refactoring instances and their context information.

[0099] Step S4-2: For the identified high-technical debt code regions, extract their code representations and context information.

[0100] Step S4-3: Apply a similarity matching algorithm to search for cases semantically similar to the current code in the historical refactoring case library. The similarity matching algorithm can be expressed as:

[0101]

[0102] where Sim(C1,C2) represents the overall similarity between code snippets C1 and C2, and sim syn 、sim sem 、sim int and sim dom represent the similarities at the syntax layer, semantic layer, intention layer, and domain layer respectively, and θ1, θ2, θ3, and θ4 are the corresponding weight parameters.

[0103] Optionally, in some embodiments, similarity matching can use a graph neural network-based method to convert the code representation into a graph structure, and then use a graph matching algorithm to calculate the similarity. For example, a graph isomorphism network (GIN) or a graph attention network (GAT) can be used to process the graph representation of the code, and then a graph-level similarity metric function is used to calculate the similarity between cases.

[0104] In a specific application scenario, for example, for a system with a microservices architecture, the similarity of service interdependencies can be additionally considered, and the similarity calculation formula is modified as:

[0105]

[0106] Among them, sim srv represents the similarity of service dependency relationships, and T srv represents the service dependency representation, and θ5 is the corresponding weight parameter.

[0107] Step S4-4: Generate customized refactoring suggestions based on the characteristics of the matched historical cases and the current code, including refactoring types, scopes, and specific implementation steps.

[0108] Step S4-5: Verify and optimize the generated refactoring suggestions to ensure their feasibility and effectiveness. The verification process includes static analysis and simulation execution to evaluate whether the refactoring will introduce new problems or break existing functions.

[0109] Through step S4, this application can provide developers with specific and actionable refactoring suggestions. Therefore, these suggestions not only consider the syntactic and semantic characteristics of the code but also the original intentions of the developers and business domain concepts, making them more accurate and relevant. As Figure 5 shown, step S4 includes sub-steps such as building a historical refactoring case library, extracting code representations and context information, applying a similarity matching algorithm, generating refactoring suggestions, and verifying and optimizing the suggestions.

[0110] Step S5: Optimize the refactoring path planning.

[0111] In this step, according to another embodiment of this application, considering resource constraints and business value, plan the optimal refactoring execution path for the development team. Specifically, this step includes the following sub-steps:

[0112] Step S5-1: Evaluate the implementation cost of each refactoring suggestion, including the required time, human resources, and technical difficulty.

[0113] Step S5-2: Evaluate the business value of each refactoring suggestion, including the improvement of code maintainability, system stability, and development efficiency.

[0114] Step S5-3: Calculate the refactoring benefit prediction value, which is jointly determined by the implementation cost and business value and can be expressed as:

[0115]

[0116] Among them, ROI refactor represents the refactoring benefit prediction value, V business represents the business value, P success represents the probability of success, and C implementation represents the implementation cost.

[0117] In some embodiments, the business value V business can be further broken down into multiple sub-factors, for example:

[0118] V business = w1·V maintenance + w2·V performance + w3·V reliability + w4·V security

[0119] wherein, V maintenance represents the value of reduced maintenance cost, V performance represents the value of performance improvement, V reliability represents the value of reliability improvement, V security represents the value of security improvement, and w1, w2, w3, and w4 are the weights of each sub-factor.

[0120] Optionally, the present application may also consider the factor of the accumulation of technical debt over time, and predict the refactoring benefits for different time windows, such as short-term (1 - 3 months), medium-term (3 - 12 months), and long-term (more than 1 year) benefit predictions. In this case, a time decay factor can be used to adjust the weights of each time window.

[0121] Step S5-4: Consider the dependencies between refactoring suggestions and construct a refactoring dependency graph.

[0122] Step S5-5: Apply an optimization algorithm to generate a refactoring execution path that maximizes the overall benefit under given resource constraints. The optimization problem can be expressed as:

[0123]

[0124] x i ∈ {0, 1}, i = 1, 2,..., n

[0125] wherein, x i represents whether to execute the i-th refactoring suggestion, ROI i represents the predicted benefit value of the i-th refactoring suggestion, C i represents the implementation cost of the i-th refactoring suggestion, B represents the total resource budget, and n represents the total number of refactoring suggestions.

[0126] Through step S5, according to the technical solution of the present application, scientific refactoring decision support is provided for the development team. It should be noted that this ensures the maximum reduction of technical debt and the improvement of business value under limited resources. As Figure 6 shown, step S5 includes sub-steps such as evaluating implementation costs and business value, calculating predicted benefit values, constructing dependency graphs, and generating optimal execution paths.

[0127] Comprehensively Figure 1As shown, the method of the present application includes the complete technical solutions of the above steps S1 to S5. These steps are organically combined to form a complete software development method based on multi-dimensional analysis and self-supervised learning.

[0128] According to another embodiment of the present application, there is provided a software development device based on multi-dimensional analysis and self-supervised learning. It should be understood that the device includes:

[0129] A technical debt quantification module, configured to build a multi-dimensional technical debt quantification model and calculate a technical debt score according to factors such as code complexity, change frequency, and defect correlation;

[0130] A code representation construction module, configured to build a multi-level code semantic representation model, and build a multi-level representation of the code syntax layer, semantic layer, intention layer, and domain layer through combined analysis of abstract syntax trees, data flow graphs, and control flow graphs;

[0131] A self-supervised learning module, configured to apply self-supervised learning algorithms to train code representations, learn the mapping relationships between levels, and combine code-natural language contrast learning to identify the development intentions and business concepts contained in the code;

[0132] A refactoring suggestion generation module, configured to automatically generate refactoring suggestions and implementation steps for specific code snippets based on the historical refactoring case library and the learned semantic representations;

[0133] A path planning module, configured to calculate the predicted refactoring benefit value by combining business value assessment and technical debt quantification results, and generate an optimal refactoring execution path considering resource constraints.

[0134] It should be noted that the above modules can be implemented using hardware circuits or a combination of software and a processor.

[0135] Optionally, in some embodiments, the software development device of the present application may further include the following additional modules:

[0136] A code clone detection module, configured to identify duplicate code snippets in the code library and provide additional basis for refactoring;

[0137] A change impact analysis module, configured to predict the code areas that may be affected by refactoring operations and evaluate refactoring risks;

[0138] An automatic test generation module, configured to generate verification tests for refactoring operations to ensure the functional correctness of the refactored code;

[0139] A user feedback learning module, configured to collect developers' feedback on refactoring suggestions and continuously improve the accuracy of the model.

[0140] In some other embodiments, the above-mentioned modules can be implemented as independent microservice components, communicating with each other through standard API interfaces, which facilitates the extension and maintenance of the system. For example, the technical debt quantification module can be implemented as an independent service to regularly analyze the codebase and update the technical debt score; the code representation construction module and the self-supervised learning module can be combined into a code understanding service to provide code semantic representations for other modules.

[0141] The present application also provides an embodiment of a computer device for implementing the above-mentioned software development method based on multi-dimensional analysis and self-supervised learning.

[0142] The computer device includes a central processing unit (CPU), a memory, a network interface, and an input / output interface. It should be understood that the CPU can be a multi-core processor or a distributed processor cluster for executing program instructions to implement the processing logics of the above-mentioned steps. In addition, the memory can include a random access memory (RAM) and a read-only memory (ROM) for storing program instructions and processing data. The network interface is used to communicate with other computer devices through a network. The input / output interface is used to connect external devices such as a display, a keyboard, a mouse, etc.

[0143] According to another embodiment of the present application, a computer program is stored on the computer device, and when the computer program is executed by the central processing unit, the functions of the above-mentioned steps S1 to S5 are implemented. Specifically, the computer program includes:

[0144] The program code of the technical debt quantification module for implementing the function of step S1;

[0145] The program code of the code representation construction module for implementing the function of step S2;

[0146] The program code of the self-supervised learning module for implementing the function of step S3;

[0147] The program code of the refactoring suggestion generation module for implementing the function of step S4;

[0148] The program code of the path planning module for implementing the function of step S5.

[0149] Optionally, in some embodiments, the computer device can be implemented using a distributed computing architecture, where different functional modules are deployed on different computing nodes. For example, the technical debt quantification module and the path planning module can be deployed on a central server, while the code representation construction module and the self-supervised learning module, due to high computing requirements, can be deployed on computing nodes with GPU acceleration capabilities.

[0150] In some other embodiments, the computer device can be integrated with a version control system (such as Git) to automatically trigger the analysis process when a developer submits code, providing instant refactoring suggestions to the developer. Optionally, the present application can also be integrated with a continuous integration / continuous deployment (CI / CD) pipeline, as part of the code quality check, to ensure that the technical debt does not exceed a preset threshold.

[0151] It should be specifically noted that the above embodiments are only the preferred embodiments of the present application, and those skilled in the art can make various other changes or modifications that do not violate the concept of the present application based on the above embodiments. Therefore, these changes or modifications are all within the protection scope of the present application within the scope of the appended claims.

[0152] In addition, it should be understood that the steps of the present application can be flexibly configured according to actual needs. For example, in some embodiments, only the technical debt quantification function of step S1 can be used without implementing other steps; in some other embodiments, steps S2 and S3 can be combined to implement the code semantic understanding function; in still some other embodiments, the technical solution of the present application can be integrated with existing code analysis tools to provide more comprehensive software development support.

Claims

1. A software development method, characterized in that, Including the following steps: Construct a multi-dimensional technical debt quantification model, which calculates a technical debt score value based on factors such as code complexity, change frequency, defect correlation, and dependency complexity; Construct a multi-level code semantics representation model, which constructs a multi-level representation of the code syntax layer, semantics layer, intention layer, and domain layer by combining and analyzing the abstract syntax tree, data flow graph, and control flow graph; Apply a self-supervised learning model to train the code representation. The self-supervised learning model learns the mapping relationship between levels and identifies the development intentions and business concepts contained in the code by minimizing a loss function that includes reconstruction loss, contrast loss, and domain consistency loss; Generate context-related refactoring suggestions. The refactoring suggestions provide the refactoring type, refactoring scope, and specific implementation steps for a specific code snippet based on the historical refactoring case library and the learned semantic representation; Optimize the refactoring path planning. The optimization calculates a refactoring benefit prediction value by combining business value assessment and technical debt quantification results, and generates an optimal refactoring execution path considering resource constraints.

2. The software development method according to claim 1, wherein The steps of constructing the multi-dimensional technical debt quantification model include: Collect historical data from the code repository. The historical data includes code version history, defect records, and change frequency information; Calculate code complexity metrics, which include cyclomatic complexity, inheritance depth, method length, and class cohesion; Analyze the code change pattern and identify code areas that are frequently changed but structurally unstable; Calculate the correlation between code areas and historical defects; Construct a technical debt scoring algorithm, which synthesizes multi-dimensional metrics into a single technical debt score value.

3. The software development method according to claim 1, wherein The steps of constructing the multi-level code semantics representation model include: Analyze the source code to generate an abstract syntax tree representation and capture the syntax structure information of the code; Based on the abstract syntax tree, construct a data flow graph to represent the flow and transformation relationship of data in the program; Based on the abstract syntax tree, construct a control flow graph to represent the execution path and conditional branches of the program; Integrate the abstract syntax tree, data flow graph, and control flow graph to construct a unified code representation model; Utilize code comments, variable naming, and documentation information to extract developer intentions and business domain concepts.

4. The software development method according to claim 1, characterized in that The steps of applying the self-supervised learning model to train the code representation include: Construct self-supervised learning tasks, which include code completion tasks, context prediction tasks, and syntax structure reconstruction tasks; Establish a loss function for the self-supervised learning model. The loss function includes reconstruction loss, contrast loss, and domain consistency loss; Use code snippets in the code repository as training data and optimize the model parameters by minimizing the loss function; Generate a deep semantic representation of the code and improve the intention layer representation and domain layer representation; Establish a code-natural language mapping model to align the code representation with the natural language description.

5. The software development method according to claim 1, wherein The steps of generating context-related refactoring suggestions include: Construct a historical refactoring case library and collect previous successful refactoring instances and their context information; For the identified high-technical debt code areas, extract their code representations and context information; Apply the similarity matching algorithm to find cases in the historical refactoring case library that are semantically similar to the current code; Generate customized refactoring suggestions based on the characteristics of the matched historical cases and the current code; Verify and optimize the generated refactoring suggestions to ensure the feasibility and effectiveness of the suggestions.

6. The software development method according to claim 1, characterized in that The steps for optimizing the refactoring path planning include: Evaluate the implementation cost of each refactoring suggestion, where the implementation cost includes the required time, human resources, and technical difficulty; Evaluate the business value of each refactoring suggestion, where the business value includes the improvement of code maintainability, system stability, and development efficiency; Calculate the predicted refactoring benefit value, which is jointly determined by the implementation cost and the business value; Consider the dependencies between refactoring suggestions and construct a refactoring dependency graph; Apply an optimization algorithm to generate a refactoring execution path that maximizes the overall benefit under given resource constraints.

7. The software development method according to claim 4, characterized in that The loss function of the self-supervised learning model also includes a structure-preserving loss, which is used to ensure that the representations learned by the model retain the structural information of the code.

8. The software development method according to claim 5, characterized in that, The similarity matching algorithm uses a graph neural network-based method to convert the code representation into a graph structure and then uses a graph matching algorithm to calculate the similarity between cases.

9. The software development method according to claim 6, wherein The business value is further broken down into maintenance cost reduction value, performance improvement value, reliability improvement value, and security improvement value.

10. A software development system, characterized in that, Include: A technical debt quantification module for constructing a multi-dimensional technical debt quantification model and calculating the technical debt score based on factors such as code complexity, change frequency, and defect correlation; A code representation construction module for constructing a multi-level code semantic representation model and constructing a multi-level representation of the code syntax layer, semantic layer, intention layer, and domain layer through a combined analysis of the abstract syntax tree, data flow graph, and control flow graph; A self-supervised learning module for training the code representation using a self-supervised learning algorithm, learning the mapping relationship between levels, and combining code-natural language contrast learning to identify the development intentions and business concepts contained in the code; A refactoring suggestion generation module for automatically generating refactoring suggestions and implementation steps for specific code fragments based on the historical refactoring case library and the learned semantic representations; A path planning module for calculating the predicted refactoring benefit value by combining the business value assessment and the technical debt quantification results and generating an optimal refactoring execution path considering resource constraints.

Citation Information

Cited By

  • Data internal feature vector generation method, medium and system

    CN120596896A