Deep learning-based method and system for timely defect prediction of software

A deep learning-based method for software defect prediction in edge computing applications addresses the challenge of timely defect detection in autonomous driving systems by utilizing advanced algorithms and models, enhancing defect identification and quality improvement.

WO2025150599A1PCT designated stage expired Publication Date: 2025-07-17IND COOP FOUND CHONBUK NAT UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/000881
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2024-01-18
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing technologies face challenges in predicting and responding to software defects in edge computing applications, particularly in autonomous driving systems, which are critical for passenger safety, in a timely manner.

Method used

A deep learning-based method involving data collection, labeling, embedding, learning, and evaluation is employed to predict software defects, utilizing algorithms like SZZ and models like UniXCoder and Bi-LSTM, focusing on commit-level defect prediction to enhance defect identification during the development stage.

Benefits of technology

The method effectively identifies and predicts software defects at a commit level, reducing code inspection efforts and improving software quality by enabling timely defect detection and fixation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024000881_17072025_PF_FP_ABST
    Figure KR2024000881_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A method for timely defect prediction of software according to an embodiment of the present invention may comprise: a data collection step of collecting data; a data labeling step of performing labeling on data in which a defect is likely to occur by using an identification algorithm on the collected data; an embedding step of receiving an input of the data labeled in the data labeling step and performing embedding; a learning step of learning context and meaning of the data embedded in the embedding step on the basis of deep learning; and an evaluation step of evaluating a learning result on the basis of training data learned in the learning step.
Need to check novelty before this filing date? Find Prior Art

Description

Deep Learning-Based Software Defect Prediction Method and System

[0001] The present invention relates to a deep learning-based software timely defect prediction method and system.

[0002] Edge computing is applied to a variety of applications, most notably autonomous driving software. Autonomous driving operates by computing data collected from traffic information and providing it to other vehicles. Autonomous driving systems utilize edge computing technology based on 5G networks to collect and process massive amounts of sensor data, which is then transmitted to other vehicles to provide AI-based services.

[0003] The increasing complexity and software-intensive nature of autonomous driving systems is leading to a variety of software defect incidents. However, effective solutions for predicting and responding to software defects in edge computing applications in a timely manner remain elusive.

[0004] Because autonomous driving software is directly linked to passenger safety, identifying and improving defects to ensure reliability is crucial. Therefore, to address the challenges of identifying defects in autonomous driving systems, effective software defect prediction methods are essential.

[0005] The problem to be solved by the present invention is to provide a software timely defect prediction method and system based on deep learning that can effectively predict software defects in a timely manner.

[0006] A software timely defect prediction method according to an embodiment of the present invention may include a data collection step of collecting data about software, a data labeling step of labeling data with a possibility of defect occurrence using an identification algorithm on the collected data, an embedding step of receiving the labeled data in the data labeling step and performing embedding, a learning step of learning context and meaning for the embedded data in the embedding step based on deep learning, and an evaluation step of evaluating a learning result based on learning data learned in the learning step.

[0007] In the data labeling step according to one embodiment of the present invention, the collected data may include commit data and code change data.

[0008] In the data labeling step according to one embodiment of the present invention, the identification algorithm may be an algorithm that automatically identifies change data that causes a defect.

[0009] The identification algorithm according to one embodiment of the present invention may include a step of searching change data for a keyword to identify a commit corresponding to the keyword, a step of identifying a changed code line of a previous version and a modified version in the commit corresponding to the keyword, a step of generating a change that performs at least one of modifying and deleting a code line in a previous modification for the last commit of the commits corresponding to the identified code line, and a step of labeling the commit data for the change among the commit data as a defect and labeling other changes as defect-free.

[0010] The identification algorithm according to one embodiment of the present invention may be an SZZ algorithm.

[0011] The embedding step according to one embodiment of the present invention can perform embedding considering the context and meaning, the hierarchical structure and meaning information of the source code, and the relationship between the commit message and the code change using a pre-trained embedding model based on the commit data and the code change data.

[0012] The embedding model according to one embodiment of the present invention may be UniXCoder.

[0013] The software timely defect prediction method according to one embodiment of the present invention may further include a preprocessing step of performing tokenization on the commit data and the code change data prior to the embedding step.

[0014] The learning step according to one embodiment of the present invention can learn the context and meaning of the embedded data based on deep learning using a bidirectional learning model capable of learning the relationship between previous data and subsequent data for the embedded data.

[0015] The bidirectional learning model according to one embodiment of the present invention may be a Bi-LSTM model.

[0016] The evaluation step according to one embodiment of the present invention may include a step of generating new learning data that can be expressed by combining the learned commit data and the code change data into one, a classification learning step of performing learning for classification using the new learning data, and a classification learning evaluation step of performing an evaluation on the learning for classification using a loss function.

[0017] The software timely defect prediction method according to one embodiment of the present invention may include a step of outputting the evaluated final data.

[0018] In the software timely defect prediction method according to one embodiment of the present invention, the software may be characterized as being an edge computing application.

[0019] According to embodiments of the present invention, defects in software can be effectively identified during the development stage.

[0020] According to embodiments of the present invention, defects in software can be effectively predicted.

[0021] According to embodiments of the present invention, the quality of software can be effectively improved.

[0022] FIG. 1 is a block diagram showing a software timely defect prediction system according to one embodiment of the present invention.

[0023] FIG. 2 is a flowchart illustrating a software timely defect prediction method according to one embodiment of the present invention.

[0024] FIG. 3 is a conceptual diagram specifically illustrating a part of a software timely defect prediction method according to one embodiment of the present invention.

[0025] FIG. 4 is a conceptual diagram specifically illustrating another part of a software timely defect prediction method according to one embodiment of the present invention.

[0026] A software timely defect prediction method according to an embodiment of the present invention may include a data collection step of collecting data about software, a data labeling step of labeling data with a possibility of defect occurrence using an identification algorithm on the collected data, an embedding step of receiving the labeled data in the data labeling step and performing embedding, a learning step of learning context and meaning for the embedded data in the embedding step based on deep learning, and an evaluation step of evaluating a learning result based on learning data learned in the learning step.

[0027] Hereinafter, with reference to the attached drawings, preferred embodiments of the present invention will be described in detail so that those skilled in the art can easily implement the invention. However, the present invention may be implemented in various different forms and is not limited or restricted by the following examples.

[0028] In order to clearly explain the present invention, a detailed description of a part that is irrelevant to the description or a related known technology that may unnecessarily obscure the gist of the present invention has been omitted, and when adding reference signs to components of each drawing in this specification, the same or similar reference signs are attached to the same or similar components throughout the specification.

[0029] In addition, terms and words used in this specification and claims should not be interpreted as limited to their usual or dictionary meanings, but should be interpreted as meanings and concepts that conform to the technical idea of ​​the present invention based on the principle that the inventor can appropriately define the concept of the term to explain his or her own invention in the best way.

[0030] FIG. 1 is a block diagram showing a software timely defect prediction system according to one embodiment of the present invention.

[0031] The software just-in-time defect prediction system (1) is a system capable of identifying software defects during the software development phase. The software just-in-time defect prediction system (1) can perform just-in-time defect prediction in edge computing applications.

[0032] The software timely defect prediction system (1) may include a data processing module (10).

[0033] The data processing module (10) can collect data for the software from outside and perform labeling on the collected data.

[0034] The software just-in-time defect prediction system (1) may include a JIT (Just In Time) defect prediction module (20).

[0035] The JIT defect prediction module (20) can perform embedding on labeled data, perform deep learning on the embedded data, and evaluate the learning results by deep learning.

[0036] The software timely defect prediction system (1) may include a control module (30).

[0037] The control module (30) can be electrically connected to the data processing module (10) and the JIT defect prediction module (20). The control module (30) can perform control over the data processing module (10) and the JIT defect prediction module (20). It can be understood that the operation steps of the data processing module (10) and the JIT defect prediction module (20) described below are controlled by the control module (30).

[0038] FIG. 2 is a flowchart illustrating a software timely defect prediction method according to one embodiment of the present invention.

[0039] The software timely defect prediction system (1) can identify software defects and perform timely defect prediction through the software timely defect prediction method described below.

[0040] The software timely defect prediction method may include a data collection step according to S100.

[0041] The data collection phase may involve collecting data from an external source. The data may be software-related data. Specifically, the software may be, but is not limited to, an edge computing application.

[0042] The data collection step can be performed by the data processing module (10), but is not limited thereto.

[0043] The software timely defect prediction method may include a data labeling step according to S200.

[0044] The data labeling step may be a step of labeling data that may contain defects using an identification algorithm on the collected data.

[0045] The data labeling step may be performed by the data processing module (10), but is not limited thereto.

[0046] The software timely defect prediction method may include a data embedding step according to S300.

[0047] The embedding step may be a step that receives labeled data from the data labeling step and performs embedding.

[0048] The embedding step may be performed by the JIT defect prediction module (20), but is not limited thereto.

[0049] The software timely defect prediction method may include a deep learning training step according to S400.

[0050] The learning phase may be a phase in which context and meaning for embedded data in the embedding phase are learned based on deep learning.

[0051] The learning phase may be performed by the JIT defect prediction module (20), but is not limited thereto.

[0052] The software timely defect prediction method may include an evaluation step according to S500.

[0053] The evaluation step may be a step that evaluates the learning results based on the learning data learned in the learning step.

[0054] The evaluation step may be performed by the JIT defect prediction module (20), but is not limited thereto.

[0055] More specific operations of the above-described S100 to S500 will be described later with reference to FIGS. 3 and 4.

[0056] FIG. 3 specifically illustrates some of the software timely defect prediction methods (S100 to S200) according to one embodiment of the present invention.

[0057] In the data collection step according to S100, the data processing module (10) can collect commit data from open source autonomous driving software. The data processing module (10) can collect commit messages and code changes using PYDRILLER.

[0058] The data processing module (10) can collect data from autonomous driving software.

[0059] In the data labeling step according to S200, the collected data of the data processing module (10) may include commit data and code change data.

[0060] In the data labeling step, the identification algorithm may be an algorithm that automatically identifies change data that causes defects.

[0061] Specifically, the identification algorithm may include the steps described below.

[0062] The identification algorithm may include a step of searching change data for keywords and identifying commits corresponding to the keywords.

[0063] The identification algorithm may include a step of identifying changed code lines of a previous version and a modified version in a commit corresponding to the keyword.

[0064] The identification algorithm may include a step of generating a change that performs at least one of modifying and deleting a line of code from a previous modification for the last commit of the commits corresponding to the identified line of code.

[0065] The identification algorithm may include a step of labeling commit data for the above changes among the commit data as defective and labeling other changes as non-defective.

[0066] The identification algorithm may be the SZZ algorithm. Specifically, the identification algorithm may be the MA-SZZ algorithm among the SZZ algorithms. A more specific description of the SZZ algorithm may be applied to the general contents of the SZZ algorithm.

[0067] FIG. 4 specifically illustrates another part (S300 to S500) of the software timely defect prediction method according to one embodiment of the present invention.

[0068] File-level defect prediction is typically performed prior to the integration testing phase, while commit-level defect prediction of the JIT defect prediction module (20) may be performed after each file has been changed during the implementation phase. JIT defect prediction of the JIT defect prediction module (20) may be performed at the commit level, which can predict defects at a more granular level than the file level.

[0069] Because commits are typically smaller than files, the amount of code to be inspected can be reduced by identifying defects. Developers can check for defects in their commits while submitting modified code to the repository, reducing the time and effort required to inspect code for defects.

[0070] Commit-level defect prediction can be applied during the implementation phase, and commit-level defect prediction helps developers identify defects in situations where they remember the details of the code, enabling them to predict and fix defects in a timely manner, unlike file-level defect prediction.

[0071] The JIT defect prediction module (20) can be suitable for autonomous driving software by utilizing hierarchical and semantic information (e.g., characteristics of a programming language) of the source code in terms of input data.

[0072] The embedding step according to S300 can be performed based on commit data (e.g., code change in FIG. 4) and code change data (e.g., commit change in FIG. 4), using a pre-trained embedding model to perform embedding that considers context and meaning, the hierarchical structure and semantic information of the source code, and the relationship between commit messages and code changes. Embedding can generate embedding data.

[0073] The embedding model can be UniXCoder, a pre-trained, integrated cross-modal pre-trained model capable of processing both commit messages (natural language) and code changes (programming language).

[0074] The software just-in-time defect prediction method may further include a preprocessing step that performs tokenization on commit data and code change data prior to the embedding step.

[0075] The learning step according to S400 can learn the context and meaning of embedded data based on deep learning by using a bidirectional learning model that can learn the relationship between previous and subsequent data for embedded data.

[0076] A bidirectional learning model can be a Bi-LSTM model.

[0077] A bidirectional learning model can help effectively transmit information about commit messages and code changes to a classifier. The classifier can then perform the evaluation step described below.

[0078] The evaluation steps according to S500 may include the steps described below.

[0079] The evaluation step may include a step of generating new learning data that can be represented by concatenating learned commit data (e.g., message vector in FIG. 4) and code change data (e.g., change vector in FIG. 4).

[0080] The evaluation step may include a classification learning step that performs learning for classification using new learning data.

[0081] The evaluation step may include a classification learning evaluation step that evaluates learning for classification using a loss function (e.g., sigmoid).

[0082] The software timely defect prediction method may include a step of outputting the evaluated final data (result of FIG. 4).

[0083] The above-described software timely defect prediction method and system can be effective in improving the quality of software under development by reducing the effort of developers to inspect or test code.

[0084] In this invention, we collected and labeled autonomous driving software data to perform JIT defect prediction in autonomous driving software, and inserted hierarchical structure and semantic information to perform defect prediction. This significantly improved defect prediction performance. The proposed method outperforms existing machine learning models and cutting-edge approaches, and particularly reduces code inspection effort.

[0085] Specifically, the combination of commit messages and code changes reduces code inspection effort, and the combination of commit-level features can lead to superior performance even on evaluation metrics that do not consider code inspection effort.

[0086] In addition, since the context and semantic learning steps of the method proposed in the present invention improve defect prediction performance, defects in autonomous driving software can be identified from the development stage, thereby contributing to improving software quality.

[0087] Although the present invention has been described above with reference to limited embodiments and drawings, the present invention is not limited thereto, and various embodiments are possible within the scope equivalent to the technical idea of ​​the present invention and the patent claims to be described below by a person having ordinary skill in the art to which the present invention pertains.

[0088] [Explanation of symbols]

[0089] 1: Software Defect Prediction System

[0090] 10: Data Processing Module

[0091] 20: JIT Defect Prediction Module

[0092] 30: Control module

Claims

1. Data collection stage for collecting data about software; A data labeling step that labels data with a possibility of defect occurrence using an identification algorithm on the collected data; An embedding step that receives labeled data from the above data labeling step and performs embedding; A learning step for learning the context and meaning of the embedded data in the above embedding step based on deep learning; and A software timely defect prediction method, comprising an evaluation step for evaluating learning results based on learning data learned in the above learning step.

2. In claim 1, In the above data labeling step, The above collected data is, A software timely defect prediction method including commit data and code change data.

3. In claim 2, In the above data labeling step, The above identification algorithm is, A software timely defect prediction method, which is an algorithm that automatically identifies change data that causes defects.

4. In claim 3, The above identification algorithm is, A step of searching change data by keywords to identify commits corresponding to the keywords; In a commit corresponding to the above keyword, a step of identifying changed code lines of a previous version and a modified version; A step of generating a change that performs at least one of modifying and deleting a code line in a previous modification for the last commit of the commit corresponding to the identified code line; and A software timely defect prediction method, comprising the step of labeling commit data for the above changes among commit data as a defect and labeling other changes as defect-free.

5. In claim 3, The above identification algorithm is, SZZ algorithm, a software just-in-time defect prediction method.

6. In claim 2, The above embedding step is, A software timely defect prediction method, which performs embedding considering context and meaning, hierarchical structure and meaning information of source code, and relationships between commit messages and code changes using a pre-trained embedding model based on the above commit data and the above code change data.

7. In claim 6, The above embedding model is, UniXCoder, a software just-in-time defect prediction method.

8. In claim 6, A software timely defect prediction method further comprising a preprocessing step of performing tokenization on the commit data and the code change data prior to the embedding step.

9. In claim 6, The above learning steps are: A software timely defect prediction method that learns the context and meaning of embedded data based on deep learning by using a bidirectional learning model capable of learning the relationship between previous and subsequent data for embedded data.

10. In claim 9, The above two-way learning model is, A software timely defect prediction method using the Bi-LSTM model.

11. In claim 2, The above evaluation steps are: A step of generating new learning data that can be expressed by combining the learned commit data and the code change data into one; A classification learning step for performing learning for classification using the above new learning data; A software timely defect prediction method, comprising a classification learning evaluation step for evaluating learning for the classification using a loss function.

12. In claim 1, A software timely defect prediction method, comprising a step of outputting the evaluated final data.

13. In claim 1, The above software, A software timely defect prediction method characterized by being an edge computing application.