Software development process management method and system based on artificial intelligence
By combining the WordPiece word segmenter and the bidirectional long and short-term gated loop model of context attention, the problem of poor accuracy of code defect detection in the existing technology is solved, more efficient and more accurate code defect detection is achieved, and the automation and intelligence of the software development process is improved.
Patent Information
- Application Number
- CN202510311819.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
AI Technical Summary
In the existing software development process management, the accuracy of code defect detection is poor, mainly because traditional methods only analyze from a single source code data, ignore the potential information of expert indicators, and lack the understanding of code context semantics, making it difficult to capture long-term dependencies in complex code.
The data conversion method combined with WordPiece word segmenter is adopted to convert expert indicators and source code data into feature representations suitable for model input, and a two-way long-term gated loop model is used for code defect detection, accurately capturing the context semantics and long-term dependencies of complex codes.
Improve the accuracy and efficiency of code defect detection, and more accurately detect defects caused by the code context environment, thus helping the software development team continuously improve code quality.
Smart Images

Figure CN120216374A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of software development management, and specifically refers to a software development process management method and system based on artificial intelligence. Background Art
[0002] For software development process management, artificial intelligence technology is introduced in each stage of software development (including requirements analysis, design, coding, testing, etc.). Through data-driven and automated analysis, the software development process is managed, aiming to timely discover potential defects existing in the software development process, assist the software development team to quickly respond to potential defects, and then improve the software development efficiency, and realize the automation, intelligence and high efficiency of the software development process.
[0003] However, in the existing software development process management, there are technical problems as follows: software instability is mainly related to code defects. Timely discovering code defects can significantly improve software quality. However, most of the existing code defect detections only analyze from single source code data, ignoring the potential information represented by expert metrics, resulting in poor accuracy of code defect detection; manual inspection of code defects is not only time-consuming but also prone to omissions. Moreover, traditional code defect detection models often lack the understanding of code context semantics when dealing with complex codes, and are prone to ignoring long-term dependencies, resulting in the model being unable to detect defects caused by code context environments, thus affecting the accuracy and efficiency of code defect detection. Summary of the Invention
[0004] In view of the above situation, to overcome the defects of the prior art, the present invention provides a software development process management method and system based on artificial intelligence. In the existing software development process management, software instability is mainly related to code defects. Timely detection of code defects can significantly improve software quality. However, most existing code defect detections only analyze from single source code data, ignoring the potential information represented by expert metrics, resulting in poor accuracy of code defect detection. The present solution creatively adopts a data conversion method combining the WordPiece tokenizer to convert expert metrics and source code data into feature representations suitable for model input, fully integrating the two types of data information, providing more comprehensive and multi-dimensional input data for subsequent code defect detection, thereby improving the accuracy of code defect detection. In the existing software development process management, manual inspection of code defects is not only time-consuming but also prone to omissions. Traditional code defect detection models often lack the understanding of code context semantics when dealing with complex code and are prone to ignoring long-term dependencies, resulting in the model being unable to detect defects caused by the code context environment, thus affecting the accuracy and efficiency of code defect detection. The present solution creatively adopts a context attention bidirectional long short-term memory recurrent neural network model for code defect detection, accurately capturing the context semantics and long-term dependencies of complex code, more accurately and efficiently capturing potential defects in complex code, improving the accuracy and efficiency of code defect detection, and thus helping software development teams continuously improve code quality during the development process.
[0005] The technical solution adopted by the present invention is as follows: A software development process management method based on artificial intelligence provided by the present invention includes the following steps:
[0006] Step S1: Data preparation;
[0007] Step S2: Data optimization processing;
[0008] Step S3: Data conversion;
[0009] Step S4: Code defect detection;
[0010] Step S5: Software development process management.
[0011] Further, in step S1, the data preparation is used to prepare the original data required for software development process management. Specifically, source code data is collected from software project documents, and the source code data includes defective source code data and defect-free source code data;
[0012] The defect types of the defective source code data include attribute errors, type errors, value errors, key errors, index errors, runtime errors, and other errors.
[0013] Further, in step S2, the data optimization process is used to optimize the quality of the original data and extract features. Specifically, the source code data is processed through data cleaning and code metrics to obtain source code optimized data and expert metric description text data, including the following steps:
[0014] Step S21: Data cleaning, specifically, by removing noise, imputing missing values, and filtering duplicate values from the source code data, source code optimized data is obtained;
[0015] Step S22: Code metrics, used to calculate expert metrics and textify them, including the following steps:
[0016] Step S221: Original metrics, specifically, by calculating the number of source code lines, logical code lines, and comment lines, original metrics are performed;
[0017] Step S222: Halstead metrics, used to evaluate the overall quality of the source code. Specifically, by calculating the program vocabulary, program length, program volume, program difficulty, and potential defect count, Halstead metrics are performed;
[0018] Step S223: Calculate the cyclomatic complexity;
[0019] Step S224: Evaluate the maintainability index;
[0020] Step S225: Based on the source code optimized data, through the original metrics, the Halstead metrics, the calculated cyclomatic complexity, and the evaluated maintainability index, code metrics are performed to obtain expert metrics;
[0021] Step S226: Textify the metric results. Specifically, through a rule-based text generation method, the expert metrics are converted into text descriptions to obtain expert metric description text data;
[0022] The expert metrics include the number of source code lines, logical code lines, comment lines, program vocabulary, program length, program volume, program difficulty, potential defect count, cyclomatic complexity, and maintainability index.
[0023] Further, in step S3, the data conversion is used to convert the expert metric description text data and the source code optimized data into a feature representation suitable for input into the model. Specifically, based on the expert metric description text data and the source code optimized data, a data conversion method combining a WordPiece tokenizer is adopted to obtain a source code comprehensive feature sequence, including the following steps:
[0024] Step S31: Build a tokenization model. Specifically, build a WordPiece tokenizer, including the following steps:
[0025] Step S311: Initialize the vocabulary, specifically by decomposing the input data of the WordPiece tokenizer into sub-words to obtain the vocabulary;
[0026] Step S312: Vocabulary merging, specifically by counting the co-occurrence frequencies of sub-word pairs in the vocabulary and selecting the sub-word pair with the highest co-occurrence frequency to merge into a sub-word unit;
[0027] Step S313: Update the vocabulary, specifically by repeatedly performing the vocabulary merging until the size of the vocabulary reaches a predetermined value to obtain the final vocabulary;
[0028] Step S32: Tokenization processing, specifically by using the WordPiece tokenizer to perform tokenization processing on the expert index description text data and the source code optimization data respectively to obtain the expert index feature sequence and the source code feature sequence;
[0029] Step S33: Generation of the source code comprehensive feature sequence, specifically by combining the expert index feature sequence and the source code feature sequence to generate the source code comprehensive feature sequence, and the calculation formula is:
[0030] ;
[0031] In the formula, X is the source code comprehensive feature sequence, [CLS] is the sequence start marker, Fea ind is the expert index feature sequence, [SEP] is the sequence separator, Fea code is the source code feature sequence, [EOS] is the sequence end marker, and concat(·) is the concatenation operation.
[0032] Furthermore, in step S4, the code defect detection is used to detect the defects existing in the code, specifically by using the context attention bidirectional long short-term gated recurrent model based on the source code comprehensive feature sequence to perform code defect detection to obtain the code defect information;
[0033] The context attention bidirectional long short-term gated recurrent model includes an input embedding layer, a bidirectional long short-term memory layer, a gated recurrent unit layer, a context attention block, a convolutional pooling layer, and a classification output layer;
[0034] The bidirectional long short-term memory layer is used to capture the context dependency relationship of the source code comprehensive feature sequence;
[0035] The gated recurrent unit layer is used to process the long-term dependency relationship in the source code comprehensive feature sequence and improve the model calculation efficiency;
[0036] The context attention block is used to strengthen the attention to the most relevant regions for code defect detection;
[0037] The code defect detection includes the following steps:
[0038] Step S41: Construct an input embedding layer. Specifically, through word embedding processing and positional embedding processing on the comprehensive feature sequence of the source code, a comprehensive embedding vector of the source code is obtained. The calculation formula is:
[0039] ;
[0040] In the formula, Em i is the i-th comprehensive embedding vector of the source code, specifically referring to the comprehensive embedding vector corresponding to the i-th element in the comprehensive feature sequence of the source code. i is the element index, specifically referring to the element index in the comprehensive feature sequence of the source code. WordEmbedding(·) is word embedding processing, and X i is the i-th element in the comprehensive feature sequence of the source code, and PositionalEmbedding(·) is positional embedding processing;
[0041] The calculation formula for the positional embedding processing is:
[0042] ;
[0043] ;
[0044] In the formula, p i,0 is the 0-th dimensional position encoding of the i-th element in the comprehensive feature sequence of the source code, p i,1 is the 1-st dimensional position encoding of the i-th element in the comprehensive feature sequence of the source code, is the d emb-1 -th dimensional position encoding of the i-th element in the comprehensive feature sequence of the source code, d emb is the embedding dimension, k is the embedding dimension index, is the k-th dimensional position encoding of the i-th element in the comprehensive feature sequence of the source code, sin(·) is the sine function, cos(·) is the cosine function, is the modulo operator, indicates that k is even, indicates that k is odd;
[0045] Step S42: Construct a bidirectional long short-term memory layer. Specifically, in the bidirectional long short-term memory layer, the comprehensive embedding vector of the source code is processed by the forward long short-term memory unit and the backward long short-term memory unit respectively, and then the output results of the forward long short-term memory unit and the backward long short-term memory unit are feature concatenated to obtain the bidirectional long short-term feature of the source code;
[0046] Step S43: Construct a gated recurrent unit layer. Specifically, an input gate, a forget gate, and a reset gate are set in the gated recurrent unit layer. The source code bidirectional long-term and short-term features are processed through the gated recurrent unit layer to obtain source code dependency features;
[0047] Step S44: Construct a context attention block, including the following steps:
[0048] Step S441: Concatenate the source code bidirectional long-term and short-term features and the source code dependency features to obtain source code bidirectional long-term and short-term dependency features;
[0049] Step S442: Calculate the context attention weights. The calculation formula is:
[0050] ;
[0051] In the formula, is the context attention weight at the t-th time step, t is the first index of the time step, j is the second index of the time step, exp(·) is the natural exponential function, score(·) is the attention score calculation function, h t is the source code bidirectional long-term and short-term dependency feature at the t-th time step, h j is the source code bidirectional long-term and short-term dependency feature at the j-th time step, and T is the maximum time step;
[0052] Step S443: Calculate the source code context features. The calculation formula is:
[0053] ;
[0054] In the formula, code context is the source code context feature;
[0055] Step S45: Construct a convolutional pooling layer. Specifically, in the convolutional pooling layer, convolutional operations are performed on the source code context features, and then max pooling and average pooling operations are performed to obtain source code global semantic features;
[0056] Step S46: Construct a classification output layer. Specifically, in the classification output layer, the source code global semantic features are classified through a dense layer and a Softmax activation function to obtain the model detection results;
[0057] Step S47: Model training. Specifically, through the construction of the input embedding layer, the construction of the bidirectional long short-term memory layer, the construction of the gated recurrent unit layer, the construction of the context attention block, the construction of the convolutional pooling layer, and the construction of the classification output layer, a context attention bidirectional long short-term gated recurrent model is constructed, and the context attention bidirectional long short-term gated recurrent model is trained to obtain a code defect detection model;
[0058] Step S48: Calculate the code defect detection result. Perform code defect detection through code defect detection to obtain code defect information, where the code defect information includes the code defect type, defect severity, and defect confidence level.
[0059] Further, in step S5, the software development process management is specifically as follows: Whenever a software developer submits code, code defect detection is performed, and a code defect report is generated and fed back to the software developer, thereby assisting the software developer to promptly repair code defects and improve software development efficiency and code quality.
[0060] A software development process management system based on artificial intelligence provided by the present invention includes: a data preparation module, a data optimization processing module, a data conversion module, a code defect detection module, and a software development process management module;
[0061] The data preparation module is used for data preparation. Through data preparation, source code data is obtained and sent to the data optimization processing module;
[0062] The data optimization processing module is used for data optimization processing. Through data optimization processing, source code optimized data and expert index description text data are obtained and sent to the data conversion module;
[0063] The data conversion module is used for data conversion. Through data conversion, a source code comprehensive feature sequence is obtained and sent to the code defect detection module;
[0064] The code defect detection module is used for code defect detection. Through code defect detection, code defect information is obtained and sent to the software development process management module;
[0065] The software development process management module is used for software development process management. Through software development process management, a code defect report is generated.
[0066] The beneficial effects achieved by the present invention using the above solution are as follows:
[0067] (1) In the existing software development process management, software instability is mainly related to code defects. Timely detection of code defects can significantly improve software quality. However, most existing code defect detections only analyze from single-source code data, ignoring the potential information represented by expert metrics, resulting in poor accuracy of code defect detection. This solution creatively adopts a data conversion method combining the WordPiece tokenizer to convert expert metrics and source code data into feature representations suitable for model input, fully integrating the information of the two types of data, providing more comprehensive and multi-dimensional input data for subsequent code defect detection, thereby improving the accuracy of code defect detection.
[0068] (2) In the existing software development process management, manual inspection of code defects is not only time-consuming but also prone to omissions. Traditional code defect detection models often lack the understanding of code context semantics when dealing with complex code and are prone to ignoring long-term dependencies, resulting in the model being unable to detect defects caused by the code context environment, thereby affecting the accuracy and efficiency of code defect detection. This solution creatively adopts a context attention bidirectional long short-term memory (LSTM) model for code defect detection, accurately capturing the context semantics and long-term dependencies of complex code, more accurately and efficiently capturing potential defects in complex code, improving the accuracy and efficiency of code defect detection, and thus helping software development teams continuously improve code quality during the development process. Brief Description of the Drawings
[0069] Figure 1 It is a flowchart of a software development process management method based on artificial intelligence provided by the present invention;
[0070] Figure 2 It is a schematic diagram of a software development process management system based on artificial intelligence provided by the present invention;
[0071] Figure 3 It is a flowchart of step S3;
[0072] Figure 4 It is a flowchart of step S4.
[0073] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. Detailed Embodiments
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0075] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0076] Embodiment 1, refer to Figure 1 , a software development process management method based on artificial intelligence provided by the present invention, the method includes the following steps:
[0077] Step S1: Data preparation;
[0078] Step S2: Data optimization processing;
[0079] Step S3: Data conversion;
[0080] Step S4: Code defect detection;
[0081] Step S5: Software development process management.
[0082] Embodiment 2, refer to Figure 1 , this embodiment is based on the above embodiment. In step S1, the data preparation is used to prepare the original data required for software development process management. Specifically, source code data is collected from software project documents, and the source code data includes defective source code data and defect-free source code data;
[0083] The defect types of the defective source code data include attribute errors, type errors, value errors, key errors, index errors, runtime errors, and other errors.
[0084] Embodiment 3, refer to Figure 1 , this embodiment is based on the above embodiment. In step S2, the data optimization processing is used to optimize the quality of the original data and extract features. Specifically, data optimization processing is performed on the source code data through data cleaning and code metrics to obtain source code optimization data and expert index description text data, including the following steps:
[0085] Step S21: Data cleaning, specifically, by removing noise, imputing missing values, and filtering duplicate values from the source code data, optimized source code data is obtained;
[0086] Step S22: Code metrics, used to calculate expert metrics and textify them, including the following steps:
[0087] Step S221: Original metrics, specifically, by calculating the number of source code lines, logical code lines, and comment lines, original metrics are performed;
[0088] Step S222: Halstead metrics, used to evaluate the overall quality of the source code. Specifically, by calculating the program vocabulary, program length, program volume, program difficulty, and potential defect number, Halstead metrics are performed;
[0089] The calculation formula for the program vocabulary:
[0090] ;
[0091] where N vo is the program vocabulary, n1 is the number of operators in the program, and n2 is the number of operands in the program;
[0092] The calculation formula for the program length:
[0093] ;
[0094] where N L is the program length, N1 is the total number of operators, and log2(·) is the logarithmic function with base 2;
[0095] The calculation formula for the program volume:
[0096] ;
[0097] where Vol is the program volume;
[0098] The calculation formula for the program difficulty:
[0099] ;
[0100] where Diff is the program difficulty and N2 is the total number of operands;
[0101] The calculation formula for the potential defect number:
[0102] ;
[0103] where Bugs is the potential defect number;
[0104] Step S223: Calculate the cyclomatic complexity and evaluate the source code complexity by comments. Specifically, convert the source code in the source code optimization data into a control flow graph, and calculate the cyclomatic complexity based on the structure of the control flow graph. The calculation formula is:
[0105] ;
[0106] In the formula, Cyc is the cyclomatic complexity, E cyc is the number of edges in the control flow graph, N cyc is the number of nodes in the control flow graph, P cyc is the number of connected branches of the source code;
[0107] Step S224: Evaluate the maintainability index. The calculation formula is:
[0108] ;
[0109] In the formula, MI is the maintainability index, ln(·) is the natural logarithm function, and S is the number of source code lines;
[0110] Step S225: Based on the source code optimization data, perform code metrics through the original metrics, the Halstead metrics, the calculated cyclomatic complexity, and the evaluated maintainability index to obtain expert metrics;
[0111] Step S226: Textualize the metric results. Specifically, convert the expert metrics into text descriptions through a rule-based text generation method to obtain expert metric description text data;
[0112] The expert metrics include the number of source code lines, the number of logical code lines, the number of comment lines, the program vocabulary, the program length, the program volume, the program difficulty, the number of potential defects, the cyclomatic complexity, and the maintainability index.
[0113] Example 4, refer to Figure 1 and Figure 3 , this example is based on the above example. In step S3, the data conversion is used to convert the expert metric description text data and the source code optimization data into a feature representation suitable for input to the model. Specifically, based on the expert metric description text data and the source code optimization data, a data conversion method combining the WordPiece tokenizer is adopted to obtain a comprehensive source code feature sequence, including the following steps:
[0114] Step S31: Build a tokenization model. Specifically, build a WordPiece tokenizer, including the following steps:
[0115] Step S311: Initialize the vocabulary. Specifically, obtain the vocabulary by decomposing the input data of the WordPiece tokenizer into subwords;
[0116] Step S312: Lexical merging, specifically, by counting the co-occurrence frequencies of sub-word pairs in the vocabulary, selecting the sub-word pair with the highest co-occurrence frequency and merging them into a sub-word unit;
[0117] Step S313: Update the vocabulary, specifically, repeat the above-mentioned lexical merging until the size of the vocabulary reaches a predetermined value to obtain the final vocabulary;
[0118] Step S32: Word segmentation processing, specifically, use the WordPiece tokenizer to perform word segmentation processing on the expert index description text data and the source code optimization data respectively to obtain the expert index feature sequence and the source code feature sequence;
[0119] Step S33: Generation of the source code comprehensive feature sequence, specifically, combine the expert index feature sequence and the source code feature sequence to generate the source code comprehensive feature sequence, and the calculation formula is:
[0120] ;
[0121] In the formula, X is the source code comprehensive feature sequence, [CLS] is the sequence start marker, Fea ind is the expert index feature sequence, [SEP] is the sequence separator, Fea code is the source code feature sequence, [EOS] is the sequence end marker, and concat(·) is the concatenation operation used to concatenate the sequence start marker, the expert index feature sequence, the sequence separator, the source code feature sequence, and the sequence end marker;
[0122] By performing the above operations, in the existing software development process management, software instability is mainly related to code defects. Timely detection of code defects can significantly improve software quality. However, most existing code defect detections only analyze from single source code data, ignoring the potential information represented by expert indexes, resulting in poor accuracy of code defect detection. This solution creatively adopts a data conversion method combining the WordPiece tokenizer to convert expert indexes and source code data into feature representations suitable as model inputs, fully combines the information of the two types of data, provides more comprehensive and multi-dimensional input data for subsequent code defect detection, and thus improves the accuracy of code defect detection.
[0123] Example Five, refer to Figure 1 and Figure 4 Based on the above example, in step S4, the code defect detection is used to detect the defects existing in the code. Specifically, based on the source code comprehensive feature sequence, a context attention bidirectional long short-term gated recurrent model is adopted to perform code defect detection to obtain code defect information;
[0124] The context attention bidirectional long short-term gated recurrent model includes an input embedding layer, a bidirectional long short-term memory layer, a gated recurrent unit layer, a context attention block, a convolutional pooling layer, and a classification output layer;
[0125] The bidirectional long short-term memory layer is used to capture the context dependencies of the source code comprehensive feature sequence;
[0126] The gated recurrent unit layer is used to process the long-term dependencies in the source code comprehensive feature sequence and improve the model calculation efficiency;
[0127] The context attention block is used to strengthen the attention to the regions most relevant to code defect detection;
[0128] The code defect detection includes the following steps:
[0129] Step S41: Construct an input embedding layer, specifically by performing word embedding processing on the source code comprehensive feature sequence and performing positional embedding processing to obtain the source code comprehensive embedding vector. The calculation formula is:
[0130] ;
[0131] In the formula, Em i is the i-th source code comprehensive embedding vector, specifically referring to the source code comprehensive embedding vector corresponding to the i-th element in the source code comprehensive feature sequence. i is the element index, specifically referring to the element index in the source code comprehensive feature sequence. WordEmbedding(·) is the word embedding processing, and X i is the i-th element in the source code comprehensive feature sequence, and PositionalEmbedding(·) is the positional embedding processing;
[0132] The calculation formula for the positional embedding processing is:
[0133] ;
[0134] ;
[0135] In the formula, p i,0 is the 0-th dimensional position encoding of the i-th element in the source code comprehensive feature sequence, p i,1 is the 1-st dimensional position encoding of the i-th element in the source code comprehensive feature sequence, is the d emb-1 -th dimensional position encoding of the i-th element in the source code comprehensive feature sequence, d emb is the embedding dimension, k is the embedding dimension index, is the k-th dimensional position encoding of the i-th element in the code comprehensive feature sequence, sin(·) is the sine function, and cos(·) is the cosine function, is the modulo operator, indicating that k is even, indicating that k is odd;
[0136] Step S42: Construct a bidirectional long short-term memory layer. Specifically, in the bidirectional long short-term memory layer, the source code comprehensive embedding vector is processed by the forward long short-term memory unit and the backward long short-term memory unit respectively, and then the output results of the forward long short-term memory unit and the backward long short-term memory unit are feature concatenated to obtain the source code bidirectional long short-term features;
[0137] Step S43: Construct a gated recurrent unit layer. Specifically, an input gate, a forget gate, and a reset gate are set in the gated recurrent unit layer, and the source code bidirectional long short-term features are processed by the gated recurrent unit layer to obtain the source code dependency features;
[0138] Step S44: Construct a context attention block, including the following steps:
[0139] Step S441: Feature concatenate the source code bidirectional long short-term features and the source code dependency features to obtain the source code bidirectional long short-term dependency features;
[0140] Step S442: Calculate the context attention weights, and the calculation formula is:
[0141] ;
[0142] In the formula, is the context attention weight at the t-th time step, t is the first index of the time step, j is the second index of the time step, exp(·) is the natural exponential function, score(·) is the attention score calculation function, h t is the source code bidirectional long short-term dependency feature at the t-th time step, h j is the source code bidirectional long short-term dependency feature at the j-th time step, and T is the maximum time step;
[0143] Step S443: Calculate the source code context features, and the calculation formula is:
[0144] ;
[0145] In the formula, code context is the source code context feature;
[0146] Step S45: Construct a convolutional pooling layer. Specifically, in the convolutional pooling layer, the source code context features are convolved, and then max pooling and average pooling operations are performed to obtain the source code global semantic features;
[0147] Step S46: Construct a classification output layer. Specifically, in the classification output layer, classify the global semantic features of the source code through a dense layer and a Softmax activation function to obtain the model detection result;
[0148] Step S47: Model training. Specifically, through the construction of the input embedding layer, the construction of the bidirectional long short-term memory layer, the construction of the gated recurrent unit layer, the construction of the context attention block, the construction of the convolutional pooling layer, and the construction of the classification output layer, construct a context attention bidirectional long short-term gated recurrent model, and perform model training on the context attention bidirectional long short-term gated recurrent model to obtain a code defect detection model;
[0149] Step S48: Calculate the code defect detection result. Perform code defect detection through code defect detection to obtain code defect information, where the code defect information includes the code defect type, defect severity, and defect confidence;
[0150] By performing the above operations, in the existing software development process management, manual inspection of code defects is not only time-consuming but also prone to omissions. Traditional code defect detection models often lack the understanding of code context semantics when dealing with complex code and are prone to ignoring long-term dependencies, resulting in the model being unable to detect defects caused by the code context environment, thereby affecting the accuracy and efficiency of code defect detection. The present solution creatively uses a context attention bidirectional long short-term gated recurrent model for code defect detection, accurately captures the context semantics and long-term dependencies of complex code, more accurately and efficiently captures potential defects in complex code, improves the accuracy and efficiency of code defect detection, and thus helps software development teams continuously improve code quality during the development process.
[0151] Embodiment Six. Refer to Figure 1 , based on the above embodiment, in step S5, the software development process management is specifically that whenever a software developer submits code, code defect detection is performed, and a code defect report is generated and fed back to the software developer, thereby assisting the software developer to repair code defects in a timely manner and improving software development efficiency and code quality.
[0152] Embodiment Seven. Refer to Figure 2 , based on the above embodiment, an artificial intelligence-based software development process management system provided by the present invention includes: a data preparation module, a data optimization processing module, a data conversion module, a code defect detection module, and a software development process management module;
[0153] The data preparation module is used for data preparation. Through data preparation, source code data is obtained, and the source code data is sent to the data optimization processing module;
[0154] The data optimization processing module is used for data optimization processing. Through data optimization processing, optimized source code data and expert index description text data are obtained, and the optimized source code data and the expert index description text data are sent to the data conversion module;
[0155] The data conversion module is used for data conversion. Through data conversion, a comprehensive source code feature sequence is obtained, and the comprehensive source code feature sequence is sent to the code defect detection module;
[0156] The code defect detection module is used for code defect detection. Through code defect detection, code defect information is obtained, and the code defect information is sent to the software development process management module;
[0157] The software development process management module is used for software development process management. Through software development process management, a code defect report is generated.
[0158] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0159] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention.
[0160] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A software development process management method based on artificial intelligence, characterized by: The method comprises the following steps: Step S1: Data preparation, used to prepare the original data required for software development process management and obtain source code data; Step S2: data optimization processing, specifically, performing data optimization processing on the source code data through data cleaning and code measurement to obtain source code optimization data and expert indicator description text data; Step S3: data conversion, used to convert the expert indicator description text data and the source code optimization data into feature representations suitable for the input model, specifically, based on the expert indicator description text data and the source code optimization data, a data conversion method combined with a WordPiece word segmenter is used to obtain a source code comprehensive feature sequence; Step S4: code defect detection, which is used to detect defects in the code. Specifically, based on the comprehensive feature sequence of the source code, a contextual attention bidirectional long-short term gated loop model is used to perform code defect detection to obtain code defect information; The contextual attention bidirectional long short-term gated recurrent model includes an input embedding layer, a bidirectional long short-term memory layer, a gated recurrent unit layer, a contextual attention block, a convolutional pooling layer, and a classification output layer; The bidirectional long short-term memory layer is used to capture the context dependency of the source code comprehensive feature sequence; The gated recurrent unit layer is used to process long-term dependencies in the source code comprehensive feature sequence and improve the model calculation efficiency; The context attention block is used to strengthen the attention on the most relevant area of code defect detection; Step S5: Software development process management.
2. The software development process management method based on artificial intelligence according to claim 1, characterized in that: In step S4, the code defect detection includes the following steps: Step S41: construct an input embedding layer, specifically by performing word embedding processing on the source code comprehensive feature sequence and position embedding processing to obtain a source code comprehensive embedding vector. The calculation formula is: ; In the formula, Em i is the i-th source code comprehensive embedding vector, specifically the source code comprehensive embedding vector corresponding to the i-th element in the source code comprehensive feature sequence, i is the element index, specifically the element index in the source code comprehensive feature sequence, WordEmbedding(·) is the word embedding processing, X i is the i-th element in the source code comprehensive feature sequence, PositionalEmbedding(·) is the position embedding process; The calculation formula for the position embedding process is: ; ; In the formula, p i,0 is the 0th dimension position code of the ith element in the source code comprehensive feature sequence, p i,1 is the first dimension position code of the ith element in the source code comprehensive feature sequence, is the dth element of the i-th element in the source code comprehensive feature sequence emb-1 dimensional position encoding, d emb is the embedding dimension, k is the embedding dimension index, is the k-th dimension position code of the ith element in the code comprehensive feature sequence, sin(·) is the sine function, cos(·) is the cosine function, is the remainder operator, means k is an even number, Indicates that k is an odd number; Step S42: constructing a bidirectional long short-term memory layer, specifically, in the bidirectional long short-term memory layer, processing the source code comprehensive embedding vector respectively through the forward long short-term memory unit and the backward long short-term memory unit, and then performing feature splicing on the output results of the forward long short-term memory unit and the backward long short-term memory unit to obtain the source code bidirectional long short-term feature; Step S43: constructing a gated recurrent unit layer, specifically setting an input gate, a forget gate, and a reset gate in the gated recurrent unit layer, and processing the source code bidirectional long-term and short-term features through the gated recurrent unit layer to obtain source code dependency features; Step S44: construct a context attention block, including the following steps: Step S441: concatenating the source code bidirectional long-term and short-term features and the source code dependency features to obtain the source code bidirectional long-term and short-term dependency features; Step S442: Calculate the context attention weight, the calculation formula is: ; In the formula, is the context attention weight of the t-th time step, t is the first index of the time step, j is the second index of the time step, exp(·) is the natural exponential function, score(·) is the attention score calculation function, and h t is the bidirectional long-term and short-term dependency feature of the source code at the tth time step, h j is the bidirectional long-term and short-term dependency feature of the source code at the jth time step, and T is the maximum time step; Step S443: Calculate the source code context feature, the calculation formula is: ; In the formula, code context is the source code context feature; Step S45: constructing a convolutional pooling layer, specifically, in the convolutional pooling layer, performing a convolution operation on the source code context features, and then performing maximum pooling and average pooling operations to obtain the source code global semantic features; Step S46: constructing a classification output layer, specifically, in the classification output layer, classifying the global semantic features of the source code through a dense layer and a Softmax activation function to obtain a model detection result; Step S47: model training, specifically, constructing a contextual attention bidirectional long-short-term gated recurrent model by constructing an input embedding layer, constructing a bidirectional long-short-term memory layer, constructing a gated recurrent unit layer, constructing a contextual attention block, constructing a convolutional pooling layer, and constructing a classification output layer, and performing model training on the contextual attention bidirectional long-short-term gated recurrent model to obtain a code defect detection model; Step S48: Calculate the code defect detection result, perform code defect detection through code defect detection, and obtain code defect information, wherein the code defect information includes code defect type, defect severity and defect confidence.
3. The software development process management method based on artificial intelligence according to claim 2, characterized in that: In step S3, the data conversion is used to convert the expert indicator description text data and the source code optimization data into a feature representation suitable for the input model, specifically, based on the expert indicator description text data and the source code optimization data, a data conversion method combined with a WordPiece word segmenter is used to obtain a source code comprehensive feature sequence, including the following steps: Step S31: construct a word segmentation model, specifically constructing a WordPiece word segmenter, including the following steps: Step S311: Initialize the vocabulary, specifically by decomposing the input data of the WordPiece word segmenter into subwords to obtain the vocabulary; Step S312: vocabulary merging, specifically, by counting the co-occurrence frequencies of subword pairs in the vocabulary table, selecting the subword pairs with the highest co-occurrence frequency and merging them into subword units; Step S313: updating the vocabulary, specifically, repeatedly performing the vocabulary merging until the vocabulary size reaches a predetermined value, thereby obtaining a final vocabulary; Step S32: word segmentation processing, specifically, using the WordPiece word segmenter to perform word segmentation processing on the expert indicator description text data and the source code optimization data respectively, to obtain an expert indicator feature sequence and a source code feature sequence; Step S33: Generate a source code comprehensive feature sequence, specifically, combine the expert indicator feature sequence and the source code feature sequence to generate a source code comprehensive feature sequence, and the calculation formula is: ; In the formula, X is the source code comprehensive feature sequence, [CLS] is the sequence start marker, Fea ind is the expert indicator feature sequence, [SEP] is the sequence separator, Fea code is a source code feature sequence, [EOS] is the sequence terminator, and concat(·) is a concatenation operation.
4. The software development process management method based on artificial intelligence according to claim 3 is characterized in that: In step S2, the data optimization process is used to optimize the quality of the original data and extract features, specifically, the source code data is optimized by data cleaning and code measurement to obtain source code optimization data and expert indicator description text data, including the following steps: Step S21: data cleaning, specifically, removing noise, interpolating missing values, and filtering duplicate values from the source code data to obtain source code optimization data; Step S22: Code metrics, used to calculate expert indicators and textualize them, includes the following steps: Step S221: original measurement, specifically, performing original measurement by calculating the number of source code lines, the number of logic code lines, and the number of comment lines; Step S222: Halstead metric is used to evaluate the overall quality of the source code, specifically by calculating the program vocabulary, program length, program volume, program difficulty and number of potential defects to perform Halstead metric; Step S223: Calculate cyclomatic complexity and evaluate the complexity of source code; Step S224: Evaluate the maintainability index; Step S225: According to the source code optimization data, code metrics are performed through the original metrics, the Halstead metrics, the calculation cyclomatic complexity and the evaluation maintainability index to obtain expert indicators; Step S226: Textualizing the measurement results, specifically, converting the expert indicators into text descriptions through a rule-based text generation method to obtain expert indicator description text data.
5. The software development process management method based on artificial intelligence according to claim 4, characterized in that: In step S1, the data preparation is used to prepare the original data required for software development process management, specifically collecting source code data from software project documents, the source code data including defective source code data and non-defective source code data; The defect types of the defective source code data include attribute errors, type errors, value errors, key errors, index errors, runtime errors and other errors.
6. The software development process management method based on artificial intelligence according to claim 5, characterized in that: In step S5, the software development process management is specifically to perform code defect detection and generate a code defect report every time the software developer submits the code, and feed it back to the software developer.
7. A software development process management system based on artificial intelligence, used to implement a software development process management method based on artificial intelligence as described in any one of claims 1 to 6, characterized in that: include: Data preparation module, data optimization processing module, data conversion module, code defect detection module and software development process management module.
8. The software development process management system based on artificial intelligence according to claim 7, characterized in that: The data preparation module is used for data preparation, obtains source code data through data preparation, and sends the source code data to the data optimization processing module; The data optimization processing module is used for data optimization processing, and obtains source code optimization data and expert indicator description text data through data optimization processing, and sends the source code optimization data and the expert indicator description text data to the data conversion module; The data conversion module is used for data conversion, obtains a source code comprehensive feature sequence through data conversion, and sends the source code comprehensive feature sequence to the code defect detection module; The code defect detection module is used for code defect detection, obtains code defect information through code defect detection, and sends the code defect information to the software development process management module; The software development process management module is used for software development process management, and generates a code defect report through software development process management.