A Method for Deep Neural Network Type Inference Based on Program Slicing

Through the deep neural network type derivation method based on program slices, the problems of low efficiency and low accuracy of dynamic language program type derivation are solved, efficient and accurate type derivation is achieved, and the readability and maintainability of software engineering are improved.

CN114580641BActive Publication Date: 2025-07-25NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210052087.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-17
Publication Date
2025-07-25
Estimated Expiration
2042-01-17

AI Technical Summary

Technical Problem

The existing dynamic language program type derivation methods are inefficient and have low accuracy, especially in large-scale projects, manual labeling is time-consuming and variable type derivation accuracy is not high, and the context information of variables or functions is not fully utilized.

Method used

The deep neural network type derivation method based on program slice is adopted, and the data set is constructed by obtaining items with type label information, using coding technology to embed information into vector form, training deep neural network models, predicting program variable types or function signatures, and combining program slice technology and deep learning technology.

Benefits of technology

It significantly improves the accuracy and efficiency of type derivation of dynamic language programs, reduces the time for developers and maintains project codes for understanding and maintaining projects, and improves the readability, comprehensibility and maintainability of software engineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580641B_ABST
    Figure CN114580641B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for inferring the type of a deep neural network based on program slicing. First, a large number of dynamic language program projects containing type annotation information are obtained, the type information in the projects is extracted, and a type information data set is constructed; then, the extracted type information data set is embedded in vector form by using encoding technology; finally, the embedded vector is used to train a deep neural network model, and the trained model is used to predict the type of program variables or function signatures. The purpose of the present invention is to solve the problems of low efficiency and low accuracy of type inference for dynamic language programs existing at present, improve the readability, understandability and maintainability of dynamic language programs in the production practice of software engineering, and ultimately achieve the goal of improving software quality assurance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of program analysis and type inference in software engineering, and is particularly applicable to the field of type inference for dynamic language programs. Its purpose is to automatically improve the accuracy of type inference, and it is a software quality assurance method for improving software readability, understandability, and maintainability. Background Art

[0002] In dynamic language programs, variable types, object attributes, class structures, etc. can all be changed during program execution, which makes dynamic languages highly flexible. However, the flexible syntax features of dynamic languages make programs lack strict static checking and constraints, making it difficult for developers to timely discover defects in programs and increasing the costs of program debugging and maintenance. If automatic high-precision type inference for dynamic languages can be achieved, rich type information can be obtained. Type information can not only improve the readability, understandability, and maintainability of code, but also enable high-accuracy static checking of dynamic language programs.

[0003] Taking Python as an example, there are currently many Python type inference methods. According to the differences in the technologies used by the methods, they can be roughly divided into the following two categories: The first category is type inference based on traditional methods. This type of method only focuses on the source program and infers the types of identifiers from program contexts such as data flow information. The second category is type inference based on machine learning. This type of method focuses on the learning of program internal rules and uses machine learning technologies to predict the types of identifiers. This type of method mainly focuses on type dataset construction and model training. Theoretical and empirical studies have shown that the above theories and technologies all have good type inference effects and provide certain help for developers and maintainers. However, for dynamic languages, the first category of methods requires adding type annotations to the code libraries (standard libraries and third-party libraries), and based on the existing type code libraries, infers the types of identifiers in the program. When the project scale is large, manually annotating the code libraries is very time-consuming. Therefore, the type inference technology based on traditional methods cannot be applied to actual projects; for the second category of methods, there is no need to rely on code libraries containing type annotation information, and the trained type inference model is applicable to projects of any scale, outputting the type vectors of the predicted identifiers in the project and omitting the manual annotation work. However, most of the existing methods currently perform type inference on functions, and the accuracy of variable type inference is not high; moreover, the existing methods do not combine program slicing technology and do not make full use of the context information related to variables or functions. Therefore, the accuracy of type prediction is relatively low.

[0004] In summary, the present invention proposes a method for deep neural network type inference based on program slicing, which can automatically perform type inference on dynamic language programs. This method has the following three aspects of innovation: First, this method combines program slicing technology and makes full use of the identifier-related information in the program, thereby effectively improving the accuracy of type inference; Second, this method innovates the process of data embedding. First, the text information in the data set is preprocessed into the form of substrings, and then the substrings are embedded, effectively avoiding the problem of OOV (Out Of Vocabulary, the input data is not in the vocabulary) during the embedding process; Finally, this method also combines deep learning technology, which can not only infer the types of variables in the program, but also infer function signatures. Therefore, this method can significantly improve the accuracy of type inference and the efficiency of the inference process, reduce the time spent by developers and maintainers in understanding and maintaining project code, and effectively improve the efficiency of software R & D and software maintenance. Summary of the Invention

[0005] The present invention effectively solves the existing problems of low efficiency and low accuracy of dynamic language type inference by providing a method for deep neural network type inference based on program slicing, improves the readability, comprehensibility and maintainability of dynamic language programs in the production practice of software engineering, and finally achieves the goal of improving software quality assurance.

[0006] To achieve the above goal, the present invention proposes a method for deep neural network type inference based on program slicing. This method first needs to obtain a large number of projects containing type annotation information, extract the type information in the projects, and construct a type information data set; Secondly, this method embeds the extracted type information data set into a vector form using encoding technology; Then, based on the generated vectors by embedding, a deep neural network model is trained; Finally, the trained model is used to predict the types of program variables or function signatures. Specifically, this method includes the following steps.

[0007] 1) Construction of type information data set. First, a large number of projects containing type annotation information are crawled from open source code hosting platforms such as github. For variables or functions without type annotation information in the projects, traditional type inference tools are used to infer their type information in order to construct a large-scale type information data set. Then traverse the program AST (Abstract Syntax Tree), extract the names, types, location information of variables or functions and their program slice information, and construct the type information data set DataSet Type 。

[0008] 2) Embedding of type information. First, for the type information data set DataSet TypePreprocess the names of variables or functions and program slices with BPE (Byte Pair Embedding) encoding, and process the text information into the form of substrings. Then, input the substrings generated by the preprocessing into the Word2vec model for embedding to implement the conversion of text information into the input vector form Vec required for learning by the deep neural network. unioin Conversion.

[0009] 3) Training of the deep neural network model. Combine the joint vector Vec of the variables or functions and their program slices obtained by embedding in the previous step with the type vector encoded by one-hot (one-hot encoding) as the input data Vec to train the deep neural network model. During the training process, it is necessary to adjust the hyperparameters of the model or change the model structure according to the training results to optimize the model and improve the type prediction effect of the model, and finally obtain the type inference model M. unioin , combined with the type vector encoded by one-hot (one-hot encoding) as the input data Vec input to train the deep neural network model. During the training process, it is necessary to adjust the hyperparameters of the model or change the model structure according to the training results to optimize the model and improve the type prediction effect of the model, and finally obtain the type inference model M. Optimal .

[0010] 4) Type inference. After the model is trained, use the same embedding method to input the joint vector Vec of the variables or functions to be predicted and their program slices into the trained type inference model M pred for prediction. The prediction result is the vector Vec of the variable type or function parameters and return value types. Optimal in for prediction. The prediction result is the vector Vec of the variable type or function parameters and return value types. type . Description of the Drawings

[0011] Figure 1 is a flowchart of a deep neural network type inference method based on program slices in the implementation of the present invention. Detailed Embodiments

[0012] To better understand the technical content of the present invention, some specific examples are enumerated and combined with the accompanying drawings as follows.

[0013] As Figure 1 shown, a deep neural network type inference method based on program slices includes the following four steps:

[0014] S1 Construction of the type information dataset. Obtain a large number of open-source projects containing type annotation information, and then traverse the program AST to extract type information related to variables or functions to construct a type information dataset.

[0015] The implementation process of S1 specifically includes the following three steps:

[0016] S101: Obtain a large number of dynamic language program projects Projects containing type annotation information from open-source code hosting platforms such as github open ;

[0017] S102: For variables or functions in the project without type annotation information, use traditional type inference tools to infer their type information and supplement the obtained type information to the original project.

[0018] S103: Generate and traverse the AST for each source file in the project, and extract the name, type, and location information of variables or functions from specific nodes of the AST.

[0019] S104: The program slice Slicing can be obtained according to the identifier name and location relationship obtained from the AST. id For example, in this embodiment, from the statement pytest_plugins = ["abilian.testing.fixtures"], the variable name is extracted as pytest_plugins, the variable type is list, and its program slice is pytest_plugins = ["abilian.testing.fixtures"].

[0020] S2 Type information embedding. First, use BPE to preprocess the text information in the type information dataset into the form of substrings Substring, and then input the substring Substring into the Word2vec model for embedding, so as to convert the text information into the vector form Vec required by the deep neural network model. input .

[0021] The implementation process of S2 specifically includes the following four steps:

[0022] S201: Input the names of variables or functions and their program slices in the type information dataset into the BPE model for training.

[0023] S202: Input the data of the training set, test set, and validation set into the trained BPE model, and the output result is the training set, test set, and validation set generated in the form of preprocessed substrings. For example, in this embodiment... Checkerpytest_plugins = ["abilian... After BPE preprocessing, it is... Che@@cker@@pytest_@@plugins = ["@@abilian... (@ @ represents the string delimiter);

[0024] S203: Input the preprocessed data in the form of substrings into the Word2vec model for embedding to obtain the input data Vec in the form of high-dimensional vectors required for training the deep neural network model. unioin ;

[0025] S204: Input the type of a variable or function into the one-hot model for processing, and output the type vector Vec. label For example, in this embodiment, the one-hot encoding vector corresponding to the list type is [0, 1, 0, 0, 0,..., 0] (the second bit is 1, and the rest of the encoding bits are all 0).

[0026] S3 Training of the deep neural network model. The combined vector of the embedded variable or function and its program slice, along with the type vector encoded by one-hot (one-hot code), is used as the input data to train and optimize the deep neural network model, obtaining the finally trained type inference model.

[0027] The implementation process of S3 specifically includes the following three steps:

[0028] S301: Combine the combined vector of the embedded variable or function and its program slice with the one-hot encoded type vector Vec input as the input data to train the deep neural network model;

[0029] S302: Adjust the hyperparameters of the model according to the training results to optimize the model;

[0030] S303: Loop S301 and S302 until the model effect is optimal, obtaining the final type inference model M Optimal .

[0031] S4 Type inference. Represent the variable or function to be predicted and its program slice as a combined vector using the method of S2 type information embedding, and input the combined vector into the trained type inference model for prediction. The prediction result is the vector of the variable type, function parameters, and return value type.

[0032] The implementation process of S4 specifically includes the following two steps:

[0033] S401: Represent the variable or function to be predicted and its program slice as the combined vector Vec required by the model using the method of S2 type information embedding pred ;

[0034] S402: Input the combined vector into the trained type inference model M Optimal for prediction. The prediction result is the vector Vec of the variable type, function parameters, and return value type. type .

[0035] In summary, the present invention solves the problems of low efficiency and low accuracy in dynamic language program type inference, effectively improves the readability, comprehensibility, and maintainability of dynamic language programs in the production practice of software engineering, and better guarantees the quality of software engineering.

Claims

1. A method for type inference of deep neural networks based on program slicing, characterized in that To deduce the types of dynamic language programs, first, a large number of projects containing type annotation information need to be obtained, the type information in the projects is extracted, and a type information dataset is constructed; second, the method embeds the extracted type information dataset into a vector form using encoding techniques; then, based on the vectors generated by the embedding, a deep neural network model is trained; finally, the trained model is used to predict the types of program variables or function signatures; specifically, the method includes the following steps: 1) Construction of type information dataset: First, crawl a large number of projects containing type annotation information from the github open-source code hosting platform. For variables or functions without type annotation information in the projects, use traditional type inference tools to infer their type information in order to construct a large-scale type information dataset; then traverse the program abstract syntax tree AST, extract the names, types, location information, and program slice information of variables or functions, and construct the type information dataset DataSet Type ; 2) Type information embedding: First, perform byte pair encoding (BPE) preprocessing on the names of variables or functions and program slices in the type information data set DataSet Type to process the text information into the form of substrings; then, input the substrings generated by the preprocessing into the Word2vec model for embedding to achieve the conversion of the text information into the input vector form Vec unioin required for deep neural network learning; 3) Training of the deep neural network model: Use the joint vector Vec of the variables or functions obtained by embedding and their program slices in the previous step unioin , and combine it with the type vector encoded by one-hot as the input data Vec input to train the deep neural network model; During the training process, it is necessary to adjust the hyperparameters of the model or change the model structure according to the training results in order to optimize the model and improve the type prediction effect of the model, and finally obtain the type inference model M Optimal ; 4) Type inference: After the model is trained, use the same embedding method to obtain the joint vector Vec of the variable or function to be predicted and its program slice pred Input the trained type inference model M Optimal Perform prediction in it, and the prediction result is the vector Vec of the variable type or function parameter and return value type type .

2. The method for inferring the type of a deep neural network based on program slicing according to claim 1, characterized in that In step 1), a type information dataset is constructed; a large number of open-source projects containing type annotation information are obtained, and then the program AST is traversed to extract type information related to variables or functions, and a type information dataset is constructed.

3. The method for deriving the type of a deep neural network based on program slicing according to claim 1, wherein In step 2), vectors required to generate deep neural network vectors by embedding type information are obtained; first, the text information in the type information dataset is preprocessed into the form of substrings using BPE, and then the substrings are input into the Word2vec model for embedding, so as to convert the text information into the vector form required by the deep neural network model.

4. The method for inferring the type of a deep neural network based on program slicing according to claim 1, wherein In step 3), a deep neural network model is trained; the joint vectors of the embedded variables or functions and their program slices, combined with the type vectors encoded by the one-hot encoding, are used as input data to train and optimize the deep neural network model, and finally, a well-trained type deduction model is obtained.

5. The method for deriving the type of a deep neural network based on program slicing according to claim 1, characterized in that In step 4), the types of dynamic program variables and function signatures are deduced; the method of type information embedding is used to represent the variable or function to be predicted and its program slice as a joint vector, and the joint vector is input into the trained type deduction model for prediction, and the prediction result is the vector of the variable type or function parameters and return value types.

Citation Information

Patent Citations

  • UAF vulnerability detection method based on statement joint coding deep neural network

    CN110162972A

  • CNN and LSTM image high-level semantic understanding method based on information gain

    CN110188819A