Deep learning-assisted automatic code standardization review method and system

By constructing a multimodal fusion deep learning network architecture and adversarial training mechanism, generating a dataset of non-standardized code, performing syntax parsing and feature extraction, and designing a cross-entropy loss function for adversarial training, the problem of low efficiency in code standardization review in existing technologies is solved, and high-accuracy review of multi-language codes is achieved.

CN120085911BActive Publication Date: 2025-09-12KAIYUAN HUACHUANG TECH (GRP) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510579640.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-12
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Existing technologies are difficult to adapt to diverse code scenarios, resulting in low efficiency, strong subjectivity, and inconsistent understanding of code standardization reviews. It is especially difficult to effectively deal with codes in different languages, styles, and project structures.

Method used

By constructing a multimodal fusion deep learning network architecture and combining the adversarial training mechanism of standardized and non-standard code samples, a non-standard code dataset is generated, syntax parsing and feature extraction are performed, and a cross-entropy loss function is designed for adversarial training optimization to construct a code standardization review network.

Benefits of technology

It achieves the ability to adapt to the complex semantic features of multi-language codes and improves the accuracy of code standardization review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085911B_ABST
    Figure CN120085911B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for automatic code standardization review assisted by deep learning, which relates to the field of data processing technology, including: collecting a standard code data set, performing mutation enhancement, and generating a non-standard code data set; obtaining a standard code feature set and a non-standard code feature set; building a code annotation system, annotating the standard code feature set and the non-standard code feature set, and obtaining a code sample data set; designing a deep learning network architecture and a cross-entropy loss function; using a multimodal fusion network and a cross-entropy loss function to perform adversarial training optimization on the code sample data set, and constructing a code standardization review network; automatically reviewing the code data to be tested based on the code standardization review network, and determining the code standardization review results. The present invention solves the technical problem that the existing technology is difficult to adapt to diverse code scenarios, and achieves the technical effect of improving the accuracy of code standardization review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for automatic code standardization review assisted by deep learning. Background Art

[0002] With the ever-increasing demand for software development, code quality and standardization have become critical factors in ensuring system stability, maintainability, and security. Traditional code standardization review methods primarily rely on manual review or static rule-based review tools. These methods often suffer from low review efficiency, high subjectivity, and inconsistent understanding of standards when dealing with large-scale projects or multi-team collaborative development. This is particularly problematic when dealing with code in different languages, styles, and project structures, making them difficult to effectively address diverse code scenarios. Summary of the Invention

[0003] This application provides a deep learning-assisted automatic code standardization review method and system, which is used to solve the technical problem that existing technologies are difficult to adapt to diverse code scenarios.

[0004] In view of the above problems, this application provides a deep learning-assisted automatic code standardization review method and system.

[0005] In a first aspect of the present application, a deep learning-assisted automatic code standardization review method is provided, the method comprising:

[0006] A standardized code dataset is collected from open source projects and professional projects, and the standardized code dataset is mutated and enhanced to generate a non-standardized code dataset; the standardized code dataset and the non-standardized code dataset are subjected to grammatical parsing to obtain a standardized code feature set and a non-standardized code feature set; according to the code specification definition standard, a code annotation system is established, and positive samples and negative samples are annotated for the standardized code feature set and the non-standardized code feature set according to the code annotation system to obtain a code sample dataset; a deep learning network architecture is designed, and the deep learning network architecture is a multimodal fusion network, and a cross-entropy loss function is designed at the same time; the multimodal fusion network and the cross-entropy loss function are used to perform adversarial training optimization on the code sample dataset to construct a code standardization review network; based on the code standardization review network, the code data to be detected is automatically reviewed to determine the code standardization review results.

[0007] A second aspect of the present application provides a deep learning-assisted automatic code standardization review system, the system comprising:

[0008] A variation enhancement module is used to collect and obtain standard code datasets from open source projects and professional projects, perform variation enhancement on the standard code dataset, and generate a non-standard code dataset; a syntax parsing processing module is used to perform syntax parsing on the standard code dataset and the non-standard code dataset to obtain a standard code feature set and a non-standard code feature set; a sample annotation module is used to build a code annotation system according to the code specification definition standard, and perform positive and negative sample annotation on the standard code feature set and the non-standard code feature set according to the code annotation system to obtain a code sample dataset; a design module is used to design a deep learning network architecture, which is a multimodal fusion network, and at the same time design a cross-entropy loss function; an adversarial training optimization module is used to perform adversarial training optimization on the code sample dataset using the multimodal fusion network and the cross-entropy loss function to construct a code standardization review network; an automatic review module is used to automatically review the code data to be detected based on the code standardization review network to determine the code standardization review results.

[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0010] This application collects and obtains a standardized code dataset from open source projects and professional projects, performs mutation enhancement on the standardized code dataset, and generates a non-standardized code dataset; performs syntax parsing on the standardized code dataset and the non-standardized code dataset to obtain a standardized code feature set and a non-standardized code feature set; builds a code annotation system according to the code specification definition standard, and annotates the standardized code feature set and the non-standardized code feature set with positive samples and negative samples according to the code annotation system to obtain a code sample dataset; designs a deep learning network architecture, which is a multimodal fusion network, and designs a cross-entropy loss function; uses the multimodal fusion network and the cross-entropy loss function to perform adversarial training optimization on the code sample dataset to construct a code standardization review network; based on the code standardization review network, automatically reviews the code data to be detected and determines the code standardization review results. The present invention solves the technical problem that the existing technology is difficult to adapt to diversified code scenarios. By constructing a multimodal fusion deep learning network architecture and combining the adversarial training mechanism of standardized and non-standardized code samples, it realizes the adaptability to the complex semantic features of multi-language codes and achieves the technical effect of improving the accuracy of code standardization review. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0012] Figure 1 A flowchart of a method for automatic code standardization review assisted by deep learning provided in an embodiment of the present application;

[0013] Figure 2 Schematic diagram of the structure of the deep learning-assisted automatic code standardization review system provided in an embodiment of the present application.

[0014] Explanation of the accompanying symbols: variation enhancement module 11, syntax parsing processing module 12, sample labeling module 13, design module 14, adversarial training optimization module 15, automatic review module 16. DETAILED DESCRIPTION

[0015] This application provides a deep learning-assisted automatic code standardization review method and system to solve the technical problem that existing technologies are difficult to adapt to diverse code scenarios. By constructing a multimodal fusion deep learning network architecture and combining an adversarial training mechanism for standardized and non-standard code samples, it achieves the ability to adapt to the complex semantic features of multi-language codes, thereby achieving the technical effect of improving the accuracy of code standardization review.

[0016] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0017] It should be noted that any variations of the terms "include" and "have" are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.

[0018] Example 1, as Figure 1 As shown, the present application provides a deep learning-assisted automatic code standardization review method, the method comprising:

[0019] Step S100: acquiring a standard code dataset from open source projects and professional projects, performing mutation enhancement on the standard code dataset, and generating a non-standard code dataset.

[0020] In an embodiment of the present application, code samples that comply with coding standards are first collected from open source projects (such as public code hosting platforms such as GitHub and GitLab) and professional projects (such as high-quality code libraries within an enterprise) to construct a standard code dataset, that is, a high-quality code collection that complies with language grammar, naming standards, has a clear structure, and is free of logical errors.

[0021] Next, the standardized code dataset is mutated and enhanced. In this process, code mutation rules are first constructed, rule elements are extracted and data is retrieved for non-standard code samples, and a non-standard code dataset containing multiple types of abnormal features is constructed. The proportion of each type of problem elements is statistically analyzed to form a problem distribution reference; then, the mutation rules are weighted according to the distribution results, and the standardized code dataset is enhanced by combining various mutation methods to generate a non-standard code dataset.

[0022] Furthermore, in the method provided in the embodiment of the application, the generating of the non-standard code data set further includes:

[0023] Design code mutation rules, wherein the code mutation rules include format errors, irregular naming, logical defects, security vulnerabilities and performance issues; search for irregular code data based on each rule element in the code mutation rules to obtain an irregular rule code data set; perform statistics on the proportion of each rule element in the irregular rule code data set to obtain the proportion of irregular code rule elements; mutate and enhance the standard code data set based on the code mutation rules and the proportion of irregular code rule elements to generate an irregular code data set.

[0024] In an embodiment of the present application, code mutation rules are designed, and the code mutation rules include format errors (such as inconsistent indentation, missing brackets), non-standard naming (such as confusing variable naming, inconsistent with semantics), logical defects (such as incorrect boundary conditions, confusing conditional expressions), security vulnerabilities (such as plain text passwords, unchecked input) and performance issues (such as redundant calculations, dead loops).

[0025] We then searched for irregular code data based on the elements of each code variation rule. Using the typical problem elements defined in the rules as keywords or detection criteria, we used static code analysis tools (such as SonarQube, Pylint, and Checkstyle) to batch scan source code from open source and real-world projects, extracting irregular code snippets that exhibit these characteristics. Each type of code variation element, such as "short variable names," "unhandled exceptions," and "redundant loops within functions," was identified and located by the tool, ultimately resulting in a dataset of irregular code.

[0026] Next, we counted the proportion of each rule element in the non-standard rule code dataset, that is, we counted the frequency of occurrence of format errors, non-standard naming, logical defects, security vulnerabilities and performance issues in the entire dataset. We calculated the proportion parameters of each type of error by frequency ratio, and finally formed the proportion of non-standard code rule elements.

[0027] Finally, based on the code mutation rules and the proportion of non-standard code rule elements, the standardized code dataset is mutated and enhanced. In this process, each rule element in the code mutation rule is first weighted to form a mutation rule element weight factor. Based on this, a code element mutation strategy, including single-point mutation and combination mutation, is set to further determine the application coverage of each type of mutation rule element in the standardized code. Finally, based on the set mutation strategy and coverage requirements, differentiated mutation processing is implemented on the standardized code dataset to generate a non-standard code dataset covering multiple types of problem characteristics.

[0028] Furthermore, in the method provided in the embodiment of the application, the generating of the non-standard code data set further includes:

[0029] According to the proportion of the non-standard code rule elements, each rule element in the code mutation rule is assigned a mutation weight to obtain a mutation rule element weight factor; based on the code mutation rule, a code element mutation strategy is set, and the code element mutation strategy includes a single-point mutation strategy and a combination mutation strategy; according to the mutation rule element weight factor, the mutation rule element code coverage is determined; the code element mutation strategy and the mutation rule element code coverage are used to perform mutation enhancement processing on the standard code data set to generate the non-standard code data set.

[0030] In an embodiment of the present application, first, mutation weights are assigned to various types of variation rule elements according to the proportion of non-standard code rule elements. By using frequency normalization, the statistically obtained frequencies of occurrence of various types of non-standard elements (such as format errors, non-standard naming, logical defects, security vulnerabilities, performance issues) in the non-standard code data set are standardized into percentages as weight coefficients of each corresponding variation element, thereby obtaining a set of variation rule element weight factors.

[0031] Next, based on the code mutation rules, we set up code element mutation strategies. Specifically, for single-point mutation strategies, we define the application of one mutation rule to a single location in the code at a time, such as replacing a variable name with a nonstandard one or introducing a boundary error in a loop structure. For combined mutation strategies, we set up the simultaneous injection of multiple mutation rules of different types into the same code segment, such as introducing format errors within a function while also changing the logical structure of conditional statements.

[0032] A stratified sampling approach is then used to allocate the canonical code dataset and calculate the code coverage of the mutation rule elements. Specifically, based on the weight factor of each type of rule element, the number of samples to be enhanced is proportionally stratified from the canonical code dataset, and a "sample-mutation rule" correspondence table is generated. For example, if the weight factor for logic defects is 30%, 30% of the canonical samples are marked as targets for logic defect injection. This allocation mechanism clearly defines the sample range that each type of mutation should cover, thereby determining the code coverage of the mutation rule elements.

[0033] Finally, we employ an abstract syntax tree (AST) rewriting method to perform structural mutation operations on the canonical code samples. First, we use a language parser (such as Python's AST library) to parse the canonical code into an AST structure and locate target injection locations, such as variable declarations, expression nodes, and control flow statements. Then, based on the injection strategy and the injection type specified in the coverage table, we perform structural modifications on the syntax nodes. These modifications can include replacing non-standard names, inserting formatting errors, disrupting conditional expressions, adding unchecked inputs, or introducing invalid loops. Each mutation ensures syntactical legality and structural integrity, but may also contain standardization defects. Through this process, we ultimately generate a dataset of non-standard code.

[0034] Step S200: performing syntax parsing on the standard code data set and the non-standard code data set to obtain a standard code feature set and a non-standard code feature set.

[0035] In an embodiment of the present application, when performing syntax parsing processing on the standard code data set and the non-standard code data set, cleaning processing, programming language identification, syntax parsing and feature extraction are performed on the standard code data set and the non-standard code data set in sequence, and combined with semantic enhancement and label annotation, the structural features and semantic information of the code are extracted, and finally the standard code feature set and the non-standard code feature set are obtained.

[0036] Furthermore, in the method provided in the embodiment of the application, obtaining the standard code feature set and the non-standard code feature set further includes:

[0037] The standard code data set and the non-standard code data set are cleaned to obtain a usable standard code data set and a usable non-standard code data set; the usable standard code data set and the usable non-standard code data set are detected and classified according to the programming language to obtain code programming language information; based on the code programming language information, the usable standard code data set and the usable non-standard code data set are grammatically parsed and feature extracted to obtain standard code grammatical feature data and non-standard code grammatical feature data; semantic enhancement and labeling are performed on the standard code grammatical feature data and non-standard code grammatical feature data to obtain a standard code feature set and a non-standard code feature set.

[0038] In this example, a regular expression matching method is first used to perform preliminary cleaning of the standard code dataset and the non-standard code dataset. This method identifies and removes comments (such as comments beginning with / / or #), extra blank lines, illegal characters, and invalid formatting (such as garbled characters and missing end tags) in the source code, ensuring that each code segment has basic structural integrity and syntactically parseable, thereby obtaining a usable standard code dataset and a usable non-standard code dataset.

[0039] We then use a combined file suffix and keyword recognition method to detect and classify the programming language of the available code data. Specifically, we determine the language type based on the file name extension (such as .py, .java, .cpp) and typical syntax identifiers in the code snippet (such as def in Python). In this way, we identify the language attributes of each code segment and obtain the code programming language information.

[0040] Based on the programming language information, we then used syntax analysis tools supported by the corresponding language (such as Python's AST library, Java's JavaParser, and Clang AST for C++) to perform syntax parsing and feature extraction on the available standard and non-standard code datasets. By constructing and traversing the abstract syntax tree (AST) structure, we extracted key syntax nodes such as variable definitions, function structures, conditional statements, and loop statements from each code segment, generating structured grammatical feature data for standard and non-standard code, respectively.

[0041] The grammatical feature data is then semantically enhanced using a method that combines data flow and control flow graphs. Based on the original grammatical structure, this method analyzes contextual information such as the scope of variables in the function body, dependency paths, and call order to form a representation of the logical dependencies between variables and statements. Finally, a label injection method is used to classify and annotate the semantically enhanced feature data. Each code segment is assigned a clear category label based on its source (standard or non-standard dataset) and combined with the corresponding structural and semantic features to construct a standard code feature set and a non-standard code feature set, respectively.

[0042] Furthermore, in the method provided in the embodiment of the application, the obtaining of the standard code feature set and the non-standard code feature set further includes:

[0043] Context information embedding and semantic association expansion are performed on the standard code grammatical feature data and the non-standard code grammatical feature data to determine the standard code semantic feature data and the non-standard code semantic feature data; semantic feature generation is performed based on the standard code semantic feature data and the non-standard code semantic feature data to obtain standard code semantic enhancement features and non-standard code semantic enhancement features; the standard code grammatical feature data and the non-standard code grammatical feature data are fused with the standard code semantic enhancement features and the non-standard code semantic enhancement features to obtain the standard code feature set and the non-standard code feature set.

[0044] In an embodiment of the present application, context information is first embedded and semantic associations are expanded for the syntax feature data of the standard code and the syntax feature data of the non-standard code, and a static graph construction method based on control dependency and data dependency analysis is adopted. Specifically, compilation analysis technology is used, for example, using a compiler framework such as LLVM or Clang, or using static analysis tools such as Understand or Joern, to extract control dependencies between statements from the syntax tree, such as the path association between conditional branches and loop statements, and data dependencies, such as the chain of effects of variables in the definition and use process. These dependencies are organized into structured graph data to represent the execution logic order and variable transfer path in the code. Through this process, the semantic feature data of the standard code and the semantic feature data of the non-standard code are obtained respectively.

[0045] Semantic features are then generated based on the aforementioned semantic feature data. Manually coded rules and structural classification methods are used to categorize and identify various control structures and data dependency patterns, and convert them into quantifiable semantic description fields. For example, for deeply nested loop structures, their nesting levels can be recorded, and the use of variables across function calls can be identified as cross-scope dependencies. At the same time, structural attributes such as the maximum branch width, longest path depth, and the number of repeated variable assignments in a statement block are counted. These semantic attributes are generated through manual rule design and code structure analysis tools, forming semantic enhancement features for both standard and non-standard code.

[0046] Finally, the grammatical feature data is fused with the semantic enhancement features. Using field concatenation and structural merging, the structural information in the grammatical features and the contextual relationships in the semantic enhancement features are standardized. The two are then merged based on a preset feature template. For example, grammatical node type information and control structure depth information are concatenated into a unified structural description field, and a structural semantic fusion expression is formed through logical concatenation. After the fusion is complete, both standard and non-standard code feature sets are obtained, encompassing both grammatical structure, contextual relationships, and behavioral features.

[0047] Step S300: Building a code annotation system according to the code specification definition standard, annotating the standard code feature set and the non-standard code feature set with positive samples and negative samples according to the code annotation system to obtain a code sample data set.

[0048] In the embodiment of the present application, first, according to the definition standard of the code specification, a judgment basis for identifying and classifying code quality is constructed. This standard is derived from existing mainstream coding specification documents, including but not limited to PEP8, CERT secure coding specifications and industry internal R&D specifications, covering naming rules, format specifications, annotation requirements, logical integrity, exception handling, security policies and performance requirements. By structurally combing the items of these standards, they are converted into a set of rule items with judgment conditions, and a rule logic for automatic judgment and execution is formed, and finally a code annotation system is built.

[0049] Subsequently, the completed standard code feature sets and non-standard code feature sets were annotated according to the code annotation system. Using rule matching and structural comparison analysis methods, the standard rules in the code annotation system were loaded one by one and applied to each sample in the code feature set in turn, combining their grammatical features and semantic enhancement information for rule checking. For naming standard rules, for example, string pattern matching was used to check whether variable, function, and class names met the naming style requirements. For structural standard rules, the Abstract Syntax Tree (AST) structure was used to determine whether functions were too long, had empty function bodies, or repeated logic blocks. For security standard rules, static code checking methods were used to identify issues such as unchecked input, hard-coded passwords, and external command execution.

[0050] During the annotation process, if a sample fully complies with all the rules in the annotation system, it is marked as a positive sample, indicating that it meets the defined standards of the code specification. If a sample violates any of the rules, it is marked as a negative sample, indicating that it contains non-standard content in terms of structure, semantics, or security. The annotation process uses a logging and tagging recording mechanism to ensure that the annotation results and corresponding triggering rules of each sample can be tracked, reviewed, and iteratively updated. Ultimately, all annotated sample data is organized into a code sample dataset according to a unified structure format.

[0051] Step S400: Design a deep learning network architecture, which is a multimodal fusion network, and design a cross entropy loss function.

[0052] In an embodiment of the present application, a deep learning network architecture is first designed. Specifically, by defining the code input forms of grammatical modality, semantic modality and text modality, a corresponding multimodal sub-network is constructed, and the input layer embedding is completed; on this basis, splicing fusion or attention mechanism is selected as the feature fusion method to construct a fusion layer; finally, the code compliance classifier is embedded into the output layer, and the input layer, fusion layer and output layer are connected in series to complete the deep learning network architecture of multimodal fusion.

[0053] While building the network architecture, we designed a cross-entropy loss function as a supervisory signal during training. The cross-entropy loss function is used in both binary and multi-classification tasks. Its core concept is to measure the difference between the predicted probability distribution and the true label distribution. In code compliance review tasks, labeling samples as "compliant" or "non-compliant" is a typical binary classification problem. Using the cross-entropy loss function effectively penalizes the model's incorrect label predictions. It performs particularly well in tasks involving unbalanced datasets and those requiring precise probability outputs, improving both the model's convergence speed and classification accuracy.

[0054] Furthermore, in the method provided in the embodiment of the application, the design of the deep learning network architecture further includes:

[0055] Define code input modalities, which include grammatical modalities, semantic modalities, and textual modalities, and design an input multimodal subnetwork based on the code input modalities; embed the input multimodal subnetwork into an input layer to obtain a network input layer; obtain a feature fusion mechanism, which includes multimodal feature splicing and fusion and attention mechanism feature fusion, and construct a feature fusion layer based on the feature fusion mechanism; embed a code compliance classifier into the network output layer, and build the deep learning network architecture based on the network input layer, the feature fusion layer, and the network output layer, which are serially merged.

[0056] In the embodiment of the present application, the code input modality is first defined, and the information sources of the code samples are divided into three categories: syntactic modality, semantic modality, and textual modality using a feature partitioning method. Syntactic modality is obtained through the abstract syntax tree (AST), including features such as the code's structural hierarchy, node types, and statement composition; semantic modality uses the control flow graph (CFG) and data flow graph (DFG) to extract control dependencies and variable dependencies, expressing the logical structure of the code at the execution level; textual modality is input through the original code character stream, retaining surface information such as keywords, punctuation, and indentation format. Through this step, the multi-source structure of the model input features is clarified.

[0057] Next, we design a multimodal input subnetwork based on the code input modality, using a modular neural network design approach to construct independent subnetworks for each modality. Syntactic modality input is vectorized using an AST (Addressed Object Format) structure encoding network, typically employing a positional encoding tree-based expansion approach. Semantic modality input is encoded using a sequence network or adjacency matrix input structure to encode dependency paths. Text modality is transformed into word vectors using an embedding layer for code characters or word segmentation results. The results of each of the three subnetworks are used as intermediate feature outputs to provide input data for subsequent fusion.

[0058] The input layer of the multimodal subnetwork is then embedded, and the outputs of the three subnetworks are mapped into a unified representation space using feature embedding mapping. By setting a unified dimensional space, the results of the separate processing of grammatical, semantic, and textual features are converted into fixed-length vectors and normalized to make the features of different modalities comparable. Ultimately, a structurally consistent network input layer is formed, completing the transformation from raw features to a unified embedded representation.

[0059] Next, a feature fusion mechanism is identified, and a fusion strategy selection method is used to determine whether to use multimodal feature concatenation or attention-based feature fusion. Concatenation directly concatenates the input vectors of each modality to form a joint feature without altering their original content. Attention-based fusion, by introducing a weight calculation module, assigns dynamic weight coefficients to features of different modalities, enabling modeling of differences in importance between modalities. Based on the selected mechanism, a feature fusion layer is further constructed to comprehensively represent the features of the three modalities, improving the model's ability to integrate complex code features.

[0060] Finally, the code compliance classifier is embedded in the network output layer, and a fully connected neural network is used to construct the output module. This classifier performs nonlinear transformation and classification on the fused multimodal feature input, outputting a predicted probability of each code sample being "compliant" or "non-compliant." Ultimately, the network input layer, the feature fusion layer, and the network output layer are serially combined to create a complete deep learning network architecture.

[0061] Step S500: Adopting the multimodal fusion network and the cross entropy loss function to perform adversarial training optimization on the code sample dataset to construct a code standardization review network.

[0062] In an embodiment of the present application, a multimodal fusion network is used to perform compliance classification training on a code sample dataset to construct a basic compliance review network; then the cross-entropy loss function is used to perform loss assessment to obtain initial network loss data; based on the loss data, an adversarial mechanism is introduced to determine the strength of the adversarial sample; finally, the adversarial sample strength is combined with the original sample data to perform adversarial training optimization on the basic network to construct a code compliance review network.

[0063] Furthermore, in the method provided in the embodiment of the application, the construction of the code standardization review network further includes:

[0064] The multimodal fusion network is used to perform compliance classification training on the code sample dataset to obtain a basic compliance review network; the cross-entropy loss function is used to perform loss assessment on the basic compliance review network to obtain initial network loss data; an adversarial mechanism is introduced to determine the strength of the adversarial sample based on the initial network loss data; based on the adversarial sample strength and the code sample dataset, the basic compliance review network is optimized through adversarial training to construct a code standardization review network.

[0065] In an embodiment of the present application, a supervised training method is first used to perform compliance classification training on a code sample dataset using a multimodal fusion network. Specifically, the feature inputs containing grammatical modality, semantic modality, and textual modality are encoded through corresponding sub-networks respectively, and the multimodal features are integrated using splicing or attention mechanisms in the fusion layer. The fused features are then fed into a fully connected classifier for training. By continuously iterating the model parameters, the model is able to accurately output the classification results of whether each piece of code is standardized or non-standard, thereby obtaining a basic compliance review network with preliminary discrimination capabilities.

[0066] Next, the basic compliance review network is evaluated using a cross-entropy loss calculation method. The model's classification output is compared with the true label, and the prediction error for each sample is calculated using the cross-entropy function. This function amplifies the prediction error term using a logarithmic function, making the model more sensitive to classification boundaries. The average loss value across all samples forms the target evaluation metric during training, ultimately generating initial network loss data that reflects the model's learning performance.

[0067] Then, using a gradient-based adversarial perturbation generation method, we introduce an adversarial mechanism and determine the strength of adversarial examples based on the initial network loss data. Using the Fast Gradient Sign Method (FGSM), we add a small perturbation along the positive gradient direction of the cross-entropy loss function to the input sample, and use the magnitude of this perturbation as the strength of the adversarial example.

[0068] Finally, adversarial training optimization is performed on the basic compliance review network based on the strength of the adversarial examples and the code sample dataset. Specifically, the code sample dataset is perturbed based on the strength of the adversarial examples to generate adversarial sample data, which is then combined with the original code sample dataset to construct code training sample data. Subsequently, the basic compliance review network is evaluated for training updates using the cross-entropy loss function to obtain network update loss data. Based on this loss data, model parameters are continuously adjusted, and iterative adversarial optimization is performed to ultimately complete enhanced training of the basic network and construct the code compliance review network.

[0069] Furthermore, in the method provided in the embodiment of the application, the construction of the code standardization review network further includes:

[0070] The adversarial sample strength is used to perform adversarial perturbations on the code sample dataset to obtain adversarial sample data, and code training sample data is obtained based on the code sample dataset and the adversarial sample data; the cross-entropy loss function is used to perform training update evaluation on the basic compliance review network based on the code training sample data to obtain network update loss data; based on the network update loss data, the basic compliance review network is iteratively adversarially optimized to obtain a code standardization review network.

[0071] In this embodiment, the code sample dataset is first perturbed using adversarial sample strength. The Fast Gradient Signed Method (FGS) adds small perturbations based on the gradient direction to the input features in the code sample dataset. The perturbation amplitude is controlled by a preset adversarial sample strength to ensure that the generated adversarial samples retain their legal grammatical structure while being highly confusing to the underlying model. This results in a batch of adversarial sample data that is both aggressive and structurally controllable.

[0072] The generated adversarial sample data is then fused with the original code sample dataset. Through sample splicing and reconstruction, a new training data set is formed. That is, code training sample data that contains both the original samples and their corresponding perturbation samples. This is used to improve the model's generalization ability when facing boundary samples and uncertain inputs.

[0073] Next, the cross-entropy loss function is used to evaluate the training update of the basic compliance review network based on the code training sample data. In this step, the code training sample data is input into the constructed basic compliance review network, the model output is obtained, and the cross-entropy loss function is used to calculate the difference between the predicted result and the true label to evaluate the current model performance. The smaller the loss value, the closer the model prediction is to the true label. The obtained loss value is the network update loss data for the current round, which is used to quantify the model's learning effect under the interference of adversarial samples.

[0074] Finally, the basic compliance review network is iteratively optimized based on the network update loss data. In this process, the adversarial sample strength is first dynamically adjusted based on the network update loss data to obtain the adversarial sample update strength. Based on this, the basic compliance review network is then trained through multiple rounds of iterative adversarial training until the preset convergence conditions are met, thus constructing the initial compliance review network. Finally, the initial network is tested, evaluated, and optimized, ultimately resulting in a stable and discriminative code compliance review network.

[0075] Furthermore, in the method provided in the embodiment of the application, obtaining the code standardization review network further includes:

[0076] The adversarial sample strength is dynamically updated using the network update loss data to obtain the adversarial sample update strength; the basic compliance review network is iteratively adversarially trained based on the adversarial sample update strength until a preset convergence condition is met to obtain an initial compliance review network; the initial compliance review network is tested, optimized and updated to obtain the code standardization review network.

[0077] In the embodiment of the present application, the loss trend driven adjustment method is first used to dynamically update the strength of the adversarial sample. Specifically, after each round of training, the cross entropy loss value of the basic compliance review network in the current training round is recorded and compared with the loss of the previous round to calculate the loss change rate. If the loss decreases slowly or tends to be stable, the amplitude of the adversarial perturbation is increased through a linear enhancement strategy. Otherwise, the perturbation intensity is appropriately weakened to avoid training instability. In this way, the strength of the adversarial sample is updated.

[0078] The basic compliance review network is then trained and optimized over multiple rounds using a fixed-batch adversarial training method. Specifically, before each round of training, the original training data is perturbed using the updated adversarial sample strength to generate new adversarial samples, which are then mixed with the unperturbed samples to form a mixed training batch. Each round of training uses a fixed-size batch of samples as input to the model, and standard forward and backpropagation algorithms are used to update parameters while monitoring changes in loss values. When the loss fluctuations for multiple consecutive rounds are less than a set threshold, or the validation set accuracy tends to stabilize, the preset convergence conditions are considered to have been met, and the training process is terminated, resulting in the initial compliance review network.

[0079] Finally, the initial compliance review network was tested and optimized using a validation set performance feedback optimization method. Pre-prepared independent validation samples were fed into the network, and key performance indicators such as precision, recall, and false positive rate were measured on both standard and non-standard code samples. Based on the evaluation results, parameter fine-tuning was used to make small adjustments to the fusion layer connection structure, classifier output thresholds, or weights of each modal feature to improve the model's ability to discriminate borderline samples and overall stability. After optimization, the final code compliance review network was obtained.

[0080] Step S600: Automatically review the code data to be inspected based on the code standardization review network to determine the code standardization review result.

[0081] In an embodiment of the present application, the code data to be tested is first input into a trained code compliance review network. The network automatically identifies and classifies the input data and outputs a corresponding prediction result. Based on this result, it is determined whether the code complies with the specification. If the prediction result is "compliant", a compliance mark is output; if it is "non-compliant", a non-compliance mark is output. Finally, the automatic review of the code to be tested is completed, and its corresponding code compliance review result is determined.

[0082] In the embodiments of the present application, in summary, the embodiments of the present application have at least the following technical effects:

[0083] This application collects and obtains a standardized code dataset from open source projects and professional projects, performs mutation enhancement on the standardized code dataset, and generates a non-standardized code dataset; performs syntax parsing on the standardized code dataset and the non-standardized code dataset to obtain a standardized code feature set and a non-standardized code feature set; builds a code annotation system according to the code specification definition standard, and annotates the standardized code feature set and the non-standardized code feature set with positive samples and negative samples according to the code annotation system to obtain a code sample dataset; designs a deep learning network architecture, which is a multimodal fusion network, and designs a cross-entropy loss function; uses the multimodal fusion network and the cross-entropy loss function to perform adversarial training optimization on the code sample dataset to construct a code standardization review network; based on the code standardization review network, automatically reviews the code data to be detected and determines the code standardization review results. The present invention solves the technical problem that the existing technology is difficult to adapt to diversified code scenarios. By constructing a multimodal fusion deep learning network architecture and combining the adversarial training mechanism of standardized and non-standardized code samples, it realizes the adaptability to the complex semantic features of multi-language codes and achieves the technical effect of improving the accuracy of code standardization review.

[0084] Example 2, based on the same inventive concept as the deep learning-assisted automatic code standardization review method in the previous embodiment, Figure 2 As shown, this application provides a deep learning-assisted automatic code standardization review system. The system and method embodiments in the embodiments of this application are based on the same inventive concept. The system includes:

[0085] A variation enhancement module 11 is used to collect and obtain a standard code data set from open source projects and professional projects, perform variation enhancement on the standard code data set, and generate a non-standard code data set; a syntax parsing processing module 12 is used to perform syntax parsing on the standard code data set and the non-standard code data set to obtain a standard code feature set and a non-standard code feature set; a sample annotation module 13 is used to build a code annotation system according to the code specification definition standard, and perform positive sample and negative sample annotation on the standard code feature set and the non-standard code feature set according to the code annotation system to obtain a code sample data set; a design module 14 is used to design a deep learning network architecture, which is a multimodal fusion network, and at the same time design a cross-entropy loss function; an adversarial training optimization module 15 is used to perform adversarial training optimization on the code sample data set using the multimodal fusion network and the cross-entropy loss function to construct a code standardization review network; an automatic review module 16 is used to automatically review the code data to be detected based on the code standardization review network to determine the code standardization review result.

[0086] Furthermore, the system is also used to implement the following functions:

[0087] Design code mutation rules, wherein the code mutation rules include format errors, irregular naming, logical defects, security vulnerabilities and performance issues; search for irregular code data based on each rule element in the code mutation rules to obtain an irregular rule code data set; perform statistics on the proportion of each rule element in the irregular rule code data set to obtain the proportion of irregular code rule elements; mutate and enhance the standard code data set based on the code mutation rules and the proportion of irregular code rule elements to generate an irregular code data set.

[0088] Furthermore, the system is also used to implement the following functions:

[0089] According to the proportion of the non-standard code rule elements, each rule element in the code mutation rule is assigned a mutation weight to obtain a mutation rule element weight factor; based on the code mutation rule, a code element mutation strategy is set, and the code element mutation strategy includes a single-point mutation strategy and a combination mutation strategy; according to the mutation rule element weight factor, the mutation rule element code coverage is determined; the code element mutation strategy and the mutation rule element code coverage are used to perform mutation enhancement processing on the standard code data set to generate the non-standard code data set.

[0090] Furthermore, the system is also used to implement the following functions:

[0091] The standard code data set and the non-standard code data set are cleaned to obtain a usable standard code data set and a usable non-standard code data set; the usable standard code data set and the usable non-standard code data set are detected and classified according to the programming language to obtain code programming language information; based on the code programming language information, the usable standard code data set and the usable non-standard code data set are grammatically parsed and feature extracted to obtain standard code grammatical feature data and non-standard code grammatical feature data; semantic enhancement and labeling are performed on the standard code grammatical feature data and non-standard code grammatical feature data to obtain a standard code feature set and a non-standard code feature set.

[0092] Furthermore, the system is also used to implement the following functions:

[0093] Context information embedding and semantic association expansion are performed on the standard code grammatical feature data and the non-standard code grammatical feature data to determine the standard code semantic feature data and the non-standard code semantic feature data; semantic feature generation is performed based on the standard code semantic feature data and the non-standard code semantic feature data to obtain standard code semantic enhancement features and non-standard code semantic enhancement features; the standard code grammatical feature data and the non-standard code grammatical feature data are fused with the standard code semantic enhancement features and the non-standard code semantic enhancement features to obtain the standard code feature set and the non-standard code feature set.

[0094] Furthermore, the system is also used to implement the following functions:

[0095] Define code input modalities, which include grammatical modalities, semantic modalities, and textual modalities, and design an input multimodal subnetwork based on the code input modalities; embed the input multimodal subnetwork into an input layer to obtain a network input layer; obtain a feature fusion mechanism, which includes multimodal feature splicing and fusion and attention mechanism feature fusion, and construct a feature fusion layer based on the feature fusion mechanism; embed a code compliance classifier into the network output layer, and build the deep learning network architecture based on the network input layer, the feature fusion layer, and the network output layer, which are serially merged.

[0096] Furthermore, the system is also used to implement the following functions:

[0097] The multimodal fusion network is used to perform compliance classification training on the code sample dataset to obtain a basic compliance review network; the cross-entropy loss function is used to perform loss assessment on the basic compliance review network to obtain initial network loss data; an adversarial mechanism is introduced to determine the strength of the adversarial sample based on the initial network loss data; based on the adversarial sample strength and the code sample dataset, the basic compliance review network is optimized through adversarial training to construct a code standardization review network.

[0098] Furthermore, the system is also used to implement the following functions:

[0099] The adversarial sample strength is used to perform adversarial perturbations on the code sample dataset to obtain adversarial sample data, and code training sample data is obtained based on the code sample dataset and the adversarial sample data; the cross-entropy loss function is used to perform training update evaluation on the basic compliance review network based on the code training sample data to obtain network update loss data; based on the network update loss data, the basic compliance review network is iteratively adversarially optimized to obtain a code standardization review network.

[0100] Furthermore, the system is also used to implement the following functions:

[0101] The adversarial sample strength is dynamically updated using the network update loss data to obtain the adversarial sample update strength; the basic compliance review network is iteratively adversarially trained based on the adversarial sample update strength until a preset convergence condition is met to obtain an initial compliance review network; the initial compliance review network is tested, optimized and updated to obtain the code standardization review network.

[0102] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0103] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0104] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.

Claims

1. A deep learning-assisted automatic code standardization review method, characterized by: The method comprises: Acquire a standardized code dataset from open source projects and professional projects, perform mutation enhancement on the standardized code dataset, and generate a non-standard code dataset; Performing syntax parsing on the standard code data set and the non-standard code data set to obtain a standard code feature set and a non-standard code feature set; According to the code specification definition standard, a code annotation system is established, and positive and negative samples of the standard code feature set and the non-standard code feature set are annotated according to the code annotation system to obtain a code sample data set; Designing a deep learning network architecture, wherein the deep learning network architecture is a multimodal fusion network, and designing a cross entropy loss function; Adopting the multimodal fusion network and the cross entropy loss function to perform adversarial training optimization on the code sample dataset to construct a code standardization review network; Automatically reviewing the code data to be inspected based on the code standardization review network to determine the code standardization review result; The generating of the non-standard code data set includes: Design code mutation rules, including format errors, naming irregularities, logic defects, security vulnerabilities, and performance issues; Searching for irregular code data based on each rule element in the code variation rule to obtain an irregular regular code data set; Performing statistics on the proportion of each rule element in the irregular rule code data set to obtain the proportion of irregular code rule elements; Performing mutation enhancement on the standard code dataset based on the code variation rule and the proportion of the non-standard code rule elements to generate a non-standard code dataset; Assigning a mutation weight to each rule element in the code mutation rule according to the proportion of the non-standard code rule elements to obtain a mutation rule element weight factor; Based on the code mutation rule, a code element mutation strategy is set, wherein the code element mutation strategy includes a single point mutation strategy and a combination mutation strategy; Determining the variation rule element code coverage rate according to the variation rule element weight factor; The code element variation strategy and the variation rule element code coverage are used to perform variation enhancement processing on the standard code data set to generate the non-standard code data set.

2. The deep learning-assisted automatic code standardization review method according to claim 1, characterized in that: The obtaining of the standard code feature set and the non-standard code feature set includes: Cleaning the standard code dataset and the non-standard code dataset to obtain a usable standard code dataset and a usable non-standard code dataset; Performing detection and classification processing on the available standard code data set and the available non-standard code data set according to programming language to obtain code programming language information; Performing syntax parsing and feature extraction on the available standard code dataset and the available non-standard code dataset based on the code programming language information to obtain standard code syntax feature data and non-standard code syntax feature data; Semantic enhancement and labeling are performed on the standard code grammatical feature data and the non-standard code grammatical feature data to obtain a standard code feature set and a non-standard code feature set.

3. The deep learning-assisted automatic code standardization review method according to claim 2, characterized in that: The obtaining of the standard code feature set and the non-standard code feature set includes: Performing context information embedding and semantic association expansion on the standard code grammatical feature data and the non-standard code grammatical feature data to determine the standard code semantic feature data and the non-standard code semantic feature data; Generate semantic features based on the standard code semantic feature data and the non-standard code semantic feature data to obtain standard code semantic enhancement features and non-standard code semantic enhancement features; The standard code grammatical feature data and the non-standard code grammatical feature data are fused with the standard code semantic enhancement features and the non-standard code semantic enhancement features to obtain the standard code feature set and the non-standard code feature set.

4. The deep learning-assisted automatic code standardization review method according to claim 1, characterized in that: The design of the deep learning network architecture includes: Defining code input modalities, wherein the code input modalities include grammatical modalities, semantic modalities, and textual modalities, and designing an input multimodal subnetwork based on the code input modalities; Embedding the input multimodal sub-network into an input layer to obtain a network input layer; Acquire a feature fusion mechanism, the feature fusion mechanism including multimodal feature splicing fusion and attention mechanism feature fusion, and construct a feature fusion layer according to the feature fusion mechanism; The code compliance classifier is embedded into the network output layer, and the network input layer, the feature fusion layer and the network output layer are combined in series to build the deep learning network architecture.

5. The deep learning-assisted automatic code standardization review method according to claim 1, characterized in that: The construction of the code standardization review network includes: Using the multimodal fusion network to perform compliance classification training on the code sample dataset to obtain a basic compliance review network; Performing loss assessment on the basic compliance review network using the cross entropy loss function to obtain initial network loss data; Introducing an adversarial mechanism to determine the strength of the adversarial sample based on the initial network loss data; Based on the strength of the adversarial samples and the code sample data set, the basic compliance review network is optimized through adversarial training to construct a code compliance review network.

6. The deep learning-assisted automatic code standardization review method according to claim 5, characterized in that: The construction of the code standardization review network includes: Performing adversarial perturbations on the code sample dataset using the adversarial sample strength to obtain adversarial sample data, and obtaining code training sample data based on the code sample dataset and the adversarial sample data; Using the cross entropy loss function to perform training update evaluation on the basic compliance review network based on the code training sample data, to obtain network update loss data; The basic compliance review network is iteratively optimized based on the network update loss data to obtain a code standardization review network.

7. The deep learning-assisted automatic code standardization review method according to claim 6, characterized in that: The code standardization review network includes: Dynamically updating the adversarial sample strength using the network update loss data to obtain the adversarial sample update strength; Iteratively adversarial training is performed on the basic compliance review network based on the updated strength of the adversarial sample until a preset convergence condition is met, thereby obtaining an initial compliance review network; The initial compliance review network is tested, optimized and updated to obtain the code standardization review network.

8. Deep learning-assisted automatic code standardization review system, characterized by: The system is used to execute the deep learning-assisted automatic code standardization review method according to any one of claims 1 to 7, and the system includes: A variation enhancement module is used to collect and obtain standard code datasets from open source projects and professional projects, perform variation enhancement on the standard code datasets, and generate non-standard code datasets; A syntax parsing processing module, configured to perform syntax parsing on the standard code data set and the non-standard code data set to obtain a standard code feature set and a non-standard code feature set; A sample annotation module is used to build a code annotation system according to the code specification definition standard, and annotate the standard code feature set and the non-standard code feature set with positive samples and negative samples according to the code annotation system to obtain a code sample dataset; A design module is used to design a deep learning network architecture, wherein the deep learning network architecture is a multimodal fusion network and a cross entropy loss function; An adversarial training optimization module, configured to perform adversarial training optimization on the code sample dataset using the multimodal fusion network and the cross entropy loss function to construct a code standardization review network; The automatic review module is used to automatically review the code data to be detected based on the code standardization review network and determine the code standardization review result.

Citation Information

Patent Citations

  • Code review method and device, electronic equipment and medium

    CN117648931A