An intelligent code rewriting and verification method for the trust innovation platform

Through the intelligent code rewriting method based on the big model, combined with source code labeling, functional module division, code analysis and static verification, the problem of converting image processing code from different programming languages ​​into Java code is solved, and efficient and accurate code conversion and verification is achieved, improving the portability and maintainability of the code.

CN119025091BActive Publication Date: 2025-05-13NINGBO JINWANG INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411505085.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-05-13
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately convert image processing code written in languages ​​such as Python, C++ and MATLAB into Java code, and ensure that the converted code is free of potential errors, and is suitable for the image processing type of the target information innovation platform.

Method used

Through source code labeling, functional module division, code analysis, transformation model generation, static verification and automated testing, a transformation model based on large models is built to achieve efficient conversion and verification of code. The specific steps include labeling and functional module division of the source code, determining the module functions using the clustering algorithm, performing code parsing and conversion, generating Java code sequences, and ensuring the correctness and performance of the code through static verification and automated testing.

Benefits of technology

It realizes efficient conversion of image processing algorithms to Java code from different programming languages, ensuring that the converted code is suitable for the image processing type of the target information innovation platform, improving the portability and maintainability of the code, and reducing development time and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119025091B_ABST
    Figure CN119025091B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent code rewriting and verification method for an information and innovation platform, which belongs to the technical field of computer code writing, and includes labeling source code; dividing source code function modules according to the label annotation so that each function module corresponds to an image processing type, and obtaining a target source code module according to a target information and innovation platform; parsing the source code module of the target image processing type, inputting the parsed code into a conversion model, performing code conversion, and generating a corresponding Java code sequence; statically verifying the Java code sequence generated by the conversion model, and if the static verification passes, performing automated testing according to the target image processing type until outputting Java code that can execute the target image processing type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer code writing technology, and more specifically to an intelligent code rewriting and verification method for an information innovation platform. Background Art

[0002] In today's digital age, image processing technology has become an indispensable part of scientific research and industrial applications. Programming languages ​​such as Python, C++, and MATLAB have been widely used in the field of image processing due to their powerful library support and efficient performance. However, with the rise of the information technology application innovation (ICT) platform, languages ​​such as Java have gradually become the mainstream programming languages ​​of the target platform. Due to the significant differences between Python, C++, and MATLAB and Java in terms of syntax, data types, and execution environment, there are many challenges in directly converting image processing codes written in these languages ​​into Java codes. Traditional code conversion methods often rely on manual translation, which is not only time-consuming and labor-intensive, but also prone to errors. In addition, due to the complexity and diversity of image processing algorithms, manually converted codes often require a lot of debugging and verification work to ensure their correctness and performance. Therefore, how to efficiently and accurately convert image processing codes written in languages ​​such as Python, C++, and MATLAB into the editing language of the target ICT platform and verify them has become an urgent problem to be solved. In recent years, the development of large model technology has provided a new idea for solving this problem. By training large-scale neural network models, large models can learn the semantic and structural differences between different programming languages ​​and automatically complete the code conversion task. At the same time, the large model can also perform preliminary verification on the converted code to ensure its logical correctness and performance. Therefore, the study of image processing code conversion and verification methods and systems based on large models has important practical significance and application value.

[0003] The invention patent with publication number CN115145574A discloses a code generation method, device, storage medium and server, wherein the method includes: obtaining at least one code file in the source code information corresponding to the service interface; performing syntax parsing on each code file in the at least one code file to generate a syntax tree corresponding to each code file; traversing the syntax tree to collect the interface attribute information set in each code file; using the set interface specification, and generating the target code information corresponding to the source code information based on the interface attribute information set.

[0004] Although the existing technology can reduce the communication cost between users when generating target code information, reduce the errors caused by manual writing, and improve the accuracy and convenience of code generation, it still fails to solve the problem of converting image processing algorithms written in different programming languages ​​into Java code, and ensuring that the converted code is free of potential errors, while being able to effectively process the image processing type of the target trusted innovation platform. Therefore, in order to overcome these limitations, the present invention proposes an intelligent code rewriting and verification method for trusted innovation platforms. Summary of the invention

[0005] In view of the deficiencies in the prior art, the purpose of the present invention is to provide an intelligent code rewriting and verification method for the trusted innovation platform, which solves the problem of converting image processing algorithms written in different programming languages ​​into Java code and ensures that the converted code is free of potential errors, while being able to effectively handle the image processing type of the target trusted innovation platform. Through source code labeling, functional module division, code parsing, conversion model generation, static verification and automated testing, not only the portability and maintainability of the code are improved, but also the development time and cost are reduced, providing an efficient and reliable code conversion and verification solution for the trusted innovation platform.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] An intelligent code rewriting and verification method for an information innovation platform, characterized by comprising:

[0008] Label the source code and use natural language processing technology to extract text information from code comments and document strings for auxiliary labeling;

[0009] Divide the source code function modules according to the label annotations so that each function module corresponds to an image processing type, and obtain the source code module of the target image processing type according to the target information innovation platform;

[0010] Parse the source code module of the target image processing type, input the parsed code into the conversion model, perform code conversion, and generate a corresponding Java code sequence;

[0011] The Java code sequence generated by the conversion model is statically verified. If the static verification passes, automated testing is performed according to the target image processing type until Java code that can execute the target image processing type is output.

[0012] Specifically, the division of the source code function modules includes:

[0013] Text search is used to extract tags related to image processing in source code modules;

[0014] Use clustering algorithms to analyze text data on the extracted tags and determine the function of each module;

[0015] According to the function of the module, determine the image processing type corresponding to each module;

[0016] Group modules according to their functions and create file storage modules.

[0017] Specifically, the clustering algorithm determines the module function specific steps include:

[0018] Preprocess the extracted tags;

[0019] For each tag, count the frequency of each word, that is, ,in, It is The number of times a word appears, is the total number of words in the label and measures the importance of each word in the label, i.e. ,in, is the total number of tags, It includes The number of tags for each word;

[0020] The feature vector is constructed by combining the frequency of each word and the importance of each word in the label, that is, ,in, Indicates of the tags The feature vector elements of words, and ;

[0021] Select a clustering algorithm, cluster the labels according to the feature vector to group similar labels together to form different categories, count the number of labels in each category, and configure the clustering threshold , obtain the categories whose number of labels is greater than the clustering threshold as key categories;

[0022] Based on the keywords of the key categories, the key categories are mapped to an image processing function category according to a mapping table.

[0023] Specifically, the specific steps of the code parsing include:

[0024] Preprocess the source code, including removing comments and blank lines;

[0025] Get all tags in the source code, construct a tag sequence, and convert the tag sequence into an abstract syntax tree;

[0026] Perform semantic analysis on the abstract syntax tree to extract semantic information from the source code, including variable types, function signatures, and control flows;

[0027] Serialize the abstract syntax tree into input format and convert the abstract syntax tree into a linear data structure, including lists and arrays.

[0028] Specifically, the specific steps of converting the abstract syntax tree into a data structure for input include:

[0029] Start recursively traversing the root node of the abstract syntax tree, and visit each node according to the breadth-first search process;

[0030] For each node, extract node information, which includes the type of node , the value of the node , child node list ;

[0031] According to the extracted node information, construct the input sequence , .

[0032] Specifically, the specific steps of the code conversion include:

[0033] Build a conversion model based on the large model, construct a loss function according to the target image processing type, input the input sequence of the source code module into the conversion model, and convert the source code programming language into Java through the conversion model;

[0034] Perform code refactoring and code optimization, optimize code structure and logic, improve readability and maintainability, improve performance, and reduce resource consumption.

[0035] Specifically, the construction of the loss function includes:

[0036] Construct the cross entropy loss to measure the difference between the predicted result and the true label. The cross entropy loss calculation formula is as follows: ,in, is the number of code nodes, is the number of categories, is the true label, is the probability predicted by the transformation model;

[0037] The code edit distance loss is constructed to measure the similarity between the converted code and the original code. The code edit distance loss calculation formula is as follows: ,in, is the sample size, is the original node code, is the code of the converted node, is a function that calculates the edit distance between code snippets.

[0038] Specifically, the construction of the loss function also includes:

[0039] Construct an image type constraint loss to ensure that the converted code is suitable for the target image processing type. The image type constraint loss calculation formula is as follows: ,in, is an indicator function, if If the image processing code is included in the required library, it returns 1, otherwise it returns 0, which is used to evaluate whether the converted code meets the requirements of the image processing type;

[0040] Combining cross entropy loss, code edit distance loss and image type constraint loss, the total loss function is constructed: ,in, , is the weight parameter.

[0041] Specifically, the static verification step includes:

[0042] The structural complexity of the code is obtained by the number and nesting depth of conditional statements. The calculation formula of code complexity is as follows: ;in, is the code complexity, is the number of edges, is the number of nodes, is the number of connected components, is the number of conditional statements, is the nesting depth;

[0043] The redundant parts in the code are identified and quantified by measuring the average size and similarity of the code blocks. The formula for calculating the code repetition rate is as follows: ,in, is the code repetition rate, is the number of distinct code blocks, is the total number of lines of code, is the average size of a code block, is the similarity measure of code blocks;

[0044] Combining branch coverage and function coverage to measure the coverage of the code, the calculation formula of code coverage is as follows: ,in, is the code coverage, TC is the number of test cases executed, is the total number of test cases, is the branch coverage, is the functional coverage.

[0045] Specifically, the static verification step also includes:

[0046] Use exponential function and weight coefficient to increase code complexity , code repetition rate and code coverage Combining the quality scores into an overall score, the code quality score is calculated as follows:

[0047]

[0048] in, is the code quality score, , and is the weight coefficient;

[0049] Configuring scoring thresholds ,If the code quality score is higher than the score threshold, it passes the static verification, otherwise the static verification fails, and it is necessary to reselect the conversion model architecture and re-convert the code.

[0050] Beneficial effects of the present invention:

[0051] 1. By building a conversion model based on a large model, the conversion of image processing algorithms written in different programming languages ​​to Java code is realized. In order to ensure that the converted code is suitable for the target image processing type, the cross entropy loss, code edit distance loss and image type constraint loss are used to construct the total loss function. The comprehensive application of these loss functions helps to improve the accuracy and robustness of the conversion model and ensure that the generated Java code can meet the expected functional requirements.

[0052] 2. By calculating the structural complexity of the code, we can understand the overall complexity level of the code, use the average size and similarity measurement of code blocks to identify and quantify the redundancy in the code, combine branch coverage and function coverage to measure the coverage of the code, and use exponential functions and weighted coefficients to combine code complexity, code repetition rate and code coverage into an overall quality score, providing an objective and comprehensive indicator that helps to quickly judge the quality level of the code and improve the efficiency and quality of the entire development process.

[0053] 3. This method makes the code structure clearer by labeling the source code and dividing it into functional modules. The code is reconstructed and optimized during the conversion process, which improves readability and maintainability and reduces the difficulty of subsequent maintenance. The natural language processing technology is used to extract comments and document string information, making the division of functional modules more accurate and enhancing the code conversion model's understanding of the source code functions. This helps generate Java code that is highly matched with the target image processing type. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a structural diagram of an intelligent code rewriting and verification method for a trusted innovation platform;

[0055] Figure 2A schematic diagram annotating the source code;

[0056] Figure 3 Flowchart for the division of source code functional modules;

[0057] Figure 4 Flowchart to determine module functions for clustering algorithm;

[0058] Figure 5 Flowchart of static verification steps. DETAILED DESCRIPTION

[0059] See also Figure 1 This embodiment introduces an intelligent code rewriting and verification method for a trusted innovation platform, including:

[0060] Step S1: label the source code and use natural language processing technology to extract text information from code comments and document strings for auxiliary label labeling;

[0061] Step S2: Divide the source code function modules according to the label annotations so that each function module corresponds to an image processing type, and obtain the source code module of the target image processing type according to the target information innovation platform;

[0062] Step S3: parsing the source code module of the target image processing type, inputting the parsed code into the conversion model, performing code conversion, and generating a corresponding Java code sequence;

[0063] Step S4: statically verify the Java code sequence generated by the conversion model. If the static verification passes, perform automated testing according to the target image processing type until Java code that can execute the target image processing type is output.

[0064] In this embodiment, an intelligent code rewriting and verification method for a xinchuang platform is provided, which converts image processing algorithms written in different programming languages ​​into Java code, and ensures its correctness and reliability through static verification and automated testing. To improve the portability and maintainability of the code, while reducing development time and cost. First, the source code in the data set is annotated, including the programming language type, source code structure, source code function, input and output parameters. In the source code annotation process, natural language processing technology is used to extract code comments and document strings for text analysis, and extract keywords, phrases and other information. And according to the extracted text information, the source code is divided into different functional modules, such as image loading, image scaling, image rotation, image enhancement, etc. Determine the target image processing type, such as image graying, image edge detection, image denoising, etc., and then obtain the source code module of the target function according to the functional requirements of the target xinchuang platform; parse the obtained source code, involving converting the source code into an abstract syntax tree or other intermediate representation for subsequent processing and conversion. Provide a structured way to understand the structure of the source code, so that code conversion can be performed more accurately. After obtaining the parsed result of the source code, it is serialized into a format suitable for the input of the conversion model. So that the source code can be understood and processed by the conversion model while maintaining the structure and semantic information of the code. The serialized code input is provided to the conversion model so that the conversion model can receive standardized code input, so that accurate code conversion can be performed. After the conversion model receives the input, the code conversion begins. Including operations such as programming language conversion, code refactoring, code optimization, and adding new functions. To generate optimized and improved code, improve the performance, readability and maintainability of the code. Finally, the converted code sequence is output to make it logically consistent with the original code. It is a compilable Java code that can be directly used for further development or deployment. Finally, the static code analysis tool is used to statically verify the converted Java code sequence to check whether the code complies with the Java coding specifications and potential errors. If the static verification passes, write automated test cases according to the target image processing type to perform functional testing on the converted Java code to ensure that it correctly implements the expected functions.

[0065] See also Figure 2 The source code annotation in step S1 includes:

[0066] Annotate the programming language used in the source code, including Python, Java, C++, etc., to determine the syntax and semantic characteristics of the code. This helps to understand the specifications and usage of different programming languages, and provides a basis for subsequent source code analysis. For example, Python's dynamic type characteristics and concise syntax are significantly different from Java's strong type checking and object-oriented programming style. By clarifying the programming language type, you can choose the appropriate parsing tools and techniques to perform more accurate static and dynamic code analysis.

[0067] Marking source code structural elements, including functions, classes, and modules, helps understand the organization and modular design of the code; the structural elements of the source code are the basic components of the code logic. By marking these elements, you can clearly see the hierarchy and module division of the code. For example, a complex source code may be decomposed into multiple classes and modules, each module is responsible for a specific function. Marking these structural elements not only helps understand the function of a single file or module, but also helps understand the architecture and design patterns of the entire system.

[0068] The function of annotating source code is used to describe the specific function or purpose of the code, including sorting algorithms, data structure implementations, etc. This helps to understand the main role and application scenarios of the code. By describing the function of the code in detail, its role and value in the project can be better evaluated. For example, a sorting algorithm may be used for data processing, while a data structure implementation is used to efficiently store and retrieve information. This kind of function annotation can also provide a reference for code reuse and optimization.

[0069] Annotate the source code input parameters and their data types, as well as the returned output results and their data types, including integers, floating-point numbers, and strings. This helps you understand how data is stored and processed in the source code. By annotating the input and output parameters and their data types, you can clearly see the expected input and actual output of a function or method. For example, a function may accept two integers as input and return a Boolean value, which can help you understand the purpose and boundary conditions of the function and avoid type errors and runtime errors.

[0070] The natural language processing technology described in step S1 extracts text information, including:

[0071] Clean and standardize the text data in the source code for subsequent natural language processing tasks. This includes removing punctuation, converting to lowercase letters, and segmenting words. This ensures the consistency and availability of text data. These preprocessing steps can effectively reduce noise data and improve the accuracy of subsequent analysis. For example, removing punctuation can avoid misjudgments caused by special characters, while converting to lowercase letters ensures a unified representation of words and avoids matching problems caused by differences in uppercase and lowercase letters.

[0072] Identify named entities in code comments and document strings, including variable names, function names, and class names. This helps you understand the structure and usage of the code, and better analyze the code's functions and logic. Named entity recognition can extract key elements from the code, quickly locate and understand the core parts of the code. For example, by identifying all variable names and function names, you can quickly understand the input-output relationship and main operation objects of the code, providing a basis for further code analysis and refactoring.

[0073] Extract the relationship between entities from code comments and document strings, including function calls, parameter passing, etc., to understand the logical flow and functionality of the code, as well as the interaction between different parts. Build a dependency graph between code entities through relationship extraction, display function call chains and data flows, so that you can see the code execution path and the interaction between modules more clearly.

[0074] See also Figure 3 The division of the source code function modules in step S2 includes:

[0075] Text search is used to extract tags related to image processing in source code modules;

[0076] Use clustering algorithms to analyze text data on the extracted tags and determine the function of each module;

[0077] According to the function of the module, determine the image processing type corresponding to each module;

[0078] Group modules according to their functions and create file storage modules.

[0079] In this embodiment, a text search tool or a parsing library of a programming language, such as the re library in Python, is used. Function calls, method names, etc. related to image processing in the source code can be efficiently identified. A set of keywords or patterns are defined to match function calls, method names, etc. related to image processing. It helps to quickly locate relevant code snippets and improve search efficiency. Traverse the source code file and search for these keywords or patterns line by line. The matched tags are stored in a list or set to facilitate subsequent analysis and processing. A clustering algorithm is selected, including K-means, DBSCAN, etc., which can effectively group similar tags together to form different categories. The extracted tags are converted into vector representations, and word embedding techniques such as Word2Vec, GloVe, or other feature extraction methods can be used to enable the clustering algorithm to better understand and compare the similarities between tags. The clustering algorithm is applied to cluster the tag vectors to obtain different categories, which helps to discover potential functional modules and image processing types. The features of each category are analyzed to determine the functions represented by each category, providing a basis for subsequent functional division. For each category obtained by clustering, the function represented by the category is determined according to the tags it contains. Map each category to a specific image processing type, such as edge detection, filtering, segmentation, etc. Create a dictionary or mapping table to associate each category with its corresponding image processing type for easy subsequent management and reference. Traverse the source code files and assign code blocks to corresponding files based on the previously determined categories and image processing types. Create a separate file for each image processing type and write the relevant code blocks into it. Add appropriate comments to each file to explain the functions and uses contained in the file. Divide the different functional modules in the source code and organize and manage them according to the image processing type. This embodiment improves the readability and maintainability of the code, facilitating subsequent development and maintenance work. At the same time, by classifying and organizing the source code, the structure of the code can be more clearly understood.

[0080] See also Figure 4 Specifically, the functions of the clustering algorithm determination module include:

[0081] Preprocess the extracted tags, including removing stop words, punctuation, lowercase, and stemming or word form restoration to reduce noise and redundant information and improve the effect of subsequent feature extraction;

[0082] For each tag, count the frequency of each word, that is, ,in, It is The number of times a word appears, is the total number of words in the label and measures the importance of each word in the label, i.e. ,in, is the total number of tags, It includes The number of labels for each word. The denominator is increased by 1 to avoid the denominator being zero.

[0083] The feature vector is constructed by combining the frequency of each word and the importance of each word in the label, that is, ,in, Indicates of the tags The feature vector elements of words, and ; To convert labels into vectors in high-dimensional space to facilitate processing by clustering algorithms;

[0084] Select a clustering algorithm to cluster the labels according to the feature vector. Clustering algorithms include K-means and DBSCAN to group similar labels together to form different categories. Count the number of labels in each category and configure the clustering threshold , obtain the categories whose number of tags is greater than the clustering threshold as key categories to extract keywords with high frequency and large weight;

[0085] According to the keywords of category clustering, a mapping table is created to map them to a specific module function, including edge detection, filtering, segmentation, etc. In this way, the label data can be combined with the actual image processing tasks to provide strong support for subsequent processing.

[0086] Specifically, the specific steps of code parsing in step S3 include:

[0087] Preprocess the source code, including removing irrelevant information such as comments and blank lines, to clean and simplify the source code and make it more suitable for subsequent parsing and conversion. For example, regular expressions or specific text processing libraries can be used to automatically identify and remove comments. At the same time, blank lines and other non-code elements are removed to ensure the continuity and readability of the code;

[0088] Get all the tags in the source code, construct a tag sequence, and convert the tag sequence into an abstract syntax tree to represent the structure of the source code, where each node represents a syntax structure in the source code, including variables, functions, statements, etc. Decompose the source code into a series of tags and arrange them into a linear sequence in the order they appear in the source code. This tag sequence will serve as the basis for constructing the abstract syntax tree;

[0089] Perform semantic analysis on the abstract syntax tree to extract semantic information from the source code, such as variable types, function signatures, control flow, etc. This ensures that the code is logically correct and complies with the language specification. By checking the usage of variables, the parameter types of function calls, and other information, potential errors or inconsistencies can be discovered;

[0090] Serialize the abstract syntax tree into a format suitable for input. Convert the abstract syntax tree into a linear data structure, such as a list or array, so that the conversion model can receive and process it. The serialization process involves traversing the abstract syntax tree and storing the information of each node in a list in a certain order. So that the conversion model can receive and process the serialized data.

[0091] Specifically, the specific steps of converting the abstract syntax tree into a data structure for input include:

[0092] Recursively traverse from the root node of the abstract syntax tree. During the traversal process, visit each node according to the breadth-first search. Use the queue structure to record the access path. Each time a node is visited, its child nodes are added to the end of the queue. The nodes in the queue are visited in sequence to ensure that all nodes are visited in hierarchical order.

[0093] For each node, extract the node information. The node information includes the type of the node , Representation Node The type of the node, such as expression, statement, function definition, etc. , Representation Node Values, such as variable names, constant values, etc., child node lists , Representation Node The collection of child nodes;

[0094] According to the extracted node information, construct the input sequence , ; Each element represents the information of a node, including type, value and child node list;

[0095] Specifically, the specific steps of the code conversion in step S3 include:

[0096] Build a conversion model based on the large model, construct a loss function according to the target image processing type, input the input sequence of the source code module into the conversion model, and convert the source code programming language into Java through the conversion model;

[0097] Perform code refactoring and code optimization, optimize code structure and logic, improve readability and maintainability, improve performance, and reduce resource consumption.

[0098] In this embodiment, a conversion model architecture is selected based on a large model, including Transformer and Seq2Seq. According to the annotation information, the encoder, decoder structure and attention mechanism of the conversion model are designed; the input sequence of the source code module is obtained and the source code is ensured to be complete and contains all necessary dependencies and references. A conversion model suitable for converting the source code programming language to Java is selected. The conversion model can be a pre-trained machine learning model or a rule-based converter. To ensure that the selected model can handle the language characteristics of the source code and can accurately convert it to Java (a description can be added to this part later: an exemplary selection of Transformer to convert the source code to Java format). The input sequence of the source code module is provided to the conversion model, and the conversion process is started. During the conversion process, the conversion model input sequence analyzes the source code line by line and converts it into the corresponding Java code. Once the conversion is completed, the generated Java code is checked to ensure its correctness and completeness. This can be done by compiling the Java code and running the test. If any errors or problems are found, the conversion model needs to be adjusted and reconverted. Once the conversion result is confirmed to be correct, the generated Java code is refactored and optimized, including improving the code structure, improving readability and maintainability, and optimizing performance and resource consumption. After code refactoring and optimization, the generated Java code needs to be tested again to ensure its functionality is correct and performance is improved, including running unit tests, integration tests, and performance tests.

[0099] The construction of the loss function includes:

[0100] Cross entropy loss is used to measure the difference between the predicted result and the true label. In the code conversion task, the source code can be provided as input to the conversion model, and then the corresponding Java code is generated by the decoder. The cross entropy loss calculation formula is as follows:

[0101]

[0102] in, is the number of code nodes, is the number of categories, is the true label, is the probability predicted by the conversion model; by minimizing the cross entropy loss, the model’s prediction accuracy for the true label can be improved, thereby enhancing the accuracy of the conversion result.

[0103] Code edit distance loss is used to measure the similarity between the converted code and the original code. The edit distance algorithm is used to calculate the distance between two code snippets. The smaller the edit distance, the more similar the two code snippets are. The code edit distance loss calculation formula is as follows:

[0104]

[0105] in, is the sample size, is the original node code, is the code of the converted node, It is a function that calculates the edit distance between code snippets. By optimizing the code edit distance loss, the model can better preserve the structure and logic of the original code and ensure the quality of the generated code.

[0106] Image type constraint loss is used to ensure that the converted code is suitable for the target image processing type, requiring that the converted code must contain specific image processing libraries or function calls. The image type constraint loss calculation formula is as follows:

[0107]

[0108] in, is an indicator function, if If the image processing code is included in the required library, it returns 1, otherwise it returns 0, which is used to evaluate whether the converted code meets the requirements of the image processing type. By introducing this constraint, it can ensure that the generated code is not only correct but also meets the requirements of specific application scenarios.

[0109] Combining cross entropy loss, code edit distance loss and image type constraint loss, set the total loss function:

[0110]

[0111] in, , is a weight parameter used to balance the impact of cross entropy loss, code edit distance loss, and image type constraint loss.

[0112] The Java code sequence obtained by the conversion model is statically verified through test cases. If the static verification passes, automated testing is performed according to the target image processing type.

[0113] See also Figure 5 Specifically, the static verification step in step S4 includes:

[0114] Code complexity measures the complexity of the converted Java code structure by the number of conditional statements. and nesting depth , obtain the structural complexity of the code to identify complex code segments to avoid difficulties in maintenance and understanding. The calculation formula for code complexity is as follows:

[0115]

[0116] in, is the code complexity, is the number of edges, which is used to represent the number of connections between different nodes in the code; is the number of nodes, which represent the number of independent logic blocks or decision points in the code; is the number of connected components, which is used to represent the number of independent parts in the code; is the number of conditional statements, which is used to indicate the number of conditional branches contained in the code. is the nesting depth, which is used to indicate the maximum depth of nested structures in the code.

[0117] The redundancy of Java code is evaluated by code repetition rate. The redundant parts in the code are identified and quantified by the average size and similarity measurement of code blocks. The calculation formula of code repetition rate is as follows:

[0118]

[0119] in, is the code repetition rate, is the number of different code blocks, which is used to indicate the number of non-repeated code blocks in the code; is the total number of lines of code, which is used to indicate the total number of lines of code; is the average size of the code block, which is used to indicate the average number of lines in each code block; S is the similarity measure of the code blocks, which is used to indicate the similarity between different code blocks.

[0120] Code coverage measures the degree to which test cases cover the code. High coverage means more comprehensive testing. Combined with branch coverage B and function coverage F, the code coverage of test cases is comprehensively measured to ensure that key functions and boundary conditions are fully tested. The calculation formula for code coverage is as follows:

[0121]

[0122] in, is the code coverage, is the number of test cases executed, which is used to indicate the number of test cases actually run; is the total number of test cases, which is used to represent the number of all possible test cases; is the branch coverage, which is used to indicate the number of conditional branches covered by the test; It is the functional coverage, which is used to indicate the number of function points covered by the test.

[0123] Use exponential function and weight coefficient to increase code complexity , code repetition rate and code coverage It is combined into an overall quality score to highlight the nonlinear relationship between different indicators and provide an integrated perspective for the final quality assessment. The calculation formula of the code quality score is as follows:

[0124]

[0125] in, is the code quality score, , and It is a weight coefficient, which indicates the relative importance of each indicator in the comprehensive evaluation. If a certain indicator is given a high weight, then even if other indicators perform well, the poor performance of this indicator may significantly lower the total score. A high QS value indicates that the overall quality of the code is high, and the various indicators are well developed without obvious shortcomings. A low QS value indicates that there may be one or more serious problems, such as high complexity, high duplication, and low coverage.

[0126] Configuring scoring thresholds ,If the code quality score is higher than the score threshold, the static verification is passed, otherwise the static verification fails, and it is necessary to reselect the conversion model architecture and re-convert the code;

[0127] Specifically, the specific steps of the automated test in step S4 include:

[0128] According to the image processing type of the source code module, the target image processing type to be tested is clearly defined, including image denoising, image enhancement, and image segmentation;

[0129] Obtain test images of the target image processing type, including different image types, such as JPEG, PNG, BMP, different image sizes, and different image types, such as text images, facial images, and low-light images, to ensure that the acquired test images can cover all possible situations; through a variety of test images, the performance and stability of the image processing algorithm can be more comprehensively evaluated;

[0130] Write test cases for the target image processing type. The test cases include input data, operation steps, and expected output. The input data provides the image file or image object for testing; the operation steps describe how to apply the image processing algorithm to the input data; the expected output defines the expected result of the image processing, and the expected result of the image is measured by the output result of the source code.

[0131] Use automated testing frameworks to execute test cases and record test results. Automated testing frameworks include JUnit, TestNG, and Mockito. Automated testing can save time and labor costs while improving the accuracy and reliability of testing. By analyzing test results, problems can be quickly located and fixed.

[0132] According to the test results, the failure cases and reasons for the failure of the automated test are obtained, and the problematic code units are located according to the failure cases; the problematic code units are reconstructed according to the failure reasons using the large model; and the test cases are repeatedly executed to ensure that all problems have been resolved and no new problems have been introduced, and the Java code of the executable target image processing type is output.

[0133] Working principle and its effect:

[0134] An intelligent code rewriting and verification method for the ICT platform converts image processing algorithms in different programming languages ​​into Java code, and uses natural language processing technology to annotate and analyze the source code to ensure that the generated code meets the requirements of the target platform. The effect is to improve the portability and maintainability of the code while reducing development time and cost.

[0135] First, the source code is labeled and natural language processing technology is used to extract text information from code comments and document strings to assist in labeling. This step helps to understand the structure and function of the code and lays the foundation for subsequent steps. According to the labeling, the source code is divided into different functional modules, each corresponding to an image processing type. By analyzing the label data, the code snippets related to the specific image processing task are identified and classified. The source code module of the target image processing type is parsed, and the parsed code is input into the conversion model to generate the corresponding Java code sequence. The conversion model is built on the large model and can convert the source code into Java code. The Java code sequence generated by the conversion model is statically verified to ensure that the code complies with the Java coding specification and check for potential errors. If the static verification passes, the next step of automated testing is carried out. Automated test cases are written according to the target image processing type to perform functional testing on the Java code to ensure that it correctly implements the expected functions. After the test passes, executable Java code is output.

[0136] Through static verification and automated testing, the converted Java code is ensured to be of high quality and free of potential errors, thereby improving the reliability and stability of the code. During the code conversion process, the code was refactored and optimized to improve the readability and maintainability of the code, while improving performance and reducing resource consumption. The automated code conversion, verification, and testing process greatly reduces the time and cost of manual operations and improves development efficiency. It provides an intelligent code rewriting and verification solution for the Xinchuang platform, which helps promote technological innovation and development.

[0137] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. An intelligent code rewriting and verification method for a trusted innovation platform, characterized in that: include: Label the source code and use natural language processing technology to extract text information from code comments and document strings for auxiliary labeling; Divide the source code function modules according to the label annotations so that each function module corresponds to an image processing type, and obtain the source code module of the target image processing type according to the target information innovation platform; Parse the source code module of the target image processing type, input the parsed code into the conversion model, perform code conversion, and generate a corresponding Java code sequence; Static verification is performed on the Java code sequence generated by the conversion model. If the static verification passes, automated testing is performed according to the target image processing type until Java code that can execute the target image processing type is output; The specific steps of the code conversion include: Build a conversion model based on the large model, construct a loss function according to the target image processing type, input the input sequence of the source code module into the conversion model, and convert the source code programming language into Java through the conversion model; Perform code refactoring and code optimization to optimize code structure and logic; The construction of the loss function includes: Construct the cross entropy loss to measure the difference between the predicted result and the true label. The cross entropy loss calculation formula is as follows: ,in, is the number of code nodes, is the number of categories, is the true label, is the probability predicted by the transformation model; The code edit distance loss is constructed to measure the similarity between the converted code and the original code. The code edit distance loss calculation formula is as follows: ,in, is the sample size, is the original node code, is the code of the converted node, is a function that calculates the edit distance between code snippets.

2. The intelligent code rewriting and verification method for the credential creation platform as claimed in claim 1, characterized in that: The division of the source code function modules includes: Text search is used to extract tags related to image processing in source code modules; Use clustering algorithms to analyze text data on the extracted tags and determine the function of each module; According to the function of the module, determine the image processing type corresponding to each module; Group modules according to their functions and create file storage modules.

3. The intelligent code rewriting and verification method for the credential creation platform as described in claim 2 is characterized in that: The clustering algorithm determines the module function specifically including: Preprocess the extracted tags; For each tag, count the frequency of each word, that is, ,in, It is The number of times a word appears, is the total number of words in the label and measures the importance of each word in the label, i.e. ,in, is the total number of tags, It includes The number of tags for each word; The feature vector is constructed by combining the frequency of each word and the importance of each word in the label, that is, ,in, Indicates of the tags The feature vector elements of words, and ; Select a clustering algorithm, cluster the labels according to the feature vector to group similar labels together to form different categories, count the number of labels in each category, and configure the clustering threshold , obtain the categories whose number of labels is greater than the clustering threshold as key categories; Based on the keywords of the key categories, the key categories are mapped to an image processing function category according to a mapping table.

4. The intelligent code rewriting and verification method for the trust innovation platform as described in claim 1 is characterized in that: The specific steps of the code parsing include: Preprocess the source code, including removing comments and blank lines; Get all tags in the source code, construct a tag sequence, and convert the tag sequence into an abstract syntax tree; Perform semantic analysis on the abstract syntax tree to extract semantic information from the source code, including variable types, function signatures, and control flows; Serialize the abstract syntax tree into input format and convert the abstract syntax tree into a linear data structure, including lists and arrays.

5. The intelligent code rewriting and verification method for the credential creation platform as claimed in claim 4, characterized in that: The specific steps of converting the abstract syntax tree into a data structure for input include: Start recursively traversing the root node of the abstract syntax tree, and visit each node according to the breadth-first search process; For each node, extract node information, which includes the type of node , the value of the node , child node list ; Construct the input sequence based on the extracted node information.

6. The intelligent code rewriting and verification method for the credential creation platform as claimed in claim 1, characterized in that: The construction of the loss function also includes: Construct an image type constraint loss to ensure that the converted code is suitable for the target image processing type. The image type constraint loss calculation formula is as follows: ,in, is an indicator function, if If the image processing code is included in the required library, it returns 1, otherwise it returns 0, which is used to evaluate whether the converted code meets the requirements of the image processing type; Combining cross entropy loss, code edit distance loss and image type constraint loss, the total loss function is constructed: ,in, , is the weight parameter.

7. The intelligent code rewriting and verification method for the trust innovation platform as claimed in claim 1 is characterized in that: The static verification step includes: The structural complexity of the code is obtained by the number and nesting depth of conditional statements. The calculation formula of code complexity is as follows: ;in, is the code complexity, is the number of edges, is the number of nodes, is the number of connected components, is the number of conditional statements, is the nesting depth; The redundant parts in the code are identified and quantified by measuring the average size and similarity of the code blocks. The formula for calculating the code repetition rate is as follows: ,in, is the code repetition rate, is the number of distinct code blocks, is the total number of lines of code, is the average size of code blocks, S is the similarity measure of code blocks; Combining branch coverage and function coverage to measure the coverage of the code, the calculation formula of code coverage is as follows: ,in, is the code coverage, TC is the number of test cases executed, is the total number of test cases, B is the branch coverage, and F is the functional coverage.

8. The intelligent code rewriting and verification method for the trust innovation platform as claimed in claim 1 is characterized in that: The static verification step also includes: Use exponential function and weight coefficient to increase code complexity , code repetition rate and code coverage Combining the quality scores into an overall score, the code quality score is calculated as follows: ,in, is the code quality score, , and is the weight coefficient; Configuring scoring thresholds ,If the code quality score is higher than the score threshold, it passes the static verification, otherwise the static verification fails, and it is necessary to reselect the conversion model architecture and re-convert the code.

Citation Information

Patent Citations

  • Code generation method and device, storage medium and server

    CN115145574A

  • Intelligent transcoding adaptation method and system for credential technology and application platform

    CN118113292A