Code automatic generation and optimization system based on multiple modes
Through multimodal data fusion and optimization algorithm, the problems of implicit constraints and domain knowledge recognition in code generation are solved, and efficient and accurate code generation and optimization are achieved.
Patent Information
- Application Number
- CN202510764201.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing code generation technologies are difficult to accurately capture implicit constraints and domain knowledge, resulting in a semantic gap between the generated code and the real needs. The convergence efficiency of static analysis tools is inefficient and the code generation effect is poor.
A multimodal-based code automatic generation and optimization system is adopted, and through data processing, task analysis, data fusion, optimization decision-making, static analysis and test evaluation modules, combined with the CLIPS rule engine, knowledge graph and particle swarm optimization algorithm, it dynamically adapts multimodal data to optimize the code generation strategy.
It improves the accuracy of task intention recognition, reduces code generation bias, reduces computing resource consumption, and realizes the global optimization of code generation strategy.
Smart Images

Figure CN120276718A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic programming, and particularly to a multi-modal based code automatic generation and optimization system. Background Art
[0002] Currently, code automatic generation technologies usually adopt sequence-to-sequence models, abstract syntax tree generation models or template-based semantic transformation methods to map the requirement descriptions input by users into executable codes. Also, unstructured or semi-structured information such as log data, historical code snippets, development documents, and runtime feedback is also very important in the development process, providing great help for assisting task semantic understanding and code context construction.
[0003] However, traditional methods are also difficult to accurately capture the implicit constraints and domain knowledge in code generation requirements, which will cause some semantic gaps between the generated codes and the real requirements. In the code optimization link, existing static analysis tools detect code smells based on rules, and most dynamic optimization methods have a large search space, resulting in relatively low convergence efficiency and greatly reducing the code generation effect. Therefore, a mechanism with optimized modal data selection strategy is very important. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a multi-modal based code automatic generation and optimization system to solve the problems of insufficient dynamic adaptability of multi-modal data fusion and lack of cross-modal collaboration in code generation path optimization.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: The present invention provides a multi-modal based code automatic generation and optimization system, which includes a data processing module that collects raw data, performs data cleaning and preprocessing, and constructs a raw data set; A task analysis module that combines the CLIPS rule engine and a knowledge graph, analyzes task descriptions and constraint conditions in combination with the raw data, identifies task goals and requirements, and outputs a task risk assessment report and a task intention set; A data fusion module that selects data sources based on the task risk assessment report and the task intention set, performs data fusion using the weighted average method, and outputs a fused multi-modal data set; An optimization decision module that uses a particle swarm optimization algorithm and a path planning algorithm to select the best modal data subset, and performs optimization adjustment according to task requirements, and outputs a code generation strategy; A static analysis module that uses an abstract syntax tree and a control flow graph to perform static analysis on the code generation strategy, identifies performance bottlenecks and potential errors and performs optimization, and generates preliminary codes; The test evaluation module uses mutation testing to evaluate and improve unit tests and integration tests, and conducts quality and performance evaluations on the code through continuous integration and continuous delivery, outputting a code evaluation report and optimization suggestions.
[0007] As a preferred solution of the multimodal-based code automatic generation and optimization system of the present invention, wherein: the data processing module collects raw data, performs data cleaning and preprocessing, and constructs a raw data set. The specific steps are as follows: Collect raw data using the Representational State Transfer (REST) application programming interface, establish a data directory structure archive according to the task ID, and output a structured raw data sample set; Automatically complete missing values using placeholder diagrams, and then use Z-value standardization and Isolation Forest algorithm for data standardization and data anomaly detection, outputting a high-quality data set after cleaning; Extract text using key phrase extraction, identify function names to extract function structures from code, and perform structured extraction on log parsing event sequences, and reorganize them according to unified field specifications to construct data entries with a unified structure; Generate a unique hash identifier for the data entries, and combine rule label extraction to generate a raw data set.
[0008] As a preferred solution of the multimodal-based code automatic generation and optimization system of the present invention, wherein: the task analysis module combines the CLIPS rule engine and the knowledge graph, analyzes the task description and constraint conditions in combination with the raw data, identifies the task objectives and requirements, and outputs a task risk assessment report and a task intention set. The specific steps are as follows: Use the domain knowledge graph and the natural language processing toolkit to map text semantics to entity nodes and attribute edges in the graph, and construct a task semantic graph; Adopt the CLIPS rule template to define behavior objectives, resource requirements, and context constraints, substitute the graph entities into the preset rules for reasoning, and output a task intention set; Based on the multi-level reasoning and task dependency graph construction mechanism of CLIPS, record the trigger paths of each reasoning conclusion, determine the conflict source through backtracking, generate a dependency chain, and output a structured risk node chain graph; Combine the text generation template and the structured export tool, and comprehensively transform the task intention set and the risk node chain graph into a readable report, and output a task risk assessment report.
[0009] As a preferred solution of the multimodal-based code automatic generation and optimization system of the present invention, wherein: the data fusion module selects data sources based on the task risk assessment report and the task intention set, uses the weighted average method for data fusion, and outputs a fused multimodal data set. The specific steps are as follows: Use natural language processing methods to identify and extract task elements from task risk assessment reports and task intent sets, utilize dependency syntax analysis to process the relationships between task nodes, and combine with a knowledge graph to output a task node combination and task node dependency graph; Apply relationship extraction to analyze the association between task nodes and data sources, and perform priority matching through a weighted matching algorithm to output an accurate mapping of data requirements and data sources; Based on the task node combination, task node dependency graph, and accurate mapping of data requirements and data sources, use weighted average fusion. Through time series alignment and data preprocessing, process and fuse different modality data to generate a multi-modal data set in a unified format.
[0010] As a preferred solution of the multi-modal based code automatic generation and optimization system of the present invention, wherein: the optimization decision module uses a particle swarm optimization algorithm and a path planning algorithm to select the best modality data subset, and perform optimization adjustment according to task requirements, and output a code generation strategy. The specific steps are as follows. Use a particle swarm optimization algorithm to evaluate the fused multi-modal data set and calculate the best modality data subset that meets the code generation task; Plan the transmission path of the best modality data subset through the A algorithm, and optimize the path of the selected modality data during transmission to output an optimized data transmission path; Based on the selected best modality data subset and data transmission path, use the gradient descent method for multi-objective optimization, dynamically fine-tune the feature representation according to task requirements, and output modality data matching the task; Use a preset code generation template to convert the modality data and data transmission path matching the task into a code execution strategy for task requirements.
[0011] As a preferred solution of the multi-modal based code automatic generation and optimization system of the present invention, wherein: the static analysis module uses an abstract syntax tree and a control flow graph to perform static analysis on the code generation strategy, identify performance bottlenecks and potential errors and optimize them to generate preliminary code. The specific steps are as follows. Use a syntax parser to convert the code generation strategy into a weighted abstract syntax tree, parse the code structure, generate weights for each node, mark the data modality source, and output an abstract syntax tree with weights and modality labels; Generate a control flow graph based on the abstract syntax tree, use a control flow analysis algorithm to calculate edge weights, identify execution paths, and construct a control flow graph with edge weight markings; Based on the control flow graph, use the proximal policy optimization algorithm to optimize the code execution strategy and output an optimized action sequence; Use a code converter to reconstruct the optimized action sequence into an abstract syntax tree to generate preliminary code.
[0012] As a preferred solution of the multi-modal based code automatic generation and optimization system described in the present invention, where: the test evaluation module uses mutation testing to evaluate and improve unit testing and integration testing, evaluates the quality and performance of the code through continuous integration and continuous delivery, and outputs a code evaluation report and optimization suggestions. The specific steps are as follows: By performing syntax replacement, boundary perturbation, and control structure adjustment on the preliminary code, multiple code mutants with different structures are constructed to form test evaluation samples, and a set of code mutants with diverse structures is output; Apply the existing unit test set to the set of code mutants, record the killing effect of test cases and generate a coverage matrix, and output a reconstructed high-coverage unit test set through coverage clustering and hole analysis; Based on the reconstructed unit test set, extract module control dependencies and interface call paths, generate a new control flow graph, automatically generate an integration test set through a path-driven mechanism, and output a path-complete integration test set; Through continuous integration and continuous delivery, perform automated testing, collect mutation coverage rate, performance metrics, and resource utilization evaluation parameters, and generate a code evaluation report and optimization suggestions.
[0013] The beneficial effects of the present invention are: by semantically aligning natural language task descriptions with multi-modal data such as code, logs, and requirement documents, the accuracy of task intention recognition is improved, and the code generation deviation caused by single-modal analysis is reduced; through dynamically screening the optimal modal data subset and optimizing the data transmission path, while ensuring the code performance, the consumption of computing resources is reduced, and the global optimality of the code generation strategy is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 FIG. is a flowchart of a multi-modal based code automatic generation and optimization system.
[0016] Figure 2 FIG. is a flowchart of a data acquisition and preprocessing module.
[0017] Figure 3 FIG. is a multi-modal data fusion and task matching diagram.
[0018] Figure 4 FIG. is a flowchart of code generation and optimization testing. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings of the specification.
[0020] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0021] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.
[0022] Referring to Figures 1 to 4 , which is an embodiment of the present invention. This embodiment provides a multi-modal based code automatic generation and optimization system, including the following steps: A data processing module, which collects raw data, performs data cleaning and preprocessing, and constructs a raw data set.
[0023] Furthermore, use the Representational State Transfer (REST) application programming interface to collect raw data, establish a data directory structure for archiving according to the task ID, and output a structured raw data sample set; Specifically, use the Representational State Transfer (REST) application programming interface to call the data collection interfaces of external platforms and local services to receive data, including natural language descriptions, historical code snippets, development logs, compilation feedback information, etc. Attach a unique task identifier as a request parameter for each collection task, and output a raw multi-modal data set organized by task ID; establish a corresponding index table, record the task ID, collection time, data type and file path, and output task-level archived data with a directory structure and metadata index; parse the data content, extract fields, and fill in default values for irregular fields, and output a raw data set with a unified structure and standardized fields; Preferably, through a multi-level data collection and reorganization process, the efficiency and quality of data processing are improved. Using the Representational State Transfer (REST) application programming interface for data collection, the unified definition of data formats and the archived management of task IDs improve the standardization level of data collection and the automation and standardization of data collection; the data directory structure constructed by task ID makes each piece of data have a clear index and hierarchy, improving maintainability and flexibility, and enhancing the organization and manageability of data; Automatically complete the missing values using placeholder graphs, and then use Z - value standardization and isolation forest algorithm for data standardization and data anomaly detection to output a high - quality dataset after cleaning; Specifically, for the missing values in the high - quality dataset, use the placeholder graph technology to mark and generate a marking matrix. Use the placeholder graph and matrix decomposition algorithm to complete the data, infer a reasonable filling scheme for the missing values, and construct a high - quality dataset after completion. Perform standardization processing on the high - quality dataset, adjust the mean of the data to 0 and the standard deviation to 1, and output the standardized dataset, where the mean of each feature is 0 and the standard deviation is 1, and all features are on the same scale. Use the isolation forest algorithm to calculate the isolation depth of each data point in all trees for each data point, measure its anomaly degree through the average depth, calculate the outliers and remove them from the dataset, and output the high - quality dataset after cleaning; Preferably, automatically completing the missing values using placeholder graphs improves the continuity and integrity of the data. By using the matrix decomposition method to infer the filling of the missing values, the internal correlation between the data is retained to improve the data quality. Z - value standardization eliminates the dimensional difference between the feature values, enabling data in different dimensions to be compared under the same standard, enhancing the scientificity and reliability of data analysis. Using the isolation forest algorithm improves the reliability and accuracy of the data and enhances the stability of data analysis; Adopt key - phrase extraction to extract the text, identify function names to extract the function structure from the code, and perform structured extraction on the log to parse the event sequence, and reorganize it according to the unified field specification to construct data entries with a unified structure; Specifically, adopt key - phrase extraction to perform word segmentation, part - of - speech tagging and stop - word filtering on the task - description text, and output the corresponding words of the task - description text. Use the abstract syntax tree to perform lexical analysis and syntactic parsing on the original code text, extract the function - definition nodes in the abstract syntax tree, and output the function - level structured code containing function names, function bodies, comments and call structures. Use a sliding window to read the log text, perform line - by - line processing, apply the template mining algorithm to perform template clustering on the log format, extract the variable and constant parts, and output the formatted event - log sequence. Map the extracted keywords, function structures and log events to the unified fields respectively, perform association and integration based on the task ID, make all modal data correspond to a task unit, and then perform normalization processing on the field content to construct data entries with a unified structure; Preferably, by using the placeholder map method to complement the missing items in the original data for structural consistency, the integrity and processability of the data samples are improved; the Z-value standardization method is used to perform scale normalization on numerical data, enhancing the comparability and calculation stability between different data features; combined with the isolation forest algorithm to detect outliers in the multi-dimensional space of the data set, the ability to identify outlier samples and abnormal distributions is improved, thereby enhancing the robustness of the overall data quality; Generate a unique hash identifier for the data entry, and combine rule label extraction to generate the original data set; Specifically, splice the fields in the data entry, perform standard encoding on the spliced string, and pass it into the hash function to calculate the digest value to generate a unique hash identifier for the data entry; perform rule matching on the text field and the code field respectively according to the network configuration, interface call and business logic, and assign one or more groups of labels to the data entries that meet the conditions to generate the original data with semantic context markers; Preferably, through the hash identifier generation mechanism, the uniqueness, integrity and immutability of the data entry are ensured, and at the same time, an accurate identifier is provided for the data; through the rule label extraction mechanism, the semantic hierarchical management ability and automatic classification efficiency of the original data are improved, and the indexing and refined data scheduling speed are increased; the original data set constructed after the combination of the two has traceability, scalability and retrievability.
[0024] The task analysis module, in combination with the CLIPS rule engine and the knowledge graph, combines the original data analysis task description and constraint conditions to identify the task objectives and requirements, and outputs a task risk assessment report and a task intention set.
[0025] Use the domain knowledge graph and the natural language processing toolkit to map the text semantics to the entity nodes and attribute edges in the graph, and construct a task semantic graph; It should be noted that the domain knowledge graph and the natural language processing toolkit are used to map the text semantics to the entity nodes and attribute edges in the graph, and construct a task semantic graph; Use the natural language processing toolkit to identify the core semantic units in the task text, perform word segmentation, part-of-speech tagging, dependency syntax analysis and named entity recognition, extract the task actions, objects and conditional elements, and output a set of structured semantic units; Specifically, the text regularization module is used to clean the input task description text, including deleting invalid characters and unifying punctuation formats, converting all letters to lowercase using a format standardization tool, and unifying the expressions of numbers, units, symbols, etc.; a word segmentation engine is used to finely split the text, cutting continuous text into basic word units with semantic meanings to form a preliminary word sequence; a part-of-speech recognition tool is used to assign part-of-speech tags to each word segmentation result, combined with a syntactic structure analyzer to identify the grammatical dependency relationships between words, and an entity recognition tool is used to identify the specific proprietary entities in the text; finally, three types of key semantic elements in the task description, task actions, task objects, and conditions and constraints, are extracted, and these semantic elements are output in a structured format to construct a set of structured semantic units; Preferably, through task text cleaning and standardization processing, the quality and consistency of the input text are improved, noise symbols and non-standard expressions are eliminated, and the parsability and accuracy of the text in subsequent processing are ensured. Through word segmentation and part-of-speech tagging, the understanding ability of each word in the text is improved. Through the output of structured semantic units, the understanding depth and flexibility of complex task descriptions are improved; Combined with the domain knowledge graph and the predefined term mapping rule table, the set of semantic units is bound to the entity nodes in the graph, and a list of semantic units bound to the graph entity nodes is output; Specifically, a knowledge graph is used to build domain-specific knowledge resources to construct a domain knowledge graph. The entity nodes in the knowledge graph represent concepts or objects in the domain, and the relationships between entity nodes are represented by edges. According to the terms and concepts in the specific domain, a term mapping rule table is established, and the table lists the mapping relationships between terms and entity nodes in the knowledge graph. For the semantic units extracted from the task text, by comparing the term mapping rule table, the mapping relationships with the entity nodes in the knowledge graph are identified, and the semantic units in the task text are bound to the entity nodes in the knowledge graph, and a list of semantic units bound to the graph entity nodes is output; Preferably, through the combination of the knowledge graph and the term mapping rule table, the understanding ability of the key information in the task description is improved; by binding the semantic units to the entity nodes, the structuring and semantic layering of the task description are improved; through the application of the domain knowledge graph, the domain adaptation ability is improved, and the domain knowledge can be efficiently used to deeply analyze the task description, identify the core concepts and entities involved in related tasks, and thus enhance the ability to process complex domain tasks; through the support of the predefined term mapping rule table, the flexibility and accuracy of processing domain-specific terms are improved, making the terms in the task consistent with the entity nodes in the knowledge graph; Using the graph edge relation template library and the graph structure generation algorithm, construct semantic relation edges and generate a task semantic graph. Construct attribute edges according to the semantic dependencies between entities, and combine the graph reasoning mechanism to complete the logical relations to construct a task semantic graph; Specifically, call the edge relation template library to identify the logical relations between semantic units, construct semantic edges according to the matching rules, and output the structured description of the preliminary semantic relation edges; use the graph structure generation algorithm to organize them into a directed graph structure to generate a preliminary semantic graph, where nodes represent semantic units and edges represent semantic relations, and construct a preliminary task semantic graph structure; use a custom semantic dictionary to match and annotate the context content of each node in the preliminary semantic graph, and combine the preset conditional rule set to extract the restrictive or conditional semantics between nodes, and add the extracted conditions to the graph in the form of attribute edges to output an enhanced semantic graph structure containing condition restrictions and task attributes; use the graph traversal algorithm to traverse the entire semantic graph, identify redundant paths, duplicate nodes or redundant attribute edges in the structure, streamline multiple equivalent edges and duplicate connections that do not affect logical expression for the same semantic unit, and fold the structure of links with too large a logical level span to construct a task semantic graph with optimized structure; Preferably, through natural language processing and text preprocessing operations, the noise, redundancy and ambiguous expressions in the descriptive text are eliminated, and the purity and structuring degree of task information extraction are improved; through core semantic unit extraction and standardized expression, the unstructured elements such as actions, objects and conditions in natural language are transformed into a unified and standardized structured semantic sequence, and the consistency and operability of semantic recognition are improved; through the joint matching of the domain knowledge graph and term rules, the accuracy and domain pertinence of strong semantic understanding are improved; Use the CLIPS rule template to define the behavior goals, resource requirements and context constraints, substitute the graph entities into the preset rules for reasoning, and output a set of task intents; Specifically, use the Harbin Institute of Technology Language Cloud combined with knowledge graph query technology to analyze the task semantic graph, read the entity nodes related to the task description in the graph, and extract their attribute labels and relation edges to output a list of entity information, including the structured results of graph elements such as task behaviors, data resources, and environmental constraints; use an open-source rule engine platform as the core reasoning tool, load the preset rule files and task conditions, define task behaviors through rule templates, and use the extracted graph entities as facts to input and inject them into the CLIPS environment to output the list of activated rules and the corresponding reasoning paths; use the built-in forward chaining reasoning mechanism of CLIPS to execute the eligible rules in turn to generate new task reasoning results, and output a structured task intent set containing task goals, resource requirements and applicable conditions; It should be noted that the preset rules include behavior target rules, resource requirement rules, and context constraint rules. The behavior target rules use the condition pattern matching mechanism of CLIPS to match according to the entity node types and their semantic attributes in the semantic graph. The resource requirement rules are based on the graph relationship structure to match the edge types between task nodes and the resource entities they depend on. The context constraint rules use CLIPS to perform constraint matching on the context fields in the rules. Preferably, by using the combined mechanism of the knowledge graph and the rule engine, semantic ambiguity in the natural language parsing process is reduced, the logical consistency of task recognition and the stability of the output are enhanced, and the accuracy and consistency of task semantic recognition are improved. The explicit reasoning method based on rules is adopted to enhance the interpretability and controllability of the task reasoning process. Based on the multi-level reasoning and task dependency graph construction mechanism of CLIPS, record the trigger paths of each reasoning conclusion, determine the conflict source through backtracking, generate a dependency chain, and output a structured risk node chain graph. Specifically, through the built-in rule matching of CLIPS, according to the rule priority and the status of the fact set, multi-round forward reasoning is carried out starting from the basic entities. The reasoning process progresses in a hierarchical manner, and the output of each level is used as the input of the next level, and the task objectives, constraint relationships, dependency nodes, and corresponding triggered rule sequences at each stage are generated. Combining the graph structure construction algorithm and the dependency parsing strategy, the conclusions generated in each round of reasoning are used as nodes in the graph, and the causal relationships between different reasoning stages are used as edges to generate a directed dependency graph and a structured task dependency graph. During the rule execution process, track and log the facts, preconditions, and conclusions relied on by each rule trigger to form a complete reasoning path chain and output a reasoning path record table. Based on the backtracking analysis of the dependency path and the rule conflict detection algorithm, starting from abnormal reasoning conclusions or resource conflict nodes, trace all trigger paths backward along the dependency graph, analyze rule overlap, resource duplicate reference, or logical contradiction situations, and output a conflict node list. Combining the graph structure visualization tool and the structured data export tool, export the identified conflict paths, related nodes, and upstream and downstream dependency relationships in a structured graphic form, and generate a risk node chain graph for further viewing by users to output a structured risk node chain graph. Preferably, by recording the trigger path and its preconditions of each rule, the reasoning process is transparent and traceable, enhancing the interpretability and traceability of task reasoning. The risk node chain graph not only has a high degree of structure but also can directionally display high-risk paths or nodes, with pertinence and practicality. The rules of the CLIPS rule inference engine have strong combinability. Combining the dependency graph mechanism, it strengthens the knowledge expansion ability and rule reuse ability. The robustness and logical consistency of the dependency analysis and conflict recognition mechanism are enhanced, improving the logical consistency and fault tolerance of the overall operation. Combine text generation templates and structured export tools to transform the mission intent set and risk node chain diagram into a readable report and output the mission risk assessment report; Specifically, use the Microsoft document format processing tool to convert the task intent set and risk node chain diagram into unified data, establish field mapping relationships, including target descriptions, dependency paths, risk levels, conflict sources, etc., and output a standardized intermediate information representation structure; preset multiple text generation templates for different task types, risk levels, and dependency chain paths, automatically fill in key content through placeholders and field mapping, and preliminarily generate task risk assessment statement paragraphs and task intent natural language descriptions; use the structured export tool to combine the task intent text, risk node chain diagram visualization chart, and related annotations into a complete report template. The chart part combines the graph structure drawing tool to display the dependency relationship and risk path. The text part uses automatic insertion to complete the organization of the text content, and outputs a complete task risk assessment report containing task objectives, intent descriptions, risk path diagrams, and structured descriptions; use document conversion and archiving tools to perform version number marking and directory archiving, support retrieval, review, and version tracing, and output a structured archived task risk assessment report.
[0026] The data fusion module selects the data source based on the mission risk assessment report and the mission intent set, uses the weighted average method to perform data fusion, and outputs the fused multimodal data set.
[0027] Furthermore, natural language processing methods are used to identify and extract task elements of task risk assessment reports and task intent sets, dependency syntax analysis is used to process the relationship between task nodes, and knowledge graphs are combined to output task node combinations and task node dependency graphs; Specifically, use Jieba word segmentation, regular expressions and lexical filtering rules to sentence-break the input text, remove stop words, unify terminology, and output standardized task text corpus; design multi-layer matching templates, extract task elements based on verb phrases, noun phrases and prepositional structures, and output a set of task element triples; use the Stanford syntax analyzer to perform grammatical structure analysis on sentences, and extract subject, predicate, object, modifiers and parallel structures in task sentences. Establish node connections through syntactic relationships, generate a task node combination graph, in which nodes represent task behaviors and edges represent grammatical or semantic connections, and construct a task node combination graph; perform entity matching and relationship alignment on the extracted task elements and the predefined domain knowledge graph, use string similarity calculation and attribute mapping rules to connect semantic nodes to entities and their upper and lower hierarchies in the graph, expand the logical dependency path, construct a task node dependency graph, and output task node combinations and task node dependency graphs; Preferably, by using natural language processing methods, the ability to understand the semantics of task texts can be improved, especially achieving high-precision structured extraction in identifying task risks and intention-related elements; integrating knowledge graphs strengthens term alignment, semantic disambiguation, and context completion capabilities, enabling semantic enhancement and standardized mapping of task information; the overall composite process improves the automation of task information extraction and the docking efficiency of upstream and downstream modules, enhancing intelligent decision-making capabilities and generalization adaptability; Apply relationship extraction to analyze the association between task nodes and data sources, and perform priority matching through a weighted matching algorithm to output an accurate mapping of data requirements and data sources; Specifically, describe the corresponding rules between keywords and data source fields, parse the task description text, extract keywords and their dependency relationships, compare with the dictionary library and template rules, determine which task nodes have a call requirement for the data source, and construct a preliminary association table between task nodes and data fields; use a weighted matching algorithm to determine the priority and the optimal data source. By setting multiple matching weight factors, such as semantic similarity scores, context consistency weights, historical frequency factors, and data availability weights, etc., each pair of candidates is weighted and scored, the results are sorted, and the highest-scoring matching item corresponding to each task node is extracted to output an accurate mapping of data requirements and data sources; Preferably, through a combination of rule matching and weighted scoring, the implicit data usage intention in the task description can be captured more accurately, thus achieving an accurate connection between the data source and task requirements and realizing a high-precision static matching between task requirements and data sources; the rules and weights have high controllability and strong transparency, which are convenient for debugging and modification, enhancing interpretability and maintainability; Based on task node combination, task node dependency graphs, and the accurate mapping of data requirements and data sources, use weighted average fusion. Through time series alignment and data preprocessing, process and fuse different modality data to generate a multi-modal data set in a unified format; Specifically, the task node combination, the task node dependency graph, and the precise mapping of data requirements and data sources are uniformly transformed into key-value pairs or table forms by using format conversion and unit standardization. The numerical units are unified and normalized, and the timestamp format is standardized to construct a multi-modal intermediate data table with consistent format and standardized units. The task node dependency graph is processed by graph traversal, and the data fusion order between modalities is established according to the logical relationship between nodes. It is clear what kind of fusion operation should be performed between each pair of task nodes, such as time synchronization alignment, field merging and splicing, or redundant field elimination. A weight value is set for each fusion relationship to represent its importance in the fusion, and a fusion process control table and a fusion weight configuration table are generated. The data with timestamps is matched using the sliding window method, and the data with the closest time is selected for pairing. The points that cannot be aligned are filled in by linear interpolation, and a data sample table with consistent time dimension and cross-modal alignment is output. The data fields from different modalities are weighted and averaged according to the preset weights. For non-numerical fields, a weight-based selection mechanism is executed, and the one with the largest weight is the main field, and the others are merged as annotations. A standardized sample data set after multi-modal fusion is output, and the field values represent the fusion results. The field integrity of the fused result is checked and output according to the unified structure template to generate a multi-modal data set with a unified format. Preferably, by unifying the structure, standardizing the key-value format, and aligning the time, the structural consistency and semantic coherence between multi-modal data are improved; by using a weighted average fusion method based on task weights and data confidence, the representativeness and task adaptability of the data fusion process are enhanced; by clarifying the task node dependency relationship and the data source mapping path, the operability and controllability of data organization and processing are improved; overall, the data integration ability, time series coordination ability, and high-quality input construction ability in a multi-source heterogeneous data environment are enhanced.
[0028] Optimize the decision-making module, use the particle swarm optimization algorithm and the path planning algorithm to select the best modal data subset, and optimize and adjust according to the task requirements, and output the code generation strategy.
[0029] Furthermore, use the particle swarm optimization algorithm to evaluate the fused multi-modal data set and calculate the best modal data subset that meets the code generation task. Specifically, different types of modal data specifications are converted into a unified structured form, and a set of structured unified modal features is output as the input basis for particle swarm search; by calculating the matching scores between task requirement keywords and modal content keywords, penalties are set according to the redundancy weights between modal contents to measure the degree of support of any modal combination for the current task objective, and a quantifiable fitness function is output for ranking the advantages and disadvantages of modal combinations; each particle represents a Boolean vector encoding whether to select each type of modality, and the number of particles, the maximum number of iterations, the speed range, and the convergence threshold are set. In each iteration, the individual best and global best are updated according to the fitness of the current modal combination, and the position and speed formulas are used to iteratively search for the optimal modal combination, and the Boolean vector of the modal combination with the best fitness is output, that is, the best support modal subset for the task. It should be noted that the expression of the fitness function is as follows: ; Among them, represents the fitness of the given modality and measures the degree of matching of the modality to the target task requirements. The higher the fitness value, the better the modality meets the task requirements and the better its performance. (−∞, + ) represents the weight of the coverage-related part, which controls the influence degree of coverage in calculating the fitness. (0, + ) measures the matching or coverage degree between the modality and the target task . It indicates the extent to which the modality can meet the requirements of the task . [0, 1] represents the weight of the overlap-related part, which controls the influence degree of overlap in the fitness calculation. is just the name of the function and does not represent a value or constant by itself. It must be used in conjunction with the input variables and . (0, + ) represents the redundancy or similarity between the sub-modalities in the modality . [0, 1], where 0 means that the modality completely does not match the task , and 1 means that the modality completely covers the requirements of the task . It is just the name of a function, which does not represent a value or a constant by itself and must be used in conjunction with input variables. for use.
[0030] Plan the transmission path of the optimal modal data subset through the A algorithm, optimize the path of the selected modal data during transmission, and output the optimized data transmission path. It should be noted that the transmission path of the optimal modal data subset is planned through the A algorithm, the path of the selected modal data during transmission is optimized, and the optimized data transmission path is output. Furthermore, calculate the data priority based on the task requirements, use the particle swarm optimization algorithm to evaluate the priority of different modal data, and output the priority list. Specifically, by analyzing the task description text, identify the task objectives, constraints, and data dependencies, structurally represent the task requirements through the entities and attributes in the knowledge graph to form the semantic vector of the task, and generate the structured data of the task requirements; analyze the characteristics of each modality, encode these characteristics, and compare them with the task requirements to calculate the correlation and influence degree between the data and the task, generate the feature vector of each modality data, indicating its matching degree with the task and its role in task completion; use the particle swarm optimization algorithm to sort and evaluate the priority of the feature vectors of multi-modal data. Through multiple iterations, the particle swarm algorithm outputs the optimal data priority sorting, and finally outputs the data priority list. Preferably, by parsing the task requirements and structuring them, the task requirements are highly compatible with the data priority; the particle swarm optimization algorithm is used for data priority evaluation, and the efficient multi-objective optimization improves the execution efficiency of the overall system; comprehensive evaluation is carried out by combining various factors such as task requirements, network status, and data characteristics, which can intelligently adjust the data priority, enable the task to be completed efficiently, and improve the intelligence and flexibility of the system. Use the heuristic A algorithm for preliminary path planning, dynamically adjust the transmission path, and output the optimized data transmission path. Specifically, using the connection relationships among multi-modal data source nodes, processing nodes, and terminal nodes, a weighted directed graph is constructed. The nodes in the graph represent data generation sources, data processing, or target devices. The weights of the edges in the graph are determined by the real-time network state, including indicators such as bandwidth, latency, and packet loss rate. Each type of modal data is attached with a transmission attribute label, and a directed graph with network performance weights and modal characteristic labels is output, providing a structural input for path planning. Using the heuristic A algorithm, evaluation strategies are set according to different data types, and parameters for path search are output, enabling the algorithm to find the optimal path under performance conditions. Applying the A algorithm to search and select the path of data from the source to the target on the constructed graph structure, starting from the starting node, traversing adjacent nodes and selecting priority nodes according to the cost function, recording all feasible paths under the condition of meeting the constraints, evaluating the path performance for different modal data respectively, outputting a set of feasible data transmission paths, and sorting them in descending order according to the comprehensive cost. For the selected path, if a network state change is detected, the path is updated while recording the change frequency. The path update preferentially executes the principle of minimum range recalculation, and an optimized data transmission path is output. Preferably, the heuristic A algorithm combines multi-constraints and performance parameters, reduces the traversal of invalid paths, shortens the calculation time, and improves the path planning efficiency. In the path evaluation process, modal data characteristic labels are used, enabling different types of data to be transmitted on the path with the highest adaptability, strengthening the data type adaptation ability. Through the path real-time adjustment and local reconstruction mechanism, even in the case of partial network node failures or congestion, an effective path can be quickly reconstructed, enhancing stability, and strengthening the path robustness and recoverability. Based on the selected best modal data subset and data transmission path, multi-objective optimization is performed using the gradient descent method, and the feature representation is dynamically fine-tuned according to the task requirements, and modal data matching the task is output. Specifically, combine the content header with the structure of the metadata parsing modal data subset, extract feature vectors for each modality using the term frequency-inverse document frequency encoder, use zero-padding to uniformly map all modalities into vectors with the same feature dimension, and construct a unified feature matrix; use the batch gradient descent algorithm for iterative optimization, automatically calculate the gradient and update the feature vector matrix, and output the optimal feature matrix; adopt the early fusion method, splice the features of different modalities at the feature level to construct a unified multi-modal feature vector; use the genetic algorithm to optimize and select the feature vectors according to the task requirements, select the most relevant features by setting the optimization goal, and continuously adjust the feature combination according to the optimization algorithm to output the optimal feature subset; use a convolutional neural network to fine-tune the feature subset, and make the feature representation more in line with the specific task requirements through end-to-end training, and output the final feature representation after task fine-tuning; according to the task requirements, map the selected feature data to the target format, match the feature data with the task description, generate data that conforms to the task format, and output the modal data that matches the task. Preferably, through the iterative optimization mechanism of the gradient descent method, the accuracy and stability of the modal feature representation are improved; through the automatic parameter fine-tuning mechanism, the dynamic adaptability and intelligence of the response are enhanced; through the automatic update of the feature representation, the context relevance and semantic consistency of the final code generation strategy are enhanced; a fully automated modal data optimization process is constructed to reduce the manual intervention cost and error risk. Use a preset code generation template to convert the modal data and data transmission path that match the task into a code execution strategy required by the task. Specifically, use another language recognition tool to automatically extract the keyword fields, attributes, and context semantics in the modal data, and output a set of structured semantic representations; select the corresponding code structure template according to the task type to construct a code skeleton with placeholders; use a key-value mapping table to replace the template tags, and embed the extracted fields, functions, parameters, etc. into the corresponding positions of the template to output a complete target code segment; combine the hypertext transfer protocol request template with rule splicing, automatically insert the data acquisition logic related to the path in the generated code, and inject the target IP, port, and protocol into the data sending block in the template according to the path structure to output a complete code block with path communication statements; use code splicing and a syntax checker to check, organize, and output the complete code structure, verify whether it contains the necessary structures, and output the code execution strategy required by the task. Preferably, the consistency of the code structure and the semantic accuracy are improved. By integrating the multi-protocol template segment adaptation mechanism, the protocol adaptation ability and response efficiency of task execution are enhanced; by using the template-driven fast splicing method, the generation speed and maintainability are improved; by using the syntax tree parser-assisted verification mechanism, the code correctness and execution stability are enhanced; by constructing a structured mapping mechanism, the transparency and traceability between the task logic and the code strategy are enhanced.
[0031] The static analysis module uses the abstract syntax tree and the control flow graph to perform static analysis on the code generation strategy, identify performance bottlenecks and potential errors and optimize them, and generate preliminary code.
[0032] Furthermore, the syntax parser is used to convert the code generation strategy into a weighted abstract syntax tree, parse the code structure, generate weights for each node, mark the data modality source, and output the abstract syntax tree with weights and modality labels; Specifically, lexical analysis is used to divide the code generation strategy text into tokens, and then syntax analysis is performed to convert it into an abstract syntax tree according to the defined grammar rules, and the basic abstract syntax tree is output; the rule library is used to generate weights according to indicators such as node type, complexity, nesting depth, usage frequency, and influence range, and output the abstract syntax tree structure with a weight field; the modality mapping table is used to establish the mapping rules between the generation strategy fields and their source modalities, combined with the original code generation strategy input, match its data source through the field name, semantic annotation or context annotation, and attach the modality label as an attribute to the abstract syntax tree node. After modality fusion, multiple modality labels are allowed to be attached, and the priority is set, and the abstract syntax tree node with fields is output; tree structure encapsulation processing is used, the intermediate code structure is used to represent the whole tree, and the weight and modality information are merged into each node, and the abstract syntax tree with weights and modality labels is output; Preferably, by converting the code generation strategy into a structured abstract syntax tree, the overall perception of the potential structure of the generated code is enhanced, and the accuracy of static analysis is improved; by semantic weights, the pertinence and efficiency of the optimization strategy are improved, the convergence speed and effectiveness of the optimization algorithm are enhanced, and the weighted and modality-labeled abstract syntax tree improves the interpretability and debuggability of code generation and enhances the transparency of the system; Based on the abstract syntax tree, a control flow graph is generated. The control flow analysis algorithm is used to calculate the edge weights, identify the execution paths, and construct a control flow graph with edge weight markings; Specifically, an abstract syntax tree traversal algorithm is adopted to identify control structure nodes. On the premise of semantic preservation, sequential statements are aggregated into a single basic block, and a list of basic blocks is output. Static semantic rules are used to determine the execution dependencies and control flow connections between basic blocks, and the topological structure of the control flow graph is output. Combining static analysis metrics with heuristic strategies, probabilistic weights are added to the control flow edges to construct a basic weighted control flow graph. The basic block nodes are labeled with semantic tags, and the edge weights are added to the weight fields to output a control flow graph with edge weight markings. Preferably, by constructing a control dependence graph and analyzing edge weights, the ability to identify critical execution paths is improved; by fusing edge weights and semantic tags, the context awareness of static analysis is enhanced; by constructing an execution path map, the controllability and interpretability of code execution strategy optimization are improved; by standardizing the output of structural information, the reusability and extensibility are enhanced. Based on the control flow graph, the proximal policy optimization algorithm is used to optimize the code execution strategy, and an optimized action sequence is output. Specifically, a graph structure normalization method is used to normalize the weighted control flow graph, making the identification of nodes and edges unified and the weight information clearly accessible, and a standardized control flow graph with a clear structure and accessible edge weights is output. All feasible path segments are extracted from the control flow graph, each path segment is represented as a sequence of state nodes, a state space is constructed, and the state is represented as a numerical sequence using a unique encoding method, and a structured set of state spaces and their corresponding numerical encodings are output. For the optimizable structure types in the control flow graph, optional optimization operations are defined, each operation is assigned a unique action identifier, and rewards are set according to the policy objectives, and a discrete action set and a policy evaluation objective are output. The state space is traversed, all legal actions are attempted for each state, the control flow graph transformation is simulated, and the local structure of the transformed weighted control flow graph is evaluated. If the transformation brings benefits, it is recorded as the optimal policy action, and an optimized action sequence is output. Combining sequence encoding with abstract syntax mapping, the optimized actions are output in sequence form, each action contains fields such as node identifier, operation type, target path segment, and expected impact, and a standard structured file is constructed, and an optimized action sequence is output. Preferably, by modeling the control flow graph edge weights and path recognition, the ability to analyze code execution bottlenecks is improved; by adopting the proximal policy optimization algorithm, the stability and convergence of code optimization are enhanced; by discretizing the definition and reward of the action space, the adaptability and generalization ability to multiple optimization objectives are improved; by rule driving, the interpretability and deployment controllability of the system are enhanced. A code converter is used to reconstruct the optimized action sequence into an abstract syntax tree to generate preliminary code. Specifically, a custom instruction parser is used to read and parse the operation types included in the optimized action sequence, generating an operation instruction stream that can be mapped to an abstract syntax tree structure. The operation instruction stream is parsed using a syntax tree traversal algorithm to locate the target nodes or substructures in the syntax tree that need to be modified, generating a node mapping index table. A tree structure modification library is used to perform structural addition, deletion, and modification operations on the abstract syntax tree according to the node mapping index table, outputting an abstract syntax tree optimized structurally. The abstract syntax tree is verified in combination with the syntax specification, and a source code text is generated using a code generator, making it semantically compatible with the language specification without syntax conflicts or structural omissions, generating preliminary code. Preferably, the determinacy and stability of code generation are improved through a structured instruction mapping mechanism; the semantic fidelity is improved through an accurate syntax tree rewriting mechanism; the code generation efficiency is improved through a node positioning and matching mechanism; the code usability and compilability are improved through a syntax integrity check; and the cross-modal code understanding ability is enhanced through a modal label inheritance mechanism.
[0033] The test evaluation module uses mutation testing to evaluate and improve unit tests and integration tests, and conducts quality and performance evaluations on the code through continuous integration and continuous delivery, outputting a code evaluation report and optimization suggestions.
[0034] Furthermore, by performing syntax replacement, boundary perturbation, and control structure adjustment on the preliminary code, multiple code mutants with different structures are constructed to form test evaluation samples, outputting a set of code mutants with diverse structures. It should be noted that by performing syntax replacement, boundary perturbation, and control structure adjustment on the preliminary code, multiple code mutants with different structures are constructed to form test evaluation samples, outputting a set of code mutants with diverse structures. Furthermore, the control structures and boundary positions in the preliminary code are identified and extracted to determine the candidate points that can be mutated, outputting a structured list of candidate points. Specifically, a syntax parser is used to parse the preliminarily generated source code into an abstract syntax tree, obtaining the structural hierarchy relationship of the source code and the complete abstract syntax tree structure. A syntax tree traversal algorithm is used to extract all nodes of the control structure type, match the nodes that conform to the control structure labels in the abstract syntax tree, and output a set of control structure candidate points. Comparison operations are searched for in the control structure and expression nodes to identify the relationship between constants and variables, mark the boundary values, and output a set of boundary position candidate points. The control structure candidate point set and the boundary position candidate point set are combined to annotate meta-information such as point type, source code position, relevant variables, and operators, outputting a structured list of candidate points. Preferably, the understanding ability of the code semantic level is strengthened by constructing an abstract syntax tree, and the accuracy of code structure parsing is improved; the control structure type is extracted through static traversal and syntax node screening technology, and the recognition efficiency of candidate mutation points is improved, making the target of control structure adjustment operations clear and precise; the usability and generality of mutation point information are improved through a unified structured candidate point format output method; by uniformly abstracting control logic and boundary conditions into candidate mutation units, the construction basis of test sample diversity is improved, and the coverage rate, fault triggering ability and robustness evaluation level in automated testing are strengthened; Use mutation rules such as syntax replacement, boundary perturbation, and control structure adjustment to generate multiple code mutants with different structures and behaviors, and construct a diverse set of preliminary code mutants; Specifically, using the abstract syntax tree and regular expressions, traverse the syntax nodes in the preliminary code, identify basic syntax units through the abstract syntax tree, select syntax nodes that meet specific mutation rules, and perform replacement according to preset rules. Use regular expressions or template matching to perform syntax replacement on parts such as strings, function calls, and conditional expressions in the code to generate a set of mutants; identify the boundary conditions in the code, especially the boundary values in loop structures and conditional judgments, adjust the boundary values, perturb different types of boundaries, and output multiple perturbed code mutants; use control flow graph analysis and conditional branch rewriting to adjust the control structure of the program, analyze the control flow graph of the program, identify the control structure, mutate the control structure, change the branch execution order in the conditional judgment of the program, and adjust the iteration method in the loop structure to output multiple mutants; use metadata annotation and multi-dimensional analysis to organize and aggregate the generated mutants, uniformly organize all the generated code mutants, classify and organize the mutants according to function and execution path, and use the mutation point type to mark each mutant to construct a diverse set of preliminary code mutants; Preferably, through the structure mutation technology, a code mutant set is constructed from three levels of syntax, boundary, and control from multiple angles, which ensures the differentiated expression of functional behaviors, improves test sample diversity, expands the coverage of code logic and execution paths, and enhances the discoverability of software defects and robustness verification ability; Add metadata annotation to the code mutants to indicate the mutation type, mutation point location, and differences before and after mutation, organize the mutants according to function and execution path, and output a code mutant set with diverse structures; Specifically, use an abstract syntax tree comparison tool to perform difference analysis on the original code and mutant code. Traverse the nodes of the abstract syntax tree to compare differences, extract the mutation type, the location of the mutation point in the source code, and the changes before and after the mutation content for each mutation, and construct a standardized mutation metadata structure; use a code annotation injector to embed metadata information into the mutant code, automatically insert annotations before and after the mutation point to indicate the mutation type and number, and form a commented code mutant; use a static analysis tool to extract the function module attribution and potential impact paths of each mutant, classify the mutants according to the control flow graph, and perform hierarchical classification by module, path, and impact area to generate a structured classification index; bind and store all the commented and metadata-containing mutant codes with their index structures, and output a code mutant set with diverse structures. Preferably, by using abstract syntax tree difference analysis and a standardized metadata structure, the traceability and interpretability of code mutants are improved; the mutants are classified and organized by functional paths using the control flow, enhancing the structural organization and scheduling efficiency of test samples; combined with structured metadata management, the detection ability of the test set for potential defects and boundary condition errors is improved, and the test coverage and error detection ability are enhanced. Apply the existing unit test set to the code mutant set, record the killing effect of test cases and generate a coverage matrix, and output a reconstructed high-coverage unit test set through coverage clustering and hole analysis. Specifically, load the code mutants with diverse structures into the test environment one by one, apply the test cases in the existing unit test set to each mutant respectively, judge the killing effect of the mutant by monitoring abnormal behaviors, and generate a killing matrix; record the code coverage information of each test case when executing the original code, summarize it into a coverage matrix, where the matrix elements indicate whether the test case covers the code element, and output a structured coverage matrix; use a density clustering algorithm to perform similarity analysis on the coverage matrix, analyze the overlap degree between test cases, cluster highly redundant test cases into one category, retain the central test case, eliminate redundancy, and output the clustering division result; combine matrix sparsity analysis and Boolean vector inversion method for hole detection, extract the hole areas in the coverage matrix, locate these hole areas to control structures or logical paths, and synthesize new targeted test cases based on syntax tree analysis and symbolic execution to form a hole test case patch set; uniformly select test cases with high killing rate, high path coverage rate, and low redundancy, add hole patch test cases, verify the coverage and effectiveness of the reconstructed set on the mutant set and the original code, and output a reconstructed high-coverage unit test set. Preferably, the effectiveness and discrimination of test cases are improved through mutation testing, redundancy is compressed and test execution efficiency is improved through coverage clustering, coverage integrity and path diversity are strengthened through hole analysis, and the interpretability and visibility of test evaluation are enhanced through structured matrix analysis. Based on the refactored unit test set, extract the module control dependencies and interface call paths, generate a new control flow graph, automatically generate an integration test set through a path-driven mechanism, and output a path-complete integration test set; Specifically, perform static analysis on the tested code corresponding to the refactored unit test set, identify the control dependency edges in each module, extract all external interface calls and cross-module dependency relationships, establish a call path graph, and construct a control flow dependency graph and an interface call path graph; unify and integrate the control flow dependency graph and the interface call path graph into an extended control flow graph, add a path label mechanism to identify paths, boundary paths, and risk paths. For asynchronous calls and callback chains, adopt a delayed graph construction and back-edge reconstruction strategy to generate a path-complete control flow graph for multi-module integration; traverse all critical paths in the path-complete control flow graph that are not covered by the existing tests, use symbolic execution and constraint solving to generate test input combinations that meet the constraint conditions for each path, automatically synthesize test scripts for multi-module interaction, and inject verification assertions to output a path-complete integration test set; Through continuous integration and continuous delivery, perform automated testing, collect mutation coverage rates, performance metrics, and resource utilization evaluation parameters, and generate a code evaluation report and optimization suggestions; Specifically, use a continuous integration configuration file to automatically trigger the build and test processes when the code is committed. Define the build tasks through a class markup language configuration file, including build logs, unit test logs, and mutation test start signals; use open-source tools for mutation testing, insert a mutation testing phase in the continuous integration process, automatically perform semantic perturbation on the code and then run the test set, count the number and types of mutant test cases, generate mutation coverage rates, and output mutation test reports, coverage reports, and path kill matrices; in the continuous integration and continuous delivery pipelines, embed a performance profiler or monitoring agent when executing the test code to sample function call times, CPU usage peaks, memory allocation frequencies, etc., and output performance trend graphs, resource utilization curve graphs, and hot function call stack records; combine custom scripts to aggregate the results, aggregate multiple data sources such as test coverage rates, mutation kill rates, and performance reports, generate suggestions based on static rules, trigger suggestions according to the rules, and generate a code evaluation report and optimization suggestions; Preferably, enhance the real-time performance and consistency of code test coverage through an automated trigger mechanism, mutation test metrics, improve the lethality and defect sensitivity of the test set, generate a structured code evaluation report and optimization suggestions, and enhance the operability and visualization of test results.
[0035] In summary, the present invention: semantically aligns natural language task descriptions with multi-modal data such as code, logs, and requirement documents to improve the accuracy of task intention recognition and reduce code generation biases caused by single-modal analysis; dynamically filters the optimal subset of modal data and optimizes the data transmission path to reduce computational resource consumption while ensuring code performance, achieving a global optimum for the code generation strategy.
[0036] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A multi-modal based code automatic generation and optimization system, characterized in that: including, a data processing module that collects raw data, performs data cleaning and preprocessing, and constructs a raw data set; a task analysis module that combines the CLIPS rule engine and the knowledge graph, analyzes the task description and constraint conditions in combination with the raw data, identifies task objectives and requirements, and outputs a task risk assessment report and a task intention set; a data fusion module that selects data sources based on the task risk assessment report and the task intention set, performs data fusion using the weighted average method, and outputs a fused multi-modal data set; an optimization decision module that uses the particle swarm optimization algorithm and the path planning algorithm to select the best modal data subset, and performs optimization and adjustment according to task requirements, and outputs a code generation strategy; a static analysis module that uses the abstract syntax tree and the control flow graph to perform static analysis on the code generation strategy, identifies performance bottlenecks and potential errors and performs optimization, and generates preliminary code; a test evaluation module that uses mutation testing to evaluate and improve unit testing and integration testing, and evaluates the quality and performance of the code through continuous integration and continuous delivery, and outputs a code evaluation report and optimization suggestions.
2. The multimodal-based code automatic generation and optimization system according to claim 1, characterized in that: The data processing module collects raw data, performs data cleaning and preprocessing, and constructs a raw data set. The specific steps are as follows: Collect raw data using the Representational State Transfer (REST) application programming interface, establish a data directory structure for archiving according to the task ID, and output a structured raw data sample set; Automatically complete missing values using placeholders, and then perform data standardization and data anomaly detection using the Z-score standardization and Isolation Forest algorithms, and output a high-quality data set after cleaning; Extract text using key phrase extraction, identify function names to extract function structures from code, and perform structured extraction on log parsing event sequences, and reorganize them according to a unified field specification to construct data entries with a unified structure; The key phrase extraction is a method in the field of natural language processing, and its purpose is to perform word segmentation, part-of-speech tagging, and stop word filtering on the task description text; The key phrase extraction refers to performing word segmentation, part-of-speech tagging, and stop word filtering on the text; Generate a unique hash identifier for the data entries, and combine rule label extraction to generate a raw data set.
3. The multimodal-based code automatic generation and optimization system according to claim 2, characterized in that: The task analysis module combines the CLIPS rule engine and the knowledge graph, analyzes the task description and constraint conditions in combination with the raw data, identifies task objectives and requirements, and outputs a task risk assessment report and a task intention set. The specific steps are as follows: Use the domain knowledge graph and the natural language processing toolkit to map the text semantics to the entity nodes and attribute edges in the graph, and construct a task semantic graph; Use the CLIPS rule template to define behavior objectives, resource requirements, and context constraints, substitute the graph entities into the preset rules for reasoning, and output a task intention set; Based on the multi-level reasoning of CLIPS and the task dependency graph construction mechanism, record the trigger paths of each reasoning conclusion, determine the conflict source through the backtracking method, generate a dependency chain, and output a structured risk node chain graph; Combine the text generation template and the structured export tool, and comprehensively transform the task intention set and the risk node chain graph into a readable report, and output a task risk assessment report.
4. The multimodal-based code automatic generation and optimization system according to claim 3, wherein: The data fusion module selects data sources based on the task risk assessment report and the task intention set, uses the weighted average method for data fusion, and outputs the fused multi-modal data set. The specific steps are as follows: Use natural language processing methods to identify and extract the task elements of the task risk assessment report and the task intention set, use dependency syntax analysis to process the relationships between task nodes, and combine with the knowledge graph to output the task node combination and the task node dependency graph; Apply relation extraction to analyze the association between task nodes and data sources, and perform priority matching through a weighted matching algorithm to output the precise mapping of data requirements and data sources; Based on the task node combination, the task node dependency graph, and the precise mapping of data requirements and data sources, use weighted average fusion. Through time series alignment and data preprocessing, process and fuse different modal data to generate a multi-modal data set in a unified format.
5. The multimodal-based code automatic generation and optimization system according to claim 4, wherein: The optimization decision module uses the particle swarm optimization algorithm and the path planning algorithm to select the best modal data subset, and performs optimization and adjustment according to the task requirements, and outputs the code generation strategy. The specific steps are as follows: Adopt the particle swarm optimization algorithm to evaluate the fused multi-modal data set and calculate the best modal data subset that meets the code generation task; Plan the transmission path of the best modal data subset through the A algorithm, and optimize the path of the selected modal data during transmission to output the optimized data transmission path; Based on the selected best modal data subset and the data transmission path, use the gradient descent method for multi-objective optimization, dynamically fine-tune the feature representation according to the task requirements, and output the modal data that matches the task; Use the preset code generation template to convert the modal data and data transmission path that match the task into the code execution strategy required by the task.
6. The multimodal-based code automatic generation and optimization system according to claim 5, characterized in that: The static analysis module uses the abstract syntax tree and the control flow graph to perform static analysis on the code generation strategy, identify performance bottlenecks and potential errors and optimize them to generate preliminary code. The specific steps are as follows: Use a syntax parser to convert the code generation strategy into a weighted abstract syntax tree, parse the code structure, generate weights for each node, and mark the data modal source to output an abstract syntax tree with weights and modal labels; Generate a control flow graph based on the abstract syntax tree, use a control flow analysis algorithm to calculate edge weights, identify execution paths, and construct a control flow graph with edge weight markings; Based on the control flow graph, use the proximal policy optimization algorithm to optimize the code execution strategy and output the optimized action sequence; Use a code converter to reconstruct the optimized action sequence into an abstract syntax tree to generate preliminary code.
7. The multi-modal-based code automatic generation and optimization system according to claim 6, characterized in that: The test evaluation module uses mutation testing to evaluate and improve unit testing and integration testing, and evaluates the quality and performance of the code through continuous integration and continuous delivery, and outputs a code evaluation report and optimization suggestions. The specific steps are as follows: By performing syntax replacement, boundary perturbation, and control structure adjustment on the preliminary code, construct multiple code mutants with different structures to form a test evaluation sample, and output a set of code mutants with diverse structures; Apply the existing unit test set to the code mutant set, record the killing effect of test cases and generate a coverage matrix, and through coverage clustering and hole analysis, output the reconstructed high-coverage unit test set; Based on the reconstructed unit test set, extract module control dependencies and interface call paths, generate a new control flow graph, and automatically generate an integration test set through a path-driven mechanism, and output a path-complete integration test set; Through continuous integration and continuous delivery, execute automated tests, collect mutation coverage rate, performance metrics, and resource utilization evaluation parameters, and generate a code evaluation report and optimization suggestions.
8. The multimodal-based code automatic generation and optimization system according to claim 7, characterized in that: The best modal data transmission path is planned through the A algorithm, and the path of the selected modal data during transmission is optimized, and the optimized data transmission path is output. The specific steps are as follows. Calculate the data priority based on the task requirements, use the particle swarm optimization algorithm to evaluate the priority of different modal data, and output a priority list; Use the heuristic A algorithm for preliminary path planning, dynamically adjust the transmission path, and output the optimized data transmission path.
9. The multimodal-based code automatic generation and optimization system according to claim 8, wherein: The use of the domain knowledge graph and the natural language processing toolkit maps the text semantics to the entity nodes and attribute edges in the graph, and constructs a task semantic graph. The specific steps are as follows. Use the natural language processing toolkit to identify the core semantic units in the task text, perform word segmentation, part-of-speech tagging, dependency syntax analysis, and named entity recognition, extract task action, object, and condition elements, and output a set of structured semantic units; Combine the domain knowledge graph and the predefined term mapping rule table to bind the set of semantic units to the entity nodes in the graph, and output a list of semantic units bound to the graph entity nodes; Use the graph edge relationship template library and the graph structure generation algorithm to construct semantic relationship edges and generate a task semantic graph, construct attribute edges according to the semantic dependencies between entities, and complete the logical relationships in combination with the graph reasoning mechanism to construct a task semantic graph.
10. The multi-modal-based code automatic generation and optimization system according to claim 9, The feature lies in that: by performing syntax replacement, boundary perturbation, and control structure adjustment on the preliminary code, multiple code mutants with different structures are constructed to form a test evaluation sample, and a code mutant set with diverse structures is output. The specific steps are as follows. Identify and extract the control structures and boundary positions in the preliminary code, determine the candidate mutation points, and output a structured list of candidate points; Use mutation rules such as syntax replacement, boundary perturbation, and control structure adjustment to generate multiple code mutants with different structures and behaviors, and construct a diverse preliminary code mutant set; Add metadata annotations to the code mutants, indicate the mutation type, mutation point location, and differences before and after mutation, and sort the mutants according to functions and execution paths, and output a code mutant set with diverse structures.
Citation Information
Patent Citations
Dynamic code generation method and system based on interface document
CN118860356A
Software architecture code generation method and system based on large language model
CN119597267A
Program code intelligent conversion method and system based on knowledge graph
CN120104113A
Systems and methods for automated software code generation
WO2025099638A1
Cited By
Large model test case intelligent generation system based on big data
CN120448285A
A big data-based large model test case intelligent generation system
CN120448285B
Method and system for automatically generating front-end and back-end codes by adopting credential and credential system
CN120523466A
Work order abstract generation method based on low code configuration
CN120523939A
A method for generating work order summaries based on low-code configuration
CN120523939B