Data quality detection and improvement system for large language model

The system addresses limitations in code optimization by employing a large language model for comprehensive quality assessment and adaptive optimization, enhancing code repair effectiveness and development efficiency through real-time feedback and performance analysis.

CN120315718AInactive Publication Date: 2025-07-15北京开放传神科技有限公司
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510253772.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology cannot fully evaluate the code quality, the repair effect is limited, the repair path search efficiency is inefficient, and the lack of adaptability and continuous optimization capabilities, resulting in limited software development efficiency and quality.

Method used

The data quality detection and improvement system adopts a large language model, including data analysis and feature extraction module, multi-dimensional quality detection module, code optimization and verification module, feedback and adaptive optimization module, and adaptive optimization module. Adaptive optimization and repair are achieved through static analysis, semantic embedding, multi-dimensional evaluation, dynamic testing and feedback loops.

Benefits of technology

It realizes continuous adaptive optimization and repair, has the ability to intelligently adjust the repair path, and can continuously improve the repair effect based on error type, code structure and developer feedback, ensuring the continuous optimization and stability of the repair effect.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention provides a data quality detection and improvement system for a large language model, and relates to the technical field of code quality management. The data analysis and characteristic extraction module is used for extracting structured and semantic characteristics of codes and constructing a code distribution diagram; the multi-dimensional quality detection module comprehensively evaluates the data set quality from three aspects of semantic diversity, structural consistency and error mode; the code optimization and verification module generates an optimization suggestion based on a detection result, and verifies the effectiveness of an optimization scheme through dynamic evaluation and unit testing; the feedback and self-adaptive optimization module is combined with developer feedback and runtime performance test to dynamically adjust a repair strategy, and the complex code repair efficiency is improved through quantum computing auxiliary path search; according to the method, full-process automation from quality detection to optimization verification is achieved, the method is suitable for multi-dependency and multi-module complex code repair tasks, and the quality and efficiency of the code life cycle are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code quality management, and specifically to a data quality detection and improvement system for large language models. Background Art

[0002] In the process of software development, the evaluation and optimization of code quality are important and complex tasks. Existing code optimization and repair technologies have limitations, unable to comprehensively evaluate code quality, with limited repair effects, lacking adaptability and continuous optimization capabilities.

[0003] Traditional code optimization and repair methods mainly rely on static analysis, dynamic monitoring, and manual intervention, suffering from problems such as low efficiency, insufficient accuracy, and low automation; these methods are difficult to deeply understand the semantics and logical structure of code, unable to effectively detect and repair complex code defects; at the same time, existing technologies lack adaptability and cannot dynamically adjust optimization strategies according to actual situations, resulting in poor repair effects.

[0004] In addition, existing technologies cannot achieve all-round code quality evaluation and continuous adaptive optimization, with inaccurate repair effect evaluation, low efficiency in searching for repair paths, and lacking self-evolution and closed-loop optimization mechanisms; these problems have hindered the development of code optimization and repair technologies and affected the efficiency and quality of software development. Summary of the Invention

[0005] (I) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the existing technology, the present invention provides a data quality detection and improvement system for large language models to solve the problems in the existing technology such as unable to comprehensively evaluate code quality, limited repair effects, inaccurate repair effect evaluation, low efficiency in searching for repair paths, and closed-loop optimization mechanisms.

[0007] (II) Technical Solutions

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: A data quality detection and improvement system for large language models, including a data parsing and feature extraction module, a multi-dimensional quality detection module, a code optimization and verification module, and a feedback and adaptive optimization module; The data parsing and feature extraction module is responsible for parsing the code dataset, extracting its structural features and semantic features, extracting the syntax features of the code through a static analysis tool, generating semantic embeddings using a semantic embedding model, and constructing a code distribution map to represent the semantic and call relationships between code entities; The multi-dimensional quality detection module comprehensively evaluates the quality of the dataset from three dimensions: semantic diversity, structural consistency, and error patterns, including quantifying the diversity of function implementation methods by analyzing the semantic vector distribution, identifying unreasonable dependencies between modules through the module dependency graph, and combining pre-trained models to predict potential errors; Based on the detection results of the multi-dimensional quality detection module, the code optimization and verification module generates optimization suggestions, refactors the code, and verifies the effectiveness of the optimized code through unit tests. It also evaluates the effect of the repair solution in real time through a dynamic evaluation model to optimize the repair path and strategy; The feedback and adaptive optimization module conducts runtime performance tests on the optimized code by receiving developer feedback and combining performance analysis tools, automatically adjusts the repair strategy according to the feedback and test results, and combines a quantum search engine to accelerate the search process for complex repair paths.

[0009] Preferably, the operation process of the data parsing and feature extraction module is divided into multiple steps; first, the module parses the input code dataset through a static analysis tool; the static analysis tool constructs an abstract syntax tree (AST) of the code and extracts key information related to the code structure, including function declarations, variable definitions, dependencies, module grouping, and error patterns; for example, for the function call chain in the code, the tool parses the call stack to record the call order between functions and the data flow dependencies; after obtaining the static structure of the code, the data parsing and feature extraction module passes the data to a semantic embedding model, which uses a deep learning network based on the LLM architecture to semantically model the code fragments using the context encoding mechanism; specifically, in the implementation, by extracting the context window information of the code fragments, the logical relationships between variables and functions in the context are encoded as semantic vectors; the generated semantic vectors are mapped to a high-dimensional space for similarity analysis and classification; next, a code distribution map is constructed based on the extracted syntactic and semantic features; the nodes of the distribution map represent code entities such as functions, classes, and modules, and the edges represent the semantic or call relationships between entities; to more precisely analyze the interaction relationships between modules, graph algorithms such as depth-first search (DFS) are used to identify complex call paths; during this process, potential circular reference problems are marked; finally, the data parsing and feature extraction module stores the extracted features as a structured data file through an interface to provide basic data support for the multi-dimensional quality detection module.

[0010] Preferably, the multi-dimensional quality detection module comprehensively evaluates the code quality from three aspects: semantic diversity, structural consistency, and error patterns. In the semantic diversity analysis, the multi-dimensional quality detection module first performs clustering operations on the semantic vectors generated by all code snippets. The specific method is to divide the semantic vectors into several groups using the K-means or DBSCAN algorithm. Each group represents the logical similarity of a certain function implementation. By statistically calculating the distribution density of each group of vectors, the semantic diversity score of the code is calculated and compared with the ideal value. In the structural consistency evaluation, the multi-dimensional quality detection module uses the code distribution map to identify the dependencies between modules, uses topological sorting to judge the feasibility of the dependency paths, and at the same time detects possible loops in the graph, that is, circular dependencies, through algorithms. When unreasonable dependencies between modules are detected, they will be automatically marked and optimization suggestions will be generated. In the error pattern detection, the multi-dimensional quality detection module combines a pre-trained model to predict potential errors in the code. For example, the multi-dimensional quality detection module will scan common error patterns, including array out-of-bounds, null pointer exceptions, etc. The predicted errors are output as a probability distribution by the deep learning model, and the system decides whether to mark them as high-risk according to the set confidence threshold. Finally, a quality report is generated by integrating the detection results of the three dimensions, including the overall code score, the list of major problems, and specific optimization suggestions.

[0011] Preferably, based on the output of the quality detection module, the code optimization and verification module generates optimization suggestions and performs code refactoring. In specific operations, corresponding optimization strategies will be called according to each type of quality problem. For example, for the problem of insufficient semantic diversity, the system will analyze the code in the high-density area of the semantic vector distribution and suggest that developers introduce diverse implementation solutions through code refactoring. For structural consistency problems, the system will generate an optimized path for module dependencies and reconstruct the call relationships between modules. During the actual refactoring process, the code optimization and verification module automatically processes the code through an integrated refactoring tool, such as merging duplicate code, eliminating dead code, and optimizing function definitions. After the code optimization is completed, the optimization results are verified through a unit test framework. The test framework includes two parts: static testing and dynamic testing. The static testing verifies the syntactic correctness and logical integrity of the code, and the dynamic testing captures the runtime behavior of the code by running actual test cases. The test results are compared with the code before optimization, and the system records key metrics such as performance changes and error reduction rates and feedbacks them to the optimization engine for correction.

[0012] Preferably, the feedback and adaptive optimization module dynamically adjusts the optimization strategy by receiving developer feedback in real time and combining performance test data; developers submit feedback through an interactive interface, including satisfaction scores for the optimized code and improvement suggestions; at the same time, the built-in performance analysis tool conducts runtime performance tests on the optimized code to collect feedback data on execution time, memory consumption, and error rate; these feedback data are fed into the adaptive optimization engine as the basis for dynamically adjusting the repair strategy; when adjusting the strategy, the system uses reinforcement learning algorithms to explore the strategy space and select an optimization path that can improve performance; in addition, the feedback and adaptive optimization module combines a quantum computing-assisted adaptive repair path search engine (QC-RPS) to quickly find the optimal solution in complex dependencies through quantum amplitude estimation and quantum search algorithms; finally, the system applies the adjusted optimization strategy to the next optimization task to achieve iterative evolution of the repair solution.

[0013] Preferably, the static analysis tool and the semantic embedding model work together to achieve a comprehensive analysis of code data; the static analysis tool parses the syntax structure of the code, including extracting the abstract syntax tree (AST) and the dependency graph, and identifying explicit dependencies in the code; at the same time, the semantic embedding model generates high-dimensional semantic vectors by analyzing the code context; in the specific operation process, the static analysis tool is responsible for generating the initial features of the syntax nodes, including variable types, function definitions, and call paths; the semantic embedding model, based on the static analysis results, further uses a context encoder to capture the logical relationships of code fragments; for example, for a complex nested loop code, the static analysis tool can provide the syntax features of each layer of the loop, while the semantic embedding model can supplement the logical dependencies between variables in the loop; this combination enables the parsing module to conduct a comprehensive analysis of the code from both static and semantic features, providing more accurate basic data for subsequent modules.

[0014] Preferably, the Adaptive Program Repair and Evolution Path Optimization Algorithm (APE-RP) combines genetic algorithms with deep reinforcement learning to simulate the complete process of code repair and evolution; the genetic algorithm first generates a set of initial repair solutions, which are constructed by analyzing the types of code errors and the solutions to similar problems in historical repair data; each repair solution is represented by a sequence of repair operations, such as changing variable types, adjusting function call parameters, or restructuring module dependency paths; next, the algorithm uses a fitness function to evaluate the effectiveness of each solution, and the fitness function is calculated based on multiple metrics such as code execution performance, error repair rate, and complexity reduction rate; during the evaluation phase, the behavior of the repaired code is captured by running unit tests and compared with the original code to generate a detailed difference report; subsequently, in the deep reinforcement learning phase, the current repair path is optimized by using the repair history data as a training set; the reinforcement learning model adopts an Actor-Critic architecture, where the Actor is responsible for generating possible repair actions and the Critic evaluates the impact of each action on the final goal; APE-RP also continuously optimizes the repair strategy through a feedback mechanism. For example, if new problems are found in the code after repair, the algorithm will adjust the fitness function to avoid similar errors from occurring again; finally, after multiple rounds of genetic selection and reinforcement learning optimization, APE-RP can output an efficient repair path and comprehensively verify the repair effect.

[0015] Preferably, the Intelligent Feedback Adaptive Repair Evaluation Mechanism (ISA-FRM) is responsible for evaluating the effectiveness of the repair solution in real time and dynamically adjusting the repair strategy according to developer feedback and code behavior data; when the code is optimized, ISA-FRM first captures the behavioral differences before and after repair, including functional correctness, running performance, and code complexity changes; through pre-set test cases and dynamic performance monitoring tools, metrics such as execution time, resource occupancy, and error rate are recorded; during the test process, the system compares the recorded performance data with the benchmark data before optimization to generate quantitative improvement metrics; next, developers can submit feedback on the repair results through an interactive interface, including satisfaction scores, reasons for dissatisfaction, and further optimization suggestions; ISA-FRM trains an adaptive optimization model by combining developer feedback and performance data, enabling it to adjust the repair strategy based on historical feedback experience; in addition, this mechanism has a built-in anomaly detection module that automatically triggers a rollback process and generates new repair suggestions when it detects that the behavior of the repaired code does not match the expectation; this adaptive mechanism ensures that each optimization can be fine-tuned for the current problem while avoiding damage to existing functions.

[0016] Preferably, the quantum computing-assisted adaptive repair path search engine (QC-RPS) combines quantum amplitude estimation and quantum search algorithms to achieve efficient search in a large-scale repair path space. First, the search space of the repair task is modeled as a multi-dimensional graph structure, where nodes represent possible repair operations and edges represent the dependencies between operations. In classical search algorithms, due to the complexity and large scale of the repair path space, the search efficiency is limited. QC-RPS evaluates the feasibility of multiple repair paths simultaneously in a single operation by utilizing the principle of quantum superposition. In specific operations, the system first uses the quantum amplitude estimation algorithm to assign priorities to each path, and the priorities are based on the potential effects of the repair path, operation costs, and historical data matching. Then, it quickly locates the preferred path through quantum amplitude estimation and the Grover search algorithm. When encountering complex situations such as multi-module dependencies or circular references, QC-RPS can also combine classical algorithms to dynamically adjust the search strategy to ensure the reliability and efficiency of the repair plan. Finally, the repair plan output by the system is verified by a traditional optimization engine to ensure its stable application in the actual development environment.

[0017] Preferably, the multi-dimensional adaptive optimization framework is the overall coordination mechanism of the system, and its core task is to comprehensively analyze and optimize the code by combining multiple dimensions of static, dynamic, semantic, and execution context. In the static analysis stage, the framework extracts the basic characteristics of the code, including variable declarations, module grouping, and function call paths. In the dynamic analysis stage, the framework captures the context information during code execution through runtime tools, such as memory allocation, thread status, and error logs. In the semantic analysis stage, the framework uses a deep learning model to generate semantic embeddings to capture the logical semantic relationships of code fragments. Finally, by combining the execution context information, the framework analyzes the performance of the code in the actual running environment, including the way of processing input data and the dependencies on external resources. Based on this multi-dimensional information, the framework invokes the optimization strategies of different modules and continuously adjusts the optimization plan through a feedback loop. For example, when static analysis discovers unreasonable dependencies between modules, the framework will call the code optimization and verification module for repair. If dynamic analysis captures performance bottlenecks in the optimized code during runtime, the system will readjust the optimization path through the feedback and adaptive optimization module. This full-process closed-loop optimization ensures the efficiency and stability of the code at all stages.

[0018] (III) Beneficial Effects

[0019] The present invention provides a data quality detection and improvement system for large language models, which has the following beneficial effects: 1. The present invention has the ability of continuous adaptive optimization and repair. By combining LLM deep learning and genetic algorithms, it can simulate the repair evolution process, intelligently adjust the repair path, and continuously improve the repair effect according to different error types, code structures, development environments, and developer feedback.

[0020] 2. The present invention adopts a behavior-driven repair effect evaluation mechanism. Using behavior analysis, LLM deep learning, and reinforcement learning technologies, it can quantitatively measure the actual impact of the repair solution on code execution in real time, and adjust and optimize the strategy according to the feedback to ensure the continuous optimization of the repair effect and avoid regression problems. Detailed implementation manners

[0021] The technical solutions in the embodiments of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] Embodiment 1: An embodiment of the present invention provides a data quality detection and improvement system for large language models. In this embodiment, the specific implementation process of a data quality detection and improvement system based on large language models is mainly described. First, the code dataset input into the large language model LLM for deep learning is parsed by the data parsing and feature extraction module. At the beginning of this process, the user uploads the code dataset to the system interface, and the system calls a static analysis tool to perform a preliminary parse of the code. The tool extracts the basic structural features of the code by constructing an abstract syntax tree, including function definitions, module grouping, and dependency paths. Subsequently, the system generates a module dependency graph based on the abstract syntax tree and identifies possible circular dependencies or unreasonable call paths between modules through a depth-first search algorithm. On this basis, the static analysis tool marks potential dependency problem areas to complete the basic structure analysis of the code. Next, the parsing module calls a semantic embedding model to analyze the semantic features of code fragments. The model encodes the context window of the code fragment semantically and maps the code fragment into a high-dimensional semantic vector. These vectors are further clustered in a multi-dimensional space to capture the semantic relationships between code fragments. The generated semantic vectors are used to construct a code semantic distribution map, where the nodes represent code entities and the edges represent semantic similarity or call relationships. The parsing module stores the extracted syntactic features and semantic features in a structured database uniformly to provide basic data support for the subsequent work of the quality detection module. After the data parsing is completed, the multi-dimensional quality detection module receives the output data of the parsing module and conducts a comprehensive quality assessment from three dimensions: semantic diversity, structural consistency, and error patterns. In the semantic diversity analysis, the system calls a clustering algorithm to divide the semantic vectors into several groups with similar functions and calculates the distribution density and coverage range of each group to analyze the diversity level of different function implementation methods. For semantic groups with too high distribution density or too narrow coverage range, the system marks them as objects that need to be optimized. In the structural consistency evaluation, the system checks the call paths between modules based on the module dependency graph, identifies circular reference problems through topological sorting, and analyzes unreasonable cross-module calls through call chain depth. Subsequently, the system combines a pre-trained model to scan common error patterns, including null pointer exceptions, resource leaks, etc. The model generates a prediction probability for potential errors, and the system records the high-risk areas in the detection report. After the quality detection is completed, the code optimization and verification module starts to refactor the code according to the optimization suggestions in the detection report. The system calls a preset optimization tool to complete specific refactoring operations, such as merging code fragments with repeated functions, adjusting the module call order to eliminate circular dependencies, optimizing variable definitions and function parameter designs, etc. The optimized code will be verified for its functional correctness and performance improvement level through unit tests. The unit tests include two stages: static verification and dynamic execution. Static verification ensures the syntactic correctness and logical consistency of the code, and dynamic execution captures the code behavior through actual operation and generates detailed performance data.Finally, the optimized code will be fed back to the adaptive optimization module. The module combines the developer's feedback with the running data captured by the performance analysis tool to adjust the optimization strategy. The adjusted strategy will be used in the next code optimization task. The entire implementation process achieves continuous improvement and iterative optimization through a feedback loop mechanism, thereby comprehensively enhancing the quality and performance of the code.

[0023] Example Two: Based on Embodiment 1, this embodiment extends the optimization process of complex code by combining a quantum computing-assisted adaptive repair path search engine. Different from Embodiment 1, this embodiment specifically targets the situation where there are complex multi-module dependencies in the code. By introducing quantum computing technology, it accelerates the optimization process and designs a more refined repair strategy for the system. The implementation process first has the data parsing and feature extraction module complete the preliminary parsing of the code. The static analysis tool constructs an abstract syntax tree for the input code dataset and extracts the module dependency paths. Different from Embodiment 1, during the process of parsing and marking the dependency relationships of the modules, this embodiment calls the quantum-assisted dependency analysis module, which uses the principle of quantum superposition to simultaneously evaluate the rationality of multiple dependency paths, thereby quickly marking high-risk dependencies in the huge path space. At the same time, the semantic embedding model performs semantic analysis on the code fragments, and the generated semantic vectors are mapped into a multi-dimensional semantic space and cooperate with the quantum computing module. Through quantum amplitude estimation, priorities are assigned to each group of semantic vectors, which improves the efficiency of semantic feature analysis while dealing with complex dependency relationships. Next, the multi-dimensional quality detection module receives the output data of the parsing module and predicts the optimization path in combination with quantum computing. In this embodiment, after the module dependency graph initially identifies potential cyclic dependencies through topological sorting, the system calls the quantum search engine for a quick search of the optimization path. The specific process is as follows: The system first constructs a graph structure model of the dependency path, where the nodes represent modules and the edges represent the call relationships between modules. The quantum search engine uses the Grover algorithm to parallelly search the path space and outputs the optimization path. Subsequently, the system combines classical algorithms to verify the results of the quantum search to ensure the actual feasibility of the repair path. After generating the optimization path, the code optimization and verification module adjusts the call relationships between modules according to the path and further optimizes the code repair strategy using the APE-RP algorithm. APE-RP generates multiple groups of initial repair solutions through genetic algorithms and iteratively optimizes the repair path in combination with deep reinforcement learning. In particular, through the quantum computing-assisted dynamic path selection function, the repair efficiency is significantly improved. After the optimization is completed, the system verifies the function and performance of the code through unit tests. In the dynamic test, this embodiment specifically adds input data for complex scenarios to verify the running stability of the optimized code in a high-load environment. Finally, the adaptive optimization module adjusts the optimization path by capturing developer feedback and performance test data. In this embodiment, the adaptive module combines the quantum superposition mechanism to perform efficient pattern matching on the repair historical data, thereby generating a more intelligent optimization strategy. The adjusted strategy is stored in the optimization engine for the next optimization task. Through the synergistic effect of quantum computing and classical optimization strategies, this embodiment successfully solves the problem of low efficiency in Embodiment 1 when dealing with complex dependency relationships, thereby significantly improving the performance and quality improvement ability of the system in the multi-module complex code scenario.

[0024] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A data quality detection and improvement system for large language models, characterized in that, The system includes a data parsing and feature extraction module, a multi-dimensional quality detection module, a code optimization and verification module, and a feedback and adaptive optimization module; The data parsing and feature extraction module is responsible for parsing the code dataset, extracting its structural and semantic features, extracting the syntactic features of the code through a static analysis tool, generating semantic embeddings using the semantic embedding model of the LLM model, and constructing a code distribution map to represent the semantic and call relationships between code entities; The multi-dimensional quality detection module comprehensively evaluates the quality of the dataset from three dimensions: semantic diversity, structural consistency, and error patterns, including quantifying the diversity of function implementation methods by analyzing the semantic vector distribution, identifying unreasonable dependencies between modules through the module dependency graph, and predicting potential errors by combining a pre-trained model; Based on the detection results of the multi-dimensional quality detection module, the code optimization and verification module generates optimization suggestions, refactors the code, and verifies the effectiveness of the optimized code through unit tests. It also evaluates the effect of the repair solution in real time through a dynamic evaluation model, and optimizes the repair path and strategy; The feedback and adaptive optimization module conducts runtime performance testing on the optimized code by receiving developer feedback and combining a performance analysis tool, automatically adjusts the repair strategy according to the feedback and test results, and combines a quantum search engine to accelerate the search process for complex repair paths.

2. The data quality detection and improvement system for a large language model according to claim 1, wherein: The data parsing and feature extraction module includes a static analysis tool and a semantic embedding model. The static analysis tool is used to parse the syntactic features of the code, and the semantic embedding model is a deep learning-based LLM model, mainly used to generate semantic features of code fragments and construct a code distribution map. In the map, nodes represent code entities, and edges represent call relationships or semantic similarities. Through a multi-dimensional adaptive code optimization framework that combines multiple dimensions of static, dynamic, semantic, and execution context, comprehensive analysis is performed to automatically identify code quality problems and dynamically adjust optimization strategies to achieve all-round optimization of the code lifecycle.

3. The data quality detection and improvement system for a large language model according to claim 1, characterized in that: The multi-dimensional quality detection module analyzes the semantic vector distribution of code fragments to quantify different implementation methods of the same function, identifies unreasonable dependencies or circular references between modules by constructing a module dependency graph, and simultaneously combines a static analysis tool and a pre-trained model to predict potential errors. It automatically repairs and evolves the detected errors through the Adaptive Program Repair and Evolution Path Optimization Algorithm (APE-RP), and optimizes the repair path and strategy.

4. The data quality detection and improvement system for a large language model according to claim 1, characterized in that: The Adaptive Program Repair and Evolution Path Optimization Algorithm (APE-RP) combines a genetic algorithm and deep reinforcement learning. By simulating the code repair and evolution process, it optimizes the error repair path. The path optimization is continuously adjusted through a feedback mechanism to ensure the accuracy and efficiency of the repair strategy and avoid regression errors.

5. A data quality detection and improvement system for a large language model according to claim 1, characterized in that: After generating optimization suggestions, the code optimization and verification module immediately evaluates the effect of the repair solution using an intelligent feedback adaptive repair evaluation mechanism, captures the behavioral differences before and after the repair, and makes adaptive adjustments through developer feedback and code execution data to optimize the repair strategy.

6. The data quality detection and improvement system for a large language model according to claim 1, wherein: The said feedback and adaptive optimization module combines a quantum computing-assisted adaptive repair path search engine. Through the principles of quantum superposition and quantum interference, it can quickly find the optimal repair solution in a vast repair path space. Especially when facing complex dependency relationships and code interactions, it utilizes the quantum parallel computing ability to accelerate the search for repair paths and synergistically works with classical repair algorithms to improve the repair efficiency.

7. The data quality detection and improvement system for a large language model according to claim 1, characterized in that: The quantum computing-assisted adaptive repair path search engine uses quantum amplitude estimation and quantum search algorithms to perform parallel searches on repair paths, and combines classical repair algorithms to dynamically adjust and optimize the solution, which is particularly suitable for complex code repair tasks with multiple dependencies and multiple modules.

Citation Information

Cited By

  • Code error repairing method and system based on large model

    CN120492319A

  • Method and device for evaluating semantic quality of data set, and electronic equipment

    CN120996027A

  • Method and system for automatically adjusting parameters of power distribution network data quality improvement algorithm

    CN121031999A

  • Unmanned aerial vehicle flight log data analysis method and device, computing equipment and storage medium

    CN121233752A

  • Database execution code automatic generation method and system based on large language model

    CN121478802A