Large model testing system, method, equipment and medium

By evaluating the performance changes of large models during multiple rounds of code iteration optimization tasks and adaptively selecting the target level of feedback detail, the problems of low testing efficiency and high resource consumption of large models are solved, achieving more efficient testing and resource saving.

CN121501637APending Publication Date: 2026-02-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202511633330.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies perform poorly in multi-round iterative tests of code generation capabilities for large models and consume high computational resources.

Method used

By evaluating the performance changes of large models during multiple rounds of code iteration optimization tasks, and adaptively selecting target levels with varying levels of feedback detail, feedback analysis is conducted to help poorly performing large models overcome iteration bottlenecks, improve testing efficiency, and save computational resources.

Benefits of technology

It improves the efficiency and effectiveness of large model testing, avoids unnecessary deep feedback analysis overhead, and saves computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501637A_ABST
    Figure CN121501637A_ABST
Patent Text Reader

Abstract

The invention provides a large model testing system, method and device and a medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of machine learning, deep learning, large models and the like. The large model to be tested is configured to execute a plurality of rounds of code iterative optimization tasks. The system comprises an execution unit configured to execute a code sample generated by a large model in a current round in a sandbox environment to obtain an execution result; the index unit is configured to calculate a performance improvement index based on the execution result, and the performance improvement index represents code generation performance changes of the large model among a plurality of rounds; the feedback unit is configured to select a target level from a plurality of preset diagnosis levels based on the performance improvement index, and the plurality of diagnosis levels respectively represent different feedback information detail degrees; and feedback information is generated based on the target level and the execution result, and the large model executes the next round of code iteration optimization task based on the feedback information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of machine learning, deep learning, and large models, and specifically to a large model testing system, a large model testing method, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include natural language processing, computer vision, speech recognition, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] With the rapid development of large model technology, large models are gradually demonstrating their powerful ability to complete various tasks, including automatically generating corresponding code snippets. Large models can be integrated into various automation tools or development platforms to automate tedious coding tasks. This approach not only significantly improves the programming efficiency of software developers but also assists them in handling complex algorithmic logic and business processes, thereby improving the quality of the final software product and driving the rapid development of the software engineering industry.

[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0005] This disclosure provides a large model testing system, a large model testing method, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] According to one aspect of this disclosure, a large model testing system is provided. The large model to be tested is configured to perform multiple rounds of code iteration optimization tasks. The system includes: an execution unit configured to execute code samples generated by the large model in the current round in a sandbox environment to obtain execution results; an index unit configured to calculate performance improvement indices based on the execution results, wherein the performance improvement indices characterize the changes in code generation performance of the large model across multiple rounds; and a feedback unit configured to: select a target level from multiple preset diagnostic levels based on the performance improvement indices, wherein each of the multiple diagnostic levels characterizes a different level of detail in the feedback information; and generate feedback information based on the target level and the execution results, wherein the large model performs the next round of code iteration optimization tasks based on the feedback information.

[0007] According to another aspect of this disclosure, a method for information processing based on a large model is provided. The method includes: performing multiple rounds of code iteration optimization tasks using a large model to be tested; executing code samples generated by the large model in the current round in a sandbox environment to obtain execution results; calculating performance improvement metrics based on the execution results, whereby the performance improvement metrics characterize the code generation performance changes of the large model across multiple rounds; selecting a target level from multiple preset diagnostic levels based on the performance improvement metrics, wherein each of the multiple diagnostic levels characterizes a different level of detail in the feedback information; and generating feedback information based on the target level and the execution results, wherein the large model performs the next round of code iteration optimization tasks based on the feedback information.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described above.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described method.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described method when executed by a processor.

[0011] According to one or more embodiments of this disclosure, this disclosure evaluates the performance changes of a large model during multiple rounds of code iteration optimization tasks, and based on the performance changes, adaptively selects a more targeted target level from multiple preset diagnostic levels that characterize different levels of detail of feedback information to perform feedback analysis on the execution results of the sample code generated by the large model in the current round. This helps large models with poor performance to get rid of iteration bottlenecks, improve testing efficiency and testing results, and avoid unnecessary deep feedback analysis overhead when the large model performs well, thus saving computing resources.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown; Figure 2 A structural block diagram of a large-scale model testing system according to an embodiment of the present disclosure is shown; Figure 3 A structural block diagram of a large-scale model testing system according to an embodiment of the present disclosure is shown; Figure 4 A flowchart of a large-model testing method according to embodiments of the present disclosure is shown; and Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0016] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0017] The terminology used in the description of the various examples in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0018] In related technologies, existing methods perform poorly in performing multiple rounds of iterative testing on the code generation capabilities of large models, and consume high computational resources.

[0019] To address the aforementioned issues, this disclosure evaluates the performance changes of a large model during multiple rounds of code iteration optimization tasks. Based on these performance changes, it adaptively selects a more targeted target level from multiple pre-defined diagnostic levels representing different levels of detail in feedback information to perform feedback analysis on the execution results of the sample code generated by the large model in the current round. This helps large models with poor performance to overcome iteration bottlenecks, improve testing efficiency and effectiveness, and avoids unnecessary deep feedback analysis overhead when the large model performs well, thus saving computational resources.

[0020] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0021] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0022] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the methods of this disclosure.

[0023] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.

[0024] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0025] Users can use client devices 101, 102, 103, 104, 105, and / or 106 for human-computer interaction. The client devices provide interfaces that enable users to interact with them. The client devices can also output information to the user through these interfaces. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0026] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0027] Network 110 can be any type of network well known to those skilled in the art, and can support data communication using any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0028] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0029] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0030] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.

[0031] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0032] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0033] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0034] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0035] According to one aspect of this disclosure, a large model testing system is provided. The large model to be tested is configured to perform multiple rounds of code iteration optimization tasks. For example... Figure 2 As shown, the system 200 includes: an execution unit 210 configured to execute code samples generated by the large model in the current round in a sandbox environment to obtain execution results; an indicator unit 220 configured to calculate performance improvement indicators based on the execution results, wherein the performance improvement indicators characterize the code generation performance changes of the large model across multiple rounds; and a feedback unit 230 configured to: select a target level from multiple preset diagnostic levels based on the performance improvement indicators, wherein each of the multiple diagnostic levels characterizes a different level of detail in the feedback information; and generate feedback information based on the target level and the execution results, wherein the large model executes the code iteration optimization task for the next round based on the feedback information.

[0036] Therefore, by evaluating the performance changes of a large model during multiple rounds of code iteration optimization tasks, and based on these performance changes, the system adaptively selects a more targeted target level from multiple pre-defined diagnostic levels that characterize different levels of feedback detail to perform feedback analysis on the execution results of the sample code generated by the large model in the current round. This helps large models that are not performing well to get rid of iteration bottlenecks, improve testing efficiency and test results, and avoid unnecessary deep feedback analysis overhead when the large model is performing well, thus saving computing resources.

[0037] The large model described in this disclosure can be a large language model with code generation capabilities. Such a large model typically has billions, tens of billions, or even hundreds of billions of parameters and is trained on massive amounts of text and code data (and optionally, other modalities), enabling it to generate initial code samples based on the requirements of natural language coding tasks, while also possessing the ability to iteratively optimize the code. The ability to iteratively optimize the code refers to the large model's ability to understand and utilize feedback information from the previous iteration to generate optimized code samples.

[0038] In some embodiments, the code iterative optimization task can be performed as follows: In the initial round (e.g., the first round), the large model receives an initial code task requirement (e.g., "generate a Python function to calculate the Fibonacci sequence") and generates a first round of code samples based on that requirement. In subsequent rounds (e.g., the second round and beyond), the large model receives feedback information for the first round of code samples (e.g., generated by the feedback unit 230 mentioned above) and generates code samples for that round based on that feedback information. In this way, the large model continuously receives feedback information from the previous round and generates new rounds of optimized code samples in multiple rounds until a preset termination condition is met.

[0039] In some embodiments, the preset termination conditions may include reaching a predetermined total number of iterations, the performance improvement indicators of the aforementioned indicator unit 220 meeting preset conditions, or other termination conditions, which are not limited here.

[0040] In some embodiments, before testing using the system 200 described above, the model type, the type of code iteration optimization task, and the total number of iterations r can be determined in advance. The type of code iteration optimization task may include, for example, Python function generation, Java full-stack development, etc.

[0041] In some embodiments, execution unit 210 is used to execute the sample code generated in each round of the large model in a sandbox environment and capture the execution results.

[0042] In some embodiments, the sandbox environment can be an isolated execution environment implemented based on container technology (such as Docker containers). The sandbox environment can be configured to support the compilation and execution of multiple mainstream programming languages ​​(e.g., Python, Java, C++, etc.). The sandbox environment can also integrate automated testing frameworks (e.g., pytest, JUnit, etc.) to achieve automated code testing and is configured to capture the execution results of code samples in real time. Furthermore, the sandbox environment may include data verification and retransmission mechanisms to ensure the reliability of execution result transmission and avoid subsequent feedback and metric calculation deviations caused by data packet loss.

[0043] In some embodiments, the execution unit may initialize the sandbox environment during the system initialization phase and load the corresponding programming language compilation environment and testing framework.

[0044] In one exemplary embodiment, the sandbox environment can be designed with differentiated hardware resource configurations for different deployment scenarios. For example, in mainstream servers, each Docker container can be bound to two logical cores and initially allocated 8GB of memory. The execution unit can also be configured to automatically increase the container memory to 16GB through a dynamic memory expansion mechanism when a large code project (such as system code exceeding 5000 lines) is detected, in order to avoid execution interruption due to insufficient memory. In edge servers, each container can be bound to one ARM core and allocated 4GB of memory to adapt to the lightweight deployment requirements of edge devices.

[0045] In some embodiments, the execution result of the code sample may include a success or failure flag for the code execution, as well as a raw error log in case of execution failure. The raw error log may include various errors captured by the sandbox environment during code execution, such as syntax errors, runtime exceptions (e.g., null pointer exceptions, array out-of-bounds errors), and output mismatch issues.

[0046] In some embodiments, to ensure the reliability of the execution result transmission, the execution unit may use the CRC32 checksum algorithm to perform integrity verification on the execution result. If the verification fails, a retransmission mechanism is automatically triggered. After the verification is completed, the execution unit 210 may synchronously transmit the execution result (e.g., in JSON format) to the indicator unit 220 and the feedback unit 230.

[0047] In some embodiments, the metrics unit 220 evaluates the execution results captured by the execution unit 210 to calculate a performance improvement metric. The performance improvement metric, as a comparative measure, guides the feedback unit 230 to adaptively select a more targeted target level by comparing the code generation performance changes of a large model across multiple rounds.

[0048] In some embodiments, the performance improvement metric can be determined based on the performance of the large model in the current iteration round and its performance in another iteration round or at a baseline. The performance improvement metric can be determined based on metrics such as the error rate and number of errors in the execution results of the corresponding round. The other iteration round or baseline can be the previous round, the initial round, or other rounds.

[0049] According to some embodiments, the metric unit 220 can be configured to: determine the current round error rate of the large model based on the execution results; and obtain the first round error rate of the baseline model's execution code iterative optimization task. The performance improvement metric can characterize the reduction in the current round error rate of the large model relative to the first round error rate of the baseline model.

[0050] Therefore, by introducing a fixed baseline model's first-round error rate as a unified benchmark, it is possible to fairly and standardizedly compare the ability and efficiency of different large models under test to reduce errors using feedback information, thus solving the problem of unfair comparison caused by the different initial performance of each model.

[0051] In one exemplary embodiment, the error rate E for the current round can be obtained by dividing the number of code segments that failed to execute in the current round by the total number of code segments. i Furthermore, the first-round error rate E0 of the baseline model's execution code iterative optimization task can be obtained by dividing the number of failed code iterations in the first round by the total number of code iterations. The performance improvement metric can be expressed as: (E0 - E...) i ) / E0× 100%.

[0052] It is understood that the error rate of the current round of the large model and the error rate of the first round of the baseline model can also be determined in other ways, which will not be elaborated here. In another exemplary embodiment, the code iterative optimization task includes an embodiment of multiple subtasks, and the error rate can be the ratio of the number of subtasks that fail to execute to the total number of tasks.

[0053] In some embodiments, the model type of the baseline model can be predetermined before testing. During system initialization, the system can directly retrieve the first-round error rate (E0) of the baseline model from the database and store it in the cache.

[0054] In some embodiments, the feedback unit 230 is used to dynamically select a target level for the current round from a plurality of preset diagnostic levels based on the performance improvement index calculated by the index unit 220, and generate feedback information based on the target level.

[0055] The level of detail in the feedback information characterizes the richness or granularity of the feedback information provided by feedback unit 230 to the large model. Therefore, the richness or granularity of the feedback information obtained varies when different diagnostic levels are used to generate feedback information. Choosing simpler feedback information when the code iteration effect is good can avoid unnecessary overhead of deep feedback analysis, while choosing more complex feedback information when the code iteration effect is poor can help the large model quickly correct errors, reduce the number of test rounds, and save the overall computing resources of the test system.

[0056] According to some embodiments, multiple diagnostic levels can each be associated with one or more preset diagnostic analysis types to characterize different levels of detail in the feedback information. Generating feedback information based on the target level and the execution result may include analyzing the execution result based on one or more preset diagnostic analysis types associated with the target level to generate the feedback information.

[0057] Therefore, by using the above method, the level of detail in feedback information is concretized into one or more specific diagnostic analysis types, thus providing the feedback unit with a configurable mechanism for performing structured analysis of execution results and generating feedback information. This approach allows for the flexible invocation of different analysis logics based on the selected level during testing, thereby enabling control over the level of detail in the feedback information.

[0058] According to some embodiments, one or more preset diagnostic analysis types may include at least one of the following: error rate; error type; error location information; and input-output comparison of failure code samples.

[0059] In some embodiments, the error rate is the error rate of the large model in the current round. This type can indicate the overall performance of the large model.

[0060] In some embodiments, error types can characterize the specific exception category that occurs during the execution of a code sample. Error types can include syntax errors, runtime exceptions (such as IndexError or NullPointerException), logical errors, etc. Therefore, by providing error types to a large model, it is possible to help the large model quickly identify the nature of the error, thereby guiding the large model in the correct direction for correction.

[0061] In some embodiments, error location information can characterize the specific location in the code sample that caused the execution failure. Error location information may include the specific line number where the error occurred, the related code snippet, or the name of a function. Therefore, by providing accurate error location information to a large model, the time and computational resources required for the large model to search for the root cause of errors in lengthy code can be significantly reduced, directly guiding the large model to correct specific parts of the code.

[0062] In some embodiments, comparing the input and output of a failed code sample can characterize the difference between the actual and expected output of the code sample under a specific test case. Therefore, this input-output comparison provides crucial information for the large model to understand boundary conditions and logical flaws, thereby significantly increasing the probability of the large model successfully correcting the code in the next round.

[0063] The diagnostic analysis types described above can provide multi-dimensional correction criteria for large models. It is understood that other types of diagnostic analysis can also be included in the preset options, which are not limited here.

[0064] In one exemplary embodiment, three diagnostic levels can be set: a full level, with associated diagnostic analysis types including error rate, error type, error location information, and input-output comparison of failed code samples; a middle level, with associated diagnostic information including error rate and error type; and a lite level, with associated diagnostic information including error rate.

[0065] According to some embodiments, multiple diagnostic levels can correspond to multiple preset indicator ranges, and a diagnostic level representing a higher level of feedback detail can correspond to a lower indicator range. Selecting a target level from the preset multiple diagnostic levels based on performance improvement indicators can include: in response to determining that a performance improvement indicator falls within one of the multiple indicator ranges, selecting the diagnostic level corresponding to that indicator range as the target level.

[0066] Thus, by adopting the above method, a mapping relationship is established between the indicator range of performance improvement metrics and the diagnostic level, realizing an absolute selection control strategy. In particular, by corresponding the diagnostic level (e.g., fine level) representing a higher level of feedback detail with a lower indicator range, it is ensured that when the system detects poor model iteration performance, it can automatically force a switch to high-detail feedback to improve the efficiency and success rate of iterative optimization.

[0067] In some embodiments, the number of indicator intervals can correspond to the number of diagnostic levels. For example, N indicator intervals can correspond one-to-one with N diagnostic levels.

[0068] In one exemplary embodiment, the system presets three diagnostic levels (e.g., lightweight, medium, and fine), and sets three indicator ranges based on performance improvement metrics: range A (e.g., performance improvement metric below 30%), range B (e.g., performance improvement metric between 30% and 60%), and range C (e.g., performance improvement metric above 60%). The feedback unit can be configured to: select the corresponding fine level in response to determining that the performance improvement metric is in range A; select the corresponding medium level in response to determining that the performance improvement metric is in range B; and select the corresponding lightweight level in response to determining that the performance improvement metric is in range C.

[0069] According to some embodiments, selecting a target level from a set of preset diagnostic levels based on performance improvement metrics may include: obtaining the historical level selected in the previous round; and determining whether to select a diagnostic level that represents a higher or lower level of feedback information detail than the historical level as the target level based on the comparison results between the performance improvement metrics and preset thresholds.

[0070] Therefore, when selecting the target level, the above method not only relies on the current performance improvement indicators, but also on the historical levels selected by the system in the previous round, thereby achieving smoother level adjustments, avoiding frequent level jumps near the boundary of indicator values, and enhancing the stability of the feedback strategy.

[0071] According to some embodiments, determining whether to select a diagnostic level that represents a higher or lower level of feedback detail as the target level based on the comparison result of the performance improvement index and a preset threshold may include: in response to determining that the performance improvement index is lower than a preset first threshold, selecting a diagnostic level that represents a higher level of feedback detail than the historical level as the target level.

[0072] Thus, an iterative optimization upward floating upgrade mechanism is implemented through the above method. When the system detects that the performance improvement index of a large model is lower than the first threshold, it indicates that the model may be stuck in an iterative bottleneck or the correction effect is not good. At this time, the system automatically selects a higher-level diagnostic level (for example, upgrading from "medium level" to "fine level") to provide the large model with richer diagnostic analysis types, so as to help it get out of the bottleneck and improve the efficiency of iterative optimization.

[0073] In some embodiments, the first threshold may be preset or dynamically determined. When selecting a target level, the historical level can be increased by a preset step size (e.g., increased by X levels, where X is an integer greater than or equal to 1), or the preset highest level can be selected directly (e.g., the "fine level" can be selected directly).

[0074] According to some embodiments, determining whether to select a diagnostic level that represents a higher or lower level of feedback detail as the target level based on the comparison result of the performance improvement index and a preset threshold may include: in response to determining that the performance improvement index is higher than a preset second threshold, selecting a diagnostic level that represents a lower level of feedback detail than the historical level as the target level.

[0075] Thus, an iterative optimization mechanism for downward floating degradation is implemented through the above method. When the system detects that the performance improvement index of the large model is higher than the second threshold, it indicates that the current iteration effect of the model is good and may not require costly detailed diagnostic information. At this time, the system automatically selects a lower diagnostic level (e.g., downgrading from "medium level" to "basic level"), thereby avoiding unnecessary deep feedback analysis overhead and effectively reducing the system's computational resource consumption.

[0076] In some embodiments, the second threshold can be preset or dynamically determined. When selecting a target level, the historical level can be reduced by a preset step size (e.g., reduced by Y levels, where Y is an integer greater than or equal to 1), or the preset lowest level can be selected directly (e.g., the "lightweight level" can be selected directly).

[0077] In some embodiments, the first threshold and the second threshold can be set to be the same or different, and X and Y can be set to be the same or different.

[0078] In some embodiments, when performing analysis corresponding to the diagnostic analysis type, the feedback unit can design differentiated parsing logic for different programming languages. In one exemplary embodiment, for Python code, the feedback unit can use the ast module to construct an abstract syntax tree and efficiently extract error information through a node traversal priority strategy. For example, it can prioritize traversing function definition nodes (def) and loop nodes (for / while). This can reduce error location time from 6 seconds in the traditional approach to 1.2 seconds. In another exemplary embodiment, for Java code, the feedback unit can integrate the error diagnosis interface of the Eclipse JDT Core compiler to directly obtain information such as the error code line number and error type output by the compiler, significantly improving the location accuracy.

[0079] According to some embodiments, a code iteration optimization task may include multiple subtasks. The code iteration optimization task may be a benchmark set containing multiple independent test cases or code requirements. The multiple subtasks may correspond to each independent test case or code requirement in the benchmark set (e.g., an algorithm problem or a feature implementation).

[0080] Figure 3 A structural block diagram of a large-model testing system 300 according to an exemplary embodiment of the present disclosure is shown. Figure 3 As shown, system 300 includes an execution unit 310, an indicator unit 320, a feedback unit 330, and a memory unit 340. The memory unit 340 is configured to store the historical progress states of multiple subtasks. It is understood that the implementation details of the execution unit 310, indicator unit 320, and feedback unit 330 can be referred to the description of the execution unit 210, indicator unit 220, and feedback unit 230 in system 200 above, and will not be repeated here.

[0081] The execution unit 310 is further configured to: read the historical pass status of each of the multiple subtasks from the memory unit to determine one or more target subtasks that are still in a failed state before the current round; and prioritize the execution of the code in the code sample corresponding to one or more target subtasks.

[0082] Therefore, by caching and utilizing historical pass states in memory unit 340, an intelligent execution scheduling strategy is implemented. By prioritizing the execution of code corresponding to target subtasks that are still in a failed state, the system can avoid unnecessary and repetitive execution and verification on subtasks that have already passed the test, thereby significantly reducing invalid computation, lowering resource consumption in the sandbox environment, and improving the overall system efficiency of multi-round iterative testing.

[0083] In some embodiments, the history can be updated and stored by the feedback unit 330 or the memory unit 340 after the execution result (e.g., a success or failure flag) is captured by the execution unit 310 in the previous round.

[0084] In some embodiments, priority execution may include: the execution unit 310 reads the status list of all subtasks from the memory unit 340 in the current round, constructs an execution queue containing only the target subtasks that are still in a failed state, and prioritizes the execution of the code corresponding to the target subtask in the queue.

[0085] According to some embodiments, the historical access status of high-frequency accesses can be stored in a first type of storage medium, and the historical access status of low-frequency accesses can be stored in a second type of storage medium. The access speed of the first type of storage medium is higher than that of the second type of storage medium.

[0086] Therefore, by storing frequently accessed data in high-speed media, fast access to execution units and indicator units and low system latency are ensured, while by migrating infrequently accessed data to lower-cost media, cache overflow is avoided, and the stability and scalability of the system when processing large-scale subtasks are improved.

[0087] In some embodiments, the first type of storage medium can be server memory and can be managed via a Redis database. The second type of storage medium can be a solid-state drive (SSD). In an exemplary embodiment, the memory unit can be configured to store high-frequency access data from the most recent three rounds (e.g., historical pass status, performance improvement metrics, etc.) in memory and store low-frequency access history data from three rounds prior in the SSD.

[0088] In some embodiments, the memory unit may also be configured to store the historical level selected by the feedback unit in the previous round.

[0089] According to some embodiments, the execution unit can be further configured to: monitor the CPU utilization of the sandbox environment in real time; and, in response to the CPU utilization meeting a preset condition for a continuous preset duration, terminate the execution of the code sample and release the computing resources of the sandbox environment.

[0090] Thus, an automated exception handling mechanism is achieved through the above methods. This mechanism effectively prevents execution stalls caused by code samples under test (e.g., code containing infinite loops or infinite recursion), avoiding the entire test system from hanging or crashing due to a single exception sample. By promptly terminating exception execution and releasing computing resources, the robustness and stability of the large-scale model testing system are improved, ensuring the continuous and reliable operation of the testing process.

[0091] In one exemplary embodiment, the continuous preset duration is 5 seconds, and the preset condition is that the CPU utilization rate is 100%.

[0092] In some embodiments, terminating the execution of the code sample may include: the execution unit sending a termination signal (e.g., a SIGKILL signal) to the process in the sandbox environment (e.g., a Docker container). Releasing computing resources may be configured to complete within a preset time (e.g., within 500ms) to ensure that resources are quickly reclaimed.

[0093] According to some embodiments, the feedback unit can be further configured to generate optimized loop prompt information as feedback information in response to determining that the code sample has been terminated.

[0094] Therefore, by using the above method, an execution-level exception event is transformed into a structured feedback signal that can be used for iterative optimization. This mechanism allows the large model to detect code failures due to execution stagnation, enabling targeted optimization of loop termination conditions or recursion exits in the next round, further improving testing efficiency.

[0095] In some embodiments, the optimization loop prompt information may include specific types of feedback, such as "execution stall risk warning" or "infinite loop risk warning." This feedback information can clearly indicate that the current code of the large model has a risk of causing execution timeout, guiding it to supplement or correct the loop termination judgment logic in the next round.

[0096] According to some embodiments, the metric unit can be further configured to output performance improvement metrics as a test result of the code iterative optimization capability of a large model.

[0097] Therefore, the above method provides an objective and quantifiable evaluation standard. The test results can intuitively reflect the iterative optimization capability of the large model under test in correcting errors and improving code quality using feedback information, while also enabling testers to obtain the performance improvement status of the current round.

[0098] In some embodiments, the metrics unit can visualize the changing trends of performance improvement metrics across multiple rounds and push the visualization results to the user interface. This allows users to view evaluation trends in real time, improving the user experience.

[0099] In some embodiments, the metrics unit may also store test results in a memory unit (e.g., in a database or file) for subsequent analysis. Test results may be presented as a numerical matrix categorized by model type, subtask set, and iteration rounds, or as a visualization (e.g., a performance improvement curve).

[0100] In some embodiments, the metric unit may be configured to calculate cumulative performance metrics. The cumulative performance metric indicates the number or proportion of subtasks that have been successfully solved at least once after r rounds of iteration. Here, the large model generates k code samples for the same subtask in each round. The cumulative performance metric can quantify the ability of the large model to cumulatively complete subtasks as the number of iteration rounds increases.

[0101] In some embodiments, the metric unit may continuously monitor the cumulative performance metric. The metric unit may be configured to: in response to determining that the cumulative performance metric exceeds a preset range, verify the data integrity of the execution result of the current round. In response to determining that the data verification fails (e.g., data tampering or packet loss occurs), the metric unit may instruct the execution unit to re-execute the code samples marked as failed in the current round. By this means, computing resources can be saved. In response to determining that the data verification is successful (i.e., the data is complete but the calculation result is abnormal), the metric unit may check the calculation logic of the cumulative performance metric.

[0102] In an exemplary embodiment, when there is a situation of insufficient sample number (e.g., n - i + 1 < k) in the calculation logic, the metric unit may automatically correct the corresponding combination number to 0 and recalculate the cumulative performance metric to ensure that its value is within a reasonable preset range (e.g., 0% - 100%).

[0103] According to another aspect of the present disclosure, a method for testing a large model is provided. As Figure 4 shown, method 400 includes: step S401, performing multiple rounds of code iteration optimization tasks using the large model to be tested; step S402, executing the code samples generated by the large model in the current round in a sandbox environment to obtain an execution result; step S403, calculating a performance improvement metric based on the execution result, where the performance improvement metric characterizes the change in code generation performance of the large model among multiple rounds; step S404, selecting a target level from a preset multiple diagnostic levels, where the multiple diagnostic levels respectively characterize different degrees of detail of feedback information; and step S405, generating feedback information based on the target level and the execution result, where the large model performs the next round of code iteration optimization task based on the feedback information.

[0104] It can be understood that the operations of steps S401 to S405 in method 400 may refer to the descriptions of units 210 to 230 in system 200 and the large model to be tested above, and will not be elaborated here.

[0105] According to some embodiments, multiple diagnostic levels can each be associated with one or more preset diagnostic analysis types to characterize different levels of detail in the feedback information. Step S405, generating feedback information based on the target level and the execution result, may include: analyzing the execution result based on one or more preset diagnostic analysis types associated with the target level to generate feedback information.

[0106] According to some embodiments, multiple diagnostic levels can correspond to multiple preset indicator ranges, and a diagnostic level representing a higher level of feedback detail can correspond to a lower indicator range. Step S404, selecting a target level from the multiple preset diagnostic levels based on performance improvement indicators includes: in response to determining that the performance improvement indicator is located in one of the multiple indicator ranges, selecting the diagnostic level corresponding to that indicator range as the target level.

[0107] According to some embodiments, step S404, selecting a target level from a plurality of preset diagnostic levels based on performance improvement indicators, may include: obtaining the historical level selected in the previous round; and determining whether to select a diagnostic level that represents a higher or lower level of feedback information detail than the historical level as the target level based on the comparison result between the performance improvement indicators and preset thresholds.

[0108] According to some embodiments, the code iterative optimization task includes multiple subtasks. The large model testing method may further include: obtaining the historical pass status of each of the multiple subtasks to determine one or more target subtasks that are still in a failed state before the current round. Step S402, executing the code sample generated by the large model in the current round in the sandbox environment, and obtaining the execution result may include: prioritizing the execution of code in the code sample corresponding to one or more target subtasks.

[0109] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0110] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0111] refer to Figure 5The present invention describes a structural block diagram of an electronic device 500 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0112] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0113] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, disk and optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0114] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods, processes, and / or processes described above. For example, in some embodiments, these methods, processes, and / or processes may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods, processes, and / or processes described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform these methods, processes, and / or processes by any other suitable means (e.g., by means of firmware).

[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0116] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0117] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0120] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0121] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0122] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A large-scale model testing system, wherein, The large model to be tested is configured to perform multiple rounds of code iteration optimization tasks, and the system includes: The execution unit is configured to execute the code sample generated by the large model in the current round in a sandbox environment to obtain the execution result; The metrics unit is configured to calculate a performance improvement metric based on the execution results, the performance improvement metric representing the change in code generation performance of the large model across the multiple rounds; and The feedback unit is configured as follows: Based on the performance improvement metrics, a target level is selected from a set of pre-defined diagnostic levels, where each of the multiple diagnostic levels represents a different level of detail in the feedback information; and Based on the target level and the execution result, feedback information is generated, wherein the large model executes the next round of code iteration optimization task based on the feedback information.

2. The system according to claim 1, wherein, Each of the multiple diagnostic levels is associated with one or more preset diagnostic analysis types to characterize different levels of detail in the feedback information. The step of generating feedback information based on the target level and the execution result includes: The execution results are analyzed based on one or more preset diagnostic analysis types associated with the target level to generate the feedback information.

3. The system according to claim 2, wherein, The one or more preset diagnostic analysis types include at least one of the following: Error rate; Error type; Error location information; and Input and output comparison of failed code samples.

4. The system according to any one of claims 1 to 3, wherein, The multiple diagnostic levels correspond to multiple preset indicator ranges, and the diagnostic level representing a higher level of feedback detail corresponds to a lower indicator range. The selection of a target level from the preset multiple diagnostic levels based on the performance improvement indicators includes: In response to determining that the performance improvement metric falls within one of the plurality of metric intervals, the diagnostic level corresponding to that metric interval is selected as the target level.

5. The system according to any one of claims 1 to 3, wherein, The selection of a target level from a preset set of diagnostic levels based on the performance improvement indicators includes: Retrieve the historical level selected in the previous round; and Based on the comparison results between the performance improvement index and the preset threshold, it is determined whether to select a diagnostic level that represents a higher or lower level of feedback information detail than the historical level as the target level.

6. The system according to claim 5, wherein, The step of determining whether to select a diagnostic level representing higher or lower levels of feedback detail as the target level based on the comparison results between the performance improvement index and a preset threshold includes: In response to determining that the performance improvement metric is below a preset first threshold, a diagnostic level that represents a higher level of feedback detail than the historical level is selected as the target level.

7. The system according to claim 5, wherein, The step of determining whether to select a diagnostic level representing higher or lower levels of feedback detail as the target level based on the comparison results between the performance improvement index and a preset threshold includes: In response to determining that the performance improvement metric is higher than a preset second threshold, a diagnostic level with a lower level of feedback detail than the historical level is selected as the target level.

8. The system according to any one of claims 1 to 3, wherein, The indicator unit is configured as follows: Based on the execution results, determine the error rate of the current round of the large model; and Obtain the first-round error rate of the baseline model when performing the code iterative optimization task. The performance improvement metric represents the reduction in the error rate of the current round of the large model relative to the error rate of the first round of the baseline model.

9. The system according to any one of claims 1 to 3, wherein, The code iteration optimization task includes multiple sub-tasks, and the system also includes: The memory unit is configured to store the historical pass-through state of each of the plurality of subtasks. The execution unit is further configured as follows: Read the historical pass status of each of the plurality of subtasks from the memory unit to determine one or more target subtasks that were still in a failed state before the current round; and The code in the code sample that corresponds to the one or more target subtasks will be executed first.

10. The system according to claim 9, wherein, The history of high-frequency accesses is stored in a first type of storage medium through state storage, and the history of low-frequency accesses is stored in a second type of storage medium through state storage, wherein the access speed of the first type of storage medium is higher than that of the second type of storage medium.

11. The system according to any one of claims 1 to 3, wherein, The execution unit is further configured as follows: Real-time monitoring of CPU utilization in the sandbox environment; and In response to the CPU utilization rate meeting a preset condition for a continuous preset duration, the execution of the code sample is terminated and the computing resources of the sandbox environment are released.

12. The system according to claim 11, wherein, The feedback unit is further configured as follows: In response to determining that the code sample has been terminated, an optimized loop prompt message is generated as feedback information.

13. A large model testing method, comprising: Perform multiple rounds of code iteration optimization tasks using the large model to be tested; Execute the code sample generated by the large model in the current round in the sandbox environment to obtain the execution result; Based on the execution results, a performance improvement index is calculated, which characterizes the code generation performance change of the large model across the multiple rounds. Based on the performance improvement metrics, a target level is selected from a set of pre-defined diagnostic levels, where each of the multiple diagnostic levels represents a different level of detail in the feedback information; and Based on the target level and the execution result, feedback information is generated, wherein the large model executes the next round of code iteration optimization task based on the feedback information.

14. The method according to claim 13, wherein, Each of the multiple diagnostic levels is associated with one or more preset diagnostic analysis types to characterize different levels of detail in the feedback information. The step of generating feedback information based on the target level and the execution result includes: The execution results are analyzed based on one or more preset diagnostic analysis types associated with the target level to generate the feedback information.

15. The method according to claim 13 or 14, wherein, The multiple diagnostic levels correspond to multiple preset indicator ranges, and the diagnostic level representing a higher level of feedback detail corresponds to a lower indicator range. The selection of a target level from the preset multiple diagnostic levels based on the performance improvement indicators includes: In response to determining that the performance improvement metric falls within one of the plurality of metric intervals, the diagnostic level corresponding to that metric interval is selected as the target level.

16. The method according to claim 13 or 14, wherein, The selection of a target level from a preset set of diagnostic levels based on the performance improvement indicators includes: Retrieve the historical level selected in the previous round; and Based on the comparison results between the performance improvement index and the preset threshold, it is determined whether to select a diagnostic level that represents a higher or lower level of feedback information detail than the historical level as the target level.

17. The method according to claim 13 or 14, wherein, The code iteration optimization task includes multiple sub-tasks, and the method further includes: Obtain the historical progress status of each of the multiple subtasks to identify one or more target subtasks that are still in a failed state before the current round. Specifically, executing the code sample generated by the large model in the current round within the sandbox environment yields the following execution results: The code in the code sample that corresponds to the one or more target subtasks will be executed first.

18. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 13-17.

19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, Computer instructions are used to cause a computer to perform the method according to any one of claims 13-17.

20. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 13-17.