An extreme test method, system, device and medium of an artificial intelligence model

By building scenario templates to generate test cases and using heuristic search algorithms to determine the optimal parameter combinations for automated testing, the problem of insufficient test coverage of AI models under extreme conditions is solved, achieving efficient and comprehensive testing and model optimization, and improving the stability of game products and user experience.

CN119557215BActive Publication Date: 2025-11-28GUANGZHOU YINGFENG NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411604123.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-11-28
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing technologies cannot fully cover the testing of AI models under extreme conditions, resulting in low testing efficiency and the easy omission of important issues, making it difficult to meet the complex and ever-changing game environment and the unpredictability of player behavior.

Method used

By constructing multiple scenario templates, a set of test cases is generated. A heuristic search algorithm is used to determine the optimal combination of test parameters for automated testing. The model's running status is monitored through a game data acquisition component, and test reports are generated for optimization and retesting.

Benefits of technology

It has achieved full automation and intelligence in AI model testing, significantly improving testing efficiency and model quality, ensuring the stability of game products and user experience, and comprehensively covering performance and stability under extreme conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557215B_ABST
    Figure CN119557215B_ABST
Patent Text Reader

Abstract

The application provides a limit test method, system, equipment and medium of an artificial intelligence model, and the method specifically comprises the following steps: constructing a plurality of scene templates for different types of games, generating a test case set by combining and transforming each scene template according to a boundary condition input by a user; based on the test case set, determining an optimal test parameter combination satisfying the artificial intelligence model by using a heuristic search algorithm according to a preset optimization objective function; automatically testing the artificial intelligence model according to the optimal test parameter combination, generating a model test result; determining the performance of the artificial intelligence model under the boundary condition by analyzing the model test result, generating a model test report, and optimizing and retesting the artificial intelligence model through the model test report. The application realizes full-process automation and intelligentization of game AI model testing, and significantly improves the test efficiency and model quality.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a limit test method, system, device and medium for an artificial intelligence model. BACKGROUND

[0002] With the rapid development of artificial intelligence (AI) technology, especially the widespread application of large language models (such as Claude, GPT-4) and image processing models (such as Haiper), the game development industry is undergoing an unprecedented revolution. These advanced AI models are integrated into game customer service systems, translation platforms, art image generation, and other self-developed systems, greatly improving the efficiency and quality of game development, operation, and user experience. They can not only handle complex natural language interactions and generate high-quality image content, but also achieve seamless multilingual translation, providing strong support for the global game market.

[0003] However, although these AI models have shown excellent performance in standard test environments, their stability and reliability face serious challenges when faced with extreme or boundary conditions. The game environment is complex and variable, and player behavior is unpredictable. These extreme conditions may include high concurrency requests, abnormal inputs, resource limitations, and other scenarios, which pose higher requirements for the robustness of AI models.

[0004] Currently, boundary testing for AI models mainly relies on manually designed test cases. This method is not only time-consuming and labor-intensive, but also difficult to comprehensively cover all possible extreme conditions, resulting in low testing efficiency and the risk of missing important issues. In addition, with the iterative updates of game versions, new boundary conditions are constantly emerging, and traditional manual testing methods are difficult to keep up with this changing pace, further exacerbating the difficulty and complexity of testing work. SUMMARY

[0005] The present application aims to provide a limit test method, system, device and medium for an artificial intelligence model, which realizes the full-process automation and intelligentization of game AI model testing, significantly improves the testing efficiency and model quality, and provides a strong guarantee for the stability and user experience of game products, to solve at least one of the above-mentioned prior art problems.

[0006] In a first aspect, the present application provides a limit test method for an artificial intelligence model, which specifically comprises:

[0007] A plurality of scene templates are constructed for different types of games, and a test case set is generated by combining and transforming each scene template according to user input boundary conditions;

[0008] Determine, based on the test case set, an optimal test parameter combination satisfying the artificial intelligence model according to a preset optimization objective function by using a heuristic search algorithm.

[0009] Perform automated testing on the artificial intelligence model according to the optimal test parameter combination, generate a model test result, and monitor the running state of the artificial intelligence model in the process of automated testing by deploying a game data acquisition component.

[0010] Determine the performance of the artificial intelligence model under boundary conditions by analyzing the model test result, generate a model test report, and optimize and retest the artificial intelligence model through the model test report.

[0011] In a second aspect, the present application provides an extreme test system for an artificial intelligence model, which specifically comprises:

[0012] A first test module is configured to construct multiple scene templates for different types of games, generate a test case set by combining and transforming each scene template according to user input boundary conditions;

[0013] A second test module is configured to determine an optimal test parameter combination satisfying the artificial intelligence model based on the test case set according to a preset optimization objective function by using a heuristic search algorithm.

[0014] A third test module is configured to perform automated testing on the artificial intelligence model according to the optimal test parameter combination, generate a model test result, and monitor the running state of the artificial intelligence model in the process of automated testing by deploying a game data acquisition component.

[0015] A fourth test module is configured to determine the performance of the artificial intelligence model under boundary conditions by analyzing the model test result, generate a model test report, and optimize and retest the artificial intelligence model through the model test report.

[0016] In a third aspect, the present application provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory, and when the computer program is executed on the processor, the extreme test method for an artificial intelligence model is realized.

[0017] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and when the computer program is executed on a processor, the extreme test method for an artificial intelligence model is realized.

[0018] Compared with the prior art, the present application has at least one of the following technical effects:

[0019] 1、The application realizes the full-process automation and intelligentization of game AI model testing, significantly improves the testing efficiency and model quality, and provides a strong guarantee for the stability and user experience of game products.

[0020] 2、The application can automatically generate diversified test cases, comprehensively cover and test the stability and performance of AI models under extreme conditions, not only significantly improve the testing efficiency and quality, but also effectively guarantee the reliability and security of the game system, and provide players with a more stable and smooth game experience.

[0021] 3、The application provides an efficient and comprehensive AI model limit testing method for the game development industry, which helps to improve the reliability and security of the game system and promote the in-depth application and development of AI technology in the game field.

[0022] 4、The application automatically generates test cases, reduces manual intervention, and significantly shortens the testing period.

[0023] 5、The application generates diversified test cases by combining and transforming scene templates, and comprehensively covers possible extreme conditions.

[0024] 6、The application determines the optimal test parameter combination based on the heuristic search algorithm to ensure the maximization of the testing effect.

[0025] 7、The application deploys a game data collection component to monitor the running state of the AI model in real time, and adjusts and retests according to the test results to improve the model performance.

[0026] 8、The application combines player portrait data and deep reinforcement learning algorithm to optimize and adjust the AI model individually, and improves the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0028] Figure 1 is a flowchart of a limit testing method of an artificial intelligence model provided by an embodiment of the present application;

[0029] Figure 2 is a structural schematic diagram of a limit testing system of an artificial intelligence model provided by an embodiment of the present application;

[0030] Figure 3 is a structural schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0031] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and

[0032] It will be understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0033] It will be understood that the term "and / or," when used in the specification and in the following claims, is intended to mean one or more of the associated listed items can be present, and, further, that unless otherwise managed, one or more of such associated listed items can be present in 100% of the cases.

[0034] As used in this specification and claims, the terms "if", "for example", and "for instance" can be used interchangeably to mean "when" or "in response to a determination" or "in response to a detection." Similarly, the phrase "if determined" or "if detected" can be construed to mean "upon a determination" or "in response to a determination" or "upon a detection" or "in response to a detection."

[0035] In addition, the terms "first", "second", "third", etc. as used in the description and the claims of this specification, are used as labels for the purpose of distinguishing between items that are to be distinguished from one another. In no way are these terms to be interpreted as indicating relative importance.

[0036] Reference to "one embodiment" or "some embodiments" or "one implementation" or "some implementations" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" or "in some embodiments" or "in one implementation" or "in some implementations" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily referring to a single implementation. Furthermore, the terms "comprise," "comprising," "include," "including," "contain," "containing," "have," "having," and the like mean "including but not limited to."

[0037] In the embodiments of the present application, the execution subject of the flow includes a terminal device. The terminal device includes but is not limited to a server, a computer, a smart phone, a tablet computer and other devices capable of executing the method disclosed in the present application. Figure 1 A flowchart of a limit test method of an artificial intelligence model disclosed in an embodiment of the present application is shown, and is described in detail as follows:

[0038] S101, a plurality of scene templates are constructed for different types of games, and a test case set is generated by combining and transforming each scene template according to the boundary conditions input by the user.

[0039] In the present embodiment, according to the boundary conditions input by the user, the system can automatically select and combine corresponding scene templates, and transform the elements in the templates to meet specific needs. For example, in an RPG game, according to the condition of "player level 10, task type treasure hunting", the system selects the "dungeon" scene template, and adds monsters and treasure chests matching the player's level, while adjusting the difficulty of the maze to adapt to the task requirements. In a shooting game, if the user requires "nighttime battle on city streets", the system selects the "city street" scene template and adjusts the lighting effect to nighttime mode, adding enemy types and tactical strategies specific to nighttime.

[0040] By combining and transforming scene templates, the system can generate a series of specific test case sets. Each test case contains detailed information such as scene description, player state, task objective, and expected result. These test cases can be used for subsequent automated testing or manual testing to ensure that the game can run normally and meet user needs under different scenarios and conditions.

[0041] In the present embodiment, by constructing diverse scene templates and combining and transforming them according to the boundary conditions input by the user, a large number of test cases can be generated, significantly improving the coverage and comprehensiveness of the test.

[0042] In some embodiments, in the above step S101, the plurality of scene templates are constructed for different types of games, and a test case set is generated by combining and transforming each scene template according to the boundary conditions input by the user, specifically including:

[0043] A plurality of scene templates are constructed for different types of games to form a scene template library, and the scene templates include a plurality of key elements and a plurality of constraint conditions;

[0044] The boundary condition data input by the user is obtained, and the boundary condition data and each scene template of the scene template library are matched to extract a plurality of first scene templates;

[0045] An initial test case set is generated by traversing combinations of key elements and constraints of each first scene template using a depth-first exploration algorithm.

[0046] A greedy algorithm is used to analyze the coverage and diversity of each initial test case in the initial test case set, obtaining a target test case set.

[0047] In this embodiment, first define several different types of games, such as role-playing games (RPG), shooting games, strategy games, etc. Each game type has its unique elements and gameplay, for example: RPG scene template, key elements include character level, equipment type, monster species, map type, etc., constraints include character level needs to be higher than monster level to cause damage, equipment type affects character attributes, certain maps have specific monsters and events, etc.; shooting game scene template, key elements include weapon type, enemy type, terrain (such as open land, building), etc., constraints include different weapons have different damage effects on enemies, certain terrain will affect shooting accuracy and concealment, etc.; strategy game scene template, key elements include resource types (such as gold, wood), unit types, terrain (such as plains, mountains), etc., constraints include resource allocation affects unit production and upgrade, different terrain affects unit movement speed and combat power, etc.

[0048] The above scene templates are stored in a database to form a scene template library, each template has a unique identifier and description information.

[0049] The user inputs specific boundary conditions of the game through the interface, such as "in the RPG game, the character level is 10, the equipment is a primary sword, and it is facing a medium-level monster in the forest". Match the user input boundary conditions with each template in the scene template library, extract the scene template that meets the conditions as the first scene template. Starting from the first scene template, use the depth-first search (DFS) algorithm to traverse all possible combinations of key elements and constraints. For example, consider different combinations of character level, equipment, and monsters to generate multiple initial test cases.

[0050] Apply the greedy algorithm to the initial test case set, the goal is to maximize the coverage and diversity of test cases. Among them, coverage analysis is used to ensure that test cases can cover as much game logic and scene as possible, and diversity analysis is used to reduce redundancy between test cases, ensuring that each test case provides unique information. Finally, a concise and efficient target test case set is obtained.

[0051] In this embodiment, by constructing a scene template library and an automatic matching mechanism, test cases for specific game types and boundary conditions are quickly generated, reducing the time and cost of manually writing test cases. The depth-first exploration algorithm ensures that the test cases can cover complex game logic and scenarios, while the optimization of the greedy algorithm further improves the comprehensiveness and depth of the test. By maximizing the coverage and diversity of test cases, it ensures that the test can find more potential problems and defects, improving the quality and stability of the game.

[0052] In step S102, based on the test case set, a heuristic search algorithm is used to determine the optimal test parameter combination that satisfies the artificial intelligence model according to the preset optimization objective function.

[0053] In this embodiment, first, according to the different scenes and requirements of games or software, a test case set is constructed and generated. These test cases contain various possible input conditions and expected outputs for evaluating the performance and stability of the artificial intelligence model. In order to determine the optimal test parameter combination, a preset optimization objective function is needed. This function usually measures the model performance under different parameter combinations based on the execution results of the test cases, such as accuracy, recall rate, F1 score, execution time, etc.

[0054] The heuristic search algorithm can efficiently search in the solution space to find parameter combinations close to the optimal solution. For example, the genetic algorithm may include the following steps: 1. Initialize the population: randomly generate a set of parameter combinations as the initial population. 2. Fitness evaluation: evaluate the fitness of each parameter combination using the preset optimization objective function. 3. Selection operation: select excellent parameter combinations for genetic operation according to the fitness. 4. Crossover and mutation: perform crossover and mutation operations on the selected parameter combinations to generate new parameter combinations. 5. Iterative optimization: repeat steps 2-4 until the stopping condition is met (such as reaching the maximum number of iterations, the fitness no longer significantly improves, etc.). After multiple iterations, the algorithm outputs a set of parameter combinations with the highest fitness, which is the optimal test parameter combination that satisfies the artificial intelligence model.

[0055] In this embodiment, the heuristic search algorithm can efficiently search in the solution space, quickly find parameter combinations close to the optimal solution, and thus shorten the test period. Testing the model under different parameter combinations helps to discover potential errors and instability factors, thereby enhancing the stability and reliability of the model.

[0056] In some embodiments, in step S102, based on the test case set, a heuristic search algorithm is used to determine the optimal test parameter combination that satisfies the artificial intelligence model according to the preset optimization objective function, which specifically includes:

[0057] acquire game project requirement data and specific requirement data of the artificial intelligence model, construct an optimization objective function according to the game project requirement data and the specific requirement data, the optimization objective function including model inference speed, memory occupation performance index and resource consumption index;

[0058] divide the test case set into a plurality of parameter subspaces, each parameter subspace representing a similar game scenario;

[0059] In each of the parameter subspaces, a simulated annealing algorithm is used for search optimization to obtain a plurality of test parameter combinations;

[0060] By evaluating the model inference speed, memory occupation and resource consumption index corresponding to each test parameter combination, it is determined whether the requirements of the optimization objective function are met. If the requirements are met, the test parameter combination is added to the optimal solution set, and the optimal test parameter combination is obtained according to the optimal solution set.

[0061] In this embodiment, an artificial intelligence model is constructed for a newly developed online multiplayer shooting game for enemy behavior prediction and intelligent path planning. The game project requires the model to have high inference speed, low memory occupation and reasonable resource consumption to ensure the smoothness of the game and the experience of the players. The game project requirement data includes that the game needs to support at least 100 players online at the same time, the response time of the enemy AI needs to be less than 50 milliseconds, and the game needs to have good performance on different devices (including low-end mobile phones, high-end PCs, etc.); the specific requirement data of the artificial intelligence model includes that the model needs to be able to handle complex game environments (such as dynamic obstacles, different terrains and weather conditions), the accuracy of the model needs to be more than 90%, and the memory occupation of the model cannot exceed 10% of the total memory of the game.

[0062] Based on the above requirements, the optimization objective function is constructed as follows:

[0063]

[0064] Wherein, α, β, γ are weight coefficients, which can be adjusted according to the priority of the project, Inference Speed is the model inference speed, Memory Usage is the memory occupied by the model during running, Total Memory is the total memory of the game, ResourceConsumption is the resource consumption (such as CPU usage, network bandwidth, etc.) of the model during running, and text Objective is the total optimization objective of the model.

[0065] The test case set is divided into multiple parameter subspaces according to game scene characteristics (such as map type, number of players, enemy density, etc.). For example: parameter subspace A: small map, small number of players, low-density enemies; parameter subspace B: large map, large number of players, high-density enemies; parameter subspace C: complex terrain (such as mountains, cities), medium number of players, medium-density enemies.

[0066] In each parameter subspace, the simulated annealing algorithm is used to search for the optimal test parameter combination. The simulated annealing algorithm gradually finds the global optimal solution in the parameter space by simulating the solid annealing process in physics, and the specific steps are as follows:

[0067] Initialization: randomly select a test parameter combination as the current solution and set the initial temperature.

[0068] Iteration: Randomly generate a new solution in the neighborhood of the current solution, calculate the change of the objective function value. If the new solution is better or meets a certain probability (related to the current temperature and the quality of the solution), accept the new solution as the current solution.

[0069] Cooling: As the number of iterations increases, gradually reduce the temperature to make the algorithm gradually converge to the global optimal solution.

[0070] The test parameter combination obtained by the simulated annealing algorithm in each parameter subspace is evaluated, and the corresponding model inference speed, memory usage and resource consumption indicators are calculated to determine whether it meets the requirements of the optimization objective function. If it meets the requirements, the test parameter combination is added to the optimal solution set. Finally, the test parameter combination with the best overall performance is selected from the optimal solution set as the final solution.

[0071] In this embodiment, by constructing the optimization objective function and applying the simulated annealing algorithm, a test parameter combination that performs well in inference speed, memory usage and resource consumption can be found, thereby improving the overall performance of the model.

[0072] Further, the optimal test parameter combination is obtained from the optimal solution set, specifically including:

[0073] Using the K-means clustering algorithm, the optimal solution set is divided into K clusters;

[0074] Selecting a cluster center point from each cluster as a representative sample point to form a sample subset;

[0075] For each sample point, using the boundary value analysis method, generate corresponding boundary value test data for game state parameters and event trigger conditions;

[0076] Among them, for the numerical type parameters, the upper and lower limits are selected as the boundary values, for the Boolean type parameters, true and false are selected as the boundary values, and for the enumeration type parameters, all enumeration values are selected as the boundary values;

[0077] Using the orthogonal table design method, an orthogonal experiment group is constructed from the sample subset and the boundary value test data, and each factor of the orthogonal experiment group represents a game state parameter or an event trigger condition.

[0078] Select a test parameter combination that meets the orthogonality from the orthogonal test group, and use the test parameter combination that meets the orthogonality as the optimal test parameter combination.

[0079] In this embodiment, a set of representative and efficient test parameter combinations are generated through further clustering, sample selection, boundary value analysis and orthogonal table design, which are used to verify the stability and performance of the game under different conditions.

[0080] Each test parameter combination in the optimal solution set is regarded as a data point, and the K-means algorithm is used for clustering. The algorithm iteratively assigns data points to the nearest cluster center and updates the position of the cluster center until the stopping condition is met (such as the cluster center no longer changes or the maximum number of iterations is reached). In each cluster, the cluster center point is selected as the representative sample point of the cluster. These center points are usually located at the center of the cluster and can better reflect the common characteristics of the test parameter combinations in the cluster.

[0081] The boundary value analysis method is used to generate boundary value test data. For numerical type parameters, for example, if a parameter in the game is "player speed" and its normal range is [1, 10], then 1 (lower limit) and 10 (upper limit) are selected as boundary value test data. For Boolean type parameters, such as "whether to enable automatic aiming", true and false are directly selected as boundary values. For enumeration type parameters, such as "weapon type" contains {pistol, rifle, sniper}, all enumeration values are selected as boundary value test data.

[0082] Each cluster centroid (representing a different game scenario or configuration) and the boundary value test data obtained through boundary value analysis are considered as factors in the orthogonal experiment. The level of each factor is the different values ​​of that factor (i.e., the specific parameter combination of the cluster centroid or the boundary value). An appropriate orthogonal array is selected based on the number of factors and the number of levels for each factor. For example, if there are 3 factors, each with 2 levels (such as two cluster centroids and their boundary value tests), then L4 (2^3 represents 3 factors, each with 2 levels) can be selected. Following the arrangement of the orthogonal array, test parameter combinations are selected from the sample subset and boundary value test data to construct the orthogonal experimental group. Each row of the orthogonal array represents a test parameter combination, ensuring that the different levels of each factor appear evenly in the experiment, while reducing the number of experiments.

[0083] Since the orthogonal array design already guarantees the orthogonality of the experiments (i.e., the independence between factors), all combinations of test parameters can be directly selected from the orthogonal experimental groups as the optimal combination of test parameters. These combinations cover different game scenarios and configurations, and also consider boundary cases, thus possessing high representativeness and testing efficiency.

[0084] In this embodiment, cluster analysis and orthogonal array design ensure that the combination of test parameters covers all possible scenarios and boundary conditions of the game, improving the comprehensiveness and coverage of the test. Orthogonal array design reduces testing costs by decreasing the number of experiments while maintaining the effectiveness and representativeness of the test.

[0085] S103, the artificial intelligence model is automatically tested according to the optimal test parameter combination to generate model test results, and the running status of the artificial intelligence model during the automated testing process is monitored by deploying a game data acquisition component.

[0086] In this embodiment, the test environment is configured according to the optimal combination of test parameters, including hardware resources, software version, network conditions, etc., and an automated testing framework and tools, such as Selenium and Appium, are deployed to execute test cases. The set of test cases generated based on the optimal combination of test parameters is loaded into the automated testing framework, ensuring that the test cases cover the main functions and boundary conditions of the model. The automated testing framework is then started, and the test cases are executed according to the preset test order. The test execution process is monitored, and test logs and exception information are recorded. After the test is completed, test data is collected, including test time, execution results (success / failure), performance metrics (such as response time, accuracy), etc. Statistical analysis of the test results is performed to generate a test report.

[0087] According to the test requirements, select appropriate game data collection components such as APM (Application Performance Management) tools, log collection systems, etc. Deploy the data collection components in the test environment and configure their monitoring indicators and collection frequencies to ensure that the data collection components can collect key data in the model running process in real time and accurately. Start the data collection components and begin monitoring the running state of the model in the automated testing process. Real-time view monitoring data, including CPU usage, memory usage, response time, error rate, and other key indicators, and automatically trigger an alarm when the monitoring data exceeds the preset threshold.

[0088] In this embodiment, automated testing can greatly reduce the workload of manual testing and improve testing efficiency. At the same time, test cases based on the optimal test parameter combination can more accurately evaluate model performance. Through automated testing and detailed test result analysis, potential problems and defects in the model can be found to ensure testing quality.

[0089] In some embodiments, in step S103, the automated testing of the artificial intelligence model according to the optimal test parameter combination generates model test results, specifically including:

[0090] Use JMeter to simulate a high-concurrency request test environment, and in the high-concurrency request test environment, perform automated testing of the artificial intelligence model through the optimal test parameter combination to obtain a first model test result;

[0091] Use Selenium to simulate a user interaction test environment, and in the user interaction test environment, perform automated testing of the artificial intelligence model through the optimal test parameter combination to obtain a second model test result.

[0092] In this embodiment, JMeter and Selenium are used to simulate a high-concurrency request test environment and a user interaction test environment, respectively, and automated testing is performed through the previously determined optimal test parameter combination.

[0093] Specifically, create a new test plan in JMeter and add a thread group, set the number of threads (i.e., the number of concurrent users) and the number of loops to simulate a high-concurrency request scenario. For example, set the number of threads to 1000 and the number of loops to 5, indicating that 1000 users are simulated to send 5 requests simultaneously. Add HTTP request defaults within the thread group, setting basic information such as server name or IP address, port number, etc. Add an HTTP request, enter the chatbot API interface URL, and set the request method (such as POST), request body (including the optimal test parameter combination, such as user questions, conversation context, etc.). Add listeners such as "aggregate reports" or "graph results" to view and analyze test results after testing is complete. Then start the test plan, and JMeter will simulate a large number of users sending requests to the chatbot simultaneously. After the test is complete, view the response time, throughput, error rate, and other performance indicators through the listener as the first model test result.

[0094] Install Selenium WebDriver (such as ChromeDriver) and Selenium libraries (such as Selenium Python bindings) and prepare a test script writing environment (such as a Python IDE). Use Selenium to write automated test scripts to simulate user login, open chat windows, input questions (use user questions in the optimal test parameter combination), receive and verify responses, and other interactive processes. Selenium's WebDriver API can be used to simulate user clicks, text input, element loading, and other operations. Then execute the test script, and Selenium will automatically open the browser to simulate user interactions with the chatbot. During testing, capture and verify the chatbot's response content, response time, and other information as the second model test result. Use assertions to check if the response meets expectations and record any failures.

[0095] In this embodiment, high-concurrency testing by JMeter can evaluate the performance of the chatbot under high load, including response time, throughput, error rate, and other key indicators. User interaction testing by Selenium can simulate real user scenarios to evaluate the chatbot's dialogue fluency, accuracy, and other user experience aspects.

[0096] S104, by analyzing the model test results, determining the performance of the artificial intelligence model under boundary conditions, generating a model test report, and optimizing the artificial intelligence model through the model test report.

[0097] In this embodiment, the data collected from the automated testing process includes model output and performance indicators under boundary conditions such as extreme inputs, resource limitations, etc. Identify which boundary conditions have a significant impact on model performance, analyze the output error rate, response time, resource consumption, and other key indicators of the model under these boundary conditions, and use statistical methods and machine learning algorithms such as clustering analysis and regression analysis to understand the relationship between boundary conditions and model performance. According to the analysis results, determine which boundary conditions the model performs poorly, i.e. problem areas. Analyze the causes of problem areas, such as uneven data distribution, insufficient feature extraction, model structure defects, etc.

[0098] Organize test data, analysis results, and problem areas into structured report content. The report includes test objectives, test methods, test environments, test results (including performance under boundary conditions), problem area analysis, improvement suggestions, and other parts. Visual tools such as charts and tables can be used to clearly display test results and analysis conclusions.

[0099] According to the improvement suggestions in the test report, develop a specific model tuning plan, which may include data cleaning and preprocessing, model structure optimization, parameter adjustment, and introduction of new features. Modify and adjust the model according to the tuning plan, and re-run the test cases in the same test environment, especially focusing on the boundary conditions of problem areas. Analyze the test results after tuning to verify whether the model's performance under boundary conditions has improved. If the tuning effect is not satisfactory, continue to adjust the tuning plan and repeat the tuning and testing process.

[0100] In this embodiment, by deeply analyzing the test results and performance under boundary conditions, problems in the model can be found and targeted tuning can be performed to improve the overall performance of the model. Boundary condition testing helps identify the model's performance under extreme conditions, and tuning can enhance the model's stability and robustness.

[0101] In some embodiments, in step S104, the model is tuned and retested based on the model test report, specifically including:

[0102] According to the model test report, determine the specific layer of the artificial intelligence model that needs to be tuned;

[0103] According to the specific layer, determine the original training data set and expand and improve it through a generative adversarial network to form a target training data set;

[0104] Based on the target training data set, perform transfer learning and model fine-tuning on the artificial intelligence model.

[0105] In this embodiment, the test reports generated by JMeter and Selenium are carefully analyzed, with particular attention paid to test cases that have long response times, high error rates, or poor user feedback. By comparing the inputs and outputs of these test cases and combining this information with knowledge of the model architecture, it is determined which specific layer or layers are causing performance bottlenecks. Based on the model architecture documentation and source code, the specific layers that need to be tuned are precisely located.

[0106] Collect and organize the original data sets used to train the artificial intelligence model, such as the game guide robot. These data sets should contain various types of questions and corresponding replies. Then design a GAN model suitable for the game guide robot data set, including a generator (Generator) and a discriminator (Discriminator). The generator is responsible for generating fake question-reply pairs, which should be as close to the distribution of the real data set as possible, and the discriminator is responsible for distinguishing between generated data and real data. Train the GAN model to enable the generator to generate question-reply pairs that are increasingly close to real data. Use the new data generated by the generator to expand the original training data set, forming the target training data set. These data can increase the model's generalization ability and robustness.

[0107] Set up a new training environment, including necessary hardware resources (such as GPUs), software frameworks (such as TensorFlow or PyTorch), and dependent libraries. Load the previously trained game guide robot, including its weights and architecture. Use the target training data set for transfer learning of the model, focusing on adjusting the specific layers that have been identified as needing tuning. A smaller learning rate and more iterations can be used to fine-tune these layers to avoid damaging the model's performance in other areas. At the same time, the weights of other parts of the model can be kept unchanged or fine-tuned to maintain the overall performance stability. Use a new test set (which can be different from the original test set) to validate and evaluate the fine-tuned model, compare the performance of the model before and after fine-tuning, and ensure that the tuning effect meets expectations.

[0108] In this embodiment, expanding the training data set using a generative adversarial network can increase the diversity of the model's training samples, thereby improving the model's generalization ability and robustness. By fine-tuning rather than retraining the entire model, the risk of overfitting due to insufficient data or uneven data distribution can be reduced. Transfer learning and fine-tuning techniques can accelerate the model tuning process and reduce the time and resources required to train the model from scratch.

[0109] In some embodiments, the method further comprises:

[0110] Obtain the operation logs of players in different game scenarios, analyze the operation logs through an LSTM model, and obtain player portrait data, including player behavior patterns and player preference features.

[0111] The game level pass rate, player retention time key indicators and player experience data are used as the reward function, the player portrait data is used as the state space, the interactive behavior of the artificial intelligence model is used as the action space, and the deep reinforcement learning algorithm is used to optimize and adjust the artificial intelligence model.

[0112] In this embodiment, the operation logs of the players in different game scenarios are collected through the game client or server, which include detailed operation records such as the clicks, movements, skill releases, and item uses of the players, as well as corresponding timestamps and game scenario identifiers. The collected operation logs are converted into structured data, including fields such as player ID, timestamp, operation type, operation object, and scenario ID. Key features such as operation frequency, operation sequence, skill usage preference, and item usage habit are extracted from the operation logs. The operation logs of the players are arranged in chronological order to form time series data, which is used as the input of the LSTM model.

[0113] The LSTM model is built using the Keras or TensorFlow library of Python. The model includes an input layer, multiple LSTM layers (for capturing long-term dependencies in time series), a fully connected layer (for feature extraction and classification / regression), and an output layer. Suitable hyperparameters are set, such as the number of units in the LSTM layer, learning rate, optimizer, etc. The preprocessed operation log data is used to train the LSTM model. The performance of the model is evaluated through cross-validation and other methods to ensure that the model can accurately capture the behavior patterns and preference features of the players. The trained LSTM model is used to analyze new operation logs and generate player portrait data, which includes the behavior patterns (such as operation habits, skill usage strategies, etc.) and preference features (such as favorite game types, character positioning, etc.) of the players.

[0114] A reasonable reward mechanism is designed to encourage the artificial intelligence model to adopt behaviors that can improve the player experience and game performance. The game level pass rate, player retention time key indicators, and player experience data (such as satisfaction scores, feedback opinions, etc.) are used as components of the reward function. The player portrait data is used as the state space, including the behavior patterns and preference features of the players. The interactive behavior of the artificial intelligence model is defined as the action space, including in-game decision-making, recommendations, guidance, etc. The deep reinforcement learning algorithm (such as DQN, A3C, etc.) is used to train the artificial intelligence model. Through interaction with the simulation environment or actual game environment, the strategy and behavior of the model are continuously optimized to maximize the expected value of the reward function.

[0115] In this embodiment, the operation log of the player is analyzed by the LSTM model to generate personalized player portrait data, providing the player with a game experience that better meets their interests and preferences. The artificial intelligence model can adjust the game difficulty, recommend content, and other aspects based on the player's real-time behavior, improving the player's satisfaction and retention rate. The deep reinforcement learning algorithm can continuously optimize the strategy and behavior of the artificial intelligence model, improving key indicators such as game level pass rate and player experience data. Through continuous training and optimization, the artificial intelligence model can gradually adapt to the needs and changes of different players, maintaining the competitiveness and appeal of the game. The artificial intelligence model can intelligently interact and guide players based on their behavior patterns and preference characteristics, enhancing the interactivity and interest of the game. Through real-time interaction and feedback loops with players, the artificial intelligence model can continuously learn and evolve, providing players with a more rich and engaging gaming experience.

[0116] Referring to Figure 2 , an embodiment of the present application provides an extreme test system 2 of an artificial intelligence model, which specifically comprises:

[0117] The first test module 201 is configured to construct a plurality of scene templates for different types of games, and generate a test case set by combining and transforming each scene template according to the boundary conditions input by the user;

[0118] The second test module 202 is configured to determine an optimal test parameter combination that satisfies the artificial intelligence model based on the test case set and according to a preset optimization objective function by using a heuristic search algorithm;

[0119] The third test module 203 is configured to perform automated testing on the artificial intelligence model according to the optimal test parameter combination, generate a model test result, and monitor the running state of the artificial intelligence model during the automated testing by deploying a game data acquisition component;

[0120] The fourth test module 204 is configured to determine the performance of the artificial intelligence model under the boundary conditions by analyzing the model test result, generate a model test report, and perform retesting and optimization of the artificial intelligence model based on the model test report.

[0121] It can be understood that the contents in the extreme test method embodiment of the artificial intelligence model as shown in Figure 1 are applicable to the extreme test system embodiment of the artificial intelligence model, the functions specifically realized by the extreme test system embodiment of the artificial intelligence model are the same as those of the extreme test method embodiment of the artificial intelligence model as shown in Figure 1 , and the beneficial effects achieved are the same as those of the extreme test method embodiment of the artificial intelligence model as shown in Figure 1 .

[0122] It should be noted that the information interaction between the above systems, the execution process and the like, since based on the same concept as the method embodiments of the present application, the specific functions and the technical effects brought about, can be specifically referred to the method embodiments part, and will not be repeated here.

[0123] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific name of each functional unit and module is only for easy distinction, and does not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0124] With reference to Figure 3 The embodiment of the present application also provides a computer device 3, comprising a memory 302 and a processor 301 and a computer program 303 stored in the memory 302, when the computer program 303 is executed on the processor 301, the limit test method of the artificial intelligence model is realized as any one of the above methods.

[0125] The computer device 3 can be a desktop computer, a notebook computer, a palm computer and a cloud server and the like. The computer device 3 can include, but is not limited to, a processor 301, a memory 302. Those skilled in the art can understand that, Figure 3 It is only an example of the computer device 3, and does not constitute a limitation on the computer device 3, and can include more or fewer components than the illustration, or combine certain components, or different components, for example, it can also include input and output devices, network access devices and the like.

[0126] The processor 301 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0127] The memory 302 can be an internal storage unit of the computer device 3 in some embodiments, for example, a hard disk or a memory of the computer device 3. The memory 302 can also be an external storage device of the computer device 3 in other embodiments, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 302 can include both the internal storage unit and the external storage device of the computer device 3. The memory 302 is used to store an operating system, an application program, a boot loader, data, and other programs, for example, program codes of the computer program, etc. The memory 302 can also be used to temporarily store data that has been output or is to be output.

[0128] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is run by a processor, the limit test method of the artificial intelligence model is implemented.

[0129] In this embodiment, the integrated unit, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the computer program for instructing the relevant hardware to complete all or part of the processes in the above-described embodiment methods can be stored in a computer readable storage medium. The computer program can be executed by a processor to implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer readable medium at least includes any entity or device capable of carrying the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium can not be an electrical carrier signal and a telecommunication signal.

[0130] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0131] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solutions. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0132] In the embodiments disclosed in the present application, it should be understood that the disclosed apparatus / terminal equipment and methods can be implemented in other ways. For example, the apparatus / terminal equipment embodiments described above are only schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0133] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may also be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

Claims

1. A method for limit testing of an artificial intelligence model, characterized in that, The method specifically includes: Multiple scenario templates are built for different types of games. Based on the boundary conditions input by the user, a set of test cases is generated by combining and transforming the various scenario templates. Specifically, this includes: Multiple scene templates are built for different types of games to form a scene template library. The scene template includes several key elements and several constraints. Obtain boundary condition data input by the user, match the boundary condition data with each scene template in the scene template library, and extract several first scene templates; A depth-first search algorithm is used to traverse the combination schemes between key elements and constraints of each first scenario template to generate an initial test case set. A greedy algorithm is used to analyze the coverage and differences of each initial test case in the initial test case set to obtain the target test case set; Based on the aforementioned test case set, and according to a preset optimization objective function, a heuristic search algorithm is used to determine the optimal combination of test parameters that satisfies the artificial intelligence model, specifically including: Acquire game project requirement data and specific requirement data of the artificial intelligence model, and construct an optimization objective function based on the game project requirement data and the specific requirement data. The optimization objective function includes model inference speed, memory usage performance index and resource consumption index. The test case set is divided into multiple parameter subspaces, each parameter subspace representing a similar game scenario; Within each parameter subspace, a simulated annealing algorithm is used for search optimization to obtain multiple combinations of test parameters; By evaluating the model inference speed, memory usage, and resource consumption indicators corresponding to each of the test parameter combinations, it is determined whether the requirements of the optimization objective function are met. If the requirements are met, the test parameter combination is added to the optimal solution set, and the optimal test parameter combination is obtained based on the optimal solution set. The AI ​​model is automatically tested based on the optimal combination of test parameters to generate model test results. At the same time, the running status of the AI ​​model during the automated testing process is monitored by deploying a game data acquisition component. By analyzing the test results of the model, the performance of the artificial intelligence model under boundary conditions is determined, a model test report is generated, and the artificial intelligence model is optimized and retested based on the model test report.

2. The method according to claim 1, characterized in that, The step of obtaining the optimal test parameter combination based on the optimal solution set specifically includes: The optimal solution set is divided into K clusters using the K-means clustering algorithm; Cluster centers are selected from each cluster as representative sample points to form a sample subset; For each of the sample points, boundary value analysis is used to generate corresponding boundary value test data for the game state parameters and event triggering conditions; Specifically, for numeric parameters, upper and lower limits are selected as boundary values; for boolean parameters, true and false are selected as boundary values; and for enumeration parameters, all enumeration values ​​are selected as boundary values. Using the orthogonal array design method, an orthogonal experimental group is constructed from the sample subset and the boundary value test data, where each factor of the orthogonal experimental group represents a game state parameter or an event triggering condition. Select test parameter combinations that satisfy orthogonality from the orthogonal experimental groups, and take the test parameter combinations that satisfy orthogonality as the optimal test parameter combinations.

3. The method according to claim 1, characterized in that, The step of automatically testing the artificial intelligence model based on the optimal combination of test parameters and generating model test results specifically includes: Using JMeter to simulate a high-concurrency request test environment, the artificial intelligence model is automatically tested in the high-concurrency request test environment using the optimal test parameter combination to obtain the first model test result; Using Selenium to simulate a user interaction testing environment, the artificial intelligence model is automatically tested in the user interaction testing environment using the optimal combination of test parameters to obtain the test results of the second model.

4. The method according to claim 1, characterized in that, The process of fine-tuning and retesting the artificial intelligence model using the model test report specifically includes: Based on the model test report, the specific layers of the artificial intelligence model that need to be optimized are determined; The original training dataset is determined based on the specific layer, and the original training dataset is extended and improved by a generative adversarial network to form the target training dataset. Based on the target training dataset, the artificial intelligence model is subjected to transfer learning and model fine-tuning.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: The system acquires player operation logs in different game scenarios, analyzes the operation logs using an LSTM model to obtain player profile data, which includes player behavior patterns and player preference characteristics. Using key metrics such as game level completion rate, player retention time, and player experience data as reward functions, player profile data as state space, and the interactive behavior of the artificial intelligence model as action space, the artificial intelligence model is personalized and optimized through deep reinforcement learning algorithms.

6. A limit testing system for an artificial intelligence model, characterized in that, The system specifically includes: The first testing module is used to build multiple scene templates for different types of games. Based on the boundary conditions input by the user, it generates a set of test cases by combining and transforming various scene templates. Specifically, it includes: Multiple scene templates are built for different types of games to form a scene template library. The scene template includes several key elements and several constraints. Obtain boundary condition data input by the user, match the boundary condition data with each scene template in the scene template library, and extract several first scene templates; A depth-first search algorithm is used to traverse the combination schemes between key elements and constraints of each first scenario template to generate an initial test case set. A greedy algorithm is used to analyze the coverage and differences of each initial test case in the initial test case set to obtain the target test case set; The second testing module is used to determine the optimal combination of test parameters that satisfies the artificial intelligence model based on the set of test cases and according to a preset optimization objective function, using a heuristic search algorithm. Specifically, it includes: Acquire game project requirement data and specific requirement data of the artificial intelligence model, and construct an optimization objective function based on the game project requirement data and the specific requirement data. The optimization objective function includes model inference speed, memory usage performance index and resource consumption index. The test case set is divided into multiple parameter subspaces, each parameter subspace representing a similar game scenario; Within each parameter subspace, a simulated annealing algorithm is used for search optimization to obtain multiple combinations of test parameters; By evaluating the model inference speed, memory usage, and resource consumption indicators corresponding to each of the test parameter combinations, it is determined whether the requirements of the optimization objective function are met. If the requirements are met, the test parameter combination is added to the optimal solution set, and the optimal test parameter combination is obtained based on the optimal solution set. The third testing module is used to automatically test the artificial intelligence model according to the optimal combination of test parameters, generate model test results, and monitor the running status of the artificial intelligence model during the automated testing process by deploying game data acquisition components. The fourth testing module is used to analyze the model test results, determine the performance of the artificial intelligence model under boundary conditions, generate a model test report, and optimize and retest the artificial intelligence model based on the model test report.

7. A computer device, characterized in that, include: A memory and a processor, and a computer program stored in the memory, which, when executed on the processor, implements the extreme testing method for the artificial intelligence model as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the extreme testing method for the artificial intelligence model as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intelligent driving vehicle autonomous ability test method based on scene confrontation

    CN113158560A

  • Automatic test method for reducing test technical threshold of intelligent terminal system

    CN117931620A