A test case generation method, device and equipment and storage medium

CN122817097APending Publication Date: 2026-09-25SANGFOR TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611134594.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-25

Smart Images

  • Figure CN122817097A_ABST
    Figure CN122817097A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a test case generation method and device, equipment and a storage medium, relating to the technical field of software test automation, which analyzes multiple source files such as scheme design documents and functional brain maps, constructs structured information recognizable by a model, effectively solves the problem of incomplete and unclear structure of single input source information, and guarantees comprehensive and accurate understanding of requirements. A first large language model is used to perform semantic analysis on the structured information to generate a test point list matching a business scenario, replacing manual process analysis and improving extraction integrity and systematicness. A second large language model is used to perform reasoning analysis on the test points to batch generate corresponding test cases in a standard format, covering multiple types of test scenarios and improving test coverage and test case quality. A double large language model collaborative architecture is adopted to automatically generate test cases with high coverage and high quality by combining text and structure dual guidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software test automation technology, and in particular to a method, apparatus, device and storage medium for generating test cases. Background Technology

[0002] Software testing is a critical activity in the software development lifecycle, ensuring that software products meet customer expectations and design requirements through systematic verification and validation. In current software testing processes, testers typically need to manually analyze requirements and write all possible test scenarios and test cases after developers provide detailed design solutions. This manual writing process is not only time-consuming but also prone to overlooking key logic due to oversights or insufficient personal experience, leading to incomplete test coverage. Traditional test documentation also suffers from being bloated and inefficient, lacking flexibility to adapt to changes and often reducing the agility of testing efforts.

[0003] To improve the efficiency and consistency of test design, the industry has begun to explore methods for automatically generating test cases using artificial intelligence technology. However, most existing research on automatically generating test cases from large language models focuses on generating test cases from code or improving code quality through test cases during the code generation process. There is still a lack of sufficient exploration into enabling large language models (LLMs) to independently generate high-quality business scenario test cases based on business needs.

[0004] It is evident that how to automatically generate high-coverage, high-quality test cases is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this application is to provide a method, apparatus, device, and storage medium for generating test cases, which can automatically generate high-coverage, high-quality test cases.

[0006] This application provides a method for generating test cases, including: The multi-source files are parsed to construct structured information that the model can recognize; the multi-source files include solution design documents and functional mind maps corresponding to business scenarios; The first type of large language model is used to perform semantic parsing of structured information to generate a list of test points that match the business scenario; The second type of large language model is used to perform reasoning analysis on the test point list to generate a test case set that conforms to the test case format; wherein, the test case set includes the test cases corresponding to each test point.

[0007] On the one hand, multi-source files are parsed to construct structured information that the model can recognize, including: The chapters and items in the scheme design document are broken down to determine the paragraph content of each functional point; Extract the hierarchical relationships of each functional point from the functional mind map; The hierarchical relationship between each functional point is integrated with the paragraph content of each functional point to establish a function-sub-function mapping table and its corresponding set of requirement specification fragments; wherein, the set of requirement specification fragments includes the text content corresponding to each functional point in the function-sub-function mapping table.

[0008] On the one hand, the hierarchical relationships of each functional point are integrated with the paragraph content of each functional point to establish a function-sub-function mapping table and its corresponding set of requirement specification fragments, including: Add the hierarchical relationship of each function point to the corresponding paragraph content to determine the sub-function list that matches each paragraph content; Based on the paragraph content of each function point, generate test check items for each function point; The sub-function list matching the content of each paragraph and the test check items of each function point are integrated to obtain the function-sub-function mapping table and its corresponding set of requirement specification fragments.

[0009] On the one hand, the first type of large language model is used to perform semantic parsing of structured information to generate a list of test points that match the business scenario, including: The first type of large language model is used to perform contextual analysis on structured information in order to deduce test points one by one and identify implicit test requirements; Based on the test points and implicit test requirements, a test point list is generated; the test point list includes the test scenarios that need to be verified for each function and a summary of the corresponding expected behavior.

[0010] On the one hand, the second type of large language model is used to perform reasoning analysis on the test point list to generate a set of test cases that conform to the test case format, including: The set test case format and test point list, including the test scenarios to be verified for each function point and the corresponding expected behavior summary, are input into the second type of large language model to output test cases that match each test scenario; the test case format includes test case number, test case name, preconditions, test steps and expected results.

[0011] On the one hand, after using the second type of large language model to perform reasoning analysis on the test point list and generate a set of test cases that conform to the test case format, it also includes: Verify each test case in the test case set and obtain the verification results; If the validation results include target test cases that fail validation, display the validation results of the target test cases so that users can adjust the test cases based on the validation results of the target test cases; After adjusting the test cases, output the adjusted test case set.

[0012] On the one hand, each test case in the test case set is validated to obtain the validation results, including: Determine whether each test case in the test case set conforms to the template format requirements; If the first test case does not meet the template format requirements, output the validation result that the first test case failed the format validation. Determine whether the field content of each test case in the test case set is complete; If the field content of the second test case is incomplete, output the validation result that the content validation of the second test case failed; Determine whether the test case logic of each test case in the test case set meets the consistency requirements; If the logic of the third test case does not meet the consistency requirements, output the verification result that the logic consistency check of the third test case fails; Determine whether there is redundancy or conflict among the test cases contained in the test case set; If there is redundancy or conflict between the fourth and fifth test cases, output the verification results of the redundancy or conflict between the fourth and fifth test cases.

[0013] This application also provides a test case generation apparatus, including a construction unit, a first generation unit, and a second generation unit; The building unit is used to parse multiple source files to construct structured information that the model can recognize; among them, multiple source files include solution design documents and functional mind maps corresponding to business scenarios; The first generation unit is used to perform semantic parsing of structured information using the first type of large language model to generate a test point list that matches the business scenario. The second generation unit is used to perform reasoning analysis on the test point list using the second type of large language model to generate a set of test cases that conforms to the test case format; wherein, the set of test cases includes the test cases corresponding to each test point.

[0014] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of the test case generation method described above when executing the computer program.

[0015] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described test case generation methods.

[0016] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described test case generation methods.

[0017] As can be seen from the above technical solutions, parsing multiple source files constructs structured information that the model can recognize. These multiple source files include solution design documents and functional mind maps corresponding to business scenarios, effectively solving the problems of incomplete and unclear information from a single input source, ensuring the comprehensiveness and accuracy of requirement understanding. The first type of large language model is used to semantically parse the structured information, generating a test point list matching the business scenario. This replaces the tedious process of manually compiling test points, significantly reducing manual analysis costs and avoiding omissions due to differences in individual experience, thus improving the completeness and systematic nature of test point extraction. The second type of large language model is used to reason and analyze the test point list, generating a set of test cases conforming to the test case format. This set includes test cases corresponding to each test point. This ensures that each test point generates corresponding adapted test cases, achieving automated batch production of test cases, significantly shortening the test case development cycle, and ensuring that the generated test cases cover multiple test scenarios, effectively improving test coverage and test case quality. In this technical solution, the illusion phenomenon when generating content by large language models is greatly reduced through dual guidance of text and structure. The dual large language model collaborative architecture is adopted. The first type of large language model solves the problem of omissions in understanding long documents, while the second type of large language model solves the problem of test case generation accuracy. This division of labor and collaboration mode has more advantages in engineering implementation than single model-driven or algorithm iteration, and can complement each other's capabilities, significantly improving the coherence and accuracy of test case logic. Attached Figure Description

[0018] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a method for generating test cases provided in an embodiment of this application; Figure 2 A flowchart illustrating a method for constructing model-recognizable structured information, as provided in this application embodiment; Figure 3This is a schematic diagram of a test case generation device provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0021] The terms "comprising" and "having," and any variations thereof, in the specification and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.

[0022] Software testing is a critical activity in the software development lifecycle, ensuring that software products meet customer expectations and design requirements through systematic verification and validation. To improve the efficiency and consistency of test design, the industry has begun exploring methods for automatically generating test cases using artificial intelligence technologies. Large language models (such as the GPT series) have become powerful tools for automated test case generation due to their strong text understanding and generation capabilities. LLMs can automatically generate test cases from sets of requirements specification fragments or user stories, achieving more comprehensive test coverage than manual methods and reducing human error.

[0023] While large language models show promise for generating test cases, current research and applications remain limited. Most current research on automatically generating test cases using large models focuses on generating unit tests from code or improving code quality during the code generation process using test cases. When test scenarios are complex, even state-of-the-art LLMs often struggle to generate completely correct test cases due to limitations in their reasoning and computational capabilities. Therefore, how to more effectively utilize large language models to automatically generate high-coverage, high-quality business logic test cases remains a technical problem to be solved.

[0024] Compared to cumbersome linear documents, mind maps are lean documents that can intuitively display system functions and testing ideas through keywords and hierarchical structures, and their creation and updates are much faster.

[0025] Therefore, this application provides a test case generation method, apparatus, device, and storage medium that simultaneously parses natural language-described solution design documents and structured functional mind maps. By combining text with graphical mind map content through LLM, the understanding of system functions is deepened, ensuring that no requirement details provided by different representations are overlooked. The mind map provides an intuitive hierarchical structure, upon which LLM can systematically traverse functional nodes to generate test cases covering various test scenarios, significantly improving test coverage.

[0026] Compared to manual design, this solution more systematically identifies boundary values ​​and extreme cases, providing a comprehensive set of test scenarios. The generated test cases strictly adhere to the defined test case format, ensuring consistent document style. This solution transforms test case design from manual writing to automated generation, significantly reducing manpower and accelerating the test preparation cycle. The automated process reduces repetitive work, allowing testers to focus more on test case review and test execution, thereby shortening the overall product delivery cycle. Furthermore, this solution easily integrates with existing test management tools; for example, generated test cases can be directly imported into test management systems such as JIRA and TestLink via API, integrating into existing development processes. In an agile development environment, it can quickly regenerate updated test cases after requirement or design changes, maintaining consistency between test cases and the latest requirements.

[0027] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] Next, a method for generating test cases provided in the embodiments of this application will be described in detail. Figure 1 A flowchart illustrating a test case generation method provided in this application embodiment, the method comprising: S101: Parse multiple source files to construct structured information that the model can recognize.

[0029] In this embodiment of the application, the multi-source files may include solution design documents and functional mind maps corresponding to the business scenario.

[0030] The solution design document is a detailed design document provided to developers, including descriptions of key functions and process specifications. It can be in Markdown format.

[0031] Functional mind maps can be drawn by testers using XMind tools and exported as a text outline-structured visual functional summary file.

[0032] Structured information is the conversion of unstructured or semi-structured scheme design documents and functional mind maps into standardized data that can be read and calculated by the model, including but not limited to function-subfunction mapping tables and their corresponding sets of requirement specification fragments.

[0033] In the implementation, the solution design document in Markdown format can be read first, and then broken down by chapter and item to extract the functional titles and corresponding descriptive paragraphs of each functional module. Next, the text outline of the functional mind map is parsed, and the mind map node hierarchy is converted into tree-structured data such as JSON. The content of each node in this tree-structured data includes the functional point name and a brief description. Finally, the decomposed solution design document and the converted mind map information are merged, the hierarchical relationship of the mind map is appended to the corresponding paragraphs, the sub-functional list corresponding to a certain description is clarified, and a test checklist is generated for each functional point, thus outputting a structured function-sub-functionality mapping table and its corresponding set of requirement specification fragments.

[0034] By performing multi-source analysis on solution design documents and functional mind maps and constructing structured information, we can unify the data representation of different input formats, transforming scattered and unstructured requirement information into standardized data with clear hierarchy and complete connections. This provides an unambiguous and highly complete input foundation for subsequent large language model processing, effectively avoiding the problems of requirement misunderstanding and information gaps caused by a single information source. At the same time, it improves the efficiency and accuracy of the large language model in parsing requirements, laying a reliable data foundation for subsequent test point extraction and test case generation.

[0035] S102: Use the first type of large language model to perform semantic parsing on structured information and generate a list of test points that match the business scenario.

[0036] In this embodiment, a first-class large language model can be used to perform contextual analysis on structured information to deduce test points one by one and identify implicit test requirements. Based on the test points and implicit test requirements, a list of test points is generated.

[0037] Implicit test requirements refer to constraints and exception handling logic that are not explicitly stated in the design document but are necessarily included in the business logic, such as amount limits, network interruption, permission verification, and timeout handling.

[0038] The test checklist can include the test scenarios that need to be verified for each function and a summary of the expected behavior.

[0039] The test scenarios to be verified can include normal scenarios, abnormal scenarios, and boundary conditions. Boundary conditions usually refer to the extreme values ​​of input parameters (such as maximum length, minimum value, null value, etc.). For example, "amount limit" is a typical boundary condition. The expected total daily transfer limit is 20,000. Transfers of 20,000 or less will be successful, while transfers exceeding 20,000 will fail.

[0040] The expected behavior is used to characterize what the test results should be.

[0041] The first type of large language model is a large language model with a long context window, strong document comprehensive understanding ability and cross-document reasoning ability. In the embodiments of this application, Claude2 can be selected as the first type of large language model. It can support input of about 100K words and can parse the complete design scheme and mind map structure as a whole.

[0042] In practice, structured information can be input into the Claude2 model, and prompts can be configured to instruct the model to extract all functional points, test targets, and verification scenarios.

[0043] Test objectives are used to characterize what needs to be tested. Test objectives can be the verification standards that the system should meet under a specific function or the business rules that need to be covered. For example, for a money transfer function, the test objective might be "verify the limit control for large money transfers".

[0044] The Claude2 model reads the complete design scheme in context and synthesizes the mind map structure, reasoning out possible test points one by one like a human test analyst.

[0045] Because Claude can reason across long documents, it can identify implicit business rules and exceptions in the design. For example, in the description of the "transfer function" in the design document, Claude2 identifies implicit requirements involving amount limits and network interruption handling, and, with the help of mind map prompts, focuses on the fact that the "transfer function" belongs to the "account management" module, thus ensuring the comprehensiveness and hierarchical relevance of the extracted test points. Finally, the module outputs a list of test points.

[0046] Each function in the function-subfunction mapping table can be viewed as a function point. A function point typically corresponds to multiple test points. The first type of large language model identifies multiple test scenarios (such as normal and abnormal) that need to be verified for each function point. These test scenarios constitute the test points for that function point.

[0047] The sub-functions in the mind map are the primary source of test scenarios. The Type I Large Language Model (TLM) identifies implicit business rules and exceptions in the design, meaning that even exceptions not depicted in the mind map will be generated by the TLM based on the requirements document. Therefore, a test scenario is the union of the sub-functions and the implicit test requirements inferred from the TLM.

[0048] In practical applications, users can choose the appropriate model based on their needs and data privacy requirements: for example, Tongyi Qianwen can be prioritized for Chinese systems, GPT-4 can be selected for those requiring English comprehension or code analysis, and Claude 2 can be used for analyzing extremely long documents. Multiple models can also be combined to leverage their respective strengths. Through proper configuration, the models can be fully utilized to improve performance, ensuring accuracy while enhancing the efficiency of test case generation.

[0049] Employing a first-class large language model with long context and global understanding capabilities, it performs deep semantic parsing of structured information. This simulates the analytical logic of a senior test analyst, systematically mining explicit and implicit test requirements within the entire document. Combining brain-layer hierarchical relationships, it achieves full-featured functional point extraction, automatically generating a comprehensive, hierarchical, and complete list of test points. This significantly reduces the cost of manual sorting, avoids missing test points due to insufficient personal experience or oversight, and significantly improves the completeness, systematicness, and accuracy of test point extraction.

[0050] S103: Use the second type of large language model to perform reasoning analysis on the test point list and generate a set of test cases that conform to the test case format.

[0051] The test case set includes the test cases corresponding to each test point.

[0052] The set test case format and test point list, including the test scenarios to be verified for each function point and the corresponding expected behavior summary, are input into the second type of large language model to output test cases that match each test scenario; the test case format includes test case number, test case name, preconditions, test steps and expected results.

[0053] The second type of large language model is a large language model with high-precision text generation, strong logical expression and strict format compliance. In the embodiments of this application, GPT-4 can be selected as the second type of large language model. GPT-4 can generate standardized, detailed and highly consistent test cases.

[0054] Test case format can be a predefined standard template, including test case number, test case name, preconditions, test steps, and expected results.

[0055] The test case format can also be adjusted according to the company's testing specifications.

[0056] The test case generation feature offers flexible customization capabilities, allowing templates and output formats to be adjusted according to the specifications of different testing teams. For example, fields can be added or removed, format layouts or language styles can be configured to align with internal enterprise test case management standards.

[0057] In practical implementation, a dedicated prompt word template can be built, and the set test case format and test point list can be input into the GPT-4 model. The GPT-4 model generates test cases that match each test scenario according to the format requirements, ensuring that the test case steps are clear, the expected results are clear, the fields are complete, and the numbers are unique. Finally, a standardized test case set covering all test scenarios is output.

[0058] The following is a sample Prompt input to the second type of large language model: Please generate test cases based on the following test criteria. Each test case should include: test case number, test case name, preconditions, test steps, and expected results. Test criteria: 1. User Login - Normal: Successfully log in using a valid username and correct password. 2. User Login - Incorrect Password: Enter a valid username but the password is incorrect; the system displays an error message indicating that login is not allowed. (Further details omitted).

[0059] Once GPT-4 receives the above prompts and key points list, it will generate corresponding test cases one by one. For example, for the "User Login - Normal" scenario, the test case generated by GPT-4 might be: Test Case Number: TC001. Test Case Name: User successfully logs in using correct credentials. Preconditions: The system server is running normally; the user has registered in the system and has a valid account. Test Steps: 1. Open the login page; 2. Enter a valid username (e.g., User123) in the username field; 3. Enter the correct password for the user in the password field; 4. Click the "Login" button. Expected Results: 1. The system verifies the username and password; 2. After successful verification, the user successfully logs in, and the page redirects to the homepage; the user's username is displayed in the upper right corner of the page, indicating a normal login status.

[0060] For example, in the "user login - incorrect password" scenario, GPT-4 generates test steps and expected results to verify the system's response when the password is incorrect. The test case text generated by GPT-4, after being formatted, is almost indistinguishable from professionally written test cases.

[0061] GPT-4 can provide highly detailed and context-dependent test case descriptions, and its powerful reasoning capabilities are suitable for complex application scenarios requiring meticulous testing. Through model generation, a complete software test case document can be obtained within minutes, covering both positive and negative test scenarios for each functional module involved in the design.

[0062] This application leverages a second-class large language model with high-precision text generation and strong formatting compliance capabilities to automatically generate test cases based on a test point list and a standard template. This enables batch and rapid production of test cases, ensuring that each test point corresponds to a standardized, complete, and logically rigorous test case, covering normal, abnormal, and boundary scenarios. A unified test case writing style and field specifications significantly shorten the test case development cycle, improve test coverage and test case quality, reduce manual writing costs, facilitate test case management and tool integration, and effectively improve the efficiency, standardization, and agility of software testing.

[0063] As can be seen from the above technical solutions, parsing multiple source files constructs structured information that the model can recognize. These multiple source files include solution design documents and functional mind maps corresponding to business scenarios, effectively solving the problems of incomplete and unclear information from a single input source, ensuring the comprehensiveness and accuracy of requirement understanding. The first type of large language model is used to semantically parse the structured information, generating a test point list matching the business scenario. This replaces the tedious process of manually compiling test points, significantly reducing manual analysis costs and avoiding omissions due to differences in individual experience, thus improving the completeness and systematic nature of test point extraction. The second type of large language model is used to reason and analyze the test point list, generating a set of test cases conforming to the test case format. This set includes test cases corresponding to each test point. This ensures that each test point generates corresponding adapted test cases, achieving automated batch production of test cases, significantly shortening the test case development cycle, and ensuring that the generated test cases cover multiple test scenarios, effectively improving test coverage and test case quality. In this technical solution, the illusion phenomenon when generating content by large language models is greatly reduced through dual guidance of text and structure. The dual large language model collaborative architecture is adopted. The first type of large language model solves the problem of omissions in understanding long documents, while the second type of large language model solves the problem of test case generation accuracy. This division of labor and collaboration mode has more advantages in engineering implementation than single model-driven or algorithm iteration, and can complement each other's capabilities, significantly improving the coherence and accuracy of test case logic.

[0064] Figure 2 A flowchart illustrating a method for constructing model-recognizable structured information, as provided in this application embodiment, is included in the following method: S201: Decompose the chapters and items contained in the scheme design document to determine the paragraph content of each functional point.

[0065] Chapter and item breakdown can refer to dividing the entire document into independent content units corresponding to each functional point according to the document's title hierarchy, paragraph identifiers, item numbers, and other structures.

[0066] The paragraph content of a function point refers to the descriptive text that directly corresponds to a certain function. It may include the implementation logic of the function, usage scenarios, input and output requirements, and related business rules.

[0067] In the specific implementation, the design document in Markdown format can be read, and the document can be structurally split according to the chapter title, level number, and segment identifier. The descriptive paragraphs corresponding to each function point such as login, registration, data query, data modification, deletion, and export can be extracted one by one. The requirements, implementation methods, and constraints of each function point are clarified, forming a set of plain text content that corresponds one-to-one with each function point.

[0068] By breaking down the solution design document into chapters and items and determining the paragraph content of each functional point, the lengthy and holistic requirements document can be divided into text units with appropriate granularity and clear functions. This makes the requirements information of each functional point independent, clear, and accurately identifiable by the model, avoiding misunderstandings caused by mixed document content. It also provides a clear and accurate text foundation for subsequent mind map information fusion and test point extraction.

[0069] S202: Extract the hierarchical relationships of each functional point from the functional mind map.

[0070] The hierarchical relationship of functional points refers to the parent-child relationship, subordinate relationship, branch structure, etc. among various functions, such as the hierarchical affiliation and arrangement order between main functions, sub-functions, and lower-level functional points.

[0071] In practical implementation, the mind map text outline structure exported by the XMind tool can be read, the mind map nodes can be parsed according to parent-child level and hierarchical depth, the mind map can be converted into tree structure data such as JSON, the root node and the subordinate relationship between functional nodes at all levels can be extracted, forming complete functional hierarchical topology information, and clarifying the position and relationship of each functional point in the overall system.

[0072] By extracting the hierarchical relationships of functional points from functional mind maps, the graphical and visual functional structure can be transformed into machine-recognizable structured relational data, preserving the hierarchical logic and branch coverage relationships of requirements, making up for the lack of a structured framework in plain text design documents, and providing clear structural guidance for subsequent multi-source information fusion.

[0073] S203: Integrate the hierarchical relationship of each functional point with the paragraph content of each functional point to establish a function-sub-function mapping table and its corresponding set of requirement specification fragments.

[0074] The set of requirement specification fragments may include the text content corresponding to each function point in the function-subfunction mapping table.

[0075] The function-subfunction mapping table is a structured data used to record the relationship between the main function and the corresponding subfunction, which can clearly show the hierarchical structure and coverage of functions.

[0076] The set of requirement specification fragments includes the text content corresponding to each function point in the function-subfunction mapping table, which is used to fully express the business meaning, implementation logic and constraint rules of the function point.

[0077] In this embodiment, the hierarchical relationship of each functional point can be added to the corresponding paragraph content to determine the sub-functional list that matches each paragraph content; based on the paragraph content of each functional point, test check items are generated for each functional point; the sub-functional list that matches each paragraph content and the test check items of each functional point are integrated to obtain a function-sub-functional mapping table and its corresponding set of requirement specification fragments.

[0078] A sub-function list refers to a collection of subordinate functions that belong to a certain main function.

[0079] Test check items are preliminary test clues generated for each function point in the mapping table.

[0080] In the specific implementation, the extracted hierarchical relationships of each functional point are attached to the paragraph content that matches each functional point, and the corresponding list of subordinate sub-functions is clearly matched for each main functional paragraph; at the same time, according to the paragraph content of each functional point, corresponding test check items are generated for each functional point; finally, the sub-function list matched by each paragraph and the test check items of each functional point are integrated to form a function-sub-function mapping table, and a set of requirement specification fragments matching the mapping table are generated simultaneously.

[0081] A collection of requirement specification fragments contains multiple fragments that need to be specified. A requirement specification fragment can be a section title corresponding to a specific functional point and the descriptive paragraphs associated with it, extracted from a Markdown document. For example, if a design document has a section on "transfer functionality" that describes amount limits and transaction fee logic, then all the Markdown text in that section constitutes a "requirement specification fragment."

[0082] For example, if the mind map shows that the "User Login" function has sub-nodes such as "Enter Correct Information to Login," "Enter Incorrect Password," and "Password Reset Process," then the module will list these sub-functional points after the "User Login" design description, providing clues for the subsequent processing of the second type of large language model. After this processing, a structured function-sub-function mapping table and a corresponding set of requirement specification fragments are output.

[0083] By bidirectionally integrating functional hierarchy relationships with functional paragraph content and establishing a function-subfunction mapping table and a corresponding set of requirement description fragments, it is possible to achieve an organic combination of textual requirements and structural requirements. This ensures that the requirement information retains both complete semantic descriptions and clear hierarchical relationships, comprehensively covering explicit requirements and structurally implicit requirements, avoiding omissions or incomplete understanding of requirements. At the same time, it provides standardized and highly available structured input for subsequent semantic parsing and test case generation, improving the accuracy and coverage of subsequent processing flows.

[0084] In this embodiment, by decomposing the chapters and entries of the design document to determine the paragraph content of each functional point, the lengthy and unstructured design document can be broken down into text units with clear functions and boundaries, providing a precise and unambiguous text foundation for subsequent information fusion and model processing. By extracting the hierarchical relationships of each functional point from the functional mind map, the graphical functional structure can be transformed into machine-recognizable hierarchical relationships, compensating for the lack of a structured framework in pure text requirements and ensuring complete functional coverage and logical clarity. By deeply integrating hierarchical relationships with paragraph content, a function-sub-function mapping table and a set of supporting requirement description fragments can be established, enabling the standardized integration of multi-source requirement information. This organically combines the hierarchical structure of the mind map with the semantic content of the design document, effectively avoiding the problems of requirement omissions, misunderstandings, and incomplete coverage caused by a single information source. It significantly improves the completeness, accuracy, and usability of structured information, providing stable and reliable data support for subsequent semantic parsing, test point extraction, and automatic test case generation for large language models. At the same time, it simplifies the test preparation process, reduces manual sorting costs, and improves the overall test design efficiency.

[0085] After using the second type of large language model to perform reasoning analysis on the test point list and generate a test case set that conforms to the test case format, the test cases contained in the test case set can be verified to obtain the verification results.

[0086] Validation refers to the automated checking of the format, content, logic, and relevance of test cases according to preset rules. Validation results can include whether a test case passed validation and the reasons for failure.

[0087] In this embodiment, it can be determined whether each test case in the test case set conforms to the template format requirements. These template format requirements may include test case numbering rules, field order, statement specifications, uniqueness of numbers, naming conventions, etc. If the first test case does not conform to the template format requirements, a validation result indicating that the first test case failed the format validation is output.

[0088] By setting template format requirements, we can ensure that all test cases contain necessary fields, have unique and logically ordered numbering, and that the statements of steps and expected results conform to the template requirements.

[0089] Determine if the fields of each test case in the test case set are complete. These fields can include required fields such as test case number, test case name, preconditions, test steps, and expected results. If the fields of the second test case are incomplete, output a validation result indicating that the second test case failed.

[0090] Determine whether the logic of each test case in the test case set meets consistency requirements. Consistency requirements can include matching test steps with expected results, consistency between expected results and requirements, and reasonable operational logic. If the logic of the third test case does not meet consistency requirements, output a validation result indicating that the consistency check of the third test case logic failed.

[0091] In practical applications, the system can check the consistency of use case logic based on simple rules or additional large language models. Once a missing field or obviously unreasonable content is found, the use case will be marked as requiring manual review.

[0092] Simple rule checks can include keyword feature pool comparison and state machine path verification.

[0093] Keyword feature pool comparison: Extract verbs and their objects (such as "jump", "display", "report error") from the "requirement specification fragment". Search the generated "expected result" field for corresponding synonyms or feature words.

[0094] State machine path verification: Utilize the hierarchical logic defined in the generated "Function-Sub-Function Mapping Table" to verify whether the last step of the "Test Step" and the "Expected Result" belong to the same sub-node range at the functional level.

[0095] Determine whether there is redundancy or conflict among the test cases in the test case set. Redundancy refers to using multiple test cases to repeatedly verify the same test scenario, and conflict refers to different test cases yielding contradictory expected results for the same function. If there is redundancy or conflict between the fourth and fifth test cases, output the verification results for redundancy or conflict between the fourth and fifth test cases.

[0096] In practice, the test case set can be checked one by one by the verification program. First, the format specification and field integrity are verified. Then, the logical consistency is verified by rule matching or large language model. Finally, redundancy and conflict are identified by semantic similarity detection. Unqualified test cases are marked and verification results containing unqualified items and reasons are generated.

[0097] In addition to the above implementation methods, other methods can be used, such as regular expression to verify format, keyword matching to verify content integrity, vector similarity to verify redundancy and conflict, and full verification of large language models. The verification dimensions can also include project specification fields such as test priority, test type, and associated requirement number.

[0098] If the validation results include target test cases that failed validation, the validation results for those target test cases are displayed to allow users to adjust their test cases based on those results. After the test cases are adjusted, the adjusted test case set is output.

[0099] Adjustments to test cases can include user modifications, additions, deletions, and merging of test cases.

[0100] In practice, testers can view and edit tagged test cases through an interactive interface. For example, if the preconditions of a test case are unclear, testers can supplement and improve them; if a crucial business branch is missing test cases, they can manually add them. The entire fine-tuning process can also progressively optimize the model: tester feedback can be used to further adjust the prompts or training for the next LLM generation, enabling the model to generate more accurate test cases under similar requirements. After review, the final confirmed test case set can be exported as a standard format document (such as Excel, CSV, or a customized internal format) with one click, or directly synchronized to the test management platform via a tool interface.

[0101] In addition to the above implementation methods, the verification results can also be displayed through pop-up reminders, document annotations, email notifications, and interface callbacks. The adjustment methods support batch editing, one-click replacement, and automatic field completion.

[0102] By displaying the verification results of the target test cases and providing an entry point for manual adjustment, machine verification can be combined with human experience. This can compensate for possible inference biases in the model based on automated generation, ensuring that the test cases fit the actual business scenario and improving the accuracy and usability of the test cases.

[0103] In this embodiment, after automatically generating test cases from the large language model, the test case set undergoes multi-dimensional automated verification, including format compliance, field completeness, logical consistency, and redundancy / conflict checks. This quickly identifies issues such as format errors, missing content, logical contradictions, and duplicate conflicts, effectively improving the standardization and accuracy of test cases. By displaying target test cases that fail verification and providing a manual adjustment interface, deviations in model generation can be corrected based on human experience, further ensuring that test cases align with actual business and testing needs. After adjustment, a standardized test case set is output. This approach retains the efficiency of automated generation while ensuring test case quality through manual verification, preventing unqualified test cases from entering the test execution phase, reducing testing risks, improving the stability of the testing process, and facilitating test case management and tool integration. It significantly improves the standardization, reliability, and execution efficiency of the overall testing work.

[0104] Figure 3 A schematic diagram of a test case generation device provided in an embodiment of this application includes a construction unit 31, a first generation unit 32, and a second generation unit 33; The construction unit 31 is used to parse multi-source files to construct structured information that the model can recognize; among them, the multi-source files include solution design documents and functional mind maps corresponding to business scenarios; The first generation unit 32 is used to perform semantic parsing of structured information using the first type of large language model to generate a test point list that matches the business scenario. The second generation unit 33 is used to perform reasoning analysis on the test point list using the second type of large language model to generate a test case set that conforms to the test case format; wherein, the test case set includes the test cases corresponding to each test point.

[0105] In some embodiments, the building unit includes a disassembly subunit, an extraction subunit, and a fusion subunit; Sub-unit decomposition is used to break down the chapters and items contained in the scheme design document and determine the paragraph content of each functional point; Extracting sub-units is used to extract the hierarchical relationships of each functional point from the functional mind map; The fusion sub-unit is used to merge the hierarchical relationship of each functional point with the paragraph content of each functional point to establish a function-sub-function mapping table and its corresponding set of requirement specification fragments; wherein, the set of requirement specification fragments includes the text content corresponding to each functional point in the function-sub-function mapping table.

[0106] In some embodiments, the fusion subunit is used to add the hierarchical relationship of each functional point to the corresponding paragraph content to determine the sub-functional list that matches each paragraph content; generate test check items for each functional point based on the paragraph content of each functional point; and integrate the sub-functional list that matches each paragraph content and the test check items of each functional point to obtain a function-sub-functional mapping table and its corresponding set of requirement specification fragments.

[0107] In some embodiments, the first generation unit is configured to perform contextual analysis on structured information using a first type of large language model to deduce test points one by one and identify implicit test requirements; and generate a test point list based on the test points and implicit test requirements; wherein the test point list includes the test scenarios to be verified for each function point and the corresponding expected behavior summary.

[0108] In some embodiments, the second generation unit is used to input the set test case format and test point list, including the test scenarios to be verified for each function point and the corresponding expected behavior summary, into the second type of large language model to output test cases that match each test scenario; wherein, the test case format includes test case number, test case name, preconditions, test steps and expected results.

[0109] In some embodiments, after using the second type of large language model to perform reasoning analysis on the test point list and generate a test case set that conforms to the test case format, the system further includes a verification unit, a display unit, and an output unit. The verification unit is used to verify each test case in the test case set and obtain the verification result. The display unit is used to display the verification results of the target test cases when the verification results include target test cases that fail verification, so that users can adjust the test cases based on the verification results of the target test cases; The output unit is used to output the adjusted set of test cases after the test cases have been adjusted.

[0110] In some embodiments, the verification unit includes a first judgment subunit, a first output subunit, a second judgment subunit, a second output subunit, a third judgment subunit, a third output subunit, a fourth judgment subunit, and a fourth output subunit; The first judgment subunit is used to determine whether each test case in the test case set meets the template format requirements; The first output subunit is used to output the verification result of the first test case failing the format verification if the first test case does not meet the template format requirements. The second judgment subunit is used to determine whether the field content of each test case contained in the test case set is complete; The second output subunit is used to output the verification result that the content verification of the second test case fails when the field content of the second test case is incomplete. The third judgment subunit is used to determine whether the test case logic of each test case contained in the test case set meets the consistency requirements; The third output subunit is used to output the verification result that the logic consistency check of the third test case fails when the test case logic of the third test case does not meet the consistency requirements. The fourth judgment subunit is used to determine whether there is redundancy or conflict among the test cases contained in the test case set; The fourth output subunit is used to output the verification results of the redundancy or conflict between the fourth and fifth test cases if there is redundancy or conflict between them.

[0111] Figure 3 For a description of the features in the corresponding embodiments, please refer to Figure 1 The relevant descriptions of the corresponding embodiments will not be repeated here.

[0112] As can be seen from the above technical solutions, parsing multiple source files constructs structured information that the model can recognize. These multiple source files include solution design documents and functional mind maps corresponding to business scenarios, effectively solving the problems of incomplete and unclear information from a single input source, ensuring the comprehensiveness and accuracy of requirement understanding. The first type of large language model is used to semantically parse the structured information, generating a test point list matching the business scenario. This replaces the tedious process of manually compiling test points, significantly reducing manual analysis costs and avoiding omissions due to differences in individual experience, thus improving the completeness and systematic nature of test point extraction. The second type of large language model is used to reason and analyze the test point list, generating a set of test cases conforming to the test case format. This set includes test cases corresponding to each test point. This ensures that each test point generates corresponding adapted test cases, achieving automated batch production of test cases, significantly shortening the test case development cycle, and ensuring that the generated test cases cover multiple test scenarios, effectively improving test coverage and test case quality. In this technical solution, the illusion phenomenon when generating content by large language models is greatly reduced through dual guidance of text and structure. The dual large language model collaborative architecture is adopted. The first type of large language model solves the problem of omissions in understanding long documents, while the second type of large language model solves the problem of test case generation accuracy. This division of labor and collaboration mode has more advantages in engineering implementation than single model-driven or algorithm iteration, and can complement each other's capabilities, significantly improving the coherence and accuracy of test case logic.

[0113] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described test case generation method embodiments.

[0114] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described test case generation method embodiments at runtime.

[0115] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0116] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described test case generation method embodiments.

[0117] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described test case generation method embodiments.

[0118] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0119] The foregoing has provided a detailed description of a test case generation method, apparatus, electronic device, computer-readable storage medium, and computer program product provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for generating test cases, characterized in that, include: The multi-source files are parsed to construct structured information that the model can recognize; wherein, the multi-source files include solution design documents and functional mind maps corresponding to business scenarios; The structured information is semantically parsed using a first-class large language model to generate a list of test points that match the business scenario; The second type of large language model is used to perform reasoning analysis on the test point list to generate a test case set that conforms to the test case format; wherein, the test case set includes the test cases corresponding to each of the test points.

2. The test case generation method according to claim 1, characterized in that, The multi-source files are parsed to construct structured information that the model can recognize, including: The chapters and items contained in the design document are broken down to determine the paragraph content of each functional point; Extract the hierarchical relationships of each functional point from the functional mind map; The hierarchical relationship of each functional point is integrated with the paragraph content of each functional point to establish a function-sub-function mapping table and its corresponding set of requirement description fragments; wherein, the set of requirement description fragments includes the text content corresponding to each functional point in the function-sub-function mapping table.

3. The test case generation method according to claim 2, characterized in that, The hierarchical relationships between functional points are integrated with the paragraph content of each functional point to establish a function-sub-function mapping table and its corresponding set of requirement specification fragments, including: Add the hierarchical relationship of each function point to the corresponding paragraph content to determine the sub-function list that matches each paragraph content; Based on the paragraph content of each function point, generate test check items for each function point; The sub-function list matching the content of each paragraph and the test check items of each function point are integrated to obtain the function-sub-function mapping table and its corresponding set of requirement specification fragments.

4. The test case generation method according to claim 1, characterized in that, The structured information is semantically parsed using a first-class large language model to generate a list of test points matching the business scenario, including: The first type of large language model is used to perform contextual analysis on the structured information in order to deduce the test points one by one and identify the implicit test requirements; Based on the test points and the implicit test requirements, a test point list is generated; wherein, the test point list includes the test scenarios to be verified for each function point and the corresponding expected behavior summary.

5. The test case generation method according to claim 1, characterized in that, The second type of large language model is used to perform reasoning analysis on the test point list to generate a set of test cases that conform to the test case format, including: The set test case format and the test point list, including the test scenarios to be verified for each function point and the corresponding expected behavior summary, are input into the second type of large language model to output test cases that match each test scenario; wherein, the test case format includes test case number, test case name, preconditions, test steps and expected results.

6. The method for generating test cases according to any one of claims 1 to 5, characterized in that, After using the second type of large language model to perform reasoning analysis on the test point list and generate a test case set that conforms to the test case format, the process also includes: Each test case in the test case set is validated to obtain the validation results; If the verification result includes a target test case that fails verification, the verification result of the target test case is displayed so that the user can adjust the test case based on the verification result of the target test case; After adjusting the test cases, output the adjusted test case set.

7. The test case generation method according to claim 6, characterized in that, The test cases in the test case set are validated to obtain validation results, including: Determine whether each test case in the test case set conforms to the template format requirements; If the first test case does not meet the template format requirements, output the validation result that the first test case failed the format validation; Determine whether the field content of each test case in the test case set is complete; If the field content of the second test case is incomplete, output the validation result that the content validation of the second test case failed; Determine whether the test case logic of each test case in the test case set meets the consistency requirements; If the test case logic of the third test case does not meet the consistency requirements, output the verification result that the logic consistency check of the third test case fails; Determine whether there is redundancy or conflict among the test cases contained in the test case set; If there is redundancy or conflict between the fourth and fifth test cases, output the verification results of the redundancy or conflict between the fourth and fifth test cases.

8. A test case generation device, characterized in that, It includes a building unit, a first generation unit, and a second generation unit; The construction unit is used to parse multi-source files to construct structured information that the model can recognize; wherein, the multi-source files include solution design documents and functional mind maps corresponding to business scenarios; The first generation unit is used to perform semantic parsing on the structured information using a first type of large language model to generate a test point list that matches the business scenario; The second generation unit is used to perform reasoning analysis on the test point list using a second type of large language model to generate a test case set that conforms to the test case format; wherein, the test case set includes test cases corresponding to each of the test points.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the test case generation method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the test case generation method as described in any one of claims 1 to 7.