Data generation system, data generation method, and data generation program

WO2026203711A1PCT designated stage Publication Date: 2026-10-01HITACHI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/001524
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-01-19
Publication Date
2026-10-01

Smart Images

  • Figure JP2026001524_01102026_PF_FP_ABST
    Figure JP2026001524_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a data generation system, a data generation method, and a data generation program with which it is possible to obtain appropriate and comprehensive test data. The present invention comprises: an input / output pattern identification unit that generates an input / output pattern of test data; a test data candidate generation unit that generates a test data candidate including a plurality of pieces of input / output data on the basis of software specification data; and a pattern collation unit that associates individual patterns included in the input / output pattern with individual pieces of input / output data included in the test data candidate and generates test data including a plurality of pieces of input / output data associated with the patterns. The input / output pattern identification unit is provided with: a data generation unit that generates a plurality of pieces of input / output data on the basis of the specification data; a data classification unit that classifies the pieces of input / output data into a plurality of clusters according to similarity of values; a condition analysis unit that extracts classification conditions under which the pieces of input / output data are classified into the clusters; and a pattern generation unit that generates an input / output pattern from a combination of the plurality of extracted classification conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Data generation system, data generation method, and data generation program

[0001] The present invention relates to a data generation system, a data generation method, and a data generation program for generating test data for software testing.

[0002] To ensure software quality, it is necessary to conduct software testing to verify that the software operates according to predetermined specifications. In conducting software testing, it is crucial to comprehensively prepare test data, which consists of pairs of inputs and their expected outputs, in order to verify that the software produces the specified output for each input.

[0003] Prior art for creating test data for such software testing includes Patent Documents 1 and 2. Patent Document 1 discloses a test input data generation process in which an information processing device is made to execute a test case generation process in which an information processing device is made to obtain an arbitrary existing test case from a test case storage unit that stores an existing test case consisting of test input data which is data necessary for a series of operations of the program and test output data obtained as a result of executing the program according to the test input data, and generates new test input data which is different from the test input data of the existing test case; a program execution process which provides the new test input data to the program and makes it execute the series of operations; and a test case generation process which obtains new test output data from the execution result of the program, compares it with the test output data of the existing test case, and if a predetermined difference is detected in the result of the comparison, generates a new test case consisting of the new test input data and the new test output data, and stores the test case in the test case storage unit.

[0004] Patent Document 2 discloses a method for "receiving software specifications, generating a test case from these software specifications that includes test input values ​​for the software and expected output values ​​that are expected to be obtained when the software is executed with the test input values ​​as input, checking whether the values ​​that can be output according to the software specifications are included in the expected output values, and if, as a result of the check, it is determined that the values ​​that can be output according to the software specifications are not included in the expected output values, generating a test case consisting of the values ​​that can be output according to the software specifications and the corresponding test input values, and adding it to the generated test case."

[0005] Japanese Patent Publication No. 2009-163609 Japanese Patent Publication No. 2014-186407

[0006] According to Patent Document 1, by adjusting the input values ​​of existing test data to generate new test data, it is possible to generate diverse test data from a minimum amount of test data. Furthermore, according to Patent Document 2, by expanding the test data based on the output values ​​that can actually occur, it is possible to generate test data that covers all the values ​​that can be output according to the software specifications.

[0007] However, with any of the above conventional technologies, it is difficult to obtain comprehensive test data when the distinction between input and output values ​​is not clearly defined beforehand. For example, when input and output values ​​are continuous, it is often impractical to prepare test data that covers all values. As a countermeasure, one can divide the range of values ​​that the software can output into multiple regions and generate test data for each divided region to cover all output values ​​to be tested. However, if it is unclear how to divide the regions and generate test data, it becomes difficult to obtain appropriate divided regions, and therefore difficult to generate comprehensive test data.

[0008] This invention has been made in view of the above problems, and aims to provide a data generation system, data generation method, and data generation program that can obtain valid and comprehensive test data by appropriately dividing the range of values ​​that input and output values ​​can take from the distribution of input and output values ​​based on the specifications of the software to be tested, even when the division of input and output values ​​is not clearly defined in advance, and generating test data belonging to each range.

[0009] The present invention includes multiple means for solving at least part of the above problems, one example being as follows: A data generation system for generating test data for software testing, comprising: an input / output pattern identification unit that generates an input / output pattern including a plurality of patterns consisting of combinations of classification conditions for the test data; a test data candidate generation unit that generates a plurality of test data candidates including a plurality of input / output data which are pairs of input values ​​and expected output values, based on specification data of the software to be tested; and a pattern matching unit that links each of the patterns included in the input / output pattern with each of the input / output data included in the test data candidate and generates the test data including the plurality of input / output data linked to each of the patterns, wherein the input / output pattern identification unit comprises: a data generation unit that generates a plurality of the input / output data based on the specification data; a data classification unit that classifies each of the input / output data into a plurality of clusters based on the similarity between the input value and the expected output value included in the input / output data; a condition analysis unit that extracts a plurality of classification conditions for classifying each of the input / output data into each of the clusters from the results of the classification; and a pattern generation unit that generates the input / output pattern from a brute-force combination of the extracted plurality of classification conditions.

[0010] According to the present invention, a data generation system, a data generation method, and a data generation program are provided that can obtain valid and comprehensive test data by appropriately dividing the range of possible input and output values ​​from the distribution of input and output values ​​based on the specifications of the software to be tested, and generating test data belonging to each range.

[0011] Other issues, configurations, and effects not mentioned above will be clarified by the following description of the embodiments.

[0012] This figure shows an example of the configuration of the data processing system in Example 1. This figure shows an example of the device configuration of the test target information management system, data generation system, and test execution system included in the data processing system in Example 1. This is a functional block diagram showing an example of the configuration of the data generation system in Example 1. This is a flowchart showing an example of the test data generation process by the data generation system in Example 1. This is a functional block diagram showing an example of the configuration of the input / output pattern identification unit in Example 1. This is an explanatory diagram of the overview of the processing of the input / output pattern identification unit in Example 1. This is a functional block diagram showing an example of the configuration of the test data candidate generation unit in Example 1. This is an explanatory diagram of the overview of the processing of the test data candidate generation unit in Example 1. This is a functional block diagram showing an example of the configuration of the pattern matching unit in Example 1. This is an explanatory diagram of the overview of the processing of the pattern matching unit in Example 1. This figure shows an example of the display by the output device of the result of linking input / output patterns stored in the memory area with test data candidates in Example 1. This is a figure showing the range of possible values ​​for the two output items of the test target in a two-dimensional Cartesian coordinate system. This is a figure showing the range of possible values ​​for the two output items of the test target in a two-dimensional Cartesian coordinate system. This is a functional block diagram showing an example of the configuration of the input / output pattern identification unit in Example 2. This is an explanatory diagram of the overview of the processing of the pattern generation unit in Example 2. This is a functional block diagram showing an example of the configuration of the data generation system in Example 3. This is an explanatory diagram outlining the processing of the feasibility estimation unit in Example 3. This is a functional block diagram showing an example of the configuration of the test data candidate generation unit in Example 4. This is a functional block diagram showing an example of the configuration of the data generation system in Example 5. This is a functional block diagram showing an example of the configuration of the data generation system in Example 6. This is a functional block diagram showing an example of the configuration of the data generation system in Example 7. This is a functional block diagram showing an example of the configuration of the input / output pattern identification unit in Example 8.

[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The embodiments are examples for explaining the present invention, and omissions and simplifications are appropriately made for clarifying the description. The present invention can also be implemented in various other forms. Unless particularly limited, each constituent element may be either singular or plural.

[0014] The position, size, shape, range, etc. of each constituent element shown in the drawings may not represent the actual position, size, shape, range, etc. in order to facilitate understanding of the invention. Therefore, the present invention is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings. When there are a plurality of constituent elements having the same or similar functions, description may be given by adding different subscripts to the same reference numeral. Further, when there is no need to distinguish between the plurality of constituent elements, the subscript may be omitted in the description.

[0015] In the embodiments, processing performed by executing a program may be described. Here, a computer executes a program by a processor (e.g., CPU, GPU), and performs processing defined by the program while using storage resources (e.g., memory), interface devices (e.g., communication ports) and the like. Therefore, the processor may be the subject of processing performed by executing the program. Similarly, the subject of processing performed by executing a program may be a controller, an apparatus, a system, a computer, or a node including a processor.

[0016] The subject of processing performed by executing a program only needs to be an arithmetic unit, and may include a dedicated circuit that performs specific processing. Here, the dedicated circuit is, for example, FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit), CPLD (Complex Programmable Logic Device), or the like.

[0017] The program may be installed into a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable storage medium. When the program source is a program distribution server, the program distribution server includes a processor and a storage resource that stores the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. Further, in the embodiments, two or more programs may be implemented as one program, or one program may be implemented as two or more programs.

[0018] (1) Configuration of Data Processing System Fig. 1 is a diagram showing an example of the configuration of the data processing system according to the first embodiment. The data processing system 1 includes a test target information management system 2, a data generation system 3, a test execution system 4, and a communication path 5.

[0019] The communication path 5 is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network), and is a communication path that communicatively connects each system (devices and terminals constituting each system) of the data processing system 1 to each other.

[0020] The test target information management system 2 is a system operated by a test target information manager. The test target information manager is, for example, a software vendor or a developer of software used in an own organization, and is an entity that manages information related to software used by the own organization, another organization, or an individual. The information related to software includes at least information on items input to the software, and also includes specification information such as output items from the software, types and ranges of values that can be taken by each item of input and / or output, and the content of processing executed by the software. Hereinafter, the information related to software is referred to as software specification data or simply specification data. The test target information management system 2 stores software specification data, and transmits and updates information as necessary.

[0021] The data generation system 3 is a system used by data generation system users. Data generation system users are entities that obtain specification data of the software under test (hereinafter referred to as the software under test or simply the software under test) from the software under test information management system 2 and use the data generation system 3 to generate test data for software testing (hereinafter referred to as testing).

[0022] The data generation system 3 generates test data for testing using the specification data of the software under test. Test data is data consisting of a combination of input values ​​for testing the software under test and the expected output values ​​for those input values.

[0023] The test execution system 4 is a system used by the test executor. The test executor is the entity that uses the test data generated by the data generation system 3 to execute tests using the test execution system 4.

[0024] The test execution system 4 uses test data to perform tests on the software under test. During the test, for example, it inputs test input values ​​defined in the test data into the software and determines whether the output result matches the expected output value.

[0025] Furthermore, some or all of the users of the data generation system, the administrator of the information under test, and the test executors may be the same entity. Also, some or all of the data generation system 3, the information under test management system 2, and the test execution system 4 may operate on the same information processing device or system. In this case, data transmission and reception between each system may be carried out via independent paths without going through the communication path 5.

[0026] (2) Diagram 2 of the device configuration for each system shows an example of the device configuration of the test target information management system, data generation system, and test execution system included in the data processing system in Example 1.

[0027] The test target information management system 2 consists of at least a communication device 21 and a storage device 22. The storage device 22 has a storage area 221 for storing specification data 221A of the software under test. As described above, the specification data 221A includes at least information on input items to the software, as well as output items from the software, the types and ranges of values ​​that each input or output or both (input / output) item can take, the content of processing performed by the software, etc. (Hereafter, input items, output items, and input / output items are simply referred to as input, output, and input / output, respectively). The specification data 221A is created by the test target information manager and stored in the storage device 22.

[0028] The data generation system 3 consists of a CPU (Central Processing Unit) 31, an input device 32, an output device 33, a communication device 34, and a storage device 35. The data generation system 3 is an information processing device or system such as a personal computer, a server computer, or a handheld computer.

[0029] The input device 32 consists of a keyboard or mouse, and the output device 33 consists of a display or printer. The communication device 34 is equipped with a NIC (Network Interface Card) for connecting to a wireless LAN or wired LAN. The storage device 35 is a storage medium such as RAM (Random Access Memory) or ROM (Read Only Memory). Output results and intermediate results from each processing unit may be output as appropriate via the output device 33.

[0030] The storage device 35 stores programs that implement the input / output pattern identification unit 351, the test data candidate generation unit 352, and the pattern matching unit 353, which will be described later, and also has a storage area 354 for storing the test data generation results. When each program stored in the storage device 35 is executed by the CPU 31, the input / output pattern identification unit 351, the test data candidate generation unit 352, and the pattern matching unit 353, which will be described later, are implemented in the data generation system 3.

[0031] The test execution system 4 consists of at least a communication device 41 and a test execution device 42. The test execution device 42 is an information processing device or system such as a personal computer, server computer, or handheld computer.

[0032] (3) Configuration and Processing Flow of the Data Generation System The configuration and processing flow of the data generation system 3 in Example 1 will be described using diagrams 3 and 4. Figure 3 is a functional block diagram showing an example of the configuration of the data generation system in Example 1. The data generation system 3 consists of an input / output pattern identification unit 351, a test data candidate generation unit 352, a pattern matching unit 353, and a storage area 354.

[0033] Figure 4 is a flowchart showing an example of the test data generation process by the data generation system in Embodiment 1. The processing content of each component of the data generation system 3 will be explained using Figure 4. The test data generation process shown in Figure 4 is initiated when the data generation system 3 receives specification data 221A from the test target information management system 2, or when a user of the data generation system performs an input operation via the input device 32, and the data generation system 3 executes the processes from steps S601 to S606.

[0034] The data generation system 3 acquires the specification data 221A of the test target from the storage area 221 of the storage device 22 provided by the test target information management system 2 (S601). The specification data 221A is input to the input / output pattern identification unit 351 and the test data candidate generation unit 352, respectively, as shown in Figure 3.

[0035] The input / output pattern identification unit 351 uses the input specification data 221A to identify input / output patterns that are represented as combinations of classification conditions for values ​​that the input / output under test can take (hereinafter, combinations of classification conditions are referred to as patterns), as described later (S602).

[0036] The test data candidate generation unit 352 uses the input specification data 221A to generate test data candidates, which consist of a combination of a test input value and an expected output value for that input value (expected output value) (hereinafter referred to as a set of input value and expected output value, or input / output data), as described later (S603).

[0037] The pattern matching unit 353 receives the input / output pattern identified by the input / output pattern identification unit 351 and the test data candidate generated by the test data candidate generation unit 352 as input, and matches each input / output data included in the test data candidate to which pattern included in the input / output pattern, thereby linking the input / output pattern and the test data candidate (S604).

[0038] Next, the pattern matching unit 353 calculates the coverage rate as the ratio of the total number of input / output patterns to the number of patterns for which corresponding input / output data existed (associated), and determines whether the coverage rate is equal to or greater than a predetermined threshold (S605). The threshold is predetermined and set in the pattern matching unit 353.

[0039] If the coverage rate is determined to be less than a predetermined threshold, the pattern matching unit 353 causes the test data candidate generation unit 352 to re-execute the test data candidate generation process (S603), and repeats the processes of S604 and S605.

[0040] If the coverage rate is determined to be above a predetermined threshold, the pattern matching unit 353 completes the test data generation process and outputs the result of associating input / output patterns with test data candidates. The result of associating input / output patterns with test data candidates output from the pattern matching unit 353, that is, the test data 354A composed of multiple input / output data associated with each pattern, is stored in the storage area 354 and output to the test execution system 4 (S606). The test data 354A stored in the storage area 354 may also be output to the output device 33 of the data generation system 3 as appropriate (for example, by displaying it on a screen).

[0041] The test execution system 4 uses the test data 354A input from the data generation system 3 to actually execute the tests under test.

[0042] Furthermore, if the pattern matching unit 353 determines in S605 that the coverage rate does not exceed a predetermined threshold and the processing of S603 to S605 is repeated, the processing time and number of processing steps for S603 to S605 are measured. If the processing time reaches a predetermined time (for example, 30 minutes or 1 hour, or any arbitrarily set time) or the number of processing steps reaches a predetermined number (for example, 5 times, or any arbitrarily set number) due to the multiple repetitions of S603 to S605, the pattern matching unit 353 may output the result of linking the input / output patterns with the test data candidates at that time as test data 354A and complete the test data generation process. In this case, the measurement of processing time and number of processing steps may be performed by the pattern matching unit 353, the test data candidate generation unit 352, or the CPU 31, etc.

[0043] (4) Details of each component of the data generation system (4-1) Input / Output Pattern Identification Unit Details of the input / output pattern identification unit 351 will be explained using Figures 5 and 6. Figure 5 is a functional block diagram showing an example of the configuration of the input / output pattern identification unit in Embodiment 1. The input / output pattern identification unit 351 consists of a data generation unit 3511, a data classification unit 3512, a condition analysis unit 3513, and a pattern generation unit 3514.

[0044] The data generation unit 3511 receives the specification data 221A of the test target as input from the test target information management system 2. From the received specification data 221A, the data generation unit 3511 generates multiple pairs of values ​​that simulate the input given to the test target (test input values) and the expected output values ​​for those input values ​​(input / output data).

[0045] The data generation unit 3511 generates values ​​that simulate inputs for each input item included in the specification data 221A. As for the method of generating values ​​that simulate inputs, a method of random generation within an arbitrary range or within a type or range defined in the specification data 221A may be adopted, or a method of random generation within a predetermined search range for input values ​​may be adopted.

[0046] Furthermore, the data generation unit 3511 calculates expected output values ​​for values ​​that simulate the input. Expected output values ​​are the values ​​of each output item that are expected to be output if the system under test operates according to the specification data. The calculation method may involve using an emulator or simulator created in advance to simulate the processing content of the system under test, or by inputting information about the processing content executed by the system under test into a large-scale language model to generate output values.

[0047] The data generation unit 3511 outputs the multiple input and output data generated by the above process to the data classification unit 3512.

[0048] The data classification unit 3512 receives multiple input and output data generated by the data generation unit 3511 as input. The data classification unit 3512 classifies each input and output data into multiple clusters (also called clustering) according to the degree of similarity between them.

[0049] For classification, one may employ hierarchical clustering methods such as Ward's method, single-link method, full-link method, and centroid method; clustering methods that optimize the nearest neighbor, such as k-means, EM algorithm, and spectral clustering; or clustering methods that optimize the discriminant boundary, such as unsupervised SVM (Support Vector Machine), VQ algorithm, and SOM (Self-Organizing Maps).

[0050] Furthermore, the number of clusters to be classified may be determined by a number arbitrarily decided in advance, by the number of classifications that yields the best information criteria such as AIC (Akaike Information Criterion) or BIC (Bayesian Information Criterion), or by the number of classifications that ensures the sum squared deviation, determined by the representative value of each classification (such as the cluster centroid) and the value belonging to that classification, falls within a specific range.

[0051] The data classification unit 3512 outputs multiple input and output data and their classification results (also called clustering results) to the condition analysis unit 3513.

[0052] The condition analysis unit 3513 receives multiple input / output data and their classification results as input from the data classification unit 3512. The condition analysis unit 3513 constructs a classification model that determines which cluster each input / output data belongs to. The construction of the classification model may employ methods based on arbitrary evaluation metrics such as maximizing information gain, minimizing information entropy, the Gini coefficient, or classification error, and may take the form of branching from a parent node to any number of child nodes. Furthermore, it extracts multiple classification conditions from each parent node to the child nodes from the constructed classification model. At this time, weighting and filtering of each classification condition may be performed according to the frequency of appearance of a particular classification condition in a single classification model, the frequency of appearance of a particular classification condition in multiple classification models when multiple classification models are constructed, or the speed at which each classification condition appears when the classification conditions are extracted from the root node. The condition analysis unit 3513 outputs the extracted multiple classification conditions to the pattern generation unit 3514.

[0053] The pattern generation unit 3514 receives multiple classification conditions as input from the condition analysis unit 3513. The pattern generation unit 3514 generates a brute-force combination of each classification condition. The pattern generation unit 3514 then aggregates the brute-force combinations of classification conditions generated by the above process and outputs them to the pattern matching unit 353 as an input / output pattern.

[0054] Figure 6 is an explanatory diagram illustrating the overview of the processing of the input / output pattern identification unit in Embodiment 1. The processing content of the input / output pattern identification unit 351 will be explained in more detail using Figure 6. In the example shown in Figure 6, each input item to be tested is IN A IN B And so on, and each output item is OUT A OUT B It is indicated as such. Furthermore, each input / output item is assumed to take a continuous value.

[0055] The data generation unit 3511 first randomly generates input values ​​3511A for each input item. Furthermore, using an emulator created by simulating the processing content under test, it generates expected output values ​​3511B corresponding to each input value 3511A for each output item. In the example shown in Figure 6, sequential numbers starting from 1 are assigned to each generated input and output data for explanatory purposes, but such numbering is not necessary, and if numbers are assigned, some kind of identification symbol can be used instead of sequential numbers.

[0056] The data classification unit 3512 groups each input and output data generated by the data generation unit 3511 together with the same items (for example, OUT) A The data is compared with other data (e.g., with each other), and pairs with similar values ​​are classified into the same cluster. Figure 6 shows an example where the data classification unit 3512 classifies each pair into a total of four clusters, Clusters 1 to 4. In this example, for example, input and output data with serial numbers #2, 3, 12, and 14 are classified into Cluster 1 because they are similar, and similarly, input and output data with serial numbers #1, 6, 9, and 11 are classified into Cluster 2, input and output data with serial numbers #5, 8, 10, and 16 are classified into Cluster 3, and input and output data with serial numbers #4, 7, 13, and 15 are classified into Cluster 4. As a result, a clustering result 3512A is obtained in which each input and output data is classified into one of four clusters.

[0057] The condition analysis unit 3513 constructs a classification model 3513A that determines which of the four clusters (Cluster 1 to 4) included in the clustering result 3512A each input / output data belongs to. In FIG. 6, the condition analysis unit 3513 employs a binary tree model, uses the four clusters (classification labels) included in the clustering result 3512A as objective variables, and uses, as feature quantities, the value of each item of the expected output value 3511B and the magnitude relationship between each item of the expected output value 3511B and each item of the input value 3511A to construct the classification model 3513A. An example of this construction is shown. Regarding the magnitude relationship, for example, OUT B has a value that is IN A can be expressed by generating a binary value that is 0 if it is smaller than the value of, and 1 otherwise. Also, in the example shown in FIG. 6, the item of the expected output value 3511B used as a feature quantity is OUT A and OUT B only, and the items of the input value 3511A are also IN A and IN B only, and the classification model 3513A constructed by the condition analysis unit 3513 is shown in a simplified manner.

[0058] As shown in FIG. 6, in the classification model 3513A constructed using such objective variables and feature quantities, OUT A includes one classification condition for binary splitting based on the value of OUT B includes two classification conditions for binary splitting based on the value of, and the condition analysis unit 3513 extracts these three classification conditions.

[0059] Note that, in the example shown in FIG. 6, values of items of the input value 3511A other than those described above, the magnitude relationship between items of the input value 3511A, and the magnitude relationship between items of the expected output value 3511B are not used as feature quantities, but a classification model may be constructed using these values, magnitude relationships, etc. as feature quantities.

[0060] The pattern generation unit 3514 generates an input / output pattern 3514A by generating all brute-force combinations of a total of three classification conditions extracted from the classification model 3513A. In the example shown in FIG. 6, OUT A is a condition for binary splitting the value with a constant 30, OUT B is IN AThe condition for dividing into two is OUT. B IN B By exhausting all possible combinations of the conditions for dividing into two, 2 × 2 × 2 = 8 possible input / output patterns 3514A are generated.

[0061] (4-2) Test Data Candidate Generation Unit Details of the test data candidate generation unit 352 will be explained using Figures 7 and 8. Figure 7 is a functional block diagram showing an example of the configuration of the test data candidate generation unit 352 in Embodiment 1. The test data candidate generation unit 352 consists of an input value generation unit 3521 and an assumed output calculation unit 3522.

[0062] The input value generation unit 3521 receives specification data 221A as input from the test target information management system 2. The input value generation unit 3521 generates one or more values ​​that simulate the inputs given to the test target from the received specification data 221A. In this case, the input value generation unit 3521 generates values ​​that simulate the inputs for each input item included in the specification data 221A. As for the generation method, a method of random generation within an arbitrary range or a type or range defined in the specification data 221A may be adopted, or a method of random generation within each range defined in advance as an input value search range. The input value generation unit 3521 outputs the generated one or more input values ​​to the assumed output calculation unit 3522. Note that, as described in the explanation of Figure 4, the input value generation unit 3521 may also execute the above process triggered by an instruction from the pattern matching unit 353.

[0063] The expected output calculation unit 3522 receives one or more input values ​​generated by the input value generation unit 3521 and calculates one or more expected output values ​​from each input value. The calculation method may include using an emulator or simulator created in advance to simulate the processing content of the system under test, or inputting information about the processing content executed by the system under test into a large-scale language model to generate output values. When using the latter method, the expected output calculation unit 3522 is configured to also receive specification data 221A as input. Specifically, the expected output calculation unit 3522 is configured to acquire necessary information regarding the output items, the range of possible values, and the processing content executed by the system under test.

[0064] The expected output calculation unit 3522 outputs multiple input / output data sets, each combining an input value and each expected output value, to the pattern matching unit 353 as test data candidates. Note that at either the input value generation unit 3521 or the expected output calculation unit 3522, unnecessary input values ​​or input / output data may be deleted according to predefined conditions.

[0065] Figure 8 is an explanatory diagram illustrating the overview of the processing of the test data candidate generation unit in Embodiment 1. The processing content of the test data candidate generation unit 352 will be explained in more detail using Figure 8. In the example shown in Figure 8, similar to the example shown in Figure 6, each input item of the test target is IN A IN B And so on, and each output item is OUT A OUT B This indicates, for example.

[0066] The input value generation unit 3521 generates multiple input values ​​3521A for each input item included in the specification data 221A. The input value generation unit 3521 may, for example, randomly generate multiple values ​​for each input item. Alternatively, the input value generation unit 3521 may be configured to receive only the input item information for the test target from the specification data 221A.

[0067] The assumed output calculation unit 3522 generates assumed output values ​​3522A corresponding to each of the input values ​​3521A. The assumed output calculation unit 3522 automatically extracts each output item using an emulator created by simulating the processing content of the test target, and calculates the assumed output value for each output item. The assumed output calculation unit 3522 also combines the corresponding input values ​​and assumed output values ​​in the input values ​​3521A and assumed output values ​​3522A to generate test data candidate 3522B containing multiple input / output data (in the example in Figure 8, input and output data are shown row by row), as shown in Figure 8.

[0068] (4-3) Pattern Matching Section The details of the pattern matching section 353 will be explained using Figures 9, 10, and 11. Figure 9 is a functional block diagram showing an example of the configuration of the pattern matching section in Embodiment 1.

[0069] The pattern matching unit 353 consists of a data merging unit 3531 and a comprehensive determination unit 3532.

[0070] The data merging unit 3531 receives input / output patterns from the input / output pattern identification unit 351 and test data candidates from the test data candidate generation unit 352 as input. The data merging unit 3531 determines which of the patterns included in the input / output pattern each input / output data in the test data candidate corresponds to (which of which it satisfies) and associates the input / output data with the corresponding pattern. In this case, if multiple input / output data are associated with one pattern in the input / output pattern, the unit may associate only one arbitrarily selected pair, associate all input / output data, or arbitrarily select and associate a predetermined number of input / output data. Also, if a single input / output data is associated with multiple patterns in the input / output pattern, the unit may associate that input / output data with all corresponding patterns, or associate it only with the pattern that does not yet have any input / output data associated with it. The data merging unit 3531 outputs the result of associating the input / output pattern with the test data candidate to the coverage determination unit 3532.

[0071] The coverage determination unit 3532 receives the result of linking input / output patterns with test data candidates from the data linking unit 3531 as input. The coverage determination unit 3532 calculates the coverage rate as the ratio of the number of patterns for which corresponding input / output data exists (is linked) to the total number of input / output patterns. Next, the coverage determination unit 3532 determines whether the calculated coverage rate is equal to or greater than a predetermined threshold. If the coverage rate is equal to or greater than the predetermined threshold, it completes the test data generation process. If the coverage rate is less than the predetermined threshold, it sends an instruction to the test data candidate generation unit 352 to re-execute the test data candidate generation process.

[0072] As described in the explanation of Figure 4, the pattern matching unit 353 outputs the result of linking the input / output patterns with the test data candidates as test data 354A when the coverage determination unit 3532 determines that the coverage rate is above a predetermined threshold. However, the timing of outputting the linking result is not limited to this. For example, the linking result may be output each time the above processing is executed by the data merging unit 3531 and the coverage determination unit 3532 (regardless of the determination result by the coverage determination unit 3532), or each time the above processing is re-executed (repeated) by the test data candidate generation unit 352, the data merging unit 3531, and the coverage determination unit 3532 for a predetermined number of times or periods of time.

[0073] If output is generated each time the above processing is performed by the data merging unit 3531 and the coverage determination unit 3532, the pattern matching unit 353 outputs the linking result as an intermediate result while the coverage rate is below a predetermined threshold, and outputs the linking result as the final result, i.e., test data 354A, when the coverage rate becomes equal to or greater than the predetermined threshold. Also, if output is generated each time the above processing is re-executed (repeated) by the test data candidate generation unit 352, the data merging unit 3531, and the coverage determination unit 3532 for a predetermined number of times or periods of time, the pattern matching unit 353 outputs the linking result as the final result, i.e., test data 354A, even if the coverage rate is below a predetermined threshold.

[0074] In this way, the intermediate and final results (test data 354A) output from the pattern matching unit 353 are stored in the storage area 354. As described in the explanation of Figure 4, the final result, test data 354A, is output to the test execution system 4. The intermediate and final results stored in the storage area 354 may also be output to the output device 33 of the data generation system 3 as appropriate (for example, by displaying them on a screen).

[0075] Figure 10 is an explanatory diagram illustrating the overview of the pattern matching unit's processing in Embodiment 1. The processing content of the pattern matching unit 353 will be explained in more detail using Figure 10. In the example shown in Figure 10, similar to the examples shown in Figures 6 and 8, each input item of the test target is entered as IN A IN B etc., output each output item. AOUT B The items are shown as such, and each item is assumed to take a continuous value. Furthermore, as a prerequisite for processing, the pattern matching unit 353 receives the input / output pattern 3514A exemplified in Figure 6 from the input / output pattern identification unit 351 and the test data candidate 3522B exemplified in Figure 8 from the test data candidate generation unit 352 as input.

[0076] The data merging unit 3531 determines which of the patterns #1 to #8 included in the input / output pattern 3514A each input / output data included in the test data candidate 3522B (shown row by row in Figure 10) corresponds to, and associates each input / output data with the corresponding pattern. In Figure 10, the determination result is shown as 3531A. In this example, for example, the input / output data shown in the first row of the test data candidate 3522B is determined to correspond to pattern #2, the input / output data shown in the second row is determined to correspond to pattern #4, the input / output data shown in the fourth row is determined to correspond to pattern #8, the input / output data shown in the fifth row is determined to correspond to pattern #5, and the input / output data shown in the sixth row is determined to correspond to pattern #6, and they are associated accordingly.

[0077] In the example shown in Figure 10, the system ensures that no more than two sets of input / output data are associated with each pattern included in the input / output pattern 3514A. Specifically, if multiple input / output data points correspond to a single pattern, only the first input / output data point determined (the input / output data point appearing in a higher row of the test data candidate 3522B in Figure 10) is associated. For example, the input / output data shown in the third row of test data candidate 3522B corresponds to pattern #2 of input / output pattern 3514A, but since the input / output data shown in the first row is already associated with pattern #2, this input / output data is not associated with pattern #2. On the other hand, for patterns #1, #3, and #7 of input / output pattern 3514A, there is no corresponding input / output data in test data candidate 3522B, so none of the input / output data points are associated.

[0078] The coverage determination unit 3532 calculates the coverage rate 3532A of the input / output patterns according to the result of the linking of the input / output patterns 3514A and the test data candidates 3522B in the data linking unit 3531. In the example shown in Figure 10, the total number of input / output patterns 3514A is "8", and the number of patterns #2, #4, #5, #6, and #8 to which input / output data are linked is "5", so the coverage rate 3532A calculated by the coverage determination unit 3532 is 5 / 8 = 62.5%.

[0079] The coverage determination unit 3532 determines whether the coverage rate 3532A calculated in this way is equal to or greater than a predetermined threshold. If the coverage rate is equal to or greater than the predetermined threshold, it completes the test data generation process. If the coverage rate is less than the predetermined threshold, it sends an instruction to the test data candidate generation unit 352 to re-execute the test data candidate generation process.

[0080] Figure 11 shows an example of the display by the output device of the association results between input / output patterns and test data candidates stored in the memory area in Embodiment 1. Using Figure 11, the contents of the intermediate or final results of the association results between input / output patterns and test data candidates will be explained. As a premise, it is assumed that test data 354A, which is the intermediate or final result of the association results between input / output pattern 3514A and test data candidate 3522B, is stored in the memory area 354.

[0081] The output device 33 displays the test data 354A, which is an intermediate or final result of the association between the input / output pattern 3514A stored in the memory area 354 and the test data candidate 3522B, on the display as, for example, display content 33A. In display content 33A, all the contents of the input / output pattern 3514A are shown on the left, and for patterns in the input / output pattern 3514A to which input / output data is associated, the associated input / output data is shown on the right. The blank area (row) on the right side of display content 33A means that there is no input / output data in the test data candidate 3522B corresponding to the pattern shown on the left, and no input / output data is associated. When there is no blank area on the right side of display content 33A, that is, when input / output data is associated with all patterns, it can be considered that comprehensive test data has been obtained.

[0082] Furthermore, the display content 33A of the test data 354A, which is an intermediate or final result, is not limited to the example shown in Figure 11, but may be displayed in any other manner, and may also include information other than the intermediate result or test data 354A. For example, the display content 33A may include information on the coverage rate of the intermediate result or test data 354A to be displayed, or information on input / output data from the test data candidate 3522B that were not associated with the input / output pattern 3514A.

[0083] (5) Operation and Effects of the Data Generation System The mechanism for generating valid and comprehensive test data using the data generation system will be explained using Figures 12A and 12B.

[0084] Figures 12A and 12B show the two output items (OUT) being tested. A OUT B This figure shows the range of possible values ​​for ) in a two-dimensional Cartesian coordinate system, and all are OUT. A The value of the X axis, OUT B The value of OUT is used as the Y-axis. A and OUT B The expected output values ​​for each are shown in a two-dimensional Cartesian coordinate system with each data point 1200. Figures 12A and 12B show, as an example, 15 sets of OUT calculated by the data generation unit 3511. A and OUT B The expected output values ​​are shown in a two-dimensional Cartesian coordinate system.

[0085] As described in the explanation of Figure 1, in testing, it is necessary to generate test data that covers as many expected output values ​​as possible in order to determine whether the output value of the test target for various test input values ​​is the same as the expected output value. However, since it is not practical to generate test data that covers all expected output values ​​in the execution of the test, there is a method to generate test data that covers the expected output values ​​in a realistic way by dividing the range of values ​​that each output item of the test target can take into multiple regions and generating at least one test data for each divided region. Figures 12A and 12B show the output items of the test target, OUT, based on this method. A and OUTB This example shows how to divide the range of possible values ​​into multiple regions.

[0086] Figure 12A shows OUT A and OUT B Figure 12B shows an example where the range of possible values ​​is divided into equal intervals (in this example, into 9 regions enclosed by dotted square frames), and test data (i.e., pairs of input values ​​and expected output values) belonging to each divided region is generated. Figure 12B shows an example where the division is based on the input / output pattern identified by the input / output pattern identification unit 351 (in this example, into 6 regions), and test data belonging to each divided region is generated. Note that when dividing at equal intervals as in Figure 12A, minimizing the intervals to the minimum possible size of each divided region corresponds to a full test that covers all expected output values.

[0087] Incidentally, the operation of the system under test may involve typical input values ​​producing typical output values, or rare input values ​​producing rare output values. Due to such operation, the distribution of expected output values ​​for the system under test is divided into dense parts representing typical cases and sparse parts representing rare cases. For example, in the example shown in Figure 12A, parts 1210 and 1230 enclosed by dotted circles are dense parts, and parts 1220 and 1240 enclosed by dotted circles are sparse parts. In each of these dense parts 1210 and 1230 and sparse parts 1220 and 1240, the one or more expected output values ​​contained in each part will be similar to each other. Therefore, it is not necessarily required to generate test data for each expected output value contained in each part; for example, generating at least one test data for each part is sufficient to cover all expected output values ​​in that part.

[0088] However, in the example shown in Figure 12A, OUT is generated regardless of the density of the distribution of the assumed output values ​​as described above. A and OUT BBecause the range of possible values ​​is divided into equal intervals, even if the assumed output values ​​are similar, such as in dense areas 1210 and 1230 or sparse areas 1240, they will be divided into two or more regions. Test data will be generated for each of these divided regions, resulting in redundant testing for each portion. Conversely, dense and sparse regions may be included in the same region, and only typical or rare test data may be generated for that divided region, potentially leading to insufficient testing.

[0089] In contrast, in the example shown in Figure 12B, the density of the distribution of the assumed output values ​​as described above is taken into account, and OUT A and OUT B By dividing the range of possible values, the resulting division will be in a reasonable form (a reasonable division region will be obtained).

[0090] Let me explain in detail. First, as described in the explanation of Figure 6, the data classification unit 3512 compares each input and output data generated by the data generation unit 3511 with those of the same item, and classifies them into Clusters 1 to 4 according to the degree of similarity of their values. As a result, as shown in Figure 12B, each assumed output value (each data point 1200) is classified into Clusters 1 to 4 (regions enclosed by dashed rectangular frames in Figure 12B) according to its degree of similarity. In Figure 12B, the regions of Clusters 1 and 3 correspond to the denser parts, respectively, and the regions of Clusters 2 and 4 correspond to the sparser parts, respectively. In this way, the density of the distribution of assumed output values ​​is taken into account by the classification according to the degree of similarity of the values, and the OUT A and OUT B The range of possible values ​​will be divided into regions 1 through 4.

[0091] Next, as described in the explanation of Figure 6, the condition analysis unit 3513 extracts multiple classification conditions from the classification model 3513A for classifying each input / output data into Clusters 1 to 4, and the pattern generation unit 3514 generates an input / output pattern 3514A from all possible combinations of the multiple classification conditions. This input / output pattern 3514A is not only a combination of classification conditions for Clusters 1 to 4, but also conditions for further classifying Clusters 1 to 4 (i.e., conditions for further dividing each region of Clusters 1 to 4).

[0092] For example, in the example shown in Figure 6, patterns #3 and #4 of the input / output pattern 3514A are both combinations of classification conditions for Cluster1, A The value of IN B If the value is greater than this, it becomes a condition to further classify Cluster1 into two (divide the area of ​​Cluster1 into two). Similarly, patterns #1 and #2 are both combinations of classification conditions for Cluster2, but IN A The value of IN B If the value is smaller than this, it becomes a condition to classify (split) Cluster 2 into two. Patterns #5 and #7 are both combinations of classification conditions for Cluster 3, but IN A The value of IN B If the value is greater than this, it becomes a condition to classify (split) Cluster 3 into two. Patterns #6 and #8 are both combinations of classification conditions for Cluster 4, but IN A The value of IN B If the value is smaller than this, it is a condition to classify (divide) Cluster 4 into two. Note that Figure 12B is IN A The value of IN B This example shows how, when the value is greater than [a certain value], each region of Cluster1 and Cluster3 is divided into two by patterns #3, 4 and #5, 7.

[0093] In this way, by applying input / output patterns to further classify Clusters 1 to 4, the classification based on the degree of similarity of values ​​can be further refined, that is, it becomes possible to classify by more strictly determining the degree of similarity of values. Then, according to this refined classification, OUT A and OUT B By dividing the range of possible values, each divided region better reflects the density of the distribution of the expected output values, thereby improving the validity of the division.

[0094] Furthermore, even if the number of input / output data generated by the data generation unit 3511 is insufficient and the degree of similarity of the values ​​cannot be sufficiently determined (i.e., the tendency of density in the distribution of expected output values ​​cannot be fully reflected), resulting in insufficient classification into clusters, applying input / output patterns to supplement the classification enables detailed classification according to the degree of similarity of the values, and according to this classification, OUT A and OUT B By dividing the range of possible values, the validity of the division can be improved.

[0095] Furthermore, by generating test data that covers the reasonable division regions described above, it becomes possible to generate comprehensive test data that enables thorough testing without omissions while suppressing redundant test execution for multiple similar expected output values.

[0096] In Example 1, an example was described in which the input / output pattern identification unit 351 identifies input / output patterns using only specification data 221A. However, the input / output pattern identification unit 351 is not limited to this, and may also identify input / output patterns using predetermined correction data. Therefore, Example 2 describes an example of an input / output pattern identification unit configured in this way. In the following description, explanations that overlap with Example 1 will be omitted, and only the differences will be described.

[0097] Figure 13 is a functional block diagram showing an example of the configuration of the input / output pattern identification unit in Embodiment 2, and components identical to those in the input / output pattern identification unit 351 shown in Figure 5 are denoted by the same reference numerals. The input / output pattern identification unit 1351 in Embodiment 2 has substantially the same configuration as the input / output pattern identification unit 351 shown in Figure 5, differing only in that it includes a pattern generation unit 3515 instead of a pattern generation unit 3514.

[0098] The pattern generation unit 3515 receives multiple classification conditions from the condition analysis unit 3513 as input, and also receives correction data 221B from the test target information management system 2. The correction data 221B is data that specifies additional conditions to be included in the classification conditions and deletion conditions to be excluded from the classification conditions (also called correction content), and is created in advance by the test target information manager and stored in the storage area 221 of the storage device 22 of the test target information management system 2.

[0099] The pattern generation unit 3515 generates input / output patterns by applying the correction content specified in the correction data 221B to the multiple classification conditions received from the condition analysis unit 3513. Specifically, if one or more additional conditions specified in the correction data 221B are not included in the multiple classification conditions received, the pattern generation unit 3515 adds those one or more additional conditions as new classification conditions. Also, if one or more deletion conditions specified in the correction data 221B are included in the multiple classification conditions received, the pattern generation unit 3515 deletes the classification conditions corresponding to those one or more deletion conditions. The pattern generation unit 3515 outputs the input / output patterns generated by the multiple classification conditions corrected in this way to the pattern matching unit 353.

[0100] Figure 14 is an explanatory diagram illustrating the overview of the processing of the pattern generation unit in Embodiment 2. The processing contents of the pattern generation unit 3515 will be explained in more detail using Figure 14. Figure 14 shows an example in which the input / output pattern identification unit 1351 generates input / output patterns with software that implements a market order management system (hereinafter referred to as the order management system) as the test subject.

[0101] The trade management system under test accepts one or more sell bids and buy bids, and executes a trade processing that determines which sell bid will win against which buy bid. Each sell bid can specify the maximum and minimum quantity to be won (hereinafter referred to as the sell bid quantity). The trade management system receives the sell bid quantities (maximum and minimum quantities) of one or more sell bids and the buy bid quantities of one or more buy bids as input, and calculates the winning quantity of each sell bid for each buy bid as output. Figure 14 shows an example of input / output pattern generation when the trade management system calculates the winning quantity of a single sell bid.

[0102] As a premise of the trade execution management system, in trade execution processing, due to market specifications, the trade quantity cannot fall below the minimum quantity, and if it does fall below, the trade quantity will always be 0. Therefore, whether or not the trade quantity is 0 becomes one of the important classification conditions that constitute the input / output pattern, and conversely, other constants are often not given particular importance. Under this premise, as an example, let's assume that correction data 221B specifying three additional conditions 221B1 (#1 to #3) and one deletion condition 221B2 (#1) has been created in advance, as shown in Figure 14.

[0103] As a result of the processing from the data generation unit 3511 to the condition analysis unit 3513, if the condition analysis unit 3513 extracts the four classification conditions 3513A1 to 3513A4 shown in Figure 14, the pattern generation unit 3515 receives these four classification conditions 3513A1 to 3513A4 from the condition analysis unit 3513. The multiple input and output data generated by the data generation unit 3511 do not necessarily reflect all of the behavior of the test target, so, for example, a value or condition that is not actually important, such as classification condition 3513A3, may be extracted as one of the classification conditions.

[0104] The pattern generation unit 3515 receives correction data 221B from the test target information management system 2. Since condition #1 of the additional conditions #1 to #3 specified in the correction data 221B1 is not included in the classification conditions 3513A1 to 3513A4, it adds a new classification condition 3515A according to condition #1 of additional condition 221B1. Conditions #2 and #3 of additional condition 221B1 are already included as classification conditions 3513A1 and 3513A2, so no additional processing is performed. However, if there are differences in operators (for example, the difference between "<" and "≦") between the existing classification conditions and the additional conditions, the pattern generation unit 3515 may perform the correction.

[0105] Furthermore, the pattern generation unit 3515 deletes the classification condition 3513A3 according to the condition #1 of the deletion condition 521B2 specified in the correction data 221B, as this condition is included as the classification condition 3513A3.

[0106] The pattern generation unit 3515 generates input / output patterns consisting of 16 different patterns by generating a brute-force combination of the four classification conditions 3513A1, 3513A2, 3513A4, and 3515A that have been corrected in this manner, similar to the example shown in Figure 6.

[0107] As explained above, by configuring the pattern generation unit to correct the multiple classification conditions extracted by the condition analysis unit using pre-created correction data, it becomes possible to reduce the omission of necessary classification conditions, the extraction and inclusion of unnecessary classification conditions, and improve the accuracy of the input / output patterns. Furthermore, by dividing the range of values ​​that each output item of the test target can take according to these improved input / output patterns, it is expected that the validity of the division will be further improved.

[0108] In Example 1, we described an example in which the input / output pattern identified by the input / output pattern identification unit 351 is used as is, and the pattern matching unit 353 performs actions such as linking the input / output pattern with test data candidates to generate and output test data 354A. However, the data generation system 3 is not limited to this, and may also estimate the feasibility of the identified input / output pattern and correct the input / output pattern according to the estimation result. Therefore, Example 3 describes an example of a data generation system with such a configuration. Note that in the following description, explanations that overlap with Example 1 will be omitted, and only the differences will be described.

[0109] Figure 15 is a functional block diagram showing an example of the configuration of the data generation system in Embodiment 3, and components identical to those in the data generation system 3 shown in Figure 3 are denoted by the same reference numerals. The data generation system 153 in Embodiment 3 has substantially the same configuration as the data generation system 3 shown in Figure 3, but differs in that it newly includes a feasibility estimation unit 355.

[0110] The feasibility estimation unit 355 receives the result of associating the input / output patterns with test data candidates output by the pattern matching unit 353 (intermediate or final result) as input, and estimates whether there is input / output data corresponding to each pattern included in the input / output pattern. Specifically, for each pattern, when the number of re-executions (repetitions) of the test data candidate generation process by the test data candidate generation unit 352 reaches a predetermined number (also called a threshold), the pattern matching unit 353 estimates that there is no corresponding test data candidate for that pattern, i.e., an impossible pattern to realize. According to the estimation result, the feasibility estimation unit 355 corrects the input / output patterns, i.e., deletes the corresponding pattern or displays the pattern estimated to be impossible to realize via the output device 33, or both.

[0111] The input / output pattern correction may also be performed by the input / output pattern identification unit 351 (for example, the pattern generation unit 3514) based on the notification or instruction of the estimation result from the feasibility estimation unit 355 (Figure 15 shows an example of this case). The corrected input / output pattern is input again to the pattern matching unit 353 from the input / output pattern identification unit 351 or the feasibility estimation unit 355, and processing is performed by the pattern matching unit 353.

[0112] The threshold for the number of re-executions of the test data candidate generation process by the test data candidate generation unit 352 (also called the number of test data candidate generations), which serves as the basis for estimation by the feasibility estimation unit 355, may be a predefined value or determined from the trend of the actual coverage rate. As an example of determining the threshold from the trend of the actual coverage rate, a method can be applied in which the relationship between the number of re-executions and the coverage rate value is modeled and the number of re-executions required to reach a predetermined coverage rate is predicted. For example, a predetermined value can be set in the feasibility estimation unit 355 for the default coverage rate. By using a default coverage rate, the degree to which the test data candidates cover input and output patterns can be maintained at an appropriate level, and as a result, comprehensive test data can be generated.

[0113] Furthermore, the modeling method described above may include approximation techniques using functions such as Nth-degree functions, exponential functions, or logarithmic functions where N is an integer greater than or equal to 1; linear models using multiple regression models or generalized additive models; nonlinear models using polynomial representations, Gaussian process regression, or neural networks; or machine learning techniques that construct predictive models using sequential value analysis models such as moving averages, AR (Auto Regressive) models, or ARIMA (Auto Regressive Integrated Moving Average) models. In addition, the predicted number of retries obtained by the above method may be used as a threshold for the number of retries by adding the expected prediction error range or a predefined value.

[0114] Furthermore, the basis for estimation by the feasibility estimation unit 355 is not limited to the number of times the test data candidate generation process by the test data candidate generation unit 352 is re-executed, but may also be the total processing time when the processes S603 to S605 described in the explanation of Figure 4 are repeated. In this case, a threshold for processing time is determined in advance, and the above estimation by the feasibility estimation unit 355 is performed when the processing time reaches that predetermined threshold.

[0115] Figure 16 is an explanatory diagram illustrating the overview of the processing of the feasibility estimation unit in Embodiment 3. The processing contents of the feasibility estimation unit 355 will be explained in more detail using Figure 16. First, the feasibility estimation unit 355 determines a threshold number of times the test data candidate generation process by the test data candidate generation unit 352 will be re-executed. Figure 16 shows an example in which the feasibility estimation unit 355 predicts the number of times it will take to reach a predetermined coverage rate based on the actual trend of the coverage rate, and determines a threshold number of times the test data candidate generation process will be re-executed.

[0116] Specifically, as shown in the graph at the top of Figure 16, the coverage rate increases more slowly as the number of test data candidate generation cycles increases. This trend is approximated by a curve, and the number of test data candidate generation cycles at the point where the approximation curve intersects with a predetermined coverage rate is determined. In this case, a logarithmic curve may be used as the approximation curve, for example. If the value of the number of test data candidate generation cycles indicated by the intersection point is a decimal, it is converted to an integer value by rounding up to the first decimal place. The feasibility estimation unit 355 uses the number of test data candidate generation cycles obtained in this way as the re-execution count threshold 355A.

[0117] Furthermore, the threshold for the number of retry attempts is not limited to the feasibility estimation unit 355, but may also be determined by other components within the data generation system 153, such as the pattern matching unit 353, or a new component for determining the threshold for the number of retry attempts may be provided within the data generation system 153. Alternatively, it may be determined by, for example, an external information processing device or a user of the data generation system, and then input and set in the data generation system 153.

[0118] In the data generation system 153, the test data candidate generation process by the test data candidate generation unit 352 and the process of linking input / output patterns with test data by the pattern matching unit 353 (processes S603 to S605 shown in Figure 4) are repeated until the re-execution threshold 355A is reached. The lower part of Figure 16 shows an example of the linking result (intermediate or final result) between input / output patterns and test data candidates when the number of test data candidate generation counts reaches the re-execution threshold 355A. This example shows the same linking result as shown in Figures 10 and 11.

[0119] When the number of test data candidate generation cycles reaches the re-execution threshold 355A, the feasibility estimation unit 355 receives the association results (intermediate or final results) between input / output patterns and test data candidates from the pattern matching unit 353, as shown in the figure, and identifies the input / output patterns that were not associated with test data candidates (pairs of input values ​​and expected output values). In the example shown in Figure 16, patterns #3 and #7 among the association results are the ones in question, and the feasibility estimation unit 355 identifies these patterns.

[0120] The feasibility estimation unit 355 either removes the identified patterns (patterns #3 and #7 in Figure 16) from the input / output patterns, displays them via the output device 33 as estimated results of unfeasible patterns, or does both. When displayed via the output device 33, the fields for test data candidates corresponding to each identified pattern may be displayed blank, as shown in the display example in Figure 11, or they may be shaded, filled, or highlighted with an arbitrary color (black in Figure 16), or an arbitrary comment ("No corresponding test data candidates found" in Figure 16) may be added, or a combination of these may be used in the display.

[0121] An example of an impossible pattern being included in the input / output pattern will be explained using the input / output pattern generation example for the market execution management system shown in Figure 14.

[0122] As described in the explanation of Figure 14, the pattern generation unit 3515 generates input / output patterns by generating all possible combinations of the four classification conditions 3515A, 3513A1, 3513A2, and 3513A4. Among these input / output patterns, for example, the pattern "winning bid amount > 0 and winning bid amount ≥ minimum amount and winning bid amount < purchase bid amount and winning bid amount ≤ maximum amount" can actually be realized. On the other hand, for patterns that include the condition "winning bid amount > maximum amount", it is impossible to have a winning bid that satisfies such a pattern due to the specifications of the market. For example, in the calculation of assumed output values ​​using an emulator or simulator by the assumed output calculation unit 3522, assumed output values ​​that satisfy such a pattern are not calculated. Therefore, no matter how many times the test data candidate generation process by the test data candidate generation unit 352 is executed, test data candidates associated with such patterns cannot be generated. Accordingly, such patterns can be deleted as they are impossible to realize.

[0123] As explained above, by providing a feasibility estimation unit in the data generation system, the feasibility of input / output patterns identified by the input / output pattern identification unit can be estimated, and the input / output patterns can be corrected according to the estimation results, thereby improving the accuracy of the input / output patterns. In addition, patterns estimated to be unfeasible may include rare cases in which the corresponding input values ​​and expected output values ​​are extremely small, even though they are actually feasible. Therefore, by displaying the estimation results via the output device, users of the data generation system can discover or recognize such cases, which is expected to contribute to the generation of comprehensive test data.

[0124] In Example 1, the input value generation unit 3521 of the test data candidate generation unit 352 was described as generating input values ​​using a method of random generation within an arbitrary range or a type or range defined in the specification data 221A based on the specification data 221A, or by random generation within each range defined in advance as the search range for input values. However, the method of generating input values ​​is not limited to these, and input values ​​may be generated by dividing the search area according to the distribution of the input and output values ​​of the test target. Therefore, Example 4 will describe an example of a test data candidate generation unit configured in this way. Note that in the following description, explanations that overlap with Example 1 will be omitted, and only the differences will be described.

[0125] Figure 17 is a functional block diagram showing an example of the configuration of the test data candidate generation unit in Embodiment 4, and components identical to those in the test data candidate generation unit 352 shown in Figure 7 are denoted by the same reference numerals. The test data candidate generation unit 1352 in Embodiment 4 has substantially the same configuration as the test data candidate generation unit 352 shown in Figure 7, but differs in that it newly includes an input pattern identification unit 3523.

[0126] Although not shown in the diagram, the input pattern identification unit 3523 has the same configuration as the input / output pattern identification unit 351 shown in Figure 5, namely, a data generation unit, a data classification unit, a condition analysis unit, and a pattern generation unit. It generates input and output data, classifies it into clusters, constructs a classification model and extracts classification conditions, and generates patterns by exhaustively trying all combinations of classification conditions, and outputs them to the input value generation unit 3521. However, the condition analysis unit constructs a classification model and extracts classification conditions using the values ​​of each item in the input values ​​and the relative magnitudes of the inputs as features, and the pattern generation unit generates input patterns based on these classification conditions.

[0127] Similar to the example shown in Figure 12B, the range of possible values ​​for each input item under test can be divided according to the input pattern generated by the pattern generation unit. The input value generation unit 3521 receives the input pattern from the input pattern identification unit 3523 and generates at least one input value belonging to each region divided by the input pattern.

[0128] As explained above, by providing an input pattern identification unit in the test data candidate generation unit, it becomes possible to generate input values ​​that conform to the distribution of input and output values ​​of the test target. As a result, the test data candidate generation unit can efficiently generate input values ​​for both typical and rare cases, thereby reducing the number of times the test data candidate generation process is re-executed by the test data candidate generation unit, and a reduction in the processing time required for comprehensive test data generation can be expected.

[0129] In Example 1, we described an example in which the data generation system 3 outputs the finally generated test data 354A to the test execution system 4 to complete the test data generation process (no further processing is performed thereafter). However, the system is not limited to this example; the data generation system 3 may also receive the test execution results from the test execution system 4, generate and output additional test data based on those results, and have the test execution system 4 run the test again. Therefore, Example 5 describes an example of a data generation system with such a configuration. Note that in the following description, explanations that overlap with Example 1 will be omitted, and only the differences will be explained.

[0130] Figure 18 is a functional block diagram showing an example of the configuration of the data generation system in Embodiment 5, and components identical to those in the data generation system 3 shown in Figure 3 are denoted by the same reference numerals. The data generation system 183 in Embodiment 5 has substantially the same configuration as the data generation system 3 shown in Figure 3, but differs in that it is equipped with a test data candidate generation unit 2352 instead of a test data candidate generation unit 352, and receives test execution results from the test execution system 4.

[0131] The test data candidate generation unit 2352 receives test execution results as feedback from the test execution system 4. The feedback includes, for example, the input / output data (hereinafter referred to as "defect detection location") in the test data 354A where a value different from the expected output value is output from the test target, or the value output by the test target. The test execution system 4 may select and output such feedback, or the test execution system 4 may output all test execution results and the test data candidate generation unit 2352 may extract the above-mentioned defect detection locations as feedback from the test execution results it receives.

[0132] The test data candidate generation unit 2352 identifies one or more input / output data points indicated as defect detection locations in the received feedback, and generates one or more new input / output data points by adding a disturbance to the input value of each of those input / output data points and calculating a new assumed output value, thereby generating new test data candidates composed of these new input / output data points. In this embodiment, adding a disturbance means adjusting the input value to a value before or after it (more specifically, shifting the input value forward or backward), for example, if the input value is "0", it means changing that value to "1" or "-1".

[0133] The test data candidate generation unit 2352 may generate multiple input values ​​by setting the input value items for the defect detection location to be disturbed to one or more items at a time (while fixing the values ​​of other items). The disturbance value may be a value obtained by multiplying the smallest step size to which the input value to be disturbed can be changed by an arbitrary constant, and the sign may be positive or negative (for example, if it is an integer value, the smallest step size will be "1", and the constant multiple of that value will be "1", "-1", "2", or "-2"). The number of newly generated input and output data may be an arbitrarily predefined number, a number adjusted so that the test execution time by the test execution system 4 falls within a predetermined value, or the number of times to change one or more arbitrary or specific input values ​​by the above step size until those or any of the input values ​​reach a specific size (or smallness).

[0134] As explained above, the test data candidate generation unit receives information on the location of defects detected in the test execution results from the test execution system and generates new test data candidates. The data generation system then generates and outputs additional test data, causing the test execution system to run the test again. This makes it possible to understand in detail the behavior (changes in output values, etc.) near the location of the defect detected under test, and is expected to shorten the time required for the test executor to investigate the cause.

[0135] In Example 1, an example was described in which the input / output pattern identification unit 351 identifies the input / output pattern based on the specification data 221A and outputs it to the pattern matching unit 353 to complete the process. However, the system is not limited to this, and the results of test data candidate generation by the test data candidate generation unit 352 may be added to the data for input / output pattern identification, and the input / output pattern identification unit 351 may update the input / output pattern. Therefore, Example 6 describes an example of a data generation system with such a configuration. Note that in the following description, explanations that overlap with Example 1 will be omitted, and only the differences will be described.

[0136] Figure 19 is a functional block diagram showing an example of the configuration of the data generation system in Embodiment 6, and components identical to those in the data generation system 3 shown in Figure 3 are denoted by the same reference numerals. The data generation system 193 in Embodiment 6 has substantially the same configuration as the data generation system 3 shown in Figure 3, but differs in that it is equipped with an input / output pattern identification unit 2351 instead of the input / output pattern identification unit 351, and receives test data candidates from the test data candidate generation unit 352.

[0137] Although not shown in the diagram, the input / output pattern identification unit 2351 has the same configuration as the input / output pattern identification unit 351 shown in Figure 5 and performs the same input / output pattern identification process. Then, the data generation unit in the input / output pattern identification unit 2351 receives test data candidates (multiple input / output data) from the test data candidate generation unit 352 and adds the received test data candidates to the multiple input / output data it has generated. Based on the multiple input / output data to which the test data candidates have been added, the data classification unit, condition analysis unit, and pattern generation unit perform the series of processes again to identify input / output patterns using the test data candidates and update the input / output patterns.

[0138] Furthermore, if there is a difference (such as an increase or decrease in patterns or a change in conditions within a pattern) between the input / output patterns using the test data candidates and the previously identified input / output patterns (input / output patterns identified without using the test data candidates), the pattern generation unit may correct the previously identified input / output patterns based on that difference and update the input / output patterns.

[0139] Furthermore, in the example shown in Figure 19, as described above, the input / output pattern identification unit 2351 is configured to receive test data candidates generated by the test data candidate generation unit 352, but it may also receive test data 354A and perform the above processing.

[0140] As explained above, the input / output pattern identification unit can improve the accuracy of the input / output pattern by updating the input / output pattern using test data candidates or test data. Furthermore, by dividing the range of values ​​that each input / output under test can take according to the input / output pattern with improved accuracy in this way, it is expected that the validity of the division will be further improved. Note that a configuration similar to this embodiment, that is, a configuration that updates the input pattern using test data candidates or test data, may also be applied to the input pattern identification unit 3523 in Example 4 (shown in Figure 17).

[0141] In Example 1, the data generation system 3 output test data 354A by directly linking input / output patterns with test data candidates. However, it is not limited to this; it is also possible to generate explanatory text (for example, item names, sentences, etc. in natural language, hereinafter referred to as explanatory text) to describe the content of the input / output patterns and add it to the test data 354A before outputting it. Example 7 describes an example of a data generation system with such a configuration. Note that in the following description, explanations that overlap with Example 1 will be omitted, and only the differences will be described.

[0142] Figure 20 is a functional block diagram showing an example of the configuration of the data generation system in Embodiment 7, and components identical to those in the data generation system 3 shown in Figure 3 are denoted by the same reference numerals. The data generation system 203 in Embodiment 7 has substantially the same configuration as the data generation system 3 shown in Figure 3, but differs in that it newly includes an input / output pattern explanation unit 356.

[0143] The input / output pattern description unit 356 receives input / output patterns from the input / output pattern identification unit 351 and description data 221C from the test target information management system 2 as input. Description data 221C is data that associates each input / output item with its semantic information (item name expressed in natural language, etc.). In Figure 20, the description data 221C is shown separately from the specification data 221A in the storage area 221 of the test target information management system 2, but the contents of the description data 221C may be included in the specification data 221A (even if the description data 221C is part of the specification data 221A). If the description data 221C is data separate from the specification data 221A, the description data 221C is defined or created in advance by the test target information manager, etc., and stored in the storage area 221 of the storage device 22 provided in the test target information management system 2.

[0144] The input / output pattern description unit 356 describes each input / output item in the input / output pattern (for example, OUT in the input / output pattern 3514A shown in Figure 6) based on the description data 221C received by the input / output pattern description unit 356. A OUT BThe system performs processes such as replacing terms (etc.) with semantic information in the explanatory data 221C, and converting descriptions of operators or mathematical formulas into natural language, thereby generating a natural language-based explanatory text about the contents of the input / output pattern. Alternatively, the system may use a natural language handling model, such as a large-scale language model, to perform the above-mentioned conversion processes and generate the explanatory text. The input / output pattern explanatory unit 356 adds the generated explanatory text to the test data 354A.

[0145] In Figure 20, an example is shown where the input / output pattern explanation unit 356 receives input / output patterns from the input / output pattern identification unit 351. However, the system is not limited to this, and input / output patterns may also be received from the pattern matching unit 353 or the storage area 354.

[0146] As explained above, by providing an input / output pattern description section in the data generation system and generating natural language descriptions of the input / output patterns and attaching them to the test data, for example, users of the data generation system can easily understand the content of the input / output patterns from the descriptions and determine the feasibility of each pattern. This is expected to shorten the time required to generate test data and improve the accuracy of the input / output patterns.

[0147] In Example 1, an example was described in which the data generation unit 3511 of the input / output pattern identification unit 351 generates multiple input / output data. However, the example is not limited to this, and the input / output pattern identification unit 351 may identify an input / output pattern using multiple predefined input / output data. Therefore, Example 8 describes an example of an input / output pattern identification unit configured in this way. Note that in the following description, explanations that overlap with Example 1 will be omitted, and only the differences will be described.

[0148] Figure 21 is a functional block diagram showing an example of the configuration of the input / output pattern identification unit in Embodiment 8, and components identical to those in the input / output pattern identification unit 351 shown in Figure 5 are denoted by the same reference numerals. The input / output pattern identification unit 3351 in Embodiment 8 has substantially the same configuration as the input / output pattern identification unit 351 shown in Figure 5, but differs in that it does not include a data generation unit 3511.

[0149] The data classification unit 3512 in the input / output pattern identification unit 3351 receives input / output examples 221D along with specification data 221A from the test target information management system 2 as input. Input / output examples 221D are multiple input / output data defined or created in advance by the test target information manager, etc., and are stored in the storage area 221 of the storage device 22 provided in the test target information management system 2. The data classification unit 3512 performs a cluster classification process according to the degree of similarity of each input / output data in the received input / output examples 221D. The processing in the input / output pattern identification unit 3351 from this point onward is the same as the processing in the input / output pattern identification unit 351 described in the explanation of Figure 5.

[0150] As explained above, the input / output pattern identification unit generates input / output patterns using input / output examples defined or created in advance by the test target information manager, etc., thereby preventing the risk of operational errors in emulators, etc., used in the data generation unit 3511 for calculating expected output values, etc. Furthermore, since the input / output example 221D is created by the test target information manager, etc., it can be expected to be highly reliable input / output data, and an improvement in the accuracy of the input / output patterns can be expected.

[0151] Although embodiments of the present invention have been described in detail above, the present invention is not limited to the embodiments described above, and various design modifications can be made without departing from the spirit of the invention as described in the claims. For example, each of the embodiments described above is described in detail in order to explain the present invention in an easy-to-understand manner, and is not necessarily limited to having all of the described configurations. Furthermore, it is possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add a configuration from another embodiment to the configuration of one embodiment. In addition, it is possible to add, delete, or replace a part of the configuration of each embodiment with other configurations.

[0152] 1...Data processing system 2...Test target information management system 3, 153, 183, 193, 203...Data generation system 4...Test execution system 5...Communication path 21, 34, 41...Communication device 22, 35...Storage device 31...CPU 32...Input device 33...Output device 42...Test execution device 221, 354...Storage area 351, 1351, 2351, 3351...Input / output pattern identification unit 352, 1352, 2352...Test data candidate generation unit 353...Pattern matching unit 355...Feasibility estimation unit 356...Input / output pattern explanation unit 3511...Data generation unit 3512...Data classification unit 3513...Condition analysis unit 3514, 3515...Pattern generation unit 3521...Input value generation unit 3522...Expected output calculation unit 3523...Input pattern identification unit 3531...Data merging unit 3532...Comprehensive determination unit

Claims

1. A data generation system for generating test data for software testing, comprising: an input / output pattern identification unit that generates input / output patterns including a plurality of patterns consisting of combinations of classification conditions for the test data; a test data candidate generation unit that generates test data candidates including a plurality of input / output data which are pairs of input values ​​and expected output values, based on specification data of the software to be tested; a pattern matching unit that associates each of the patterns included in the input / output pattern with each of the input / output data included in the test data candidate, and generates test data including the plurality of input / output data associated with each of the patterns, wherein the input / output pattern identification unit comprises: a data generation unit that generates a plurality of input / output data based on the specification data; a data classification unit that classifies each of the input / output data into a plurality of clusters based on the similarity between the input value and the expected output value included in the input / output data; a condition analysis unit that extracts a plurality of classification conditions for classifying each of the input / output data into each of the clusters from the results of the classification; and a pattern generation unit that generates the input / output pattern from a brute-force combination of the extracted plurality of classification conditions.

2. A data generation system according to claim 1, wherein the test data candidate generation unit comprises: an input value generation unit that generates a plurality of input values ​​for software testing for each of a plurality of input items of the software based on the specification data; and an assumed output calculation unit that calculates a plurality of assumed output values ​​for each of the input values ​​of each of the input items based on the specification data.

3. A data generation system according to claim 1, comprising: a pattern matching unit that determines which of the patterns included in the input / output pattern each of the input / output data included in the test data candidate corresponds to, and a data linking unit that links each of the input / output data to the corresponding pattern; and a coverage determination unit that calculates a coverage rate that shows the ratio of the number of patterns to which the input / output data is linked among the input / output patterns to the total number of patterns in the input / output pattern, and determines whether or not the test data candidate generation unit needs to regenerate the test data candidate according to the coverage rate.

4. A data generation system according to claim 1, wherein the pattern generation unit corrects the extracted plurality of classification conditions using correction data that specifies one or more necessary or unnecessary classification conditions, and generates the input / output pattern from a brute-force combination of the corrected plurality of classification conditions.

5. A data generation system according to claim 3, further comprising a feasibility estimation unit, wherein the pattern matching unit causes the test data candidate generation unit to repeatedly generate the test data candidates if the coverage rate is less than a predetermined threshold, and the feasibility estimation unit estimates that one or more of the patterns to which none of the input / output data are associated are impossible to realize when the number of times the test data candidates are generated by the test data candidate generation unit reaches a predetermined number.

6. A data generation system according to claim 2, wherein the test data candidate generation unit further comprises an input pattern identification unit, the input pattern identification unit is configured in the same manner as the input / output pattern identification unit, and generates an input pattern from a brute-force combination of a plurality of classification conditions for classifying each of the input / output data based on the input value contained in each of the input / output data.

7. A data generation system according to claim 2, wherein the test data candidate generation unit receives the results of the software test from an external source, and generates one or more new input / output data by adding a disturbance to the input value and calculating a new expected output value with respect to the input / output data for which a value different from the expected output value for the input value shown in the results of the software test has been output, thereby generating new test data candidates.

8. A data generation system according to claim 1, wherein the data generation unit receives the candidate test data or the test data, and adds a plurality of the input / output data contained in the candidate test data or the test data to a plurality of the input / output data generated based on the specification data to generate a plurality of the input / output data.

9. A data generation system according to claim 1, further comprising an input / output pattern explanation unit, wherein the input / output pattern explanation unit generates a natural language explanation of the content of the input / output pattern and adds it to the test data.

10. A data generation system according to claim 1, wherein the data classification unit receives an input / output example from an external source that includes a plurality of predefined input / output data, and classifies each of the input / output data included in the input / output example.

11. A data generation method for generating test data for software testing, comprising: an input / output pattern identification step of generating an input / output pattern including a plurality of patterns consisting of combinations of classification conditions for the test data; a test data candidate generation step of generating a test data candidate including a plurality of input / output data which are pairs of input values ​​and expected output values, based on specification data of the software to be tested; and a pattern matching step of associating each of the patterns included in the input / output pattern with each of the input / output data included in the test data candidate, and generating the test data including the plurality of input / output data associated with each of the patterns, wherein the input / output pattern identification step generates a plurality of input / output data based on the specification data, classifies each of the input / output data into a plurality of clusters based on the similarity between the input value and the expected output value included in the input / output data, extracts a plurality of classification conditions for classifying each of the input / output data into each of the clusters from the result of the classification, and generates the input / output pattern from a brute-force combination of the extracted plurality of classification conditions.

12. A data generation method according to claim 11, wherein in the test data candidate generation step, a plurality of input values ​​for software testing are generated for each of the plurality of input items of the software based on the specification data, and a plurality of assumed output values ​​are calculated for each of the input values ​​of each of the input items based on the specification data.

13. A data generation method according to claim 11, wherein in the pattern matching step, it is determined which of the patterns included in the input / output pattern each of the input / output data included in the test data candidate corresponds to; each of the input / output data is associated with the corresponding pattern; a coverage rate is calculated, which is the ratio of the number of patterns to which the input / output data is associated among the input / output patterns to the total number of patterns in the input / output pattern; and a data generation method is determined according to the coverage rate, wherein it is necessary to re-execute the test data candidate generation step.

14. A data generation program that causes an information processing system, comprising at least a CPU and a memory device, to execute a process for generating test data for software testing, comprising: an input / output pattern identification process that generates an input / output pattern including a plurality of patterns consisting of combinations of classification conditions for the test data; a test data candidate generation process that generates a plurality of test data candidates including a plurality of input / output data which are pairs of input values ​​and expected output values, based on specification data of the software to be tested; a pattern matching process that associates each of the patterns included in the input / output pattern with each of the input / output data included in the test data candidate, and generates the test data including the plurality of input / output data associated with each of the patterns; further, in the input / output pattern identification process, a process that generates a plurality of input / output data based on the specification data; a process that classifies each of the input / output data into a plurality of clusters based on the similarity between the input value and the expected output value included in the input / output data; a process that extracts a plurality of classification conditions from the results of the classification for classifying each of the input / output data into each of the clusters; and a process that generates the input / output pattern from a brute-force combination of the extracted plurality of classification conditions, the data generation program that causes the information processing system to execute these processes.