Test data generation method and device, computer equipment and storage medium

By receiving user requests and analyzing rule list data, the data simulation device is generated, and the problems of low generation efficiency, poor security and insufficient flexibility in the prior art are solved, and fast, safe and easy-to-use test data generation is realized, and a variety of database environments are adapted to.

CN120448275APending Publication Date: 2025-08-08YGSOFT INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510643604.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing test data generation methods and technologies have problems such as low generation efficiency, poor security and insufficient flexibility, which are difficult to meet the needs of agile development and diversified testing, and have high requirements for users' technical background.

Method used

It provides a test data generation method, which can analyze rule list data by receiving requests from user terminals, generate data simulation devices, supports multiple database types, provides intuitive configuration interface and intelligent rule analysis, reduces manual configuration workload, and supports data rollback, automatic backup and concurrent simulation and other auxiliary functions.

Benefits of technology

It realizes the rapid generation of test data that meets business needs, improves generation efficiency and security, lowers the threshold for user use, supports multiple database environments, and enhances the reusability and ease of use of configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448275A_ABST
    Figure CN120448275A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the technical field of software testing of artificial intelligence, and relates to a test data generation method and device, computer equipment and a storage medium, the method comprises the steps that a test data generation request sent by a user terminal is received, the test data generation request comprises table list data and rule list data, and the table list data and the rule list data are sent to the user terminal; the rule list data is used for guiding generation of simulation data; analyzing the rule list data to obtain test data generation logic; generating a logic generation data simulation device according to a configured simulation environment, the table list data and the test data; and executing the data simulation device to obtain target test data. The problems that an existing test data generation method and technology are low in generation efficiency, poor in safety and insufficient in flexibility are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence software testing technology, and in particular to a test data generation method, apparatus, computer equipment, and storage medium. Background Art

[0002] In the field of software testing, test data is the most basic and core input element. Its effectiveness, diversity and data size directly affect the efficiency and quality of software testing.

[0003] However, existing test data generation methods and technologies have many problems and limitations in practical applications, as shown below:

[0004] 1. Data security risks

[0005] Directly using real user data for testing carries a high data security risk, especially when sensitive information (such as personal privacy and business secrets) is involved. For software vendors, directly using customer data may not only violate relevant laws and regulations (such as GDPR and CCPA), but may also lead to data leaks, causing immeasurable losses to businesses and users.

[0006] 2. Difficulty in sourcing test data

[0007] In actual testing, it's often difficult to obtain ideal test data. For example, data in certain business scenarios may involve complex logical relationships or specific data distributions, and existing technologies struggle to quickly generate test data that meets these requirements. This results in insufficient test coverage and compromises the accuracy and reliability of test results.

[0008] 3. Heavy R&D workload

[0009] Existing test data generation solutions typically require the development of dedicated data simulation devices for each data table or business scenario. This not only increases the workload for R&D personnel, but also extends development cycles and increases costs. Furthermore, this customized development approach struggles to adapt to rapidly changing business needs.

[0010] 4. Diversified application scenarios

[0011] With the increasing complexity of software systems and the diversification of business scenarios, test data needs to be highly flexible and configurable to meet the testing requirements of different scenarios. However, existing technologies often lack flexible configuration mechanisms and are difficult to quickly adapt to diverse test scenarios.

[0012] 5. Fast delivery requirements

[0013] In the context of agile development and continuous integration / continuous delivery (CI / CD), rapidly building simulation test environments and generating batch test data have become important requirements. However, existing technologies often fail to meet these requirements, resulting in low testing efficiency and hindering overall development progress.

[0014] 6. High technical threshold

[0015] Existing test data generation tools often require users to have a certain technical background and even require secondary development or customized configuration. This increases the difficulty and learning cost for non-technical users or testers, and reduces the universality and ease of use of the tools.

[0016] It can be seen that the existing test data generation methods and technologies have the problems of low generation efficiency, poor security and insufficient flexibility. Summary of the Invention

[0017] The purpose of the embodiments of the present application is to propose a test data generation method, apparatus, computer equipment and storage medium to solve the problems of low generation efficiency, poor security and insufficient flexibility in existing test data generation methods and technologies.

[0018] In order to solve the above technical problems, the present application provides a test data generation method, which adopts the following technical solutions:

[0019] Receiving a test data generation request sent by a user terminal, wherein the test data generation request includes table list data and rule list data, and the rule list data is used to guide the generation of simulation data;

[0020] Performing a parsing operation on the rule list data to obtain test data generation logic;

[0021] generating a data simulation device according to the configured simulation environment, the table list data, and the test data generation logic;

[0022] The data simulation device is executed to obtain target test data.

[0023] In order to solve the above technical problems, the embodiment of the present application further provides a test data generation device, which adopts the following technical solution:

[0024] a request receiving module, configured to receive a test data generation request sent by a user terminal, wherein the test data generation request includes table list data and rule list data, and the rule list data is used to guide the generation of simulation data;

[0025] A rule parsing module, configured to parse the rule list data to obtain test data generation logic;

[0026] A simulation device generation module, configured to generate a data simulation device according to the configured simulation environment, the table list data and the test data generation logic;

[0027] The simulation device execution module is used to execute the data simulation device to obtain target test data.

[0028] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0029] The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the test data generating method described above.

[0030] The present application provides a test data generation method, comprising: receiving a test data generation request sent by a user terminal, wherein the test data generation request includes table list data and rule list data, the rule list data being used to guide the generation of simulation data; performing a parsing operation on the rule list data to obtain test data generation logic; generating a data simulation device according to a configured simulation environment, the table list data, and the test data generation logic; and executing the data simulation device to obtain target test data. Compared with the prior art, the present application uses an intuitive configuration interface and intelligent rule parsing to enable users to easily generate test data without having a deep technical background. In addition, the present application uses automated rule parsing and device generation to quickly generate test data that meets business needs. The present application supports multiple database types (such as MySQL, Oracle, MongoDB, etc.), and users only need to configure once to use it in multiple database environments, significantly improving the reusability of the configuration. The present application automatically generates about 80% of the configuration data through model information, reducing the workload and error rate of manual configuration. The present application provides auxiliary functions such as data rollback, automatic backup, concurrent simulation, and data refresh, greatly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0032] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0033] Figure 2This is a flow chart of the test data generation method provided in the embodiment of the present application;

[0034] Figure 3 It is a structural diagram of the test data generating device provided in an embodiment of the present application;

[0035] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0037] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0038] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0039] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0040] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0041] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.

[0042] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .

[0043] It should be noted that the test data generating method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the test data generating device is generally set in the server / terminal device.

[0044] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0045] Continue to refer Figure 2 , shows a flow chart of an embodiment of a test data generation method according to the present application. The test data generation method includes: step S201, step S202, step S203, step S204 and step S205.

[0046] In step S201 , a test data generation request sent by a user terminal is received, wherein the test data generation request includes table list data and rule list data, and the rule list data is used to guide the generation of simulation data.

[0047] In the embodiments of the present application, the user terminal refers to a terminal device used to execute the image processing method for preventing document abuse provided by the present application. The user terminal can be a mobile terminal such as a mobile phone, a smart phone, a laptop computer, a digital broadcast receiver, a PDA (personal digital assistant), a PAD (tablet computer), a PMP (portable multimedia player), a navigation device, etc., as well as a fixed terminal such as a digital TV, a desktop computer, etc. It should be understood that the examples of user terminals here are only for convenience of understanding and are not used to limit the present application.

[0048] In the embodiments of this application, this application supports a variety of mainstream database systems such as MySQL, Oracle, and DAMO, meeting the data simulation needs of different users in diverse database environments. Furthermore, to ensure data security and privacy, this application features encrypted storage of data source configuration information, ensuring the secure storage of user configuration information, preventing the risk of sensitive data leakage, and providing reliable protection for user data management.

[0049] In the embodiment of the present application, the deployment simulation list configuration table includes: table name and expected data size, specifically:

[0050] Table name: the physical name of the table to be simulated;

[0051] Expected data size: The amount of data that the current table will reach.

[0052] In the embodiment of the present application, the deployment simulation rule configuration table includes: name, field name, rule type and rule definition, specifically:

[0053] Name: The physical name of the table to be simulated;

[0054] Field Name: The physical name of the field for which simulation rules are to be defined;

[0055] Rule type: supports three rule types: unique key, foreign key, and constant;

[0056] Rule definition: Unique keys do not need to be defined, foreign keys need to specify dependent foreign key information, and constants need to specify constant generation rules.

[0057] In the embodiment of the present application, configuration management allows users to flexibly adjust the configuration according to the simulation target and simulation environment. Specifically:

[0058] Simulation target management: Users can define the simulation target database, table structure, and data size according to actual needs. The system supports multiple database types and complex data models to ensure the applicability of simulation tasks;

[0059] Simulation environment configuration: users can specify the running environment of simulation tasks, including hardware resource allocation, network configuration, execution priority, etc., to optimize task execution efficiency;

[0060] Simulation list management: Users can dynamically add and delete simulation lists, flexibly defining which tables need to participate in simulation, and which tables need to pause or remove simulation to adapt to different business scenarios;

[0061] Simulation rule management: Users can configure field generation rules based on business needs, including unique keys, foreign keys, constants, random values, data formats, etc. The system supports dynamic adjustment of rules to ensure the accuracy and consistency of simulation data.

[0062] In step S202, the rule list data is parsed to obtain test data generation logic.

[0063] In step S203, a data simulation device is generated according to the configured simulation environment, table list data, and test data generation logic.

[0064] In step S204, the data simulation device is executed to obtain target test data.

[0065] In the embodiments of this application, the technical solution of this application mainly includes six core modules: data simulation list configuration, data simulation rule configuration, data simulation rule parsing, data simulation device generation, data simulation device execution, and auxiliary functions. The specific implementation and function of each module are described in detail below:

[0066] 1. Data simulation list configuration module

[0067] This module allows users to enter a list of tables to be simulated and specify the expected data size (e.g., number of rows, amount of data, etc.). This module clarifies the scope and objectives of test data generation, providing the foundation for subsequent rule configuration and data generation.

[0068] 2. Data simulation rule configuration module

[0069] This module allows users to self-configure field generation rules, including unique key rules, foreign key constraints, constant value settings, data type definitions, and business logic rules. This module ensures that the generated test data accurately reflects the data characteristics and logical relationships in real business scenarios.

[0070] 3. Data simulation rule parsing module

[0071] This module automatically parses the business rules defined by users in the rule configuration module and converts them into computer-executable logical instructions. During the parsing process, the system performs rule validation and conflict detection to ensure the legitimacy and consistency of the rules. The module's role is to transform user-configured rules into actionable data generation logic.

[0072] 4. Data simulation device generation module

[0073] Based on the user-configured simulation environment, target data size, and parsed rules, this module automatically generates a data simulation device. This data simulation device includes a data generation device, data insertion logic, and a data validation mechanism. This module automates and standardizes test data generation, reducing manual intervention.

[0074] 1) This device can automatically collect unique keys, foreign keys, constants, existing data size and other information based on the data model information of the simulation environment as input for data simulation rules;

[0075] 2) This device can automatically calculate the difference in data volume based on the user-configured data scale and the existing data scale, which is used as the quantitative target for this data simulation. Therefore, users only need to pay attention to the expected data scale and do not need to pay attention to the actual data scale. At the same time, it also avoids the phenomenon of repeated simulation data in the case of multiple users;

[0076] 3) Based on the above information and user configuration information, the device automatically generates data simulation rules, and automatically generates executable data simulation units based on the rules, which are then handed over to the execution unit for execution, ultimately generating simulation data.

[0077] 5.Data simulation device execution module

[0078] This module automatically executes the generated data simulation device and records detailed logs during the execution process (such as data generation progress, error messages, etc.). During execution, the system supports concurrent data generation to improve efficiency. The role of this module is to ensure that test data is generated and stored in the target database efficiently and accurately.

[0079] 6. Auxiliary function module

[0080] This module provides a series of user-friendly auxiliary functions, including simulated data rollback, automatic backup, concurrent simulation, and data refresh. The data rollback function restores the database to its initial state after testing; the automatic backup function prevents data loss; the concurrent simulation function supports multi-threaded data generation, improving efficiency; and the data refresh function quickly updates test data to adapt to new testing requirements. This module aims to enhance the system's usability and flexibility to meet diverse testing needs.

[0081] 1) Simulation data rollback means that if the user finds that the simulation data is no longer needed due to configuration errors or after the test is completed, the simulation data can be rolled back with one click. That is, the relevant tables can be rolled back to the state before the simulation data was added.

[0082] 2) Automatic backup means that if the user wants to back up the simulated table before simulation, the user does not need to back up the table by himself. He only needs to input the backup signal and the device will automatically back up the relevant table during the simulation process;

[0083] 3) Concurrent simulation means that if the user simulates multiple tables at one time, a concurrent strategy can be enabled to perform multi-threaded parallel simulation, thereby greatly reducing simulation time and ensuring delivery progress;

[0084] Data refresh means that during concurrent simulation, if the user wishes to modify the simulation rules for a table that has not yet been simulated, they can directly modify and issue a refresh configuration command without stopping the ongoing simulation. The simulation device will automatically read the latest rules for the relevant table and perform data simulation without affecting the simulation of other tables. If the user still wants to simulate according to the old rules after modifying the configuration rules, the simulation can proceed according to the old rules without issuing a refresh command.

[0085] In the embodiments of this application, modular design and intelligent rule parsing are used to achieve efficient generation and management of simulation data. Compared with the existing technology, this application has the following significant advantages:

[0086] 1) Efficiency: Through automated rule parsing and device generation, the efficiency of test data generation is significantly improved.

[0087] 2) Flexibility: Supports user-defined rules and can adapt to a variety of complex business scenarios.

[0088] 3) Security: Avoid using real data to ensure data security.

[0089] 4) Ease of use: Provide user-friendly auxiliary functions to lower the user threshold.

[0090] In an embodiment of the present application, a test data generation method is provided, comprising: receiving a test data generation request sent by a user terminal, wherein the test data generation request includes table list data and rule list data, and the rule list data is used to guide the generation of simulation data; parsing the rule list data to obtain test data generation logic; generating a data simulation device according to the configured simulation environment, the table list data, and the test data generation logic; executing the data simulation device to obtain target test data. Compared with the prior art, the present application uses an intuitive configuration interface and intelligent rule parsing to enable users to easily generate test data without having a deep technical background. In addition, the present application uses automated rule parsing and device generation to quickly generate test data that meets business needs; the present application supports multiple database types (such as MySQL, Oracle, MongoDB, etc.), and users only need to configure once to use it in multiple database environments, significantly improving the reusability of the configuration; the present application automatically generates about 80% of the configuration data through model information, reducing the workload and error rate of manual configuration; the present application provides auxiliary functions such as data rollback, automatic backup, concurrent simulation, and data refresh, which greatly improves the user experience.

[0091] In some optional implementations of the embodiments of the present application, after the step of receiving the test data generation request sent by the user terminal and before the step of parsing the rule list data to obtain the test data generation logic, the following steps are further included:

[0092] Perform data cleaning, data standardization, and data formatting operations on the rule list data.

[0093] In the embodiments of this application, data cleaning is a key step in preprocessing the rule list data, aiming to remove noise, errors, and inconsistencies in the data to improve data quality and the accuracy of subsequent processing. The following are the main contents of the data cleaning operation:

[0094] 1. Remove duplicate data

[0095] Note: Duplicate rule records may exist in the rule list data. Duplicate data increases the complexity of data processing and may cause redundancy during test data generation. Deduplication ensures the uniqueness of each rule.

[0096] Example: Suppose there are two identical rules in the rule list, "User login password length must be at least 8 characters", and only one is retained after deduplication.

[0097] 2. Handling Missing Values

[0098] Operational instructions: Some fields may be missing in the rule data, such as rule descriptions, constraints, etc. Missing values will affect the model's understanding and parsing of the rules and need to be processed.

[0099] Treatment method:

[0100] Deleting missing records: If the missing fields are critical to rule parsing and the missing percentage is small, you can consider deleting records containing missing values.

[0101] Filling missing values: Based on business rules and context, missing values can be filled using methods such as default values, mean, median, or mode. For example, for missing numeric ranges, a reasonable default range can be filled based on common business knowledge.

[0102] 3. Correcting incorrect data

[0103] Operation instructions: The rule data may contain spelling errors, grammatical errors, logical errors, etc. These errors will affect the model's correct interpretation of the rules and need to be corrected.

[0104] Example: The rule "user age must be greater than 18 years old and less than or equal to 60 years old" is incorrectly written as "user age must be greater than 18 years old and less than 60 years old" (missing the equals case) and needs to be corrected.

[0105] 4. Dealing with Outliers

[0106] Operational Note: The rule data may contain some abnormal values that do not conform to business logic or actual conditions. For example, the rule may contain unreasonable value ranges or invalid dates.

[0107] Treatment method:

[0108] Deleting outliers: If outliers have a significant impact on the overall data and cannot be corrected, consider deleting them.

[0109] Correcting outliers: Correcting outliers based on business rules and actual conditions. For example, adjusting an unreasonable value range to a reasonable range.

[0110] In the embodiment of this application, data standardization is to convert the rule list data into a unified standard format to facilitate model processing and analysis. The following are the main contents of the data standardization operation:

[0111] 1. Unified terminology and nomenclature

[0112] Note: In rule list data, different terms and naming methods may be used to describe the same concept. For example, "user ID" may be written as "user number" or "user ID." Unifying terminology and naming can avoid ambiguity in the model's understanding of concepts.

[0113] Example: All fields representing user identities are uniformly named "user_id".

[0114] 2. Standardize numerical format

[0115] Operation Instructions: For numerical data in rules, you need to standardize its format, such as the number of decimal places, units, etc. For example, all currency amounts can be standardized to "yuan" as the unit and retain two decimal places.

[0116] Example: The rule "Product price does not exceed 1000" can be standardized as "Product price does not exceed 1000.00 yuan."

[0117] 3. Standardize date and time formats

[0118] Operation Note: If the rule involves date and time, they need to be standardized into a unified format, such as "YYYY-MM-DD" or "YYYY-MM-DD HH:MM:SS".

[0119] Example: Normalize "October 1, 2023" to "2023-10-01".

[0120] 4. Unified logical expression

[0121] Operation Instructions: The logical expressions in the rules may exist in various forms, such as "and", "and", "or", "or", etc. Unifying the logical expressions can avoid confusion when the model parses the logical relationships.

[0122] Example: All logical "ands" are represented by "AND", and all logical "or" are represented by "OR".

[0123] In the embodiment of the present application, data formatting is to convert the cleaned and standardized rule list data into a specific format to facilitate model input and subsequent processing. The following are the main contents of the data formatting operation:

[0124] 1. Convert to structured data format

[0125] Operation Instructions: Convert the rules described in natural language into a structured data format, such as JSON, XML, or CSV. Structured data formats facilitate model parsing and processing.

[0126] 2. Serialize data

[0127] Operation instructions: For some data that needs to be input into a deep learning model, serialization operations may be required to convert it into a vector or tensor form that the model can process.

[0128] Example: Use word embedding technology to convert each word in the rule text into a vector, and then represent the entire rule text as a vector sequence.

[0129] 3. Add metadata

[0130] Operation Instructions: Add metadata to the formatted data, such as the source, creation time, and modification time of the rule. Metadata can help track and manage rule data.

[0131] 4. Data Blocking and Encoding

[0132] Operational Note: For large-scale rule list data, it may be necessary to block it so that it can be input into the model in batches. At the same time, for categorical variables or text data, encoding operations such as one-hot encoding and label encoding are required.

[0133] Example: Divide the rule data into blocks according to certain rules, and perform one-hot encoding on the categorical fields in the rules (such as rule type).

[0134] In an embodiment of the present application, by performing data cleaning, standardization, and formatting operations on the rule list data, the data quality can be improved, so that the model can parse and understand the rules more accurately, thereby generating test data that meets the requirements.

[0135] In some optional implementations of the embodiments of the present application, the above-mentioned step of parsing the rule list data to obtain the test data generation logic specifically includes the following steps:

[0136] Read the system database and obtain the ambiguous text list data in the system database;

[0137] Determining whether the rule list data contains ambiguous description text related to the ambiguous text list data;

[0138] If there is ambiguous description text, the trained semantic recognition model is called, and the rule list data is input into the semantic recognition model for semantic recognition operation to obtain the true semantic description text;

[0139] Perform semantic correction operations on the rule list data based on the real semantic description text.

[0140] In the embodiments of this application, the system database typically stores a large amount of business data, rule data, and potentially ambiguous text information. Ambiguous text list data refers to a collection of text that has multiple semantic interpretations and is prone to misunderstanding or confusion. This ambiguous text may originate from historical rule records, user feedback, system logs, and so on.

[0141] In the embodiments of this application, rule list data is an important basis for test data generation and may contain some descriptive text related to the ambiguous text list data. These descriptive texts may be semantically ambiguous due to unclear expression, inappropriate wording, etc. Determining whether such ambiguous descriptive text exists in the rule list data is necessary to enable accurate semantic recognition and correction.

[0142] In the embodiments of this application, when ambiguous descriptive text is detected in the rule list data, a trained semantic recognition model is needed to accurately understand the true semantics of the text. Semantic recognition models are typically based on deep learning technologies such as BERT and GPT. These models are trained on large amounts of annotated data and are able to understand the semantics and context of natural language.

[0143] In an embodiment of the present application, after obtaining the true semantic description text, it is necessary to apply it to the rule list data to correct the original ambiguous description text, so that the semantics of the rule list data is more accurate and clear, and to avoid test data generation errors caused by ambiguity.

[0144] In the embodiment of the present application, the ambiguous description text in the rule list data can be effectively processed to improve the accuracy and reliability of test data generation.

[0145] In some optional implementations of the embodiments of the present application, before the step of calling a trained semantic recognition model if the ambiguous description text exists and inputting the rule list data into the semantic recognition model for semantic recognition operation to obtain the true semantic description text, the following steps are also included:

[0146] Reading a local database, obtaining a sample text from the local database, and determining each word segment included in the sample text;

[0147] Determine the word vector corresponding to each word segment based on the semantic analysis model to be trained;

[0148] Acquire semantic attributes in the local database, and determine a first feature representation vector of the sample text related to the semantic attributes based on an attention matrix corresponding to the semantic attributes in the semantic analysis model to be trained and a word vector corresponding to each word segment;

[0149] Determining a second feature representation vector of the sample text related to the semantic attribute based on a self-attention matrix for representing correlations between different semantic attributes included in the semantic analysis model to be trained and the first feature representation vector;

[0150] Determining a classification result output by the semantic training model to be trained according to the semantic analysis model to be trained and the second feature representation vector, the classification result including the semantic attribute to which the sample text belongs and the sentiment polarity corresponding to the semantic attribute to which the sample text belongs;

[0151] According to the classification result and the preset annotations of the sample text, the model parameters in the semantic analysis model are adjusted to obtain the semantic analysis model.

[0152] In an embodiment of the present application, a plurality of texts may be first obtained from the local database, and a training set consisting of the obtained plurality of texts may be determined. Then, for each text in the training set, the text may be used as a sample text.

[0153] In the embodiment of the present application, when determining the segmented words contained in the sample text, the sample text may be first segmented to obtain each segmented word contained in the sample text. When segmenting the sample text, any segmentation method may be used. Of course, each character in the sample text may also be processed as a segmented word. It should be understood that the example of segmentation processing here is only for convenience of understanding and is not intended to limit the present application.

[0154] In an embodiment of the present application, the semantic analysis model may include at least four layers, namely: a semantic representation layer, an attribute representation layer, an attribute correlation representation layer, and a classification layer.

[0155] In an embodiment of the present application, the semantic representation layer includes at least a sub-model for outputting a bidirectional semantic representation vector, such as a BERT (Bidirectional Encoder Representations from Transformers) model. Each word segmentation can be input into the semantic representation layer in the semantic analysis model to obtain a bidirectional semantic representation vector corresponding to each word segmentation output by the semantic representation layer as the word vector corresponding to each word segmentation. It should be understood that the model for outputting a bidirectional semantic representation vector includes other models in addition to the above-mentioned BERT model. The examples of the model for outputting a bidirectional semantic representation vector here are only for the convenience of understanding and are not used to limit this application.

[0156] In an embodiment of the present application, the word vector corresponding to each word segmentation can be input into the attribute representation layer in the semantic analysis model, and the word vector corresponding to each word segmentation can be attention-weighted through the attention matrix corresponding to the semantic attribute contained in the attribute representation layer. Based on the word vector corresponding to each word segmentation after attention weighting, the first feature representation vector of the sample text involving the semantic attribute is determined.

[0157] In an embodiment of the present application, the first feature representation vector of each semantic attribute involved in the sample text can be input into the attribute correlation representation layer in the speech analysis model, and the first feature representation vector of each semantic attribute involved in the sample text can be self-attention weighted through the above-mentioned self-attention matrix contained in the attribute correlation representation layer. Based on each first feature representation vector after self-attention weighting, the second feature representation vector of each semantic attribute involved in the sample text is determined.

[0158] In the embodiment of the present application, the classification layer includes at least a hidden layer, a fully connected layer and a softmax layer.

[0159] In an embodiment of the present application, the second feature representation vector of each semantic attribute of the sample text can be input into the hidden layer, fully connected layer and softmax layer in the classification layer in sequence. According to each second feature representation vector and the classification parameters corresponding to each semantic attribute contained in the hidden layer, fully connected layer and softmax layer of the classification layer, the sample text is classified to obtain the classification result output by the classification layer.

[0160] In the embodiment of the present application, the classification result includes at least the semantic attribute to which the sample text belongs and the sentiment polarity corresponding to the semantic attribute to which the sample text belongs.

[0161] In an embodiment of the present application, the emotional polarity can be quantified using a numerical value. For example, the closer the numerical value is to 1, the more positive the emotional polarity is; the closer the numerical value is to -1, the more negative the emotional polarity is; and the closer the numerical value is to 0, the more neutral the emotional polarity is.

[0162] In an embodiment of the present application, the model parameters that need to be adjusted include at least the above-mentioned classification parameters, and may also include the above-mentioned attention matrix and self-attention matrix. The model parameters in the semantic analysis model can be adjusted using traditional training methods. That is, directly based on the above-mentioned classification results and the preset annotations for the sample text, the loss corresponding to the classification result (hereinafter referred to as the first loss) is determined, and the model parameters in the semantic analysis model are adjusted with the minimization of the first loss as the training goal to complete the training of the semantic analysis model.

[0163] In an embodiment of the present application, since the self-attention matrix for representing the correlation between different semantic attributes has been added to the above-mentioned semantic analysis model, the semantic analysis model trained using the above-mentioned traditional training method can more accurately analyze the semantics of the text to be analyzed.

[0164] In some optional implementations of the embodiments of the present application, the test data generation logic includes intent recognition rules, and the step of inputting the rule list data into a natural language processing model for rule parsing to obtain the test data generation logic specifically includes the following steps:

[0165] An intent recognition operation is performed on the rule list data according to the natural language processing model to obtain the test data generation logic.

[0166] In the embodiments of the present application, the natural language processing (NLP) model has the ability to understand and parse human natural language text. Rule list data is usually presented in natural language form and contains various business rules, constraints, and logical relationships. The NLP model can conduct in-depth analysis of these rule texts and mine the semantic information contained therein, thereby providing a basis for the derivation of test data generation logic.

[0167] In the embodiment of the present application, the specific process of the intention recognition operation may be:

[0168] (1) Text preprocessing

[0169] Word Segmentation and Part-of-Speech Tagging: First, perform word segmentation on the rule list data, splitting continuous text into meaningful lexical units. For example, for the rule "The user's age must be greater than 18 years old", after word segmentation, we get "user", "age", "must", "be greater than", "18", "years old". At the same time, perform part-of-speech tagging on each word to clarify its grammatical role. For example, "user" is a noun, and "be greater than" is a verb, etc. This helps the model better understand the structure and semantics of the text.

[0170] Stop Word Removal: Stop words refer to words that frequently appear in text but contribute little to semantic understanding, such as "of", "is", "in", etc. Removing these stop words can reduce data noise and improve the processing efficiency and accuracy of the model.

[0171] Text Vectorization: Convert the preprocessed text into a numerical vector form that can be processed by a computer. Common methods include the Bag of Words model, Word Embedding, etc. Word embedding methods such as Word2Vec and GloVe can map words into a low-dimensional vector space, making words with similar semantics closer in the vector space, thus better preserving the semantic information of words.

[0172] (2) Application of the Intent Recognition Model

[0173] Model Selection and Loading: According to the complexity of the rules and domain characteristics, select an appropriate NLP intent recognition model. Common models include recurrent neural networks (RNN) and their variants (such as LSTM, GRU) based on deep learning, convolutional neural networks (CNN), and pre-trained language models (such as BERT, GPT, etc.). Load the already trained model, which has learned a large number of semantic and intent features of natural language text during the training process.

[0174] Feature Extraction and Matching: Input the preprocessed rule text vector into the intent recognition model. The model will extract the features of the text and match them with the predefined intent category features. For example, for the rule "The order amount cannot exceed 1000 yuan", the model may recognize that its intent is related to "amount limit".

[0175] Intent Classification and Output: Based on the results of feature matching, the model classifies the rule text into the corresponding intent categories and outputs the intent recognition results. Intent categories can include data range constraints, data format requirements, logical relationship judgments, etc.

[0176] In the embodiments of this application, the specific process from intent recognition to the derivation of the test data generation logic can be:

[0177] (1) Extraction of Logic Rules Based on Intent

[0178] Logical analysis for different intents: Based on the intent recognition results, further logical analysis is performed on the rule list data. For example, if the intent is "data range constraint," information such as the data field, minimum value, and maximum value must be extracted from the rule text. For the rule "User age must be between 18 and 60," the extracted data field is "age," with a minimum value of 18 and a maximum value of 60.

[0179] Logical relationship analysis: Analyze the logical relationships within the rule text, such as those expressed by logical conjunctions like "and," "or," and "not." For example, in the rule "Users must be members and over 18 years old," the logical relationship of "and" exists, indicating that both conditions must be met simultaneously.

[0180] (2) Construction of generation logic

[0181] Conditional judgment logic: Build conditional judgment logic for test data generation based on information such as the extracted data range and format requirements. For example, for an age range constraint, you can build the following conditional judgment logic: If the generated user data contains an age field, check whether the age value is between 18 and 60. If not, regenerate the data.

[0182] Data Generation Strategy: Determine the test data generation strategy based on the rule requirements. For example, if the rule requires a specific date format (such as "YYYY-MM-DD"), the test data must be generated in that format. For logical relationship judgments, such as the "and" relationship mentioned above, ensure that both the membership and age requirements of over 18 are met when generating test data.

[0183] Exception handling logic: Consider possible exceptions within the rules and build corresponding exception handling logic. For example, if the rules specify data generation rules for certain special circumstances, such as "When the user is a VIP, the upper limit can be increased to 2,000 yuan," you need to add VIP user judgment and corresponding upper limit adjustment logic to the generation logic.

[0184] In an embodiment of the present application, an intent recognition operation can be performed on the rule list data based on a natural language processing model, and accurate and reasonable test data generation logic can be derived, thereby improving the quality and effectiveness of the test data.

[0185] In some optional implementations of the embodiments of the present application, after the step of executing the data simulation device to obtain target test data, the following steps are further included:

[0186] Performing a data verification operation on the target test data to obtain a data verification result;

[0187] If the data verification result contains abnormal test data that does not meet the requirements of the original rules, the natural language processing model is optimized based on the abnormal test data, and the rule list data is re-performed with intent recognition based on the optimized natural language processing model to generate target test data that meets the requirements of the original rules.

[0188] In the embodiments of the present application, after generating target test data based on the rule list data, data validation is performed to ensure that the generated test data strictly complies with the original rule requirements. This is a key step in ensuring test data quality, as test data that does not comply with the rules may lead to inaccurate test results and an inability to effectively evaluate the performance and functionality of the system.

[0189] In an embodiment of the present application, first, it is necessary to extract clear rule requirements from the original rule list data. These rules may include data formats (such as date format, numerical precision, etc.), data ranges (such as minimum and maximum values ​​of numerical values), logical relationships (such as relationships expressed by logical conjunctions such as "and", "or", and "not") and other specific business constraints. For example, the original rule may stipulate that "the order amount must be a positive number and not exceed 10,000 yuan", then it is necessary to extract the two conditions "order amount>0" and "order amount≤10,000" from the rule.

[0190] In this embodiment of the present application, the target test data is input into the validation rules one by one for verification. For each test data, it is checked whether it meets all the defined validation rules. The verification result of each test data is recorded, including whether it passed the verification and the specific reason for failing the verification. For example, if the test data of an order amount is -500, the verification result will show that the data does not meet the rule that "the order amount must be a positive number."

[0191] In the embodiment of the present application, the verification results of all test data are summarized and analyzed to generate a data verification result report. The report content may include information such as the number of test data that passed the verification, the number of test data that failed the verification, and the distribution of rule types that failed the verification.

[0192] In the embodiments of the present application, when abnormal test data that does not meet the original rule requirements is discovered, it is necessary to conduct in-depth analysis of the reasons for this data. This may be due to the natural language processing model's deviation in understanding the rule text, resulting in the generated test data not meeting expectations. For example, if the rule text uses ambiguous expressions such as "large amount", the model may not accurately understand the specific scope of "large", thereby generating amount data that does not meet the rules. All abnormal test data and their corresponding rule text are collected to form an abnormal sample set. These abnormal samples will serve as an important basis for model optimization.

[0193] In this embodiment of the present application, if the abnormal test data is caused by the model's lack of learning of certain rule situations, you can consider adding relevant training data. For example, collect more sample data containing fuzzy expression rules and corresponding correct test data, add this data to the model's training set, and retrain the model to improve the model's understanding of fuzzy rules.

[0194] In the embodiments of the present application, natural language processing models typically have some adjustable parameters, such as learning rate, batch size, and number of iterations. By adjusting these parameters, the model training process can be optimized and the model performance can be improved. For example, appropriately reducing the learning rate can make the model more stable during training and avoid overfitting or underfitting.

[0195] In the embodiments of this application, if the existing model architecture cannot effectively handle the complex semantics and logical relationships in the rule text, it is possible to consider improving the model architecture. For example, more advanced pre-trained language models such as BERT and GPT-3 can be used. These models are pre-trained on large-scale text data and have stronger semantic understanding and generation capabilities. Alternatively, the existing model can be improved, such as by adding an attention mechanism to enable the model to better focus on key information in the rule text.

[0196] In the embodiments of this application, domain expertise is incorporated into the model optimization process. For example, for financial rules, financial experts' interpretations and explanations of the rules can be incorporated. This domain knowledge can be incorporated into the model as additional features or constraints, improving the model's ability to understand the rules and generate test data that complies with the rules.

[0197] In the embodiment of the present application, the target test data can be effectively verified, and the natural language processing model can be optimized based on the verification results, so as to generate target test data that meets the original rule requirements, thereby improving the quality of the test data and the accuracy of the test.

[0198] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0199] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0200] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0201] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0202] Further references Figure 3 , as a response to the above Figure 2 In order to realize the method shown in FIG, the present application provides an embodiment of a test data generating device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0203] like Figure 3 As shown, the test data generating device 200 of the embodiment of the present application includes:

[0204] The request receiving module 210 is configured to receive a test data generation request sent by a user terminal, wherein the test data generation request includes table list data and rule list data, and the rule list data is used to guide the generation of simulation data;

[0205] The rule parsing module 220 is used to parse the rule list data to obtain the test data generation logic;

[0206] A simulation device generation module 230 is used to generate a data simulation device according to the configured simulation environment, table list data and test data generation logic;

[0207] The simulation device execution module 240 is used to execute the data simulation device to obtain target test data.

[0208] In an embodiment of the present application, a test data generation device 200 is provided, including: a request receiving module 210, used to receive a test data generation request sent by a user terminal, wherein the test data generation request includes table list data and rule list data, and the rule list data is used to guide the generation of simulation data; a rule parsing module 220, used to parse the rule list data to obtain test data generation logic; a simulation device generation module 230, used to generate a data simulation device according to the configured simulation environment, table list data and test data generation logic; a simulation device execution module 240, used to execute the data simulation device to obtain target test data. Compared with the existing technology, this application uses an intuitive configuration interface and intelligent rule parsing, allowing users to easily generate test data without having a deep technical background. In addition, this application can quickly generate test data that meets business needs through automated rule parsing and device generation; this application supports multiple database types (such as MySQL, Oracle, MongoDB, etc.), and users only need to configure it once to use it in multiple database environments, significantly improving the reusability of the configuration; this application automatically generates about 80% of the configuration data through model information, reducing the workload and error rate of manual configuration; this application provides auxiliary functions such as data rollback, automatic backup, concurrent simulation, data refresh, etc., which greatly improves the user experience.

[0209] In some optional implementations of the embodiments of the present application, the test data generating device 200 further includes:

[0210] The data preprocessing module is used to perform data cleaning, data standardization and data formatting operations on the rule list data.

[0211] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device according to an embodiment of the present application.

[0212] The computer device 300 includes a memory 310, a processor 320, and a network interface 330 that are interconnected through a system bus. It should be noted that the figure only shows the computer device 300 having components 310-330, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0213] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0214] The memory 310 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disk, optical disk, etc. In some embodiments, the memory 310 may be an internal storage unit of the computer device 300, such as a hard disk or memory of the computer device 300. In other embodiments, the memory 310 may also be an external storage device of the computer device 300, such as a plug-in hard disk equipped on the computer device 300, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 310 may also include both the internal storage unit of the computer device 300 and its external storage device. In the embodiment of the present application, the memory 310 is generally used to store the operating system and various application software installed on the computer device 300, such as computer-readable instructions for the test data generation method. In addition, the memory 310 can also be used to temporarily store various data that has been output or is about to be output.

[0215] In some embodiments, the processor 320 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 320 is generally used to control the overall operation of the computer device 300. In the embodiment of the present application, the processor 320 is used to execute computer-readable instructions stored in the memory 310 or process data, such as computer-readable instructions for executing the test data generation method.

[0216] The network interface 330 may include a wireless network interface or a wired network interface. The network interface 330 is generally used to establish a communication connection between the computer device 300 and other electronic devices.

[0217] The computer equipment provided by this application allows users to easily generate test data without having a deep technical background through an intuitive configuration interface and intelligent rule parsing. In addition, this application can quickly generate test data that meets business needs through automated rule parsing and device generation; this application supports multiple database types (such as MySQL, Oracle, MongoDB, etc.), and users only need to configure it once to use it in multiple database environments, which significantly improves the reusability of the configuration; this application automatically generates about 80% of the configuration data through model information, reducing the workload and error rate of manual configuration; this application provides auxiliary functions such as data rollback, automatic backup, concurrent simulation, data refresh, etc., which greatly improves the user experience.

[0218] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the test data generation method as described above.

[0219] The computer-readable storage medium provided by this application allows users to easily generate test data without having a deep technical background through an intuitive configuration interface and intelligent rule parsing. In addition, this application can quickly generate test data that meets business needs through automated rule parsing and device generation; this application supports multiple database types (such as MySQL, Oracle, MongoDB, etc.), and users only need to configure it once to use it in multiple database environments, significantly improving the reusability of the configuration; this application automatically generates about 80% of the configuration data through model information, reducing the workload and error rate of manual configuration; this application provides auxiliary functions such as data rollback, automatic backup, concurrent simulation, data refresh, etc., which greatly improves the user experience.

[0220] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0221] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the scope of application of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of protection of the present application.

Claims

1. A test data generation method, characterized in that: The steps include: Receiving a test data generation request sent by a user terminal, wherein the test data generation request includes table list data and rule list data, and the rule list data is used to guide the generation of simulation data; Performing a parsing operation on the rule list data to obtain test data generation logic; generating a data simulation device according to the configured simulation environment, the table list data, and the test data generation logic; The data simulation device is executed to obtain target test data.

2. The test data generation method according to claim 1, wherein: After the step of receiving the test data generation request sent by the user terminal and before the step of parsing the rule list data to obtain the test data generation logic, the following steps are further included: Perform data cleaning, data standardization, and data formatting operations on the rule list data.

3. The test data generation method according to claim 1, wherein: The step of parsing the rule list data to obtain the test data generation logic specifically includes the following steps: Reading a system database, and obtaining ambiguous text list data in the system database; Determining whether there is ambiguous description text related to the ambiguous text list data in the rule list data; If the ambiguous description text exists, calling the trained semantic recognition model, and inputting the rule list data into the semantic recognition model to perform semantic recognition operation to obtain the true semantic description text; A semantic correction operation is performed on the rule list data according to the true semantic description text.

4. The test data generation method according to claim 3, wherein: Before the step of calling a trained semantic recognition model if the ambiguous description text exists and inputting the rule list data into the semantic recognition model to perform a semantic recognition operation to obtain a true semantic description text, the following steps are also included: Reading a local database, obtaining a sample text from the local database, and determining each word segment included in the sample text; Determine the word vector corresponding to each word segment based on the semantic analysis model to be trained; Acquire semantic attributes in the local database, and determine a first feature representation vector of the sample text related to the semantic attributes based on an attention matrix corresponding to the semantic attributes in the semantic analysis model to be trained and a word vector corresponding to each word segment; Determining a second feature representation vector of the sample text related to the semantic attribute based on a self-attention matrix for representing correlations between different semantic attributes included in the semantic analysis model to be trained and the first feature representation vector; Determining a classification result output by the semantic training model to be trained according to the semantic analysis model to be trained and the second feature representation vector, the classification result including the semantic attribute to which the sample text belongs and the sentiment polarity corresponding to the semantic attribute to which the sample text belongs; According to the classification result and the preset annotations of the sample text, the model parameters in the semantic analysis model are adjusted to obtain the semantic analysis model.

5. The test data generation method according to claim 1, wherein: The test data generation logic includes intent recognition rules. The step of inputting the rule list data into a natural language processing model for rule parsing to obtain the test data generation logic specifically includes the following steps: An intent recognition operation is performed on the rule list data according to the natural language processing model to obtain the test data generation logic.

6. The test data generating method according to claim 5, characterized in that: After the step of executing the data simulation device to obtain target test data, the following steps are also included: Performing a data verification operation on the target test data to obtain a data verification result; If the data verification result contains abnormal test data that does not meet the requirements of the original rules, the natural language processing model is optimized based on the abnormal test data, and the rule list data is re-performed with intent recognition based on the optimized natural language processing model to generate target test data that meets the requirements of the original rules.

7. A test data generating device, characterized in that: include: a request receiving module, configured to receive a test data generation request sent by a user terminal, wherein the test data generation request includes table list data and rule list data, and the rule list data is used to guide the generation of simulation data; A rule parsing module, configured to parse the rule list data to obtain test data generation logic; A simulation device generation module, configured to generate a data simulation device according to the configured simulation environment, the table list data and the test data generation logic; The simulation device execution module is used to execute the data simulation device to obtain target test data.

8. The test data generating device according to claim 7, wherein: The device further comprises: The data preprocessing module is used to perform data cleaning operations, data standardization operations and data formatting operations on the rule list data.

9. A computer device comprising a memory and a processor, characterized in that: The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the test data generating method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the test data generating method according to any one of claims 1 to 6.