Test data generation method and related device

By using the knowledge base and component library of intelligent agents, the production process of automatically generating test data solves the problem of low test data generation efficiency in existing technologies, and realizes efficient and intelligent test data generation.

CN121524040APending Publication Date: 2026-02-13BEIJING AUTONAVI YUNMAP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511434761.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

In existing technologies, generating test data manually is inefficient and cannot efficiently meet the testing needs of diverse and complex business scenarios.

Method used

By employing a knowledge base and component library of intelligent agents, production process information is generated by determining test requirements, and components in the component library are called to automatically generate test data.

Benefits of technology

It improves the efficiency and accuracy of test data generation, and can intelligently match test data requirements to adapt to complex and ever-changing business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524040A_ABST
    Figure CN121524040A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a test data generation method and a related device. The method comprises the following steps: determining test demand information; inputting the test demand information into an intelligent agent, so that the intelligent agent generates production flow information corresponding to the test demand information according to a knowledge base; wherein the knowledge base comprises description information of a plurality of components in a preset component library, the components are used for obtaining at least one type of test data in a business scene corresponding to the test demand information, and the production process information comprises at least one step of generating the test data corresponding to the test demand information; and calling at least part of components in the component library according to the production process information, and generating test data corresponding to the test demand information. Through the method provided by the embodiment of the invention, the efficiency of generating the test data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of testing, in particular to a test data generation method and related device. BACKGROUND

[0002] With the development of technology, the types of data in some business scenarios are constantly enriched. For example, in the out-of-home business scenario, the application of travel services can generate multiple large categories of data such as traffic signal data, vehicle flow dynamic data, and route planning data. Under these large categories of data, there are also small categories of data such as traffic signal lamp position data, traffic signal lamp real-time state data, and vehicle flow data. In addition, the application of various subdivided businesses also increases the complexity of interaction between data.

[0003] Therefore, based on the influence of the growth of data types and the complexity of data, it is difficult to efficiently generate test data when generating test data through artificial construction when testing business capabilities in different business scenarios. SUMMARY

[0004] The present application provides a test data generation method and related device to improve the efficiency of generating test data.

[0005] In a first aspect, the present application provides a test data generation method, which comprises:

[0006] determining test requirement information;

[0007] inputting the test requirement information into an agent to enable the agent to generate production process information corresponding to the test requirement information according to a knowledge base; wherein the knowledge base comprises description information of a plurality of components in a preset component library, the components are used to obtain at least one type of test data in a business scenario corresponding to the test requirement information, and the production process information comprises at least one step of generating the test data corresponding to the test requirement information;

[0008] generating the test data corresponding to the test requirement information according to the production process information and by calling at least part of the components in the component library.

[0009] In a second aspect, the present application provides a test data generation device, which comprises:

[0010] a determination module configured to determine test requirement information;

[0011] The processing module is configured to input the test requirement information into the agent, so that the agent generates production process information corresponding to the test requirement information according to a knowledge base; wherein the knowledge base comprises description information of a plurality of components in a preset component library, the components are used to obtain at least one kind of test data in a business scenario corresponding to the test requirement information, and the production process information comprises at least one step of generating the test data corresponding to the test requirement information.

[0012] The processing module is further configured to generate the test data corresponding to the test requirement information according to the production process information and by invoking at least part of the components in the component library.

[0013] In a third aspect, the present application provides an electronic device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method provided in the first aspect of the present application.

[0014] In a fourth aspect, the present application provides a computer readable storage medium, and the computer readable storage medium stores computer execution instructions, and when a processor executes the computer execution instructions, the method provided in the first aspect of the present application is implemented.

[0015] In a fifth aspect, the present application provides a computer program product, comprising a computer program, and when the computer program is executed by a processor, the method provided in the first aspect of the present application is implemented.

[0016] The test data generation method and related device provided by the present application determine test requirement information; input the test requirement information into an agent, so that the agent generates production process information corresponding to the test requirement information according to a knowledge base; since the knowledge base of the agent comprises description information of a plurality of components in a preset component library, each component has the ability to obtain at least one kind of test data in a business scenario corresponding to the test requirement information, therefore, in the case that the agent understands the test requirement, the whole process of generating test data can be intelligently streamlined in combination with the ability of the components and the demand of the test activity for the test data, and then the production process information corresponding to the test requirement information is obtained, and the production process information comprises steps of generating the test data corresponding to the test requirement information. Therefore, at least part of the components in the component library can be reasonably invoked according to the production process information, and the test data corresponding to the test requirement information is generated, and then the efficiency of generating the test data is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate one embodiment consistent with the present application and, together with the specification, serve to explain the principles of the application.

[0018] Figure 1A schematic diagram of a test activity provided for an embodiment of the present application;

[0019] Figure 2 A schematic diagram of a test data generation method provided for an embodiment of the present application;

[0020] Figure 3 A schematic diagram of a platform system for implementing test data service provided for an embodiment of the present application;

[0021] Figure 4 A schematic diagram of a multi-level intention recognition intelligent agent provided for an embodiment of the present application;

[0022] Figure 5 A schematic diagram of a platform system for implementing test data service provided for an embodiment of the present application;

[0023] Figure 6 A schematic diagram of retrieval enhancement generation provided for an embodiment of the present application;

[0024] Figure 7 A schematic diagram of an associated scene recommendation capability implementation provided for an embodiment of the present application;

[0025] Figure 8 A schematic diagram of a test data generation apparatus provided for an embodiment of the present application;

[0026] Figure 9 A schematic diagram of an electronic device provided for an embodiment of the present application.

[0027] The specific embodiments of the present application have been shown through the above-described drawings, and will be described in more detail hereinafter. These drawings and written descriptions are not intended to limit the scope of the present application concept in any way, but to illustrate the present application concept to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0028] The exemplary embodiments will be described in detail herein below with reference to the drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.

[0029] It should be noted that the user information (including but not limited to user equipment information, user attribute information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards, and provide corresponding operation portal for user to choose authorization or refusal.

[0030] The test data can also be understood as test scenarios or test cases and their corresponding data used for program testing in business scenarios. The test data is the specific data preparation used in the test activity, which can be a single data or a combination of multiple data.

[0031] Figure 1 The schematic diagram of the test activity provided by the embodiments of the present application is shown in Figure 1 As shown in the figure, taking the test activity of an outbound business scenario as an example, when testing the platform program related to the outbound business scenario, a combination scenario can be obtained by customizing the combination of dynamic, static and route type data, simulating the external factors that may be encountered in the outbound process from the starting point to the ending point and the behaviors that may occur in the outbound process. The combination scenario can be a test scenario, and all data corresponding to the test scenario form a set of test data. The dynamic type data can be understood as real-time changing data, such as the reading of the remaining time of the green light in a traffic light, etc. The static type data can be understood as non-real-time changing data, such as the total number of lanes in a road, the position coordinates or total number of traffic light devices in a section of road, etc.

[0032] As shown in Figure 1 The test data can constitute a test scenario when simulating the outbound process using the platform program. By calling the test data through the service interface of the platform program, the platform program can respond to each event triggered in the test scenario and obtain the response results such as the decisions or actions made by the platform program when responding to each event. By analyzing the response results, the test results of whether the platform program responds to each event as expected can be obtained, and according to the test results, the defects of the platform program can be understood, and the platform program can also be optimized through the test results.

[0033] Due to the diversified development of various business scenarios, the breadth of data categories included in many businesses and the complexity of data interaction are in a significant upward trend. The complexity and richness of business scenarios have evolved test data from relatively single category data to multi-category data, and a test scenario may be a complex scenario generated by relying on multi-dimensional data fusion, which undoubtedly brings great difficulty to the efficiency of generating test data in the test activity.

[0034] In the existing test data generation method, a test requirement is analyzed by an experienced test personnel manually, and steps for generating various types of data are determined according to the test requirement, and the required data is generated step by step according to the steps, and finally the generated data is combined to obtain a set of test data for testing.

[0035] In this process, the understanding of the business functions that can be implemented by the platform or program and the familiarity with the generation logic of various types of data required in each business function are highly dependent on the test personnel. Due to the complexity of business functions and the breadth of data categories, and factors such as different business functions usually being developed by different test personnel, it is difficult for test personnel to comprehensively master the implementation of each business function and the generation logic of each type of data, and it is difficult to manually construct test data. Therefore, the efficiency of the existing test data generation is low.

[0036] Therefore, in order to improve the efficiency of test data generation, an embodiment of the present application provides a test data generation method. The method determines test requirement information and inputs the test requirement information into an agent, so that the agent generates production process information corresponding to the test requirement information according to a knowledge base. The knowledge base of the agent includes description information of a plurality of components in a preset component library. Each component has the ability to generate at least one type of test data in a business scenario corresponding to the test requirement information. Therefore, when the agent understands the test requirement, it can intelligently analyze the matching between the data generation capabilities of each component in the component library and the requirements of the test activity on the test data, make a decision to obtain production process information representing each step of generating test data, and then generate test data corresponding to the test requirement information according to the production process information, thereby improving the efficiency of generating test data.

[0037] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The embodiments described below and the features in the embodiments can be combined with each other without conflict. The sequence of steps in each method embodiment described below is only an example, not a strict limitation.

[0038] Figure 2 A flowchart of a test data generation method provided by an embodiment of the present application is shown. The method can be implemented by an electronic device with corresponding data processing capability, such as a computer, a server or a server cluster. As shown in the figure, the method includes: Figure 2

[0039] S201, determining test requirement information.

[0040] ​The test requirement information can be information for characterizing specific requirements of the test data for the test activity. The test requirement information can be determined by receiving input query information, etc.

[0041] For example, an electronic device can be loaded with a platform system for testing data service, which can provide services for generating test data, and realize service of test data. The service of test data can be understood as providing test data required in the test process in the form of service, so as to facilitate testers to obtain and apply more flexibly and efficiently. For example, the platform system for realizing the service of test data can support various productized use modes such as conversational, web user interface (WEB UI), open application programming interface (Open API) and agent.

[0042] For example, the test activity needs to test the platform program of the out-of-industry business, and specifically needs to test the guiding function related to path navigation. The platform system for realizing the service of test data can receive input query information, which can include keywords and other information representing test requirements, such as road type and road name. The platform system can determine the test requirement information after receiving the query information.

[0043] In one possible implementation, when determining the test requirement information, an agent can be used to determine the test requirement information, specifically by obtaining natural language query information; inputting the natural language query information into an intent recognition agent to obtain an intent recognition result representing an intent, the intent recognition agent being an agent for analyzing natural language query information through a natural language model to recognize an intent; and in the case that the intent recognition result includes an intent of generating test data, determining the test requirement information according to the natural language query information through the intent recognition agent.

[0044] For example, the intent recognition agent can be an agent loaded into a platform system that implements test data service. An agent can be understood as an intelligent system capable of autonomous thinking and action. For instance, an agent can automatically process questions and requests through learning and reasoning, thereby providing more efficient and accurate services. The agent can include machine learning-based network models such as large language models (LLMs). Large language models, often simply referred to as big language models or big models, are a natural language processing technology based on deep learning. They can understand natural language and generate natural language text. In the embodiments of this application, the large language model can be any large model capable of processing natural language; it can be a pre-trained large model or a large model obtained through autonomous training based on business scenario requirements.

[0045] Figure 3 This is a flowchart illustrating the platform system for implementing test data service provided in the embodiments of this application, such as... Figure 3 As shown, the platform system includes an intent recognition agent, which includes an LLM (Limited Language Management) module. The intent recognition agent has the ability to recognize the intent of natural language query information, and is used to better generate test data that meets the expectations of the testing activities. Through this intent recognition agent, the intent of natural language query information can be recognized.

[0046] For example, the intent recognition agent in the platform system can receive natural language query information through text input or voice input. This natural language query information can be any supported language, and its length can be any length within a preset maximum length. For instance, the language and the preset maximum length can be preset according to the corresponding support range of the LLM. For ease of description, natural language query information can also be simply referred to as natural language or query information.

[0047] Intent recognition results can characterize the intent of natural language. These results can include at least one of multiple intents, such as intents for personalized interaction, scope retrieval, and test scenario generation. To achieve high-accuracy natural language intent recognition, the intent recognition agent in this embodiment can be a multi-level intent recognition agent designed using the concept of thought chains. Through the classification and recognition of natural language by this multi-level intent recognition agent, it is possible to achieve generalized to refined classification and recognition of natural language intents, thereby enhancing the recognition.

[0048] Figure 4 This is a schematic diagram of a multi-level intent recognition agent provided in the embodiments of this application, such as... Figure 4As shown, query information can be received through a multi-level intent recognition agent. After receiving the query information, the multi-level intent recognition agent can classify and recognize the intent to obtain the corresponding intent recognition result.

[0049] For example, a three-layer classification mechanism can be designed to align with potential behaviors, covering three types of intents: personalized interaction, range retrieval, and test scenario generation. Multiple specialized knowledge bases can be created to house knowledge content from different intent domains, forming a knowledge base recall sequence under a thought chain mechanism, enabling effective indexing of relevant knowledge. When natural language is received, i.e., after a query is input, the query can be analyzed layer by layer to accurately recall corresponding knowledge. If the intent recognition result determines that it matches one of the three intent categories, further matching is unnecessary. This mechanism enables the intent recognition agent to demonstrate high adaptability and accurate service capabilities in complex and varied queries, while also achieving improved response speed by adaptively adjusting to the complexity of the instructions.

[0050] For example, if the received query is "What time is it?", the multi-level intent recognition agent can analyze the query to determine its corresponding personalized interactive intent. Therefore, it can combine knowledge such as template dialogues, QA dialogues, and interactive dialogues in knowledge base A with LLM to generate content and output the corresponding answer. For example, for the query "What time is it?", the output answer could be "It's XX o'clock now. Are you tired from work? Let me tell you a joke."

[0051] For example, if the received query is "What types of construction events are there?", the multi-level intent recognition agent can analyze the query to determine the intent of the corresponding scope of retrieval. Therefore, it can combine the capability scope of each component, attribute value list, and data-related business knowledge in knowledge base B with the LLM to generate content and output the corresponding answer. For example, corresponding to the query "What types of construction events are there?", the output answer could be "The required output scope retrieval has been identified; construction events include XXX." For instance, content can also be generated based on prompts.

[0052] In one possible implementation, the intent recognition agent includes an intent knowledge base for intent recognition. The intent recognition agent determines test requirement information based on natural language query information, including: retrieving knowledge content related to the natural language query information from the intent knowledge base to obtain target knowledge content associated with the test requirement information; inputting the target knowledge content and corresponding prompt words into the natural language model of the intent recognition agent to obtain the test requirement information, wherein the prompt words are used to guide the natural language model to generate content in combination with the target knowledge content.

[0053] The intent knowledge base can include a specialized knowledge base that matches various intents in the intent recognition agent, including professional knowledge in business scenarios. For example Figure 4 corresponding to the personalized interaction knowledge base A and the knowledge base B corresponding to the business scenario.

[0054] When the intent recognition result hits the test scenario generation class of intent, the intent can be refined and split. For example, first inject the capability range of each component, the strategy used between components, generate disassembled positive examples, generate disassembled negative examples, limit and simplify examples, etc. in the corresponding knowledge base B; and set the role, strategy, skill, output, etc. in the prompt word, so as to effectively guide the LLM thinking through the prompt word. The LLM will recall the knowledge in the knowledge base B and the corresponding prompt word for enhanced learning to generate an intent recognition result. The intent recognition result can include test requirement information.

[0055] The prompt word can be a prompt input information used to guide and constrain the content production process of the large model. Through the prompt word, the model can be guided to focus on a specific topic, adopt a specific style or format, or follow a specific logical structure, etc. The prompt word can be provided at the time of input, or can be generated according to the knowledge content retrieved in the context or knowledge base. The prompt word can contain explicit instructions or questions to help the large model understand the intent and expected output.

[0056] As shown in Figure 4 For unfeasible test requirement information or insufficient condition test requirement information, the process can be interrupted, and the analysis result can be fed back. The feasible test requirement information can output enhanced enhanced query information. The enhanced query information can be understood as information obtained after the input natural language is accurately disassembled and enhanced, which can interpret the test requirement information. The enhanced query information can include: original query, test requirement feasibility conclusion, sufficiency analysis conclusion and complete process of test data generation, wherein the complete process of test data generation can include components, attributes and step sequence to be used, etc. The enhanced query information can be used as the input of the subsequent content generation agent and scenario recommendation agent, to realize query information intent enhancement, and effectively improve the accuracy of the knowledge base recall of the agent in the next step of cooperation.

[0057] For example, if the received query is "Hello, I need to generate XXX", the multi-level intent recognition agent can analyze the query and determine that it corresponds to the test scenario generation class of the test scene. Therefore, the enhanced query information can be obtained by combining the component capability range, attribute value list, data related business knowledge, generation and disassembly positive examples and generation and disassembly negative examples in the knowledge base B, and the corresponding prompt words, and the content generation of the LLM. During the generation process, the multi-level intent recognition agent can perform feasibility analysis, sufficiency analysis, and scene generation process design on the test requirement information. The feasibility analysis can be understood as matching the test requirement with the generation capability of the agent to determine whether the test data of the test scene can be generated. The sufficiency analysis can be understood as sufficiency analysis on the received test requirement to determine whether the received information is missing conditions, etc. The scene generation process design can be understood as designing or describing the complete process of test data generation. Corresponding to the query, the generated enhanced query information can include the conclusion of the feasibility, the original query, the key components and keywords needed in the production process, the specific steps and key attributes, etc. The enhanced query information can be input into the content generation agent and the scene recommendation agent, and the generated answer result can be "Recognize your requirement as scene generation, generate step XXX".

[0058] Taking the travel scenario as an example, the knowledge base recall is usually performed according to the received query information, and the completeness of the query information has a great influence on the similarity calculation in the recall process. If the received query information is a simple natural language, for example, "Query the walking route with left turn as the main navigation action", in the content generation link, the driving, riding and walking query components all have the attribute of main navigation action, and the strategies of the components corresponding to the three are different in content generation. If the received query information is directly used for retrieval, a large amount of irrelevant content will be recalled because of the "main navigation action" matching, and even the increase of invalid content may lead to the output of incorrect or invalid results without recalling valid content. As described in the previous intent enhancement, the received query information is enhanced by the preposed intent recognition agent, and the output result with explicit information, for example, "Key words: walking attribute query component", is used as the enhanced query information input, which can retrieve knowledge content that is more consistent with the intent. This method of deep analysis and expansion of the original query information by splitting complex problems can improve the reliability and efficiency of the query.

[0059] S202, input the test requirement information to the agent, so that the agent generates production process information corresponding to the test requirement information according to the knowledge base; wherein the knowledge base includes description information of a plurality of components in a preset component library, the component is used to obtain at least one kind of test data in a business scenario corresponding to the test requirement information, and the production process information includes at least one step of generating the test data corresponding to the test requirement information.

[0060] For example, the agent can be a content generation agent. The test requirement information can be enhanced query information. As shown in Figure 4 The enhanced query information can be input into the content generation agent, so that the content generation agent generates production process information corresponding to the test requirement information according to the knowledge base of the content generation agent. The knowledge base of the content generation agent can be as shown in Figure 4 The knowledge base C includes description information of component functions, business data knowledge, component protocol description, content generation positive and negative examples, and standardized process content generation strategy. The production process information can be information indicating how the agent or the data generation engine produces the test data. For example, the production process information can be information in a standardized protocol format.

[0061] The component library can be a collection including a plurality of components, each component can generate at least one kind of test data, and the test data can be any type of data in a business scenario, such as static, dynamic and other types of data in an industry business scenario. A unit with the ability to generate any one or more types of data in a business scenario can be regarded as a component, and the components in the component library can be easily combined. The component library can also be understood as a collection of data generation capabilities. The component can be a module or device including a computer program, such as an application programming interface (application programming interface, API) and the like.

[0062] Figure 5 The architecture of the platform system for realizing the service of test data provided by the embodiments of the present application is shown in Figure 5 As shown in the figure, the architecture of the platform system for realizing the service of test data can adopt a layered architecture, including an application layer, a service layer, an engine and a model layer, and a data layer.

[0063] The application layer is responsible for external interaction and provides natural language to data test scenario generation. Natural language to data can be understood as intelligently identifying the needs of test activities for test scenarios according to the received natural language, shielding the scene generation method, and providing the test data required by the test. In addition to realizing test scenario generation, the application layer can also provide test scenario analysis recommendation and individualized scenario adaptation capabilities, and each function can support multiple productized usage modes such as conversation, WEB UI, Open API and agent.

[0064] The service layer is responsible for processing and providing core logic. The method provided in the embodiments of the present application can be implemented based on a multi-agent collaborative architecture. The multi-agent collaborative architecture can also be understood as a multi-agent system, which can be a system composed of multiple agents, each agent has its own goals and behaviors, and interacts and cooperates with each other, so as to better handle complex problems and requirements, and improve the overall efficiency and performance of the system. Based on an artificial intelligence (AI) development framework, a multi-thought chain mechanism can be designed, a multi-agent system can be built, and intelligent collaboration and asynchronous independence between agents can be realized.

[0065] In the multi-agent collaborative architecture, an intent recognition agent, a content generation agent, and a scene recommendation agent can be included. Each agent can be constructed based on an LLM, retrieval-augmented generation (RAG), a prompt word, and a tool mode to enhance the generation framework to build three agents of intent enhancement, content generation, and scene analysis recommendation. A multi-chain interaction mechanism is designed to realize the linkage between the multi-agent and the underlying engine. The retrieval-augmented generation can be a technology that uses information from private or proprietary data sources to assist text generation, which can be implemented through a retrieval model. The retrieval-augmented generation can also be understood as enhanced retrieval. By combining the retrieval model and the LLM, the retrieved knowledge can be used to enhance the output content of the generation model.

[0066] As in the above embodiments, based on the intent recognition agent, the ability of high-accuracy natural language intent recognition can be realized. The thought chain mechanism is constructed to split the recall sequence of multiple specialized knowledge bases, and a multi-level intent recognition model (multi-level intent recognition agent) is embedded to realize the classification and enhancement of natural language intent from generalization to refinement, bring high adaptability and accurate service ability of the intent recognition agent, and effectively improve the accuracy of the content generation agent in knowledge base recall in the next step of collaboration.

[0067] The intent recognition agent can perform semantic analysis and enhanced retrieval on the received query information, and then can realize data generation to obtain corresponding enhanced query information after feasibility analysis and data generation flow disassembly, to be transmitted to the scene recommendation agent and the content generation agent.

[0068] The scene recommendation agent can implement traffic-based scene analysis. Specifically, the scene distribution and attribute distribution can be analyzed based on expert experience, and then intent-based scene recommendation and coverage-based scene recommendation can be implemented. In the intent-based scene recommendation, the recommendation information associated with the test data can be obtained through intent extraction and associated scene mining. In the coverage-based scene recommendation, the recommendation information quality can be improved through coverage evaluation and scene supplementation, and then test data with higher coverage can be obtained.

[0069] The content generation agent can be used to generate production flow information according to enhanced query information including test requirement information. The content generation agent can generate test data through enhanced retrieval, flow content generation, and tools. The tools can include real-time information processing, basic information processing, and production flow creation tools. The production flow creation tool can interact with the task management, and can create scenes through API and / or user interface (UI). The task management can be a module responsible for managing execution tasks, and can assist the content generation agent in implementing production flow scheduling and result collection.

[0070] The engine layer and the model layer are responsible for complex calculation and model derivation processing. The engine layer can include data analysis engines, data processing engines, data production flow engines, and flow orchestration engines. Each engine focuses on the implementation and scheduling of componentized data capabilities and data analysis algorithms, and the model layer focuses on the training and inference of large models. Both complement each other and support the implementation of upper layer functions.

[0071] The data layer is responsible for the interaction, fusion, and sedimentation of business data in the field. It can interface with multi-format business libraries and real-time data scene repositories to sediment data content, and can build multiple specialized knowledge bases to sediment platform and business knowledge, providing bottom layer content support for upper layer data capabilities and enhanced retrieval.

[0072] S203, according to the production flow information, calling at least part of the components in the component library, generating test data corresponding to the test requirement information.

[0073] For example, the agent can call at least part of the components in the component library to generate test data corresponding to the test requirement information according to the production flow information. Alternatively, the agent can input the production flow information into an engine with data processing capability to call each component in the component library. The test data corresponding to the test requirement information can be one type of test data or multiple types of test data. The generated test data can include a single data or multiple data. The test data corresponding to the test requirement information can be understood as a set of test data composed of multiple test data.

[0074] In a possible implementation, the production flow information includes a production flow construction protocol, the production flow construction protocol is used to indicate at least one step of generating test data corresponding to the test requirement information in a standardized protocol manner; and according to the production flow information, at least part of the components in the component library are invoked to generate the test data corresponding to the test requirement information, including: inputting the production flow information into a data production flow engine to obtain calling code of the at least part of the components to be invoked; and according to the calling code, at least part of the components in the component library are invoked to generate the test data corresponding to the test requirement information.

[0075] For example, the production flow construction protocol can be a protocol file in a self-defined protocol format, the agent can generate the production flow construction protocol, and indicate each step of generating the test data corresponding to the test requirement information to the tool through the production flow construction protocol. The data production flow engine can be a preset software tool, and can include a computer program. The production flow construction protocol can be parsed through the data production flow engine to obtain calling code capable of invoking the components, and the calling code is a computer program. The API and other components can be invoked through the calling code to query or generate the test data.

[0076] As shown in Figure 3 For the platform system implementing the service of the test data, the underlying capability support can be established in advance. The underlying data capability is incrementally expanded according to the business scenarios, and a complete component library of data capability is constructed. For example, for the out-of-town business, the time and space complexity problem of the travel scenario dependent data is focused on, the complete data capability is constructed, the dynamic, static, map, lane level and route data generation capabilities are supported, and the component library of lane level event construction, lane level attribute query and multi-traffic flow state construction capability is constructed. The standardized protocol for production flow construction can be defined in advance to flexibly adapt the combination of components.

[0077] When the test data corresponding to the test requirement information is generated, the components in the component library can be invoked through the preset data production flow engine to generate the test data corresponding to the test requirement information. Taking the travel scenario as an example, the production flow information generated by the agent can include a standardized production flow construction protocol. Through the underlying integrated flow engine, the components are bridged to make the data generated in the test scenario consistent and coherent, and the flexible adaptation and scheduling of the complex scenario are realized. Not only the efficient creation and intelligent scheduling of the production flow task of the complex scenario are realized, but also the task is fully supported from multiple dimensions to ensure the agile response and high flexibility to diversified requirements.

[0078] The standardized production flow construction protocol can be understood as abstracting the generation process of the test data of the test scene into the concept of "flow", and pre-defining the standardized production flow construction protocol to flexibly combine the capabilities of corresponding components to complete the construction of a specific production task. On this basis, the frequently used travel scene templates can be extracted from the flow to promote the efficient construction of the platform templated scene task flow and improve the flexibility and speed of scene building.

[0079] For example, the scheduling process of the test data generation of the entire test scene can be completed by the pre-constructed flow orchestration engine, thereby supporting the flexible, simple, efficient and realistic restoration of the complex test data of the test scene from the industry business starting point to the end point.

[0080] In a possible implementation, according to the calling code, at least part of the components in the component library are called to generate the test data corresponding to the test requirement information, including: inputting the calling code into the flow orchestration engine to obtain the calling sequence corresponding to at least one step; and calling at least part of the components in the component library in the calling sequence through the flow orchestration engine to generate the test data corresponding to the test requirement information.

[0081] As shown in Figure 3 and Figure 5 , the flow orchestration engine can run the calling code to reasonably arrange and schedule each component in the component library according to each step indicated by the production flow construction protocol to generate the test data. The test data can be one group or multiple groups, and each group of test data corresponds to one test scene or one test case.

[0082] For example, after the calling code is input into the flow orchestration engine, the program of the flow orchestration engine can run the calling code, and can arrange the component calling sequence according to the input parameters, output parameters and the timing relationship between different types of parameters in the calling code, and then can call each component to produce test data according to the calling sequence.

[0083] Taking the industry business scene as an example, complete data capability is the basis for supporting flexible construction of the travel scene. The embodiments of the present application can first focus on solving the space-time complexity of the travel scene dependent data, expand dozens of data capabilities such as dynamics, routes, lane levels, static attributes and differentiation, and construct a component library for generating data of hundreds of data categories and thousands of attribute types. The flow orchestration engine can realize the interactive processing and protocol adaptation of each category of data through custom post-processing. The flow orchestration engine can also set construction engines for data interpolation / thinning, data matching and attribute processing, and can also set query engines for protocol packaging, scene rule mapping and path mapping.

[0084] Exemplarily, the platform system implementing the test data service can be enabled through an open interface, allowing flexible combination of custom components to build personalized tasks, which can simplify business automation integration processes. And to solve the problem of complex interface protocol organization, the interface-to-data converter is proposed in the embodiments of the present application to generate preset format data from WEB UI to object notation. The interface-to-data converter has the conversion capability from interface to API body data, which can quickly generate a complex API request body based on an easily understandable interface operation method and seamlessly connect to the Open API, making the automation workflow more smooth and efficient. Among them, the interface-to-data converter can be understood as a tool that can generate test data for input operation objects based on input operations from the interface.

[0085] Exemplarily, the platform system implementing the test data service can provide a platformized operable session page, which can reduce the cost of data construction, realize easy operation, and realize the convenience of data creation.

[0086] In the embodiments of the present application, the method determines the test requirement information, inputs the test requirement information into the agent, so that the agent generates production process information corresponding to the test requirement information according to the knowledge base. Since the knowledge base of the agent includes the description information of a plurality of components in the preset component library, each component has the ability to obtain at least one test data in the business scenario corresponding to the test requirement information, therefore, in the case that the agent understands the test requirement, the whole process of generating test data can be intelligently streamlined by combining the ability of the component and the demand of the test activity for the test data, and then the production process information corresponding to the test requirement information is obtained, which includes the steps of generating the test data corresponding to the test requirement information. Therefore, according to the production process information, at least part of the components in the component library can be reasonably called to generate the test data corresponding to the test requirement information, and then the efficiency of generating the test data is improved.

[0087] It can be seen that the embodiments of the present application provide a natural language to test scene generation solution based on a large language model. The test data service capability is realized through the large model technology, and the data preparation efficiency and coverage problem in the test activity are solved.

[0088] The method of the embodiment can realize content generation and associated scene recommendation of natural language to scene production flow construction protocol. Around the chain thinking, three independent agents of intent recognition, content generation and scene recommendation are designed, each agent constructs a specialized knowledge base, and the RAG scheme is used to realize the enhancement of domain knowledge. By building a multi-agent system architecture, the intelligent agents can work together to generate content, and can call the data production flow engine to complete the generation of the overall test scene. For scene analysis, the core scene set is filtered based on business traffic, and then the scene recommendation is completed by mining the semantic and scene correlation of the large model.

[0089] The embodiment of the present application can realize one-key generation from natural language to complex test scene. Through a new process: starting from the intent recognition agent, passing through the double-agent of content generation and scene recommendation, and then to the scene generation engine, finally calling the rich component library, the core intent of the natural language instruction can be intelligently parsed, multiple data generation capabilities can be linked, and high-accuracy natural language intent recognition and second-level complex test data production flow construction can be realized.

[0090] Illustratively, the recall accuracy of the knowledge base is a key factor in determining the performance of large language models. The method provided by the embodiment of the present application can optimize the quality of knowledge content when knowledge is entered into the database through a knowledge base segmentation strategy, thereby improving the recall accuracy of the knowledge base.

[0091] In one possible implementation, the knowledge base and / or the intent knowledge base stores knowledge content in the following way: for any content to be entered into the database, if it is determined that the content to be entered into the database includes multiple knowledge related to semantics within the business scene, the multiple knowledge is determined as one knowledge content and stored in the knowledge base and / or the intent knowledge base; for any content to be entered into the database, if it is determined that the content to be entered into the database includes a single knowledge that is not related to semantics within the business scene, the single knowledge is determined as one knowledge content and stored in the knowledge base and / or the intent knowledge base; and a deduplication operation is performed on each stored knowledge content to remove knowledge contents with the same or similar content.

[0092] Illustratively, the knowledge base and / or the intent knowledge base can use a segmented, short segmented and refined knowledge base segmentation strategy when storing knowledge content. The content to be entered into the database can be any content related to the business scene, for example, it can be reference literature related to the business scene, etc. Taking the travel scene as an example, it can be a description file of the component capability range related to the travel scene, an attribute value list file, travel or road related data business knowledge, etc.

[0093] Figure 6 The schematic diagram of the retrieval enhanced generation provided by the embodiment of the present application is as follows: Figure 6As shown, when writing in the knowledge base, content quality operations can be performed first to determine high-quality content to be stored in the accumulated and precipitated data. In addition, key summaries, content classification and segment marker operations can be performed to pre-process the data of the knowledge to be stored. When the data is segmented and sliced to form multiple knowledge contents, the above-mentioned knowledge base segmentation strategy can be used to segment long related content, short independent content, and similar content.

[0094] Long segmentation can be understood as a strategy of not separating related content. The classification mechanism is used to write knowledge questions and answers, and the length of the segment is increased as much as possible to ensure that the content with nested relationship and strong correlation is in one slice to form a complete knowledge content. If the content is long, the connection between segments can be enhanced by adding a summary before the segment.

[0095] Short segmentation can be understood as a strategy of independent content with less coupling. For example, regional codes, data attributes, and other unique and independent content, the length of the segment is reduced to avoid multiple independent and irrelevant information in one knowledge content after slicing, which affects the recall accuracy.

[0096] Precise preparation can be understood as a strategy of similar content with less segmentation. The knowledge base content takes precision and not quantity. Similarity analysis can be used to remove duplicate knowledge contents with the same or similar content, which can effectively utilize each recall opportunity and avoid unnecessary segments being removed without recall due to too similar recall content.

[0097] Since the length of the segment and the completeness of the sliced content have a significant impact on recall during the construction of the knowledge base, in this embodiment, the knowledge content storage of the knowledge base and / or the intent knowledge base is performed by using the above-mentioned knowledge base segmentation strategy, which can effectively recall the knowledge segments in different application scenarios, improve the accuracy of knowledge recall, and further improve the accuracy of the retrieval and enhancement generation.

[0098] For example, in order to improve the accuracy of knowledge recall, a multi-path recall strategy can be used to recall in multiple ways.

[0099] In one possible implementation, the knowledge content associated with the natural language query information is retrieved to obtain the target knowledge content associated with the test requirement information, including: performing word segmentation on the natural language query information to obtain keywords, and retrieving in the intent knowledge base based on the keywords to obtain keyword retrieval results; performing semantic feature extraction on the natural language query information to obtain semantic information, and retrieving in the intent knowledge base based on the semantic information to obtain semantic retrieval results; and obtaining the target knowledge content associated with the test requirement information based on the keyword retrieval results and the semantic retrieval results.

[0100] For example, Figure 5As shown, the agent can adopt a multi-path recall strategy in parallel with vector recall and keyword recall in the retrieval enhancement generation process. Vector recall and keyword recall are two effective ways to recall knowledge content in the knowledge base, and both have advantages in different scenarios due to differences in underlying implementation principles. In the specific application scenario of natural language to data, for the accurate retrieval of some specific attributes or specific attributes in natural language, keyword recall is more optimal; for the retrieval of similar semantics, vector recall is more intelligent. Therefore, the method combines the two to form a multi-path recall strategy.

[0101] As an example, key words can be obtained by word segmentation based on natural language query information, and key word retrieval results can be obtained by retrieval in the intent knowledge base based on the key words. Semantic features can be extracted based on the natural language query information, such as embedding operation on the natural language query information to obtain corresponding semantic information, and semantic retrieval results can be obtained by retrieval in the intent knowledge base based on the semantic information.

[0102] When the target knowledge content associated with the test requirement information is obtained based on the key word retrieval results and the semantic retrieval results, the key word retrieval results and the semantic retrieval results can be first subjected to a re-ranking operation. For example, N knowledge contents are taken from the key word retrieval results and the semantic retrieval results respectively, and the distance score is calculated to remove duplicates and sort by size, and the top M results are taken as the target knowledge content associated with the test requirement information, thereby realizing segmentation and cleaning of the recalled key word retrieval results and semantic retrieval results, and obtaining the recalled knowledge content. Wherein, N is any positive integer, M is a positive integer less than or equal to N, and M and N can be preset according to the business scenario. Based on this, the target knowledge content with high accuracy and high diversity can be returned.

[0103] As an example, as shown in Figure 6 The pre-agent can receive query information and output enhanced query information, and the pre-agent can be an intent recognition agent. When the enhanced query information is input into the content generation agent for retrieval enhancement generation, a multi-path recall strategy in parallel with vector recall and keyword recall can also be adopted to improve the accuracy and diversity of the recall results. If other agents are included in the platform system that implements the test data service, a multi-path recall strategy can also be used for knowledge content recall.

[0104] In a possible implementation, the method further includes: inputting the test requirement information into a scene recommendation agent, and searching for target associated content that has content association with the test requirement information in a scene set through the scene recommendation agent, the scene set being a set including a plurality of test data; inputting the target associated content and a prompt word corresponding to the target associated content into a natural language model of the scene recommendation agent for content generation to obtain recommendation information related to the content of the test requirement information; and returning the recommendation information.

[0105] As shown in Figure 3 and Figure 4 , the scene recommendation agent can include an LLM, and inputting enhanced query information including the test requirement information into the scene recommendation agent can enable the scene recommendation agent to search for target associated content that has content association with the test requirement information in a scene set, and then input the target associated content in combination with a prompt word into the LLM to generate recommendation information related to the content of the test requirement information. The recommendation information can be a keyword, answer content and / or test data associated with the test requirement information.

[0106] By way of example, scene recommendation can refer to recommending relevant test scenes in combination with real-world factors to improve test scene coverage. The scene set can be understood as a scene denominator for providing a plurality of test data.

[0107] In a possible implementation, the scene set is obtained in the following manner: obtaining a plurality of business data, the business data being data obtained when the business scene is actually applied; standardizing the plurality of business data into a plurality of standard business data in a unified data format; performing data analysis and processing on the plurality of standard business data to obtain tree structure data; performing business feature extraction on the tree structure data to obtain a plurality of business features, the business features being used to represent at least one feature in the business scene; performing feature clustering according to the plurality of business features, and obtaining the scene set based on the feature clustering result.

[0108] Figure 7 A schematic diagram of an implementation of the associated scene recommendation capability provided by the embodiments of the present application is shown in Figure 3 to Figure 7 , by establishing a heterogeneous data integration capability and a traffic tree analysis model, an interactive intent recognition, feature analysis and associated scene analysis are completed in combination with a large model, and finally a scene recommendation agent is encapsulated to provide a natural language to associated scene recommendation capability.

[0109] In the case where the format and content of the business data are specific, the heterogeneous data integration can be performed on each business data to obtain the scene set.

[0110] Two aspects of data analysis can be performed on the data. One aspect is to extract a non-navigation point announcement tone (NDT) from the data, obtain corresponding keywords, and perform data mapping and weight ordering on each keyword to obtain an ordered data result. Another aspect is to support access to various data formats such as structured or semi-structured data for data storage, transmission or exchange through multi-format input compatibility. Data standardization preprocessing can be completed through data cleaning, standardization and conversion. After preprocessing, feature extraction processing can be performed on the preprocessed data through a traffic tree analysis model.

[0111] As shown in Figure 5 , through the data analysis engine, full table analysis, feature matching and traffic tree analysis algorithms can be performed on the data. Through the data processing engine, heterogeneous data integration, data preprocessing and feature extraction operations can be performed on the data. As shown in Figure 7 , the preprocessed data can be subjected to full table analysis to convert complex data structures into data structures suitable for analysis, for example, into a syntax tree type structure. Then, feature extraction can be performed from the full-leaf nodes to obtain business feature variables.

[0112] For example, data structure conversion can be the conversion of nested, hierarchical data structures into a flat, single-layer data structure. In processing more complex object notation text type semi-structured data, text type structured data or other nested format data, through data structure conversion, data query, analysis and visualization can be more easily performed. For example, for an object notation text type semi-structured data object containing nested arrays or objects, through data structure conversion, the object notation text type semi-structured data object is converted into a simple key-value pair list or table form. In the process of implementing data structure conversion, the following operations can be implemented, such as an unfolding nested structure operation to extract elements in nested objects or arrays, a merging field operation to merge fields in nested structures into top-level structures, and a redundancy removal operation to remove duplicate or redundant data.

[0113] Through the operations of business experience input, business configuration, feature extraction and feature matching, the ordered data result and the obtained business feature variables are combined with experience to supplement the association logic between key features in the business data, and a specific business scenario set is constructed. For example, in the obtained scenario set, data that can be accessed through a distributed data layer can be obtained through data filtering and standardization operations.

[0114] As shown in Figure 4 and Figure 7As shown, for the outbound service scenario, the scene recommendation agent can guide the large model to think through the relevance analysis of the prompt words when performing associated scene recommendation, extract geographic information, traffic information, traffic events, dynamic data, map data, and other key information related to travel data from the received natural language test requirements, perform relevance calculation in combination with the foregoing scene set, and finally select the top (TOP) result that meets the test activity requirements after filtering. When performing relevance analysis, at least one of information matching, association rules, weight sorting, and difference set sorting can be used to achieve this. According to the scene application, the TOP mechanism, standardized output, data construction, and interactive feedback can be used to obtain the output recommendation information.

[0115] In the embodiments of the present application, relying on the fusion exploration of the traffic tree analysis algorithm and the data neural network, a scene recommendation mechanism capable of deeply understanding the customized business scene coverage requirements is proposed. Specifically, by identifying the test requirements corresponding to the natural language, analyzing in combination with the historical data mode and the business logic, the accurate matching of the associated test data and the minimum test data recommendation based on coverage can be realized, which can improve the test efficiency and the coverage breadth of the test.

[0116] In a possible implementation, the method further includes: storing the test requirement information and the corresponding test data into a scene library; and in a subsequent round of test data generation, in a case where the test requirement information determined in the current round is the same as any test requirement information stored in the scene library, determining the test data corresponding to any test requirement information as the test data corresponding to the test requirement information in the current round.

[0117] The scene library can be a collection of multiple test requirement information and corresponding test data. After storing multiple test requirement information and corresponding test data, when there is subsequent query information inputting the same test requirement information, the stored test data can be returned to improve the speed of outputting test data.

[0118] In a possible implementation, the scene library further stores a label corresponding to the test data, and the label is used to indicate that the corresponding test data is a positive sample or a negative sample; the method further includes: iteratively optimizing the agent by supervised training according to at least one test data and the label corresponding to the test data in the scene library.

[0119] As shown in Figure 3 and Figure 5 Based on the generated test data, a scene library can be built. For example, the natural language is stored together with the generated test data and the generated recommendation information in the scene library, the positive and negative samples are retained, which can be used for subsequent rapid feedback of the same scene and iterative optimization of the overall effect of natural language to data scene generation.

[0120] Exemplarily, the iterative optimization can be performed by using an iterable agent effect evaluation mechanism. For example, a training set is constructed based on high-frequency scenarios to evaluate the iteration effect of different versions. In the agent debugging process, the effect of a single iteration cannot effectively reflect the overall effect change. Therefore, the following two mechanisms can be used to effectively measure the iteration effect of different versions.

[0121] The first mechanism is an A / B comparison mechanism. The high-frequency scenario set can be extracted based on the use of traffic and the aforementioned expert experience scenario set, and converted into true values and natural language query information, which is used as a training set to perform A / B comparison for each iteration to analyze bad cases and determine the optimization direction. Among them, A and B can be understood as the performance of the agent before and after iteration.

[0122] The second mechanism is an interactive feedback mechanism. When generating test data, an interactive portal is opened for each test data, and information interaction can be performed on each test data. For example, if the test activity triggers positive interaction on the test data, it can be indicated that the test data meets the needs of the test activity, and the test data can be marked as a positive sample. If the test activity triggers negative interaction on the test data, it can be indicated that the test data does not meet the needs of the test activity, and the test data can be marked as a negative sample. Through the interactive feedback mechanism, positive and negative samples based on the generated test data can be collected to obtain a scenario library for iterative optimization.

[0123] Exemplarily, the supervised training can be any training method that can train a model based on positive sample test data and negative sample test data. For example, the agent can be trained by a contrastive learning method. Exemplarily, at least one of the content generation agent, the intent recognition agent, and the scenario recommendation agent can be individually or jointly supervised trained by a similar method to iteratively optimize the model parameters or structure of the agent.

[0124] Based on this, the positive and negative samples are accumulated through interactive feedback, and the optimization strategy is reciprocated, and the training data set is optimized and enriched, laying a data foundation for subsequent intent enhancement fine-tuning stage. By establishing an effect evaluation mechanism, the iteration effect of different versions can be effectively measured and personalized scenario adaptation capability can be provided.

[0125] In a possible implementation, the method further includes: performing a post-processing operation on the test data to obtain post-processed test data and output, the post-processing operation including at least one of a data format conversion operation, a data compression operation, a data encryption operation, and a data protocol encapsulation operation.

[0126] Based on the adaptation paradigm of LLM, custom test data post-processing capabilities such as format conversion, compression encryption, protocol encapsulation, and general interaction functions such as HTTP sending can be supported, realizing intelligent linkage of the SaaTS layer.

[0127] For example, in a programming language, data format conversion can be implemented using libraries in the programming language. For example, in a programming language, data processing and analysis open source libraries can be used to convert between multiple preset data formats, such as between structured and semi-structured data formats; or custom scripts can be used to parse and reconstruct data to achieve data format conversion; or other methods can be used.

[0128] For example, symmetric encryption algorithms or asymmetric encryption algorithms can be used to implement data encryption operations on test data; or other methods can be used. Network protocol stacks, application layer protocols, or custom protocols can be used to implement data protocol encapsulation operations on test data.

[0129] Large model capability-driven test data post-processing, test case generation and execution can realize operation intelligent linkage. This process not only covers in-depth understanding and prediction of specific operation logic, but also realizes intelligent linkage between each link, ensuring that the test scheme can be highly matched with personalized needs and business specificity, thereby improving the effectiveness and pertinence of testing.

[0130] For example, in existing technologies, test data generation products mainly provide single-dimensional data queries in the form of APIs or pages, and the data richness is not enough, and multi-environment, custom construction, and combined scenario generation are not supported. As the complexity and richness of business scenarios improve, on the one hand, test scenarios have evolved from relatively single data scenarios to complex scenarios that rely on multi-dimensional data generation and fusion, which poses challenges to the efficiency and difficulty of test scenario preparation in test activities. On the other hand, to ensure sufficient test scenario coverage, the ability to discover the relevance between scenarios is increasingly important.

[0131] The method provided by the above embodiments can provide test data service capability, realize more flexible and harmonious combination of large models and test scenario generation capability, and thus solve the data preparation efficiency and coverage problems in test activities. The method aims to flexibly assemble the test case generation capability domain platform and simplify the use of complex scenarios, realizes natural language to test scenario generation capability, and has high adaptability in the process of test intelligentization.

[0132] The method provided by the embodiments of the present application can support the construction of natural language-based test scenario generation capability, which can be combined with large models to more flexibly generate test scenarios and improve test efficiency and coverage. This capability can be applied to the test case generation field to help testers generate high-quality test cases faster. In addition, this capability can also be applied to various quality domain platforms to help related fields generate the required test scenarios faster and improve test efficiency and quality.

[0133] For example, compared with a multi-agent system, when using LLM+prompt+tool to generate test data for an out-of-town business scenario based on a single agent, there are some deficiencies. For example, it is difficult for a single agent to meet the high requirements of complex scenario generation in terms of scalability and stability. Assuming that the out-of-town scenario involves dynamic data such as events, weather, road conditions, and traffic lights, and that there are flexible combinations of hundreds of static attributes and routes in driving, cycling, and walking scenarios, the number of underlying component capabilities is dozens, the interaction methods are asynchronous and synchronous, and there are strong requirements for the calling order between components. The function call method of external capability calling will have instability and asynchronous response time, which will increase the timeout probability of large model interaction. For another example, the architecture of a single agent has at least two limitations when facing such complex tasks: first, the accuracy of knowledge base retrieval is poor, and it is difficult to effectively cover all related knowledge; second, large models are prone to omissions when processing highly interwoven information, and are prone to assumptions, false guidance, and other situations, which affect the overall results.

[0134] The multi-agent system provided by the embodiments of the present application can better make up for the above deficiencies.

[0135] Taking the out-of-industry business scenario as an example, in terms of efficiency improvement, for the natural language to data function, the artificial step-by-step data construction can be converted to natural language driven one-key task generation, and the production flow of a complex scenario can be quickly constructed. Through the improvement of the one-key generation capability of the complex scenario, the test data generation efficiency can be improved in batches through the batch coverage of the test scenario, for example, multiple versions of an application program can be updated in a unit of time. The changed data between the versions of the application program is the part that needs to be checked every time the application program is released, and by providing a data version difference link query capability combination route generation capability, the release check period can be shortened when applied to the client. For another example, by providing a route generation capability, test expansion automation can be provided by applying to the test case client, and the development cycle can be shortened.

[0136] The method provided by the embodiments of the present application can be applied to a client automation platform, a test automation platform, or in combination with other intelligent agents to realize test case generation.

[0137] Figure 8 A structural diagram of a test data generation device provided by the embodiments of the present application is shown in FIG. 8, and the embodiments of the present application further provide a test data generation device 800, which comprises: Figure 8

[0138] A determination module 801 is configured to determine test requirement information.

[0139] A processing module 802 is configured to input the test requirement information into an intelligent agent, so that the intelligent agent generates production flow information corresponding to the test requirement information according to a knowledge base. The knowledge base comprises description information of a plurality of components in a preset component library, the components are used to obtain at least one kind of test data in a business scenario corresponding to the test requirement information, and the production flow information comprises at least one step of generating the test data corresponding to the test requirement information.

[0140] The processing module 802 is further configured to call at least part of the components in the component library according to the production flow information, and generate the test data corresponding to the test requirement information.

[0141] Optionally, the production flow information comprises a production flow construction protocol, and the production flow construction protocol is used to indicate the at least one step of generating the test data corresponding to the test requirement information in a standardized protocol manner.

[0142] The processing module 802 is specifically configured to input the production flow information into a data production flow engine to obtain calling code of the at least part of the components to be called, and call at least part of the components in the component library according to the calling code to generate the test data corresponding to the test requirement information.

[0143] Optionally, the processing module 802 is specifically configured to:​

[0144] The calling code is input into the process arrangement engine to obtain a calling sequence corresponding to at least one step;

[0145] At least part of the components in the component library are called by the process arrangement engine according to the calling sequence, and test data corresponding to the test requirement information is generated.

[0146] Optionally, the determining module 801 is specifically configured to: acquire natural language query information; input the natural language query information into an intent recognition intelligent agent to obtain an intent recognition result representing an intent, the intent recognition intelligent agent being an intelligent agent that analyzes the natural language query information through a natural language model to recognize the intent; and in a case where the intent recognition result includes an intent of generating test data, determine test requirement information according to the natural language query information through the intent recognition intelligent agent.

[0147] Optionally, the intent recognition intelligent agent includes an intent knowledge base for intent recognition, and the determining module 801 is specifically configured to: retrieve, through the intent recognition intelligent agent, knowledge content associated with the natural language query information in the intent knowledge base according to the natural language query information to obtain target knowledge content associated with the test requirement information; and input the target knowledge content and a prompt word corresponding to the target knowledge content into the natural language model of the intent recognition intelligent agent to obtain the test requirement information, the prompt word being used to guide the natural language model to generate content in combination with the target knowledge content.

[0148] Optionally, the knowledge base and / or the intent knowledge base stores the knowledge content in the following manner: for any to-be-stored content, in a case where it is determined that the to-be-stored content includes a plurality of knowledge that are semantically related within a business scenario, the plurality of knowledge are determined as one knowledge content and are stored in the knowledge base and / or the intent knowledge base; for any to-be-stored content, in a case where it is determined that the to-be-stored content includes a single knowledge that is semantically unrelated within a business scenario, the single knowledge is determined as one knowledge content and is stored in the knowledge base and / or the intent knowledge base; and a de-duplication operation is performed on each stored knowledge content to remove knowledge contents that are the same or similar.

[0149] Optionally, the determining module 801 is specifically configured to: perform word segmentation on the natural language query information to obtain a keyword, and perform retrieval in the intent knowledge base based on the keyword to obtain a keyword retrieval result; perform semantic feature extraction on the natural language query information to obtain semantic information, and perform retrieval in the intent knowledge base based on the semantic information to obtain a semantic retrieval result; and obtain target knowledge content associated with the test requirement information according to the keyword retrieval result and the semantic retrieval result.

[0150] Optionally, the processing module 802 is further configured to input the test requirement information into the scene recommendation agent, and search for target associated content that has content correlation with the test requirement information in a scene set through the scene recommendation agent, the scene set being a set including a plurality of test data; input the target associated content and prompt words corresponding to the target associated content into a natural language model of the scene recommendation agent for content generation, to obtain recommendation information related to the content of the test requirement information; and return the recommendation information.

[0151] Optionally, the scene set is obtained by: obtaining a plurality of business data, the business data being data obtained when the business scene is actually applied; standardizing the plurality of business data into a plurality of standard business data in a unified data format; performing data analysis and processing on the plurality of standard business data to obtain tree structure data; performing business feature extraction on the tree structure data to obtain a plurality of business features, the business features being used to represent at least one feature in the business scene; performing feature clustering according to the plurality of business features, and obtaining the scene set based on the feature clustering result.

[0152] Optionally, the processing module 802 is further configured to store the test requirement information and the corresponding test data into a scene library; and in a subsequent round of test data generation process, in a case where the test requirement information determined in the current round is the same as any test requirement information stored in the scene library, determine the test data corresponding to any test requirement information as the test data corresponding to the test requirement information in the current round.

[0153] Optionally, the scene library further stores a label corresponding to the test data, the label being used to indicate that the corresponding test data is a positive sample or a negative sample; and the processing module 802 is further configured to iteratively optimize the agent in a supervised training manner according to at least one test data and the label corresponding to the test data in the scene library.

[0154] Optionally, the processing module 802 is further configured to perform post-processing operation on the test data to obtain post-processed test data and output, the post-processing operation including at least one of data format conversion operation, data compression operation, data encryption operation and data protocol encapsulation operation.

[0155] The test data generation apparatus provided by the embodiment of the present application can be used to execute the technical solutions of the test data generation method in any of the above embodiments of the present application, and has similar implementation principles and technical effects, which will not be described here again.

[0156] Figure 9 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1. Figure 9As shown, the electronic device of the embodiment can include at least one processor 901, and a memory 902 connected with the at least one processor in communication; wherein the memory 902 stores instructions executable by the at least one processor 901, and the instructions are executed by the at least one processor 901 to enable the electronic device to perform the method of any of the preceding embodiments.

[0157] Optionally, the memory 902 can be independent or integrated with the processor 901.

[0158] The implementation principle and technical effects of the electronic device provided by the embodiment can be referred to the foregoing embodiments, and will not be described here.

[0159] The embodiment of the application further provides a computer readable storage medium, and the computer readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the method of any of the preceding embodiments is implemented.

[0160] The embodiment of the application further provides a computer program product, and the computer program product includes a computer program, and when the computer program is executed by the processor, the method of any of the preceding embodiments is implemented.

[0161] In several embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative, for example, the division of the modules is only a logical function division, and actual implementation can have another division manner, for example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0162] The integrated modules implemented in the form of software function modules can be stored in a computer readable storage medium. The software function modules described above are stored in a storage medium, and include instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute part of the steps of the method described in each embodiment of the application.

[0163] It should be appreciated that referenced processors above can be a central processing unit (CPU), but also other general purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), and so forth. The general purpose processor can be a microprocessor or the processor can be any conventional processor. Steps of methods disclosed in connection with the present application can be directly implemented in hardware processor, executed in a processor using a hardware and software module combination, or implemented in a combination of the two. The memory can include a RAM (random access memory) and can also include a NVM (non-volatile memory), such as at least one disk storage, and can also be a U disk, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk, and so forth.

[0164] The storage medium described above can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general purpose or special purpose computer.

[0165] An exemplary storage medium is coupled to the processor so that the processor can read information from, and write information to, the storage medium. Of course, the storage medium can be part of the processor. The processor and the storage medium can be located in an application specific integrated circuits (ASIC). Of course, the processor and the storage medium can exist as discrete components in an electronic device or host device.

[0166] It should be noted that, in this document, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0167] The above-mentioned sequence numbers of embodiments of the present application are only for description, and do not represent advantages or disadvantages of the embodiments.

[0168] Those skilled in the art can clearly understand the above-mentioned embodiment method from the description of the above embodiments, which can be realized by software and necessary general hardware platform, of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, or optical disk) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device) to execute the methods described in the various embodiments of the present application.

[0169] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for generating test data, characterized in that, The method includes: Determine test requirements; The test requirement information is input into the intelligent agent, so that the intelligent agent generates production process information corresponding to the test requirement information according to the knowledge base; wherein, the knowledge base includes description information of multiple components in a preset component library, the components are used to obtain at least one test data in the business scenario corresponding to the test requirement information, and the production process information includes at least one step of generating test data corresponding to the test requirement information. Based on the production process information, at least some components in the component library are called to generate test data corresponding to the test requirement information.

2. The method according to claim 1, characterized in that, The production process information includes a production flow construction protocol, which is used to indicate at least one step in generating test data corresponding to the test requirement information in a standardized protocol manner. The step of generating test data corresponding to the test requirement information by calling at least some components in the component library based on the production process information includes: The production process information is input into the data production process engine to obtain the calling code of at least some of the components to be called; Based on the calling code, at least some components in the component library are invoked to generate test data corresponding to the test requirement information.

3. The method according to claim 2, characterized in that, The step of calling at least some components in the component library according to the calling code to generate test data corresponding to the test requirement information includes: The calling code is input into the flow orchestration engine to obtain the calling order corresponding to the at least one step; The process orchestration engine invokes at least some components in the component library according to the invocation order to generate test data corresponding to the test requirement information.

4. The method according to any one of claims 1-3, characterized in that, The determination of test requirement information includes: Obtain natural language query information; The natural language query information is input into the intent recognition agent to obtain the intent recognition result representing the intent. The intent recognition agent is an agent that parses the natural language query information through a natural language model to recognize the intent. If the intent recognition result includes the intent to generate test data, the intent recognition agent determines the test requirement information based on the natural language query information.

5. The method according to claim 4, characterized in that, The intent recognition agent includes an intent knowledge base for intent recognition. The step of determining the test requirement information based on the natural language query information using the intent recognition agent includes: The intent recognition agent retrieves knowledge content related to the natural language query information from the intent knowledge base based on the natural language query information, thereby obtaining target knowledge content related to the test requirement information. The target knowledge content and the corresponding prompt words are input into the natural language model of the intent recognition agent to obtain the test requirement information. The prompt words are used to guide the natural language model to generate content in combination with the target knowledge content.

6. The method according to claim 5, characterized in that, The knowledge base and / or the intent knowledge base stores knowledge content in the following manner: For any content to be added to the database, if it is determined that the content to be added to the database includes multiple semantically related knowledge within the business scenario, the multiple knowledge is determined as a single knowledge content and stored in the knowledge base and / or the intent knowledge base. For any content to be added to the database, if it is determined that the content to be added to the database includes a single piece of knowledge that is semantically unrelated within the business scenario, the single piece of knowledge is determined as a knowledge content and stored in the knowledge base and / or the intent knowledge base; Perform deduplication on the stored knowledge content to remove identical or similar content.

7. The method according to claim 5, characterized in that, The retrieval of knowledge content associated with the natural language query information, to obtain target knowledge content associated with the test requirement information, includes: Keywords are obtained by word segmentation based on the natural language query information, and the keywords are then used to search the intent knowledge base to obtain keyword search results; Semantic information is obtained by extracting semantic features from the natural language query information, and semantic retrieval results are obtained by searching the intent knowledge base based on the semantic information. Based on the keyword search results and the semantic search results, the target knowledge content associated with the test requirement information is obtained.

8. The method according to any one of claims 1-3, characterized in that, The method further includes: The test requirement information is input into the scene recommendation agent, and the scene recommendation agent retrieves target related content that is related to the test requirement information in the scene set. The scene set is a collection of multiple test data. The target-related content and the corresponding prompt words are input into the natural language model of the scene recommendation agent to generate content and obtain recommendation information related to the test requirement information content. Return the recommended information.

9. The method according to claim 8, characterized in that, The set of scenarios is obtained in the following way: Acquire multiple business data, which are data obtained during actual application in the business scenario; The multiple business data are standardized into multiple standard business data in a unified data format; Data analysis and processing are performed on the multiple standard business data to obtain tree-structured data; Business features are extracted from the tree-structured data to obtain multiple business features, which are used to characterize at least one feature in the business scenario. The scenario set is obtained by performing feature clustering based on the multiple business characteristics and the feature clustering results.

10. The method according to claim 8, characterized in that, The method further includes: The test requirements information and the corresponding test data are stored in the scenario library; In the subsequent test data generation process, if the test requirement information determined in the current round is the same as any test requirement information already stored in the scenario library, the test data corresponding to that test requirement information will be determined as the test data corresponding to the test requirement information of the current round.

11. A test data generation device, characterized in that, The device includes: The determination module is used to determine test requirement information; A processing module is used to input the test requirement information into an intelligent agent, so that the intelligent agent generates production process information corresponding to the test requirement information based on a knowledge base; wherein, the knowledge base includes description information of multiple components in a preset component library, the components are used to generate at least one type of test data in the business scenario corresponding to the test requirement information, and the production process information includes at least one step of generating the test data corresponding to the test requirement information. The processing module is further configured to, based on the production process information, call at least some components in the component library to generate test data corresponding to the test requirement information.

12. An electronic device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the electronic device to perform the method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-10.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-10.