Dynamic model-based test big data optimization method and system
Through the test big data optimization method based on dynamic model, test data with similarity to real data is generated with less than the threshold, and the test environment is dynamically adjusted, and data quality detection is carried out in combination with knowledge graph and blockchain technology, which solves the problems of poor authenticity of test data and bottlenecks in the data management platform performance in the existing technology, and realizes efficient and accurate test data generation and system performance evaluation.
Patent Information
- Application Number
- CN202510577973.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The test data generated by existing big data testing methods are poor in authenticity and difficult to reflect the performance of the system in the real environment. The data management platform has a longer response time when storing and retrieving data, which affects the testing efficiency.
The test big data optimization method based on dynamic models is adopted, and the test data with a similarity of less than the threshold to the real data is generated by the access to the optimized generator and discriminator, and the test environment is dynamically adjusted through the user operation behavior model and the human-computer interaction model to match the production environment. Data quality detection and problem traceability are used to use knowledge graph-based data quality testing rule base and blockchain technology.
Generate highly realistic test data, improve test accuracy, accurately evaluate system performance, quickly locate data quality problems, optimize test data quality, and ensure that the big data system operates stably, reliably and efficiently in complex task scenarios.
Smart Images

Figure CN120104978A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of data testing technology, and specifically to a test big data optimization method and system based on a dynamic model. Background Art
[0002] With the continuous development of science and technology, the current big data testing field often uses big data generation tools such as DataFactory, combined with algorithms such as random number generation, to generate large-scale and diversified test data based on the business rules and data characteristics of the software. However, the authenticity of the data generated by this method is poor, and there is still a gap between the generated data and the actual business data, which is difficult to reflect the performance of the system in the real environment. Moreover, as the amount of test data continues to grow, performance bottlenecks may occur in the data management platform. When storing and retrieving data, especially when performing complex queries, the response time will be significantly extended, affecting the test efficiency. In addition, there is a deviation between the test environment and the production environment, which affects the accurate evaluation of system performance. At the same time, the existing performance indicators have limitations. Commonly used indicators such as response time and throughput cannot fully and accurately measure some complex business scenarios and user experience indicators, such as the interactive fluency of the system and the loading speed of data visualization. In addition, the test rules also have limitations. They are usually formulated based on existing business needs and known data problems, and it is difficult to cover all potential data quality risks. Therefore, there is an urgent need for a test big data optimization method based on a dynamic model that can optimize the quality of test data, accurately evaluate system performance, and efficiently ensure data quality. Summary of the invention
[0003] The embodiment of this specification provides a test big data optimization method and system based on a dynamic model, and its technical solution is as follows: In the first aspect, the embodiments of the present specification provide a test big data optimization method based on a dynamic model, including: accessing an optimized generator and a discriminator, and generating test data whose similarity with real data is less than a similarity threshold through the optimized generator and the discriminator; accessing a user operation behavior model and a human-computer interaction model, and dynamically adjusting the test environment to match the production environment through the user operation behavior model and the human-computer interaction model; based on the test environment and the test data, performing data quality detection and problem tracing through a data quality test rule base based on a knowledge graph and blockchain technology to obtain optimized test data.
[0004] On the second aspect, the embodiments of the present specification provide a test big data optimization system based on a dynamic model, including: a test data generation module, which is used to access the optimized generator and discriminator, and generate test data whose similarity with real data is less than a similarity threshold through the optimized generator and discriminator; an adaptive performance test environment simulation module, which is used to access the user operation behavior model and the human-computer interaction model, and dynamically adjust the test environment to match the production environment through the user operation behavior model and the human-computer interaction model; an intelligent dynamic data quality testing module, which is used to perform data quality detection and problem tracing based on the test environment and test data, through a data quality testing rule base based on a knowledge graph and blockchain technology, to obtain optimized test data.
[0005] The beneficial effects brought by the technical solutions provided by some embodiments of this specification include at least: The embodiments of this specification can be connected to the optimized generator and discriminator, and generate test data with a similarity with the real data less than a similarity threshold through the optimized generator and discriminator, thereby providing data close to the real task scenario for performance and quality testing; then, the embodiments of this specification can also be connected to the user operation behavior model and the human-computer interaction model, and dynamically adjust the test environment to match the production environment through the user operation behavior model and the human-computer interaction model. The embodiments of this specification dynamically adjust the performance test environment according to the production environment, thereby providing a precise testing environment for other modules; in addition, the embodiments of this specification can also use the test environment and test data to perform data quality detection and problem tracing through a data quality test rule library based on a knowledge graph and blockchain technology.
[0006] The embodiments of this specification can not only generate highly realistic test data and improve test accuracy, but also accurately evaluate system performance and quickly locate data quality issues. The embodiments of this specification have the functions of optimizing test data quality, accurately evaluating system performance, and efficiently ensuring data quality. They are suitable for various scenarios with high requirements for big data processing and analysis, such as e-commerce platform big data analysis systems, financial institution data processing systems, etc., effectively ensuring that big data systems can run stably, reliably and efficiently in complex task scenarios, and meeting the quality requirements of big data systems in different industries. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0008] Figure 1This is a schematic diagram of an application scenario of a test big data optimization method based on a dynamic model provided in this specification.
[0009] Figure 2 This is a flowchart of a test big data optimization method based on a dynamic model provided in this specification.
[0010] Figure 3 It is a flowchart of the generator and discriminator optimization process provided in this manual.
[0011] Figure 4 This is a flowchart of the dynamic adjustment of the test environment provided in this manual.
[0012] Figure 5 This manual provides a flowchart for establishing a human-computer interaction model based on production environment data.
[0013] Figure 6 This is a flowchart of establishing a data quality testing rule base based on a knowledge graph provided in this manual.
[0014] Figure 7 This is a flowchart of problem tracing through blockchain technology provided in this manual.
[0015] Figure 8 This is a structural diagram of a test big data optimization system based on a dynamic model provided in this specification.
[0016] Fig. 9 This is a schematic diagram of the structure of an electronic device provided in this manual. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of this specification will be described clearly and completely below in conjunction with the drawings in the embodiments of this specification.
[0018] The terms "first", "second", etc. in the description and claims of this specification and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variation thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices.
[0019] The data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions.
[0020] Multiple embodiments of this specification provide a test big data optimization method based on a dynamic model. The executor of the test big data optimization method based on a dynamic model can be a test big data optimization system based on a dynamic model provided by an embodiment of the present invention.
[0021] Before this specification elaborates on a test big data optimization method based on a dynamic model in combination with one or more embodiments, it first introduces the application scenario of the test big data optimization method based on a dynamic model.
[0022] See also Figure 1 , Figure 1 A schematic diagram of an application scenario of a test big data optimization method based on a dynamic model provided in an embodiment of the present invention. In this embodiment, a test big data optimization system 100 based on a dynamic model may include a big data test optimization device 110, a user terminal 120, etc. The big data test optimization device 110 and the user terminal 120 are connected in communication.
[0023] In this embodiment, the user terminal 120 may be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC), etc. The user terminal 120 may upload real data such as user operation behavior to the big data test optimization device 110, etc.
[0024] The big data test optimization device 110 of the embodiment of this description can be a single server or a server cluster composed of multiple servers, and the multiple servers are used to implement the dynamic model-based test big data optimization method of the present application.
[0025] In this embodiment, the big data test optimization device 110 can be connected to the optimized generator and discriminator, and generate test data whose similarity with the real data is less than the similarity threshold through the optimized generator and discriminator; connect to the user operation behavior model and the human-computer interaction model, and dynamically adjust the test environment to match the production environment through the user operation behavior model and the human-computer interaction model; based on the test environment and test data, data quality detection and problem tracing are performed through a data quality test rule base based on the knowledge graph and blockchain technology to obtain optimized test data, etc.
[0026] It should be noted that Figure 1The scenario diagram of the dynamic model-based test big data optimization system shown is only an example. The dynamic model-based test big data optimization system and scenario described in the embodiment of the present invention are intended to more clearly illustrate the technical solution of the embodiment of the present invention, and do not constitute a limitation on the technical solution provided by the embodiment of the present invention. Ordinary technicians in this field can know that with the evolution of the dynamic model-based test big data optimization system and the emergence of new scenarios, the technical solution provided by the embodiment of the present invention is also applicable to similar technical problems.
[0027] See also Figure 2 , Figure 2 is a flow chart of a test big data optimization method based on a dynamic model provided by an embodiment of the present invention. The test big data optimization method based on a dynamic model can be Figure 1 The test big data optimization system 100 based on the dynamic model is shown. The test big data optimization method based on the dynamic model may at least include the following steps: 200. Connect the optimized generator and discriminator, and generate test data whose similarity with the real data is less than a similarity threshold through the optimized generator and discriminator.
[0028] In some embodiments, see Figure 3 , Figure 3 : is a flow chart of the generator and discriminator optimization process provided by an embodiment of the present invention. Before connecting the optimized generator and discriminator, it includes: 2000. Obtain real data and random noise based on the test scenario, and establish a generator and a discriminator based on the deep learning network; 2010. Determine the objective function between the generator and the discriminator; 2020. Solve the objective function based on real data and random noise to obtain the optimized generator and the optimized discriminator.
[0029] In this embodiment, the real data of the test scenario is the data of the actual production environment, which is used to truly reflect the task scenario and user behavior. The big data test optimization device 110 can first send an inquiry message about obtaining real data to the user terminal 120. After the user terminal 120 receives the inquiry message, after the user agrees to upload the real data, the user terminal 120 can send the real data of the test scenario to the big data test optimization device 110. The random noise of this embodiment is a random vector, whose elements can be extracted from a probability distribution (such as a standard normal distribution or a uniform distribution).
[0030] In this embodiment, the goal of the generator is to generate data similar to real data. In this embodiment, the generator can be established through a deep learning network, and the generator includes but is not limited to a fully connected neural network, a convolutional neural network, a recurrent neural network, a variational autoencoder, a Transformer model, etc. The goal of the discriminator is to distinguish between generated data and real data. In this embodiment, the discriminator can be established through a deep learning network, and the discriminator includes but is not limited to a fully connected neural network, a convolutional neural network, a recurrent neural network, a variational autoencoder, a Transformer model, etc.
[0031] In the process of constructing the objective function of this embodiment, the objective function needs to satisfy the following requirements: when the sample input to the discriminator comes from the distribution of real data, the output value of the discriminator is the largest; when the sample input to the discriminator comes from the distribution of random noise, the output value of the discriminator is the smallest; when the generated sample of the generator is input to the discriminator, the probability that the discriminator determines that the generated sample is true is the largest.
[0032] In this embodiment, the objective function is: , G is the generator, which is used to map random noise z to generated data G(z), D is the discriminator, x is the real data, P data is the distribution of the real data x, P z is the distribution of random noise z, and E is the expected value.
[0033] This embodiment introduces a generator and a discriminator to generate test data. For example, in an e-commerce platform scenario, this embodiment can first collect a large amount of real user behavior data, transaction data, and product data as a training set. The generator learns the distribution and characteristics of real data, and the discriminator distinguishes between generated data and real data. Through adversarial training between the generator and the discriminator, test data that is closer to the real task scenario is generated. For example, when generating user purchase behavior data, this embodiment can not only simulate the purchase time, purchase type and quantity of different users, but also generate a behavior sequence that conforms to their purchase logic based on the user's consumption level and preferences.
[0034] In this embodiment, when generating user purchase behavior data, in order to more realistically simulate the behavior of real users, in addition to considering the purchase time, types and quantities of purchased goods, and consumption levels and preferences of different users, this embodiment can also introduce more auxiliary data representing the user to generate intermediate variables, including but not limited to data such as the user's payday, family members, social status, occupation, geographic location, life stage, season and climate, marketing activity factors, etc., so as to simulate more realistic user behavior data.
[0035] For example, this embodiment can use payday as auxiliary data of users. Since different users have different paydays, this will directly affect their spending power and purchase timing. With the consent of the users, this embodiment can collect payday data of a large number of users, analyze their distribution patterns, and build a payday model, so as to generate payday data in the test data through the payday model.
[0036] For another example, the number and structure of family members have an important impact on user purchasing behavior. Families with children will have a significantly increased demand for maternal and child products, educational products, etc.; families with elderly people may pay more attention to health products, medical supplies, etc. This embodiment can simulate different purchasing scenarios based on the user's family structure information. For example, a family with young children can purchase milk powder, toys, books, etc. for the corresponding age groups at different stages of the child's growth; a family with elderly people can regularly purchase health and medical supplies such as antihypertensive drugs and blood glucose meters.
[0037] For another example, a user's social activities can also affect their purchasing decisions. Socially active users may purchase fashion clothing, cosmetics, gifts, etc. because they attend parties and events, while users who are relatively less social may spend less on these aspects. This embodiment can construct a social status model by analyzing the user's social platform data, social activity frequency, and other information. For socially active users, their purchase behavior of related products is simulated when social activities are approaching; for users who are less social, the purchase simulation of such products is reduced.
[0038] After generating simulated user purchase behavior data, this embodiment can compare it with the user's existing real shopping data, and evaluate the credibility of the simulated data by calculating indicators such as data similarity and deviation rate. For simulated data below the credibility threshold, further adjust the simulation parameters and optimize the generation model. For example, if the frequency of users purchasing a certain type of goods in the simulated data is significantly different from the real data, the reason may be that the user's preferences are not accurately grasped during the simulation process, or certain influencing factors are not fully considered. This embodiment can make targeted adjustments and regenerate data until simulated intermediate data with higher credibility is obtained.
[0039] 210. Access the user operation behavior model and the human-computer interaction model, and dynamically adjust the test environment to match the production environment through the user operation behavior model and the human-computer interaction model.
[0040] In some embodiments, see Figure 4 , Figure 4 The present invention provides a flowchart of dynamically adjusting the test environment. The user operation behavior model and the human-computer interaction model are accessed, and the test environment is dynamically adjusted to match the production environment through the user operation behavior model and the human-computer interaction model, including: 2100. Accessing a user operation behavior model, where the user operation behavior model is used to simulate operation behavior data of a real user; 2110. Access the human-computer interaction model, which is used to predict production environment resource usage data and network status; 2120. Based on the user operation behavior model, the eye tracking method and user behavior simulation tool are used to simulate the operation behavior of real users when using the big data analysis system; 2130. Based on the human-computer interaction model, the hardware configuration and network parameters in the test environment are dynamically adjusted through automated scripts to keep the test environment consistent with the production environment.
[0041] This embodiment can use cloud computing technology and intelligent monitoring systems to detect parameters such as hardware resource usage, network traffic, and load changes in the production environment in real time. In the test environment, the hardware configuration (such as the number of CPU cores and memory size of the virtual machine) and network parameters (such as network latency and bandwidth) are dynamically adjusted through automated scripts to keep the test environment highly consistent with the production environment. At the same time, this embodiment can also introduce eye tracking technology and user behavior simulation tools to simulate the operation behavior of real users when using the big data analysis system, such as clicking, sliding, inputting, etc., to more realistically test the performance of the system in actual use.
[0042] In some embodiments, before the user operation behavior model is formed, the steps include: obtaining the operation behavior data of the real user; extracting the behavior features corresponding to the operation behavior data of the real user by a feature extraction method; and establishing the user operation behavior model according to the behavior features.
[0043] This embodiment can collect the operation behavior data of real users and obtain the corresponding behavior characteristics by actually detecting and recording the operations of a large number of real users when using the big data analysis system, thereby establishing a user operation behavior model. For example, this embodiment can analyze the clicking habits of different users on different pages, count the sliding operation frequency of specific functional modules, etc., and use this as a basis to drive the user behavior simulation tool to perform simulation operations.
[0044] In some embodiments, before accessing the human-computer interaction model, it includes: acquiring production environment data; establishing a human-computer interaction model according to the production environment data, and the human-computer interaction model includes a hardware usage relationship model and a network usage model.
[0045] In this embodiment, the production environment data involved in establishing the human-computer interaction model include but are not limited to: the hardware resource usage of the production environment (such as CPU utilization, memory usage), network traffic (data transmission volume, number of requests, etc.), load changes (number of concurrent users, task processing volume, etc.), etc.; the dynamic adjustment parameters involved in the test environment of this embodiment include but are not limited to: hardware configuration parameters that need to be adjusted in the test environment (number of CPU cores, memory size) and network parameters (network latency, bandwidth), etc.
[0046] In some embodiments, see Figure 5 , Figure 5 The present invention provides a flowchart of establishing a human-computer interaction model based on production environment data. The production environment data includes the processor usage rate, the number of concurrent users, and the fluctuation data of network traffic in different time periods. The human-computer interaction model is established based on the production environment data. The human-computer interaction model includes a hardware usage model and a network usage model, including: 2112. Determine, based on the processor utilization rate and the number of concurrent users, a changing trend of the processor utilization rate in the production environment as the number of concurrent users increases within a preset time period; 2114. According to the change trend, a relationship model between the number of processors and the number of users in the production environment is established, and the relationship model is a hardware usage model; 2116. Based on the fluctuation data of network traffic in different time periods, a time series model of network traffic is established. The time series model is a network usage model.
[0047] This embodiment can establish a model of production environment resource usage and network status, that is, a hardware usage relationship model and a network usage model, by analyzing the collected production environment data. For example, a relationship model between the two is established based on the changing trend of CPU usage in the production environment as the number of concurrent users increases over a period of time; this embodiment can also establish a time series model of network traffic based on the fluctuation of network traffic in different business periods. In the test environment, the automated script calculates the hardware configuration and network parameter values that need to be adjusted based on these models and the real-time monitored production environment data. For example, if the CPU usage in the production environment reaches 80%, and the model calculates that the corresponding reasonable number of CPU cores is 4 cores, the automated script adjusts the number of CPU cores of the virtual machine in the test environment to 4 cores; if the network traffic in the production environment increases and causes the network delay to reach 50ms, the network delay of the test environment should also be adjusted to 50ms according to the model, and this embodiment makes corresponding adjustments through the automated script.
[0048] 220. Based on the test environment and test data, data quality detection and problem tracing are carried out through the data quality test rule base based on the knowledge graph and blockchain technology to obtain optimized test data.
[0049] In some embodiments, see Figure 6 , Figure 6 This is a flow chart of establishing a data quality test rule base based on a knowledge graph provided by an embodiment of the present invention. Based on the test environment and test data, data quality detection and problem tracing are performed through a data quality test rule base based on a knowledge graph and blockchain technology to obtain optimized test data, including: 2200, obtain task knowledge, the relationship between data and data quality problem information; 2210. Construct task knowledge, the relationship between data, and data quality problem information into a knowledge graph; 2220. When the task changes or a new data format appears, the knowledge graph is inferred and updated to automatically generate or adjust data quality test rules to obtain a data quality test rule library based on the knowledge graph.
[0050] This embodiment can establish a data quality test rule base based on the knowledge graph, and construct the task knowledge, the relationship between data, and common data quality problems into a knowledge graph. When the task changes or a new data format appears, the big data test optimization device 110 of this embodiment can automatically generate or adjust the data quality test rules by reasoning and updating the knowledge graph. For example, in an e-commerce platform, when the introduction of new promotional activities causes changes in the data format and task logic, the knowledge graph can quickly identify related data entities and relationships and update the test rules.
[0051] In some embodiments, see Figure 7 , Figure 7 This is a flow chart of problem tracing through blockchain technology provided by an embodiment of the present invention. Based on the test environment and test data, data quality detection and problem tracing are performed through a data quality test rule base based on a knowledge graph and blockchain technology to obtain optimized test data, including: 2230. Use blockchain technology to record the entire life cycle of test data, which includes the collection, transmission, storage and processing of test data; 2240. When there are quality problems in the test data, the problem can be traced through the chain structure of the blockchain to obtain optimized test data.
[0052] In terms of tracing data problems, this embodiment can use blockchain technology to record the entire life cycle of data, including data collection, transmission, storage and processing, etc. The operations and data changes in each link are recorded on the blockchain. When data quality problems are detected, this embodiment can quickly and accurately trace the root cause of the problem through the chain structure of the blockchain.
[0053] The embodiments of this specification can be connected to the optimized generator and discriminator, and generate test data with a similarity with the real data less than a similarity threshold through the optimized generator and discriminator, thereby providing data close to the real task scenario for performance and quality testing; then, the embodiments of this specification can also be connected to the user operation behavior model and the human-computer interaction model, and dynamically adjust the test environment to match the production environment through the user operation behavior model and the human-computer interaction model. The embodiments of this specification dynamically adjust the performance test environment according to the production environment, thereby providing a precise testing environment for other modules; in addition, the embodiments of this specification can also use the test environment and test data to perform data quality detection and problem tracing through a data quality test rule library based on a knowledge graph and blockchain technology.
[0054] The embodiments of this specification can not only generate highly realistic test data and improve test accuracy, but also accurately evaluate system performance and quickly locate data quality issues. The embodiments of this specification have the functions of optimizing test data quality, accurately evaluating system performance, and efficiently ensuring data quality. They are suitable for various scenarios with high requirements for big data processing and analysis, such as e-commerce platform big data analysis systems, financial institution data processing systems, etc., effectively ensuring that big data systems can run stably, reliably and efficiently in complex task scenarios, and meeting the quality requirements of big data systems in different industries.
[0055] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0056] See also Figure 8 , Figure 8 A structural diagram of a dynamic model-based test big data optimization system provided in an embodiment of this specification.
[0057] like Figure 8As shown, the dynamic model-based test big data optimization system may include at least a test data generation module 800, an adaptive performance test environment simulation module 810, and an intelligent dynamic data quality test module 820, wherein: A test data generation module 800 is used to access the optimized generator and discriminator, and generate test data whose similarity with real data is less than a similarity threshold through the optimized generator and discriminator; The adaptive performance test environment simulation module 810 is used to access the user operation behavior model and the human-computer interaction model, and dynamically adjust the test environment to match the production environment through the user operation behavior model and the human-computer interaction model; The intelligent dynamic data quality testing module 820 is used to perform data quality detection and problem tracing based on the test environment and test data through a data quality testing rule base based on a knowledge graph and blockchain technology to obtain optimized test data.
[0058] In some embodiments, the big data testing optimization system also includes a model optimization module, which is used to: obtain real data and random noise based on the test scenario, and establish a generator and a discriminator based on a deep learning network; determine the objective function between the generator and the discriminator, in which, when the sample input to the discriminator comes from the distribution of real data, the output value of the discriminator is the largest; when the sample input to the discriminator comes from the distribution of random noise, the output value of the discriminator is the smallest; when the generated sample of the generator is input to the discriminator, the probability that the discriminator determines that the generated sample is true is the largest; solve the objective function based on real data and random noise to obtain an optimized generator and an optimized discriminator.
[0059] In some embodiments, the objective function is , G is the generator, which is used to map random noise z to generated data G(z), D is the discriminator, x is the real data, P data is the distribution of the real data x, P z is the distribution of random noise z, and E is the expected value.
[0060] In some embodiments, the adaptive performance test environment simulation module 810 includes a model access module, which is used to: access the user operation behavior model, which is used to simulate the operation behavior data of real users; access the human-computer interaction model, which is used to predict the resource usage data and network status of the production environment; based on the user operation behavior model, simulate the operation behavior of real users when using the big data analysis system through eye tracking methods and user behavior simulation tools; based on the human-computer interaction model, dynamically adjust the hardware configuration and network parameters in the test environment through automated scripts to keep the test environment consistent with the production environment.
[0061] In some embodiments, the adaptive performance test environment simulation module 810 also includes a first model building module, which is used to: obtain the operation behavior data of real users; extract the behavior features corresponding to the operation behavior data of real users through a feature extraction method; and establish a user operation behavior model based on the behavior features.
[0062] In some embodiments, the adaptive performance test environment simulation module 810 also includes a second model building module, which is used to: obtain production environment data; and build a human-computer interaction model based on the production environment data, wherein the human-computer interaction model includes a hardware usage relationship model and a network usage model.
[0063] In some embodiments, the production environment data includes processor usage, the number of concurrent users, and network traffic fluctuation data in different time periods. The second model building module includes a model building sub-module, and the model building sub-module is used to: determine the changing trend of processor usage in the production environment as the number of concurrent users increases within a preset time period based on the processor usage and the number of concurrent users; establish a relationship model between the number of processors and the number of users in the production environment based on the changing trend, and the relationship model is a hardware usage model; establish a time series model of network traffic based on the fluctuation data of network traffic in different time periods, and the time series model is a network usage model.
[0064] In some embodiments, the big data testing optimization system includes a rule base establishment module, which is used to: obtain task knowledge, the relationship between data, and data quality problem information; construct task knowledge, the relationship between data, and data quality problem information into a knowledge graph; when the task changes or a new data format appears, the knowledge graph is inferred and updated, and data quality testing rules are automatically generated or adjusted to obtain a data quality testing rule base based on the knowledge graph.
[0065] In some embodiments, the intelligent dynamic data quality testing module 820 includes a problem tracing module, which is used to: use blockchain technology to record the entire life cycle of test data, the entire life cycle includes the collection, transmission, storage and processing of test data; when quality problems are detected in the test data, the problem is traced through the chain structure of the blockchain to obtain optimized test data.
[0066] Based on the content of the dynamic model-based test big data optimization system in multiple embodiments of this specification, it can be known that the big data test optimization system of the embodiment of this specification can be connected to the optimized generator and discriminator, and generate test data with a similarity with the real data less than a similarity threshold through the optimized generator and discriminator, thereby providing data close to the real task scenario for performance and quality testing; then, the embodiment of this specification can also be connected to the user operation behavior model and the human-computer interaction model, and dynamically adjust the test environment to match the production environment through the user operation behavior model and the human-computer interaction model. The embodiment of this specification dynamically adjusts the performance test environment according to the production environment, thereby providing a precise testing environment for other modules; in addition, the embodiment of this specification can also use the test environment and test data to perform data quality detection and problem tracing through a data quality test rule library based on a knowledge graph and blockchain technology.
[0067] The embodiments of this specification can not only generate highly realistic test data and improve test accuracy, but also accurately evaluate system performance and quickly locate data quality issues. The embodiments of this specification have the functions of optimizing test data quality, accurately evaluating system performance, and efficiently ensuring data quality. They are suitable for various scenarios with high requirements for big data processing and analysis, such as e-commerce platform big data analysis systems, financial institution data processing systems, etc., effectively ensuring that big data systems can run stably, reliably and efficiently in complex task scenarios, and meeting the quality requirements of big data systems in different industries.
[0068] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the test big data optimization system based on a dynamic model, since it is basically similar to an embodiment of a test big data optimization method based on a dynamic model, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0069] See also Fig. 9 A schematic diagram of the structure of an electronic device of a dynamic model-based test big data optimization system provided in an embodiment of this specification is shown.
[0070] like Fig. 9As shown, the electronic device 900 may include: at least one processor 910 , at least one network interface 940 , a user interface 930 , a memory 950 , and at least one communication bus 920 .
[0071] The communication bus 920 may be used to realize the connection and communication among the above-mentioned components.
[0072] The user interface 930 may include buttons, and the optional user interface may also include a standard wired interface or a wireless interface.
[0073] The network interface 940 may include, but is not limited to, a Bluetooth module, an NFC module, a ZigBee module, and a UWB module.
[0074] Among them, the processor 910 may include one or more processing cores. The processor 910 uses various interfaces and lines to connect various parts within the entire electronic device 900, and executes various functions and processes data of the electronic device 900 by running or executing instructions, programs, code sets or instruction sets stored in the memory 950, and calling data stored in the memory 950. Optionally, the processor 910 can be implemented in at least one hardware form of DSP, FPGA, and PLA. The processor 910 can integrate one or a combination of CPU and GPU. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content that needs to be displayed on the display screen.
[0075] Among them, the memory 950 may include RAM or ROM. Optionally, the memory 950 includes a non-transitory computer-readable medium. The memory 950 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 950 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 950 may also be at least one storage device located away from the aforementioned processor 910. The memory 950 as a computer storage medium may include an operating system, a communication module, a user interface module, and a test big data optimization application based on a dynamic model. The processor 910 can be used to call the test big data optimization application based on a dynamic model stored in the memory 950 about the big data test optimization system, and execute the steps in the test big data optimization method based on a dynamic model mentioned in the aforementioned embodiment.
[0076] The embodiment of the present specification also provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a computer or a processor, the computer or the processor executes the above Figure 2 to Figure 7 One or more steps in the illustrated embodiment. If the components of the electronic device described above are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0077] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of this specification is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by the computer or a data storage device such as a server or data center that contains one or more available media integrations. Available media may be magnetic media (eg, floppy disks, hard disks, tapes), optical media (eg, digital versatile discs (DVD)), or semiconductor media (eg, solid state drives (SSD)).
[0078] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes. In the absence of conflict, the technical features in this embodiment and the implementation scheme can be combined arbitrarily.
[0079] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Without departing from the design spirit of this specification, various modifications and improvements made to the technical solutions of this specification by ordinary technicians in this field should fall within the scope of protection determined by the claims of this specification.
Claims
1. A test big data optimization method based on a dynamic model, characterized in that: include: Connecting the optimized generator and discriminator, and generating test data whose similarity with the real data is less than a similarity threshold through the optimized generator and discriminator; Accessing the user operation behavior model and the human-computer interaction model, and dynamically adjusting the test environment to match the production environment through the user operation behavior model and the human-computer interaction model; Based on the test environment and the test data, data quality detection and problem tracing are performed through a data quality test rule base based on a knowledge graph and blockchain technology to obtain optimized test data.
2. The test big data optimization method based on dynamic model according to claim 1, characterized in that: Before the access to the optimized generator and discriminator, it includes: Obtain real data and random noise based on the test scenario, and build a generator and discriminator based on the deep learning network; Determine an objective function between the generator and the discriminator, in which, when the sample input to the discriminator comes from the distribution of real data, the output value of the discriminator is maximum; when the sample input to the discriminator comes from the distribution of random noise, the output value of the discriminator is minimum; when the generated sample of the generator is input to the discriminator, the probability that the discriminator determines that the generated sample is true is maximum; The objective function is solved based on the real data and the random noise to obtain an optimized generator and an optimized discriminator.
3. The test big data optimization method based on dynamic model according to claim 2 is characterized in that: The objective function is , G is a generator, which is used to map random noise z to generated data G(z), D is a discriminator, x is real data, Pdata is the distribution of real data x, Pz is the distribution of random noise z, and E is the expected value.
4. The test big data optimization method based on dynamic model according to claim 1, characterized in that: Accessing the user operation behavior model and the human-computer interaction model, and dynamically adjusting the test environment to match the production environment through the user operation behavior model and the human-computer interaction model, including: Accessing a user operation behavior model, wherein the user operation behavior model is used to simulate the operation behavior data of a real user; Accessing a human-computer interaction model, wherein the human-computer interaction model is used to predict production environment resource usage data and network status; Based on the user operation behavior model, the operation behavior of real users when using the big data analysis system is simulated by using an eye tracking method and a user behavior simulation tool; Based on the human-computer interaction model, the hardware configuration and network parameters in the test environment are dynamically adjusted through automated scripts to keep the test environment consistent with the production environment.
5. The test big data optimization method based on dynamic model according to claim 4 is characterized in that: The user operation behavior model includes: Obtain real user operation behavior data; Extracting the behavior features corresponding to the operation behavior data of the real user by a feature extraction method; A user operation behavior model is established according to the behavior characteristics.
6. The test big data optimization method based on dynamic model according to claim 4 is characterized in that: Before accessing the human-computer interaction model, the following steps are included: Obtain production environment data; A human-computer interaction model is established according to the production environment data, wherein the human-computer interaction model includes a hardware usage relationship model and a network usage model.
7. The test big data optimization method based on dynamic model according to claim 6 is characterized in that: The production environment data includes the processor usage rate, the number of concurrent users and the fluctuation data of network traffic in different time periods. The human-computer interaction model is established according to the production environment data. The human-computer interaction model includes a hardware usage model and a network usage model, including: Determine, according to the processor usage rate and the number of concurrent users, a change trend of the processor usage rate in the production environment as the number of concurrent users increases within a preset time period; According to the change trend, a relationship model between the number of processors and the number of users in the production environment is established, wherein the relationship model is a hardware usage model; A time series model of network traffic is established according to the fluctuation data of the network traffic in different time periods, and the time series model is the network usage model.
8. The test big data optimization method based on dynamic model according to claim 1, characterized in that: Based on the test environment and the test data, data quality detection and problem tracing are performed through a data quality test rule base based on a knowledge graph and blockchain technology to obtain optimized test data, including: Acquire task knowledge, relationships between data, and information about data quality issues; Constructing the task knowledge, the relationship between the data and the data quality problem information into a knowledge graph; When the task changes or a new data format appears, the knowledge graph is inferred and updated, and data quality testing rules are automatically generated or adjusted to obtain a data quality testing rule library based on the knowledge graph.
9. The test big data optimization method based on dynamic model according to claim 8, characterized in that: Based on the test environment and the test data, data quality detection and problem tracing are performed through a data quality test rule base based on a knowledge graph and blockchain technology to obtain optimized test data, including: Use blockchain technology to record the entire life cycle of test data, including the collection, transmission, storage and processing of test data; When quality problems are detected in the test data, the problem is traced through the chain structure of the blockchain to obtain optimized test data.
10. A test big data optimization system based on a dynamic model, comprising: A test data generation module, used to access the optimized generator and discriminator, and generate test data whose similarity with real data is less than a similarity threshold through the optimized generator and discriminator; An adaptive performance test environment simulation module is used to access the user operation behavior model and the human-computer interaction model, and dynamically adjust the test environment to match the production environment through the user operation behavior model and the human-computer interaction model; The intelligent dynamic data quality testing module is used to perform data quality detection and problem tracing based on the test environment and the test data through a data quality testing rule base based on a knowledge graph and blockchain technology to obtain optimized test data.
Citation Information
Patent Citations
Method and system for monitoring test process of computer network application program
CN118897809A
Mirror image type industry linkage engine system and method based on knowledge graph
CN119537864A
Automatic test evaluation construction method of multi-modal large model and related device
CN119577420A
Method and system for generating behavioral studies of an individual
US20130022950A1