Performance optimization method and system of big data medium table
By using large language models to generate test data sets and simulate test environments, the problem that traditional testing methods cannot fully cover the complex data and hardware environment of the big data middle platform is solved, and efficient performance optimization and stability improvement are achieved.
Patent Information
- Application Number
- CN202510412498.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional static testing methods cannot efficiently process the growing complex data of the big data middle platform and cannot test the hardware level, resulting in incomplete test coverage and difficulty in assessing stability and reliability.
Use pre-trained large language models to analyze business requirements information, generate multiple sets of test data sets, build simulated test environments, perform performance tests, and dynamically adjust the operating environment according to test results.
It improves the test coverage and accuracy of the big data middle platform, ensures stability and reliability in complex and changing environments, shortens the problem discovery and resolution cycle, and improves testing efficiency.
Smart Images

Figure CN120336176A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of software testing. Specifically, it relates to a method and system for optimizing the performance of a big data platform. Background Art
[0002] In the wave of the digital age, the big data platform has become the core asset for enterprises to enhance their competitiveness. However, the traditional method of manually writing test users currently has the following problems when dealing with the development and application of the big data platform:
[0003] Problem 1: When the big data platform processes massive and complex data from different business systems, due to issues such as slight differences in data formats, data missing, or outliers, the test cases designed based on known business scenarios and preset data patterns cannot fully cover these problems.
[0004] Problem 2: The big data platform usually needs to closely cooperate with various hardware devices and complex network environments, such as servers, storage devices, network switches, etc. However, traditional test cases mainly focus on functional testing at the software level, so they cannot simulate the overall impact of real physical environments such as hardware failures and network delays on the system, making it difficult to comprehensively evaluate the stability and reliability of the big data platform in a complex environment.
[0005] Problem 3: The business logic and data processing flow of the big data platform are constantly updated and optimized like the rapidly changing fashion trend. Therefore, traditional static test cases have high maintenance costs and slow response speeds, and it is difficult to quickly adapt to the changes in the system.
[0006] Therefore, there is an urgent need for a more flexible, efficient, and comprehensive testing method to meet the growing testing requirements of the big data platform. Summary of the Invention
[0007] The embodiments of this application provide a method and system for optimizing the performance of a big data platform to at least solve the technical problems that traditional static testing cannot efficiently process the increasingly complex data of the big data platform and cannot test the hardware level.
[0008] According to one aspect of the embodiments of the present application, a method for optimizing the performance of a big data middle platform is provided, including: obtaining business requirement information of the big data middle platform, where the business requirement information is used to reflect the business functions and business objectives of the big data middle platform within a target time period; analyzing the business requirement information by using a pre-trained target large language model to obtain a first data attribute set for processing business data by the big data middle platform within the target time period; generating multiple groups of test data sets based on the first data attribute set, and respectively constructing multiple simulated test environments according to the multiple groups of test data sets; performing performance tests on the big data middle platform in the multiple simulated test environments respectively to obtain multiple performance test results, and adjusting the operating environment of the big data middle platform according to the multiple performance test results.
[0009] Optionally, the training process of the target large language model includes: obtaining multiple groups of training sample data, where each group of training sample data includes: historical business requirement information of the big data middle platform and a second data attribute set for processing historical business data by the big data middle platform within a specific historical time period; obtaining a general large language model; and iteratively training the general large language model by using the multiple groups of training sample data to obtain the target large language model.
[0010] Optionally, the historical business requirement information includes at least one of the following: data format, data type, data processing logic, data storage method, data quality requirements, and data organization form for the big data middle platform to process business data, where the data format includes at least one of the following: CSV format, JSON format, XML format; the data type includes at least one of the following: integer, text, floating point number; the data processing logic includes at least one of the following: data cleaning, data conversion, data clustering, data analysis; the data storage method includes at least one of the following: relational database, non-relational database, file system; the data quality requirements include at least one of the following: fields cannot be empty, data value range constraints; the data organization form includes at least one of the following: table, document, image, audio.
[0011] Optionally, the second data attribute set includes at least one of the following: data type, data scale, data structure, data format, data boundary value, and data quality.
[0012] Optionally, there is test data in multiple dimensions within each group of test data sets. Multiple groups of test data sets are generated based on the first data attribute set, including: when the data type is included in the first data attribute set, using the first generation strategy to generate test data corresponding to the data type dimension, where the first generation strategy includes at least one of the following: st.integers strategy, st.text strategy; when the data scale is included in the first data attribute set, using the second generation strategy to generate test data corresponding to the data scale dimension, where the second generation strategy includes at least one of the following: st.lists strategy, st.data strategy; when the data structure is included in the first data attribute set, using the third generation strategy to generate test data corresponding to the data structure, where the third generation strategy includes at least one of the following: st.dictionaries strategy, st.tuples strategy; when the data boundary values are included in the first data attribute set, using the fourth generation strategy to generate test data corresponding to the data structure dimension, where the fourth generation strategy includes at least one of the following: st.floats(min_value,max_value) strategy, st.integers(min_value,max_value) strategy; when the data quality is included in the first data attribute set, using the fifth generation strategy to generate test data corresponding to the data quality dimension, where the fifth generation strategy includes at least one of the following: st.none strategy, st.just(none), st.integers(min_size,max_size) strategy, st.floats(min_size,max_size) strategy.
[0013] Optionally, performance tests are respectively conducted on the big data middle platform in multiple simulated test environments to obtain multiple performance test results, and the operating environment of the big data middle platform is adjusted according to the multiple performance test results, including: for each simulated test environment, conducting a performance test on the big data middle platform in the simulated test environment to obtain the corresponding performance test result, where the performance test result includes test results in multiple metric dimensions, and the metric dimensions include at least one of the following: response time, throughput, resource utilization rate, data loss rate, data consistency; judging whether the test results corresponding to each metric dimension in the performance test result meet the corresponding metric expected requirements; when the test results corresponding to each metric dimension in the performance test result all meet the corresponding metric expected requirements, adjusting the operating environment of the big data middle platform according to the test data set corresponding to the simulated test environment; when the test result corresponding to any metric dimension in the performance test result does not meet the corresponding metric expected requirements, adjusting the test data set corresponding to the simulated test environment, and constructing a new simulated test environment based on the adjusted test data set, and obtaining the test performance result of the big data middle platform in the new simulated test environment again.
[0014] Optionally, adjust the test data set corresponding to the simulation test environment, including: determining the metric dimension corresponding to the test result not meeting the expected metric requirements, and determining at least one target test data in the test data set related to the metric dimension; adjusting each target test data according to a preset parameter adjustment strategy set, where the parameter adjustment strategy set includes: adjustment strategies corresponding to test data for different metric dimensions.
[0015] According to another aspect of the embodiments of the present application, there is also provided a performance optimization system for a big data middle platform, including: an acquisition module, configured to acquire the business requirement information of the big data middle platform, where the business requirement information is used to reflect the business functions and business objectives of the big data middle platform during a target time period; a prediction module, configured to analyze the business requirement information by using a pre-trained target large language model to obtain a first data attribute set of the big data middle platform for processing business data during the target time period; a construction module, configured to generate multiple groups of test data sets based on the first data attribute set, and respectively construct multiple simulation test environments based on the multiple groups of test data sets; a tuning module, configured to perform performance tests on the big data middle platform in multiple simulation test environments respectively, obtain multiple performance test results, and adjust the operating environment of the big data middle platform according to the multiple performance test results.
[0016] According to another aspect of the embodiments of the present application, there is also provided a computer program product, which includes: a computer program, where when the computer program is executed by a processor, it implements the above-mentioned performance optimization method for the big data middle platform.
[0017] According to another aspect of the embodiments of the present application, there is also provided an electronic device, which includes: a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the above-mentioned performance optimization method for the big data middle platform through the computer program.
[0018] In the embodiments of the present application, the business requirement information of the big data middle platform is analyzed by using a target large language model to obtain a first data attribute set; multiple groups of test data sets are generated based on the first data attribute set, and multiple simulation test environments are constructed based on this; performance tests are performed on the big data middle platform in multiple simulation test environments respectively, multiple performance test results are obtained, and the operating environment of the big data middle platform is adjusted according to the multiple performance test results. The technical effects of systematically and intelligently improving the test coverage, accuracy, and efficiency of the big data middle platform are achieved, ensuring the stability and reliability of the big data middle platform in a complex and changing environment. Furthermore, the technical problem that traditional static tests cannot efficiently process the increasing complex data of the big data middle platform and cannot test the hardware level is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0020] Figure 1 is a schematic flowchart of an optional method for optimizing the performance of a big data middle platform according to an embodiment of the present application;
[0021] Figure 2 is a schematic structural diagram of an optional system for optimizing the performance of a big data middle platform according to an embodiment of the present application;
[0022] Figure 3 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0023] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0025] In order to better understand the embodiments of the present application, some nouns or terms that appear in the description process of the embodiments of the present application are translated and explained as follows:
[0026] Big Data Middle Platform: It refers to the collection, calculation, storage, and processing of massive amounts of data through data technologies, while unifying standards and calibers. After unifying these data, the Big Data Middle Platform will form standard data, which is then stored to form a big data asset layer, and further provide efficient services for customers. These services are strongly related to the business of the enterprise, are unique to the enterprise and can be reused. They can not only reduce duplicate construction and the cost of chimney-style collaboration, but also are the source of the enterprise's differential competitive advantage.
[0027] Hypothesis: It is a property-based testing library in Python, which is used to support test-driven development and property-based testing. Hypothesis helps developers write more comprehensive and accurate test cases by automatically generating random test data that conforms to specific specifications, thereby improving the quality and coverage of the code. The core idea of Hypothesis is to use assumptions to infer the behavior of the code and generate test data based on these assumptions. By performing intelligent search on the possible input space, it can discover error conditions and boundary conditions that cannot be covered by many manually written test cases.
[0028] CSV (Comma-Separated Value) format: It is a simple tabular data format that separates fields by commas, suitable for storing large amounts of tabular data and is easy to process and import / export.
[0029] JSON (JavaScript Object Notation) format: It is a lightweight data interchange format, easy to read and write, and is often used to transmit data in web development.
[0030] XML (eXtensible Markup Language) format: It is a markup language used to describe and transmit data, and is often used in web services and data storage.
[0031] Embodiment 1
[0032] According to the embodiments of the present application, a method for optimizing the performance of a big data middle platform is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0033] Figure 1 It is a schematic flowchart of a method for optimizing the performance of a big data middle platform provided according to the embodiments of the present application. As Figure 1 shown, the method includes the following steps:
[0034] Step S102, obtain the business requirement information of the big data middle platform.
[0035] In the technical solution provided in the above step S102, the system can extract and analyze the business functions, performance indicators, data processing requirements, and related business scenarios and goals that the big data middle platform needs to achieve during the target time period from project documents, user feedback, and market trends, so as to obtain the business requirement information.
[0036] Step S104, use the pre-trained target large language model to analyze the business requirement information, and obtain the first data attribute set for the big data middle platform to process business data during the target time period.
[0037] In the technical solution provided in the above step S104, the above target large language model is a deep learning model. After being trained with a large amount of text data, it can analyze and refine the business scenarios and business rules described in the business requirement information, identify the relevance between data, the logical process of data processing, and possible abnormal situations and boundary conditions, and obtain the key attributes of the business data that the big data middle platform needs to process during the target time period. Therefore, in the embodiment of the present application, the system can call the target large language model to analyze the business requirement information and identify the data attribute set of the business data that the big data middle platform needs to process during the target time period. This step can convert unstructured business requirements into structured data feature descriptions, thus facilitating subsequent automated test operations.
[0038] Step S106, generate multiple groups of test data sets based on the first data attribute set, and respectively construct multiple simulated test environments according to the multiple groups of test data sets.
[0039] In the technical solution provided in the above step S106, the system can use the Hypothesis test framework or a similar parameterized test tool to use each first data feature in the first data attribute set obtained in the above step S104 as a parameter seed to generate test data sets covering various business scenarios and possible abnormal situations. Then, respectively construct multiple simulated test environments according to the generated multiple groups of test data sets. This means that each simulated test environment will reflect specific data characteristics and workloads, including but not limited to the software configuration, hardware specifications, network conditions, etc. of the big data middle platform.
[0040] Step S108, perform performance tests on the big data middle platform in multiple simulated test environments respectively, obtain multiple performance test results, and adjust the operating environment of the big data middle platform according to the multiple performance test results.
[0041] In the technical solution provided in step S108 above, by running the automated performance test script in multiple simulated test environments, a systematic performance evaluation of the big data middleware platform is conducted, and the test results in each simulated test environment are collected. Then, based on the obtained performance test results, the performance bottlenecks of the big data middleware platform's operating environment are identified and adjusted accordingly to improve the overall performance of the big data middleware platform. During the above process, the test parameter set and the operating configuration of the big data middleware platform are continuously adjusted until the predetermined performance target is reached or the performance cannot be further improved through adjustment. This process ensures the comprehensiveness of the test and the continuity of the performance optimization of the big data middleware platform.
[0042] Based on the solution defined in steps S102 to S108 above, it can be known that in the embodiment of this application, the business requirement information of the big data middleware platform is analyzed using the target large language model to obtain the first data attribute set; multiple groups of test data sets are generated based on the first data attribute set, and multiple simulated test environments are constructed based on this; the big data middleware platform is performance-tested in multiple simulated test environments respectively to obtain multiple performance test results, and the operating environment of the big data middleware platform is adjusted according to the multiple performance test results. Thus, the technical effects of systematically and intelligently improving the test coverage rate, accuracy, and efficiency of the big data middleware platform are achieved, ensuring the stability and reliability of the big data middleware platform in a complex and changeable environment.
[0043] The following describes each step of the performance optimization method for the big data middleware platform in combination with a specific implementation process.
[0044] As an alternative implementation, in the technical solution provided in step S102 above, the training process of the target large language model may include:
[0045] The first step: Obtain multiple groups of training sample data.
[0046] Specifically, the system can obtain the historical business requirement information of the big data middleware platform and the second data attribute set of the big data middleware platform for processing historical business data within a specific historical time period from the historical operation records of the big data middleware platform, and through preprocessing these data information, such as extracting key information, data cleaning, format conversion, etc., multiple groups of training sample data are obtained to ensure that these data can be effectively utilized by the model. Among them, the above-mentioned second data attribute set can be obtained through statistical analysis of the actual processed data.
[0047] Specifically, the above historical business requirement information includes but is not limited to: the data format, data type, data processing logic, data storage method, data quality requirements, data organization form, etc. of the business data processed by the big data middle platform. Among them: the above data format includes but is not limited to: CSV format, JSON format, XML format, etc.; the above data type includes but is not limited to: integer, text, floating point number, etc.; the above data processing logic includes but is not limited to: data cleaning, data conversion, data clustering, data analysis, etc.; the above data storage method includes but is not limited to: relational database, non-relational database (such as NoSQL), file system, etc.; the above data quality requirements include but are not limited to: fields cannot be empty, data value range constraints, etc.; the above data organization form includes but is not limited to: tables, documents, images, audio, etc.
[0048] In addition, the above second data attribute set includes but is not limited to: data type, data scale, data structure, data format, data boundary value, data quality, etc. Among them: the above data type includes but is not limited to: numerical value, text, date, boolean value, etc.; the above data scale is used to reflect the data volume of business data, and its range can be a small order of magnitude or a PB (petatype, which is equal to 1000TB) level; the above data structure reflects the organization and association method between business data, such as arrays, tables, tree structures, etc.; the above data boundary value emphasizes the possible maximum value, minimum value and special values of business data, etc.; the above data quality involves issues such as the integrity, accuracy, consistency of business data, and whether there are outliers and missing values.
[0049] The second step: Obtain a general large language model. Among them, the above general large model can be a deep learning model pre-trained on a large amount of text data, with extensive language understanding and generation capabilities, and can process and understand complex and rich text information.
[0050] The third step: Iteratively train the general large language model with multiple groups of training sample data to obtain a target large language model. Among them, during the training process, the model can learn how to extract data features from the historical business requirement information. By continuously optimizing the model parameters, it enables the model to have the ability to more accurately predict and generate data attribute sets related to business requirements.
[0051] Therefore, the system can analyze the business requirement information of the big data middle platform by calling the target large language model trained by the above method, and obtain the first data attribute set of the business data processed by the big data middle platform during the target time period.
[0052] Furthermore, since there are test data in multiple dimensions in each group of test data sets. Therefore, when the system generates multiple groups of test data sets according to the following steps, it includes:
[0053] When the data type is included in the first data attribute set, the system can use the first generation strategy to generate test data corresponding to the corresponding data type. Among them, the first generation strategy includes but is not limited to: st.integers strategy, st.text strategy, etc. For example, if the data type is an integer, a series of random integers can be generated using st.integers; if the data type is text, a random string can be generated using st.text. This can ensure that the test data covers various data types, facilitating subsequent verification of the data processing capabilities of the big data platform for different types of data.
[0054] When the data scale is included in the first data attribute set, the system can use the second generation strategy to generate test data of a specified length or quantity. Among them, the second generation strategy includes but is not limited to: st.lists strategy, st.data strategy, etc. For example, to test the performance of the big data platform in processing a large amount of data, a test data list containing tens of thousands or even millions of records can be generated using st.lists, or a more complex data structure can be generated using the st.data strategy to simulate a large-scale data input scenario and evaluate the response time and resource consumption of the big data platform.
[0055] When the data structure is included in the first data attribute set, the system can use the third generation strategy to generate test data corresponding to the data structure. Among them, the third generation strategy includes but is not limited to: st.dictionaries strategy, st.tuples strategy, etc. For example, the st.dictionaries strategy can be used to create data with a key-value pair structure, or the st.tuples strategy can be used to generate data in the form of tuples to simulate the behavior of the big data platform when processing structured data, such as testing scenarios like database queries and data indexing.
[0056] When the data boundary values are included in the first data attribute set, the system can use the fourth generation strategy to generate test data corresponding to the data structure dimension. Among them, the fourth generation strategy includes but is not limited to: st.floats(min_value,max_value) strategy, st.integers(min_value,max_value) strategy, etc. The maximum and minimum values of different types of data can be generated through these strategies.
[0057] When data quality is included in the first data attribute set, the system can use the fifth generation strategy to generate test data corresponding to the data quality dimension. The fifth generation strategy includes, but is not limited to: st.none strategy, st.just(none), st.integers(min_size, max_size) strategy, st.floats(min_size, max_size) strategy, etc. It can also be obtained by combining the strategies in the first generation strategy, the second generation strategy, the third generation strategy, and the fourth generation strategy. For example, the system can generate outliers in the following way: first use the st.just(value) strategy to generate a fixed value, and then use the st.floats strategy or the st.integers strategy, and set the allow_infinity and allow_nan parameters to True to allow the generation of infinite, infinitesimal, and NaN values. For example:
[0058]
[0059] In addition, missing values usually appear as empty values or None in data fields. To generate such data, st.just(none) can be used. For structured data (such as dictionaries or lists), the st.dictionaries() strategy or st.list() can be used and st.none() can be embedded to generate data containing missing values. For example:
[0060]
[0061] Furthermore, duplicate values can be generated by using the st.lists strategy on the existing strategy, setting the min_size and max_size parameters to control the size of the data set, and using the st.sampled_from() method of the st.integer strategy or the st.floats strategy to sample data from a predefined list. For example:
[0062] duplicate_data_strategy = st.lists(
[0063] st.integers(min_value = 1, max_value = 100),
[0064] min_size = 10, max_size = 100
[0065] ).map(lamba 1 st :1strsandom.choices(1 st, k = 10)). Add 10 randomly repeated elements.
[0066] As an alternative implementation, in the technical solution provided in the above step S108, the system can optimize the performance of the big data middle platform according to the following method, including:
[0067] The first step: For each simulation test environment, perform a performance test on the big data middle platform under the simulation test environment to obtain the corresponding performance test results. Among them, the performance test results include test results corresponding to multiple indicator dimensions such as response time, throughput, resource utilization rate, data loss rate, data consistency, etc., to comprehensively evaluate the behavior and performance of the big data middle platform under different conditions.
[0068] The second step: Determine whether the test results corresponding to each indicator dimension in the performance test results meet the corresponding indicator expected requirements. Among them, the indicator expected requirements corresponding to each indicator dimension are set by developers based on historical experience.
[0069] In the case where the test results corresponding to each indicator dimension in the performance test results all meet the corresponding indicator expected requirements, this means that the big data middle platform performs well in the current simulation test environment. At this time, the system can adjust the operating environment of the big data middle platform according to the test data set corresponding to the simulation test environment to optimize the big data middle platform and ensure that it can also maintain good performance in the actual application environment.
[0070] In the case where the test result corresponding to any indicator dimension in the performance test results does not meet the corresponding indicator expected requirements, the system can trigger an automatic feedback mechanism at this time. The system can analyze the test results corresponding to each indicator dimension, identify the root cause of the problem, adjust the test data set corresponding to the simulation test environment, and build a new simulation test environment based on the adjusted test data set, and obtain the test performance results of the big data middle platform in the new simulation test environment again.
[0071] Specifically, the system can perform automatic feedback according to the following process to accurately identify and solve the performance bottlenecks and potential problems of the big data middle platform, including: First, determine the metric dimension corresponding to the test result not meeting the expected requirements (such as too long response time, low throughput, high resource utilization, increased data loss rate, etc.), and determine at least one target test data in the test data set related to the metric dimension. For example, for the problem of too long response time, the system can focus on the processing of large-scale data sets or the parsing of specific format data; for the problem of data consistency, the system can review the data cleaning and transformation processes, etc. Then, adjust each target test data according to the preset parameter adjustment strategy set. Among them, the parameter adjustment strategy set covers a series of adjustment strategies corresponding to the test data corresponding to different metric dimensions, such as adjusting the data scale, modifying the data format, introducing more outliers, or changing the hardware device simulation parameters, etc. These adjustment strategies can be formulated based on the historical experience of developers, software development, hardware device usage parameters, and other knowledge.
[0072] In the above embodiment, the above optimization process is a closed-loop iterative process. The system continuously monitors and analyzes the test results. Once it finds that the performance does not meet the expected performance, it will immediately take actions to adjust the test strategy to optimize the simulation test environment until the test results corresponding to all metric dimensions reach a satisfactory level. This iterative test and optimization process ensures the stability and efficiency of the big data middle platform in the face of various data and hardware environments.
[0073] Therefore, it is not difficult to see from the above content that the performance optimization method of the big data middle platform provided by the embodiments of the present application has the following technical advantages compared with the traditional method of manually constructing test cases to perform performance testing on the big data middle platform:
[0074] (1) The embodiments of the present application use a neural network model to analyze the business requirement information, identify the relevant data attribute set when the big data middle platform processes the corresponding business data, and can generate high-quality and diverse test data in combination with the Hypothesis test framework or similar parameterized test tools, so as to cover various hardware environments and software environments to construct a test environment closer to the real scenario, thereby significantly improving the comprehensiveness and accuracy of the test.
[0075] (2) The embodiments of the present application can dynamically adjust the test strategy and direction according to the performance test results in different simulation test environments, dynamically generate new test data sets, and construct new simulation test environments. Thus, it can deeply cover the potential problem areas of the big data middle platform to quickly locate and fix problems, greatly shortening the problem discovery and resolution cycle and improving the performance test efficiency of the big data middle platform.
[0076] (3) The embodiments of this application adopt a closed-loop iterative testing mechanism. By continuously and cyclically executing the processes of test data generation, test execution, result analysis, and test data adjustment, it can accurately cover and test the performance of the big data platform under various conditions. This mechanism ensures the depth and breadth of testing, and at the same time can continuously optimize the test strategy according to test feedback, significantly improving the efficiency and effectiveness of testing.
[0077] (4) The embodiments of this application can simulate multiple hardware environments and network conditions on a single system through virtualization and software simulation technologies, reducing the demand for actual hardware resources. At the same time, through automated testing, a large amount of manual intervention is avoided, thus achieving efficient utilization of resources.
[0078] Embodiment 2
[0079] According to the embodiments of this application, there is also provided a performance optimization system for a big data platform for implementing the performance optimization method of the big data platform in Embodiment 1. As Figure 2 shown, the performance optimization system for the big data platform at least includes: an acquisition module 22, a prediction module 24, a construction module 26, and a tuning module 283, where:
[0080] The acquisition module 22 is used to acquire the business requirement information of the big data platform, where the business requirement information is used to reflect the business functions and business objectives of the big data platform within the target time period;
[0081] The prediction module 24 is used to analyze the business requirement information by using a pre-trained target large language model to obtain a first data attribute set for the big data platform to process business data within the target time period;
[0082] The construction module 26 is used to generate multiple groups of test data sets based on the first data attribute set and respectively construct multiple simulated test environments according to the multiple groups of test data sets;
[0083] The tuning module 28 is used to perform performance tests on the big data platform in multiple simulated test environments respectively, obtain multiple performance test results, and adjust the operating environment of the big data platform according to the multiple performance test results.
[0084] Optionally, the performance optimization system for the big data platform further includes a model training module, which is used to train the target large language model according to the following steps, including:
[0085] The first step: Obtain multiple groups of training sample data.
[0086] Specifically, the model training module can obtain the historical business requirement information of the big data platform and the second data attribute set of the big data platform for processing historical business data within a specific historical time period from the historical operation records of the big data platform. By preprocessing these data information, such as extracting key information, data cleaning, format conversion, etc., multiple groups of training sample data are obtained to ensure that these data can be effectively utilized by the model. Among them, the above-mentioned second data attribute set can be obtained through statistical analysis of the actual processed data.
[0087] Specifically, the above-mentioned historical business requirement information includes but is not limited to: the data format, data type, data processing logic, data storage method, data quality requirements, data organization form, etc. of the big data platform for processing business data. Among them: the above-mentioned data format includes but is not limited to: CSV format, JSON format, XML format, etc.; the above-mentioned data type includes but is not limited to: integer, text, floating point number, etc.; the above-mentioned data processing logic includes but is not limited to: data cleaning, data conversion, data clustering, data analysis, etc.; the above-mentioned data storage method includes but is not limited to: relational database, non-relational database (such as NoSQL), file system, etc.; the above-mentioned data quality requirements include but are not limited to: fields cannot be empty, data value range constraints, etc.; the above-mentioned data organization form includes but is not limited to: tables, documents, images, audio, etc.
[0088] In addition, the above-mentioned second data attribute set includes but is not limited to: data type, data scale, data structure, data format, data boundary value, data quality, etc. Among them: the above-mentioned data type includes but is not limited to: numerical value, text, date, boolean value, etc.; the above-mentioned data scale is used to reflect the data volume size of business data, and its range can be small-scale or PB (petatype, which is equal to 1000TB) level; the above-mentioned data structure reflects the organization and association method between business data, such as array, table, tree structure, etc.; the above-mentioned data boundary value emphasizes the possible maximum value, minimum value and special value, etc. of business data; the above-mentioned data quality involves issues such as the integrity, accuracy, consistency of business data, and whether there are outliers and missing values.
[0089] Step 2: Obtain a general large language model. Among them, the above-mentioned general large model can be a deep learning model pre-trained on a large amount of text data, with extensive language understanding and generation capabilities, and can process and understand complex and rich text information.
[0090] Step 3: Iteratively train the general large language model with multiple groups of training sample data to obtain the target large language model. Among them, during the training process, the model can learn how to extract data features from the historical business requirement information. By continuously optimizing the model parameters, it has the ability to more accurately predict and generate data attribute sets related to business requirements.
[0091] Therefore, the prediction module 24 can analyze the business requirement information of the big data middle platform by invoking the target large language model trained by the above method, and obtain the first data attribute set for processing business data by the big data middle platform within the target time period.
[0092] Furthermore, due to the test data of multiple dimensions in each group of test data sets. Therefore, the construction module 26 can generate multiple groups of test data sets according to the following steps, including:
[0093] When the data type is included in the first data attribute set, the construction module 26 can use the first generation strategy to generate test data corresponding to the corresponding data type, where the first generation strategy includes but is not limited to: st.integers strategy, st.text strategy, etc.
[0094] When the data scale is included in the first data attribute set, the construction module 26 can use the second generation strategy to generate test data of a specified length or quantity, where the second generation strategy includes but is not limited to: st.lists strategy, st.data strategy, etc.
[0095] When the data structure is included in the first data attribute set, the construction module 26 can use the third generation strategy to generate test data corresponding to the data structure, where the third generation strategy includes but is not limited to: st.dictionaries strategy, st.tuples strategy, etc.
[0096] When the data boundary values are included in the first data attribute set, the construction module 26 can use the fourth generation strategy to generate test data corresponding to the data structure dimension, where the fourth generation strategy includes but is not limited to: st.floats(min_value,max_value) strategy, st.integers(min_value,max_value) strategy, etc. The maximum and minimum values of different types of data can be generated through these strategies.
[0097] When the data quality is included in the first data attribute set, the construction module 26 can use the fifth generation strategy to generate test data corresponding to the data quality dimension, where the fifth generation strategy includes but is not limited to: st.none strategy, st.just(none), st.integers(min_size,max_size) strategy, st.floats(min_size,max_size) strategy, etc., and can also be obtained by combining the strategies in the first generation strategy, the second generation strategy, the third generation strategy, and the fourth generation strategy.
[0098] Optionally, the tuning module 28 can optimize the performance of the big data middle platform according to the following method, including:
[0099] Step 1: For each simulated test environment, perform performance testing on the big data middleware platform in the simulated test environment to obtain the corresponding performance test results. Among them, the performance test results include test results corresponding to multiple indicator dimensions such as response time, throughput, resource utilization rate, data loss rate, data consistency, etc., so as to comprehensively evaluate the behavior and performance of the big data middleware platform under different conditions.
[0100] Step 2: Determine whether the test results corresponding to each indicator dimension in the performance test results meet the corresponding indicator expectation requirements. Among them, the indicator expectation requirements corresponding to each indicator dimension are set by developers based on historical experience.
[0101] In the case where the test results corresponding to each indicator dimension in the performance test results all meet the corresponding indicator expectation requirements, this means that the big data middleware platform performs well in the current simulated test environment. At this time, the tuning module 28 can adjust the operating environment of the big data middleware platform according to the test data set corresponding to the simulated test environment to optimize the big data middleware platform and ensure that it can also maintain good performance in the actual application environment.
[0102] In the case where the test result corresponding to any indicator dimension in the performance test results does not meet the corresponding indicator expectation requirements, at this time, the tuning module 28 can trigger an automatic feedback mechanism. The system can analyze the test results corresponding to each indicator dimension, identify the root cause of the problem, adjust the test data set corresponding to the simulated test environment, and construct a new simulated test environment based on the adjusted test data set, and obtain the test performance results of the big data middleware platform in the new simulated test environment again.
[0103] Specifically, the tuning module 28 can perform automatic feedback according to the following process to accurately identify and solve the performance bottlenecks and potential problems of the big data middleware platform, including: First, determine the indicator dimension corresponding to the test result that does not meet the indicator expectation requirements (such as too long response time, too low throughput, too high resource utilization rate, increased data loss rate, etc.), and determine at least one target test data in the test data set related to the indicator dimension. For example, for the problem of too long response time, the system can focus on the processing of large-scale data sets or the parsing of specific format data; for the problem of data consistency, the system can review the data cleaning and conversion processes, etc. Then, adjust each target test data according to the preset parameter adjustment strategy set. Among them, the parameter adjustment strategy set covers a series of adjustment strategies corresponding to the test data corresponding to different indicator dimensions, such as adjusting the data scale, modifying the data format, introducing more outliers or changing the simulation parameters of hardware devices, etc. These adjustment strategies can be formulated based on the historical experience of developers, software development, and the usage parameters of hardware devices, etc.
[0104] In the above embodiments, the optimization process is a closed-loop iterative process. The optimization module 28 continuously monitors and analyzes the test results. Once it finds that the performance does not meet the expectations, it will immediately take actions to adjust the test strategy to optimize the simulation test environment until the test results corresponding to all metric dimensions reach a satisfactory level. This iterative test and optimization process ensures the stability and efficiency of the big data middle platform in the face of various data and hardware environments.
[0105] It should be noted that each module in the performance optimization system of the big data middle platform in the embodiments of the present application corresponds one by one to each implementation step of the performance optimization method of the big data middle platform in Embodiment 1. Since detailed descriptions have been made in Embodiment 1, the details not shown in this embodiment can be referred to Embodiment 1 and will not be elaborated here.
[0106] Embodiment 3
[0107] According to an embodiment of the present application, there is also provided a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the performance optimization method of the big data middle platform in Embodiment 1.
[0108] According to an embodiment of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program. The device where the non-volatile storage medium is located executes the performance optimization method of the big data middle platform in Embodiment 1 by running the computer program.
[0109] According to an embodiment of the present application, there is also provided a processor, which is used to run a computer program. When the computer program runs, it executes the performance optimization method of the big data middle platform in Embodiment 1.
[0110] According to an embodiment of the present application, there is also provided an electronic device, which includes: a memory and a processor. The memory stores a computer program, and the processor is configured to execute the performance optimization method of the big data middle platform in Embodiment 1 through the computer program.
[0111] Specifically, when the computer program runs, it executes the following steps: obtaining the business requirement information of the big data middle platform, where the business requirement information is used to reflect the business functions and business objectives of the big data middle platform in the target time period; analyzing the business requirement information by using a pre-trained target large language model to obtain a first data attribute set of the big data middle platform for processing business data in the target time period; generating multiple groups of test data sets based on the first data attribute set, and respectively constructing multiple simulation test environments according to the multiple groups of test data sets; respectively performing performance tests on the big data middle platform in the multiple simulation test environments to obtain multiple performance test results, and adjusting the operating environment of the big data middle platform according to the multiple performance test results.
[0112] As an alternative embodiment, the above-mentioned electronic device may exist in the form of a mobile terminal, a computer terminal, or a similar computing device. Figure 3 The following shows a hardware structure block diagram of an electronic device for implementing a performance optimization method for a big data middle platform. As Figure 3 shown, the electronic device 30 may include one or more processors 302 (shown as 302a, 302b, ……, 302n in the figure) (the processor 302 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission device 306 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 3 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the electronic device 30 may further include more or fewer components than those Figure 3 shown, or have a different configuration from those Figure 3 shown.
[0113] It should be noted that the above-mentioned one or more processors 302 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the electronic device 30. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0114] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the performance optimization method of the big data middle platform in the embodiments of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, implements the vulnerability detection method of the above-mentioned application program. The memory 304 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 304 may further include a memory remotely set relative to the processor 302, and these remote memories may be connected to the electronic device 30 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0115] The transmission device 306 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the electronic device 30. In one example, the transmission device 306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 306 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0116] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the electronic device 30.
[0117] The above-mentioned serial numbers of the embodiments are only for description and do not represent the superiority or inferiority of the embodiments.
[0118] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0119] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0120] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0121] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0122] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0123] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for optimizing the performance of a big data middle platform, characterized in that, Including: Obtain the business requirement information of the big data middle platform, where the business requirement information is used to reflect the business functions and business objectives of the big data middle platform within the target time period; Analyze the business requirement information by using a pre-trained target large language model to obtain a first data attribute set for the big data middle platform to process business data within the target time period; Generate multiple groups of test data sets based on the first data attribute set, and respectively construct multiple simulated test environments according to the multiple groups of test data sets; Perform performance tests on the big data middle platform in multiple simulated test environments respectively to obtain multiple performance test results, and adjust the operating environment of the big data middle platform according to the multiple performance test results.
2. The method according to claim 1, characterized in that, The training process of the target large language model includes: Obtain multiple groups of training sample data, where each group of training sample data includes: the historical business requirement information of the big data middle platform and a second data attribute set for the big data middle platform to process historical business data within a specific historical time period; Obtain a general large language model; Iteratively train the general large language model by using the multiple groups of training sample data to obtain the target large language model.
3. The method according to claim 2, wherein The historical business requirement information includes at least one of the following: the data format, data type, data processing logic, data storage method, data quality requirements, and data organization form for the big data middle platform to process business data, where The data format includes at least one of the following: Comma Separated Values (CSV) format, JavaScript Object Notation (JSON) format, Extensible Markup Language (XML) format; The data type includes at least one of the following: integer, text, floating point number; The data processing logic includes at least one of the following: data cleaning, data transformation, data clustering, data analysis; The data storage method includes at least one of the following: relational database, non-relational database, file system; The data quality requirements include at least one of the following: the field cannot be empty, data value range constraint; The data organization form includes at least one of the following: table, document, image, audio.
4. The method according to claim 2, wherein The second data attribute set includes at least one of the following: data type, data scale, data structure, data format, data boundary value, data quality.
5. The method according to claim 4, wherein The test data in multiple dimensions in each group of test data sets, where generating multiple groups of test data sets based on the first data attribute set includes: When the data type is included in the first data attribute set, use a first generation strategy to generate test data corresponding to the data type dimension, where the first generation strategy includes at least one of the following: st.integers strategy, st.text strategy; When the data scale is included in the first data attribute set, use a second generation strategy to generate test data corresponding to the data scale dimension, where the second generation strategy includes at least one of the following: st.lists strategy, st.data strategy; When the data structure is included in the first data attribute set, test data corresponding to the data structure is generated using a third generation strategy, where the third generation strategy includes at least one of the following: st.dictionaries strategy, st.tuples strategy; When the data boundary values are included in the first data attribute set, test data corresponding to the data structure dimension is generated using a fourth generation strategy, where the fourth generation strategy includes at least one of the following: st.floats(min_value,max_value) strategy, st.integers(min_value,max_value) strategy; When the data quality is included in the first data attribute set, test data corresponding to the data quality dimension is generated using a fifth generation strategy, where the fifth generation strategy includes at least one of the following: st.none strategy, st.just(none), st.integers(min_size,max_size) strategy, st.floats(min_size,max_size) strategy.
6. The method according to claim 1, wherein Performance tests are respectively conducted on the big data middle platform in multiple simulation test environments to obtain multiple performance test results, and the operating environment of the big data middle platform is adjusted according to the multiple performance test results, including: For each simulation test environment, a performance test is conducted on the big data middle platform in the simulation test environment to obtain the corresponding performance test result, where the performance test result includes test results of multiple metric dimensions, and the metric dimensions include at least one of the following: response time, throughput, resource utilization rate, data loss rate, data consistency; Judge whether the test results corresponding to each metric dimension in the performance test result meet the corresponding metric expectation requirements; When the test results corresponding to each metric dimension in the performance test result all meet the corresponding metric expectation requirements, the operating environment of the big data middle platform is adjusted according to the test data set corresponding to the simulation test environment; When the test result corresponding to any metric dimension in the performance test result does not meet the corresponding metric expectation requirement, the test data set corresponding to the simulation test environment is adjusted, and a new simulation test environment is constructed based on the adjusted test data set, and the test performance result of the big data middle platform in the new simulation test environment is obtained again.
7. The method according to claim 6, characterized in that Adjusting the test data set corresponding to the simulation test environment includes: Determine the metric dimension corresponding to the test result that does not meet the metric expectation requirement, and determine at least one target test data in the test data set related to the metric dimension; Adjust each target test data according to a preset parameter adjustment strategy set, where the parameter adjustment strategy set includes: adjustment strategies corresponding to test data of different metric dimensions.
8. A performance optimization system for a big data middle platform, characterized in that, Include: An acquisition module, configured to acquire business requirement information of a big data middle platform, wherein the business requirement information is used to reflect the business functions and business objectives of the big data middle platform within a target time period; A prediction module, configured to analyze the business requirement information by using a pre-trained target large language model to obtain a first data attribute set for the big data middle platform to process business data within the target time period; A construction module, configured to generate multiple groups of test data sets based on the first data attribute set, and respectively construct multiple simulated test environments according to the multiple groups of test data sets; An optimization module, configured to perform performance tests on the big data middle platform in multiple simulated test environments respectively to obtain multiple performance test results, and adjust the operating environment of the big data middle platform according to the multiple performance test results.
9. A computer program product, characterized in that, Comprising: A computer program, wherein when the computer program is executed by a processor, it implements the performance optimization method of the big data middle platform according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Comprising: A memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the performance optimization method of the big data middle platform according to any one of claims 1 to 7 through the computer program.
Citation Information
Cited By
Data processing method and device, nonvolatile storage medium and electronic equipment
CN120780573A