Automatic test method and device of big data, electronic equipment and storage medium

Through automated testing methods, we can obtain the table structure and field information of the big data platform, generate distributed data query statements, execute scheduling tasks and detect anomalies. This solves the problem of low efficiency of traditional big data testing and improves test quality and stability, especially in data verification scenarios in the financial and medical industries.

CN120743897APending Publication Date: 2025-10-03CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510860751.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional big data testing solutions are inefficient and have unstable test quality. They are unable to cope with high-frequency data iteration requirements and have incomplete coverage, especially in data verification scenarios in the financial and medical industries, exposing data discrepancies. Human factors also lead to unstable test quality.

Method used

This paper provides an automated testing method for big data. It obtains the table structure and field information of distributed storage, generates distributed data query statements, executes scheduling tasks and determines data anomalies. It combines static parsing and dynamic tracking parameters to detect code compliance and dependencies and generate automated test results.

Benefits of technology

It achieves full-link quality assurance, improves the testing efficiency and stability of the big data platform, reduces the failure rate in the production environment, and ensures the accuracy and consistency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743897A_ABST
    Figure CN120743897A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic testing, can be applied to the field of science and technology finance / digital medical treatment, and discloses an automatic testing method and device for big data, electronic equipment and a storage medium. The method comprises the following steps: acquiring a table structure and field information of distributed storage, and generating a distributed data query statement by using a screened result; generating a scheduling task by using the distributed data query statement, executing the scheduling task, and judging whether data exception exists or not according to an execution result; when a first submission request of a target code is received, performing static analysis on a big data processing script contained in the request, and checking the compliance of grammar and parameters; in the running stage of the target code, parameters of the big data processing script are dynamically analyzed, and transmission logic of the parameters in distributed nodes is tracked; and when a second submission request of the big data scheduling task is received, detecting whether the parameters and the dependency relationship of the big data scheduling task are correct or not. According to the method, the big data platform test efficiency and stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automated testing technology and can be applied to the fields of science and technology finance / digital medicine, and in particular to an automated testing method, device, electronic device and storage medium for big data. Background Art

[0002] Traditional data testing relies heavily on manual labor, exposing multiple technical bottlenecks. First, testing efficiency is low. For example, in an industry processing tens of millions of data items daily, traditional methods for verifying field integrity and data consistency can take hours and struggle to cope with the high frequency of data iteration. Second, self-test case coverage is incomplete. Traditional design cases often miss key scenarios, such as horizontal comparisons of cross-system data (for example, matching records between system A and system B) or vertical consistency of historical data (for example, connecting states of the same business at different time points), leading to data discrepancies in production environments. Third, coverage of special scenarios is insufficient. Traditional testing often oversimplifies handling of extreme conditions (for example, concurrent data reading and writing under peak traffic) or complex data formats (for example, unstructured records containing special symbols). For example, in one field, rare data formats are not validated, leading to storage system anomalies. Finally, human factors dominate test quality. Manual code review can easily overlook static syntax defects (for example, function misuse leading to data truncation) or dynamic parameter errors (for example, fixed-time parameters causing logical confusion). When configuring scheduled tasks, manual errors can also lead to missing dependencies, resulting in task execution failures. The above problems are particularly prominent in scenarios such as transaction data verification in the financial industry and medical record data verification in the medical industry. Traditional manual models can no longer meet the requirements of the big data era for testing efficiency, coverage accuracy and result stability. Summary of the Invention

[0003] The main technical problem solved by the implementation method of this application is that traditional big data testing solutions are inefficient and have unstable test quality.

[0004] To solve the above technical problems, the first technical solution adopted by the embodiment of the present application is: to provide an automated testing method for big data, including: obtaining the table structure and field information of distributed storage based on the metadata of the target big data platform, receiving the user's screening of the table structure and field information, and using the screening result to generate a corresponding distributed data query statement; using the distributed data query statement to generate a scheduling task, executing the scheduling task through a preset task scheduling platform, and judging whether there is a data anomaly based on the execution result, to obtain a first test result; when receiving a first submission request of the target code, scanning the big data processing script contained in the first submission request, statically parsing the big data processing script, verifying the syntax and parameter compliance, and obtaining a second test result; during the running phase of the target code, monitoring the big data task log, dynamically parsing the parameters of the big data processing script, tracking the parameter transmission logic in the distributed nodes, and obtaining a third test result; when receiving a second submission request for the big data scheduling task, detecting whether the parameters and dependencies of the big data scheduling task are correct according to the preset task scheduling specification, and obtaining a fourth test result; using the first test result, the second test result, the third test result, and the fourth test result to generate an automated test result corresponding to the target big data platform.

[0005] Optionally, after the step of executing the scheduling task through a preset task scheduling platform and judging whether there is a data anomaly based on the execution result, it also includes: constructing a task time transmission model based on the big data link topology, configuring the time threshold of each big data node, collecting the execution time of the scheduling task on each big data node in real time, comparing the execution time with the time threshold, and obtaining the big data node with task time anomaly; calculating the fluctuation percentage of the number of rows in the big data table, verifying whether the storage size, partition row number and partition storage size of the big data table are within the preset threshold range; counting the null value rate and unique value ratio of the fields in the big data table, detecting whether the null value rate and the unique value ratio are within the preset threshold range, and verifying the business validity of the enumeration value.

[0006] Optionally, after the step of executing the scheduling task through a preset task scheduling platform and judging whether there is data anomaly based on the execution result, the step further includes: performing a horizontal consistency check of the field values ​​on the associated big data table to detect whether the matching degree of the cross-table field data is within a preset threshold range; analyzing the longitudinal fluctuation trend of the time series data in the big data table, and identifying abnormal values ​​and mutation time data points based on the longitudinal fluctuation trend; comparing the changes in the preset big data table indicator data before and after the data in the big data table is modified, and judging whether there is modified abnormal data based on the change results; verifying whether the data processing logic in the big data table is correct based on the time boundaries set by the business rules in the business system; and detecting the processing accuracy of the data under the extreme value boundaries of the maximum and minimum values ​​of the big data table field value range based on the extreme value boundaries.

[0007] Optionally, the step of statically parsing the big data processing script and verifying the compliance of syntax and parameters includes: statically scanning the big data processing script to detect whether the big data processing script contains undefined variables, undefined first-trial variables, and time parameters set to fixed values; performing syntax detection on the big data processing script to check whether the truncation operation syntax in the script is correct, verifying whether the syntax of the data definition statements in the script is correct, and detecting whether the case conversion operation in the script is correct; obtaining the association logic and association conditions of multiple data objects in the big data processing script, and detecting whether there are association errors in the association logic and the association conditions.

[0008] Optionally, the step of dynamically parsing the parameters of the big data processing script and tracing the parameter transfer logic in the distributed nodes includes: intercepting the parameter transfer data stream of the big data processing script, traversing the parameter assignment status in the parameter transfer data stream, and identifying the traversed abnormal parameters that are not assigned or have default null values; marking the parameter transfer link according to the traversed abnormal parameters, and associating them with the code position in the big data processing script; constructing verification rules through preset parameter constraints, and judging whether the parameter assignment exceeds the constraint range according to the verification rules, and if it exceeds the constraint range, judging that the corresponding parameter assignment has an overflow exception.

[0009] Optionally, the step of detecting whether the parameters and dependencies of the big data scheduling task are correct according to the preset task scheduling specifications includes: traversing the parameter configuration items of the big data scheduling task, and identifying errors in filling in basic parameters and dependent parameters according to preset parameter verification rules; recording the task node, parameter type and verification failure reason corresponding to the erroneous parameter; extracting the execution logic, parameter configuration and dependency link of the big data scheduling task in the dual-transmission environment as dual-transmission detection information, and detecting whether there is inconsistency in the dual-transmission detection information.

[0010] Optionally, before the step of using the first test result, the second test result, the third test result and the fourth test result to generate the automated test result corresponding to the target big data platform, it also includes: traversing the associated code of the big data scheduling task and the version identifier of the configuration file of the target big data platform; identifying whether there are abnormal errors that overlap each other in different versions of the associated code and the configuration file through hash verification and text difference comparison algorithm; monitoring the change records and deployment logs of the big data scheduling task, and verifying whether the submitted code and configuration changes are completely synchronized to the target operating environment based on the change records and deployment logs.

[0011] In order to solve the above technical problems, the second technical solution adopted by the embodiment of the present application is: to provide an automated testing device for big data, including: a test query statement generation module, which is used to obtain the table structure and field information of distributed storage according to the metadata of the target big data platform, receive the user's screening of the table structure and field information, and use the screening results to generate corresponding distributed data query statements; a first test result module, which is used to use the distributed data query statement to generate a scheduling task, execute the scheduling task through a preset task scheduling platform, and judge whether there is a data anomaly based on the execution result to obtain a first test result; a second test result module, which is used to scan the big data processing script contained in the first submission request when receiving the first submission request of the target code, and perform the first test on the first submission request. The big data processing script is statically parsed to verify the compliance of syntax and parameters to obtain a second test result; a third test result module is used to monitor the big data task log during the running phase of the target code, dynamically parse the parameters of the big data processing script, and track the parameter transfer logic in the distributed nodes to obtain a third test result; a fourth test result module is used to detect whether the parameters and dependencies of the big data scheduling task are correct according to the preset task scheduling specifications when a second submission request for the big data scheduling task is received to obtain a fourth test result; an automated test result module is used to use the first test result, the second test result, the third test result and the fourth test result to generate an automated test result corresponding to the target big data platform.

[0012] To solve the above technical problems, the third technical solution adopted in the embodiment of the present application is: to provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the automated testing method for big data as described above.

[0013] In order to solve the above technical problems, the fourth technical solution adopted in the embodiment of the present application is: to provide a non-volatile computer-readable storage medium, characterized in that the non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by an electronic device, the electronic device executes the automated testing method for big data as described above.

[0014] Different from the related technologies, this application achieves full-link quality assurance through a multi-dimensional automated testing mechanism. The data layer generates distributed queries based on metadata and verifies the execution results, identifying data anomalies in real time. The code layer combines static parsing and dynamic tracking to cover parameter verification in the development and operation stages. The scheduling layer pre-detects task configurations and dependencies to avoid scheduling failures due to configuration errors. Finally, the multi-stage test results are integrated to build an automated test closed loop, improve the testing efficiency and stability of the big data platform, and reduce the failure rate in the production environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] One or more embodiments are exemplarily illustrated by corresponding drawings, which do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0016] Figure 1 This is a schematic diagram of the operating environment of the automated testing method for big data provided in an embodiment of the present application.

[0017] Figure 2 This is a schematic diagram of the execution flow of the automated testing method for big data provided in an embodiment of the present application.

[0018] Figure 3 This is a schematic diagram of the execution flow of data automation testing in the big data automation testing method provided in the embodiment of the present application.

[0019] Figure 4 This is another execution flow diagram of data automation testing in the big data automation testing method provided in the embodiment of the present application.

[0020] Figure 5 This is a schematic diagram of the system structure of the automated testing device for big data provided in an embodiment of the present application.

[0021] Figure 6 This is a schematic diagram of the hardware structure of an electronic device that performs an automated testing method for big data provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0023] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematics and the logical order is shown in the flow charts, in some cases, the steps shown or described can be performed in a different order than the module division in the device schematics or the order in the flow charts.

[0024] Unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification are intended only to describe specific embodiments and are not intended to limit this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the relevant listed items.

[0025] To facilitate understanding of this embodiment, first, a large data automation testing method disclosed in the embodiment of this application is introduced in detail. Figure 1 , Figure 1 Schematic diagram of the operating environment of the automated testing method for big data provided in the embodiment of the present application. Figure 1 As shown, the execution subject of the automated testing method for big data provided in the embodiment of the present application is generally an electronic device with certain computing capabilities, such as a computer device. In some possible implementations, the automated testing method for big data can be implemented by a processor calling computer-readable instructions stored in a memory. Figure 1 The computer device in the above can be a server. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. It can be understood that Figure 1 The number of computer devices in the figure is only for reference and can be expanded to any number according to actual needs.

[0026] Please continue reading Figure 2 , Figure 2 This is a schematic diagram of the execution flow of the automated testing method for big data provided by the embodiment of the present application, such as Figure 2 As shown, the following steps are included:

[0027] S1. According to the metadata of the target big data platform, obtain the table structure and field information of the distributed storage, receive the user's screening of the table structure and field information, and use the screening results to generate the corresponding distributed data query statement.

[0028] Based on the metadata management mechanism of the target big data platform, an interactive filtering engine is constructed by parsing the table structure metadata (including field names, data types, partition information, etc.) and field-level description information in the distributed storage system. Users can perform multi-dimensional conditional filtering based on field attributes (such as data type constraints, business tags, update timestamps, etc.) and inter-table relationships (such as foreign key constraints and subject domain affiliation) through a visual interface or script input. The filtering conditions are automatically converted into distributed query logic that conforms to the syntax specifications of the target platform, generating a query plan with parallel execution capabilities, and finally outputting standardized query statements suitable for distributed computing frameworks such as Hive and Spark SQL.

[0029] S2. Use a distributed data query statement to generate a scheduling task, execute the scheduling task through a preset task scheduling platform, and determine whether there is a data anomaly based on the execution result to obtain a first test result.

[0030] Among them, based on the task scheduling mechanism of the distributed computing framework, the standardized query statements are encapsulated into schedulable task units, and the task scheduling API of the integrated target platform (for example, Apache Oozie, Azkaban workflow definition interface) is used to generate task instances according to the preset scheduling strategy (timing trigger, event trigger or dependency trigger). The scheduling task allocates computing resources through the resource manager (for example, YARN), executes parallel data queries in the distributed cluster, and collects task execution indicators (such as the amount of data processed, execution time, node load, etc.) in real time. The execution results are verified through the preset anomaly detection rules (such as the result set is empty, the number of data rows fluctuates beyond the threshold, the field value range exceeds the business scope, etc.), and the records that meet the abnormal conditions are marked as data anomalies. Finally, the first test result report containing the anomaly type, the occurrence node and the timestamp is generated.

[0031] As an optional implementation, see Figure 3 , Figure 3 FIG. 1 is a schematic diagram of the execution flow of data automation testing in the big data automation testing method provided in the embodiment of the present application, such as Figure 3 As shown, after the above step S2, the timeliness and integrity of the data in the big data table may be tested, which may specifically include the following steps S21 to S23.

[0032] S21. Build a task timeliness transmission model based on the big data link topology, configure the timeliness threshold of each big data node, collect the execution time of the scheduled task on each big data node in real time, compare the execution time with the timeliness threshold, and obtain the big data node with abnormal task timeliness.

[0033] For example, a big data link topology model is constructed based on graph theory, and tasks are broken down into nodes such as data collection, cleaning, conversion, and storage, and directed edges are defined to represent data flow relationships. A time-sensitive transmission model is established through a time series database (for example, InfluxDB), and a time-sensitive threshold based on historical mean, quantile, or business SLA is configured for each node (for example, ETL node ≤ 30 minutes). Using APM tools (for example, Pinpoint) or task scheduling platform hook functions, the start / end timestamps of tasks at each node are captured in real time, the execution time is calculated and dynamically compared with the threshold, and the early warning mechanism is triggered to mark the timed-out node. At the same time, the upstream and downstream influencing links are traced through the path search algorithm.

[0034] S22. Calculate the row number fluctuation percentage of the big data table to verify whether the storage size, partition row number, and partition storage size of the big data table are within a preset threshold range.

[0035] For example, table-level statistics are obtained through the metadata interface of a big data computing engine (e.g., Spark), and the percentage of row number fluctuations is calculated in parallel using a distributed computing framework. Based on the file metadata of storage systems such as HDFS / NAS, the table partition files are recursively traversed to calculate the storage size (number of bytes) and the number of partition rows. Pre-trained anomaly detection models (e.g., isolation forest, time series prediction models) are used to determine whether indicators exceed the range or custom thresholds, and alarms are generated for abnormal scenarios such as sudden changes in storage size (e.g., growth exceeding 50%) and partition data skew (e.g., the number of rows in a single partition exceeds 80%).

[0036] S23. Count the null value rate and unique value ratio of the fields in the big data table, detect whether the null value rate and unique value ratio are within the preset threshold range, and verify the business validity of the enumeration value.

[0037] For example, use the aggregation functions of Spark SQL / Hive SQL (such as COUNT_NULLS, APPROX_COUNT_DISTINCT) to calculate the null value rate (number of null values / total number of rows) and unique value ratio (number of unique values / total number of rows) of fields in a distributed manner. Parse the enumerated value fields through UDF functions (such as dictionary table association) to verify business validity (such as whether the gender field value is within the set {'male', 'female'}). Compare the calculation results with the threshold rules in the configuration center (such as Apollo) (such as null value rate ≤ 5%, unique value ratio ≥ 80%), generate a quality report for fields outside the range, and locate the abnormal data source nodes in combination with data lineage (such as missing field values in the upstream collection interface).

[0038] As another alternative implementation, please refer to Figure 4 , Figure 4 which is another schematic diagram of the execution process for data automated testing in the big data automated testing method provided by the embodiments of this application. As shown in Figure 4 , after the above step S2, it is also possible to test the accuracy of the data in the big data table, which can specifically include the following steps S24 to S28.

[0039] S24. Perform horizontal consistency verification of field values on the big data tables with an association relationship, and detect whether the matching degree of cross-table field data is within the preset threshold range.

[0040] For example, build a cross-table association rule engine based on data lineage, and obtain relationship definitions such as foreign key associations and business key mappings between tables through the metadata management system (such as the order table and the user table are associated through the user ID). Use distributed join calculation (such as Spark Shuffle Join) to align the field values of the associated tables, and calculate the field matching degree through similarity algorithms such as the Jaccard coefficient and edit distance (such as the text similarity between the user address field in the order table and the address field in the user table); compare the matching degree with the preset threshold (such as greater than or equal to 95%), generate a difference report for the associated field combinations below the threshold, and locate the data inconsistent nodes (such as field mapping errors in the ETL process).

[0041] S25. Analyze the longitudinal fluctuation trend of the time series data in the big data table, and identify outliers and mutation time data points based on the longitudinal fluctuation trend.

[0042] For example, time series analysis algorithms (such as ARIMA and Prophet) are used to model time series data (e.g., daily active users, transaction amounts). Sliding windows are used to calculate statistics such as rolling mean and variance to identify outliers that deviate from the 3σ range. Change point detection algorithms (such as the Pettitt test and the CUSUM algorithm) are used to detect sudden changes in data distribution and locate timestamps of trend reversals or mean jumps in the time series (e.g., a sharp drop in data after a promotional period). Data version management systems (such as Hudi and Delta Lake) are used to track data changes before and after the sudden change point and analyze the causes of fluctuations (e.g., data source switching, changes in computational logic).

[0043] S26. Compare the changes in the preset big data table indicator data before and after the data in the big data table is modified, and determine whether there is abnormal modification data based on the change results.

[0044] For example, through CDC (change data capture) technology (such as Debezium), the INSERT / UPDATE / DELETE operations of the data table are monitored, and the old version snapshot before the data modification and the new version data after the modification are captured within the transaction boundary. Based on the predefined indicator calculation logic (such as the total number of rows, field mean, unique value count, etc.), indicator snapshots are generated for the old and new versions of the data respectively (stored in a time series database or cache). The difference between the indicators before and after is compared through vector similarity calculation (such as cosine similarity), and an absolute threshold (for example, the number of rows changes > 1000) or a relative threshold (for example, the mean fluctuation > 20%) is set to trigger anomaly detection. Modification operations that exceed the threshold are marked as suspicious changes, and the source of the operation (for example, incorrect operation, program bug) is traced in combination with the audit log.

[0045] S27. Verify whether the data processing logic in the big data table is correct based on the time boundaries set by the business rules in the business system.

[0046] For example, time boundary constraints are extracted from the business rule engine (for example, transaction data must be paid within 24 hours of order creation) and converted into logical expressions based on timestamp fields (for example, payment time ≤ order creation time + 24 hours). Using stream computing frameworks such as Spark Streaming or Flink, real-time incoming data is divided into time windows, and the CEP (complex event processing) engine is used to detect event sequences that violate time order (for example, abnormal processes such as refunding first and then paying). For offline data, time logic is batch-checked through partition pruning (for example, filtering by date partitions) and vectorized calculations (for example, Pandas UDFs), generating diagnostic reports containing violation record IDs and time deviation values.

[0047] S28. According to the extreme value boundaries of the maximum value and the minimum value of the value range of the field in the big data table, the processing accuracy of the data under the extreme value boundaries is detected.

[0048] For example, the extreme value boundaries of the field value range can be obtained through the aggregation functions (e.g., MIN / MAX) of the big data computing engine (e.g., the minimum value of the age field is 0 and the maximum value is 150), and the processing logic for extreme value scenarios can be defined in combination with business rules (e.g., marking age = 0 as test data and triggering manual review when age > 120). A test data set containing extreme boundary values ​​can be constructed (e.g., equivalence class partitioning in boundary value analysis) and injected into the data processing pipeline (e.g., ETL jobs, data warehouse models) for penetration testing. The processing results can be verified through the assertion mechanism (e.g., whether age = 0 is written to a specific partition and whether age > 150 is filtered). Coverage analysis tools (e.g., JaCoCo) can be used to detect code path coverage in extreme value scenarios to ensure the robustness of the data processing logic under boundary conditions.

[0049] As an example, in a bank's credit risk control data chain, a time-sensitive transmission model is constructed based on a big data chain topology, encompassing nodes such as customer data acquisition, risk score calculation, and credit report extraction. A 10-second time-sensitive threshold is configured for the customer data acquisition node, and the execution time of this node in the distributed cluster is collected in real time. If the execution time suddenly increases to 30 seconds and exceeds the threshold, an alert is triggered, indicating a possible data interface delay or anomaly. A horizontal consistency check is performed on the loan application form and the customer information table, checking the matching degree of cross-table fields such as customer ID number and mobile phone number. If the matching degree is less than 90%, a data synchronization anomaly is reported. Time series data of daily transaction amounts is analyzed, and the Prophet algorithm is used to identify data points where an account experiences unusually large transfers during non-business hours. Anti-money laundering rules are then used to verify whether the funds flow complies with business logic. After modifying the customer risk level calculation logic in the data warehouse, indicators such as the proportion of high-risk customers and the average risk score are compared before and after the modification. If the proportion of high-risk customers suddenly drops by 40%, anomaly detection is triggered to investigate whether the risk misjudgment was caused by a logical vulnerability. Verify that POS transaction data is completed within 30 minutes of terminal check-in, based on regulatory transaction time limits. Any transactions exceeding this time limit will be marked as suspicious. For the account balance field, construct extreme value test data, including a balance of 0 (minimum) and the system's maximum value (e.g., 9999999999). This test verifies that the core system correctly triggers overdraft warnings when the balance is 0 and that no precision loss or calculation errors occur when extreme value amounts are recorded.

[0050] As another example, in a hospital's laboratory data processing chain, a time-sensitive transmission model is constructed based on a big data chain topology, encompassing nodes such as sample collection time synchronization, test result parsing, and report generation. A 15-minute time-sensitive threshold is set for the test result parsing node, and the execution time of this node in the Hadoop cluster is collected in real time. If a batch of samples takes 45 minutes to parse, exceeding the threshold, it indicates possible congestion in the test instrument interface or data format anomalies. A horizontal consistency check is performed on field values ​​in the patient's electronic medical record and the test result table, checking the matching degree of cross-table fields such as patient ID and test submission time. If the matching degree is less than 95%, it indicates an incorrect association between the sample and the medical record. Time series data of blood glucose levels of diabetic patients is analyzed, and a sliding window algorithm is used to identify an abnormal value, such as a patient's blood glucose level suddenly dropping below 3.9 mmol / L at 3:00 AM. Based on clinical guidelines, this value is determined to be a critical hypoglycemic value and trigger an alert. After optimizing the imaging data archiving logic, the image file storage size and retrieval time are compared before and after the modification. A 50% increase in retrieval time indicates a possible flaw in the index construction. Based on the business rule that prescriptions must be issued after the patient's registration time, the logical relationship between prescription order and registration time in prescription data is verified, and records that violate the chronological order are marked as data anomalies. For the white blood cell count field in routine blood test reports, extreme value data of 0 (minimum value) and 50×10^9 / L (clinically abnormally high value) are constructed to test whether the inspection system correctly indicates a specimen dilution error when the white blood cell count is 0, and whether unit conversion errors occur when entering extreme value data.

[0051] Through the above steps S21 to S28, by building a time transmission model and configuring thresholds, the node time consumption is monitored in real time to quickly locate the time abnormality nodes and ensure the timeliness of the data processing process. Calculate the row number fluctuation percentage and verify the storage size, partition row number and other indicators to effectively identify abnormal data volume fluctuations and storage skew problems. Count the field null rate and unique value ratio and verify the business validity of the enumeration value to improve data integrity and business compliance. Perform horizontal consistency check of field values ​​on related tables to ensure accurate association of cross-table data. Analyze the vertical fluctuation trend of time series data, identify outliers and mutation points, and provide early warning for business decisions. Compare the changes in indicators before and after data modification to accurately locate data modification anomalies. Verify the processing logic based on business time boundaries and field extreme value boundaries to enhance the robustness of the system in extreme scenarios and comprehensively improve the quality and efficiency of big data processing.

[0052] S3. When the first submission request of the target code is received, the big data processing script contained in the first submission request is scanned, the big data processing script is statically parsed, the syntax and parameter compliance are verified, and a second test result is obtained.

[0053] When the first commit request for the target code is received, the code quality check process is automatically triggered. First, by integrating Git Hooks or a CI / CD pipeline (e.g., Jenkins or GitLab CI), the code commit event is intercepted and the big data processing script (e.g., PySpark or HiveQL file) is extracted. Then, using a parser like ANTLR, an abstract syntax tree (AST) is constructed for the script, and static code analysis is performed on the script. This verifies syntax compliance, for example, by checking whether SQL statements conform to Hive syntax specifications and whether Python code contains syntax errors such as unclosed parentheses. It also verifies parameter configuration compliance, for example, by checking whether partition fields conform to the table structure definition and whether dynamic parameters use placeholders rather than hard-coded values. Furthermore, a custom rule engine (e.g., a SonarQube plugin) is used to detect potential risks, such as identifying undefined variables and hard-coded timestamps. This entire process eliminates code execution and quickly identifies code defects through static scanning. Finally, a second test result report is generated, containing the error type, line number, and recommended fixes, ensuring that the submitted big data processing script meets quality standards.

[0054] As an optional implementation, the process of scanning the big data processing script contained in the first submission request in the above step S3, statically parsing the big data processing script, and verifying the compliance of syntax and parameters can specifically include the following steps S31 to S33.

[0055] S31. Perform a static scan on the big data processing script to detect whether the big data processing script contains undefined variables, undefined variables that are used first, and time parameters set to fixed values.

[0056] Among them, a big data script static scanner is built based on the lexical analysis and symbol table mechanism in compilation principles. All variable reference nodes are traversed through the abstract syntax tree, and a symbol table is established to track the order of variable declaration and use. Variables that are not declared in the scope or variables that are used first and then defined are marked. At the same time, regular expressions are used to match time format strings, and semantic analysis is combined to determine whether they are hard-coded time values. For literals that conform to the ISO time format, the configuration center is used to check whether there is a corresponding dynamic parameter configuration, and the problem of script non-reusability caused by fixed time parameters is identified.

[0057] S32. Perform syntax checking on the big data processing script to check whether the syntax of the truncation operation in the script is correct, verify whether the syntax of the data definition statement in the script is correct, and check whether the case conversion operation in the script is correct.

[0058] Among them, fine-grained syntax verification is implemented based on a specific syntax rule engine. The ANTLR parser is used to identify function calls such as TRUNCATE and SUBSTRING in SQL scripts to verify whether the parameter type and quantity comply with the syntax specifications. The boundary legality of string slicing operations in Python scripts is checked. Statements such as CREATE TABLE and ALTER TABLE are parsed to verify the field type and platform compatibility, whether the partition key complies with the table design, and whether the constraints are legal. Function calls such as UPPER and LOWER are analyzed to check whether the parameters are string types to avoid invalid conversions of numeric type fields. For case-sensitive databases, the case consistency of table names and column names in quotation marks is specially checked.

[0059] S33. Obtain the association logic and association conditions of multiple data objects in the big data processing script, and detect whether there are association errors in the association logic and association conditions.

[0060] Among them, association logic errors are identified through data lineage analysis and pattern matching, query rewriting technology is used to parse JOIN, UNION and other operations, data object association graphs are constructed and association fields and conditional expressions are extracted, the data types of association keys are checked to see if they are consistent and whether there are spelling errors in field names, true / false conditions and contradictory conditions are identified and the validity of conditions is evaluated through a symbolic calculation engine, the risk of data expansion caused by circular dependencies is detected in complex association relationships, and the association logic is compared with the business requirements document to verify whether it meets expectations.

[0061] As an example, in a bank loan approval big data processing script, a static scan discovered that a SQL script referenced the undefined variable "overdue_days" (used to calculate the number of days a customer is overdue) and used it in a WHERE condition without declaring it, resulting in a logical error. The script also detected that the time parameter was hard-coded as "'2024-01-01'," making it unable to dynamically adapt to monthly data processing needs. Syntax checking revealed that the TRUNCATE statement was incorrectly written as "TRUNCAT table loan_application," lacking the keyword "E" and resulting in a syntax error. Furthermore, the data definition statement set the "loan_amount" field to CHAR type instead of NUMERIC, causing anomalies in the amount calculation. Association logic testing revealed that the loan application form and the credit report form were linked through "customer_no" and "cust_id," but these two fields were actually the same business identifier but were not uniformly named, resulting in a cross-table data matching failure. The association conditions needed to be corrected to ensure data consistency.

[0062] As another example, in a hospital drug inventory management big data script, static scanning identified the Python script's use of an undefined variable, "stock_warning" (used to trigger inventory warnings), which was prematurely called in a conditional judgment, resulting in a runtime error. Furthermore, the time parameter was fixed at "18:00:00" as the inventory cutoff time, failing to account for business scenarios in different time zones. Syntax checking revealed that the SUBSTRING function parameter in HiveQL was written as "drug code, 3" instead of "drug code, 3, 5," resulting in inaccurate results due to a lack of truncation. Furthermore, the field "expiration_date" was incorrectly defined as an INT type instead of a DATE type, making it impossible to correctly verify the drug's expiration date. Association logic checking revealed that the drug details table and the supplier table were linked via "supplier_code" and "supplier_id," with the former being a character type and the latter a numeric type. This data type mismatch caused the implicit conversion to fail, necessitating the unification of field types to ensure correct association logic.

[0063] Through the above steps S31 to S33, by building a static code detection system, comprehensive quality control of big data processing scripts is achieved, and undefined variables, incorrect variable usage order and hard-coded time parameters are automatically identified to avoid runtime exceptions and improve script reusability. Based on the grammar rule engine, grammatical details such as truncation operations, data definitions and case conversion are checked to ensure that the script complies with platform specifications. Through data lineage analysis and pattern matching, field inconsistencies, invalid conditions and logical loops in multi-table associations are detected to improve the accuracy and efficiency of cross-table data processing. The three work together to discover potential defects early in the development stage, reduce testing costs and production environment failures, and ensure the correctness and stability of big data processing logic.

[0064] S4. During the running phase of the target code, the big data task log is monitored, the parameters of the big data processing script are dynamically parsed, and the parameter transmission logic in the distributed nodes is tracked to obtain the third test result.

[0065] During the target code's runtime, dynamic parameter monitoring is achieved by building a distributed log collection and analysis system. First, Flume or Logstash is used to collect task logs from each computing node (e.g., Spark Executor, HadoopTaskTracker) in real time and aggregate them into a Kafka message queue to form a unified data stream. Then, script parameters are tracked using custom annotations (e.g., @TrackParameter) or bytecode enhancement techniques (e.g., AspectJ). Tracing logic is injected into key parameter transfer nodes (e.g., function calls, cross-node RPCs), generating timestamped parameter flow links. Next, the Flink stream processing engine parses the parameter context in the logs, constructs a directed graph model to represent the parameter transfer path in the DAG, and applies a pattern matching algorithm to identify anomalies such as parameter type mismatches, null value transfers, and range violations. Finally, Prometheus monitoring metrics (e.g., task execution time, memory usage) are combined with parameter flow data to generate a third test result report that includes parameter lifecycle, abnormal node location, and performance impact analysis. This process achieves a deep understanding of the execution logic of big data processing scripts through runtime parameter transparency, and discovers hidden defects in the parameter passing process in advance.

[0066] As an optional implementation, the process of dynamically parsing the parameters of the big data processing script and tracking the parameter transmission logic in the distributed nodes in the above step S4 can specifically include the following steps S41 to S43.

[0067] S41. Intercept the parameter transfer data stream of the big data processing script, traverse the parameter assignment status in the parameter transfer data stream, and identify the traversed abnormal parameters that are not assigned or have default null values.

[0068] Among them, a parameter data stream interception system is built based on bytecode enhancement and stream processing technology. An interceptor is implanted in the key path of parameter transfer (for example, Task serialization, RPC call) of the big data processing framework (for example, Spark, Flink), and the bytecode is modified through ASM or Byte Buddy to capture the assignment status during parameter transfer. A sliding window algorithm is used to traverse the parameter key-value pairs in the data stream, and a parameter state machine is established to track the assignment process. Combined with taint analysis technology, parameters that are not explicitly assigned or initialized to NULL are marked, and abnormal parameters caused by missing configuration, uncovered conditional branches, etc. are identified, and an exception list containing parameter name, type, and transfer stage is generated.

[0069] S42. Mark the parameter transmission link according to the traversed abnormal parameters and associate the code location in the big data processing script.

[0070] Among them, abnormal parameters are located by constructing a parameter lineage map. Based on the marked abnormal parameters, the MDC (Mapped Diagnostic Context) technology is used to carry context information (such as Task ID, function call stack) during the parameter passing process to build a directed acyclic graph (DAG) of parameter passing. By parsing the stack information in the log, the abnormal parameters are associated with the source code line number (for example, through Log4j's LocationInfo), and combined with the version control system (Git) to map to the specific code submission. For complex parameters (for example, nested objects), path expressions (for example, $.config.timeout) are used to accurately locate the abnormal fields, and finally a visual parameter transmission chain diagram is generated to intuitively show the complete path of the abnormal parameters from the entry to the error point.

[0071] S43. Construct verification rules based on preset parameter constraints, and determine whether the parameter assignment exceeds the constraint range according to the verification rules. If it exceeds the constraint range, it is determined that the corresponding parameter assignment has an overflow exception.

[0072] Among them, parameter constraint verification is implemented based on the rule engine, and parameter constraint rules (such as numerical ranges, regular expressions, business dictionaries) are loaded from the configuration center (such as Apollo, Nacos), and dynamically compiled into verification logic through the OGNL expression language. Parameter assignments are calculated in real time (for example, calculating array lengths, parsing JSON structures), and the results are compared with the constraint rules, such as checking whether the numerical parameters exceed the upper and lower bounds, and whether the strings meet the format requirements. Partitioned parallel verification is used for composite parameters (such as lists, maps), and distributed cache (such as Redis) is combined to optimize rule matching efficiency. When an overflow exception is detected, the circuit breaker mechanism is triggered (for example, pausing the task, issuing an alarm notification), and an exception report containing the illegal parameter value, constraint conditions, and impact range is generated.

[0073] As an example, while a big data processing script in a bank's funds clearing system was running, it intercepted the parameter flow for a cross-border remittance transaction and discovered that the "exchange_rate" parameter, used to calculate the handling fee, was unassigned. Investigation revealed that this parameter should have been obtained from the foreign exchange market interface, but the value had not been assigned correctly due to a failed interface call. Subsequently, based on the abnormal parameter "exchange_rate," the parameter transmission link was identified, and the transmission path was found to be "data collection module → handling fee calculation function → clearing result generation module." The anomaly was pinpointed at the line of code in the big data processing script responsible for calling the foreign exchange market interface. Finally, a validation rule was constructed based on the preset "exchange_rate" parameter constraints (the value range should be between 0.01 and 100). Since the parameter was null, clearly exceeding the constraint range, an overflow exception was determined for the parameter assignment. This immediately triggered an alarm and suspended the relevant business process, preventing financial losses caused by incorrect handling fee calculations.

[0074] As another example, during the execution of the hospital's test report generation big data processing script, the parameter transfer data flow was intercepted, and it was found that the "patient_age" parameter was in the default empty value state. This parameter is used to assist in the diagnosis of disease risk levels. Through tracing, the parameter transfer link was marked according to the abnormal parameter "patient_age", and it was clarified that its transfer link involved "patient information entry module → test data integration function → report generation module", and was associated with the code position in the script that read the age field from the patient information table. The data missing problem occurred. Then, according to the preset "patient_age" parameter constraint (the value range should be 0-150 years old), the verification rule was constructed. Since the parameter was empty and did not meet the constraint conditions, it was determined that an overflow exception occurred. The report generation process was immediately stopped, and the medical staff was notified to verify the patient information to prevent the diagnosis result from being biased due to the missing age parameter.

[0075] Through the above steps S41 to S43, by building a runtime parameter monitoring and verification system, the full-link quality control of the parameter transmission of the big data processing script is achieved, and abnormal parameters that are not assigned or default to null values ​​are dynamically identified to solve the hidden parameter problems caused by missing configuration, interface call failures, etc. Based on the parameter lineage map, the position of the abnormal code is accurately located, and a fast traceability channel from runtime exceptions to source code is established; the assignment compliance is verified in real time through parameter constraint rules, and overflow exceptions such as numerical out-of-bounds and format errors are intercepted in advance to prevent erroneous parameters from flowing into the downstream processing links. The three work together to transform traditional post-debugging into in-process prevention, significantly improve the stability of the big data system, reduce the data processing error rate caused by parameter anomalies, shorten the troubleshooting cycle, and ensure the accuracy and reliability of data in key business scenarios such as financial transaction settlement and medical diagnosis reports.

[0076] S5. When the second submission request of the big data scheduling task is received, whether the parameters and dependencies of the big data scheduling task are correct are detected according to the preset task scheduling specification to obtain a fourth test result.

[0077] Among them, in the big data scheduling system, when the second submission request is received, the automatic verification of parameters and dependencies is achieved by building a task scheduling verification engine. First, the configuration parser driven by annotations (for example, Apache Airflow's DAG file parser) reads the scheduling task metadata and extracts task parameters (for example, execution frequency, resource quota) and dependencies (for example, upstream task ID, data dependency path). Then, parameter verification is performed through a rule engine (for example, Drools), for example, checking whether the time parameter complies with the Cron expression specification and whether the memory allocation parameter exceeds the cluster limit. At the same time, a directed acyclic graph (DAG) model is constructed to represent the task dependency relationship, and a topological sorting algorithm is used to detect circular dependencies (for example, task A depends on task B, and task B depends on task A). The integrity of the data dependency is verified by a path search algorithm (for example, whether the input table exists and whether the partition is ready). For complex dependencies (for example, dynamically generated subtasks), a simulated execution technology is used to generate a virtual scheduling plan to verify the correctness of the dependency parsing. Finally, a fourth test result report containing parameter violations, dependency conflict paths, and repair suggestions is generated to ensure that the scheduling task complies with the platform specifications and business logic.

[0078] As an optional implementation, the process of detecting whether the parameters and dependencies of the big data scheduling task are correct according to the preset task scheduling specification in the above step S4 may specifically include the following steps S51 to S53.

[0079] S51. Traverse the parameter configuration items of the big data scheduling task and identify errors in filling in basic parameters and dependent parameters according to preset parameter verification rules.

[0080] For example, a parameter verification system is built based on the rule engine. Verification rules are configured for each parameter of the scheduling task through custom annotations (for example, @Required, @Range), and the parameter configuration items are traversed using the reflection mechanism. Format verification is performed on basic parameters (for example, execution time, resource allocation) (for example, the legality of Cron expressions, consistency of memory units), and logical verification is performed on dependent parameters (such as input table path, output partition) (for example, checking whether the path complies with the HDFS naming specification and whether the partition date exists). The responsibility chain model is used to perform null value checks, type conversions, and business rule verification in sequence. For example, for the "retry_times" parameter, first check whether it is a positive integer, and then check whether it exceeds the maximum retry limit of the cluster. Dynamic updates of verification rules are supported through the rule hot loading mechanism (for example, monitoring changes in the configuration center) to ensure parameter compliance in new business scenarios.

[0081] S52. Record the task node, parameter type, and verification failure reason corresponding to the error parameter.

[0082] Among them, an abnormal parameter metadata tracking system is built. During the parameter verification process, the current verification context (including task ID, parameter name, and verifier type) is stored through ThreadLocal. When an error is detected, the MDC (Mapped Diagnostic Context) technology is used to associate the error information with the specific task node. For complex parameter structures (for example, nested JSON), JSONPath expressions are used to accurately locate the error field, for example, "$.config.advanced.timeout". Combined with log link tracking (for example, Zipkin), the timestamp and call chain of the error are recorded, and temporary error snapshots are stored through distributed cache (for example, Redis) to support subsequent aggregate analysis. Finally, a multi-dimensional error matrix is ​​generated, including the task node to which the error parameter belongs, the parameter type (for example, basic parameter / dependent parameter), and the reason for the verification failure (for example, "value out of range", "format mismatch").

[0083] S53: Extract the execution logic, parameter configuration, and dependency links of the big data scheduling task in the dual-transmission environment as dual-transmission detection information, and determine whether there is inconsistency in the dual-transmission detection information.

[0084] Among them, a dual-environment consistency detection mechanism is implemented. The execution context of the scheduled tasks in the test environment and the production environment is captured through the AOP interceptor, and the execution logic (for example, SQL statements, Python scripts), parameter configuration (for example, parallelism, timeout period) and dependency links (for example, upstream task ID, data input source) are extracted. The hash algorithm is used to generate environmental feature fingerprints, and the feature values ​​of the test / production environments are compared to identify logical differences (for example, inconsistent SQL versions), parameter drift (for example, the memory configuration of the production environment is doubled), and dependency breaks (for example, the test environment uses a simulated data source). For dynamic dependencies (for example, partition paths generated based on dates), a time window sliding comparison is used to verify the consistency of dependency resolution at the same time point. When inconsistency is detected, the degree of difference is quantified through similarity calculation (for example, edit distance), and a dual-release detection report containing the difference points, impact range, and risk level is generated to provide a decision-making basis for grayscale release.

[0085] As an example, in a big data scheduling task in a bank's risk control system, when traversing parameter configuration items, it was discovered that the "risk_level_threshold" parameter was incorrectly entered as the string "medium." While the preset validation rules require this parameter to be a numeric type (for example, 0.5), this is a basic parameter entry error. Furthermore, the "historical_data_path" parameter was configured as "HDFS: / / clusterA / data," but the actual production cluster was "clusterB," causing the dependent parameter to point to the incorrect data source. The incorrect parameter was recorded: the task node to which "risk_level_threshold" belongs is "Risk Score Calculation," the parameter type is numeric, and the failure reason is "Type Mismatch." The node to which "historical_data_path" belongs is "Historical Data Import," the parameter type is path dependency, and the failure reason is "Cluster name inconsistent with production environment." During dual-release detection, we extracted the execution logic of the test and production environments and found that the SQL script used in the test environment contained a hard-coded condition "WHERE risk_level = 'medium'", while the production environment script used a parameterized query "WHERE risk_level > ?". This inconsistent execution logic could lead to deviations in risk assessment results, and the discrepancy was promptly marked as a high-risk item.

[0086] As another example, in a hospital laboratory data processing scheduling task, verification revealed that the parameter value for "sample_retention_days" was -7, violating the preset positive integer constraint and representing a basic parameter error. The parameter "lab_equipment_id" was configured as "EQP-2023," but the actual dependent equipment ID format should be "LAB-EQP-YYYY," resulting in an incorrect dependency parameter format. The task node for "sample_retention_days" was recorded as "Sample Expiration Cleanup," with a parameter type of time period and a failure reason of "Value less than the minimum value of 1." The node for "lab_equipment_id" was recorded as "Equipment Status Synchronization," with a parameter type of external resource dependency and a failure reason of "Format does not conform to business specifications." During dual-release testing, a comparison of the dependency links between the test and production environments revealed that the test environment relied on a simulated device status API, while the production environment relied on a real device interface. This dependency discrepancy could result in interface call failures in the production environment. This was marked as a medium-risk item, and a phased testing of the production environment was recommended.

[0087] Through the above steps S51 to S53, a scheduling task parameter and environment consistency verification system is constructed to achieve full-link guarantee of big data scheduling quality, automatically identify format errors, logical conflicts and environment mismatches in parameter configuration, and avoid task failures due to parameter anomalies. Establish multi-dimensional tracking capabilities for erroneous parameters, accurately locate problem nodes and failure causes, and shorten the troubleshooting cycle. Through the dual-environment detection mechanism, the execution logic differences, parameter drift and dependency breaks between the test and production environments are discovered in advance, ensuring the consistency and stability of scheduling tasks when deployed across environments. The three work together to transform traditional manual experience verification into systematic intelligent detection, significantly improving the reliability of the big data scheduling system, reducing the failure rate of the production environment, and ensuring the continuity and accuracy of key business scenarios such as financial transaction clearing and medical data processing.

[0088] S6. Generate an automated test result corresponding to the target big data platform using the first test result, the second test result, the third test result, and the fourth test result.

[0089] As an optional implementation, before the above step S6, the associated codes and configuration files of the scheduling tasks in the big data platform may be automatically tested, which may specifically include the following steps S61 to S63.

[0090] S61. Traverse the associated code of the big data scheduling task and the version identifier of the configuration file of the target big data platform.

[0091] For example, a version identification scanner is built based on metadata indexes: custom annotations (e.g., @VersionTag) are used to mark the version fields in associated code (e.g., Python scripts, SQL files) and configuration files (e.g., YAML, properties), and a file system listener (e.g., Apache Commons VFS) is used to traverse the task directory. For Git repositories, the Git command-line tool (e.g., git log --pretty=format:%H) is called to obtain the latest commit hash value as the code version identifier. For configuration centers (e.g., Nacos), the MD5 digest of configuration items is obtained through the API as the version identifier. A parallel scanning strategy is used to improve performance, and large projects are partitioned by module. Ultimately, a multidimensional index table containing file paths, version numbers, and timestamps is constructed, providing basic data for subsequent difference analysis.

[0092] S62. Identify whether there are any abnormal errors in the mutual coverage of associated codes and configuration files of different versions through hash verification and text difference comparison algorithms.

[0093] Among them, a multi-dimensional content comparison engine is implemented. For code files, the Google Diff-Match-Patch library is used to generate text difference patches (for example, diff files in a unified format). The AST parser (for example, ANTLR) is used to identify changes in grammatical structure, such as the addition or reduction of function parameters and the modification of SQL statement table names. For configuration files, JSONPath / XPath technology is used to extract key configuration items (for example, cluster address, parallelism parameters), calculate their hash values ​​and compare them with the baseline version, for example, to check whether the "spark.executor.memory" parameter is incorrectly overwritten. For binary files (for example, JAR packages), SHA-256 fingerprints are calculated for integrity verification. Machine learning models (for example, Siamese networks) are used to identify changes that are semantically similar but syntactically different, such as the case where the order of conditions in SQL queries is adjusted but logically equivalent. Finally, a difference report containing the change type, change location, and similarity score is generated.

[0094] S63. Monitor the change records and deployment logs of the big data scheduling tasks, and verify whether the submitted code and configuration changes are fully synchronized to the target operating environment based on the change records and deployment logs.

[0095] Among them, a change synchronization verification system is built to capture code / configuration change records and extract change sets (for example, commit ID, modified file list) by subscribing to Git Webhook or configuration center event bus. In combination with CI / CD tools (for example, Jenkins Pipeline), verification logic is injected into the deployment phase, and pre-checks are performed in the target environment (for example, querying HDFS file timestamps, checking container image versions) to verify whether the changes are fully synchronized. For incremental deployment scenarios, file system difference tools (for example, rsync--dry-run) are used to compare the source directory and the target directory to identify missed deployment files. Deployment logs are analyzed through a distributed log aggregation platform (for example, ELK Stack), matching the operators and timestamps in the change records, establishing a causal chain between changes and deployments, triggering automatic rollback or manual intervention processes for unsynchronized changes, and ensuring consistency between the production environment and the code base.

[0096] As an example, in a big data scheduling task for foreign exchange trading and clearing at a bank, the version identifiers of the foreign exchange rate calculation script and cluster configuration files were traversed. It was discovered that the Git commit hash of the Python script was abc123, while the version deployed last week was def456. At the same time, the MD5 digest of the cluster resource configuration file in the configuration center had changed. Using hash verification and text difference comparison algorithms, it was identified that the data conversion function in the exchange rate calculation script had been modified. The original script calculated using a fixed exchange rate, while the new version dynamically retrieved real-time exchange rates. However, the data source address in the configuration file continued to use the old version, potentially leading to data acquisition errors and the risk of overwriting each other's anomalies. Further monitoring of change records and deployment logs revealed that the developer had submitted a code update but had omitted to synchronize the API key in the configuration file. An immediate alert was issued to prevent foreign exchange trading and clearing errors and financial losses caused by incomplete configuration synchronization with the production environment.

[0097] As another example, in the hospital's patient medical record data analysis scheduling task, the version identifiers of the medical record data cleaning script and the Hadoop cluster configuration file were traversed, and it was found that the version number of the data cleaning script was updated from v1.0 to v2.0, and the version of Hadoop's YARN resource configuration file had also changed. Using the text difference comparison algorithm, it was found that the new version of the cleaning script added processing logic for rare disease diagnosis fields, but the memory allocation parameters of the YARN configuration file were not adjusted accordingly, and there was a risk of task failure due to insufficient resources. At the same time, monitoring the change records and deployment logs found that the developers submitted changes to the script and configuration, but due to network failure, only part of the Hadoop cluster configuration file in the production environment was updated. The task was intercepted in time and the operation and maintenance personnel were notified to redeploy to prevent errors in patient medical record analysis due to incomplete configuration, which would affect medical decision-making.

[0098] Through the above steps S61 to S63, by building a full-process control mechanism for code and configuration versions, accurate maintenance and risk prevention of the big data scheduling task environment are achieved, and the version identifications of the associated code and configuration files are automatically traversed to form a complete version asset map, providing a data basis for subsequent analysis. With the help of hash verification and text difference comparison algorithms, conflicts and coverage risks between different versions are accurately identified to avoid abnormal task operation due to incompatible code or configuration changes. Change records and deployment logs are monitored in real time to ensure that code and configuration changes are completely and accurately synchronized to the target operating environment, eliminating production failures caused by deployment omissions. The three work together to upgrade version management from passive investigation to active prevention, effectively reducing the system risks caused by version inconsistencies and configuration asynchrony in big data scheduling tasks, and improving the operational stability and reliability of core business scenarios such as financial transaction settlement and medical data processing.

[0099] The automated testing method for big data provided in the embodiment of the present application achieves full life cycle test coverage of big data processing through a full-process quality control system. At the data level, query statements are generated based on metadata, combined with timeliness models, storage verification and cross-table detection to monitor data timeliness, completeness and accuracy, identify problems such as data volume fluctuations and storage tilt, and ensure that key data such as finance and medical care meet business standards. At the code level, through static scanning and dynamic tracking, script syntax compliance is verified, variable usage and parameter passing anomalies are detected, and operational risks are moved forward to the development stage. At the scheduling and environment level, a task scheduling verification and dual-environment detection mechanism is constructed to verify parameter configuration, dependencies and cross-environment consistency to avoid task failure. At the version and deployment level, through version traversal, hash verification and change monitoring, code and configuration versions are managed, coverage conflicts are identified, and complete synchronization of changes is ensured. This method transforms the manual testing mode into an intelligent system, improves system stability and maintainability, reduces production failure rate, and provides technical support for high-demand scenarios.

[0100] Please continue reading Figure 5 , Figure 5 This is a schematic diagram of the system structure of the automated testing device for big data provided in the embodiment of the present application. Figure 5 As shown, the big data automated testing device 50 includes: a test query statement generation module 51 , a first test result module 52 , a second test result module 53 , a third test result module 54 , a fourth test result module 55 and an automated test result module 56 .

[0101] The test query statement generation module 51 is used to obtain the table structure and field information of the distributed storage according to the metadata of the target big data platform, receive the user's screening of the table structure and field information, and use the screening results to generate corresponding distributed data query statements.

[0102] The first test result module 52 is used to generate a scheduling task using the distributed data query statement, execute the scheduling task through a preset task scheduling platform, and determine whether there is a data anomaly based on the execution result to obtain a first test result.

[0103] The second test result module 53 is used to scan the big data processing script contained in the first submission request when receiving the first submission request of the target code, perform static analysis on the big data processing script, verify the compliance of the syntax and parameters, and obtain the second test result.

[0104] The third test result module 54 is used to monitor the big data task log during the running phase of the target code, dynamically parse the parameters of the big data processing script, track the parameter transfer logic in the distributed nodes, and obtain the third test result.

[0105] The fourth test result module 55 is used to detect whether the parameters and dependencies of the big data scheduling task are correct according to the preset task scheduling specification when receiving the second submission request of the big data scheduling task, and obtain a fourth test result.

[0106] The automated test result module 56 is configured to generate an automated test result corresponding to the target big data platform using the first test result, the second test result, the third test result, and the fourth test result.

[0107] As an optional implementation, the first test result module 52 is also specifically used to construct a task timeliness conduction model based on the big data link topology, configure the timeliness threshold of each big data node, collect the execution time of the scheduled task on each big data node in real time and compare the execution time with the timeliness threshold to obtain the big data node with abnormal task timeliness; calculate the fluctuation percentage of the number of rows in the big data table, verify whether the storage size, partition row number and partition storage size of the big data table are within the preset threshold range; count the null value rate and unique value ratio of the fields in the big data table, detect whether the null value rate and the unique value ratio are within the preset threshold range and verify the business validity of the enumeration value.

[0108] As an optional implementation, the first test result module 52 is also specifically used to perform horizontal consistency check of field values ​​on large data tables with associated relationships, and detect whether the matching degree of cross-table field data is within a preset threshold range; analyze the vertical fluctuation trend of time series data in the large data table and identify abnormal values ​​and mutation time data points based on the vertical fluctuation trend; compare the changes in the preset large data table indicator data before and after the data modification in the large data table and determine whether there is modified abnormal data based on the change results; verify whether the data processing logic in the large data table is correct according to the time boundaries set by the business rules in the business system; detect the processing accuracy of the data under the extreme value boundaries of the maximum and minimum values ​​of the large data table field value range.

[0109] As an optional implementation, the second test result module 53 is also specifically used to perform static scanning on the big data processing script to detect whether the big data processing script contains undefined variables, undefined first-trial variables, and time parameters set to fixed values; perform syntax detection on the big data processing script to check whether the truncation operation syntax in the script is correct, verify whether the syntax of the data definition statement in the script is correct, and detect whether the case conversion operation in the script is correct; obtain the association logic and association conditions of multiple data objects in the big data processing script, and detect whether there are association errors in the association logic and the association conditions.

[0110] As an optional implementation, the third test result module 54 is also specifically used to intercept the parameter transfer data flow of the big data processing script, traverse the parameter assignment status in the parameter transfer data flow, and identify the traversed abnormal parameters that are not assigned or have default null values; mark the parameter transfer link according to the traversed abnormal parameters, and associate the code position in the big data processing script; construct verification rules through preset parameter constraints, and determine whether the parameter assignment exceeds the constraint range according to the verification rules. If it exceeds the constraint range, it is determined that the corresponding parameter assignment has an overflow exception.

[0111] As an optional implementation, the fourth test result module 55 is also specifically used to traverse the parameter configuration items of the big data scheduling task, and identify errors in filling in basic parameters and dependent parameters according to preset parameter verification rules; record the task node, parameter type and verification failure reason corresponding to the erroneous parameter; extract the execution logic, parameter configuration and dependent link of the big data scheduling task in the dual-transmission environment as dual-transmission detection information, and determine whether there is inconsistency in the dual-transmission detection information.

[0112] As an optional implementation, the automated testing device 50 for big data also includes a version detection management module, which is specifically used to traverse the version identifiers of the associated code of the big data scheduling task and the configuration file of the target big data platform; identify whether there are abnormal errors that overlap each other in different versions of the associated code and the configuration file through hash verification and text difference comparison algorithms; monitor the change records and deployment logs of the big data scheduling task, and verify whether the submitted code and configuration changes are completely synchronized to the target operating environment based on the change records and deployment logs.

[0113] It should be noted that the aforementioned automated testing device for big data can execute the automated testing method for big data provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in the embodiments of the automated testing device for big data, please refer to the automated testing method for big data provided in the embodiments of this application.

[0114] Figure 6 This is a schematic diagram of the hardware structure of an electronic device that performs an automated testing method for big data provided in an embodiment of the present application. Figure 6 As shown, the electronic device 600 includes:

[0115] One or more processors 610 and memory 620, Figure 6 A processor 610 is taken as an example.

[0116] The processor 610 and the memory 620 may be connected via a bus or other means. Figure 6 The bus connection is taken as an example.

[0117] Memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the automated testing method for big data in the embodiments of the present application. Processor 610 executes the non-volatile software programs, instructions, and modules stored in memory 620 to execute various functional applications and data processing of the server, thereby implementing the automated testing method for big data in the above-mentioned method embodiment.

[0118] The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the automated testing device for big data, etc. In addition, the memory 620 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 620 may optionally include a memory remotely located relative to the processor 610, and these remote memories may be connected to the automated testing device for big data via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0119] The one or more modules are stored in the memory 620, and when executed by the one or more processors 610, perform the automated testing method for big data in any of the above method embodiments, for example, perform the above described Figure 2 Steps S1 to S6 of the method, Figure 3 Steps S21 to S23 of the method, Figure 4 Steps S24 to S28 of the method are implemented Figure 5 The functions of modules 51-56 in.

[0120] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.

[0121] An embodiment of the present application provides a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors, for example Figure 6 A processor 610 in the embodiment may enable the one or more processors to execute the automated testing method for big data in any of the above method embodiments, for example, executing the above described Figure 2 Steps S1 to S6 of the method, Figure 3 Steps S21 to S23 of the method, Figure 4 Steps S24 to S28 of the method are implemented Figure 5 The functions of modules 51-56 in.

[0122] The embodiment of the present application provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by the electronic device, the electronic device is enabled to perform the automated testing method for big data in any of the above method embodiments, for example, performing the above-described Figure 2 Steps S1 to S6 of the method, Figure 3 Steps S21 to S23 of the method, Figure 4 Steps S24 to S28 of the method are implemented Figure 5 The functions of modules 51-56 in.

[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course by hardware. Those skilled in the art can understand that all or part of the processes in the above embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Based on the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present application as described above. For the sake of simplicity, they are not provided in detail. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for automated testing of big data, characterized in that: include: Obtain the table structure and field information of distributed storage based on the metadata of the target big data platform, receive user screening of the table structure and field information, and use the screening results to generate corresponding distributed data query statements; Generate a scheduling task using the distributed data query statement, execute the scheduling task through a preset task scheduling platform, and determine whether there is a data anomaly based on the execution result to obtain a first test result; When a first submission request for the target code is received, scanning a big data processing script included in the first submission request, performing static parsing on the big data processing script, verifying syntax and parameter compliance, and obtaining a second test result; During the execution phase of the target code, the big data task log is monitored, the parameters of the big data processing script are dynamically parsed, and the parameter transmission logic in the distributed nodes is tracked to obtain a third test result; When receiving the second submission request for the big data scheduling task, detecting whether the parameters and dependencies of the big data scheduling task are correct according to the preset task scheduling specification, and obtaining a fourth test result; The first test result, the second test result, the third test result and the fourth test result are used to generate an automated test result corresponding to the target big data platform.

2. The automated testing method for big data according to claim 1, characterized in that: After the step of executing the scheduling task through the preset task scheduling platform and determining whether there is data anomaly based on the execution result, the method further includes: A task timeliness transmission model is constructed based on the big data link topology. A timeliness threshold is configured for each big data node. The execution time of the scheduled task on each big data node is collected in real time. The execution time is compared with the timeliness threshold to obtain the big data node with abnormal task timeliness. Calculate the row number fluctuation percentage of the big data table and verify whether the storage size, partition row number, and partition storage size of the big data table are within the preset threshold range; Count the null value rate and unique value ratio of fields in the big data table, detect whether the null value rate and the unique value ratio are within the preset threshold range, and verify the business validity of the enumeration value.

3. The automated testing method for big data according to claim 1, characterized in that: After the step of executing the scheduling task through the preset task scheduling platform and determining whether there is data anomaly based on the execution result, the method further includes: Perform horizontal consistency check of field values ​​on large data tables with association relationships to detect whether the matching degree of cross-table field data is within the preset threshold range; Analyze the longitudinal fluctuation trend of time series data in a large data table, and identify outliers and mutation time data points based on the longitudinal fluctuation trend; Compare the changes in the preset big data table indicator data before and after the data in the big data table is modified, and determine whether there is any abnormal modification data based on the change results; Verify the correctness of data processing logic in large data tables based on the time boundaries set by business rules in the business system; According to the extreme value boundaries of the maximum and minimum values ​​of the value range of the large data table field, the processing accuracy of the data under the extreme value boundaries is detected.

4. The automated testing method for big data according to claim 1, characterized in that: The step of statically parsing the big data processing script and verifying the compliance of syntax and parameters includes: Performing a static scan on the big data processing script to detect whether the big data processing script contains undefined variables, undefined variables that have been tried out first, and time parameters that are set to fixed values; Performing syntax testing on the big data processing script to check whether the truncation operation syntax in the script is correct, verifying whether the syntax of the data definition statement in the script is correct, and detecting whether the case conversion operation in the script is correct; Obtain the association logic and association conditions of multiple data objects in the big data processing script, and detect whether there are association errors in the association logic and the association conditions.

5. The automated testing method for big data according to claim 1, characterized in that: The step of dynamically parsing the parameters of the big data processing script and tracking the parameter transfer logic in the distributed nodes includes: Intercepting the parameter transfer data stream of the big data processing script, traversing the parameter assignment status in the parameter transfer data stream, and identifying the traversed abnormal parameters that are not assigned or have default null values; Marking the parameter transmission link according to the traversed abnormal parameter and associating it with the code position in the big data processing script; Verification rules are constructed through preset parameter constraints, and whether the parameter assignment exceeds the constraint range is determined according to the verification rules. If it exceeds the constraint range, it is determined that the corresponding parameter assignment has an overflow exception.

6. The automated testing method for big data according to claim 1, characterized in that: The step of detecting whether the parameters and dependencies of the big data scheduling task are correct according to the preset task scheduling specification includes: Traverse the parameter configuration items of the big data scheduling task, and identify basic parameter filling errors and dependent parameter filling errors according to preset parameter verification rules; Record the task node, parameter type, and verification failure reason corresponding to the error parameter; The execution logic, parameter configuration and dependent links of the big data scheduling task in the dual-transmission environment are extracted as dual-transmission detection information, and whether there is inconsistency in the dual-transmission detection information is determined.

7. The automated testing method for big data according to claim 1, characterized in that: Before the step of using the first test result, the second test result, the third test result, and the fourth test result to generate the automated test result corresponding to the target big data platform, the method further includes: Traversing the associated code of the big data scheduling task and the version identifier of the configuration file of the target big data platform; Identify whether there are any abnormal errors in which different versions of the associated code and the configuration file overlap each other through hash verification and text difference comparison algorithms; Monitor the change records and deployment logs of the big data scheduling tasks, and verify whether the submitted code and configuration changes are fully synchronized to the target operating environment based on the change records and deployment logs.

8. An automated testing device for big data, characterized in that: include: The test query statement generation module is used to obtain the table structure and field information of the distributed storage based on the metadata of the target big data platform, receive the user's screening of the table structure and field information, and use the screening results to generate the corresponding distributed data query statement; A first test result module is configured to generate a scheduling task using the distributed data query statement, execute the scheduling task through a preset task scheduling platform, and determine whether there is a data anomaly based on the execution result to obtain a first test result; A second test result module is configured to, upon receiving a first submission request for the target code, scan a big data processing script included in the first submission request, perform static parsing on the big data processing script, verify syntax and parameter compliance, and obtain a second test result; A third test result module is used to monitor the big data task log during the running phase of the target code, dynamically parse the parameters of the big data processing script, track the parameter transmission logic in the distributed nodes, and obtain a third test result; a fourth test result module, configured to, upon receiving the second submission request for the big data scheduling task, detect whether the parameters and dependencies of the big data scheduling task are correct according to a preset task scheduling specification, and obtain a fourth test result; An automated test result module is used to generate an automated test result corresponding to the target big data platform using the first test result, the second test result, the third test result and the fourth test result.

9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the automated testing method for big data according to any one of claims 1 to 7.

10. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by an electronic device, the electronic device executes the automated testing method for big data according to any one of claims 1 to 7.

Citation Information

Cited By

  • Data verification rule automatic generation and execution method and system based on large model

    CN121167246A

  • Industrial internet abnormal permission behavior detection method and system

    CN121356838A