Batch processing job automatic debugging method, device and equipment and storage medium

By replacing the target table in the debug copy with a temporary target table and deploying the batch script with the minimum privileges of the temporary account, the problems of user interference and security risks in Hadoop batch job debugging are solved, and a safe and stable debugging process is achieved.

CN121301170APending Publication Date: 2026-01-09CHINA MERCHANTS BANK
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511523437.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

During the debugging of Hadoop batch processing jobs, user debugging operations are prone to mutual interference, leading to inaccurate debugging results, damaging the stability of the test environment, and the existing permission management method cannot be dynamically adjusted, causing security risks and permission overlap.

Method used

By obtaining a debug copy, replacing the target table with a temporary target table, deploying batch scripts with the necessary permissions of the temporary account, and generating batch job debugging results, the system achieves on-demand allocation of permissions and operation sandbox isolation, ensuring the security and stability of the debugging process.

Benefits of technology

It significantly improves the security and stability of batch job debugging, prevents mutual interference between debugging operations, realistically simulates the permission status of the production environment, and ensures the reliability and accuracy of the debugging process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301170A_ABST
    Figure CN121301170A_ABST
Patent Text Reader

Abstract

The invention discloses a batch processing job automatic debugging method and device, equipment and a storage medium, and relates to the technical field of data processing and debugging verification, and the method comprises the following steps: obtaining a debugging copy; replacing a target table in the debugging copy, and determining debugging copy information; deploying a corresponding batch processing script based on the debugging copy information by using the allocated temporary account necessary permission, and determining a target batch processing operation script; and performing batch processing job debugging based on the target batch processing job script to generate a batch processing job debugging result. According to the method, the target table in the debugging copy is replaced, the batch processing script is deployed according to the distributed temporary account minimum permission, batch processing operation debugging is carried out on the target batch processing operation script, the batch processing operation debugging result is generated, and the temporary account only having the minimum necessary permission is dynamically generated for each time of debugging. And a special temporary table is automatically created as a safe output environment, so that authority distribution according to needs and operation sandbox isolation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data processing and debugging verification technology, and in particular to automatic debugging methods, apparatus, equipment and storage media for batch processing jobs. Background Technology

[0002] During the debugging process of Hadoop batch jobs, the debugging behavior of different users can easily interfere with each other. One user's debugging operation will affect the debugging of other users, resulting in inaccurate debugging results and even damaging the stability of the test environment. This makes it impossible for the debugging process to accurately reflect the actual operation in the production environment, reducing the reliability and effectiveness of the debugging results.

[0003] Currently, the existing practice is to use effective static strategies for access control, i.e., using personal cluster accounts for debugging. However, this management method cannot dynamically adjust the scope of permissions according to specific debugging tasks, leading to over-generalization of permissions. That is, users are granted permissions beyond what is actually needed for debugging. This over-generalization of permissions not only increases operational redundancy but also raises security risks. Users may inadvertently access or modify data and resources they should not have access to, resulting in overlapping permissions. This fails to isolate operational interference and makes it difficult to truly simulate production permissions, increasing security and stability risks. Therefore, how to perform automated debugging of intelligent batch processing jobs more securely has become an urgent problem to be solved.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this application is to provide a method, apparatus, device, and storage medium for automatically debugging batch processing jobs, aiming to solve the technical problem of how to perform intelligent batch processing job automatic debugging more safely.

[0006] To achieve the above objectives, this application proposes an automatic debugging method for batch processing jobs, the method comprising: Obtain a debug copy, replace the target table in the debug copy, and determine the debug copy information; Based on the debug copy information, deploy the corresponding batch script with the necessary permissions of the assigned temporary account to determine the target batch job script; Batch processing job debugging is performed based on the target batch processing job script, and batch processing job debugging results are generated.

[0007] In one embodiment, the step of replacing the target table in the debug copy to determine the debug copy information includes: Retrieve source table permission information; The permissions of the corresponding personal cluster account are verified using the source table permission information, and the permission verification result is determined. Based on the permission verification result, the target table in the debug copy is replaced to obtain debug copy information.

[0008] In one embodiment, the step of replacing the target table in the debug copy based on the permission verification result to obtain debug copy information includes: A corresponding temporary target table is generated based on the permission verification result; Based on the debug copy, different script types are identified, and script type information is determined. Based on the script type information, the corresponding replacement method is selected to replace the target table in the debug copy with the temporary target table, thereby obtaining the debug copy information.

[0009] In one embodiment, after the step of deploying the corresponding batch script with the necessary permissions of the allocated temporary account based on the debug copy information to determine the target batch job script, the method further includes: The syntax of the instruction statements in the target batch processing job script is validated, and the validation result is determined. Based on the verification results, the instruction statements in the target batch processing job script are input into the asynchronous large model to verify logical defects, thereby obtaining the verification instruction statements of the target batch processing job script.

[0010] In one embodiment, the step of validating the syntax of the instruction statements in the target batch processing job script and determining the validation result includes: The syntax of the instruction statements in the target batch processing job script is broken down into independent job statements; The operation statements are classified according to different types and purposes to determine the classification instruction statements; Based on the classification instruction statement, the corresponding verification strategy is executed to verify the corresponding syntax and generate execution record information; A verification result is generated based on the execution record information.

[0011] In one embodiment, the step of inputting the instruction statements in the target batch processing job script into the asynchronous large model to verify logical defects based on the verification result, and obtaining the verification instruction statements of the target batch processing job script, includes: Based on the verification results, the instruction statements in the target batch processing job script are input into the business orchestration model in the asynchronous large model to identify the statement category and statement logic, and determine the identification results. Based on the identification results, the instruction statements in the target batch processing job script are input into the generation and optimization model in the asynchronous large model for adjustment, thereby obtaining the verification instruction statements of the target batch processing job script.

[0012] In one embodiment, the step of debugging the batch processing job based on the target batch processing job script and generating the batch processing job debugging result includes: Obtain the production target table; Execute batch processing job debugging on the target batch processing job script to generate a debugging target table; Based on the debugging target table and the production target table, a corresponding comparison and verification report is generated to obtain the batch processing job debugging results.

[0013] Furthermore, to achieve the above objectives, this application also proposes an automatic batch processing job debugging device, which includes: The acquisition module is used to acquire a debug copy, replace the target table in the debug copy, and determine the debug copy information; The processing module is used to deploy the corresponding batch processing script based on the debug copy information and with the necessary permissions of the allocated temporary account, and to determine the target batch processing job script. The execution module is used to debug the batch processing job based on the target batch processing job script and generate the batch processing job debugging results.

[0014] In addition, to achieve the above objectives, this application also proposes an automatic batch job debugging device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the automatic batch job debugging method described above.

[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the batch job automatic debugging method described above.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: This embodiment proposes an automatic batch job debugging method. The method involves obtaining a debug copy, replacing the target table in the debug copy to determine debug copy information, deploying the corresponding batch script with the necessary permissions of an assigned temporary account based on the debug copy information, determining the target batch job script, and debugging the batch job based on the target batch job script to generate batch job debugging results. This application significantly improves the security and stability of the debugging process by replacing the target table in the debug copy and deploying the batch script with the minimum permissions of the assigned temporary account, debugging the target batch job script, and generating batch job debugging results. It dynamically generates temporary accounts with only the minimum necessary permissions for each debugging session and automatically creates a dedicated temporary table as a secure output environment, achieving on-demand permission allocation and operation sandbox isolation, effectively preventing mutual interference between debugging operations, realistically simulating the permission status in a production environment, and ensuring the reliability and accuracy of the debugging process. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the automatic debugging method for batch processing jobs in this application. Figure 2 This is a flowchart of the intelligent table name replacement process for the automatic debugging method of batch processing jobs in this application; Figure 3 This is a schematic diagram of a key segment of the automatic debugging method model for batch processing jobs in this application; Figure 4 A schematic diagram comparing the debugging target table and the production target table of the automatic debugging method for batch processing in this application; Figure 5 This is a schematic diagram of the comparative verification report of the automatic debugging method for batch processing jobs in this application; Figure 6 This is a flowchart illustrating Embodiment 2 of the batch processing job automatic debugging method of this application; Figure 7 This is a flowchart of the syntax verification process for the automatic debugging method of batch processing jobs in this application; Figure 8 This is a flowchart illustrating the logic verification of the automatic debugging method for batch processing jobs in this application. Figure 9 A simplified flowchart illustrating the automatic debugging method for batch processing jobs provided in this application embodiment; Figure 10 This is a schematic diagram of the module structure of the automatic debugging device for batch processing jobs according to an embodiment of this application; Figure 11 This is a schematic diagram of the hardware operating environment involved in the automatic debugging method for batch processing jobs in the embodiments of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] The main solution of this application embodiment is: to obtain a debug copy, replace the target table in the debug copy, and determine the debug copy information; to deploy the corresponding batch processing script with the necessary permissions of the allocated temporary account based on the debug copy information, and determine the target batch processing job script; to debug the batch processing job based on the target batch processing job script, and generate the batch processing job debugging result.

[0024] In this embodiment, for ease of description, the following description will focus on the automatic debugging device for identifying batch processing jobs.

[0025] Because existing technology cannot dynamically adjust the scope of permissions according to specific debugging tasks, permissions are over-generalized, meaning that users are granted permissions that exceed the actual debugging needs. This over-generalization of permissions not only increases the redundancy of operations but also raises security risks. Users may inadvertently access or modify data and resources that they should not touch, resulting in overlapping permissions. This makes it impossible to isolate operational interference and difficult to truly simulate production permissions, thus increasing security and stability risks.

[0026] This application provides a solution to obtain a debug copy, replace the target table in the debug copy, and determine the debug copy information; based on the debug copy information, deploy the corresponding batch processing script with the necessary permissions of the assigned temporary account to determine the target batch processing job script; and perform batch processing job debugging based on the target batch processing job script to generate batch processing job debugging results.

[0027] As can be seen from the above embodiments, this application significantly improves the security and stability of the debugging process by replacing the target table in the debug copy and deploying the batch script with the minimum permissions of the assigned temporary account, performing batch job debugging on the target batch job script, and generating batch job debugging results. It dynamically generates temporary accounts with only the minimum necessary permissions for each debugging session and automatically creates a dedicated temporary table as a secure output environment, realizing on-demand permission allocation and operation sandbox isolation, effectively preventing mutual interference between debugging operations, truly simulating the permission status in the production environment, and ensuring the reliability and accuracy of the debugging process.

[0028] Based on this, embodiments of this application provide an automatic debugging method for batch processing jobs, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the batch processing job automatic debugging method of this application.

[0029] In this embodiment, the automatic debugging method for batch processing jobs includes steps S10 to S40: Step S10: Obtain a debug copy, replace the target table in the debug copy, and determine the debug copy information; It should be noted that the debug copy information is a complete copy of the job configuration and content that can be directly used for the automated debugging process, formed after the original job script has been securely isolated and key content replaced.

[0030] Understandably, the system will automatically map and obtain the personal cluster account of the currently logged-in user through the identity authentication service tool integrated into the platform when the debugging process starts, to prevent unauthorized operations. The identity authentication service tool can automatically map the personal cluster account of the currently logged-in user on the platform without the need to manually enter credential information. At the same time, the platform uses scanning login technology to effectively ensure the authenticity of the user's identity, thereby avoiding unauthorized operations. Furthermore, it uses the batch processing job script in the original production environment to create a completely consistent debug copy through dynamic snapshots, ensuring zero intrusion and zero pollution of the original script.

[0031] In a specific embodiment, source table permission information is obtained; the permissions of the corresponding personal cluster account are verified using the source table permission information to determine the permission verification result; a corresponding temporary target table is generated based on the permission verification result. That is, the system will check the read and write permissions of the source table involved by the personal cluster account, obtain the permission verification result, and generate a corresponding temporary target table based on the permission verification result. If there is no permission, the process ends directly; if there is permission, a temporary target table is generated; if generation fails, the process ends directly; if generation succeeds, a debug copy is generated and the table name is replaced. A temporary target table is generated before each debugging session, and the production target table in the script is replaced with the temporary target table. The temporary target table has a unique identifier, and its table structure is completely consistent with the target table in the production environment, thereby achieving isolation from production data.

[0032] Based on the debug copy, different script types are identified, and script type information is determined. Specifically, based on the syntactic characteristics of various jobs, the debug copy is used to identify different script types, such as Hive jobs and Spark jobs. Hive jobs, since all jobs use HQL statements and have strong syntactic regularity, can be directly processed using regular expression matching. Regular expression matching is a technique for accurately searching and locating target table names in text based on predefined rule strings. For example, using regular expressions to accurately search and locate SQL scripts in text, the system pre-configures pattern rules for each script type that can identify the location of table names. When processing a script, the system scans the code line by line to automatically find all table name strings that match this pattern and need to be replaced. For example, it can accurately identify FROM prod_db.sales_fact or INSERT INTO. The tool avoids mismatching seemingly similar but actually aliased or string-based text in statements like `user_profile`, which contain table names such as `sales_fact` and `user_profile`. The Spark job addresses the issue of insufficient regular expression coverage caused by syntactic and user habit differences. The tool employs a hybrid approach of large model identification, regular expression matching, and preview modification to achieve accurate code replacement. Large model identification utilizes multiple models to identify the SQL script, while preview modification modifies the extracted SQL script after identification. Furthermore, a feedback-based continuous learning mechanism is introduced to continuously improve system performance by optimizing model coverage and accuracy.

[0033] Based on the script type information, the corresponding replacement method is selected to replace the target table in the debug copy with the temporary target table, thus obtaining the debug copy information, i.e., as follows. Figure 2 As shown, Figure 2This document presents a flowchart of the intelligent table name replacement process for the batch processing job automatic debugging method in this application. After identifying the script type information, such as Hive jobs or Spark jobs, regular expressions are used directly for replacement in Hive jobs. However, in Spark jobs, table names are typically defined using variables, and subsequent SQL statements use variable representations of the table name. The table name, the replacement table name, and the job content are passed to the model. The model automatically identifies the variable definitions related to the table name in the job and generates corresponding replacement variables based on the given replacement table name. The model returns an array of table name-related variables and replaced variables. The model prompt design uses a three-order structure that has been empirically verified through multiple rounds of testing: background information, matching rules, and few-shot COT (Chain-of-Thought). Figure 3 As shown, Figure 3 This diagram illustrates key segments of the batch processing job automatic debugging method model in this application. It comprises three core levels: task background description to build the model's cognitive framework, source table / table matching rules to constrain model behavior, and COT examples that integrate inference paths to enhance generalization ability. Among them, Few-shot COT is a large model prompting technology that combines few-shot learning and thought chain reasoning. Before presenting the large model with the actual problem to be solved, it provides several examples containing complete step-by-step reasoning processes. These examples clearly present the intermediate logical steps to arrive at the answer, i.e., the thought chain. In this way, the large model is guided to imitate the reasoning patterns in the examples. When faced with new problems, it can first perform step-by-step logical analysis and then generate the final answer, thereby significantly improving its reasoning accuracy and interpretability in complex tasks. At the same time, the tool records error cases during use and regularly reviews and analyzes these cases to identify them. It promptly adds uncovered scenarios to the Prompt rule base and case base to achieve continuous model optimization. Regular expressions are used for replacement operations. After the replacement is completed, users can preview and edit the replaced job to prevent replacement failure. If the user modifies the replaced job, it will be recorded to support subsequent model analysis and review.

[0034] In one feasible implementation, step S10 may include steps A11 to A13: Step A11: Obtain source table permission information; It should be noted that the source table permission information is a set of access permission rules for all original data tables that the batch processing job script needs to read or reference during execution.

[0035] It is understood that the source table permission information may include table identity, permission type, permission subject, and permission status. The table identity is the full name of the source table, such as `database_name.table_name`. The permission type refers to the required operation permissions for the corresponding data in the source table. For example, for batch processing jobs in the debugging phase, such as Hive SQL or Spark SQL, the primary permission type is SELECT permission, because the job needs to read data from these tables for calculation and transformation. In other cases, if the job logic involves writing to temporary tables or update operations, INSERT, UPDATE, and other permissions are also required. The permission subject is the entity accessing the table, which can be a personal cluster account. The permission status indicates whether the current account has been granted the necessary permissions.

[0036] Step A12: Verify the permissions of the corresponding personal cluster account using the source table permission information, and determine the permission verification result; It should be noted that the permission verification result is a conclusion generated by the system after performing an authorization check on the personal cluster account based on the permission information in the source table.

[0037] Understandably, the permission verification result indicates whether the personal cluster account has all the necessary operation permissions for all source tables in the job script. If the permission verification result is successful, it means that the personal cluster account has the necessary permissions. If the permission verification result is unsuccessful, it means that the personal cluster account lacks the necessary permissions. If the permission verification result is unsuccessful, it will clearly indicate which source tables the personal cluster account lacks the required permissions for. For example, it will return a list showing that user 'user_dev' does not have SELECT permission on table 'prod_db.sensitive_table'.

[0038] Step A13: Replace the target table in the debug copy based on the permission verification result to obtain debug copy information.

[0039] Understandably, the system will only generate a temporary target table to replace the target table in the debug copy if the permission verification result is passed. This ensures that any debugging operation must undergo strict data access permission verification, eliminating the possibility of unauthorized operations from the root. If the status is not passed, the entire process will terminate, and a temporary target table will not be generated to replace the target table in the debug copy.

[0040] In one feasible implementation, step A13 may include steps B11 to B13: Step B11: Generate a corresponding temporary target table based on the permission verification result; It should be noted that the temporary target table is a data table whose structure is consistent with the production environment target table, dynamically generated for each debugging task.

[0041] It is understood that the temporary target table is automatically created after the user passes the authorization verification and replaces the production target table in the script during the debugging process. This securely isolates all data writing operations in a temporary environment, effectively preventing direct modification and contamination of production data by debugging operations. This achieves operation sandbox isolation. Furthermore, by configuring necessary permissions for the table only for the temporary account, the temporary account is a unique account dynamically and automatically generated by the system when each debugging task starts by calling the platform's identity management API. Its name is automatically assigned according to predefined rules, and the lifecycle of the account is completely bound to a single debugging task. The system automatically destroys it after the task ends, fundamentally eliminating the risk of residual permissions and ensuring that permissions are allocated as needed. The minimum permission is precisely granted to the temporary account, granting only the most basic permissions necessary to complete the current debugging task. That is, it only has read permissions for the source table that the job depends on and write permissions for the dynamically generated temporary target table. By strictly limiting the data operation scope of the account, it is ensured that it cannot perform any operations beyond the necessity of debugging. This achieves secure sandbox isolation of operations at the permission level, ensuring data security while realistically simulating the operating state of the production environment.

[0042] Step B12: Identify different script types based on the debug copy and determine script type information; It should be noted that the script type information is an identifier of the specific category to which the identified script belongs.

[0043] Understandably, the system can analyze the syntax, keywords, and code patterns of the debug copy—for example, to determine whether it is a pure Hive QL statement or a Spark SQL script containing Python variable definitions—to determine the specific type of the script, such as Hive SQL or PySpark. This allows the system to decide on the corresponding intelligent table name replacement strategy and algorithm. For example, it can use regular expression matching and replacement for grammatically correct Hive jobs, while launching a large model for semantic recognition and variable replacement for more flexible Spark jobs, thereby achieving accurate and safe automated debugging.

[0044] Step B13: Based on the script type information, select the corresponding replacement method to replace the target table in the debug copy with the temporary target table to obtain the debug copy information.

[0045] It is understood that the debug copy information may include the modified job script content, script type, and unique identifier of the copy. The modified job script content is the new script after the temporary target table replacement operation has been completed. For example, all references to the production target table prod_table in the script have been replaced with the corresponding temporary target table prod_table_temp. The script type identifies whether the copy belongs to Hive SQL, Spark, or other types of jobs. The unique identifier of the copy is an identifier for multiple debug sessions to ensure that multiple debug sessions do not interfere with each other.

[0046] Step S20: Based on the debug copy information, deploy the corresponding batch script with the necessary permissions of the allocated temporary account to determine the target batch job script; It should be noted that the target batch processing job script is an executable batch processing script used to execute the corresponding batch processing job.

[0047] It is understood that all production environment sensitive information in the target batch processing job script has been securely replaced with temporary resources, such as production target table names, and configured to run under a temporary account with the minimum necessary permissions. Therefore, the target batch processing job script can be directly submitted to the rule engine for secure execution and generate the final job script for debugging results.

[0048] In a specific embodiment, after obtaining the debug copy information, the system enters the deployment phase. Utilizing a temporary account specifically allocated for this debugging task and its strictly limited necessary permissions, such as only SELECT permissions on the relevant source table and INSERT permissions on the temporary target table, the system securely publishes the debug copy script to the designated Hadoop / YARN computing cluster through an integrated deployment tool, such as Apache Airflow or a custom scheduler. This involves uploading the script file to the temporary account's dedicated HDFS directory, configuring the corresponding execution engine parameters and environment variables according to the script type, formally registering the job task to the cluster scheduler, and binding the temporary account's Kerberos credentials to ensure that all operations are executed within the authorized scope, thereby obtaining the target batch processing job script.

[0049] Step S30: Debug the batch processing job based on the target batch processing job script and generate the batch processing job debugging result.

[0050] It should be noted that the batch processing job debugging results are a comprehensive report generated by the system to fully evaluate the running status and data accuracy of this debugging job.

[0051] It is understood that the batch job debugging is achieved by automatically comparing the data differences between the debugging target table produced in this run and the production target table of the historical correct benchmark. If the key target fields in the two tables, such as SQL statements, do not match, the difference in the number of records is accumulated once and presented to the user in a clear visual form, such as a front-end report or notification message.

[0052] Additionally, it should be noted that the rule engine automatically schedules jobs to run in the computing cluster based on preset business rules, and strictly binds the minimum permission credentials of temporary accounts during this process to ensure that all data access operations are within a security sandbox. The verification instructions in the target batch processing job script are code snippets embedded in the target batch processing job script, specifically used for automatically verifying data quality and the correctness of business logic. They are not part of the original business logic, but are automatically injected by the system after analysis by the rule engine or large model. For example, after the script is executed, SQL for comparing data volume, consistency checks of key indicators, or verification of business rule constraints are inserted to generate the final debugging report, thereby realizing intelligent verification of the job output results. The batch processing job is a non-interactive, pre-arranged computing task used to automatically process massive amounts of data at once, triggered by fixed periods or events, and runs in the background without manual intervention.

[0053] In a specific embodiment, the production target table is obtained; the batch processing job script is executed for batch processing job debugging, generating a debugging target table. Specifically, based on the user-selected configuration, a rule engine automatically generates SQL statements comparing the record counts of the debugged target table and the production target table. The rule engine can also be replaced with a Hive-based implementation to adapt to different infrastructure environments. A corresponding comparison and verification report is generated based on the debugging target table and the production target table to obtain the batch processing job debugging results, i.e., as shown below. Figure 4 As shown, Figure 4 This diagram illustrates the comparison between the target table and the production target table in the automatic batch processing job debugging method of this application. It supports multi-dimensional verification. After the two tables are linked by primary key, records with inconsistent comparison values ​​are filtered by field. This includes comparisons of the total number of records at the table level and primary key relationships at the field level, thus comprehensively ensuring data integrity and accuracy. To improve verification efficiency, the high-performance Hetu computing engine can be used to execute these verification SQL statements, greatly shortening the verification time. After verification, the system automatically generates a structured debugging report, yielding the batch processing job debugging results, such as... Figure 5 As shown, Figure 5This is a schematic diagram of the comparison and verification report of the automatic debugging method for batch processing jobs in this application. The diagram clearly shows indicators such as the number of rows and percentage of inconsistent data. For example, it clearly indicates whether the number of rows in the production table and the debugging target table are consistent in the table-level verification. In the diagram, the production table has 13 rows, and the comparison debugging table also has 13 rows. At this time, the number of rows is consistent. In the field-level verification, the details of inconsistent record data for the corresponding field are listed item by item by primary key. In the diagram, there are 10 inconsistent records in the customer ID field, accounting for 0.77%. It also supports viewing the sample values ​​of specific difference data, thereby providing feedback and notification to the user.

[0054] In one feasible implementation, step S30 may include steps C11-C13: Step C11: Obtain the production target table; It should be noted that the production target table is a verified data table generated after batch processing operations have run normally in the production environment and is used to support actual business operations.

[0055] Understandably, when debugging new or modified batch jobs is required, the system will obtain the production target table that is currently running well in the production environment and serves as the correct output result of the job. This table will be called by the system during the automated verification process to accurately assess whether the output result of the debugging job meets the production readiness standard.

[0056] Step C12: Execute batch processing job debugging on the target batch processing job script to generate a debugging target table; It should be noted that the debugging target table is a temporary data result table generated after executing the target batch processing job script in the sandbox debugging environment.

[0057] It is understood that the debug target table is generated by replacing the production table name in the job script with a temporary table name and writing it into an isolated space with strict temporary account permissions. Its table structure is completely consistent with the production target table, but the data content is calculated based on the debug copy and test data.

[0058] Step C13: Generate a corresponding comparison and verification report based on the debugging target table and the production target table to obtain the batch processing job debugging results.

[0059] Understandably, during the verification phase, the system will automatically perform a comprehensive comparison of the debugging target table with the production target table as a benchmark, both in terms of record count and field level. By analyzing the differences between the two, a final debugging report will be generated, thereby objectively verifying whether the debugging of this operation has met expectations in terms of logical correctness and data accuracy.

[0060] This embodiment proposes an automatic batch job debugging method, which involves obtaining a debug copy; replacing the target table in the debug copy to determine the debug copy information; deploying the corresponding batch script with the necessary permissions of an allocated temporary account based on the debug copy information to determine the target batch job script; and debugging the batch job based on the target batch job script to generate batch job debugging results. This method solves the technical problem of how to perform intelligent automatic batch job debugging more securely. Compared with existing technologies, this application utilizes personal cluster account permissions to automatically replace the production target table in the copy with a temporary table to build a security sandbox, and dynamically allocates temporary accounts with only minimum necessary permissions to deploy scripts for each debugging session. It uses a rule engine to drive job execution and result verification, achieving automation and strict isolation in the debugging process. While ensuring data security, it generates highly reliable debugging results, significantly improving the security and efficiency of batch job debugging.

[0061] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment can be referred to the above description, and will not be repeated hereafter.

[0062] In this embodiment, refer to Figure 6 , Figure 6 This is a flowchart illustrating the second embodiment of the batch processing job automatic debugging method of this application. Step S20 specifically includes steps S21 to S22: Step S21: Verify the syntax of the instruction statements in the target batch processing job script and determine the verification result; It should be noted that the verification result is generated after automatically performing syntax checks and logical analysis on the instruction statements in the target batch processing job script.

[0063] In a specific embodiment, the syntax of the instruction statements in the target batch processing job script is broken down into independent job statements. That is, the tool automatically executes the `explain` command to scan all statements in the job for syntax checks. If syntax problems are found, a detailed list of all problematic SQL statements and their error messages is provided to the user. If no problems are found, the verification passes and subsequent operations can continue. The job statements are categorized according to different types and purposes, classifying the instruction statements into categories. Based on these categorized instruction statements, corresponding verification strategies are executed to verify the corresponding syntax, generating execution record information, i.e., as shown below. Figure 7 As shown, Figure 7The flowchart below shows the syntax validation process for the automatic debugging method of batch jobs in this application. The system reads the complete job script and breaks it down into independent SQL statements. Then, based on the type and purpose of each statement, it is intelligently categorized into one of the following three branches: temporary table statements, physical table creation statements, and execution statements. Temporary table statements can be SET statements, used to set environment parameters or create temporary views. Physical table creation statements can create physical tables or views. Execution statements can be INSERT, SELECT, etc., for example, DML statements that actually perform calculations and read / write data, as well as query statements. DML statements can be INSERT INTO, and query statements can be SELECT. Each statement is processed sequentially, and the most appropriate strategy is adopted based on its type. For temporary table statements (i.e., SET statements), they are executed directly. If execution fails, the SQL and error reason are recorded. For physical table statements, a suffix is ​​added, and statements using that table are replaced. The suffix is ​​added to avoid creating or modifying real physical tables during the validation process, thus polluting the online data environment. For example, the original statement is CREATE TABLE table_a ..., which the system will change to CREATE TABLE table_a_testsuffix. ...and the replacement statement that uses this table will replace all subsequent statements that reference the original table name, such as table_a, with the table name with a suffix, such as table_a_testsuffix, to maintain the consistency of the entire job context logic. If the execution fails, the SQL and error reason will be recorded. For the executed statements, i.e., insert, select, etc., the explain prefix is ​​added and explain verification is performed. The explain prefix is ​​added before the actual executed statement, such as INSERT INTO ... SELECT ..., for example: EXPLAIN INSERT INTO ... SELECT .... If the plan generation fails, it means that there is a problem with the statement, the SQL and error reason are recorded, and the verification result is generated based on the execution record information, that is, all erroneous SQL and reasons in the verification are integrated, and the corresponding temporary table is deleted.

[0064] In one feasible implementation, step S31 may include steps D11 to D14: Step D11: The syntax of the instruction statements in the target batch processing job script is broken down into independent job statements; It should be noted that the job statement is the smallest executable code unit with independent semantics and functions that constitutes the batch processing job script.

[0065] It is understood that the job statements are the basic processing units in the script parsing and verification process. The system breaks down the complete job script into a series of continuous job statements through syntax analysis, such as SET configuration statements, CREATETABLE table creation statements, INSERT INTO data insertion statements, and SELECT query statements. Each type of statement is classified according to its characteristics and execution risks, and different preprocessing strategies are adopted, so as to efficiently and safely complete the syntax and logic verification of the entire job script without polluting the production environment.

[0066] Step D12: Classify the job statements according to different types and purposes to determine the classification instruction statements; It should be noted that the classification instruction statement is a category identifier obtained by classifying the split independent job statements according to their syntax function and the degree of impact on system resources.

[0067] It is understood that the categorized instruction statements can include environment configuration statements, structure definition statements, and data manipulation statements. Environment configuration statements can be SET, creating temporary tables, etc., which only affect the session environment and have no persistent side effects, and can be executed directly. Structure definition statements can be CREATE TABLE and other statements that create physical tables or views, which will persistently change metadata, so a temporary suffix needs to be added before execution to avoid polluting the production environment. Data manipulation statements can be INSERT, SELECT, etc., which involve a large amount of data reading, writing, and calculation, so an EXPLAIN prefix needs to be added for execution plan analysis to identify performance bottlenecks or syntax errors in advance.

[0068] Step D13: Based on the classification instruction statement, execute the corresponding verification strategy to verify the corresponding syntax and generate execution record information; It should be noted that the execution log information is a complete process log generated by the system after performing syntax verification on each categorized instruction statement, containing execution status, output content, or error details.

[0069] It is understandable that the execution log information is real-time feedback generated by the system during the syntax verification phase when it processes various statements in sequence. For example, for a SET statement that is executed directly, it records whether it was successful; for a CREATE TABLE statement with a suffix, it records whether the temporary table was created successfully; and for a SELECT statement that is analyzed by EXPLAIN, it records whether the execution plan was generated normally or the specific error information. These are then summarized into execution log information, providing a traceable chain of evidence for judging the syntactic correctness of the entire script.

[0070] Step D14: Generate a verification result based on the execution record information.

[0071] Understandably, the verification result is a static verification of the script syntax to identify whether there are any syntax errors that could lead to execution failure. The script is then input into an asynchronous large model for deep semantic analysis to detect statements that are syntactically correct but have potential logical defects. This clearly identifies various problems in the script, such as syntax errors, Cartesian product risks, and incorrect use of aliases. This allows for the early detection of risks before the job is actually executed, ensuring the efficiency of debugging and the accuracy of the results.

[0072] Step S22: Based on the verification result, input the instruction statements in the target batch processing job script into the asynchronous large model to verify the logical defects, and obtain the verification instruction statements of the target batch processing job script.

[0073] Understandably, after confirming that the script has no basic syntax errors during the syntax verification phase, the system will input the instruction statements in the script into the asynchronous large model service for in-depth logical analysis. The asynchronous large model uses its understanding of SQL semantics and business logic to detect logical defects that are syntactically correct but have potential risks, such as Cartesian products caused by missing ON conditions in table joins, incorrect field alias references, and contradictory query conditions. For the identified problems, the system will automatically generate corrected optimized statements using the optimization model, which are then marked as reliable verification instruction statements.

[0074] In a specific embodiment, based on the verification result, the instruction statements in the target batch processing job script are input into the business orchestration model in the asynchronous large model to identify the statement category and statement logic, determine the identification result, and based on the identification result, the instruction statements in the target batch processing job script are input into the generation optimization model in the asynchronous large model for adjustment, thereby obtaining the verification instruction statements of the target batch processing job script, i.e. Figure 8 As shown, Figure 8This is a flowchart illustrating the logic verification process of the batch processing job automatic debugging method in this application. Analysis of the job writing content revealed logical defects that syntax validation could not identify, such as alias errors or Cartesian products, leading to data anomalies and performance degradation. Therefore, a large model is needed to identify logical problems and generate optimized SQL, assisting in locating issues and performance bottlenecks. However, using a single model results in high performance consumption and low accuracy. Therefore, multiple models are used for processing. Specifically, the business orchestration model within the asynchronous large model detects problems. If a problem exists, the generation and optimization model within the asynchronous large model generates optimized SQL. If no problem exists, the process ends directly. For example, the business orchestration model within the asynchronous large model can determine if the input statement is an SQL statement. If not, the process ends directly. If it is, it continues to check for Cartesian products and alias issues. A Cartesian product could be a missing ON condition in a two-table join, and an alias issue could be a misspelled alias in WHERE a.id=a.id. After the checks, a Python script is used to integrate and output the results. If the business orchestration detects a problem, the generation and optimization model within the asynchronous large model generates optimized SQL based on the identified business orchestration issues. If no problem exists, the process ends directly.

[0075] In one feasible implementation, step S22 may include steps E11-E12: Step E11: Based on the verification result, input the instruction statements in the target batch processing job script into the business orchestration model in the asynchronous large model to identify the statement category and statement logic, and determine the identification result; It should be noted that the identification result is a diagnostic conclusion about the statement logic output after performing in-depth analysis of the instruction statements of the target batch processing job script using the business orchestration model in the asynchronous large model.

[0076] It is understandable that the identification results can be based on the semantic understanding of the large model to determine whether there are hidden logical problems in the statement, such as whether there is a risk of Cartesian product due to the lack of ON condition in the table join, whether the field alias is used incorrectly, and whether the query conditions are contradictory.

[0077] Step E12: Based on the identification result, the instruction statements in the target batch processing job script are input into the generation and optimization model in the asynchronous large model for adjustment, so as to obtain the verification instruction statements of the target batch processing job script.

[0078] Understandably, the generation optimization model in the asynchronous large model can be used to accurately reconstruct and correct the instruction statements in the target batch processing job script. For example, adding ON clauses to table joins with missing join conditions and correcting erroneous field alias references can generate safe and reliable verification instruction statements.

[0079] This embodiment proposes an automatic batch job debugging method. It verifies the syntax of the instruction statements in the target batch job script and determines the verification result. Based on the verification result, the instruction statements in the target batch job script are input into an asynchronous large-scale model to verify logical defects, resulting in verification instruction statements for the target batch job script. This solves the technical problem of how to perform intelligent automatic batch job debugging more securely. Compared with existing technologies, this application utilizes multiple asynchronous large-scale models to automatically scan the script instructions for syntax checking, identifies and repairs hidden errors in scripts that pass the syntax check, and outputs highly reliable verification instruction statements. This automatically generates optimized and reliable scripts, effectively preventing invalid job execution, significantly saving debugging time and computing resources, and realizing an intelligent upgrade from passive error detection to proactive prevention in the debugging process.

[0080] For example, to help understand the implementation flow of the batch job automatic debugging method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 9 , Figure 9 A simplified flowchart of a batch job automatic debugging method is provided, specifically: Referring to Example 1, obtain a debug copy; replace the target table in the debug copy to determine the debug copy information; deploy the corresponding batch script with the necessary permissions of the allocated temporary account based on the debug copy information to determine the target batch processing job script; debug the batch processing job based on the target batch processing job script to generate the batch processing job debugging result. Referring to Example 2, verify the syntax of the instruction statements in the target batch processing job script to determine the verification result; input the instruction statements in the target batch processing job script into the asynchronous large model to verify logical defects based on the verification result to obtain the verification instruction statements of the target batch processing job script. First, automatically map the personal cluster account corresponding to the current user and determine whether it has permissions to all source tables for the current job execution. If not, terminate directly; if it has permissions, generate a temporary target table. If generation fails, terminate directly; if generation is successful, generate a debug copy and replace the table name, replacing the target table with the temporary target table. If execution fails, terminate the process; if execution is successful, deploy the job. If deployment fails, terminate the process; if deployment is successful, grant permissions to the temporary account, then perform job syntax verification. If verification passes, execute the job. After the job is completed, notify the user and revoke the temporary account permissions. If the job is executed successfully, data verification will be performed.

[0081] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the automatic debugging method for batch processing jobs in this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0082] This application also provides an automatic debugging device for batch processing jobs; please refer to [reference needed]. Figure 10 The batch processing job automatic debugging device includes: Module 10 is used to obtain a debug copy; Processing module 20 is used to replace the target table in the debug copy and determine the debug copy information; The processing module 20 is also used to deploy the corresponding batch processing script based on the debug copy information and with the necessary permissions of the allocated temporary account, and to determine the target batch processing job script; The execution module 30 is used to debug the batch processing job based on the target batch processing job script and generate the batch processing job debugging result.

[0083] The processing module 20 is also used to obtain source table permission information; The permissions of the corresponding personal cluster account are verified using the source table permission information, and the permission verification result is determined. Based on the permission verification result, the target table in the debug copy is replaced to obtain debug copy information.

[0084] The processing module 20 is also used to generate a corresponding temporary target table based on the permission verification result; Based on the debug copy, different script types are identified, and script type information is determined. Based on the script type information, the corresponding replacement method is selected to replace the target table in the debug copy with the temporary target table, thereby obtaining the debug copy information.

[0085] The execution module 30 is also used to verify the syntax of the instruction statements in the target batch processing job script and determine the verification result; Based on the verification results, the instruction statements in the target batch processing job script are input into the asynchronous large model to verify logical defects, thereby obtaining the verification instruction statements of the target batch processing job script.

[0086] The execution module 30 is further configured to split the syntax of the instruction statements in the target batch processing job script into independent job statements; The operation statements are classified according to different types and purposes to determine the classification instruction statements; Based on the classification instruction statement, the corresponding verification strategy is executed to verify the corresponding syntax and generate execution record information; A verification result is generated based on the execution record information.

[0087] The execution module 30 is further configured to input the instruction statements in the target batch processing job script into the business orchestration model in the asynchronous large model based on the verification result to identify the statement category and statement logic, and determine the identification result; Based on the identification results, the instruction statements in the target batch processing job script are input into the generation and optimization model in the asynchronous large model for adjustment, thereby obtaining the verification instruction statements of the target batch processing job script.

[0088] The execution module 30 is also used to obtain the production target table; Execute batch processing job debugging on the target batch processing job script to generate a debugging target table; Based on the debugging target table and the production target table, a corresponding comparison and verification report is generated to obtain the batch processing job debugging results.

[0089] The batch processing job automatic debugging device provided in this application, employing the batch processing job automatic debugging method in the above embodiments, can solve the technical problem of how to perform intelligent batch processing job automatic debugging more safely. Compared with the prior art, the beneficial effects of the batch processing job automatic debugging device provided in this application are the same as those of the batch processing job automatic debugging method provided in the above embodiments, and other technical features in the batch processing job automatic debugging device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0090] This application provides an automatic batch job debugging device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the automatic batch job debugging method in the first embodiment described above.

[0091] The following is for reference. Figure 11 The diagram illustrates a structural schematic suitable for implementing the batch job automatic debugging device in the embodiments of this application. The batch job automatic debugging device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11The batch job automatic debugging device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0092] like Figure 11 As shown, the batch job automatic debugging device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the batch job automatic debugging device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touch screens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the batch job auto-debugging equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows batch job auto-debugging equipment with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0093] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0094] The batch processing job automatic debugging device provided in this application, employing the batch processing job automatic debugging method in the above embodiments, can solve the technical problem of how to perform intelligent batch processing job automatic debugging more safely. Compared with the prior art, the beneficial effects of the batch processing job automatic debugging device provided in this application are the same as those of the batch processing job automatic debugging method provided in the above embodiments, and other technical features in this batch processing job automatic debugging device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0095] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0096] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0097] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the batch job automatic debugging method in the above embodiments.

[0098] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0099] The aforementioned computer-readable storage medium may be included in the batch job automatic debugging device; or it may exist independently and not be assembled into the batch job automatic debugging device.

[0100] The aforementioned computer-readable storage medium carries one or more programs. When the one or more programs are executed by the batch job automatic debugging device, the batch job automatic debugging device causes the following: to obtain a debug copy; to replace the target table in the debug copy and determine the debug copy information; to deploy the corresponding batch script based on the debug copy information with the necessary permissions of the assigned temporary account and determine the target batch job script; and to perform batch job debugging based on the target batch job script and generate batch job debugging results.

[0101] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0103] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0104] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described automatic batch job debugging method, thereby solving the technical problem of how to perform intelligent automatic batch job debugging more securely. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the automatic batch job debugging method provided in the above embodiments, and will not be repeated here.

[0105] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for automatically debugging batch processing jobs, characterized in that, The method includes: Obtain a debug copy, replace the target table in the debug copy, and determine the debug copy information; Based on the debug copy information, deploy the corresponding batch script with the necessary permissions of the assigned temporary account to determine the target batch job script; Batch processing job debugging is performed based on the target batch processing job script, and batch processing job debugging results are generated.

2. The method as described in claim 1, characterized in that, The step of replacing the target table in the debug copy and determining the debug copy information includes: Retrieve source table permission information; The permissions of the corresponding personal cluster account are verified using the source table permission information, and the permission verification result is determined. Based on the permission verification result, the target table in the debug copy is replaced to obtain debug copy information.

3. The method as described in claim 2, characterized in that, The step of replacing the target table in the debug copy based on the permission verification result to obtain debug copy information includes: A corresponding temporary target table is generated based on the permission verification result; Based on the debug copy, different script types are identified, and script type information is determined. Based on the script type information, the corresponding replacement method is selected to replace the target table in the debug copy with the temporary target table, thereby obtaining the debug copy information.

4. The method as described in claim 1, characterized in that, After the step of deploying the corresponding batch script with the necessary permissions of the allocated temporary account based on the debug copy information and determining the target batch job script, the method further includes: The syntax of the instruction statements in the target batch processing job script is validated, and the validation result is determined. Based on the verification results, the instruction statements in the target batch processing job script are input into the asynchronous large model to verify logical defects, thereby obtaining the verification instruction statements of the target batch processing job script.

5. The method as described in claim 4, characterized in that, The step of validating the syntax of the instruction statements in the target batch processing job script and determining the validation result includes: The syntax of the instruction statements in the target batch processing job script is broken down into independent job statements; The operation statements are classified according to different types and purposes to determine the classification instruction statements; Based on the classification instruction statement, the corresponding verification strategy is executed to verify the corresponding syntax and generate execution record information; A verification result is generated based on the execution record information.

6. The method as described in claim 4, characterized in that, The step of inputting the instruction statements in the target batch processing job script into the asynchronous large model to verify logical defects based on the verification result, and obtaining the verification instruction statements of the target batch processing job script, includes: Based on the verification results, the instruction statements in the target batch processing job script are input into the business orchestration model in the asynchronous large model to identify the statement category and statement logic, and determine the identification results. Based on the identification results, the instruction statements in the target batch processing job script are input into the generation and optimization model in the asynchronous large model for adjustment, thereby obtaining the verification instruction statements of the target batch processing job script.

7. The method as described in claim 1, characterized in that, The step of debugging the batch processing job based on the target batch processing job script and generating the batch processing job debugging result includes: Obtain the production target table; Execute batch processing job debugging on the target batch processing job script to generate a debugging target table; Based on the debugging target table and the production target table, a corresponding comparison and verification report is generated to obtain the batch processing job debugging results.

8. An automatic debugging device for batch processing jobs, characterized in that, The device includes: The acquisition module is used to acquire a debug copy, replace the target table in the debug copy, and determine the debug copy information; The processing module is used to deploy the corresponding batch processing script based on the debug copy information and with the necessary permissions of the allocated temporary account, and to determine the target batch processing job script. The execution module is used to debug the batch processing job based on the target batch processing job script and generate the batch processing job debugging results.

9. An automatic debugging device for batch processing jobs, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the batch job automatic debugging method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the batch job automatic debugging method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Field credibility evaluation method and system based on Spark and LLM

    CN121560872A

  • Spark and LLM-based field credibility evaluation method and system

    CN121560872B