Data processing methods, electronic devices, storage media and computer program products
Patent Information
- Application Number
- CN202610778556.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-09-18
AI Technical Summary
但是相关技术在管理SQL文件等过程中,对人工的依赖度较高,这导致SQL文件的执行效率较低
本申请实施例中,获取一个或多个结构化查询语句SQL文件;根据所述一个或多个SQL文件中的每一个SQL文件的命名规则和/或存储路径,确定第一信息,所述第一信息包括以下一项或多项:所述每一个SQL文件对应的执行环境,所述每一个SQL文件对应的目标数据中心;基于所述第一信息确定所述每一个SQL文件对应的执行计划;根据所述执行计划分发执行任务,其中,所述执行任务表征执行所述一个或多个SQL文件中对应的SQL文件。由此可见,本申请实施例中,可以根据SQL文件的命名规则和/或存储路径等自动得到第一信息,然后自动根据第一信息生成执行计划,并根据执行计划自动进行SQL文件的执行任务的分发,可见,本申请实施例能够实现自动化生成SQL文件的执行计划以及自动化进行SQL文件的分发,实现对SQL文件的自动化管理,减少对人工的依赖,从而提高SQL文件的执行效率。
Smart Images

Figure CN122777552A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a data processing method, electronic device, storage medium, and computer program product. Background Technology
[0002] With the rapid development of cloud management platforms, routine upgrades of these platforms involve an increasing number of database-related Structured Query Language (SQL) files. As data centers expand and SQL execution tasks become more complex, the requirements for database management and SQL file execution processes in data centers are also rising. However, related technologies rely heavily on manual intervention in managing SQL files, leading to low execution efficiency. Summary of the Invention
[0003] To address the related technical problems, embodiments of this application provide a data processing method, an electronic device, a storage medium, and a computer program product.
[0004] The technical solution of this application embodiment is implemented as follows: This application provides a data processing method, the method comprising: Retrieve one or more structured query statement SQL files; Based on the naming rules and / or storage path of each of the one or more SQL files, first information is determined, which includes one or more of the following: the execution environment corresponding to each SQL file, and the target data center corresponding to each SQL file. Based on the first information, determine the execution plan corresponding to each SQL file; Execution tasks are distributed according to the execution plan, wherein the execution tasks represent the execution of the corresponding SQL files in the one or more SQL files.
[0005] In the above scheme, the execution plan includes one or more of the following: The execution order of the corresponding SQL files; The dependencies between the corresponding SQL file and other SQL files; The corresponding SQL file corresponds to the execution environment; The execution time of the corresponding SQL file.
[0006] In the above scheme, determining the execution plan corresponding to each SQL file based on the first information includes: The execution environment corresponding to each SQL file mentioned in the first information is used as the execution environment corresponding to that SQL file in the execution plan; Based on the launch date and / or the readiness of the execution environment corresponding to each SQL file, determine the execution time and execution order of the SQL file.
[0007] In the above scheme, the execution environment corresponding to each SQL file is one of multiple first execution environments corresponding to that SQL file. These multiple first execution environments include one or more of the following: development environment; testing environment; live network verification environment; live network environment; and / or, The target data center corresponding to each SQL file is the data center of multiple provinces and the data center corresponding to that SQL file.
[0008] In the above scheme, after distributing the execution tasks according to the execution plan, the method further includes: If the execution of the first SQL file in one or more SQL files fails, and the execution of the first SQL file fails again, a rollback mechanism is triggered, and the reason for the failure of the first SQL execution is analyzed.
[0009] In the above scheme, after obtaining one or more structured query statement SQL files, the method further includes: Each SQL file is processed by the first model to obtain a first prediction result, which includes one or more of the following: the probability that the SQL file fails to execute; the probability that the execution time of the SQL file exceeds a first threshold. If the first prediction result satisfies the first condition, a first message is sent; wherein the first condition includes one or more of the following: the probability of the SQL file failing to execute is greater than a second threshold, and the probability of the SQL file's execution time exceeding the first threshold is greater than a third threshold; the first message includes one or more of the following: warning information, which is used to indicate that there is a risk in the execution of the SQL file; and suggestions for the execution process of the SQL file.
[0010] In the above scheme, the first model is trained using historical SQL execution data, which includes one or more of the following: the complete text of one or more historical SQL files; the execution result of each historical SQL file in the one or more historical SQL files; the execution environment context information of each historical SQL file; and the execution performance data of each historical SQL file.
[0011] This application also provides a data processing apparatus, including: The first acquisition unit is used to acquire one or more structured query statement SQL files; The first determining unit is configured to determine first information based on the naming rules and / or storage path of each of the one or more SQL files, wherein the first information includes one or more of the following: the execution environment corresponding to each SQL file, and the target data center corresponding to each SQL file; The second determining unit is used to determine the execution plan corresponding to each SQL file based on the first information; The first distribution unit is configured to distribute execution tasks according to the execution plan, wherein the execution task represents the execution of the corresponding SQL file in the one or more SQL files.
[0012] This application also provides an electronic device, including: a processor and a memory for storing a computer program capable of running on the processor. The processor is used to execute the steps of any of the above-mentioned technical solutions when running the computer program.
[0013] This application also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above methods.
[0014] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above methods.
[0015] The embodiments of this application have the following beneficial effects: In this embodiment, one or more structured query statement (SQL) files are obtained; based on the naming rules and / or storage paths of each SQL file in the one or more SQL files, first information is determined, the first information including one or more of the following: the execution environment corresponding to each SQL file, the target data center corresponding to each SQL file; based on the first information, an execution plan corresponding to each SQL file is determined; and execution tasks are distributed according to the execution plan, wherein the execution task represents the execution of the corresponding SQL file in the one or more SQL files. Therefore, in this embodiment, the first information can be automatically obtained based on the naming rules and / or storage paths of the SQL files, then an execution plan can be automatically generated based on the first information, and the execution tasks of the SQL files can be automatically distributed according to the execution plan. Thus, this embodiment can achieve automated generation of execution plans for SQL files and automated distribution of SQL files, realizing automated management of SQL files, reducing reliance on manual labor, and thereby improving the execution efficiency of SQL files. Attached Figure Description
[0016] Figure 1 A schematic diagram of the architecture of the Ceph storage system provided for the application embodiments of this application; Figure 2 A schematic flowchart illustrating the data processing method provided in the application embodiments of this application; Figure 3 A schematic diagram of the network architecture of the first model provided for the application embodiments of this application; Figure 4 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0018] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, the term "one or more" in this document is an exemplary expression and can be replaced with any possible expression, such as one or more, at least one, or at least one item.
[0019] It should also be understood that the term "instruction" mentioned in the embodiments of this application can be a direct instruction, an indirect instruction, or an indication of a relationship. For example, A instructing B can mean that A directly instructs B, such as B being able to obtain information through A; it can also mean that A indirectly instructs B, such as A instructing C, so B can obtain information through C; or it can mean that there is a relationship between A and B.
[0020] It should also be understood that the term "correspondence" mentioned in the embodiments of this application may indicate a direct or indirect correspondence between the two, or an association between the two, or a relationship of instruction and being instructed, configuration and being configured, etc.
[0021] It should be noted that terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0022] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.
[0023] The relevant technologies of the embodiments of this application are described below: With the rapid development of cloud management platforms, the number of related products has increased significantly, and the resource pools and databases of cloud management platforms have also grown substantially. As cloud management platforms develop rapidly, routine upgrades involve an increasing number of database-related SQL queries. Furthermore, the same SQL query needs to be executed in development, testing, live network verification, and the live network environment. Currently, the live network environment has over 30 resource pools, resulting in over 30 databases. The overall architecture is as follows: Figure 1 As shown. The current database update process is as follows: First, the developer writes an SQL statement and uploads it to the Gerrit code management repository. The database administrator retrieves the SQL statement from the Gerrit repository and manually executes it in the development environment. Then, during the testing phase, the SQL statement is manually executed in the test environment database. The test domain has three resource pool databases, and the SQL statement needs to be executed in each of them. Two days before the live network change, the administrator needs to execute the SQL statement in over 30 live network verification environments to verify it. If the verification is successful, finally, just before the system upgrade, the operations personnel manually execute the SQL statement in over 30 live network resource pools. For the complete SQL process, please refer to [link to documentation]. Figure 2 .
[0024] As data centers expand and SQL execution tasks become more complex, several technical shortcomings exist in multi-datacenter database management and SQL execution processes. These shortcomings directly impact system efficiency, security, and stability. For example, executing new or updated SQL files requires manual intervention, which is slow and error-prone. SQL operations in some resource pools are frequently missed, or duplicate executions of certain SQL files lead to upgrade failures. In severe cases, this can disrupt the normal operation of the production environment and cause major malfunctions.
[0025] To address the aforementioned problems, this application proposes a data processing method. In this embodiment, one or more structured query statement (SQL) files are obtained; based on the naming rules and / or storage paths of each SQL file, first information is determined, including one or more of the following: the execution environment corresponding to each SQL file, and the target data center corresponding to each SQL file; an execution plan corresponding to each SQL file is determined based on the first information; and execution tasks are distributed according to the execution plan, wherein the execution task represents the execution of the corresponding SQL file among the one or more SQL files. Therefore, in this embodiment, the first information can be automatically obtained based on the naming rules and / or storage paths of the SQL files, and then an execution plan can be automatically generated based on the first information. The execution tasks of the SQL files can then be automatically distributed according to the execution plan. Thus, this embodiment can achieve automated generation of execution plans for SQL files and automated distribution of SQL files, realizing automated management of SQL files, reducing reliance on manual intervention, and thereby improving the execution efficiency of SQL files.
[0026] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.
[0027] This application uses an electronic device as an example of an execution subject, and a data processing device as one of the manifestations of an electronic device. This application does not limit the manifestation of electronic devices.
[0028] Please see Figure 3 The data processing method provided in this application includes: Step 301: Obtain one or more structured query statement SQL files; For example, the one or more SQL files mentioned above are used for database management. For instance, the one or more SQL files mentioned above can be SQL files that have changed during the upgrade process of applications, services, or programs used to manage the database. For example, the one or more SQL files can be newly added SQL files or SQL files that have been updated.
[0029] For example, the one or more SQL files mentioned above can be SQL files uploaded by the user.
[0030] For example, after receiving one or more SQL files uploaded by the user, the first system automatically checks the format and naming conventions of each SQL file and verifies the integrity of the SQL files. Based on the user's permissions, it displays appropriate operation options and automatically categorizes the SQL files into the corresponding environment directories. For example, this environment directory could be the directory of the execution environment corresponding to the SQL file.
[0031] Exemplary embodiments of this application can be applied to related electronic devices, systems, or modules. The following example uses an embodiment of this application applied to a first system. For instance, this first system can be an intelligent SQL optimization and error prediction management system. Through automated management and intelligent optimization, this first system effectively improves the efficiency and accuracy of SQL execution and reduces the risk of human error. The first system in this embodiment integrates SQL management across multiple data centers.
[0032] The first system in the embodiments of this application is described below: The first system provides a management interface that is user-friendly and feature-rich. The interface is designed to simplify the management and execution of SQL files while ensuring operational security and controllability. The interface design not only prioritizes user experience but also fully considers system scalability and security. The first system possesses one or more of the following key features: hierarchical and decentralized management; an intuitive user interface; security control and auditing; efficient environment management; and scalability and customization.
[0033] The following describes the hierarchical and decentralized management characteristics of the first system: For example, in the first system, different types of users are configured with different permissions. The different types of users include one or more of the following: ordinary users, R&D personnel, database administrators, and system administrators. The system administrator can assign permissions appropriate to each user’s responsibilities to ensure that different roles can only access and operate functions that match their permissions.
[0034] It is understood that in this embodiment of the application, the first system supports multi-level user role configuration, with each role corresponding to a different scope of permissions.
[0035] For example, ordinary users have relatively basic permissions, which mainly include one or more of the following: uploading SQL files; viewing SQL execution status and logs. Ordinary users cannot execute or modify SQL files; this design is to prevent unauthorized operations from affecting the system.
[0036] For example, R&D personnel have the same permissions as ordinary users. In addition to these permissions, R&D personnel also have the ability to upload SQL files to directories in different environments, such as the R&D environment, testing environment, live verification environment, and production environment. The upload paths for these SQL files can be automatically matched by the first system or manually selected by the user, ensuring that the SQL files are accurately categorized and managed.
[0037] For example, a database administrator's permissions include one or more of the following: permission to execute SQL files, permission to view SQL files related to the environment under the database administrator's responsibility, and permission to execute SQL files as needed. The permission to execute SQL files includes one or more of the following: executing SQL files; verifying SQL files; and performing SQL rollback operations when necessary.
[0038] For example, each database administrator can only operate within their authorized environment, preventing erroneous operations across environments.
[0039] For example, the system administrator has the highest privileges, which include one or more of the following: managing all users; configuring system parameters; auditing operation logs; and intervening in the work of other types of users in special circumstances, such as forcibly suspending SQL execution in the event of a major error.
[0040] The following describes the intuitive user interface of the first system: For example, the first system has one or more of the following operation interfaces: an upload and management interface; an SQL execution status monitoring interface; and a notification and reminder function interface.
[0041] For example, regarding the upload and management interface, the first system provides a simple and intuitive upload interface. Users can upload SQL files by dragging and dropping or selecting SQL files. When a user uploads an SQL file, the first system automatically detects the SQL file's format and naming conventions, and verifies the file's integrity. The upload and management interface displays corresponding operation options based on the user's permissions and automatically categorizes the SQL files into the corresponding environment directories.
[0042] For example, the SQL execution status monitoring interface includes a real-time monitoring interface provided by the first system for administrator users. Through this interface, the status of SQL execution can be viewed intuitively, including the currently executing SQL file, executed SQL files, and the execution results (e.g., successful or failed). The interface also displays detailed logs for each SQL execution, including one or more of the following: execution time, execution environment, resource usage, etc. This information helps administrators quickly locate and resolve problems when they occur.
[0043] For example, the first system integrates multi-channel notification functionality, allowing users to choose how to receive notifications (such as email, SMS, or system messages) according to their personal preferences. For instance, when an SQL file is successfully uploaded, completes execution, or encounters an execution error, the first system will automatically send a notification to the relevant user. Administrators can also set up regular status reports, such as periodically reporting the execution status of SQL files, ensuring that no important information is missed during long-term operations.
[0044] For example, in this embodiment of the application, the functions of the upload and management interface and the SQL execution status monitoring interface can all be integrated into the first interface, which is then displayed to the user. The first interface can be used to upload SQL files, monitor the execution status of SQL files, and manage permissions. That is, in this embodiment of the application, the user can complete the upload, execution status monitoring, and permission management of SQL files within the same interface.
[0045] The following describes the security control and auditing of the first system: For example, the first system can implement permission isolation and verification. It employs a strict permission isolation mechanism to ensure that each user can only access the resources they are authorized to access. Before each operation, the system performs permission verification to prevent unauthorized users from accessing or modifying SQL files. Furthermore, the first system supports multi-factor authentication (MFA) to enhance user login security.
[0046] For example, the first system can record operation logs for auditing purposes. For instance, the first system can record logs for all critical operations, including one or more of the following: SQL file upload; SQL file execution; user permission modification, etc. Log records include one or more of the following: the time of the operation; the user performing the operation; the specific content of the operation; and the result of the operation. The logs recorded by the first system can be viewed or exported by the system administrator at any time for security auditing and compliance checks.
[0047] For example, the first system is capable of data encryption and protection. Uploaded SQL files are encrypted during storage to prevent unauthorized access. Simultaneously, the first system employs encryption protocols (such as Transport Layer Security (TLS)) during transmission to protect data security, ensuring that data is not tampered with or leaked during uploading, downloading, and execution.
[0048] The following describes the efficient environmental management of the first system: For example, the first system enables environment configuration and management. System administrators can configure and manage different environments, such as development environments, testing environments, live network verification environments, and production environments, through the system's interface. Each environment can independently configure its corresponding database connection information, execution strategies, resource limits, and other configuration parameters. Through this graphical management system, system administrators can easily configure and maintain multiple environments, ensuring the accuracy of each environment's configuration.
[0049] For example, the first system enables isolation and switching between environments. It employs a strict isolation mechanism to ensure that SQL files are not confused or misoperated across different environments. When a user uploads and executes an SQL file, the first system automatically matches the target environment corresponding to the SQL file and performs an environment consistency check before execution to ensure that the operating environment is consistent with the target environment of the SQL file. Simultaneously, administrators can easily switch between different environments through the interface and view the execution status of SQL files in each environment.
[0050] For example, the first system can also perform environmental health checks and alerts. The first system integrates an environmental health check function, which regularly scans the status of each environment, including the stability of database connections and the availability of resources. If an anomaly is detected (such as connection failure, insufficient resources, etc.), the first system will issue an alert in a timely manner to remind the administrator to take appropriate measures to avoid SQL file execution interruption or failure.
[0051] The following describes the scalability and customization of the first system, as detailed below: For example, the first system adopts a modular design, supporting the expansion of functional modules according to user needs. For instance, custom user interfaces can be developed for specific user groups, or specific third-party tools (such as monitoring tools and log analysis tools) can be integrated. This design ensures that the system can be flexibly expanded as user needs change.
[0052] For example, the first system supports customized interface layouts, allowing users to customize the interface layout and function display order according to their personal preferences. The first system provides a variety of layout templates and themes, allowing users to choose an interface style that suits their work habits. Furthermore, the first system also supports users creating shortcut buttons to simplify the execution steps of frequently used operations.
[0053] For example, the first system supports multiple languages, and the interface supports multilingual switching, making it convenient for users from different regions and language backgrounds. Users can select their preferred language in their personal settings, and the first system will automatically switch to the corresponding language interface, enhancing the convenience of the user experience.
[0054] Step 302: Determine first information based on the naming rules and / or storage path of each SQL file in the one or more SQL files. The first information includes one or more of the following: the execution environment corresponding to each SQL file, and the target data center corresponding to each SQL file. For example, in the embodiments of this application, the naming rules for SQL files include: the name of each SQL file includes the execution environment and / or target data center corresponding to the SQL file. For example, the name of the first SQL file among one or more SQL files can be: live network testing environment - identifier of data center 1 - serial number of the first SQL file.
[0055] In practical applications, the execution environment corresponding to each SQL file is one of multiple first execution environments corresponding to that SQL file. These multiple first execution environments include one or more of the following: development environment; testing environment; live network verification environment; live network environment; and / or, The target data center corresponding to each SQL file is the data center of multiple provinces and the data center corresponding to that SQL file.
[0056] For example, taking the first SQL file among one or more SQL files as an example, the execution environment corresponding to the first SQL file can be one or more of the following: development environment; testing environment; live network verification environment; live network environment; and the data center corresponding to the first SQL file can be one or more of the following data centers: data center in province A, data center in province B, data center in province C, etc.
[0057] It is understandable that the execution environment corresponding to an SQL file refers to the execution environment in which the SQL file needs to be executed.
[0058] For example, in this application embodiment, a mapping relationship between the storage path of the SQL file and the execution environment and / or the execution environment can be established. In this application embodiment, the storage path includes one or more of the following: root directory; one or more subdirectories connected by delimiters; and the folder corresponding to the SQL file.
[0059] For example, in this embodiment of the application, the storage directory can be divided into multiple different subdirectories.
[0060] For example, different subdirectories correspond to different execution environments. For instance, subdirectory 1 corresponds to the development environment, subdirectory 2 corresponds to the testing environment, subdirectory 3 corresponds to the live verification environment, and subdirectory 4 corresponds to the live environment. For example, when the storage path of the first SQL file is: root directory:\subdirectory 1\folder 1, the execution environment corresponding to the first SQL file is the development environment.
[0061] For example, different subdirectories correspond to data centers in different provinces. For instance, subdirectory 1 corresponds to the data center in province A, subdirectory 2 corresponds to the data center in province B, and subdirectory 3 corresponds to the data center in province C. For example, when the storage path corresponding to the first SQL file is root directory:\subdirectory 1\folder 1, the data center corresponding to the first SQL file is the data center in province A.
[0062] For example, upon determining the first piece of information, the one or more SQL files can be automatically categorized and stored to ensure that the SQL files for each execution environment are stored in an orderly manner in their respective directories. For instance, each SQL file in the one or more SQL files can be stored in the database of the data center corresponding to that SQL file. For example, in this embodiment of the application, different execution environments can correspond to different subdirectories, and each SQL file in the one or more SQL files can be saved in the subdirectory corresponding to the execution environment of that SQL file.
[0063] Step 303: Determine the execution plan corresponding to each SQL file based on the first information; It is understood that when the embodiments of this application are applied to the first system, the first system has highly automated SQL execution management functions, such as automatically obtaining first information and automatically generating execution plans, in order to ensure that the execution process of SQL files is seamless, accurate and timely.
[0064] For example, after one or more SQL files are uploaded, the first system will automatically generate an execution plan for each SQL file.
[0065] In practical applications, the execution plan includes one or more of the following: The execution order of the corresponding SQL files; The dependencies between the corresponding SQL file and other SQL files; The corresponding SQL file corresponds to the execution environment; The execution time of the corresponding SQL file.
[0066] For example, dependency means that when an SQL file is executed, it needs to rely on the execution process of other SQL files to obtain intermediate results or the final output results. In other words, when an SQL file is executed, it needs to rely on the execution of other SQL files.
[0067] For example, the execution time of an SQL file refers to the time schedule for the execution of each SQL file in one or more SQL files.
[0068] In practical applications, determining the execution plan corresponding to each SQL file based on the first information includes: The execution environment corresponding to each SQL file mentioned in the first information is used as the execution environment corresponding to that SQL file in the execution plan; Based on the launch date and / or the readiness of the execution environment corresponding to each SQL file, determine the execution time and execution order of the SQL file.
[0069] It is understood that in this embodiment of the application, the execution time and execution order of the SQL file are determined based on the online date and the readiness of the execution environment, so as to ensure that the SQL file is executed within a predetermined time and reduce the risk of delays in the online launch of related applications or programs.
[0070] For example, the launch date refers to the launch date of the application, program, system, patch, or service corresponding to the SQL file. The installation package of the application, program, system, patch, or service includes the SQL file, which can be used to manage the database of the corresponding data center.
[0071] For example, the first system automatically adjusts the execution time and execution order of each SQL file in one or more SQL files according to the set launch date and the readiness of each environment, so as to ensure that all SQL files are executed within the predetermined time.
[0072] Step 304: Distribute execution tasks according to the execution plan, wherein the execution tasks represent the execution of the corresponding SQL files in the one or more SQL files.
[0073] For example, the step of distributing execution tasks according to the execution plan includes: the first system distributing the execution tasks to the corresponding database administrators or automated execution modules according to the generated execution plan.
[0074] For example, if the first SQL file in one or more SQL files corresponds to data center A, the execution task of the first SQL file will be distributed to the automated execution module or the database administrator of data center A. The database administrator of data center A can trigger the automatic execution of the first SQL file.
[0075] For example, the first system supports parallel execution of SQL file execution tasks for different data centers and execution environments, thereby significantly improving execution efficiency and reducing time delays and errors caused by manual operations.
[0076] For example, to ensure the comprehensiveness and accuracy of the SQL execution process, the first system has a built-in automated alarm mechanism that can proactively monitor and notify relevant personnel at key stages of the execution process. Specific functions include one or more of the following: real-time monitoring and timely alarms; online date checks and alerts; multi-level alarm policies; alarm history records and analysis.
[0077] For example, regarding real-time monitoring and timely alerting, the first system will monitor the execution process of SQL files in real time, including key indicators such as execution success or failure, execution time, and resource consumption. If problems occur during execution (such as SQL execution failure, execution time exceeding expectations, abnormal resource usage, etc.), the system will immediately trigger an alert mechanism to notify relevant R&D personnel, database administrators, and operations and maintenance personnel via email, SMS, or instant messaging tools, ensuring that the problem can be handled in a timely manner.
[0078] For example, regarding the deployment date check and early warning function, the first system periodically checks the execution status of SQL files related to the deployment date. Based on the predetermined deployment plan, the first system automatically verifies whether the SQL files in each execution environment have been uploaded and executed, and checks the execution results. If any SQL files are found to be missing, have failed to execute, or have not yet been executed, the first system will issue an early warning at a pre-set time (such as 72 hours or 48 hours before deployment) to remind relevant personnel to take remedial measures, preventing problems from being discovered at the last minute and affecting the overall deployment schedule.
[0079] For example, to avoid information delays or omissions that might occur with a single alarm mechanism, the first system has designed a multi-level alarm strategy. Basic alarms address common anomalies (such as execution delays) and notify relevant personnel via email or system notifications. Intermediate alarms escalate the alarm level when serious problems are detected (such as SQL execution failures or critical SQL queries not being executed), directly notifying relevant personnel via SMS or instant messaging tools. Advanced alarms are used to handle critical issues (such as insufficient system resources or multiple failed execution attempts). The first system will directly initiate an emergency phone call and generate a detailed alarm report for management to make rapid decisions.
[0080] For alarm history recording and analysis, for example, the first system saves the history of all alarms and generates detailed alarm logs. The alarm logs include one or more of the following: the time the alarm was triggered; the reason for the alarm; and the result of the alarm handling. By analyzing the alarm history, the first system can identify common failure modes and vulnerabilities, provide improvement suggestions, and offer more accurate risk predictions and preventative measures for future SQL file execution.
[0081] For example, to further ensure the security and stability of SQL execution, the first system also provides an automated rollback and recovery mechanism, as detailed in the following practical applications: In practical applications, after distributing execution tasks according to the execution plan, the method further includes: If the execution of the first SQL file in one or more SQL files fails, and the execution of the first SQL file fails again, a rollback mechanism is triggered, and the reason for the failure of the first SQL file is analyzed.
[0082] It is understood that the embodiments of this application can ensure the security and data consistency of database operations through the automatic rollback function when execution fails. Analyzing the reasons for the failure of SQL file execution can help analyze the execution of SQL file and improve the success rate of SQL file execution.
[0083] For example, if the execution of the first SQL file in one or more SQL files fails, and the execution of the first SQL file fails again, a rollback mechanism is triggered to restore the state of the database corresponding to the first SQL file to the state before the execution of the first SQL.
[0084] Understandably, the above practical applications can achieve automatic rollback after SQL file execution failure. If the first system detects that a certain SQL execution has failed and cannot solve the problem through automatic retry, the first system will trigger the automatic rollback mechanism to restore the corresponding database state to the state before the SQL was executed, so as to avoid data inconsistency or system failure caused by SQL file execution failure.
[0085] For example, regarding automatic recovery and remediation, after the rollback is complete, the first system will automatically analyze the first SQL file that failed to execute, identify possible reasons for the failure (such as syntax errors, insufficient permissions, resource limitations, etc.), and attempt to automatically adjust the parameters of the first SQL file or the execution environment to remedy the situation. If the first system can successfully recover from the problem, the first SQL file will be re-executed; if the automatic recovery of the first SQL file fails, the first system will generate a detailed fault report and notify relevant personnel to perform manual intervention.
[0086] For example, embodiments of this application support intelligent SQL optimization and error prediction. The first system has a built-in machine learning module that performs SQL optimization and error prediction based on historical SQL execution data. Through the analysis of historical SQL execution data and the training of the machine learning model, the first system can predict the risk of SQL file execution failure in advance and provide early warnings and optimization suggestions for high-risk operations, thereby reducing errors that may occur during SQL execution and improving the overall stability and security of the system. Specific practical applications can be found below: In practical applications, after obtaining one or more structured query statement (SQL) files, the method further includes: Each SQL file is processed by the first model to obtain a first prediction result, which includes one or more of the following: the probability that the SQL file fails to execute; the probability that the execution time of the SQL file exceeds a first threshold. If the first prediction result satisfies the first condition, a first message is sent; wherein the first condition includes one or more of the following: the probability of the SQL file failing to execute is greater than a second threshold, and the probability of the SQL file's execution time exceeding the first threshold is greater than a third threshold; the first message includes one or more of the following: warning information, which is used to indicate that there is a risk in the execution of the SQL file; and suggestions for the execution process of the SQL file.
[0087] For example, in the embodiments of this application, the first model is a machine learning model, such as a decision tree model, a random forest model, or a gradient boosting decision tree, etc.
[0088] It is understood that the embodiments of this application optimize and predict based on historical SQL execution data, identify high-risk operations in advance and provide optimization suggestions, significantly improving the success rate of SQL execution and the stability of the system.
[0089] In practical applications, the first model is trained using historical SQL execution data, which includes one or more of the following: the complete text of one or more historical SQL files; the execution result of each historical SQL file; the execution environment context information of each historical SQL file; and the execution performance data of each historical SQL file.
[0090] It is understood that the historical SQL execution data in this application embodiment forms a set of structured datasets, laying the foundation for subsequent feature engineering and training of the first model.
[0091] For example, the complete text of one or more historical SQL files is the historical SQL statement itself, including: the complete text of historical SQL queries or commands.
[0092] For example, the execution result of the historical SQL file records whether the execution of the historical SQL file was successful. If it fails, it records the specific error type of the execution failure, such as syntax error, insufficient permissions, resource limitation, etc.
[0093] For example, the execution environment context information of each historical SQL file includes one or more of the following: database version information; operating system type; version; current host load (e.g., central processing unit (CPU), memory, input or output (I / O) load, etc.).
[0094] For example, performance data may include one or more of the following: SQL statement execution time; CPU consumption; memory usage; I / O operations, etc.
[0095] Understandably, the above data is collected through system log files, internal database execution logs, and external monitoring tools to form a structured dataset, laying the foundation for subsequent feature engineering and machine learning model training.
[0096] The following describes the process of collecting historical SQL execution data and feature engineering in this application embodiment: For example, the first system first collects historical SQL execution data from the database's logging system and monitoring tools (such as data center monitoring tools).
[0097] The feature engineering phase requires parsing the SQL statements, as detailed below: For example, in the feature engineering phase, the first system needs to convert the collected raw data (i.e., historical SQL execution data) into features that the first model can process. The specific steps are as follows: For example, historical SQL execution data can be tokenized: historical SQL execution data can be broken down into smaller tokens (such as keywords, operators, table names, field names, etc.) to facilitate subsequent processing.
[0098] For example, a syntax tree is generated based on the tokenized historical SQL execution data, thereby extracting the structural features of the historical SQL execution data.
[0099] The feature engineering stage also requires feature extraction, as detailed below: For example, the tokenized historical SQL execution data is processed using the Bag-of-Words Model to convert the tokenized historical SQL execution data into numerical features. Specifically, the tokenized historical SQL execution data can be represented as a high-dimensional sparse vector, with each dimension corresponding to the frequency of occurrence of a certain token.
[0100] For example, categorical feature encoding is performed on historical SQL execution data. Specifically, categorical features in the historical SQL execution data are transformed using one-hot encoding or label encoding to generate numerical features. For example, categorical features can be error type, database version, operating system type, etc.
[0101] For example, numerical features can be extracted from the execution environment context information. For instance, the database version number can be directly used as a numerical feature, and host load (CPU utilization, memory utilization) can also be used as a numerical feature.
[0102] Feature engineering also requires feature preprocessing, as follows: For example, numerical features are standardized so that their mean is 0 and their variance is 1, thereby eliminating the differences between different feature units.
[0103] For example, features (such as numerical features) can be normalized to obtain a feature dataset. Specifically, features can be scaled to a specific range (e.g., 0 to 1) to improve the convergence speed and prediction accuracy of the model.
[0104] For example, the first system uses the generated feature dataset (e.g., the feature dataset obtained after standardizing the features) for training and testing the candidate model, which is a machine learning model. The candidate model is trained and evaluated through testing to obtain the first model.
[0105] For example, for the SQL file error prediction task, the first system can choose a decision tree model as a candidate model. Decision tree models perform well when processing data containing a large number of categorical and numerical features, and have the following advantages: Easy to interpret: The decision-making process of the decision tree model is transparent, easy to understand and interpret, and facilitates users to analyze the basis of the model's predictions.
[0106] Handling non-linear relationships: Decision trees can capture complex non-linear relationships between features, making them suitable for complex scenarios in SQL error prediction.
[0107] No feature scaling required: Decision trees are unaffected by feature dimensions and naturally support mixed inputs of categorical and numerical features.
[0108] For example, during the training phase of the candidate model, the system employs the following steps: The following describes the steps for partitioning the dataset: For example, before splitting the dataset, the first system first cleans the collected historical SQL execution data to obtain the first dataset. For instance, the cleaning process may include removing missing values, outliers, and irrelevant data records. Preprocessing steps include normalizing or standardizing feature values to ensure the quality and consistency of the dataset.
[0109] For example, after cleaning the historical SQL execution data, the resulting first dataset is divided into a training set and a test set. For instance, 70% of the data in the first dataset can be used to train a candidate model (the training set), and 30% can be used to evaluate the performance of the candidate model (the test set). This ratio can be adjusted according to the specific data volume and task requirements. During the splitting of the first dataset, the consistency of data distribution should be ensured; that is, the feature distributions of the training set and the test set should be as similar as possible.
[0110] For example, if the error types in the dataset (such as the training set) are imbalanced (e.g., some error types are very rare), it may be necessary to oversampling or undersampling the training set to avoid the candidate model being biased towards common error types.
[0111] The following describes the cross-validation steps: For example, embodiments of this application can perform K-fold cross-validation on the training set. It is understood that the system employs the K-fold cross-validation method to better evaluate the generalization ability of candidate models. Specific steps include the following: Divide the training set into K subsets (usually K=5 or K=10). In each validation, a subset of the training set is selected as the validation set, and the remaining K-1 subsets are used for training the candidate model; Train candidate models and evaluate their performance on a validation set; Each subset is used as a validation set in turn, and finally the evaluation results of the performance of K candidate models are obtained; Calculate the average performance metrics (such as accuracy and F1 score) across K validations to evaluate the overall performance of the candidate model.
[0112] Understandably, cross-validation helps reduce the evaluation bias of candidate models due to the randomness of data partitioning, especially when the dataset is small or imbalanced.
[0113] The following describes the parameter tuning steps: For example, taking a decision tree model as the candidate model, several hyperparameters affect the performance of a decision tree model, such as maximum depth (max_depth), minimum number of sample splits (min_samples_split), and minimum number of samples per leaf node (min_samples_leaf). These hyperparameters need to be adjusted before training the decision tree model to find the optimal configuration. Specifically, parameter tuning can be performed in the following ways: For example, for grid search, the search range of hyperparameters can be defined, such as max_depth being selectable between [5, 10, 15] and min_samples_split being selectable between [2, 5, 10]. The system will exhaustively search all possible combinations of hyperparameters, train candidate models, and evaluate the performance of each hyperparameter combination. Finally, it will select the hyperparameter combination whose performance on the validation set meets the set conditions, such as the set condition being that the hyperparameter combination has the best performance.
[0114] For example, a search space for hyperparameters can be defined for random search, and combinations of hyperparameters can be randomly selected for training and evaluation. Generally, random search is more efficient than grid search when resources are limited. After a period of random search, the combination of hyperparameters with the best performance is selected.
[0115] The following describes error weight settings: For example, different types of SQL errors have different impacts on the system, so the system assigns different weights to each type of SQL error. More severe or high-risk error types are given higher weights.
[0116] For example, by introducing error weights into the loss function, the candidate model is made to focus more on predicting high-weight errors during training. Adjusting the loss function helps the candidate model improve its ability to identify high-risk errors on imbalanced datasets.
[0117] For example, in addition to error weights, different cost functions can be set to make the candidate model pay a greater "cost" when misjudging high-risk errors, thereby optimizing prediction accuracy.
[0118] The training phase of the model has been described above. The following section describes the model validation and evaluation phase: First, there is the performance evaluation, which can be seen in the following example: For example, the candidate model is evaluated using a test set to generate a confusion matrix. The confusion matrix shows the performance of the candidate model on the actual classification and the predicted classification, including four types of results: TruePositive (TP), False Positive (FP), True Negative (TN), and False Negative (FN). A True Positive is a sample that is actually positive but is correctly classified as positive; a False Positive is a sample that is actually negative but is incorrectly classified as positive; a True Negative is a sample that is actually negative but is correctly classified as negative; and a False Negative is a sample that is actually positive but is incorrectly classified as negative.
[0119] For example, in the embodiments of this application, the evaluation metrics for candidate models include one or more of the following: Precision measures the proportion of samples that a candidate model predicts to be positive, but which are actually positive. The formula is TP / (TP + FP).
[0120] Recall measures the proportion of a candidate model that is correctly identified as positive out of all samples that are actually positive. The formula is TP / (TP + FN).
[0121] The F1 score is the harmonic mean of precision and recall, taking both into account. The formula is 2*(Precision * Recall) / (Precision + Recall).
[0122] The Area Under Curve (AUC-ROC) curve is used to measure the classification performance of a candidate model at different thresholds. The closer the area under the ROC curve (AUC) value is to 1, the better the performance of the candidate model.
[0123] For example, the trained candidate model is evaluated using a first evaluation method to obtain the evaluation result. The first evaluation method includes one or more of the following: evaluating the trained candidate model based on evaluation metrics; evaluating the candidate model using a test set to generate a confusion matrix; and, if the evaluation result indicates that the performance of the trained candidate model meets the set conditions, using the trained candidate model as the first model.
[0124] The following describes how the first model is selected in the embodiments of this application: For example, embodiments of this application can select candidate models through comparative testing: For example, as mentioned above, a decision tree model can be selected as the first model. If the decision tree model performs poorly in the performance evaluation, the system can consider using other more complex or suitable machine learning models for comparative testing. For example, one or more of the following models can be selected as the first model: Random forests integrate multiple decision tree models, which can improve the robustness and accuracy of the model, and are particularly suitable for scenarios with a large number of features and high data noise.
[0125] Gradient Boosting Decision Tree: This method trains multiple weak models (decision trees) through stepwise optimization and then integrates them into a strong model. It is suitable for prediction tasks where the cost of errors is high.
[0126] For example, the same cross-validation and performance evaluation are performed on different models (such as decision tree models, random forest models, gradient boosting decision trees, etc.) to compare the performance of different models on the test set. The model with the higher AUC value and better F1 score is selected as the first model.
[0127] For example, based on the test results, the hyperparameters of the selected model are further adjusted (e.g., in a fine-tuning manner) to obtain a first model to achieve the best prediction performance, wherein the selected model can be a model with a high AUC value and a better F1 score.
[0128] As can be seen from the above, in order to solve the problems existing in the related technologies, this application proposes an intelligent SQL optimization and error prediction management system (i.e., the first system) that unifies and integrates multi-datacenter databases. This intelligent SQL optimization and error prediction management system solves the problems existing in the related technologies in the following ways: First, it achieves comprehensive automated management and execution. Through the system's automated management functions, users can simplify the management and execution process of SQL files, reduce manual operation steps, and thus reduce the risk of human error. The system automatically generates execution plans, distributes execution tasks, and supports parallel processing to improve execution efficiency.
[0129] Secondly, the system provides a unified management interface and strict hierarchical access control. It offers a feature-rich and secure management interface, allowing users to upload SQL files, monitor execution status, and manage permissions within a single interface. The system supports multi-level user role configuration, achieving strict hierarchical access control and ensuring that each user role can only access and operate functions matching their permissions.
[0130] Third, the system features built-in intelligent execution monitoring and multi-level alarm mechanisms. These mechanisms enable proactive detection of problems at critical stages of SQL execution and timely notification of relevant personnel. The system also supports automatic rollback and recovery mechanisms to ensure rapid recovery in the event of execution failure, preventing data corruption or system interruption.
[0131] Fourth, machine learning algorithms are introduced to achieve SQL optimization and error prediction. The system integrates a machine learning module, which can perform SQL optimization and error prediction based on historical data. Through in-depth analysis of execution data and model training, the system can identify high-risk operations in advance and provide optimization suggestions, thereby improving the success rate of SQL execution and the stability of the system.
[0132] Furthermore, this application embodiment enables intelligent SQL optimization and error prediction. By integrating a machine learning module, the system can optimize and predict based on historical SQL execution data, proactively identifying high-risk operations and providing optimization suggestions, significantly improving the success rate of SQL execution and system stability. This application embodiment enables automated execution and management. The system supports automated SQL file management, execution plan generation, task distribution, and parallel processing, thereby reducing the complexity of manual operations and lowering the risk of operational errors. This application embodiment enables multi-level alarms and automated rollback. The system has built-in real-time monitoring and alarm mechanisms, as well as automated rollback and recovery functions in case of execution failure, ensuring the security of database operations and data consistency. This application embodiment also enables a unified management interface and hierarchical access control. The system provides a unified and feature-rich management interface, supporting multi-level user role configuration, ensuring that users with different roles can only access functions matching their permissions, preventing unauthorized operations and misoperations.
[0133] This application proposes a machine learning-based SQL optimization and error prediction model and its specific implementation method, as well as the technical process for automated management functions such as SQL file uploading, classification, execution plan generation, task distribution, execution monitoring, alarms, and rollback. Furthermore, the system in this application provides interface design and implementation methods for multi-level hierarchical management, real-time monitoring, data encryption, and operation log auditing. Based on the above solutions, the automation level of SQL execution can be improved. This system automates most operations in the SQL management and execution process, reducing the need for human intervention, thereby reducing the risk of human error and significantly improving execution efficiency. This application provides intelligent error prediction and optimization functions. By introducing a machine learning module, the system can predict the risks of SQL execution in advance and provide optimization suggestions for potentially high-risk operations, a function not found in related technologies. This application strengthens the system's security and controllability. The system adopts strict access control, multi-factor authentication, data encryption, operation log auditing, and other multi-layered security measures to ensure data and operational security during SQL execution. This application embodiment achieves comprehensive real-time monitoring and multi-level alarms. The system's built-in real-time monitoring and alarm mechanisms can promptly detect and handle abnormal situations during execution, and support automatic rollback and recovery, further ensuring system stability. This application embodiment provides flexible expansion and customization support. The system adopts a modular design, supports functional expansion and interface customization, and can be flexibly adjusted according to the needs of different users, resulting in higher usability and adaptability.
[0134] The system provided in this application embodiment is applicable to large enterprises with multiple data centers, especially in situations involving frequent SQL execution, complex environments, and massive data volumes, significantly improving management efficiency and execution accuracy. Through automated management and intelligent optimization, the system significantly improves the efficiency and accuracy of SQL execution, reduces the risk of human error, and helps enterprises achieve more efficient database management in multi-data center environments. The system reduces manual operations and SQL execution process optimization, thus lowering the labor costs of database management and maintenance, while also reducing failure and repair costs due to operational errors. The system enhances data security; its multi-layered security controls effectively prevent data leakage and unauthorized operations, helping enterprises maintain the security of their data assets. The system's intelligent optimization and error prediction functions enable faster and more accurate handling of complex SQL execution tasks. The system design fully considers future expansion needs; its modular design and customizable interface provide users with greater flexibility, and the system is compatible with multiple database types and execution environments.
[0135] Based on the embodiments described above, this application also provides a data processing apparatus, see [link to previous document]. Figure 4 The data processing device includes: The first acquisition unit 401 is used to acquire one or more structured query statement SQL files; The first determining unit 402 is configured to determine first information based on the naming rules and / or storage path of each SQL file in the one or more SQL files. The first information includes one or more of the following: the execution environment corresponding to each SQL file, and the target data center corresponding to each SQL file. The second determining unit 403 is used to determine the execution plan corresponding to each SQL file based on the first information; The first distribution unit 404 is configured to distribute execution tasks according to the execution plan, wherein the execution task represents the execution of the corresponding SQL file in the one or more SQL files.
[0136] In one embodiment, the execution plan includes one or more of the following: The execution order of the corresponding SQL files; The dependencies between the corresponding SQL file and other SQL files; The corresponding SQL file corresponds to the execution environment; The execution time of the corresponding SQL file.
[0137] In one embodiment, the second determining unit 403 determines the execution plan corresponding to each SQL file based on the first information, including: The execution environment corresponding to each SQL file mentioned in the first information is used as the execution environment corresponding to that SQL file in the execution plan; Based on the launch date and / or the readiness of the execution environment corresponding to each SQL file, determine the execution time and execution order of the SQL file.
[0138] In one embodiment, the execution environment corresponding to each SQL file is an execution environment corresponding to that SQL file within a plurality of first execution environments. The plurality of first execution environments includes one or more of the following: a research and development environment; a testing environment; a live network verification environment; a live network environment; and / or, The target data center corresponding to each SQL file is the data center of multiple provinces and the data center corresponding to that SQL file.
[0139] In one embodiment, the apparatus further includes: a first processing unit, wherein after the first distribution unit 404 distributes the execution tasks according to the execution plan, the first processing unit is configured to: If the execution of the first SQL file in one or more SQL files fails, and the execution of the first SQL file fails again, a rollback mechanism is triggered, and the reason for the failure of the first SQL execution is analyzed.
[0140] In one embodiment, the apparatus further includes a second processing unit, wherein after the first acquisition unit 401 acquires one or more structured query statement (SQL) files, the second processing unit is configured to: Each SQL file is processed by the first model to obtain a first prediction result, which includes one or more of the following: the probability that the SQL file fails to execute; the probability that the execution time of the SQL file exceeds a first threshold. If the first prediction result satisfies the first condition, a first message is sent; wherein the first condition includes one or more of the following: the probability of the SQL file failing to execute is greater than a second threshold, and the probability of the SQL file's execution time exceeding the first threshold is greater than a third threshold; the first message includes one or more of the following: warning information, which is used to indicate that there is a risk in the execution of the SQL file; and suggestions for the execution process of the SQL file.
[0141] In one embodiment, the first model is trained using historical SQL execution data, which includes one or more of the following: the complete text of one or more historical SQL files; the execution result of each of the one or more historical SQL files; the execution environment context information of each historical SQL file; and the execution performance data of each historical SQL file.
[0142] In practical applications, the first acquisition unit 401, the first determination unit 402, the second determination unit 403, the first distribution unit 404, the first processing unit, and the second processing unit can be implemented by the processor in the communication device.
[0143] It should be noted that the data processing apparatus provided in the above embodiments is only illustrated by the division of the above program modules. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the apparatus can be divided into different program modules to complete all or part of the processing described above. In addition, the data processing apparatus and data processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0144] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiments of this application, this application also provides an electronic device, see [link to relevant documentation]. Figure 5 The electronic device includes: a first communication interface 1, a first processor 2, a first memory 3, and a bus system 4.
[0145] The first processor 2 is configured to acquire one or more structured query statement (SQL) files; determine first information based on the naming rules and / or storage path of each SQL file in the one or more SQL files, the first information including one or more of the following: the execution environment corresponding to each SQL file, the target data center corresponding to each SQL file; determine the execution plan corresponding to each SQL file based on the first information; and distribute execution tasks according to the execution plan, wherein the execution task represents the execution of the corresponding SQL file in the one or more SQL files.
[0146] In one embodiment, the execution plan includes one or more of the following: The execution order of the corresponding SQL files; The dependencies between the corresponding SQL file and other SQL files; The corresponding SQL file corresponds to the execution environment; The execution time of the corresponding SQL file.
[0147] In one embodiment, the first processor 2 determines the execution plan corresponding to each SQL file based on the first information, including: The execution environment corresponding to each SQL file mentioned in the first information is used as the execution environment corresponding to that SQL file in the execution plan; Based on the launch date and / or the readiness of the execution environment corresponding to each SQL file, determine the execution time and execution order of the SQL file.
[0148] In one embodiment, the execution environment corresponding to each SQL file is an execution environment corresponding to that SQL file within a plurality of first execution environments. The plurality of first execution environments includes one or more of the following: a research and development environment; a testing environment; a live network verification environment; a live network environment; and / or, The target data center corresponding to each SQL file is the data center of multiple provinces and the data center corresponding to that SQL file.
[0149] In one embodiment, after distributing the execution tasks according to the execution plan, the first processor 2 is further configured to: If the execution of the first SQL file in one or more SQL files fails, and the execution of the first SQL file fails again, a rollback mechanism is triggered, and the reason for the failure of the first SQL execution is analyzed.
[0150] In one embodiment, after obtaining one or more structured query statement (SQL) files, the first processor 2 is configured to: Each SQL file is processed by the first model to obtain a first prediction result, which includes one or more of the following: the probability that the SQL file fails to execute; the probability that the execution time of the SQL file exceeds a first threshold. If the first prediction result satisfies the first condition, a first message is sent; wherein the first condition includes one or more of the following: the probability of the SQL file failing to execute is greater than a second threshold, and the probability of the SQL file's execution time exceeding the first threshold is greater than a third threshold; the first message includes one or more of the following: warning information, which is used to indicate that there is a risk in the execution of the SQL file; and suggestions for the execution process of the SQL file.
[0151] In one embodiment, the first model is trained using historical SQL execution data, which includes one or more of the following: the complete text of one or more historical SQL files; the execution result of each of the one or more historical SQL files; the execution environment context information of each historical SQL file; and the execution performance data of each historical SQL file.
[0152] It should be noted that the specific processing procedure of the first communication interface 1 can be understood by referring to the above method.
[0153] Of course, in practical applications, the various components in an electronic device are coupled together through bus system 4. It can be understood that bus system 4 is used to achieve communication and connection between these components. In addition to the data bus, bus system 4 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 5 The general will label all buses as Bus System 4.
[0154] The first memory 3 in this embodiment is used to store various types of data to support operation in the electronic device. Examples of such data include any computer program used to operate on the electronic device.
[0155] The methods disclosed in the embodiments of this application can be applied to the first processor 2, or implemented by the first processor 2. The first processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in the first processor 2. The first processor 2 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The first processor 2 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the first memory 3. The first processor 2 reads the information in the first memory 3 and completes the steps of the aforementioned method in combination with its hardware.
[0156] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned methods.
[0157] It is understood that the first memory 3 in this application embodiment can be a volatile memory pool or a non-volatile memory pool, or both. The non-volatile memory pool can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory pool can be a disk storage pool or a magnetic tape storage pool. The volatile memory pool can be a random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The storage pools described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of storage pools.
[0158] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a first memory 3 storing a computer program, which can be executed by a first processor 2 to complete the steps described in the aforementioned method.
[0159] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by a first processor 2 to perform the steps described in the foregoing method.
[0160] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.
Claims
1. A data processing method, characterized by, The data processing method includes: Retrieve one or more structured query statement SQL files; Based on the naming rules and / or storage path of each of the one or more SQL files, first information is determined, which includes one or more of the following: the execution environment corresponding to each SQL file, and the target data center corresponding to each SQL file. Based on the first information, determine the execution plan corresponding to each SQL file; Execution tasks are distributed according to the execution plan, wherein the execution tasks represent the execution of the corresponding SQL files in the one or more SQL files.
2. The method of claim 1, wherein, The execution plan includes one or more of the following: The execution order of the corresponding SQL files; The dependencies between the corresponding SQL file and other SQL files; The corresponding SQL file corresponds to the execution environment; The execution time of the corresponding SQL file.
3. The method of claim 2, wherein, The step of determining the execution plan corresponding to each SQL file based on the first information includes: The execution environment corresponding to each SQL file mentioned in the first information is used as the execution environment corresponding to that SQL file in the execution plan; Based on the launch date and / or the readiness of the execution environment corresponding to each SQL file, determine the execution time and execution order of the SQL file.
4. The method of claim 1, wherein, Each SQL file corresponds to an execution environment that is an execution environment within a plurality of first execution environments. These plurality of first execution environments include one or more of the following: a development environment; a testing environment; a live network verification environment; a live network environment; and / or, The target data center corresponding to each SQL file is the data center of multiple provinces and the data center corresponding to that SQL file.
5. The method of claim 1, wherein, After distributing the execution tasks according to the execution plan, the method further includes: If the execution of the first SQL file in one or more SQL files fails, and the execution of the first SQL file fails again, a rollback mechanism is triggered, and the reason for the failure of the first SQL execution is analyzed.
6. The method of claim 1, wherein, After obtaining one or more structured query statement SQL files, the method further includes: Each SQL file is processed by the first model to obtain a first prediction result, which includes one or more of the following: the probability that the SQL file fails to execute; the probability that the execution time of the SQL file exceeds a first threshold. If the first prediction result satisfies the first condition, a first message is sent; wherein the first condition includes one or more of the following: the probability of the SQL file failing to execute is greater than a second threshold, and the probability of the SQL file's execution time exceeding the first threshold is greater than a third threshold; the first message includes one or more of the following: warning information, which is used to indicate that there is a risk in the execution of the SQL file; and suggestions for the execution process of the SQL file.
7. The method of claim 6, wherein, The first model is trained using historical SQL execution data, which includes one or more of the following: the complete text of one or more historical SQL files; the execution result of each historical SQL file; the execution environment context information of each historical SQL file; and the execution performance data of each historical SQL file.
8. An electronic device, comprising: include: The processor and the memory used to store computer programs that can run on the processor. When the processor is used to run the computer program, it performs the steps of the method according to any one of claims 1 to 7.
9. A storage medium having stored thereon a computer program, characterized in that When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.