Automatic data comparison method and device based on jenkins and medium

Through the jenkins tool and automated python scripts, efficient and accurate data comparison between new and old systems or frameworks is achieved, solving the problems of inefficiency and error-prone in the existing technology, and improving the automation and security of data comparison.

CN120277050APending Publication Date: 2025-07-08AULTON NEW ENERGY AUTOMOBILE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411684261.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-11-22
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

During the process of database migration and reporting system computing framework replacement, existing manual SQL comparison methods are inefficient and prone to errors, making it difficult to efficiently automate data comparison between new and old systems or frameworks.

Method used

The automated data comparison method based on jenkins is used to build automated comparison python scripts and use jenkins tools to compare data comparison between new and old systems or frameworks, including encapsulating database connection methods, writing public method classes, configuring SQL files and placing data-driven files, combining jenkins' task management and permission control.

Benefits of technology

Improves the efficiency and accuracy of data comparison, reduces human errors, ensures data consistency and security, supports flexible expansion and management, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277050A_ABST
    Figure CN120277050A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic data comparison method and device based on jenkins and a medium, and relates to the technical field of automatic data processing. The method comprises the following steps: constructing an automatic comparison python script based on a data comparison task demand between databases to be compared, and installing the automatic comparison python script in a comparison server; wherein a to-be-applied jenkins tool is preset in the comparison server; constructing an automation task based on the IP address of the comparison server and the open port number of the jenkins tool to be applied; and calling an automatic comparison python script through the to-be-applied jenkins tool based on the target triggering condition of the automatic task so as to realize a data comparison task between the to-be-compared databases. By means of the method, efficient and automatic comparison of data between new and old systems or frames is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority based on the invention patent application filed with the China National Patent Office on December 29, 2023, with the application number 2023118678812 and the invention title "An Automated Data Comparison Method, Device, and Medium Based on Jenkins". This application incorporates the entire text of the above-mentioned Chinese patent application by reference. Technical Field

[0002] This application relates to the technical field of automated data processing, and particularly to an automated data comparison method, device, and medium based on Jenkins. Background Art

[0003] During the migration of a database scheduling system and the replacement of a report system calculation framework, it is very important to ensure data consistency between the new system and the old system. To verify the data accuracy of the new system, it is usually necessary to compare the data of the new system with that of the old system. Similarly, a comparison is also needed between the calculation results of the new framework and those of the old framework.

[0004] However, this comparison workload is huge, especially when the number of tables involved is large. Each migration or framework replacement may involve data comparison of hundreds of tables, and this comparison needs to be carried out over multiple days and multiple dimensions to ensure data integrity and accuracy.

[0005] The traditional manual SQL comparison method is inefficient in this case because it requires manual writing and execution of a large number of SQL queries, which is not only time-consuming but also error-prone. In addition, for a large amount of data comparison, this method may be limited by the database performance. Therefore, how to efficiently and automatically compare the data between the new and old systems or frameworks has become a technical problem to be solved urgently. Summary of the Invention

[0006] Embodiments of this application provide an automated data comparison method, device, and medium based on Jenkins to solve the following technical problem: how to efficiently and automatically compare the data between the new and old systems or frameworks.

[0007] In a first aspect, an embodiment of the present application provides an automated data comparison method based on Jenkins. The method is characterized in that it includes: constructing an automated comparison Python script based on the data comparison task requirements between databases to be compared, and installing the automated comparison Python script on a comparison server; wherein, the Jenkins tool is pre-installed in the comparison server; constructing an automated task based on the IP address of the comparison server and the open port number of the Jenkins tool to be applied; and calling the automated comparison Python script through the Jenkins tool to be applied based on the target trigger condition of the automated task, so as to implement the data comparison task between the databases to be compared.

[0008] The automated data comparison method based on Jenkins provided by the embodiment of the present application can quickly compare the data between different databases through an automated script and the Jenkins tool, and can be run regularly or on demand, greatly improving the efficiency and accuracy of data comparison. Since the data comparison is automated, errors caused by human operations can be reduced, thereby reducing the error rate. Since the data comparison is based on a script and the Jenkins tool, it can be run repeatedly at any time as long as the script is okay, which is very useful for situations where data comparison needs to be performed regularly. Through the automated script and the Jenkins tool, the links of manual participation in data comparison can be reduced, thereby saving labor costs. The Jenkins tool usually provides powerful monitoring and reporting functions, which enables us to easily understand the progress, results and abnormal situations of data comparison. The automated method based on Jenkins can flexibly adapt to different data comparison requirements, and only the corresponding script needs to be modified. In addition, if more databases or data sources need to be compared, only the corresponding script needs to be added for expansion. By executing the data comparison task on the server, the security and privacy of the data can be better protected. At the same time, the access rights to the data comparison task can also be controlled through the permission management function provided by the Jenkins tool.

[0009] In an implementation manner of the present application, constructing an automated comparison Python script based on the data comparison task requirements between databases to be compared specifically includes: encapsulating the connection method corresponding to the databases to be compared; writing a common method class for requirements; configuring a requirements SQL file; placing a data-driven file to be applied; and writing data comparison rules and comparison result processing rules.

[0010] In the embodiments of the present application, by encapsulating the connection method corresponding to the database to be compared, the steps of database connection can be simplified, and the reusability and maintainability of the code can be improved. At the same time, encapsulating the connection method can also better manage the details of database connection and reduce the possibility of database connection errors. Writing a common method class for requirements can provide a set of common methods, which can be used to handle common tasks in data comparison tasks, such as reading configuration files and handling exceptions. This can make the data comparison script more modular and extensible. By configuring the requirement SQL file, the SQL query statement can be separated from the data comparison task. This can improve the readability and maintainability of the code. At the same time, it is also convenient to modify the SQL query statement without modifying the data comparison script. Placing the data-driven file to be applied can facilitate the management of data-driven files. At the same time, it can also make the data comparison script more modular and extensible, so that the data-driven files can be reused in different data comparison tasks, improving the reusability of the code. By writing data comparison rules and comparison result processing rules, the task requirements of data comparison can be defined more clearly. At the same time, it can also make the data comparison script more flexible and extensible, so that different rules can be written according to different requirements to meet the needs of different data comparison tasks.

[0011] In one implementation of the present application, encapsulating the connection method corresponding to the database to be compared specifically includes: determining the database type of the database to be compared, and determining the connection parameters of the database to be compared according to the database type; based on the database type and connection parameters, determining the python connection module to be applied; encapsulating the operation instructions corresponding to the python connection module to be applied and the database to be compared; where the operation instructions include but are not limited to at least one of the following: an instruction to execute a query statement, an instruction to obtain a query result.

[0012] In the embodiments of the present application, by encapsulating the connection method of the database to be compared, the connection processes of different databases can be uniformly managed, and the repeated writing of connection code in each data comparison script can be avoided. Encapsulating the connection method can simplify the steps of database connection, so that only the connection method needs to be called in the data comparison script without caring about the specific connection details. By encapsulating the connection method, these methods can be reused in different data comparison scripts, improving the reusability of the code. When a new database type needs to be connected, only the corresponding connection method needs to be modified without modifying other parts of the data comparison script, which is convenient for maintenance and extension. Encapsulating the connection method can make the data comparison script more flexible, and different database types and connection parameters can be selected according to actual needs. When encapsulating the connection method, the logic of error handling and logging can be added to better monitor and manage the database connection process. By encapsulating the connection method, the access rights to the database can be better controlled, thus enhancing the security of data.

[0013] In an implementation manner of the present application, a public method class to be demanded is written, specifically including: based on the requirements of the data comparison task, determining the function type of the public method class to be demanded, and based on the function type, determining the corresponding function parameters; wherein, the public method class to be demanded includes but is not limited to at least one of the following: a data reading method class, a data format processing method class; determining the corresponding Python library for implementing the public method class to be demanded, and based on the function parameters and the Python library, writing the public method class to be demanded.

[0014] By writing the public method class to be demanded in the embodiments of the present application, the public functions in the data comparison task can be modularly designed, making the code clearer, easier to understand and maintain. The functions in the public method class can be reused by multiple data comparison tasks, thereby reducing the repeated writing of code and improving the code reusability. The public method class can provide a unified processing logic to ensure the consistency and accuracy of data reading and format processing. This can avoid repeating the same logic in each data comparison task. When new functions or processing logics need to be added, only the corresponding public method class needs to be modified without modifying the code in other parts. This can facilitate expansion and upgrade. Through modular design and unified processing logic, the errors and inconsistencies caused by repeated code writing can be reduced, and it is also easier to perform error troubleshooting and repair. The public method class usually has clear interfaces and documentation, making it easier for other developers to understand and use these methods, which can reduce the maintenance cost and improve the development efficiency.

[0015] In an implementation manner of the present application, a data-driven file to be applied is placed, specifically including: based on the requirements of the data comparison task, determining the type of the data-driven file to be applied, and writing the content of the data-driven file to generate the data-driven file to be applied; placing the data-driven file to be applied in a specified directory or file.

[0016] In the embodiments of the present application, by placing the data-driven files to be applied, the data-driven files can be uniformly managed, avoiding the repeated creation and placement of the driver files in each data comparison script. Placing the driver files in a specified directory or file can facilitate the access to these driver files in the data comparison script, improving the readability and maintainability of the code. By placing the driver files, these driver files can be reused in different data comparison scripts, improving the code reusability. When new data drivers need to be added or existing drivers need to be upgraded, only the corresponding driver files need to be modified, without modifying the code in other parts, which can facilitate the expansion and upgrade. By placing the driver files in a specified directory or file, the access rights to the data can be better controlled, thereby enhancing the data security. Placing the driver files in a specific directory or file can make the backup and recovery operations easier to prevent data loss or damage.

[0017] In one implementation manner of the present application, data comparison rules and comparison result processing rules are written, specifically including: based on the requirements of the data comparison task, writing data comparison rules; based on the data type of the data to be compared, determining the corresponding comparison difference threshold, and based on the comparison difference threshold, writing comparison result processing rules.

[0018] In the embodiments of the present application, by writing data comparison rules, the goals and scopes of data comparison can be clarified, ensuring the accuracy and effectiveness of data comparison. By determining the corresponding comparison difference threshold, the differences between data can be quickly identified, improving the efficiency of data comparison. By writing comparison result processing rules, the results of data comparison can be uniformly processed, ensuring the accuracy and consistency of the data. When the data comparison rules need to be modified or expanded, only the corresponding rules need to be modified, without modifying the code in other parts, which can facilitate the maintenance and expansion. By writing data comparison rules and comparison result processing rules, the accuracy and consistency of the data can be ensured, thereby improving the data quality. By writing comparison result processing rules, corresponding monitoring reports can be generated to better understand the progress and results of data comparison.

[0019] In one implementation manner of the present application, a requirements SQL file is configured, specifically including: based on the requirements of the data comparison task, determining the data comparison requirements: among them, the data comparison requirements include: comparison indicators, comparison dimensions; based on the data comparison requirements, writing the corresponding SQL query statements and storing the SQL query statements in the corresponding requirements SQL file.

[0020] In the embodiments of the present application, by configuring the requirements SQL file, the requirements of data comparison can be clarified, including comparison indicators and comparison dimensions. This can ensure the accuracy and effectiveness of data comparison.

[0021] Improve efficiency: By writing the corresponding SQL query statements and storing the SQL query statements in the corresponding required SQL files, the data to be compared can be quickly obtained, improving the efficiency of data comparison. By configuring the required SQL files, the requirements for data comparison and the corresponding SQL query statements can be uniformly managed, avoiding repeated writing and storage of SQL query statements in each data comparison task. When the same data comparison task needs to be executed, the configured required SQL files can be directly used without having to rewrite and store the SQL query statements again, which can improve the code reusability. When the data comparison requirements need to be modified or extended, only the corresponding SQL query statements need to be modified and the required SQL files need to be reconfigured, without having to modify other parts of the code, which can facilitate maintenance and extension. By configuring the required SQL files, corresponding monitoring reports can be generated to better understand the progress and results of data comparison.

[0022] In an implementation manner of the present application, after the automated comparison python script is installed on the comparison server, the method further includes: determining the read permission of the automated comparison python script and authorizing the jenkins tool to be applied.

[0023] By ensuring the read permission of the automated comparison python script in the embodiments of the present application, the security of the script and data can be better protected. Only authorized users or roles can access and execute the script, thus preventing unauthorized access and data leakage. By authorizing the jenkins tool to be applied, the access permission of the automated comparison python script can be more finely managed. Different read permissions can be granted according to the needs of different users or roles, thus better controlling the access to data. In some industries or organizations, there may be compliance and auditing requirements that require strict control over the access permission of the automated comparison python script. By determining the read permission of the automated comparison python script and authorizing it, these requirements can be met and compliance can be ensured. By authorizing the jenkins tool to be applied, the addition, modification, or deletion of permissions can be conveniently performed. When new users or roles need to be added, simple configuration can be performed in the jenkins tool, which can improve the maintainability and scalability of the system.

[0024] In an implementation manner of the present application, an automated task is constructed based on the IP address of the comparison server and the open port number of the jenkins tool to be applied, specifically including: in the jenkins tool to be applied, a new initialization automated task is created; the source code management address of the initialization automated task is configured; wherein, the source code management address is associated with an automated comparison python script; in the automated task, a build trigger is configured, and a target trigger condition corresponding to the build trigger is set; the task execution node of the automated task is configured, and execution parameters corresponding to the task execution node are set; wherein, the execution parameters include but are not limited to at least one of the following: the IP address of the comparison server, the open port number of the jenkins tool to be applied, and the execution path of the automated comparison python script.

[0025] In an implementation manner of the present application, the automated comparison python script is called through the jenkins tool to be applied to implement the data comparison task between the databases to be compared, specifically including: according to the target trigger condition, the automated task is automatically triggered, so that the jenkins tool pulls the automated comparison python script according to the configured source code management address; the automated comparison python script is executed, and the data between the databases to be compared is compared according to the data comparison rules and comparison result processing rules written in the automated comparison python script, and a comparison result is generated.

[0026] In a second aspect, an automated data comparison device based on jenkins is further provided in an embodiment of the present application, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: construct an automated comparison python script based on the data comparison task requirements between the databases to be compared, and install the automated comparison python script on the comparison server; wherein, the comparison server is preconfigured with the jenkins tool to be applied; construct an automated task based on the IP address of the comparison server and the open port number of the jenkins tool to be applied; based on the target trigger condition of the automated task, call the automated comparison python script through the jenkins tool to be applied to implement the data comparison task between the databases to be compared.

[0027] In a third aspect, an embodiment of the present application also provides a non-volatile computer storage medium for automated data comparison based on Jenkins, storing computer-executable instructions, characterized in that the computer-executable instructions are set as follows: Based on the data comparison task requirements between the databases to be compared, construct an automated comparison Python script and install the automated comparison Python script on the comparison server; wherein, the Jenkins tool is pre-set in the comparison server; Based on the IP address of the comparison server and the open port number of the Jenkins tool to be applied, construct an automated task; Based on the target trigger condition of the automated task, call the automated comparison Python script through the Jenkins tool to be applied to implement the data comparison task between the databases to be compared.

[0028] An automated data comparison method, device and medium based on Jenkins provided by an embodiment of the present application have the following beneficial effects:

[0029] Automation and efficiency improvement: Through the automated data comparison method based on Jenkins, the data comparison task can be automatically executed, avoiding the cumbersome processes of manual comparison and manual analysis, and greatly improving the efficiency of data comparison.

[0030] Consistency and accuracy: By writing and configuring the automated comparison Python script, the consistency and accuracy of data comparison can be ensured. The execution of the script is not affected by human factors, thus avoiding the possibility of human errors.

[0031] Flexibility and scalability: This method can be customized and extended according to different data comparison task requirements. For example, corresponding scripts and rules can be written and configured according to different database types, data types, comparison metrics, etc.

[0032] Security and compliance: By determining the read permissions of the automated comparison Python script and using the Jenkins tool for authorization, the security and compliance of data can be ensured. Only authorized users or roles can access and execute the script, thus preventing unauthorized access and data leakage.

[0033] Maintainability and manageability: Using the Jenkins tool for authorization and automated task management improves the maintainability and manageability of the system. The execution status, results and logs of automated tasks can be easily viewed and managed.

[0034] Improving the work process: This method can be combined with the existing work process to improve and optimize the work process of data comparison and analysis. For example, the automated data comparison task can be integrated into the daily data analysis and report generation work to improve work efficiency and quality. Description of the Drawings

[0035] The drawings described herein are provided to further understand the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0036] Figure 1 It is a flowchart of an automated data comparison method based on Jenkins provided by an embodiment of the present application;

[0037] Figure 2 It is another flowchart of an automated data comparison method based on Jenkins provided by an embodiment of the present application;

[0038] Figure 3 It is a schematic diagram of the internal structure of an automated data comparison device based on Jenkins provided by an embodiment of the present application. Detailed Embodiments

[0039] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0040] The embodiments of the present application provide an automated data comparison method, device, and medium based on Jenkins to solve the following technical problem: how to efficiently and automatically compare the data between the old and new systems or frameworks.

[0041] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the drawings.

[0042] Figure 1 It is a flowchart of an automated data comparison method based on Jenkins provided by an embodiment of the present application. As Figure 1 shown, an automated data comparison method based on Jenkins provided by an embodiment of the present application specifically includes the following steps:

[0043] Step 101: Based on the data comparison task requirements between the databases to be compared, construct an automated comparison Python script and install the automated comparison Python script on the comparison server.

[0044] In an embodiment of the present application, to implement the automated data comparison based on Jenkins, first, the automated data comparison based on Jenkins.

[0045] Specifically, encapsulate the connection method corresponding to the database to be compared; write the common method class for the requirements; configure the requirement SQL file; place the data-driven file to be applied; write the data comparison rules and the comparison result processing rules.

[0046] In an embodiment of the present application, encapsulating the connection method corresponding to the database to be compared specifically includes: determining the database type of the database to be compared, and based on the database type, determining the connection parameters of the database to be compared; based on the database type and the connection parameters, determining the Python connection module to be applied; encapsulating the operation instructions corresponding to the Python connection module to be applied and the database to be compared; where the operation instructions include but are not limited to at least one of the following: execute query statement instruction, obtain query result instruction.

[0047] Further elaboration: Determine the database type of the database to be compared: Determine whether the database to be compared is a relational database (such as MySQL, PostgreSQL, etc.) or a non-relational database (such as MongoDB, Redis, etc.). Based on the database type, determine the connection parameters of the database to be compared: For a relational database, the hostname, port number, username, password of the database, and the name of the database to be connected need to be provided; for a non-relational database, usually the hostname, port number of the database, and the name of the database or collection to be connected need to be provided. Based on the database type and the connection parameters, determine the Python connection module to be applied: For a relational database, modules such as pymysql and psycopg2 in Python can be used for connection; for a non-relational database, modules such as pymongo and redis-py in Python can be used for connection. Encapsulate the operation instructions corresponding to the Python connection module to be applied and the database to be compared: Execute query statement instruction: Use the method provided by the Python connection module to be applied to send the SQL query statement to the database for execution. Obtain query result instruction: Use the method provided by the Python connection module to be applied to obtain the query result and return it.

[0048] In an embodiment of the present application, writing the common method class for the requirements to be met specifically includes: based on the requirements of the data comparison task, determining the function type of the common method class for the requirements to be met, and based on the function type, determining the corresponding function parameters; where the common method class for the requirements to be met includes but is not limited to at least one of the following: data reading method class, data format processing method class; determining the Python library corresponding to the common method class for the requirements to be met, and based on the function parameters and the Python library, writing the common method class for the requirements to be met.

[0049] Further details: For writing the common method class for requirements: First, we need to determine the function types and corresponding function parameters of the common method class for requirements based on the data comparison task requirements. In this example, we assume that the data comparison task requirements include reading data, formatting data, and comparing data. Next, we can create two common method classes: the data reading method class and the data format processing method class. The data reading method class is used to read data from different databases or data sources, while the data format processing method class is used to format the read data for data comparison. Then, we need to determine the Python libraries required to implement these common method classes. For example, we can use the pandas library to read and process data, the numpy library for numerical calculations and statistical analysis, the matplotlib library for data visualization, etc. Finally, we can write the common method class based on the function parameters and the selected Python libraries.

[0050] In an embodiment of the present application, configure the requirements SQL file, specifically including: Based on the data comparison task requirements, determine the data comparison requirements: Among them, the data comparison requirements include but are not limited to at least one of the following: comparison metrics, comparison dimensions; Based on the data comparison requirements, write the corresponding SQL query statement and store the SQL query statement in the corresponding requirements SQL file.

[0051] Further details: For configuring the requirements SQL file: First, we need to clarify the requirements of the data comparison task, including comparison metrics and comparison dimensions. For example, we may need to compare sales data in two different databases. The comparison metrics can be sales amount, sales volume, etc., and the comparison dimensions can be time, product category, etc. Next, we need to write the corresponding SQL query statement to extract the required data from the database. The following is an example SQL query statement:

[0052] SELECT product_category,SUM(sales_amount)AS total_sales

[0053] FROM sales_data

[0054] WHERE date>='2022-01-01'AND date<='2022-12-31'

[0055] GROUP BY product_category;

[0056] The above SQL query statement will extract product category and sales amount data from the sales_data table and group them by product category. Finally, we will store the above SQL query statement in the corresponding required SQL file. The following is the content of an example required SQL file:

[0057] -- Compare sales data

[0058] SELECT product_category,SUM(sales_amount)AS total_sales

[0059] FROM sales_data_db1

[0060] WHERE date>='2022-01-01'AND date<='2022-12-31'

[0061] GROUP BY product_category;

[0062] -- Compare sales data

[0063] SELECT product_category,SUM(sales_amount)AS total_sales

[0064] FROM sales_data_db2

[0065] WHERE date>='2022-01-01'AND date<='2022-12-31'

[0066] GROUP BY product_category;

[0067] The above required SQL file contains two SQL query statements, which extract sales data from two different databases respectively and make comparisons. By configuring the required SQL file, we can conveniently execute the data comparison task and extract the required data for further analysis and processing.

[0068] In an embodiment of the present application, placing the data-driven file to be applied specifically includes: determining the type of the data-driven file to be applied based on the requirements of the data comparison task, and writing the content of the data-driven file to generate the data-driven file to be applied; placing the data-driven file to be applied in a specified directory or file.

[0069] Further details: For placing the data-driven file to be applied: First, we need to determine the type of the driver file for the data-driven file to be applied based on the requirements of the data comparison task. For example, if the data comparison task requires using a Python script for data reading and processing, then we can choose Python as the driver file type. Next, we need to write the content of the driver file to generate the data-driven file to be applied. The following is the content of the driver file for an example Python script:

[0070] import pandas as pd

[0071] # Read data from the database

[0072] data = pd.read_sql('SELECT*FROM sales_data', source)

[0073] # Process and format the data

[0074] formatted_data = data.fillna(0).astype(np.float64)

[0075] # Save the processed data to a specified file

[0076] formatted_data.to_csv('formatted_sales_data.csv', index = False)

[0077] The above driver file content uses the pandas library in Python to read sales data from the database, processes and formats it, and finally saves the processed data to a specified CSV file. Then, we save the above driver file content as a Python script file, such as data_driver.py. Finally, we place the data-driven file data_driver.py to be applied in a specified directory or file. For example, we can place it in the root directory of the project or upload it to a cloud storage platform. By placing the data-driven file to be applied, we can conveniently execute the data comparison task and use the code in the driver file to read and process the data.

[0078] In an embodiment of the present application, data comparison rules and comparison result processing rules are written, specifically including: writing data comparison rules based on the requirements of the data comparison task; determining the corresponding comparison difference threshold based on the data type of the data to be compared, and writing comparison result processing rules based on the comparison difference threshold.

[0079] Further details: For example: Requirements of the data comparison task: An e-commerce platform plans to migrate from an old database system to a new one and needs to ensure the accuracy and integrity of order data during the migration process. Therefore, it is necessary to compare the order data between the old and new systems. Analysis of the actual situation: The old system uses a MySQL database, and the new system uses a PostgreSQL database. The order data includes fields such as order number, user ID, product ID, order amount, and order placement time. The amount of data to be compared is large, involving millions of order records.

[0080] The data comparison rules are formulated as follows: Rule 1: Comparison of order number uniqueness: Since the order number is unique, the data consistency can be checked by comparing the order numbers in the old and new systems. If there is an order number in the new system that does not exist in the old system, or an order number in the old system that does not exist in the new system, it is considered that the data is inconsistent. Rule 2: Comparison of key fields: In addition to the order number, it is also necessary to compare key fields such as user ID, product ID, and order amount. The values of these fields should be consistent in the old and new systems. Rule 3: Comparison of order placement time: Since the order placement time is of the timestamp type, direct comparison may be misjudged as inconsistent due to millisecond-level differences. Therefore, the order placement time can be converted to the date format before comparison, and only the date-level accuracy is concerned. Rule 4: Comparison of data volume: Count the total amount of order data in the old and new systems to ensure that the quantities are the same. If the quantities are inconsistent, it means that there may be data loss or duplicate insertion during the migration process.

[0081] The rules for processing comparison results are formulated as follows: Rule 1: Recording of discrepant data: For the discrepant data identified according to the above rules, the details of the discrepancies need to be recorded, including information such as order number, discrepant fields, and specific discrepant values. Rule 2: Visualization of discrepant data: Use charts or reports to display the statistical results of discrepant data, such as the quantity and proportion of discrepant data, to facilitate the project team's quick understanding of the accuracy of data migration. Rule 3: Analysis of reasons for discrepancies: For the recorded discrepant data, it is necessary to further analyze the reasons for the discrepancies. It may be errors during the data migration process, changes in source data, etc. Corresponding handling measures should be taken according to different reasons, such as fixing the migration script, communicating and confirming with the business team, etc. Rule 4: Result feedback and reporting: Feedback the comparison results and the handling of discrepancies to the project team and relevant responsible persons in the form of a report to ensure that the project team has a comprehensive understanding of the accuracy of data migration and decides whether further data repair or verification work is required according to the actual situation.

[0082] In an embodiment of the present application, after constructing the automated comparison Python script, the automated comparison Python script is installed in the comparison server. It should be noted that the comparison server is pre-configured with the to-be-applied Jenkins tool for implementing automated data comparison based on Jenkins.

[0083] In an embodiment of the present application, after installing the automated comparison Python script in the comparison server, the method further includes: determining the read permission of the automated comparison Python script and authorizing it using the to-be-applied Jenkins tool.

[0084] Step 102: Construct an automated task based on the IP address of the comparison server and the open port number of the to-be-applied Jenkins tool.

[0085] In an implementation manner of the present application, constructing an automated task based on the IP address of the comparison server and the open port number of the to-be-applied Jenkins tool specifically includes: creating a new initialization automated task in the to-be-applied Jenkins tool; configuring the source code management address of the initialization automated task; wherein the source code management address is associated with the automated comparison Python script; in the automated task, configuring a build trigger and setting the corresponding target trigger condition for the build trigger; configuring the task execution node of the automated task and setting the corresponding execution parameters for the task execution node; wherein the execution parameters include but are not limited to at least one of the following: the IP address of the comparison server, the open port number of the to-be-applied Jenkins tool, and the execution path of the automated comparison Python script.

[0086] Details are described through the following embodiments: Based on the IP address of the comparison server and the open port number of the Jenkins tool to be applied, an automated task is constructed: Determine the IP address of the comparison server and the open port number of the Jenkins tool to be applied. For example, the IP address of the comparison server is 192.168.0.1, and the open port number of the Jenkins tool to be applied is 8080. Write an automated script to obtain the IP address of the comparison server and the open port number of the Jenkins tool to be applied. Appropriate code can be written using a scripting language (such as Python, Shell, etc.) to obtain the required information from the comparison server or the Jenkins tool. In the automated script, use network tools (such as ping, telnet, etc.) to detect whether the comparison server can access the open port number of the Jenkins tool to be applied. For example, the ping command can be used to detect whether the comparison server can access the IP address of the Jenkins tool, and the telnet command can be used to detect whether the open port number of the Jenkins tool is reachable. According to the detection results, write the corresponding automated task logic. For example, if the comparison server can access the open port number of the Jenkins tool to be applied, perform the corresponding operations (such as triggering a Jenkins build task, sending a notification, etc.). If the comparison server cannot access the open port number of the Jenkins tool to be applied, record the error information and send an alert. Set the corresponding running time and scheduling plan (such as a scheduled task, Cron, etc.). Ensure that the automated task is automatically executed when needed and adjusted and optimized according to actual requirements.

[0087] Step 103: Based on the target trigger condition of the automated task, call the automated comparison Python script through the Jenkins tool to be applied to implement the data comparison task between the databases to be compared.

[0088] In an embodiment of the present application, calling the automated comparison Python script through the Jenkins tool to be applied to implement the data comparison task between the databases to be compared specifically includes: Automatically trigger the automated task according to the preset target trigger condition, so that the Jenkins tool pulls the automated comparison Python script according to the configured source code management address; Execute the automated comparison Python script, and compare the data between the databases to be compared according to the data comparison rules and comparison result processing rules written in the automated comparison Python script, and generate a comparison result.

[0089] Figure 2 Another flowchart of the automated data comparison method based on Jenkins provided by the embodiments of the present application. As Figure 2As shown in the figure, another automated data comparison method based on Jenkins provided by an embodiment of the present application specifically includes the following steps:

[0090] Step 201: Based on the data comparison task requirements between the databases to be compared, construct an automated comparison Python script and install the automated comparison Python script on the comparison server; among them, the Jenkins tool to be applied is pre-installed in the comparison server.

[0091] Step 202: Based on the IP address of the comparison server and the open port number of the Jenkins tool to be applied, construct an automated task.

[0092] Step 203: In the case of needing to automatically compare data, call the automated comparison Python script through the Jenkins tool to be applied to implement the data comparison task between the databases to be compared.

[0093] The above is the method embodiment proposed by the present application. Based on the same inventive concept, an embodiment of the present application also provides an automated data comparison device based on Jenkins, and its structure is as Figure 3 shown.

[0094] Figure 3 It is a schematic diagram of the internal structure of an automated data comparison device based on Jenkins provided by an embodiment of the present application. As Figure 3 shown, the device includes:

[0095] At least one processor 301;

[0096] And a memory 302 communicatively connected to at least one processor;

[0097] Among them, the memory 302 stores instructions executable by at least one processor, and the instructions are executed by at least one processor 301 so that at least one processor 301 can:

[0098] Based on the data comparison task requirements between the databases to be compared, construct an automated comparison Python script and install the automated comparison Python script on the comparison server; among them, the Jenkins tool to be applied is pre-installed in the comparison server.

[0099] Based on the IP address of the comparison server and the open port number of the Jenkins tool to be applied, construct an automated task.

[0100] Based on the target trigger condition of the automated task, call the automated comparison Python script through the Jenkins tool to be applied to implement the data comparison task between the databases to be compared.

[0101] Some embodiments of the present application provide a non-volatile computer storage medium for automated data comparison based on Jenkins, storing computer-executable instructions, and the computer-executable instructions are set as follows: Figure 1 Based on the data comparison task requirements between the databases to be compared, construct an automated comparison Python script and install the automated comparison Python script on the comparison server; wherein, the Jenkins tool is pre-installed in the comparison server.

[0102] Based on the IP address of the comparison server and the open port number of the Jenkins tool to be applied, construct an automated task.

[0103] Based on the target trigger condition of the automated task, call the automated comparison Python script through the Jenkins tool to be applied to implement the data comparison task between the databases to be compared.

[0104] Each embodiment in the present application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0105] The systems and media provided by the embodiments of the present application correspond one by one to the methods. Therefore, the systems and media also have beneficial technical effects similar to those of the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be elaborated here.

[0106] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0107] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Thus, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Also, the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0108] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the specified functions in a block or multiple blocks.

[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the specified functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or a block or multiple blocks.

[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or a block or multiple blocks.

[0111] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0112] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0113] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0114] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0115] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An automated data comparison method based on Jenkins, characterized in that, The method includes: Based on the data comparison task requirements between databases to be compared, construct an automated comparison Python script and install the automated comparison Python script on the comparison server; wherein, the Jenkins tool is pre-installed in the comparison server. Based on the IP address of the comparison server and the open port number of the Jenkins tool to be used, construct an automated task. Based on the target trigger condition of the automated task, call the automated comparison Python script through the Jenkins tool to be used to implement the data comparison task between the databases to be compared.

2. The automated data comparison method based on Jenkins according to claim 1, wherein, Based on the data comparison task requirements between databases to be compared, constructing an automated comparison Python script specifically includes: Encapsulate the connection method corresponding to the database to be compared. Write a common method class for the requirements. Configure the requirements SQL file. Place the data-driven file to be used. Write the data comparison rules and the comparison result processing rules.

3. The automated data comparison method based on Jenkins according to claim 2, wherein Encapsulating the connection method corresponding to the database to be compared specifically includes: Determine the database type of the database to be compared, and based on the database type, determine the connection parameters of the database to be compared. Based on the database type and the connection parameters, determine the Python connection module to be used. Encapsulate the operation instructions corresponding to the Python connection module to be used and the database to be compared; wherein, the operation instructions include but are not limited to at least one of the following: execute query statement instruction, obtain query result instruction.

4. An automated data comparison method based on Jenkins according to claim 2, characterized in that Writing a common method class for the requirements specifically includes: Based on the data comparison task requirements, determine the function type of the common method class for the requirements, and based on the function type, determine the corresponding function parameters; wherein, the common method class for the requirements includes but is not limited to at least one of the following: data reading method class, data format processing method class. Determine the Python library corresponding to implementing the common method class for the requirements, and based on the function parameters and the Python library, write the common method class for the requirements.

5. The automated data comparison method based on Jenkins according to claim 2, characterized in that, Placing the data-driven file to be used specifically includes: Based on the data comparison task requirements, determine the type of the data-driven file to be used, and write the content of the driver file to generate the data-driven file to be used. Place the data-driven file to be used in a specified directory or file.

6. The automated data comparison method based on Jenkins according to claim 2, wherein Writing the data comparison rules and the comparison result processing rules specifically includes: Based on the data comparison task requirements, write the data comparison rules. Based on the data type of the data to be compared, determine the corresponding comparison difference threshold, and based on the comparison difference threshold, write the comparison result processing rules. And / or Configuring the requirements SQL file specifically includes: Based on the data comparison task requirements, determine the data comparison requirements: wherein, the data comparison requirements include but are not limited to at least one of the following: comparison indicators, comparison dimensions. Based on the data comparison requirements, write the corresponding SQL query statement and store the SQL query statement in the corresponding requirements SQL file.

7. A method for automated data comparison based on Jenkins according to claim 1, characterized in that, Based on the IP address of the comparison server and the open port number of the Jenkins tool to be applied, an automated task is constructed, specifically including: In the Jenkins tool to be applied, a new initialization automated task is created; Configure the source code management address of the initialization automated task, where the source code management address is associated with the automated comparison Python script; In the automated task, configure a build trigger and set the corresponding target trigger condition for the build trigger; Configure the task execution node of the automated task and set the corresponding execution parameters for the task execution node; where the execution parameters include but are not limited to at least one of the following: the IP address of the comparison server, the open port number of the Jenkins tool to be applied, the execution path of the automated comparison Python script.

8. The automated data comparison method based on Jenkins according to claim 1, characterized in that Through the Jenkins tool to be applied, call the automated comparison Python script to implement the data comparison task between the databases to be compared, specifically including: According to the target trigger condition, automatically trigger the automated task, so that the Jenkins tool pulls the automated comparison Python script according to the configured source code management address; Execute the automated comparison Python script, and compare the data between the databases to be compared according to the data comparison rules and comparison result processing rules written in the automated comparison Python script, and generate a comparison result.

9. An automated data comparison device based on Jenkins, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an automated data comparison method according to any one of claims 1-8.

10. A non-volatile computer storage medium, characterized in that, Stores computer-executable instructions, and when the computer-executable instructions are executed by a computer, an automated data comparison method according to any one of claims 1-8 can be implemented.