Method for cleaning multi-source data based on data processing platform and related equipment
By providing multi-source data cleaning methods on the data processing platform, users can perform data processing through visual interfaces and drag controls, and use Spark engine to automatically perform tasks, solving the problems of complex operation, insufficient automation and limitations in the existing technology, realizing efficient and easy-to-use data cleaning and processing.
Patent Information
- Application Number
- CN202510108987.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
The existing data cleaning platform is complex to operate, requires users to master a lot of technical knowledge, and the degree of automation is insufficient, the processing of multiple data sources is limited, the user interface and operation are complex, and the traceability and interpretability of the impact after data cleaning are insufficient.
Provide a multi-source data cleaning method based on a data processing platform. By connecting multiple data sources into the data platform, converting them into standard formats, and providing a visual data preprocessing operation page, users can use the Spark computing engine to automatically perform data preprocessing tasks, and generate high-quality data sources.
It significantly reduces the time for manually writing code, and non-technical users can easily use the platform for data processing, improves data processing efficiency, reduces technical thresholds, and enhances data quality and traceability.
Smart Images

Figure CN120067083A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data cleaning, and in particular, to a method for cleaning multi-source data based on a data processing platform and related devices. Background Art
[0002] Data cleaning is a key step in the processes of data science and machine learning. Many studies are continuously developing their algorithm model platforms to improve the efficiency and effectiveness of data cleaning. The goal of data cleaning is to ensure the accuracy, integrity, and consistency of data, so as to provide a high-quality foundation for subsequent analysis and modeling.
[0003] In the related art, there are problems such as complex operations and the need for users to master a large amount of technical knowledge in data cleaning platforms. Summary of the Invention
[0004] In view of this, the purpose of this application is to propose a method for cleaning multi-source data based on a data processing platform and related devices.
[0005] Based on the above purpose, this application provides a method for cleaning multi-source data based on a data processing platform, including:
[0006] Connect a first data source to the data platform, and the type of the first data source is selected from at least one of MySQL, Oracle, PostgreSQL, HDFS, Hive, and HBase;
[0007] Convert the connected first data source into a second data source in a standard format in the data platform and store it in the distributed file system HDFS;
[0008] Provide a first page in the data platform, where the first page includes a plurality of selectable second data source controls and a plurality of selectable data preprocessing controls; the data source types corresponding to the plurality of selectable second data source controls are different; the data preprocessing types corresponding to the plurality of selectable data preprocessing controls are different;
[0009] In response to a selection instruction for a target second data source control among the plurality of selectable second data source controls, and, for a first target data preprocessing control among the plurality of selectable data preprocessing controls, perform corresponding preprocessing on the second data source to generate a third data source, and generate a control for the third data source in the first page.
[0010] In some embodiments, the first page further includes a canvas; the selection instruction for the target second data source control among the plurality of selectable second data source controls includes:
[0011] Click on the target second data source control and drag the target second data source control to the first target area on the canvas;
[0012] Click on the first target data preprocessing control and drag the first target data preprocessing control to the second target area on the canvas; wherein, the second target area is set at an interval from the first target area.
[0013] In some embodiments, after dragging the first target data preprocessing control to the second target area on the canvas, the method further includes: in response to a trigger instruction for the first target data preprocessing control, providing a window for editing the processing parameters of the first target data preprocessing control;
[0014] In response to a trigger instruction for the execution control in the window, perform corresponding preprocessing on the second data source.
[0015] In some embodiments, the method further includes: displaying a control of the third data source in a third target area on the canvas; the third target area is set at an interval from the second target area;
[0016] Performing the corresponding preprocessing on the second data source includes: automatically converting the operation corresponding to the target data preprocessing control into a Spark job and executing it through the computing engine of Spark.
[0017] In some embodiments, the multiple data preprocessing controls to be selected include at least two of a missing value processing control, a duplicate data detection control, an outlier detection control, a data standardization control, a data cleaning control, a data conversion control, a data aggregation control, a data quality processing control, a data integration control, a data anonymization control, a data storage control, and a data export control;
[0018] The method further includes generating a processing log of the second data source, and the processing log includes all the processing of the second data source.
[0019] In some embodiments, the method further includes:
[0020] Click on the second target data preprocessing control among the multiple data preprocessing controls to be selected, and drag the second target data preprocessing control to the fourth target area on the canvas; the fourth target area is set at an interval from the third target area;
[0021] Perform corresponding preprocessing on the third data source to generate a fourth data source, and generate a control of the fourth data source on the first page; the control of the fourth data source is set at an interval from the control of the third data source and is displayed simultaneously.
[0022] In some of these embodiments, accessing a first data source in the data platform includes: loading a driver corresponding to the data source, accessing a library specified by the data source according to the IP address and port of the data source, and creating a connection to the data source;
[0023] The method further includes providing a first window in response to a trigger operation on the target second data source control, where the first window is used to display a plurality of operation controls for the second data source;
[0024] In response to a trigger operation on a target operation control among the plurality of operation controls, a third window is generated in the first page, and the third window is used to display data information corresponding to the target operation control in the target second data source control.
[0025] An embodiment of the present application further provides a system for cleaning multi-source data based on a data processing platform, including:
[0026] An access module configured to access a first data source in the data platform, where the type of the first data source is selected from at least one of MySQL, Oracle, PostgreSQL, HDFS, Hive, and HBase;
[0027] A conversion module configured to convert the accessed first data source into a second data source in a standard format in the data platform and store it in the distributed file system HDFS;
[0028] A first preprocessing module configured to provide a first page in the data platform, where the first page includes a plurality of second data source controls to be selected and a plurality of data preprocessing controls to be selected; the data source types corresponding to the plurality of second data source controls to be selected are different; the data preprocessing types corresponding to the plurality of data preprocessing controls to be selected are different;
[0029] A second preprocessing module configured to perform corresponding preprocessing on the second data source in response to a selection instruction for a target second data source control among the plurality of second data source controls to be selected and a selection instruction for a first target data preprocessing control among the plurality of data preprocessing controls to be selected, generate a third data source, and generate a control for the third data source in the first page.
[0030] An embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method described in any one of the preceding items when executing the program.
[0031] An embodiment of the present application further provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described in any of the foregoing.
[0032] An embodiment of the present application further provides a computer program product including computer program instructions which, when running on a computer, cause the computer to execute the method as described in any of the foregoing.
[0033] As can be seen from the above, the method and related devices for cleaning multi-source data based on a data processing platform provided by the present application include connecting a first data source to the data platform, where the type of the first data source is selected from at least one of MySQL, Oracle, PostgreSQL, HDFS, Hive, and HBase; converting the connected first data source into a second data source in a standard format and storing it in the distributed file system HDFS in the data platform; providing a first page in the data platform, where the first page includes a plurality of selectable second data source controls and a plurality of selectable data preprocessing controls; the data source types corresponding to the plurality of selectable second data source controls are different; the data preprocessing types corresponding to the plurality of selectable data preprocessing controls are different; in response to a selection instruction for a target second data source control among the plurality of selectable second data source controls and a selection instruction for a first target data preprocessing control among the plurality of selectable data preprocessing controls, performing corresponding preprocessing on the second data source to generate a third data source and generating a control for the third data source in the first page; enabling visual processing operations on data sources, allowing users to quickly build a data processing flow, significantly reducing the time for manually writing code. Therefore, non-technical users can also easily use this platform to participate in data processing, thereby enhancing the team's collaboration ability. At the same time, by converting the data source into a standard format, an in-built data quality management function is realized, and thus the data of the processed data source has accuracy and consistency, providing a solid foundation for subsequent data analysis. Therefore, the method and related devices for cleaning multi-source data based on a data processing platform provided by the embodiments of the present application can improve the processing efficiency of multi-source data, lower the technical threshold, and improve data quality, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0035] Figure 1Schematic flowchart of the method for cleaning multi-source data based on a data processing platform according to an embodiment of the present application;
[0036] Figure 2 Schematic diagram of the initial page of the canvas according to an embodiment of the present application;
[0037] Figure 3 Schematic diagram of dragging the first data source to the canvas according to an embodiment of the present application;
[0038] Figure 4 Schematic diagram of the target data preprocessing control to the canvas according to an embodiment of the present application;
[0039] Figure 5 Schematic diagram of the window for editing the processing parameters of the target data preprocessing control according to an embodiment of the present application;
[0040] Figure 6 Schematic diagram of the execution result of the target data preprocessing control according to an embodiment of the present application;
[0041] Figure 7 Schematic diagram of the data export result according to an embodiment of the present application;
[0042] Figure 8 Block diagram of the structure of the system for cleaning multi-source data based on a data processing platform according to an embodiment of the present application;
[0043] Figure 9 Schematic diagram of the structure of the electronic device according to an embodiment of the present application.
[0044] Explanation of reference numerals:
[0045] 1 - Canvas; 2 - Second data source control; 3 - Data preprocessing control, 4 - Operation control, 5 - Window, 6 - Control of the third data source, 7 - Control of the fourth data source, 8 - First window, 9 - Third window, 10 - Preview area, 11 - Second page, 810 - Access module, 820 - Conversion module, 830 - First preprocessing module, 840 - Second preprocessing module. Detailed implementation manners
[0046] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0047] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of this application should have the ordinary meanings understood by those of ordinary skill in the field to which this application belongs. The "first", "second" and similar terms used in the embodiments of this application do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0048] The development of algorithm model platforms in the field of data cleaning shows trends of intelligence, automation and cloudification. In the field of data cleaning, technology companies and startups have launched various data cleaning tools and platforms. Platforms based on big data technology usually provide powerful data integration and processing capabilities and can process data from different sources. For example, frameworks such as Apache Spark and Apache Hadoop are widely used in large-scale data processing and cleaning, supporting batch processing and real-time stream processing.
[0049] Machine learning technologies have also begun to be introduced in some studies to automate the data cleaning process. By building intelligent algorithms, errors in data such as missing values, outliers and duplicate data can be automatically identified and corrected. These algorithms usually rely on statistical analysis, rule engines or deep learning models to improve the accuracy and efficiency of data cleaning. For example, some platforms can automatically clean and standardize text data through natural language processing technology, thereby improving data consistency.
[0050] The rise of cloud computing has also provided new opportunities for the development of data cleaning. Many cloud service providers such as AWS, Google Cloud and Microsoft Azure provide data cleaning and preparation tools that can be seamlessly integrated into data analysis and machine learning workflows. Users can conveniently perform data cleaning in the cloud and enjoy the advantages of elastic computing and storage resources.
[0051] In addition, with the increasing attention to data privacy and security, data cleaning platforms have also begun to focus on compliance and data protection. For example, some platforms have built-in data anonymization and encryption functions to ensure the protection of user privacy during the cleaning process. These measures not only comply with laws and regulations, but also enhance users' trust in the platform.
[0052] Although data cleaning platforms have improved the efficiency of data processing, there is still much room for improvement in aspects such as automation, multi-source data processing, user interface and operation complexity, as well as traceability and interpretability. Among them, the lack of automation is mainly reflected in that although some data cleaning platforms have introduced automation functions, fully automated data cleaning remains a challenge. Data cleaning often involves complex domain knowledge and context, and it is difficult for automated tools to comprehensively understand the nuances in the data and replace the judgment of human experts, especially in scenarios involving semantics, implicit associations, or industry-specific rules.
[0053] Among them, the limitations of multi-data source processing are mainly reflected in that when processing data from different data sources, problems such as inconsistent data formats, units, and naming are often faced. Existing platforms still have limitations in data fusion, standardization, and consistency processing, especially when data sources are extremely diverse and of uneven quality, it is difficult for the platform to automatically generate appropriate cleaning solutions.
[0054] Among them, the processing of user interface and operation complexity is mainly reflected in that the user interfaces and operation processes of some data cleaning platforms are relatively complex, especially for users without a programming background. Users need to master a large amount of technical knowledge to use the platform efficiently, which limits the participation of non-technical users. In addition, the platform lacks simple and easy-to-use interaction tools, making the threshold of data cleaning relatively high.
[0055] Among them, the lack of traceability and interpretability of the impact after data cleaning is mainly reflected in that during the data cleaning process, the platform usually automatically or semi-automatically processes a large number of data problems, but the interpretability of the cleaning operations is not strong. Especially in complex data processing, it is difficult for users to track which data have been modified, why they have been modified, and the impact of the modification on subsequent analysis and modeling. This lack of transparency will affect the trust in the cleaned data.
[0056] Based on this, the embodiments of the present application provide an integrated platform that supports the access and visual preprocessing operations of multi-source data. This integrated design enables users to complete data acquisition, storage, processing, and output in a unified environment, greatly simplifying the data processing flow. The method and related devices for cleaning multi-source data based on a data processing platform can, through the unified processing of multi-source data by the data platform, solve to a certain extent the problems existing in the prior art, such as insufficient automation, limitations in processing multi-data sources, complexity of user interfaces and operations, and insufficient traceability and interpretability of the impacts after data cleaning. Among them, by usually introducing higher-level automation functions in the data platform, such as automatic data cleaning, pattern recognition, and intelligent data conversion, the platform can use machine learning and rule engines to automatically identify data quality problems, reduce manual intervention, and improve data processing efficiency, thus solving to a certain extent the problem of insufficient automation. By integrating multiple data source connectors (such as JDBC, API, etc.), unified access to different types and formats of data can be achieved. This integration ability enables users to seamlessly connect and process multi-source data from MySQL, Oracle, HDFS, Hive, and HBase, etc., thus avoiding the complexity of manual migration and format conversion, and solving to a certain extent the problem of limitations in processing multi-data sources. By providing an intuitive and user-friendly graphical user interface (GUI) and a low-code or no-code operation method, a visual workflow and a drag-and-drop interface are provided, reducing the usage threshold for non-technical users and making data cleaning and processing simpler, thus solving to a certain extent the problem of complexity of user interfaces and operations. Through detailed logging and auditing functions, every step of data processing can be traced. Through data version control and visualization tools, users can clearly view the changes during the data cleaning process and understand the impacts of each step on the data, thereby enhancing transparency and trust, and solving to a certain extent the problem of insufficient traceability and interpretability of the impacts after data cleaning.
[0057] As Figure 1 shown, the embodiments of the present application provide a method for cleaning multi-source data based on a data processing platform. The method for cleaning multi-source data based on a data processing platform may include:
[0058] S100, connect a first data source in the data platform, and the type of the first data source is selected from at least one of MySQL, Oracle, PostgreSQL, HDFS, Hive, and HBase;
[0059] S200, convert the connected first data source into a second data source in a standard format in the data platform and store it in the distributed file system HDFS;
[0060] S300. Provide a first page in the data platform. The first page includes multiple second data source controls 2 to be selected and multiple data preprocessing controls 3 to be selected. The data source types corresponding to the multiple second data source controls 2 to be selected are different. The data preprocessing types corresponding to the multiple data preprocessing controls 3 to be selected are different.
[0061] S400. In response to a selection instruction for a target second data source control among the multiple second data source controls 2 to be selected and a selection instruction for a first target data preprocessing control among the multiple data preprocessing controls 3 to be selected, perform corresponding preprocessing on the second data source to generate a third data source, and display a control 6 of the third data source on the first page.
[0062] The method and related device for multi-source data cleaning based on a data processing platform provided by the embodiments of this application can perform visual processing operations on data sources, enabling users to quickly construct a data processing flow and significantly reducing the time for manually writing code. Therefore, non-technical users can also easily use this platform to participate in data processing, thereby enhancing the team's collaboration ability. At the same time, by converting the data source into a standard format, the built-in data quality management function is realized, and thus the data of the processed data source has accuracy and consistency, providing a solid foundation for subsequent data analysis. Therefore, the method and related device for multi-source data cleaning based on a data processing platform provided by the embodiments of this application can improve the processing efficiency of multi-source data, lower the technical threshold, and improve data quality, etc.
[0063] Among them, in step S100, the first data sources connected in the data platform may include data sources of types such as MySQL, Oracle, PostgreSQL, HDFS, Hive, and HBase. Among them, MySQL is a popular relational database management system (RDBMS) widely used in web applications. It supports multiple operating systems and programming languages. Oracle Database is an enterprise-level RDBMS that provides high performance, reliability, and security and is commonly used in large enterprises and mission-critical applications. PostgreSQL (PG) is an open-source RDBMS with advanced features and compliance with the SQL standard. HDFS (Hadoop Distributed File System) is a distributed file system designed to store large-scale data sets and is part of the Hadoop ecosystem. Hive is a data warehouse tool used to query and analyze large-scale data sets stored in the Hadoop file system. HBase is a distributed and scalable big data storage system that is built on top of the Hadoop file system and provides random real-time read / write access to large-scale data sets.
[0064] In some of these embodiments, accessing the first data source in the data platform may specifically include: loading the driver corresponding to the data source, accessing the library specified by the data source according to the IP address and port of the data source, so as to create a connection to the data source. In specific implementation, for the access to a specific first data source, the driver corresponding to the first data source can be loaded, the account password for accessing the first data source can be used, the IP address and port of the first data source can be accessed, and the library specified by the first data source can be accessed to create a JDBC connection.
[0065] In practical applications, the driver corresponding to the first data source can be a data source driver program exclusive to the first data source, can be a JDBC driver Driver program package provided by the data source manufacturer, and can be set in the loading directory where the data source access service is started and started together with the data source access service.
[0066] In some of these embodiments, in step S200, the format of the second data source can be the Parquet format. By converting the first data source into a second data source in a standard format and efficiently storing it in HDFS, the data after the conversion of the first data source can be made reliable and scalable, which is convenient for subsequent processing and analysis.
[0067] In some of these embodiments, in step S300, there is usually a canvas 1 on the page of the data platform. The data platform usually has an initial page, which can be as Figure 2 shown. Among them, there is usually also a canvas 1 in the initial page. Generally, on the first page, there are multiple second data source controls 2 to be selected and multiple data preprocessing controls 3 to be selected, such as Figure 3 shown. Among them, the multiple second data source controls 2 to be selected can respectively correspond to different types of data sources, such as MySQL, Oracle, PostgreSQL, HDFS, Hive, or HBase, etc. The multiple data preprocessing controls 3 to be selected can respectively correspond to different types of data preprocessing, for example, it can include at least two of a missing value processing control, a duplicate data detection control, an outlier detection control, a data normalization control, a data cleaning control, a data conversion control, a data aggregation control, a data quality processing control, a data integration control, a data anonymization control, a data storage control, and a data export control.
[0068] In some of these embodiments, as Figure 4As shown, the missing value processing control can include various types of operators, each of which can separately process the data source. For example, it can include the mean filling operator, median filling operator, and mode filling operator, etc. The duplicate data detection control is used to identify and remove duplicate entries in the dataset. Usually, its corresponding operator can include a deduplication algorithm based on feature matching. The outlier detection control is mainly used to identify outliers in the data using statistical analysis or machine learning algorithms to determine whether they need to be corrected or removed. The data standardization control is mainly used to convert data in different formats and units into a unified standard, such as date format, currency unit, etc., to improve data consistency. The data conversion control is mainly used for data type conversion (such as converting strings to numerical values) or feature scaling (such as normalization, standardization), etc., to facilitate subsequent analysis and modeling. The data cleaning control is mainly used to process text data using natural language processing techniques, including removing stop words, punctuation marks, stemming, and lemmatization, etc. The data quality processing control is mainly used to automatically check and correct data according to predefined rules, such as validating data range, format, and type, etc. The data aggregation control is mainly used to identify and correct errors in the data through model learning. For example, clustering algorithms can be used to discover data patterns or anomalies. The data integration control is mainly used to integrate data from different data sources to solve the problem of data inconsistency, usually involving techniques such as data merging and deduplication. The data anonymization control is mainly used to remove or replace personal identity information through technical means when processing sensitive data to ensure data privacy and compliance. The data storage control is mainly used to store data, for example, it can be stored in a specified location. The processing parameters of the data storage control can include the storage path and storage format. For example, it can be stored in a specified format at a specified location. The data export control is mainly used to export data. For example, it can be exported in a specified format to a specified location. The processing parameters of the data export control can include the export path and export format. The export format can be exported as CSV, JSON, or other formats, and the export path can be written to the target database.
[0069] In some of these embodiments, the processing parameters of the missing value processing control may include the data type for filling in missing values. For example, it may include missing value filling processes such as mean / median filling and mode filling. Among them, mean / median filling specifically may be: for numerical data, the missing values can be filled with the mean or median of the column. Mode filling specifically may be: for categorical data, the missing values can be filled with the value that appears most frequently (the mode). When the duplicate data detection control performs duplicate data detection, it specifically may be content-based comparison: detailed comparison of each field of the record to identify duplicate records. When the constant value detection control performs outlier detection, it specifically may use statistical methods: using box plots or standard deviation methods to identify outliers. Quantile analysis can also be used: by analyzing the distribution of the data, identifying those values that fall outside the normal range.
[0070] In some of these embodiments, the selection instruction for the target second data source control among the multiple second data source controls 2 to be selected may include:
[0071] Click on the target second data source control and drag the target second data source control to the first target area in the canvas 1, as Figure 3 shown. Usually, the first target area can be any area in the canvas 1. The second data source control 2 can be the data source that needs to perform data processing, and can be specifically determined according to the user's needs.
[0072] Click on the first target data preprocessing control and drag the first target data preprocessing control to the second target area in the canvas 1, as Figure 4 shown; wherein, the second target area is set at an interval from the first target area. Usually, the second target area can be set on the right side (i.e., the back side) of the first target area, for example Figure 4 shown on the left side (i.e., the back side). Usually, the target data preprocessing control 3 can be determined according to the processing to be performed, and can be specifically determined according to the user's needs.
[0073] In some of these embodiments, the method may further include: after dragging the first target data preprocessing control to the second target area in the canvas 1, in response to a trigger instruction for the first target data preprocessing control, providing a window 5, as Figure 5 shown, and the window 5 is used to edit the processing parameters of the first target data preprocessing control. Usually, the window 5 corresponds to the specific control type, and for different types of controls, the content displayed in the window 5 is usually different.
[0074] In response to a trigger instruction for the execution control in the window 5, corresponding preprocessing is performed. Generally, execution controls for triggering the operation of the window 5 can be set in the window 5, for example, it can include a control for saving parameters and a control for saving and executing parameters. The trigger instruction can be generated by clicking on the execution control in the window 5.
[0075] In some embodiments, the method may further include: displaying a control 6 of the third data source in a third target area in the canvas 1; the third target area is set at an interval from the second target area. For example, the third target area may be located on the right side of the second target area, forming a form of data flow with the second target area.
[0076] In some embodiments, the performing corresponding preprocessing on the second data source may include: automatically converting the operation corresponding to the target data preprocessing control 3 into a Spark job and executing it through the computing engine of Spark. In this way, the user does not need to write code, and complex data processing processes can be achieved by simple dragging. Specifically, when executing, that is, when the operator corresponding to the target data preprocessing control 3 runs, first, combining the implementation function and algorithm characteristics of the operator, and combining the input data scale (that is, the data scale of the second data source), evaluate how much CPU and memory resources are required, and apply to the Spark cluster to create the required resources, and then run the operator calculation task logic. After the operator calculation is completed, recycle the CPU and memory resources applied before this task.
[0077] In some embodiments, the method further includes making multiple selections among the multiple data preprocessing controls 3 to be selected for performing multiple data processes on the second target data source; where each data process can be respectively performed based on the result obtained from the previous data process, and corresponding data source controls can be formed respectively after each process, and are located after the data source control formed by the previous process, forming a data processing data flow, as Figure 6 shown.
[0078] In some embodiments, the method may further include: clicking on the second target data preprocessing control among the multiple data preprocessing controls 3 to be selected, and dragging the second target data preprocessing control to a fourth target area in the canvas 1; the fourth target area is set at an interval from the third target area;
[0079] Performing corresponding preprocessing on the third data source to generate a fourth data source, and generating a control 7 of the fourth data source in the first page; the control 7 of the fourth data source is set at an interval from the control 6 of the third data source and is displayed simultaneously.
[0080] It should be understood that the type of the second target data preprocessing control here can be the same as or different from that of the aforementioned first target data preprocessing control. After dragging the second target data preprocessing control to the canvas 1, it can also be edited in the same way as the window 5 of the first target data preprocessing control, and corresponding preprocessing can be performed on the third data source. That is, after dragging the second target data preprocessing control to the fourth target area of the canvas 1, in response to a trigger instruction for the second target data preprocessing control, a window 5 is provided, and the window 5 is used to edit the processing parameters of the second target data preprocessing control; in response to a trigger instruction for the execution control in the window 5, corresponding preprocessing is performed on the third data source. The operations for subsequent third target preprocessing controls, fourth target preprocessing controls, fifth target preprocessing controls,..., and the nth target preprocessing controls are the same as those of the aforementioned first target preprocessing control, and specific descriptions are not provided here.
[0081] In some embodiments, for the data export control, the result after export can be displayed in the form of a pop-up window, for example Figure 7 as shown. After executing the operator corresponding to the export control, corresponding touch operations can be performed on the generated corresponding data source control, so as to display the corresponding second page 11 with a preview area 10.
[0082] In some embodiments, as Figure 6 shown, the method may further include: in response to a trigger operation for the target second data source control, providing a first window 8, and the first window 8 is used to display a plurality of operation controls 4 for the second data source. Exemplarily, the operation control 4 may include controls such as preview data.
[0083] In response to a trigger operation for the target operation control 4 among the plurality of operation controls 4, a third window 9 is generated in the first page, and the third window 9 is used to display the data information corresponding to the target operation control 4 in the target second data source control. In this way, it is convenient to view the corresponding data information in the selected second data source.
[0084] In some embodiments, it further includes generating a processing log for the second data source, and the processing log may include all processes for the second data source. For example, it may include information such as the processing type and processing time, the data storage information before processing, and the data storage information after processing.
[0085] The method for cleaning multi-source data based on a data processing platform according to the embodiments of the present application can be applied to multiple industries. For example, it can be applied to industries such as the financial industry, healthcare, retail, manufacturing, energy industry, transportation, and environmental monitoring. Among them, when applied to the financial industry, the platform can be used to access bank transaction data, stock market data, user credit records, etc., and conduct business such as risk assessment, fraud detection, and investment strategy analysis. Among them, when applied to healthcare, by integrating medical device data, electronic health records (EHR), genomic data, etc., the platform can be used for disease diagnosis, patient monitoring, personalized medical plan design, etc. When applied to the retail industry, the platform can integrate point-of-sale (POS) data, inventory information, customer behavior data, etc., and be used for inventory management, customer segmentation, sales trend analysis, and market prediction. When applied to the manufacturing industry, by accessing production line data, supply chain information, product quality inspection data, etc., the platform can be used for production monitoring, quality control, supply chain optimization, and predictive maintenance. When applied to the energy industry, the platform can be used to integrate energy consumption data, equipment operating status, weather information, etc., and conduct energy demand forecasting, power grid management, renewable energy integration, etc. When applied to transportation, by accessing traffic flow data, vehicle tracking data, accident reports, etc., the platform can be used for traffic management, route optimization, accident prevention, and emergency response. When applied to environmental monitoring, by accessing weather station data, pollution monitoring data, satellite images, etc., the platform can be used for environmental change monitoring, disaster warning, and ecosystem management.
[0086] The method for cleaning multi-source data based on a data processing platform provided by the embodiments of the present application can bring various beneficial effects through a unified multi-source data processing method, specifically including improving data quality, accelerating the data processing process, reducing the technical threshold, enhancing data traceability, supporting decision-making, and enhancing data security, etc. Among them, improving data quality is mainly reflected in the unified data processing method. Through automated data cleaning and quality detection, errors and inconsistencies can be effectively reduced. This enables subsequent analysis and decision-making to be based on high-quality data, thereby improving the accuracy and reliability of the business. Accelerating the data processing process is mainly reflected in that through integrating multiple data sources and providing automated functions, the unified data processing significantly shortens the data preparation time. Users can quickly obtain the required data, reduce manual processing steps, and improve work efficiency. Especially when facing large-scale data sets, real-time or near-real-time data analysis can be achieved. Reducing the technical threshold is mainly reflected in that the friendly user interface and low-code operation method enable non-technical users to easily participate in data processing. In this way, business personnel and data analysts can more conveniently perform data cleaning and analysis, promote cross-departmental cooperation, and improve the data-driven ability of the overall organization. Enhancing data traceability is mainly reflected in that through a complete logging and auditing function, users can clearly track each step in the data cleaning and processing process. This transparency can enhance the trust in the data processing results, enabling better reasonable judgment based on data changes in decision-making. Among them, supporting decision-making is mainly reflected in that through high-quality and real-time updated data, more accurate insights can be provided for users (such as enterprises), supporting more effective decision-making. Whether it is market analysis, business optimization, or risk management, the unified data processing can provide strong data support for users (such as enterprises). Enhancing data security is mainly reflected in that through the integrated data governance and security management mechanism, the unified data processing method can better protect sensitive data and comply with relevant regulatory requirements (such as GDPR, etc.). This can enhance the compliance of data management and reduce the risk of data leakage.
[0087] It can be understood that before using the technical solutions of the various embodiments in the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to users in an appropriate manner, and user authorization will be obtained.
[0088] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0089] As an optional but non-limiting implementation, in response to receiving an active request from a user, the way to send a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0090] It can be understood that the above notification and user authorization acquisition process is only illustrative and does not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0091] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.
[0092] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0093] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a system for cleaning multi-source data based on a data processing platform.
[0094] Referring to Figure 8 , the system 800 for cleaning multi-source data based on a data processing platform may include:
[0095] A conversion module 820, configured to convert the accessed first data source into a second data source in a standard format in the data platform and store it in the distributed file system HDFS;
[0096] A first preprocessing module 830, configured to provide a first page in the data platform, the first page including a plurality of second data source controls 2 to be selected and a plurality of data preprocessing controls 3 to be selected; the data source types corresponding to the plurality of second data source controls 2 to be selected are different; the data preprocessing types corresponding to the plurality of data preprocessing controls 3 to be selected are different;
[0097] The second preprocessing module 840 is configured to, in response to a selection instruction for a target second data source control among the multiple second data source controls to be selected, and a selection instruction for a first target data preprocessing control among the multiple data preprocessing controls to be selected, perform corresponding preprocessing on the second data source, generate a third data source, and generate a control 6 for the third data source in the first page.
[0098] In some embodiments, the first page further includes a canvas 1; the selection instruction for the target second data source control among the multiple second data source controls to be selected includes:
[0099] Click on the target second data source control and drag the target second data source control to a first target area in the canvas 1.
[0100] Click on the first target data preprocessing control and drag the first target data preprocessing control to a second target area in the canvas 1; wherein, the second target area is spaced apart from the first target area.
[0101] In some embodiments, the system further includes a processing parameter editing module: configured to, in response to a trigger instruction for the first target data preprocessing control, provide a window 5, and the window 5 is used to edit the processing parameters of the first target data preprocessing control.
[0102] In response to a trigger instruction for an execution control in the window 5, perform corresponding preprocessing on the second data source.
[0103] In some embodiments, the system further includes a first display module, configured to display the control 6 for the third data source in a third target area in the canvas 1; the third target area is spaced apart from the second target area.
[0104] In some embodiments, the second preprocessing module 840 is configured to automatically convert the operations corresponding to the target data preprocessing control 3 into Spark jobs and execute them through the computing engine of Spark.
[0105] In some embodiments, the multiple data preprocessing controls 3 to be selected include at least two of a missing value processing control, a duplicate data detection control, an outlier detection control, a data standardization control, a data cleaning control, a data conversion control, a data aggregation control, a data quality processing control, a data integration control, a data anonymization control, a data storage control, and a data export control.
[0106] In some of these embodiments, the system further includes a processing log generation module, and the method further includes generating a processing log of the second data source, where the processing log includes all processing of the second data source.
[0107] In some of these embodiments, the system further includes a third preprocessing module configured to click on a second target data preprocessing control among the multiple data preprocessing controls 3 to be selected, and drag the second target data preprocessing control to a fourth target area in the canvas 1; the fourth target area is set at an interval from the third target area;
[0108] Perform corresponding preprocessing on the third data source to generate a fourth data source, and generate a control 7 of the fourth data source in the first page; the control 7 of the fourth data source is set at an interval from the control 6 of the third data source and is displayed simultaneously.
[0109] In some of these embodiments, the access module 810 is configured to: load a driver corresponding to the data source, access a library specified by the data source according to the IP address and port of the data source, so as to create a connection to the data source.
[0110] In some of these embodiments, the system further includes a second display module configured to, in response to a trigger operation on the target second data source control, provide a first window 8, where the first window 8 is used to display a plurality of operation controls 4 for the second data source;
[0111] In response to a trigger operation on a target operation control 4 among the plurality of operation controls 4, a third window 9 is generated in the first page, and the third window 9 is used to display data information corresponding to the target operation control 4 in the target second data source control.
[0112] For convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present application, the functions of each module can be implemented in one or more software and / or hardware.
[0113] The device in the above embodiments is used to implement the corresponding method for cleaning multi-source data based on a data processing platform in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0114] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the method for cleaning multi-source data based on a data processing platform in any of the above embodiments.
[0115] Figure 9Fig. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0116] The processor 1010 may be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0117] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0118] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0119] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0120] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0121] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0122] The electronic device of the above embodiment is used to implement the corresponding method for cleaning multi-source data based on the data processing platform in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0123] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method for cleaning multi-source data based on the data processing platform as described in any of the foregoing embodiments.
[0124] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0125] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the method for cleaning multi-source data based on the data processing platform as described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0126] Based on the same inventive concept, corresponding to the method for cleaning multi-source data based on a data processing platform described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to execute the method for cleaning multi-source data based on a data processing platform. Corresponding to the execution subject of each step in each embodiment of the method for cleaning multi-source data based on a data processing platform, the processor that executes the corresponding step can belong to the corresponding execution subject.
[0127] The computer program product of the above embodiment is used to cause the computer and / or the processor to execute the method for cleaning multi-source data based on a data processing platform described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0128] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.
[0129] In addition, for simplicity of description and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device can be shown in block diagram form to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application will be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0130] Although the present application has been described in connection with specific embodiments of the present application, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used in the embodiments discussed.
[0131] Embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A method for cleaning multi-source data based on a data processing platform, characterized in that: include: Accessing a first data source in the data platform, wherein the type of the first data source is selected from at least one of MySQL, Oracle, PostgreSQL, HDFS, Hive, and HBase; In the data platform, the first data source connected is converted into a second data source in a standard format and stored in the distributed file system HDFS; Providing a first page in the data platform, the first page comprising a plurality of second data source controls to be selected and a plurality of data preprocessing controls to be selected; The data source types corresponding to the multiple second data source controls to be selected are different; The multiple data preprocessing controls to be selected correspond to different data preprocessing types; In response to a selection instruction for a target second data source control among the multiple second data source controls to be selected, and a selection instruction for a first target data preprocessing control among the multiple data preprocessing controls to be selected, corresponding preprocessing is performed on the second data source to generate a third data source, and a control for the third data source is generated in the first page.
2. The method for cleaning multi-source data based on a data processing platform according to claim 1, characterized in that: The first page also includes a canvas; the selection instruction for a target second data source control among the plurality of second data source controls to be selected includes: Click the target second data source control, and drag the target second data source control to the first target area in the canvas; Click the first target data preprocessing control, and drag the first target data preprocessing control to a second target area in the canvas; wherein the second target area is spaced apart from the first target area.
3. The method for cleaning multi-source data based on a data processing platform according to claim 2 is characterized in that: After dragging the first target data preprocessing control to the second target area in the canvas, the method further includes: in response to a trigger instruction for the first target data preprocessing control, providing a window for editing processing parameters of the first target data preprocessing control; In response to a trigger instruction for the execution control in the window, corresponding preprocessing is performed on the second data source.
4. The method for cleaning multi-source data based on a data processing platform according to claim 3 is characterized in that: The method further includes: displaying the control of the third data source in a third target area in the canvas; the third target area is spaced apart from the second target area; The performing corresponding preprocessing on the second data source includes: automatically converting the operation corresponding to the target data preprocessing control into a Spark job and executing it through the Spark computing engine.
5. The method for cleaning multi-source data based on a data processing platform according to claim 1 is characterized in that: The multiple data preprocessing controls to be selected include at least two of missing value processing controls, duplicate data detection controls, outlier detection controls, data standardization controls, data cleaning controls, data conversion controls, data aggregation controls, data quality processing controls, data integration controls, data anonymization controls, data storage controls and data export controls; The method further includes generating a processing log of the second data source, the processing log including all processing of the second data source.
6. The method for cleaning multi-source data based on a data processing platform according to claim 4 is characterized in that: The method further comprises: Clicking a second target data preprocessing control among the multiple data preprocessing controls to be selected, and dragging the second target data preprocessing control to a fourth target area in the canvas; the fourth target area is spaced apart from the third target area; Perform corresponding preprocessing on the third data source to generate a fourth data source, and generate a control of the fourth data source in the first page; the control of the fourth data source is spaced apart from the control of the third data source and displayed at the same time.
7. The method for cleaning multi-source data based on a data processing platform according to claim 6 is characterized in that: The accessing of the first data source in the data platform includes: loading a driver corresponding to the data source, accessing a library specified by the data source according to an IP address and a port of the data source, so as to create a connection to the data source; The method further includes, in response to a trigger operation on the target second data source control, providing a first window, wherein the first window is used to display a plurality of operation controls on the second data source; In response to a trigger operation on a target operation control among the multiple operation controls, a third window is generated in the first page, and the third window is used to display data information in the target second data source control corresponding to the target operation control.
8. A system for cleaning multi-source data based on a data processing platform, characterized in that: include: An access module is configured to access a first data source in the data platform, wherein a type of the first data source is selected from at least one of MySQL, Oracle, PostgreSQL, HDFS, Hive, and HBase; A conversion module is configured to convert the first data source connected to the data platform into a second data source in a standard format and store it in a distributed file system HDFS; A first preprocessing module is configured to provide a first page in the data platform, wherein the first page includes a plurality of second data source controls to be selected and a plurality of data preprocessing controls to be selected; The data source types corresponding to the multiple second data source controls to be selected are different; The multiple data preprocessing controls to be selected correspond to different data preprocessing types; The second preprocessing module is configured to respond to a selection instruction for a target second data source control among the multiple second data source controls to be selected and a selection instruction for a first target data preprocessing control among the multiple data preprocessing controls to be selected, perform corresponding preprocessing on the second data source, generate a third data source, and generate a control of the third data source in the first page.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program. 10 . A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to claim 1 .