Data processing method and device, equipment and storage medium

CN117421352BActive Publication Date: 2026-09-18INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311482352.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2026-09-18
Estimated Expiration
2043-11-08

AI Technical Summary

Technical Problem

但通过这种方式对数据信息进行挖掘的效率低

Benefits of technology

[0010] Fifthly, embodiments of this application provide a computer program product, including a computer program, which, when executed, implements the data processing method of the first aspect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117421352B_ABST
    Figure CN117421352B_ABST
Patent Text Reader

Abstract

The application provides a data processing method and device, equipment and a storage medium, relates to the field of big data, the field of financial technology or other related fields. The method comprises the following steps: acquiring a data table comprising to-be-processed data and table metadata of the data table; determining business-related information of the to-be-processed data by using a text-to-text engine on the table metadata; determining data relationships of the to-be-processed data according to the data table, wherein the data relationships comprise at least one of inter-table data relationships, intra-table column data relationships and intra-column data relationships; determining data quality of the to-be-processed data according to the data relationships; and performing information mining on the to-be-processed data according to the business-related information, the data relationships and the data quality. The method provided in the application can improve the efficiency of data information mining.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of big data, fintech, or other related fields, and in particular to a data processing method, apparatus, device, and storage medium. Background Technology

[0002] A data middle platform refers to an intermediate support platform that consolidates business and data from existing or newly built information systems, enabling data to empower new businesses and / or applications. The purpose of building a data middle platform is to rapidly release the value of data. However, due to barriers in data business knowledge, lack of modeling information for middle platform construction, and low data quality, data mining is becoming increasingly difficult. Therefore, reducing the difficulty of data mining is particularly important.

[0003] In related technologies, data mining is achieved through business domain knowledge exploration, modeling exploration, and data quality exploration. Specifically, this involves manually querying the business domain knowledge corresponding to the data, studying the ER diagram model of the data to determine the modeling relationships between data, and ensuring data quality through data cleaning. However, this method of data mining is inefficient.

[0004] Therefore, there is an urgent need for a solution that can improve the efficiency of data mining. Summary of the Invention

[0005] This application provides a data processing method, apparatus, device, and storage medium to improve the efficiency of data mining.

[0006] In a first aspect, this application provides a data processing method, comprising: acquiring a data table including data to be processed and table metadata of the data table; using a text-to-text engine to determine business-related information of the data to be processed from the table metadata; determining data relationships of the data to be processed based on the data table, wherein the data relationships include at least one of inter-table data relationships, intra-table column data relationships, and intra-column data relationships; determining data quality of the data to be processed based on the data relationships; and performing information mining on the data to be processed based on the business-related information, the data relationships, and the data quality.

[0007] Secondly, this application provides a data processing apparatus, comprising: an acquisition module for acquiring a data table including data to be processed and table metadata of the data table; a first determination module for using a text-to-text engine to determine business-related information of the data to be processed from the table metadata; a second determination module for determining data relationships of the data to be processed based on the data table, wherein the data relationships include at least one of inter-table data relationships, intra-table column data relationships, and intra-column data relationships; a third determination module for determining the data quality of the data to be processed based on the data relationships; and a mining module for performing information mining on the data to be processed based on the business-related information, the data relationships, and the data quality.

[0008] Thirdly, this application provides an electronic device, including: a processor and a memory connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the data processing method of the first aspect.

[0009] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed, are used to implement the data processing method as described in the first aspect.

[0010] Fifthly, embodiments of this application provide a computer program product, including a computer program, which, when executed, implements the data processing method of the first aspect.

[0011] The data processing method, apparatus, equipment, and storage medium provided in this application can automatically generate business-related information about the data to be processed by inputting the table metadata of the data table into a text-to-text engine. This solves the problems of low efficiency and low accuracy in manually querying business-related information. Furthermore, by analyzing the data table of the data to be processed, at least one of the following data relationships can be determined: inter-table data relationships, intra-table column data relationships, and intra-column data relationships. This addresses the problem of low efficiency in determining data relationships due to the high learning and research costs of modeling the Entity-Relationship Model (ER). This further improves the efficiency of data mining. In addition, after determining the data relationships of the data to be processed, the data quality of the data is also determined through these relationships. This solves the problem that data cleaning alone cannot guarantee high-quality data, further ensuring the accuracy of data mining. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0013] Figure 1 A schematic diagram of the structure of a data processing system provided in an embodiment of this application;

[0014] Figure 2 A schematic flowchart of a data processing method provided in an embodiment of this application;

[0015] Figure 3 A schematic diagram of a scatter plot provided in an embodiment of this application;

[0016] Figure 4 A schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application;

[0017] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0018] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0020] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0021] It should be noted that the data processing method, apparatus, equipment, and storage medium of this application can be used in the fields of big data and fintech, as well as in any field other than big data. The application fields of the data processing method, apparatus, equipment, and storage medium of this application are not limited.

[0022] First, let me explain the terms used in this application:

[0023] Data middle platform: refers to the intermediate support platform that consolidates the business and data of existing or newly built information systems and enables data to empower new businesses and / or new applications.

[0024] A data table is a grid-based virtual table that temporarily stores data (representing a table of data in memory).

[0025] Enumeration type: Used to declare a set of named constants. When a variable has several possible values, it can be defined as an enumeration type.

[0026] Scatter plot: In regression analysis, a scatter plot is a distribution of data points on a Cartesian coordinate plane. The scatter plot shows the general trend of the dependent variable changing with the independent variable, and based on this, an appropriate function can be selected to fit the data points.

[0027] In view of the problems existing in the related technologies provided in the background, this application proposes a data processing method that can improve the efficiency of data mining.

[0028] Specifically, by inputting the table metadata of the data tables to be processed into the text-to-text engine, business-related information about the data can be automatically generated. This solves the problems of low efficiency and low accuracy in manually querying business-related information, thereby further improving the efficiency of data mining. Furthermore, by performing data similarity analysis on the data tables to be processed, at least one of the following data relationships—inter-table data relationships, intra-table column data relationships, and intra-column data relationships—can be determined. This addresses the problem of low efficiency in determining data relationships due to the high learning and research costs of ER diagram modeling, further improving the efficiency of data mining. In addition, after determining the data relationships, the data quality of the data to be processed is also determined based on these relationships. This solves the problem that data cleaning alone cannot guarantee high-quality data, further improving the accuracy of data mining.

[0029] In one embodiment, the data processing method can be applied in an application scenario. Figure 1 This is a schematic diagram of the structure of a data processing system provided in an embodiment of this application, such as... Figure 1 As shown, the data processing system may include a text-to-text engine, a data relationship determination module, a data quality determination module, and an information mining module.

[0030] In this application scenario, data mining of the data to be processed requires understanding its business-related information, data relationships, and data quality. Improving the efficiency of determining these factors, while reducing their difficulty, will enhance the efficiency of data mining.

[0031] In the above application scenario, the data to be processed is stored in data tables. When performing data mining on the data to be processed, the table metadata of the data tables can be input into a text-to-text engine to obtain business-related information about the data. The data tables are then input into a data relationship determination module for data similarity analysis to determine the data relationships, which can include at least one of inter-table relationships, intra-table column relationships, and intra-column relationships. Finally, the data relationships are input into a data quality determination module for data quality analysis to determine the data quality of the data to be processed.

[0032] In the above application scenarios, after determining the business-related information, data relationships, and data quality of the data to be processed, these information-related information, data relationships, and data quality can be input into the information mining module for information mining of the data to be processed.

[0033] In light of the above scenarios, the technical solutions of this application and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0034] This application provides a data processing method. Figure 2 A flowchart illustrating a data processing method provided in an embodiment of this application is shown below. Figure 2 As shown, the data processing method includes the following steps:

[0035] S201: Obtain the data table containing the data to be processed and the table metadata of the data table.

[0036] In this step, table metadata can represent the metadata of the data table. For example, table metadata can include the table name, column names, data types, and field names for specific business functions.

[0037] S202: For table metadata, use a text-to-text engine to determine the business-related information of the data to be processed.

[0038] In this step, when performing information mining on the data to be processed, it is necessary to determine the business-related information of the data, such as the business domain and related introductory information. Therefore, table metadata can be input into the text-to-text engine to generate business-related information about the data to be processed.

[0039] S203: Determine the data relationships of the data to be processed based on the data table.

[0040] In this step, data relationships include at least one of the following: inter-table data relationships, intra-table column data relationships, and intra-column data relationships. That is, by exploring the reverse data relationship model of the data tables to be processed, the data relationships of the data to be processed can be determined.

[0041] Specifically, the data table to be processed can be one or multiple tables. Therefore, the data relationship of the data to be processed can be the inter-table data relationship between multiple data tables, the intra-table column data relationship between multiple columns in a single data table, or the intra-column data relationship of a single column in a single data table.

[0042] S204: Determine the data quality of the data to be processed based on the data relationships.

[0043] In this step, if there are relationships between the data to be processed, that is, at least one of the following: inter-table data relationships, intra-table column data relationships, and intra-column data relationships, it can be considered that the data to be processed does not have data quality problems, or that data quality problems can be ignored.

[0044] S205: Based on business-related information, data relationships, and data quality, perform information mining on the data to be processed.

[0045] In this step, once the business-related knowledge, data relationships, and data quality of the data to be processed are all determined, information mining of the data can be performed efficiently and with high accuracy. If the business-related knowledge and / or data relationships of the data to be processed are not determined, or if the data to be processed has data quality issues, then information mining of the data to be processed can be omitted.

[0046] The data processing method provided in this embodiment automatically generates business-related information about the data to be processed by inputting the table metadata of the data table into a text-to-text engine. This solves the problems of low efficiency and low accuracy in manually querying business-related information. Furthermore, by analyzing the data table of the data to be processed, at least one of the following data relationships—inter-table data relationships, intra-table column data relationships, and intra-column data relationships—is determined. This addresses the problem of low efficiency in determining data relationships due to the high learning and research costs of ER diagram modeling. This further improves the efficiency of data mining. In addition, after determining the data relationships of the data to be processed, the data quality is also determined through these relationships. This solves the problem that data cleaning alone cannot guarantee high-quality data, further ensuring the accuracy of data mining.

[0047] In one embodiment, the data table includes a first data table and a second data table. Determining the data relationships of the data to be processed based on the data tables includes: performing enumeration-based filtering on the first data in each column of the first data table and on the second data in each column of the second data table to obtain first enumeration-based data, and performing enumeration-based filtering on the second data to obtain second enumeration-based data; comparing the similarity between the first enumeration-based data and the second enumeration-based data to determine a first similarity between them; and determining the inter-table data relationships of the data to be processed based on the first similarity.

[0048] In this embodiment, the data tables to be processed may include multiple tables. To address the problem of low efficiency in determining data relationships due to the high learning and research costs of modeling ER diagrams for data, any two data tables from the multiple tables can be analyzed when determining the data relationships between the tables to be processed, i.e., the first data table and the second data table, to further improve the efficiency of data information mining.

[0049] Optionally, the first data table can be denoted as Table A, and the second data table as Table B. When determining the data relationship between Table A and Table B, an enumeration-type filtering process can be performed on the data in the i-th column of Table A, i.e., `select distinct(A(i))asc`, to obtain the first enumeration-type data Ai; and an enumeration-type filtering process can be performed on the data in the j-th column of Table B, i.e., `select distinct(B(j))asc`, to obtain the second enumeration-type data Bj. Where 1 ≤ i ≤ n, 1 ≤ j ≤ n, and i and j are both integers. This avoids redundant calculations of duplicate data, preventing resource waste.

[0050] In one optional implementation, after determining the first enumeration class data Ai and the second enumeration class data Bj, a similarity analysis can be performed on the first enumeration class data Ai and the second enumeration class data Bj to determine the first similarity between each column of data in table A and each column of data in table B. Thus, based on the first similarity, it can be determined whether there is an inter-table data relationship between table A and table B.

[0051] In one embodiment, determining the inter-table data relationship of the data to be processed based on a first similarity includes: determining that there is an association between the first data table and the second data table in response to the first similarity being greater than or equal to a first similarity threshold; and determining that there is no association between the first data table and the second data table in response to the first similarity being less than the first similarity threshold.

[0052] In this embodiment, if the first similarity between the first enumerated data Ai and the second enumerated data Bj is greater than or equal to the first similarity threshold, then the data similarity between the first enumerated data Ai and the second enumerated data Bj is considered to be high. Therefore, it can be determined that there is an inter-table data relationship between table A and table B, that is, there is an association between the first data table and the second data table. This can solve the problem of low efficiency in determining data relationships caused by the high learning and research cost of modeling the ER diagram of data, thereby further improving the efficiency of data information mining.

[0053] Optionally, if the first similarity between the first enumerated data Ai and the second enumerated data Bj is less than the first similarity threshold, then the data similarity between the first enumerated data Ai and the second enumerated data Bj is considered low. Furthermore, if the first similarity between each column Ai of table A and each column Bj of table B is less than the first similarity threshold, then it can be determined that there is no inter-table data relationship between table A and table B; that is, there is no association between the first data table and the second data table. Therefore, to avoid wasting resources, information mining of the data to be processed can be omitted.

[0054] Optionally, the first similarity threshold can be 100%, or other pre-defined similarity thresholds.

[0055] In one embodiment, determining the data relationship of the data to be processed based on the data table includes: performing enumeration-type filtering on the third data in each column of the data table to obtain third enumeration-type data; comparing the similarity of the third enumeration-type data corresponding to any two columns of third data to determine the second similarity between the third enumeration-type data corresponding to any two columns of third data; and determining the intra-table column data relationship of the data to be processed based on the second similarity.

[0056] In this embodiment, the data table to be processed can be one or multiple tables. In order to solve the problem of low efficiency in determining data relationships due to the high learning and research cost of modeling the ER diagram of the data, the data relationships of the columns within any data table can be determined to further improve the efficiency of data information mining.

[0057] Optionally, the data table can be denoted as C, and the third data in each column of table C can be denoted as C(i). When determining the relationship between columns in the data table, in order to avoid redundant calculations of duplicate data and waste resources, we can first perform enumeration-type filtering on the data in the i-th column of table C, that is, select C(i), count(C(i)) group by C(i), to obtain the third enumeration-type data Ci. Then, we can compare the similarity of any two columns Ci to obtain the second similarity. In this way, we can determine whether there is a relationship between columns in table C based on the second similarity.

[0058] In one alternative implementation, after performing enumeration-based filtering on the data in the i-th column of table C, a table CI including Ci can be obtained. Then, the similarity between any two columns of data in table CI is determined to obtain a second similarity.

[0059] In one embodiment, determining the intra-table column data relationship of the data to be processed based on the second similarity includes: in response to the second similarity being greater than or equal to the second similarity threshold, determining that the intra-table column data relationship is that there is an association relationship between the intra-table column data of the data table; and in response to the second similarity being less than the second similarity threshold, determining that there is no association relationship between the intra-table column data of the data table.

[0060] In this embodiment, if the second similarity between any two columns Ci is greater than or equal to the second similarity threshold, then the similarity between these two columns Ci is considered high. Therefore, it can be determined that there are intra-table column data relationships between the columns in table C, that is, there are associations between the intra-table column data. This solves the problem of low efficiency in determining data relationships caused by the high learning and research cost of modeling the ER diagram of the data, thereby further improving the efficiency of data information mining.

[0061] Optionally, if the second similarity between any two columns Ci is less than the second similarity threshold, then the similarity between the data in the columns of table C is considered low. Therefore, it can be determined that there is no relationship between the data in the columns within the table, that is, there is no association between the data in the columns within the table. In this case, to avoid wasting resources, information mining of the data to be processed can be omitted.

[0062] Optionally, the second similarity threshold can be 100%, or other pre-defined similarity thresholds.

[0063] In one embodiment, determining the data relationship of the data to be processed based on the data table includes: performing enumeration-type filtering on the fourth data in each column of the data table to obtain fourth enumeration-type data; generating a scatter plot of the fourth enumeration-type data, the scatter plot being used to represent the data dispersion of the fourth enumeration-type data; and determining the intra-column data relationship of the data to be processed based on the scatter plot.

[0064] In this embodiment, the data table to be processed can be one or multiple tables. In order to solve the problem of low efficiency in determining data relationships due to the high learning and research cost of modeling the ER diagram of the data, the data relationships within the columns of any data table can be determined to further improve the efficiency of data information mining.

[0065] In one alternative implementation, the data table can be denoted as D, and the fourth data in each column of table D can be denoted as D(i). When determining the data relationship within the columns of the data table, in order to avoid redundant calculations of duplicate data and thus waste resources, the data in the i-th column of table D can first be subjected to enumeration-type filtering, that is, select D(i), count(D(i)) groupby D(i), to obtain the third enumeration class data Di. For this column Di, data fitting can be performed on column Di to obtain a scatter plot. The data dispersion shown by the scatter plot can determine the data relationship within the columns of table D.

[0066] In one embodiment, determining the intra-column data relationship of the data to be processed based on the scatter plot includes: in response to the scatter plot indicating that the data dispersion of the fourth enumeration class data is greater than or equal to a preset value, determining that there is no correlation between the intra-column data of the data table; and in response to the scatter plot indicating that the data dispersion of the fourth enumeration class data is less than a preset value, determining that there is a correlation between the intra-column data of the data table.

[0067] In this embodiment, a scatter plot can be used to represent data dispersion. After generating the scatter plot of the fourth enumeration class data, the data dispersion of the fourth enumeration class data can be calculated. If the data dispersion is greater than or equal to a preset value, it can be considered that the data dispersion of the fourth enumeration class data in that column is high. Therefore, it can be considered that there is no correlation between the data within the column corresponding to the fourth enumeration class data in that column of the data table. If there is no correlation between the data within the columns of all columns in the data table, it can be considered that there is no correlation between the data within the columns of the data table. In this way, to avoid wasting resources, information mining of the data to be processed can be omitted.

[0068] Optionally, if the data dispersion is less than a preset value, it can be considered that the data dispersion of the fourth enumeration class in this column is low. Therefore, it can be assumed that there is a correlation between the data within the column corresponding to the fourth enumeration class in this column, that is, it can be assumed that there is an intra-column data relationship in the data table. This can solve the problem of low efficiency in determining data relationships caused by the high learning and research cost of ER diagram modeling of data, and further improve the efficiency of data information mining.

[0069] In an alternative implementation, the scatter plot of the fourth enumeration class data in this column can be as follows: Figure 3 As shown, Figure 3 This is a schematic diagram of a scatter plot provided in an embodiment of this application. Figure 3 It can be seen that some of the data in the fourth enumeration class of this column is abnormal.

[0070] In one embodiment, determining the data quality of the data to be processed based on data relationships includes: in response to the following: if the inter-table data relationship is that there is no association between the first data table and the second data table, or the intra-table column data relationship is that there is no association between the intra-table column data of the data table, or the intra-column data relationship is that there is no association between the intra-column data of the data table, determining that the data to be processed has a data quality problem.

[0071] In this embodiment, the quality of the data to be processed directly affects the accuracy of information mining; higher data quality results in higher accuracy. Therefore, when mining information from the data to be processed, in addition to determining the business-related information and data relationships, it is also necessary to determine the data quality. This allows for high-precision information mining through high-quality data.

[0072] Specifically, the data quality of the data to be processed can be determined based on the data relationships between the data. That is, if there are no relationships between the multiple tables containing the data to be processed, or if there are no relationships between the columns within any single table, then the data to be processed has a data quality problem. If there are relationships between the multiple tables containing the data to be processed, and if there are relationships between the columns within any single table, then the data to be processed does not have a data quality problem.

[0073] In one embodiment, after determining the data quality of the data to be processed based on the data relationships, the method further includes: generating and outputting a data quality analysis report of the data to be processed, wherein the data quality analysis report includes at least one of the following information: the completeness of the data to be processed, the relevance of the data to be processed, the uniqueness of the data to be processed, and the validity of the data to be processed.

[0074] In this embodiment, in order to facilitate relevant personnel to understand the data quality of the data to be processed in a timely manner, a data quality analysis report of the data to be processed can be automatically generated and output after the data quality of the data to be processed is determined. This can also effectively reduce the problems of low accuracy and low efficiency of manual analysis of the data quality of the data to be processed.

[0075] Optionally, the completeness of the data to be processed included in the data quality analysis report can be expressed as: Completeness = Number of null records / Total number of sample records * 100%; the relevance of the data to be processed can be expressed as: Relevance = Total number of sample records without corresponding primary keys / Total number of sample records * 100%; the uniqueness of the data to be processed can be expressed as: 1 - Total number of sample records with duplicate primary keys / Total number of sample records * 100%; and the validity of the data to be processed can be expressed as: 1 - Total number of abnormal sample records outside the value range / Total number of sample records * 100%.

[0076] The data processing method provided in this embodiment assists in the process of mining potential business information from data through text-to-text technology, and assists in the process of determining data quality through the data relationships existing in relational databases, namely, inter-table relationships (foreign key information), inter-column relationships (primary key) within tables, and intra-column relationships. Thus, data information mining can be achieved through business information, data relationships, and data quality, which can greatly reduce the cost of data mining, quickly and autonomously mine the potential value of data, and enable the data platform to quickly realize the value release of data.

[0077] This application also provides a data processing apparatus. Figure 4 This is a schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application, such as... Figure 4 As shown, the data processing device 400 includes:

[0078] The acquisition module 401 is used to acquire data tables including the data to be processed and table metadata of the data tables;

[0079] The first determination module 402 is used to determine the business-related information of the data to be processed by using a text-to-text engine on the table metadata.

[0080] The second determining module 403 is used to determine the data relationship of the data to be processed based on the data table. The data relationship includes at least one of the following: inter-table data relationship, intra-table column data relationship, and intra-column data relationship.

[0081] The third determining module 404 is used to determine the data quality of the data to be processed based on the data relationships.

[0082] The mining module 405 is used to perform information mining on the data to be processed based on business-related information, data relationships, and data quality.

[0083] Optionally, the data table includes a first data table and a second data table. When the second determining module 403 determines the data relationship of the data to be processed based on the data table, it is specifically used to: perform enumeration-type filtering on the first data for each column of first data in the first data table and on the second data for each column of second data in the second data table to obtain first enumeration-type data, and perform enumeration-type filtering on the second data to obtain second enumeration-type data; compare the similarity between the first enumeration-type data and the second enumeration-type data to determine a first similarity between the first enumeration-type data and the second enumeration-type data; and determine the inter-table data relationship of the data to be processed based on the first similarity.

[0084] Optionally, when the second determining module 403 determines the inter-table data relationship of the data to be processed based on the first similarity, it is specifically used to: determine that there is an association between the first data table and the second data table if the first similarity is greater than or equal to the first similarity threshold; and determine that there is no association between the first data table and the second data table if the first similarity is less than the first similarity threshold.

[0085] Optionally, when determining the data relationship of the data to be processed based on the data table, the second determining module 403 is specifically used to: perform enumeration-type filtering on the third data in each column of the data table to obtain third enumeration-type data; compare the similarity of the third enumeration-type data corresponding to any two columns of third data to determine the second similarity between the third enumeration-type data corresponding to any two columns of third data; and determine the intra-table column data relationship of the data to be processed based on the second similarity.

[0086] Optionally, when the second determining module 403 determines the relationship between the column data within the table of the data to be processed based on the second similarity, it is specifically used to: determine that there is an association between the column data within the table of the data table if the second similarity is greater than or equal to the second similarity threshold; and determine that there is no association between the column data within the table of the data table if the second similarity is less than the second similarity threshold.

[0087] Optionally, when determining the data relationship of the data to be processed based on the data table, the second determining module 403 is specifically used to: perform enumeration-type filtering on the fourth data in each column of the data table to obtain the fourth enumeration class data; generate a scatter plot of the fourth enumeration class data, which is used to represent the data dispersion of the fourth enumeration class data; and determine the intra-column data relationship of the data to be processed based on the scatter plot.

[0088] Optionally, when the second determining module 403 determines the intra-column data relationship of the data to be processed based on the scatter plot, it is specifically used to: determine that there is no correlation between the intra-column data of the data table if the scatter plot indicates that the data dispersion of the fourth enumeration class data is greater than or equal to a preset value; and determine that there is a correlation between the intra-column data of the data table if the scatter plot indicates that the data dispersion of the fourth enumeration class data is less than a preset value.

[0089] Optionally, when the third determining module 404 determines the data quality of the data to be processed based on the data relationship, it is specifically used to: determine that the data to be processed has a data quality problem in response to the following: the inter-table data relationship is that there is no association between the first data table and the second data table; or the intra-table column data relationship is that there is no association between the intra-table column data of the data table; or the intra-column data relationship is that there is no association between the intra-column data of the data table.

[0090] Optionally, the data processing device 400 further includes an output module (not shown) for generating and outputting a data quality analysis report of the data to be processed after determining the data quality of the data to be processed based on the data relationships. The data quality analysis report includes at least one of the following information: the completeness of the data to be processed, the relevance of the data to be processed, the uniqueness of the data to be processed, and the validity of the data to be processed.

[0091] The data processing device provided in this embodiment is used to execute the technical solution of the data processing method in the aforementioned method embodiment. Its implementation principle and technical effect are similar, and will not be described again here.

[0092] This application also provides an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be configured as a terminal device.

[0093] Electronic device 500 may include one or more of the following components: processing component 502, memory 504, power supply component 506, multimedia component 508, audio component 510, input / output interface 512, sensor component 514, and communication component 516. Input / output interface 512 may also be referred to as I / O interface 512.

[0094] Processing component 502 typically controls the overall operation of electronic device 500, including operations related to display, data interaction, and data analysis. Processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 502 may include one or more modules to facilitate interaction between processing component 502 and other components. For example, processing component 502 may include a multimedia module to facilitate interaction between multimedia component 508 and processing component 502.

[0095] Memory 504 is configured to store various types of data to support the operation of electronic device 500. Examples of this data include instructions for any application or method operating on electronic device 500, data tables of data to be processed, table metadata of the data tables, business-related information of the data to be processed, data relationships of the data to be processed, data quality analysis reports of the data to be processed, etc. Memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0096] Power supply component 506 provides power to various components of electronic device 500. Power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 500.

[0097] Multimedia component 508 includes a screen that provides an output interface between electronic device 500 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch pad (TP). If the screen includes a touch pad, the screen may be implemented as a touchscreen to receive input signals from the user. The touch pad includes one or more touch sensors to sense touches, swipes, and gestures on the touch pad.

[0098] Audio component 510 is configured to output and / or input audio signals. For example, audio component 510 includes a microphone (MIC) configured to receive external audio signals when electronic device 500 is in an operating mode, such as a voice recognition mode. The received audio signals may be further stored in memory 504 or transmitted via communication component 516. In some embodiments, audio component 510 also includes a speaker for outputting audio signals.

[0099] I / O interface 512 provides an interface between processing component 502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0100] Sensor assembly 514 includes one or more sensors for providing state detection of various aspects of electronic device 500. For example, sensor assembly 514 can detect the on / off state of electronic device 500, the relative positioning of components such as the display and keypad of electronic device 500, and the presence or absence of user contact with electronic device 500.

[0101] Communication component 516 is configured to facilitate wired or wireless communication between electronic device 500 and other devices. Electronic device 500 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 516 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth, and other technologies.

[0102] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processor devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0103] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, which can be executed by a processor 520 of an electronic device 500 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0104] A non-transitory computer-readable storage medium, wherein when the instructions in the storage medium are executed by the processor of an electronic device 500, the electronic device 500 is able to perform the aforementioned data processing method.

[0105] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed, are used to implement the technical solution of the data processing method provided in the foregoing method embodiments.

[0106] This application also provides a computer program product, including a computer program, which, when executed, is used to implement the technical solution of the data processing method provided in the foregoing method embodiments.

[0107] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0108] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0109] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0110] Furthermore, unless otherwise specified, the functional modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated unit / module described above can be implemented in hardware or as a software program module.

[0111] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.

[0112] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0113] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as these combinations of technical features do not contradict each other, they should be considered within the scope of this specification.

[0114] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0115] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A data processing method, characterized in that, include: Obtain the data table containing the data to be processed and the table metadata of the data table; For the table metadata, a text-to-text engine is used to determine the business-related information of the data to be processed. Based on the data table, determine the data relationship of the data to be processed, wherein the data relationship includes at least one of inter-table data relationship, intra-table column data relationship, and intra-column data relationship; The data quality of the data to be processed is determined based on whether the data relationship exists. Information mining is performed on the data to be processed based on the business-related information, the data relationships, and the data quality. The data table includes a first data table and a second data table. Determining the data relationships of the data to be processed based on the data tables includes: For each column of first data in the first data table and each column of second data in the second data table, the first data is subjected to enumeration filtering to obtain first enumeration data, and the second data is subjected to enumeration filtering to obtain second enumeration data; the first enumeration data and the second enumeration data are compared for similarity to determine the first similarity between the first enumeration data and the second enumeration data. Based on the first similarity, the inter-table data relationships of the data to be processed are determined; For each column of the third data in the data table, the third data is subjected to enumeration filtering to obtain the third enumeration data. A similarity comparison is performed on the third enumeration class data corresponding to any two columns of third data to determine the second similarity between the third enumeration class data corresponding to any two columns of third data. Based on the second similarity, the intra-table column data relationships of the data to be processed are determined; For each column of the fourth data in the data table, the fourth data is subjected to an enumeration-type filtering process to obtain the fourth enumeration-type data. Generate a scatter plot of the fourth enumeration class data, the scatter plot being used to represent the data dispersion of the fourth enumeration class data; Based on the scatter plot, determine the intra-column data relationships of the data to be processed.

2. The data processing method according to claim 1, characterized in that, Determining the inter-table data relationship of the data to be processed based on the first similarity includes: In response to the first similarity being greater than or equal to the first similarity threshold, the data relationship between the tables is determined to be an association between the first data table and the second data table; If the first similarity is less than the first similarity threshold, then the data relationship between the tables is determined to be that there is no association between the first data table and the second data table.

3. The data processing method according to claim 1, characterized in that, The step of determining the intra-table column data relationship of the data to be processed based on the second similarity includes: In response to the second similarity being greater than or equal to the second similarity threshold, the relationship between the column data in the table is determined to be that there is an association between the column data in the table. If the second similarity is less than the second similarity threshold, then it is determined that there is no association between the column data in the table.

4. The data processing method according to claim 1, characterized in that, Determining the intra-column data relationships of the data to be processed based on the scatter plot includes: In response to the scatter plot indicating that the data dispersion of the fourth enumeration class is greater than or equal to a preset value, it is determined that there is no correlation between the data in the columns of the data table; In response to the scatter plot indicating that the data dispersion of the fourth enumeration class is less than the preset value, the data relationship within the column is determined to be an association between the data within the columns of the data table.

5. The data processing method according to any one of claims 1 to 4, characterized in that, Determining the data quality of the data to be processed based on the data relationship includes: In response to the following conditions: there is no association between the first and second data tables in the inter-table data relationship; or, there is no association between the columns within the data tables; or, there is no association between the columns within the data tables, it is determined that the data to be processed has a data quality problem.

6. The data processing method according to claim 5, characterized in that, After determining the data quality of the data to be processed based on the data relationship, the process further includes: Generate and output a data quality analysis report of the data to be processed. The data quality analysis report includes at least one of the following: the completeness of the data to be processed, the relevance of the data to be processed, the uniqueness of the data to be processed, and the validity of the data to be processed.

7. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire data tables including the data to be processed and table metadata of the data tables; The first determining module is used to determine the business-related information of the data to be processed from the table metadata using a text-to-text engine. The second determining module is used to determine the data relationship of the data to be processed based on the data table, wherein the data relationship includes at least one of inter-table data relationship, intra-table column data relationship and intra-column data relationship; The third determining module is used to determine the data quality of the data to be processed based on whether the data relationship exists; The data mining module is used to perform information mining on the data to be processed based on the business-related information, the data relationships, and the data quality. The data table includes a first data table and a second data table. The second determining module is specifically used to perform enumeration-type filtering on the first data in each column of the first data table and on the second data in each column of the second data table to obtain first enumeration-type data, and to perform enumeration-type filtering on the second data to obtain second enumeration-type data; and to compare the similarity between the first enumeration-type data and the second enumeration-type data to determine a first similarity between the first enumeration-type data and the second enumeration-type data. Based on the first similarity, the inter-table data relationships of the data to be processed are determined; The second determining module is specifically used to perform enumeration-type filtering on each column of the third data in the data table to obtain the third enumeration-type data; A similarity comparison is performed on the third enumeration class data corresponding to any two columns of third data to determine the second similarity between the third enumeration class data corresponding to any two columns of third data. Based on the second similarity, the intra-table column data relationships of the data to be processed are determined; The second determining module is specifically used to perform enumeration-type filtering on the fourth data in each column of the data table to obtain the fourth enumeration-type data; Generate a scatter plot of the fourth enumeration class data, the scatter plot being used to represent the data dispersion of the fourth enumeration class data; Based on the scatter plot, determine the intra-column data relationships of the data to be processed.

8. An electronic device, characterized in that, include: A processor, and a memory connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed, are used to implement the method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, Includes a computer program, which, when executed, is used to implement the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data quality management method and device

    CN108595563A

  • Data processing method and device, equipment, medium and product

    CN114691274A

  • Method and device for generating presentation document, electronic equipment and storage medium

    CN116306492A