Data table processing method and device, equipment, medium and program product

By processing the underlying tables of data tables in both the non-data platform and the data platform, the overlap of target lineage relationships is determined, and data is deposited into the data platform when there is no complete overlap. This solves the problem of high-value data assets not being deposited and improves the asset value of the data platform.

CN121070931APending Publication Date: 2025-12-05INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511227813.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

During the digital transformation of the financial sector, some high-value data assets have not been settled in the data platform, making it difficult to meet the rapid and extensive data needs of the business.

Method used

By acquiring the underlying processing tables of data tables in both the non-data platform and the data platform, the degree of overlap in the data relationship is determined. In cases of incomplete overlap, high-value data tables are deposited into the data platform, and the consistency of field logic is handled by a large model to achieve data table unification.

Benefits of technology

It has increased the asset value of the data platform, met the data needs of the business, ensured unified data management, reduced data redundancy, and increased the asset value of the data platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121070931A_ABST
    Figure CN121070931A_ABST
Patent Text Reader

Abstract

The invention provides a data table processing method which can be applied to the technical field of big data and the technical field of financial science and technology, and relates to application of a large model in a financial science and technology scene. The data table processing method comprises the following steps: acquiring m first data tables of a non-data table and a bottom processing table of each first data table; n second data tables of the data table and a bottom processing table of each second data table are obtained, the bottom processing tables of the data tables are obtained based on processing links of the data tables, and the processing links of the data tables are used for representing the blood relationship of the data tables; for each first data table, determining a target blood relationship coincidence degree based on a bottom processing table of the first data table and a bottom processing table of each second data table; and under the condition that the target blood relationship overlap ratio is not equal to 1, depositing the first data table into a data table. The invention further provides a data table processing device and equipment, a medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data and the technical field of financial technology, relates to application of a large model in a financial technology scene, and more particularly to a data table processing method and device, equipment, medium and program product. BACKGROUND

[0002] With the deepening of the digital transformation work in the financial field, the use of data by financial related businesses is increasingly demanding. On the one hand, due to historical reasons, a part of data assets with sharing and high value are not built in the data platform, resulting in that the business cannot use data based on the data platform to meet the needs of rapid business innovation. On the other hand, there may be some data assets that have no sharing needs at the initial stage of construction, and are not built in the data platform. With the advancement of digital transformation, this part of data assets gradually becomes data assets with sharing and high value, but they are not deposited in the data platform, making it difficult to meet the rapid and extensive data needs of the business. SUMMARY

[0003] In view of the above problems, the present application provides a data table processing method, device, equipment, medium and program product for improving the value of data platform assets to meet the data needs of the business.

[0004] According to a first aspect of the present application, a data table processing method is provided, comprising: obtaining m first data tables of a non-data platform and a bottom layer processing table of each first data table, the bottom layer processing table of the first data table being obtained based on a processing link of the first data table, the processing link of the first data table being used to represent the blood relationship of the first data table, m being an integer greater than or equal to 1; obtaining n second data tables of a data platform and a bottom layer processing table of each second data table, the bottom layer processing table of the second data table being obtained based on a processing link of the second data table, the processing link of the second data table being used to represent the blood relationship of the second data table, n being an integer greater than or equal to 1; for each first data table, determining a target blood relationship coincidence degree based on the bottom layer processing table of the first data table and the bottom layer processing table of each second data table; and in the case where the target blood relationship coincidence degree is not equal to 1, depositing the first data table to the data platform.

[0005] According to an embodiment of the present application, for each first data table, determining a target blood relationship coincidence degree based on the bottom layer processing table of the first data table and the bottom layer processing table of each second data table comprises: for each first data table, determining a blood relationship coincidence degree between the first data table and each second data table based on the coincidence degree between the bottom layer processing table of the first data table and the bottom layer processing table of each second data table; sorting n blood relationship coincidence degrees; and determining a target blood relationship coincidence degree from the n blood relationship coincidence degrees based on the sorting result.

[0006] According to an embodiment of the present application, the m first data tables of the non-data middle platform are obtained by determining the number of calling parties and the calling volume of each data table in the non-data middle platform based on the log of the big data platform; and the data table is the first data table when the number of calling parties and the calling volume of the data table meet preset conditions.

[0007] According to an embodiment of the present application, the underlying processing table includes at least one source table, and the underlying processing table of each first data table is obtained by analyzing the processing link of the first data table to obtain at least one business system table or source table, and the source table is a source layer data table of the data middle platform; for the business system table, the source table corresponding to the business system table is determined based on the mapping relationship between the business system and the source table.

[0008] According to an embodiment of the present application, the data middle platform includes a source layer, an aggregation layer, and an extraction layer, the second data table is an aggregation layer data table or an extraction layer data table, and the n second data tables of the data middle platform and the underlying processing table of each second data table are obtained by analyzing the processing link of each second data table to obtain the underlying processing table of the second data table, and the underlying processing table includes at least one source table, and the source table is a source layer data table of the data middle platform.

[0009] According to an embodiment of the present application, when the target blood relationship coincidence degree is not equal to 1, the first data table is deposited into the data middle platform, including: when there is a data table of the non-data middle platform in the underlying processing table of the first data table, initiating a lake entry process for the first data table, the lake entry process representing storing the underlying processing table of the first data table into a data lake, and the data lake is used to store the data table of the data middle platform; and when there is no data table of the non-data middle platform in the underlying processing table of the first data table, depositing the first data table into the data middle platform.

[0010] According to an embodiment of the present application, initiating the lake entry process for the first data table includes: when the underlying processing table of the first data satisfies the lake entry condition, storing the underlying processing table of the first data table into the data lake and depositing the first data table into the data middle platform.

[0011] According to an embodiment of the present application, when the target blood relationship coincidence degree is equal to 1, the same processing logic fields in the first data table and the target data table are obtained by using a large model, and the field correspondence relationship between the first data table and the second data table is obtained, and the target data table is the second data table corresponding to the target blood relationship coincidence degree; for the fields with the same processing logic, the field in the first data table is switched to the corresponding field in the target data table based on the correspondence relationship; and for the fields in the first data table that do not exist in the correspondence relationship, the fields are deposited into the target data table.

[0012] The second aspect of the present application provides a data table processing apparatus, comprising: a first obtaining module configured to obtain m first data tables of a non-data middle station and an underlying processing table of each first data table, the underlying processing table of the first data table being obtained based on a processing link of the first data table, the processing link of the first data table being used to represent a blood relationship of the first data table, m being an integer greater than or equal to 1; a second obtaining module configured to obtain n second data tables of a data middle station and an underlying processing table of each second data table, the underlying processing table of the second data table being obtained based on a processing link of the second data table, the processing link of the second data table being used to represent a blood relationship of the second data table, n being an integer greater than or equal to 1; a determining module configured to, for each first data table, determine a target blood relationship coincidence degree based on the underlying processing table of the first data table and the underlying processing table of each second data table; and a sedimentation module configured to, in a case where the target blood relationship coincidence degree is not equal to 1, sediment the first data table to the data middle station.

[0013] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.

[0014] The fourth aspect of the present application further provides a computer-readable storage medium having stored thereon a computer program or instructions, which, when executed by a processor, implement the steps of the method.

[0015] The fifth aspect of the present application further provides a computer program product comprising a computer program or instructions, which, when executed by a processor, implement the steps of the method.

[0016] According to the embodiments of the present application, by comparing the blood relationship of the data tables in the non-data middle station with the blood relationship of the data tables in the data middle station, the high-value data tables in the non-data middle station that do not completely coincide with the target blood relationship of the data tables in the data middle station are sedimented to the data middle station, the asset value of the data middle station is improved, and the data demand of the related business is met. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above content of the present application and other purposes, features and advantages will be more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:

[0018] Figure 1 An application scenario diagram of the data table processing method, apparatus, device, medium and program product according to the embodiments of the present application is schematically shown;

[0019] Figure 2 A flowchart of the data table processing method according to the embodiments of the present application is schematically shown.

[0020] Figure 3 A flow chart of a data table processing method according to another embodiment of the present application is schematically shown;

[0021] Figure 4 A schematic diagram of obtaining a bottom layer processing table of a first data table according to an embodiment of the present application is schematically shown;

[0022] Figure 5 A schematic diagram of obtaining a bottom layer processing table of a second data table according to an embodiment of the present application is schematically shown;

[0023] Figure 6 A block diagram of a data table processing apparatus according to an embodiment of the present application is schematically shown; and

[0024] Figure 7 A block diagram of an electronic device suitable for implementing the data table processing method according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0025] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, the drawings are for illustration only and description purposes and are not intended to limit the scope of the application. In the following detailed description of embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the embodiments. However, it will be apparent to one skilled in the art that one or more embodiments can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring the concepts of the application.

[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "includes" and tautological equivalents thereof, means that the named feature, step, operation, and / or component is included, but not to the exclusion of the presence or addition of one or more other features, steps, operations, or components.

[0027] All terms used herein including technical and scientific terms have the same meanings as commonly understood by one of ordinary skill in the art unless otherwise defined herein. It should be noted that the terms used herein are merely specific embodiments proposed as examples in order to describe the most specific information of the present application and should not be interpreted in an idealized or overly formal manner.

[0028] In the case of using expressions similar to "at least one of A, B, and C, etc.", it is generally construed that the meaning of the expression is the same as that of the expression "one or more of A, B, and C" understood by one of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include a system having A alone, a system having B alone, a system having C alone, a system having A and B together, a system having A and C together, a system having B and C together, and / or a system having A, B, and C together, etc.).

[0029] Embodiments of the present application provide a data table processing method, comprising: obtaining m first data tables of a non-data middle platform and an underlying processing table of each first data table, the underlying processing table of the first data table being obtained based on a processing link of the first data table, the processing link of the first data table being used to represent the blood relationship of the first data table; obtaining n second data tables of a data middle platform and an underlying processing table of each second data table, the underlying processing table of the second data table being obtained based on a processing link of the second data table, the processing link of the second data table being used to represent the blood relationship of the second data table; for each first data table, determining a target blood relationship coincidence degree based on the underlying processing table of the first data table and the underlying processing table of each second data table; and in the case where the target blood relationship coincidence degree is not equal to 1, depositing the first data table to the data middle platform.

[0030] By comparing the blood relationship of the data table in the non-data middle platform with the blood relationship of the data table in the data middle platform, the high-value data table in the non-data middle platform which does not completely coincide with the target blood relationship of the data table in the data middle platform is deposited to the data middle platform, thereby improving the asset value of the data middle platform.

[0031] It should be noted that the data table processing method, device, equipment, medium and program product determined by the present application involve the application of data blood relationship in the field of financial technology, and can be used in the field of big data technology and the field of financial technology, and can also be used in various fields other than the field of big data technology and the field of financial technology. The application field of the data table processing method, device, equipment, medium and program product provided by the embodiments of the present application is not limited.

[0032] In the technical solutions of the present application, the user information (including but not limited to user personal information, user image information, user equipment information such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal.

[0033] In the scenario of making automated decisions by using personal information, the method, device and system provided by the embodiments of the present application all provide corresponding operation entrances for the user to select to agree or reject the automated decision result; if the user selects to reject, the expert decision process is entered. The expression "automated decision" herein refers to the activity of making decisions by automatically analyzing and evaluating the behavior habits, interests and hobbies, or economic, health and credit conditions of a person by using a computer program. The expression "expert decision" herein refers to the activity of making decisions by personnel who are engaged in a certain field of work, have specialized experience, knowledge and skills, and have reached a certain professional level.

[0034] Figure 1 An application scenario diagram of the data table processing method, apparatus, device, medium and program product according to the embodiments of the present application is schematically shown.

[0035] As shown in Figure 1 The application scenario 100 according to the embodiments can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0036] The user can use the first terminal device 101, the second terminal device 102 and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102 and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0037] The first terminal device 101, the second terminal device 102 and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers and desktop computers, etc.

[0038] The server 105 can be a server providing various services, such as a background management server supporting the website browsed by the user using the first terminal device 101, the second terminal device 102 and the third terminal device 103 (only as an example). The background management server can analyze and process the received user request and other data, and feed back the processing result (such as a web page, information or data generated or obtained according to the user request) to the terminal device.

[0039] It should be noted that the data table processing method provided in the embodiments of the present application can be generally executed by the server 105. Correspondingly, the data table processing apparatus provided in the embodiments of the present application can be generally arranged in the server 105. The data table processing method provided in the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Correspondingly, the data table processing apparatus provided in the embodiments of the present application can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0040] It should be understood that Figure 1 The number of terminal devices, networks and servers in the above-mentioned scenario is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks and servers.

[0041] The data table processing method according to the embodiments of the present application will be described in detail below based on the scenario described above. Figure 1 Figures 2-5 The data table processing method according to the embodiments of the present application will be described in detail below based on the scenario described above.

[0042] Figure 2 The flowchart of the data table processing method according to the embodiments of the present application is schematically shown.

[0043] As shown in Figure 2 The data table processing method of this embodiment includes operations S210-S240, and the data processing method can be executed by the server 105.

[0044] In operation S210, m first data tables of a non-data middle platform and an underlying processing table of each first data table are obtained, the underlying processing table of the first data table is obtained based on a processing link of the first data table, and the processing link of the first data table is used to represent the blood relationship of the first data table.

[0045] In the embodiments of the present application, m is an integer greater than or equal to 1.

[0046] For example, the data table of the non-data middle platform can be obtained from the big data platform, the big data platform can be a comprehensive system integrating data collection, storage, processing, analysis, mining and application functions, the first data table can be a data table with high value, and the data table with high value can be obtained by analyzing the use log of the big data platform.

[0047] ​The data table processed in the big data platform needs to be configured with a data table processing dependency, and therefore, the data source of the data table can be parsed through the processing dependency. Specifically, the processing dependency of the data table can be parsed by using a parsing tool to obtain the processing link of the first data table, and the processing link of the data table is traced layer by layer to finally locate the bottom layer data table of the first data table.

[0048] For another example, the data table of the non-data middle platform can be processed based on the data of the data middle platform, or can be processed based on the data table of the business system, or can be processed in a hybrid manner, that is, based on the data table of the business system and the data table of the data middle platform.

[0049] In operation S220, n second data tables of the data middle platform and bottom layer processing tables of each second data table are obtained, the bottom layer processing table of the second data table is obtained based on the processing link of the second data table, and the processing link of the second data table is used to represent the blood relationship of the second data table.

[0050] In the embodiments of the present application, n is an integer greater than or equal to 1.

[0051] The data middle platform can be a comprehensive data infrastructure that integrates multi-source data, realizes standardized management and assetized operation, and encapsulates data capabilities as services for rapid business calls.

[0052] The n second data tables of the data middle platform can be obtained directly from the data middle platform, for example, the second data table can be obtained from the aggregation layer for aggregating and summarizing data according to the hierarchical design of the data middle platform, or the second data table can be obtained from the extraction layer for extracting data from the aggregation layer or from other data storage.

[0053] Similarly, the processing dependency of the second data table can be parsed by using a parsing tool to obtain the processing link of the second data table, and the processing link of the second data table is traced layer by layer to finally locate the bottom layer data table of the second data table

[0054] In operation S230, for each first data table, the target blood relationship coincidence degree is determined based on the bottom layer processing table of the first data table and the bottom layer processing table of each second data table.

[0055] For example, for each first data table, the coincidence degree between the bottom layer processing table of the first data table and the bottom layer processing table of each second data table can be calculated as the blood relationship coincidence degree between the first data table and the second data table. It can be understood that for each first data table, multiple blood relationship coincidence degrees can be obtained, and then based on the value of the blood relationship coincidence degree, the target blood relationship coincidence degree can be determined from the multiple blood relationship coincidence degrees.

[0056] In operation S240, in the case where the target blood relationship coincidence degree is not equal to 1, the first data table is deposited to the data platform.

[0057] The target blood relationship coincidence degree is not equal to 1, that is, the target blood relationship does not completely coincide, which indicates that the data table is not constructed in the data platform. In this case, the first data table is deposited to the data platform, which can be understood as falling or storing the first data table into the data platform.

[0058] It can be understood that, by comparing the blood relationship of the data table in the non-data platform with the blood relationship of the data table in the data platform, the high-value data table in the non-data platform which does not completely coincide with the target blood relationship of the data table in the data platform is deposited to the data platform, thereby improving the asset value of the data platform.

[0059] According to the embodiment of the present application, for each first data table, the target blood relationship coincidence degree is determined based on the bottom processing table of the first data table and the bottom processing table of each second data table. For each first data table, the blood relationship coincidence degree between the first data table and each second data table is determined based on the coincidence degree between the bottom processing table and the bottom processing table of each second data table. The n blood relationship coincidence degrees are sorted. Based on the sorting result, the target blood relationship coincidence degree is determined from the n blood relationship coincidence degrees.

[0060] The first data table can be a high-value data table in the non-data platform. The bottom processing table of the first data table can be obtained by analyzing the blood relationship of the first data table. Similarly, the bottom processing table of the second data table can be obtained by analyzing the blood relationship of the second data table.

[0061] The blood relationship coincidence degree between the first data table and each second data table is determined based on the coincidence degree between the bottom processing table and the bottom processing table of each second data table. It can be understood that, for each first data table, since there are n second data tables, n blood relationship coincidence degrees are obtained. The n blood relationships are sorted to obtain a sorting result. Based on the sorting result, the maximum coincidence degree in the n blood relationship coincidence degrees can be determined as the target blood relationship coincidence degree.

[0062] It can be understood that, by determining the blood relationship coincidence degree between the data tables through the coincidence degree between the bottom processing table of the first data table and the bottom processing table of the second data table, it is not necessary to compare the contents of the data tables one by one, so that whether the data table is constructed in the data platform can be efficiently determined.

[0063] According to an embodiment of the present application, the m first data tables of the non-data middle platform are obtained by determining the number of calling parties and the calling quantity of each data table in the non-data middle platform based on the log of the big data platform; and in a case where the number of calling parties and the calling quantity of the data table satisfy a preset condition, the data table is a first data table.

[0064] For example, as shown in Table 1, Table 1 schematically shows the use number log of the big data platform. According to the use number log of the big data platform, the application calling the data table, the calling times and the system to which the data table belongs can be counted.

[0065] Table 1

[0066]

[0067] According to Table 1, the number of calling parties and the calling quantity of each data table can be counted, and the first data table is selected from the data tables of the plurality of non-data middle platforms according to the calling quantity and the number of calling parties.

[0068] Table 2

[0069]

[0070] For example, the preset condition can be that the number of calling parties is greater than 2 and the calling quantity is within the top 50. The number of calling parties greater than 2 indicates that the data table has sharing value, and the calling quantity within the top 50 indicates that the data table has business value. Thus, the first data table obtained is a high-value data table.

[0071] It can be understood that the purpose of building a data middle platform is data sharing, and a unified data export. The non-middle platform data table with sharing value and high business value is deposited into the data middle platform to expand the support range of the data middle platform, so as to improve the asset value of the data middle platform, and avoid the entry of data tables with no value or low value into the data middle platform to cause data redundancy.

[0072] According to an embodiment of the present application, the underlying processing table includes at least one source table, and the underlying processing table of each first data table is obtained by: analyzing the processing link of the first data table to obtain at least one business system table or source table, the source table being a source layer data table of the data middle platform; and for the business system table, determining the source table corresponding to the business system table based on the mapping relationship between the business system and the source table.

[0073] After the processing link of the first data table is analyzed by the analysis tool, the underlying processing table of the first data table can be obtained. It can be understood that, since the first data table is a data table of a non-data middle platform, the most underlying processing table obtained can be a source layer data table of the data middle platform, i.e., a source table, or can be a business system table, or can include both the source table and the business system table.

[0074] For the business system table, the corresponding source table needs to be determined based on the mapping relationship between the source table and the business system table.

[0075] It can be understood that by analyzing the first data table blood relationship and relying on the mapping relationship between the source table and the business system table, the first data table bottom processing table can be positioned to the source layer of the data center.

[0076] According to the embodiment of the application, the data center includes a source layer, an aggregation layer, and an extraction layer. The second data table is an aggregation layer data table or an extraction layer data table. The n second data tables of the data center and the bottom processing table of each second data table are obtained by: for each second data table, analyzing the processing link of the second data table to obtain the bottom processing table of the second data table, and the bottom processing table includes at least one source table, and the source table is a source layer data table of the data center.

[0077] The data center can be a full-amount data center, which can be designed as a hierarchical structure of a source layer, an aggregation layer, and an extraction layer. The source layer can be a raw data source in a data warehouse or a data lake, and the data stored therein is usually unprocessed raw data table, which needs to be processed before being used for analysis and processing. The aggregation layer can be used to aggregate and summarize data, and the extraction layer can be used to extract data from the aggregation layer or other data storage for analysis and other purposes.

[0078] After the processing link of the second data table is analyzed by the analysis tool, the bottom processing table of the second data table can be obtained. Since the second data table is a data table of the data center, the bottom processing table of the second data table is a source table.

[0079] It can be understood that by positioning the bottom processing table of the first data table and the bottom processing table of the second data table to the source layer, a unified comparison benchmark can be formed.

[0080] According to the embodiment of the application, in the case where the target blood relationship coincidence degree is not equal to 1, the first data table is deposited into the data center, including: in the case where the bottom processing table of the first data table includes a data table that is not a data table of the data center, initiating a lake flow process for the first data table, the lake flow process representing storing the bottom processing table of the first data table into a data lake, and the data lake is used to store the data table of the data center; in the case where the bottom processing table of the first data table does not include a data table that is not a data table of the data center, depositing the first data table into the data center.

[0081] In the case where the target blood relationship coincidence degree is not equal to 1, it is indicated that the first data table and the target data table are not completely coincident, and the target data table is the second data table corresponding to the target blood relationship coincidence degree, which further indicates that the first data table is not constructed in the data center.

[0082] In this case, it is further determined whether the first data table exists in the data table of the non-data center in the bottom layer processing table. If the first data table does not exist in the data table of the non-data center in the bottom layer processing table, the first data table is deposited to the data center. Specifically, a task flow can be initiated to the data center construction team, requiring the data table to be placed in the data center according to the processing logic.

[0083] If the first data table exists in the data table of the non-data center in the bottom layer processing table, it means that the first data table needs to be initiated to the lake process, that is, the bottom layer processing table of the first data table is stored to the data lake.

[0084] It can be understood that when the blood relationship between the first data table and the target data table does not completely coincide, it means that the high-value data table does not fall into the data center, and depositing the data table can improve the asset value of the data center.

[0085] According to the embodiments of the present application, initiating the lake process to the first data table includes: storing the bottom layer processing table of the first data to the data lake and depositing the first data table to the data center when the bottom layer processing table of the first data meets the lake condition.

[0086] For example, the lake condition can be that the data table meets the data format standardization, the metadata is complete, and it meets the data security compliance requirements. If the bottom layer processing table of the first data meets the lake condition, the bottom layer processing table of the first data is stored to the data lake and the first data table is deposited to the data center. If the bottom layer processing table of the first data does not meet the lake condition, it means that the first data table is a high-value data table, but it is not compliant, that is, the first data table does not have the condition to be deposited to the data center, so it cannot be deposited to the data center.

[0087] It can be understood that by judging the lake condition of the bottom layer processing table of the first data, non-compliant data can be prevented from entering the center, ensuring unified data management and reducing subsequent data governance costs.

[0088] According to the embodiments of the present application, in the case where the target blood relationship coincidence degree is equal to 1, the same processing logic fields in the first data table and the target data table are obtained by using a large model, and the field correspondence relationship between the first data table and the second data table is obtained. The target data table is the second data table corresponding to the target blood relationship coincidence degree. For the fields with the same processing logic, the fields in the first data table are switched to the corresponding fields in the target data table based on the correspondence relationship. For the fields in the first data table that do not exist in the correspondence relationship, the fields are deposited to the target data table.

[0089] The target blood relationship coincidence degree is equal to 1, indicating that the first data table and the target data table are completely coincident in the bottom processing table, that is, the data asset of the data center construction is also constructed in the non-data center, but a part of the business system uses the data asset constructed by the data center, and a part of the business system uses the data asset constructed by the non-data center, which affects the uniformity of the business caliber.

[0090] The processing logic of the first data table and the target data table can be analyzed by using a large model to obtain a field correspondence relationship. For fields with the same processing logic, the fields in the first data table are switched to the corresponding fields in the target data table based on the correspondence relationship. For fields in the first data table that do not exist in the correspondence relationship, the fields are deposited to the target data table.

[0091] It can be understood that by switching the fields, the business system can be prompted to use the indicators of the middle platform, and the consistency of the business caliber can be ensured.

[0092] Figure 3 A flowchart of a data table processing method according to another embodiment of the application is schematically shown.

[0093] As shown in Figure 3 , the data table processing method of this embodiment includes operations S301-S310, and the data processing method can be executed by a server in a data table processing device.

[0094] In operation S301, m first data tables of a non-data center and a bottom processing table of each first data table are obtained.

[0095] Wherein, m is an integer greater than or equal to 1.

[0096] Figure 4 A schematic diagram of obtaining the bottom processing table of the first data table according to an embodiment of the application is schematically shown.

[0097] As shown in Figure 4 , the source table is a data table of the source layer, the aggregation table is a data table of the aggregation layer, and the extraction table is a data table of the extraction layer.

[0098] The processing link of the data table Z, the data table E, and the data table M of the non-data center is analyzed. The bottom processing table of the data table Z is the data table L (i.e., the business system table L), the source table H, and the source table U. The bottom processing table of the data table E is the source table A, the source table B, and the source table F. The bottom processing table of the data table M is the source table F and the data table N (i.e., the business system table N).

[0099] For the data table Z, the underlying processing table of the data table Z can be positioned as the original table R, the original table H and the original table U according to the mapping relationship between the original table and the business system table (the business system table L has a mapping relationship with the original table R).

[0100] For the data table M, assuming that there is no mapping relationship of the data table N in the mapping relationship, the underlying processing table of the data table M is the original table F and the data table N.

[0101] In operation S302, n second data tables of the data center and an underlying processing table of each second data table are obtained.

[0102] Wherein, n is an integer greater than or equal to 1.

[0103] The underlying processing table of the second data table is obtained based on the processing link of the second data table, and the processing link of the second data table is used to represent the blood relationship of the second data table.

[0104] Figure 5 An illustrative diagram for obtaining the underlying processing table of the second data table according to the embodiment of the application is shown.

[0105] As Figure 5 shown, the processing link of the extraction table G and the extraction table J is parsed, and the underlying processing table of the extraction table G is obtained as the original table R, the original table H and the original table U, and the underlying processing table of the extraction table J is obtained as the original table A and the original table B.

[0106] In operation S303, for each first data table, the target blood relationship coincidence degree is determined based on the underlying processing table of the first data table and the underlying processing table of each second data table.

[0107] Specifically, based on the results of operation S301 and operation S302, for each data table, the one table with the highest blood relationship coincidence degree with the second data table is obtained.

[0108] As shown in Table 3, the underlying processing table of the first data table E has three, which are the original table A, B and F; the underlying processing table of the second data table J has two, which are the original table A and B, and the target blood relationship coincidence degree of the table E and J is 2 / 3, that is, 66%.

[0109] Table 3

[0110]

[0111] In operation S304, it is judged whether the target blood relationship coincidence degree is 100%.

[0112] In operation S305, if the target blood relationship coincidence degree is 100%, it is judged whether the field in the first data table and the field in the second data table have the same processing logic.

[0113] As shown in Table 4, Table 4 schematically shows the field correspondence between the first data table Z and the second data table G

[0114] Table 4

[0115]

[0116] In operation S306, for the fields with consistent processing logic, a switching process is initiated.

[0117] For example, for xxx_cnt in Table 4, it needs to be switched to xxx_count, and for xxx_num in Table 4, it needs to be switched to xxx_number.

[0118] In operation S307, for the fields with inconsistent processing logic, a sedimentation process is initiated.

[0119] That is, for the fields in the first data table that do not exist in the corresponding relationship, it means that the field is not built in the target data table and needs to be sedimented to the target data table.

[0120] In operation S308, if the target blood relationship coincidence degree is not 100%, it is judged whether the bottom processing table of the first data table includes a non-paste source table.

[0121] In operation S309, for the paste source table, a sedimentation process is initiated.

[0122] The target blood relationship coincidence degree is not 100%, which means that the blood relationship of the first data table and the target data table does not completely coincide, that is, the bottom processing table of the first data table and the bottom processing table of the target data table do not completely coincide. In this case, if the bottom processing table of the first data table is a paste source table, a sedimentation and switching process needs to be initiated for the first data table, as shown in Table 3, the data table E needs to be sequentially initiated for the sedimentation and switching process.

[0123] In operation S310, for the non-paste source table, in the case that the first data table satisfies the lake entry condition, a lake entry process is initiated.

[0124] In the case that the bottom processing table of the first data table includes a non-paste source table, it needs to be judged whether the non-paste source table satisfies the lake entry condition. If it satisfies the lake entry condition, the non-paste source table can enter the lake. After the lake entry process is completed, the sedimentation and switching processes are sequentially initiated for the first data table. If it does not satisfy the lake entry condition, the first data table does not have the condition for data center sedimentation, and no sedimentation is performed, and the process is ended.

[0125] Based on the above data table processing method, the application also provides a data table processing device. The following will be combined Figure 6 The device will be described in detail.

[0126] Figure 6 A structural block diagram of a data table processing apparatus according to an embodiment of the present application is shown schematically.

[0127] As shown in Figure 6 The data table processing apparatus 600 of this embodiment includes a first acquisition module 610, a second acquisition module 620, a determination module 630, and a sedimentation module 640.

[0128] The first acquisition module 610 is configured to acquire m first data tables of a non-data middle platform and a bottom processing table of each first data table, the bottom processing table of the first data table being acquired based on a processing link of the first data table, the processing link of the first data table being used to represent a blood relationship of the first data table, and m being an integer greater than or equal to 1. In an embodiment, the first acquisition module 610 can be configured to perform the operation S210 described above, and thus no further description is given here.

[0129] The second acquisition module 620 is configured to acquire n second data tables of a data middle platform and a bottom processing table of each second data table, the bottom processing table of the second data table being acquired based on a processing link of the second data table, the processing link of the second data table being used to represent a blood relationship of the second data table, and n being an integer greater than or equal to 1. In an embodiment, the second acquisition module 620 can be configured to perform the operation S220 described above, and thus no further description is given here.

[0130] The determination module 630 is configured to, for each first data table, determine a target blood relationship coincidence degree based on the bottom processing table of the first data table and the bottom processing table of each second data table. In an embodiment, the determination module 630 can be configured to perform the operation S230 described above, and thus no further description is given here.

[0131] The sedimentation module 640 is configured to, in a case where the target blood relationship coincidence degree is not equal to 1, sediment the first data table to the data middle platform. In an embodiment, the sedimentation module 640 can be configured to perform the operation S240 described above, and thus no further description is given here.

[0132] According to an embodiment of the present application, the determination module 630 further includes a blood relationship coincidence degree determination module, a sorting module, and a target blood relationship coincidence degree determination module. The blood relationship coincidence degree determination module is configured to, for each first data table, determine a blood relationship coincidence degree between the first data table and each second data table based on a coincidence degree between the bottom processing table and the bottom processing table of each second data table; the sorting module is configured to sort the n blood relationship coincidence degrees; and the target blood relationship coincidence degree determination module is configured to determine the target blood relationship coincidence degree from the n blood relationship coincidence degrees based on a sorting result.

[0133] According to an embodiment of the present application, the first obtaining module 610 comprises a calling condition determining module and a first data table determining module. The calling condition determining module is configured to determine the number of calling parties and the calling amount of each data table in the non-data middle platform based on the log of the big data platform. The first data table determining module is configured to determine the data table as the first data table when the number of calling parties and the calling amount of the data table satisfy a preset condition.

[0134] According to an embodiment of the present application, the underlying processing table comprises at least one source table. The first obtaining module 610 further comprises a first data table analyzing module and a source table determining module. The first data table analyzing module is configured to analyze the processing link of the first data table to obtain at least one business system table or source table. The source table is a source layer data table of the data middle platform. The source table determining module is configured to determine the source table corresponding to the business system table based on the mapping relationship between the business system and the source table for the business system table.

[0135] According to an embodiment of the present application, the data middle platform comprises a source layer, an aggregation layer and an extraction layer. The second data table is an aggregation layer data table or an extraction layer data table. The second obtaining module 620 comprises a second data table analyzing module. The second data table analyzing module is configured to analyze the processing link of each second data table to obtain the underlying processing table of the second data table. The underlying processing table comprises at least one source table. The source table is a source layer data table of the data middle platform.

[0136] According to an embodiment of the present application, the sedimentation module 640 comprises an entry lake module and a first sedimentation module. The entry lake module is configured to initiate an entry lake process for the first data table when the underlying processing table of the first data table comprises a data table of the non-data middle platform. The entry lake process indicates storing the underlying processing table of the first data table to a data lake. The data lake is configured to store the data table of the data middle platform. The first sedimentation module is configured to sediment the first data table to the data middle platform when the underlying processing table of the first data table does not comprise a data table of the non-data middle platform.

[0137] According to an embodiment of the present application, the entry lake module comprises a processing module. The processing module is configured to store the underlying processing table of the first data table to the data lake and sediment the first data table to the data middle platform when the underlying processing table of the first data satisfies an entry lake condition.

[0138] According to an embodiment of the present application, the data table processing apparatus 600 further comprises a third acquisition module, a switching module and a second sedimentation module. The third acquisition module is configured to, in the case that the target blood relationship coincidence degree is equal to 1, acquire, by using the large model, fields in the first data table and the target data table that have the same processing logic, and obtain a field correspondence relationship between the first data table and the second data table, the target data table being a second data table corresponding to the target blood relationship coincidence degree; the switching module is configured to, for the fields that have the same processing logic, switch the fields in the first data table to corresponding fields in the target data table based on the correspondence relationship; and the second sedimentation module is configured to, for fields in the first data table that are not in the correspondence relationship, sediment the fields to the target data table.

[0139] According to an embodiment of the present application, any one or more of the first acquisition module 610, the second acquisition module 620, the determination module 630 and the sedimentation module 640 can be combined in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of other modules, and implemented in one module. According to an embodiment of the present application, at least one of the first acquisition module 610, the second acquisition module 620, the determination module 630 and the sedimentation module 640 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of hardware or firmware that can be integrated or packaged with a circuit, or implemented in any one of software, hardware and firmware or in an appropriate combination of any of them. Alternatively, at least one of the first acquisition module 610, the second acquisition module 620, the determination module 630 and the sedimentation module 640 can be at least partially implemented as a computer program module that can perform corresponding functions when the computer program module is run.

[0140] Figure 7 A block diagram of an electronic device suitable for implementing the data table processing method according to an embodiment of the present application is schematically shown.

[0141] As Figure 7As shown, the electronic device 700 according to embodiments of the present application includes a processor 701 which can perform various appropriate actions and processes in accordance with a program stored in a read only memory (ROM) 702 or a program loaded into a random access memory (RAM) 703 from a storage section 708. The processor 701 can include, for example, a general purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chip set, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), and so on. The processor 701 can also include an on-board memory for cache use. The processor 701 can include a single processing unit or multiple processing units for executing different actions of the method processes according to embodiments of the present application.

[0142] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method processes according to embodiments of the present application by executing the programs in the ROM 702 and / or the RAM 703. Note that the programs can also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 can also perform various operations of the method processes according to embodiments of the present application by executing the programs stored in the one or more memories.

[0143] According to embodiments of the present application, the electronic device 700 can also include an input / output (I / O) interface 705 which is also connected to the bus 704. The electronic device 700 can also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as necessary. A removable medium 711 such as a magnetic disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 710 as necessary, so that a computer program read out therefrom is installed in the storage section 708 as necessary.

[0144] The application further provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments, or can exist independently without being assembled into the device / apparatus / system. The computer readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the application.

[0145] According to the embodiments of the application, the computer readable storage medium can be a non-volatile computer readable storage medium, which can include, but is not limited to, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In this application, a computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. For example, in the embodiments of the application, the computer readable storage medium can include one or more of the ROM 702 and / or the RAM 703 described above, and / or one or more memory other than the ROM 702 and the RAM 703.

[0146] The embodiments of the application also include a computer program product, which includes a computer program containing program codes for executing the method shown in the flow chart. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the data table processing method provided by the embodiments of the application.

[0147] The above functions defined in the system / apparatus of the embodiments of the application are performed when the computer program is executed by the processor 701. According to the embodiments of the application, the system, apparatus, module, unit, etc. described above can be implemented by computer program modules.

[0148] In one embodiment, the computer program can rely on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or installed from the detachable medium 711. The program codes contained in the computer program can be transmitted by any appropriate network medium, including but not limited to wireless, wired, etc., or any appropriate combination thereof.

[0149] In such embodiments, the computer program can be downloaded and installed from the network via the communication section 709, and / or installed from the removable media 711. When the computer program is executed by the processor 701, the above-described functions defined in the system of the embodiments of the present application are executed. According to the embodiments of the present application, the system, device, apparatus, module, unit, and the like described above can be realized by the computer program module.

[0150] According to the embodiments of the present application, the program code for executing the computer program provided by the embodiments of the present application can be written in any combination of one or more programming languages, and specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming language, and / or assembly / machine language. The programming language includes, but is not limited to, such as Java, C++, python, "C" language, or similar programming language. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, connected to the Internet through an Internet service provider).

[0151] The flowcharts and block diagrams in the drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a portion of code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the figures. For example, two blocks noted in succession can actually be executed substantially concurrently, or they can sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the flowcharts or block diagrams, and combinations of blocks in the flowcharts or block diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0152] Those skilled in the art can understand that the features described in various embodiments of the present application can be combined and / or integrated in various combinations, even if such combinations are not explicitly described in the present application. In particular, the features described in various embodiments of the present application can be combined and / or integrated in various combinations without departing from the spirit and teachings of the present application. All such combinations and / or integrations are within the scope of the present application.

Claims

1. A data table processing method characterized by, The method comprises: obtaining m first data tables of a non-data middle platform and a bottom processing table of each first data table, the bottom processing table of the first data table being obtained based on a processing link of the first data table, the processing link of the first data table being used to represent a blood relationship of the first data table, m being an integer greater than or equal to 1; obtaining n second data tables of a data middle platform and a bottom processing table of each second data table, the bottom processing table of the second data table being obtained based on a processing link of the second data table, the processing link of the second data table being used to represent a blood relationship of the second data table, n being an integer greater than or equal to 1; for each first data table, determining a target blood relationship coincidence degree based on the bottom processing table of the first data table and the bottom processing table of each second data table; and in the case where the target blood relationship coincidence degree is not equal to 1, precipitating the first data table to the data middle platform.

2. The method of claim 1, wherein, The method comprises: for each first data table, determining a blood relationship coincidence degree between the first data table and each second data table based on a coincidence degree between the bottom processing table of the first data table and the bottom processing table of each second data table; sorting n blood relationship coincidence degrees; based on the sorting result, determining the target blood relationship coincidence degree from the n blood relationship coincidence degrees.

3. The method of claim 1, wherein, The method comprises: based on logs of a big data platform, determining a number of calling parties and a calling amount of each data table in the non-data middle platform; in the case where the number of calling parties and the calling amount of the data table satisfy a preset condition, the data table is a first data table.

4. The method of claim 3, wherein, The bottom processing table comprises at least one source table, and the method comprises: parsing the processing link of the first data table to obtain at least one business system table or source table, the source table being a source layer data table of the data middle platform; for the business system table, determining a source table corresponding to the business system table based on a mapping relationship between a business system and a source table.

5. The method of claim 4, wherein, The data middle platform comprises a source layer, an aggregation layer, and an extraction layer, the second data table being an aggregation layer data table or an extraction layer data table, and the method comprises: for each second data table, parsing the processing link of the second data table to obtain the bottom processing table of the second data table, the bottom processing table comprising at least one source table, the source table being a source layer data table of the data middle platform.

6. The method of claim 2, wherein, The method comprises: In a case where the non-data middle platform data table exists in the bottom processing table of the first data table, an entry lake process is initiated for the first data table, the entry lake process representing storing the bottom processing table of the first data table to a data lake, the data lake being used to store data middle platform data tables; In a case where the non-data middle platform data table does not exist in the bottom processing table of the first data table, the first data table is deposited to a data middle platform.

7. The method of claim 6, wherein, The initiating the entry lake process for the first data table comprises: In a case where the bottom processing table of the first data satisfies an entry lake condition, the bottom processing table of the first data table is stored to a data lake and the first data table is deposited to a data middle platform.

8. The method of claim 2, wherein, The method further comprises: In a case where the target blood relationship coincidence degree is equal to 1, a first data table and a target data table are obtained by using a large model, the target data table being a second data table corresponding to the target blood relationship coincidence degree, and a field correspondence relationship between the first data table and the second data table is obtained; For the same processing logic fields, the fields in the first data table are switched to the corresponding fields in the target data table based on the correspondence relationship; For the fields in the first data table that do not exist in the correspondence relationship, the fields are deposited to the target data table.

9. A data table processing apparatus characterized by comprising: The apparatus comprises: A first obtaining module is configured to obtain m first data tables of a non-data middle platform and a bottom processing table of each first data table, the bottom processing table of the first data table being obtained based on a processing link of the first data table, the processing link of the first data table being used to represent a blood relationship of the first data table, and m being an integer greater than or equal to 1; A second obtaining module is configured to obtain n second data tables of a data middle platform and a bottom processing table of each second data table, the bottom processing table of the second data table being obtained based on a processing link of the second data table, the processing link of the second data table being used to represent a blood relationship of the second data table, and n being an integer greater than or equal to 1; A determining module is configured to, for each first data table, determine a target blood relationship coincidence degree based on the bottom processing table of the first data table and the bottom processing table of each second data table; and A depositing module is configured to, in a case where the target blood relationship coincidence degree is not equal to 1, deposit the first data table to the data middle platform.

10. An electronic device comprising: one or more processors; memory for storing one or more computer programs, characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1-8.

11. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instructions are executed by the processor to implement the steps of the method according to any one of claims 1-8.

12. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions are executed by the processor to implement the steps of the method according to any one of claims 1-8.