Data processing method and related device

By acquiring a data lineage diagram and calculating the allocation coefficient and cost, the problem of inaccurate data asset valuation results in traditional data processing methods is solved, achieving more accurate data asset valuation and supporting data capitalization.

CN121009213APending Publication Date: 2025-11-25CHINA CITIC BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511154988.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Traditional data processing methods in financial data processing systems result in low accuracy of data asset valuation results, failing to effectively reflect the reusability and actual contribution of data assets.

Method used

By obtaining a data lineage diagram, the allocation coefficient and allocation cost of the data tables are calculated layer by layer. The allocation coefficient of each target data table is determined based on the reference relationship, and the allocation cost is accumulated to obtain the data asset assessment result.

Benefits of technology

This improves the accuracy, objectivity, and effectiveness of data asset valuation results, enhances the accuracy of data asset valuation, and provides better support for data capitalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009213A_ABST
    Figure CN121009213A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and a related device, and relates to the technical field of data processing, and the method comprises the steps: obtaining a label set of a target data processing task, obtaining data table reference information of the target data processing task based on the label set, the data table reference information comprises a plurality of target data tables used for obtaining the label set and reference relations among the target data tables. The apportionment coefficient of each target data table is calculated layer by layer, the apportionment coefficient of each target data table is the ratio of the downstream reference times to the reference times, the downstream reference times are the sum of the apportionment coefficients of the downstream data tables, and the apportionment cost is obtained based on the apportionment coefficients of the target data tables and the data cost. And accumulating the apportioned cost of all the target data tables to obtain a data asset evaluation result of the target data processing task. And determining the data asset value of each target data table in the target data processing task based on the apportionment coefficient, thereby improving the accuracy of the data asset evaluation result of the target data processing task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a data processing method and related apparatus. Background Technology

[0002] Financial data processing systems perform data processing tasks based on large amounts of financial data. Traditional data processing methods often assess data assets by calculating the total cost required to reacquire or replicate the financial data, thereby supporting data capitalization. However, the data asset assessment results obtained by traditional data processing methods suffer from low accuracy. Summary of the Invention

[0003] In view of the above problems, this application provides a data processing method and related apparatus to improve the accuracy of data asset assessment results. The specific solution is as follows:

[0004] The first aspect of this application provides a data processing method, including:

[0005] In one possible implementation, obtaining the data table reference information of the target data processing task based on the tag set includes:

[0006] Obtain the data lineage graph, which includes nodes and directed edges. Each node represents a data table, and the directed edges point from the referenced data table to the referencing data table.

[0007] Based on the data lineage diagram, obtain the data table reference information for the target data processing task.

[0008] In one possible implementation, the amortization coefficient for each of the target data tables is calculated layer by layer, including:

[0009] The sum of the amortization coefficients of the downstream data tables of the target data table is obtained as the downstream reference count;

[0010] Obtain the reference count of the target data table;

[0011] The ratio of the downstream citation count to the total citation count is calculated as the amortization coefficient for the target data table.

[0012] In one possible implementation, obtaining the sum of the allocation coefficients of the downstream data tables of the target data table includes:

[0013] Based on the reference relationships indicated by the data table reference information, obtain the downstream data table of the target data table;

[0014] The sum of the amortization coefficients of the downstream data tables of the target data table is obtained as the downstream reference count.

[0015] In one possible implementation, the allocated cost is obtained based on the allocation coefficient and data cost of the target data table, including:

[0016] Obtain the data cost for each of the target data tables;

[0017] The product of the allocation coefficient and the data cost of the target data table is obtained to obtain the allocation cost of the target data table.

[0018] A second aspect of this application provides a data processing apparatus, comprising:

[0019] A tag acquisition unit is used to acquire a tag set for a target data processing task, the tag set including at least one target tag, the target tag being a data item used to perform the target data processing task;

[0020] The reference information acquisition unit is used to acquire data table reference information of the target data processing task based on the tag set. The data table reference information includes multiple target data tables for acquiring the tag set and the reference relationships between the target data tables.

[0021] The apportionment coefficient acquisition unit is used to calculate the apportionment coefficient of each target data table layer by layer. The apportionment coefficient of the target data table is the ratio of the number of downstream references to the number of references. The number of downstream references is the sum of the apportionment coefficients of the downstream data tables. The downstream data tables include other target data tables that reference the target data table. The number of references is the total number of times the table is referenced.

[0022] The cost allocation unit is used to obtain the cost allocation based on the allocation coefficient and data cost of the target data table.

[0023] The evaluation result acquisition unit is used to accumulate the allocated costs of all the target data tables to obtain the data asset evaluation result of the target data processing task.

[0024] In one possible implementation, when the reference information acquisition unit is used to acquire data table reference information of the target data processing task based on the tag set, it is specifically used for:

[0025] Obtain the data lineage graph, which includes nodes and directed edges. Each node represents a data table, and the directed edges point from the referenced data table to the referencing data table.

[0026] Based on the data lineage diagram, obtain the data table reference information for the target data processing task.

[0027] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the data processing method described in the first aspect or any implementation thereof.

[0028] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0029] The memory is used to store computer programs;

[0030] The processor is used to execute the computer program so that the electronic device can implement the data processing method of the first aspect or any implementation thereof.

[0031] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to perform the data processing method described in the first aspect or any implementation thereof.

[0032] This application provides a data processing method and related apparatus based on the above technical solution. The method involves: acquiring a tag set for a target data processing task, where the tag set includes at least one target tag, and the target tag is a data item used to perform the target data processing task. Based on the tag set, data table reference information for the target data processing task is acquired, including multiple target data tables used to acquire the tag set and the reference relationships between the target data tables. The allocation coefficient for each target data table is calculated layer by layer, where the allocation coefficient is the ratio of downstream reference counts to the total number of references, and the downstream reference count is the sum of the allocation coefficients of the downstream data tables, which include other target data tables that reference the target data table. The total number of references is the total number of times the target data table is referenced. Based on the allocation coefficient and data cost of the target data table, the allocation cost is obtained. The allocation costs of all target data tables are summed to obtain the data asset assessment result of the target data processing task. Therefore, this solution determines the allocation coefficient of each target data table based on reference relationships, and further determines the data asset value of each target data table in the target data processing task through the allocation coefficient, thereby improving the accuracy of the data asset assessment result of the target data processing task. Attached Figure Description

[0033] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0034] Figure 1A schematic diagram of a system architecture is provided for this application;

[0035] Figure 2 A flowchart of a data processing method provided in this application;

[0036] Figure 3 A flowchart illustrating the specific implementation of a data processing method provided in this application;

[0037] Figure 4 A schematic diagram illustrating data table reference information provided in this application;

[0038] Figure 5 A schematic diagram of the structure of a data processing device provided in this application;

[0039] Figure 6 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation

[0040] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0041] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0042] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0043] This application can be applied in the field of data processing technology, including but not limited to applications with data processing capabilities or cloud services provided by cloud-side servers, which will be described in detail below:

[0044] See Figure 1 , Figure 1 A schematic diagram of a system architecture is shown. The system may include a terminal 100 and a server 200. The server 200 can provide the methods provided in the embodiments of this application to one or more terminals.

[0045] The terminal 100 may be equipped with a data processing application. The application and webpage can provide an interface. The terminal 100 can receive relevant parameters input by the user on the data processing interface and send the parameters to the server 200. The server 200 can obtain the processing result based on the received parameters and return the processing result to the terminal 100.

[0046] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters on its own, without the need for the server to cooperate. This application embodiment is not limited to this.

[0047] The following description Figure 1 The product form of the mid-terminal 100;

[0048] The terminal 100 in this application embodiment can be a mobile phone, tablet computer, wearable device, vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.

[0049] Terminal 100 may include a radio frequency unit, memory, input unit, display unit, camera (optional), audio circuitry (optional), speaker (optional), microphone (optional), headphone jack (optional), processor, external interface, power supply, and other components. Those skilled in the art will understand that the above-mentioned components are merely examples and do not constitute a limitation on the terminal or multifunctional device; it may include more or fewer components, or a combination of certain components, or different components.

[0050] The input unit can be used to receive input numeric or character information, and to generate key signal inputs related to user settings and function control of the portable multi-functional device. Specifically, the input unit may include a touchscreen (optional) and / or other input devices. Other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0051] Among them, the input device can receive input data, etc.

[0052] The display unit can be used to display information input by the user or information provided to the user, various menus of the terminal, interactive interfaces, file display, and / or playback of any multimedia file. In the embodiments of this application, the display unit can be used to display an interface for displaying data processing results, processing results, etc.

[0053] The memory can be used to store software code related to the data processing method, the processor can execute the steps of the data processing method, and can also schedule other units (such as the input unit and display unit mentioned above) to achieve the corresponding functions.

[0054] This radio frequency unit (optional) can be used to receive and send signals during information transmission or calls.

[0055] In this embodiment of the application, the radio frequency unit can send data to the server 200 and receive the processing results sent by the server 200.

[0056] It should be understood that this radio frequency unit is optional and can be replaced with other communication interfaces, such as a network port.

[0057] Terminal 100 also includes a power source (such as a battery) for supplying power to the various components.

[0058] Terminal 100 also includes an external interface, which can be a standard Micro USB interface or a multi-pin connector, which can be used to connect terminal 100 to other devices for communication or to connect a charger to charge terminal 100.

[0059] Server 200 includes a bus, a processor, a communication interface, and memory. The processor, memory, and communication interface communicate with each other via the bus.

[0060] The memory can be used to store software code related to data processing methods, the processor can execute the steps of the chip's data processing methods, and it can also schedule other units to achieve corresponding functions.

[0061] With the development of big data technology, data is gradually evolving into an economically valuable asset in the operation and management of financial systems. As an important factor of production, the trend of data valorization and capitalization is becoming increasingly prominent. However, due to the complexity of data storage, use, and reuse, the accuracy of data asset assessment results for data processing tasks is relatively low under traditional data processing methods.

[0062] For example, using the replacement cost method to calculate the total cost required to reacquire or replicate data assets, and thus obtaining a data asset assessment result for a data processing task, ignores the reusability and actual contribution of data assets.

[0063] For example, the income approach is used, which is based on the expected returns from data. However, in banking operations, the returns from individual data assets are difficult to quantify.

[0064] It is evident that when using traditional data processing methods to obtain data asset assessment results for data processing tasks, the accuracy of the data asset assessment results is low because the unique characteristics of data assets, such as unlimited replicability, processability, and shareability, which distinguish them from traditional assets, are not taken into account.

[0065] To address the aforementioned issues, this application provides a data processing method. By accurately segmenting and estimating the value of data resources based on data lineage, the granularity of data asset assessment is improved, thereby enhancing the accuracy of data asset assessment results.

[0066] The data processing method of this application embodiment will be described in detail below with reference to the accompanying drawings.

[0067] Reference Figure 2 , Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of this application, such as... Figure 2 As shown in the figure, the data processing method provided in this application embodiment may include steps 201 to 205, which are described in detail below.

[0068] Step 201: Obtain the tag set for the target data processing task.

[0069] In this embodiment, the tag set includes at least one target tag, which is a data item used to perform a target data processing task.

[0070] Step 202: Based on the tag set, obtain the data table reference information for the target data processing task.

[0071] In this embodiment, the data table reference information includes multiple target data tables used to obtain the tag set and the reference relationships between the target data tables.

[0072] Step 203: Calculate the allocation coefficient for each target data table layer by layer.

[0073] In this embodiment, the allocation coefficient of the target data table is the ratio of the number of downstream references to the total number of references, the number of downstream references is the sum of the allocation coefficients of the downstream data tables, the downstream data tables include other target data tables that reference the target data table, and the number of references is the total number of times it is referenced.

[0074] Step 204: Obtain the allocated cost based on the allocation coefficient and data cost of the target data table.

[0075] Step 205: Accumulate the allocated costs of all target data tables to obtain the data asset assessment results of the target data processing task.

[0076] As can be seen from the above technical solution, the data processing method provided in this application provides a method for obtaining a tag set of a target data processing task. The tag set includes at least one target tag, which is a data item used to perform the target data processing task. Based on the tag set, data table reference information of the target data processing task is obtained. The data table reference information includes multiple target data tables used to obtain the tag set and the reference relationships between the target data tables. The allocation coefficient of each target data table is calculated layer by layer. The allocation coefficient of the target data table is the ratio of the number of downstream references to the total number of references. The number of downstream references is the sum of the allocation coefficients of the downstream data tables. The downstream data tables include other target data tables that reference the target data tables. The total number of references is the total number of times the target data tables are referenced. Based on the allocation coefficient of the target data table and the data cost, the allocation cost is obtained. The allocation costs of all target data tables are summed to obtain the data asset evaluation result of the target data processing task. It can be seen that this solution determines the allocation coefficient of each target data table based on the reference relationship, and further determines the data asset value of each target data table in the target data processing task through the allocation coefficient, thereby improving the accuracy of the data asset evaluation result of the target data processing task.

[0077] Reference Figure 3 , Figure 3 A flowchart illustrating a specific implementation of a data processing method provided in this application embodiment is shown below. Figure 3 As shown in the embodiment of this application, a data processing method may include steps 301 to 311, which are described in detail below.

[0078] Step 301: Obtain the task information of the target data processing task.

[0079] In this embodiment, the target data processing task can be any data processing task, and the execution progress of the data processing task is not limited.

[0080] In this embodiment, there are several methods for obtaining task information of data processing tasks. For example, obtaining task information pre-configured based on expert experience, or obtaining task information through analysis of data processing tasks.

[0081] Taking the data processing task as the marketing strategy formulation task in the retail payroll scenario of the banking industry as an example, the task of the marketing strategy formulation task is to predict the product push to customers two days before the next payroll is issued based on the customers' past payroll time. By setting up an A / B control group in this strategy, the increase in AUM (Assets Under Management) of customers reached by this strategy compared to customers who were not reached can be calculated.

[0082] It should be noted that the data items in the task information are the data basis for executing the data processing task, that is, the data items given are used directly to obtain the data processing results.

[0083] Step 302: Based on the task information, obtain the tag set of the target data processing task.

[0084] In this embodiment, the tag set includes at least one target tag, which is a data item used to perform a target data processing task.

[0085] Step 303: Obtain the data lineage diagram.

[0086] In this embodiment, the data lineage graph includes nodes and directed edges. Each node represents a data table, and the directed edges point from the referenced data table to the referencing data table. That is, the first node points to the second node, indicating that the data table identified by the second node references the data table identified by the first node.

[0087] In one alternative embodiment, starting from the data items directly used by each data processing task, the data source is traced layer by layer to analyze the relationship between data processing, transformation and integration, down to the basic data layer, thereby forming a complete data lineage diagram, which shows the objective dependencies between the data items in the data lineage diagram.

[0088] Step 304: Based on the data lineage diagram, obtain the data table reference information for the target data processing task.

[0089] Step 305: Obtain the sum of the apportionment coefficients of the downstream data tables of the target data table, and use it as the downstream reference count.

[0090] In this embodiment, the downstream data tables of the target data table are obtained based on the reference relationships indicated by the data table reference information. The sum of the amortization coefficients of the downstream data tables of the target data table is obtained as the downstream reference count.

[0091] Step 306: Obtain the reference count of the target data table.

[0092] Step 307: Calculate the ratio of downstream citation count to citation count as the allocation coefficient for the target data table.

[0093] See Figure 4 , Figure 4 This is a diagram illustrating data referencing information. Figure 4 As shown, when executing the target data processing task in business scenario Y, the label set includes two labels, A and B, each referencing multiple data tables. Since a single data table supports more than one business scenario but is used by multiple scenarios, the associated costs of the data table need to be distributed across these scenarios.

[0094] According to the appendix Figure 4 As shown, by analyzing the allocation coefficients of each data table used in scenario Y, the cost allocated to this business scenario can be calculated.

[0095] Since data tables typically support multiple business scenarios, a multi-level allocation method based on objective reference count statistics is used to determine the allocation ratio of each business scenario to the data table. That is, the allocation coefficient of the target data table is the ratio of downstream reference counts to the total number of references, the downstream reference count is the sum of the allocation coefficients of the downstream data tables, and the downstream data tables include other target data tables that reference the target data table; the total number of references is the total number of times the table is referenced.

[0096] Specifically, business scenario Y uses data labels A and B, which reference multiple data tables. The allocation ratio is automatically calculated based on the total number of times each data table is referenced and the specific number of times scenario Y is referenced.

[0097] The total number of references to data table A1 is 2, of which scenario Y is referenced 1 time. Therefore, the apportionment ratio is objectively calculated to be 1 / 2.

[0098] If the total number of references to data table A2 is 3, and scenario Y is referenced 1 time, then the apportionment ratio is 1 / 3.

[0099] If the total number of references to data table A3 is 3, and scenario Y is referenced 2 times, then the apportionment ratio is 2 / 3.

[0100] The allocation coefficients of the underlying data tables that extend further downwards are also automatically calculated layer by layer based on the number of references and allocation coefficients of their direct parent tables, avoiding human interference and ensuring that the results are objective and accurate.

[0101] Specifically, if table A1 is referenced a total of 2 times by the target data processing task, and is referenced 1 time by the target data processing task, then the amortization coefficient for scenario Y is 1 / 2.

[0102] Table A2 is referenced a total of 3 times, with an allocation factor of 1 / 3. Table A3 is referenced by both tags, with a total reference count of 3, therefore its allocation factor is 2 / 3.

[0103] Table B1 has a total of 9 references and an amortization factor of 1 / 9.

[0104] Table B2 has a total of 3 references and an allocation factor of 1 / 3.

[0105] Table A1-1 is referenced 2 times and is referenced by A1. Therefore, the allocation factor of table A1-1 is the allocation factor of A1 × 1 / the number of references, which is 1 / 2 × 1 / 2 = 1 / 4.

[0106] Table A1-2 has 3 references and is referenced by both A1 and A2. Therefore, its allocation factor is A1's allocation factor × 1 / number of references + A2's allocation factor × 1 / number of references, which is 1 / 2 × 1 / 3 + 1 / 3 × 1 / 3 = 5 / 18.

[0107] Table A2-1 has 2 references and is referenced by A2. Therefore, its allocation factor is the allocation factor of A2 × 1 / number of references, which is 1 / 3 × 1 / 2 = 1 / 6.

[0108] Table A3-1 has 2 references and is referenced by both A3 and B1. Therefore, its allocation factor is the allocation factor of A3 × 1 / number of references + the allocation factor of B1 × 1 / number of references, which is 2 / 3 × 1 / 2 + 1 / 9 + 1 / 2 = 7 / 18.

[0109] Similarly, the apportionment factor for B1-1 is 1 / 27.

[0110] The apportionment factor for B1-2 is 2 / 9.

[0111] The apportionment factor for B2-1 is 1 / 15.

[0112] Step 308: Obtain the data cost for each target data table.

[0113] Step 309: Obtain the product of the allocation coefficient and the data cost of the target data table to get the allocation cost of the target data table.

[0114] Step 310: Accumulate the allocated costs of all target data tables to obtain the data asset assessment results of the target data processing task.

[0115] As can be seen from the above technical solutions, the data processing method provided in this application embodiment determines the allocation coefficient entirely based on objective statistical data of the number of times the data table is referenced. It is automatically generated and calculated by the system, which significantly improves the objectivity and effectiveness of the valuation, avoids the influence of human subjective factors, and enhances the possibility of passing the review. Furthermore, it enables accurate calculation of the data asset cost of data processing tasks in complex business scenarios, providing support for data capitalization.

[0116] The above describes a data processing method provided by an embodiment of this application. The following describes an apparatus for performing the above data processing method.

[0117] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Figure 5 As shown, the data processing device 500 includes:

[0118] The tag acquisition unit 501 is used to acquire a tag set for a target data processing task, the tag set including at least one target tag, the target tag being a data item used to perform the target data processing task;

[0119] The reference information acquisition unit 502 is used to acquire data table reference information of the target data processing task based on the tag set. The data table reference information includes multiple target data tables for acquiring the tag set and the reference relationships between the target data tables.

[0120] The apportionment coefficient acquisition unit 503 is used to calculate the apportionment coefficient of each target data table layer by layer. The apportionment coefficient of the target data table is the ratio of the number of downstream references to the number of references. The number of downstream references is the sum of the apportionment coefficients of the downstream data tables. The downstream data tables include other target data tables that reference the target data table. The number of references is the total number of times the table is referenced.

[0121] The cost allocation unit 504 is used to obtain the cost allocation based on the allocation coefficient and data cost of the target data table.

[0122] The evaluation result acquisition unit 505 is used to accumulate the allocated costs of all the target data tables to obtain the data asset evaluation result of the target data processing task.

[0123] In one possible implementation, when the reference information acquisition unit is used to acquire data table reference information of the target data processing task based on the tag set, it is specifically used for:

[0124] Obtain the data lineage graph, which includes nodes and directed edges. Each node represents a data table, and the directed edges point from the referenced data table to the referencing data table.

[0125] Based on the data lineage diagram, obtain the data table reference information for the target data processing task.

[0126] In one possible implementation, when the amortization coefficient acquisition unit calculates the amortization coefficient for each of the target data tables layer by layer, it is specifically used for:

[0127] The sum of the amortization coefficients of the downstream data tables of the target data table is obtained as the downstream reference count;

[0128] Obtain the reference count of the target data table;

[0129] The ratio of the downstream citation count to the total citation count is calculated as the amortization coefficient for the target data table.

[0130] In one possible implementation, when the amortization coefficient acquisition unit is used to obtain the sum of the amortization coefficients of the downstream data tables of the target data table, it is specifically used for:

[0131] Based on the reference relationships indicated by the data table reference information, obtain the downstream data table of the target data table;

[0132] The sum of the amortization coefficients of the downstream data tables of the target data table is obtained as the downstream reference count.

[0133] In one possible implementation, the cost allocation acquisition unit, when acquiring the cost allocation based on the allocation coefficient and data cost of the target data table, specifically performs the following:

[0134] Obtain the data cost for each of the target data tables;

[0135] The product of the allocation coefficient and the data cost of the target data table is obtained to obtain the allocation cost of the target data table.

[0136] This application also provides an electronic device in its embodiments. (See reference...) Figure 6 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0137] like Figure 6 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0138] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0139] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the data processing methods provided in this application.

[0140] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the data processing methods provided in this application.

[0141] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0143] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0144] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A data processing method, characterized in that, include: Obtain a tag set for a target data processing task, the tag set including at least one target tag, the target tag being a data item used to perform the target data processing task; Based on the tag set, data table reference information for the target data processing task is obtained. The data table reference information includes multiple target data tables used to obtain the tag set and the reference relationships between the target data tables. The allocation coefficient of each target data table is calculated layer by layer. The allocation coefficient of the target data table is the ratio of the number of downstream references to the number of references. The number of downstream references is the sum of the allocation coefficients of the downstream data tables. The downstream data tables include other target data tables that reference the target data table. The number of references is the total number of times the table is referenced. Based on the allocation coefficient and data cost of the target data table, obtain the allocation cost; The data asset assessment result of the target data processing task is obtained by summing up the allocated costs of all the target data tables.

2. The data processing method according to claim 1, characterized in that, The step of obtaining data table reference information for the target data processing task based on the tag set includes: Obtain the data lineage graph, which includes nodes and directed edges. Each node represents a data table, and the directed edges point from the referenced data table to the referencing data table. Based on the data lineage diagram, obtain the data table reference information for the target data processing task.

3. The data processing method according to claim 1, characterized in that, The step-by-step calculation of the allocation coefficient for each of the target data tables includes: The sum of the amortization coefficients of the downstream data tables of the target data table is obtained as the downstream reference count; Obtain the reference count of the target data table; The ratio of the downstream citation count to the total citation count is calculated as the amortization coefficient for the target data table.

4. The data processing method according to claim 3, characterized in that, The step of obtaining the sum of the allocation coefficients of the downstream data tables of the target data table includes: Based on the reference relationships indicated by the data table reference information, obtain the downstream data table of the target data table; The sum of the amortization coefficients of the downstream data tables of the target data table is obtained as the downstream reference count.

5. The data processing method according to claim 4, characterized in that, The process of obtaining the allocated cost based on the allocation coefficient and data cost of the target data table includes: Obtain the data cost for each of the target data tables; The product of the allocation coefficient and the data cost of the target data table is obtained to obtain the allocation cost of the target data table.

6. A data processing apparatus, characterized in that, include: A tag acquisition unit is used to acquire a tag set for a target data processing task, the tag set including at least one target tag, the target tag being a data item used to perform the target data processing task; The reference information acquisition unit is used to acquire data table reference information of the target data processing task based on the tag set. The data table reference information includes multiple target data tables for acquiring the tag set and the reference relationships between the target data tables. The apportionment coefficient acquisition unit is used to calculate the apportionment coefficient of each target data table layer by layer. The apportionment coefficient of the target data table is the ratio of the number of downstream references to the number of references. The number of downstream references is the sum of the apportionment coefficients of the downstream data tables. The downstream data tables include other target data tables that reference the target data table. The number of references is the total number of times the table is referenced. The cost allocation unit is used to obtain the cost allocation based on the allocation coefficient and data cost of the target data table. The evaluation result acquisition unit is used to accumulate the allocated costs of all the target data tables to obtain the data asset evaluation result of the target data processing task.

7. The data processing apparatus according to claim 6, characterized in that, The reference information acquisition unit, when acquiring data table reference information for the target data processing task based on the tag set, is specifically used for: Obtain the data lineage graph, which includes nodes and directed edges. Each node represents a data table, and the directed edges point from the referenced data table to the referencing data table. Based on the data lineage diagram, obtain the data table reference information for the target data processing task.

8. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to perform the data processing method as described in any one of claims 1 to 5.

9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the data processing method as described in any one of claims 1 to 5.

10. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the data processing method as described in any one of claims 1 to 5.