Data management method and device, equipment and storage medium

By obtaining data characteristics for estimating value points and updating the value points of related assets based on data blood relationships, the problem of low efficiency in identification of data assets in the existing technology is solved, and a more accurate and automated evaluation of data asset value is achieved.

CN120197822APending Publication Date: 2025-06-24TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510273290.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The importance of prior art inaccurately automated identification of data assets has resulted in inefficient data processing and maintenance.

Method used

By obtaining the characteristics of the data, estimating the value scores, and updating the value scores of the parent or sub-category assets according to the data blood relationship, the automatic transmission of the importance of data assets is achieved.

Benefits of technology

It improves the accuracy of the evaluation of the value of data assets, reduces the need for manual annotation, and reduces the duration of partitioning important data assets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197822A_ABST
    Figure CN120197822A_ABST
Patent Text Reader

Abstract

The invention provides a data management method and device, equipment and a storage medium, and is applied to the technical field of big data. The method comprises the steps of obtaining at least one data feature of first data; performing value score estimation according to the at least one data feature to obtain at least one first value score corresponding to the first data; under the condition that the first value score is not lower than a preset value, parent data or child data corresponding to the first data are obtained according to the data blood relationship corresponding to the first data, and second data are obtained; and updating the value score corresponding to the second data according to the first value score to obtain at least one second value score of the second data. According to the method, the importance of the data assets can be automatically identified, and the accuracy of value score evaluation of the data assets can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of big data technology, and in particular, to a data management method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of the big data era, data has become an indispensable core asset for enterprises. Effectively collecting, storing, analyzing, and presenting data has become the key for enterprises to enhance their competitiveness. A data platform can provide comprehensive data management and analysis capabilities for enterprises or individuals, helping to efficiently collect, store, process, analyze, and utilize data assets. The evaluation of the application value of data assets in a data platform can assess the importance of data assets, which is of great significance in complex data application scenarios and indicates the direction for data maintenance and quality assurance.

[0003] In related solutions, the manifestation form of the importance level of data assets is obtained from the business through manual research. Subsequently, professionals divide a large range of important data assets and provide them to the business side for manual annotation. After obtaining the results, data processing and maintenance are carried out. This implementation process mainly consumes manpower and cannot accurately display the importance degree of data assets. Summary of the Invention

[0004] The present application provides a data management method, apparatus, device, and storage medium, which can automatically identify the importance of data assets and help improve the accuracy of data asset value evaluation.

[0005] In a first aspect, an embodiment of the present application provides a data management method, including:

[0006] Obtaining at least one data feature of first data;

[0007] Estimating a value score based on the at least one data feature to obtain at least one first value score corresponding to the first data;

[0008] When the first value score is not lower than a preset value, obtaining the parent data or child data corresponding to the first data according to the data lineage corresponding to the first data to obtain second data;

[0009] Updating the value score corresponding to the second data according to the first value score to obtain at least one second value score corresponding to the second data.

[0010] In a second aspect, an embodiment of the present application provides a data management apparatus, including:

[0011] A first acquisition unit, configured to obtain at least one data feature of first data;

[0012] A scoring unit, configured to estimate a value score according to the at least one data feature, so as to obtain at least one first value score corresponding to the first data;

[0013] A second obtaining unit, configured to, when the first value score is not lower than a preset value, obtain the parent data or child data corresponding to the first data according to the data lineage corresponding to the first data, so as to obtain second data;

[0014] An updating unit, configured to update the value score corresponding to the second data according to the first value score, so as to obtain at least one second value score of the second data.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementation manners.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, used to store a computer program, and the computer program enables a computer to execute the method in the first aspect or its various implementation manners.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer program instructions, and the computer program instructions enable a computer to execute the method in the first aspect or its various implementation manners.

[0018] In a sixth aspect, an embodiment of the present application provides a computer program, and the computer program enables a computer to execute the method in the first aspect or its various implementation manners.

[0019] Through the technical solution provided by the present application, when the first data is a high-value data asset, it is possible to update the value score of the parent asset or child asset corresponding to the first data, that is, the second data, according to the first value score of the first data, so as to make up for the value score of the parent asset or child asset of the first data, and realize the transfer of the importance of the data asset to the upper-layer parent asset or the lower-layer child asset, that is, realize the cross-layer transfer of the value score of the data asset, making the evaluation of the data importance more perfect, fair and reasonable, thereby improving the accuracy of the evaluation of the data asset value score. In addition, the embodiment of the present application can automatically identify the importance of the data asset without manually marking the scope of important assets, thereby reducing the time for dividing important data assets. Description of the Drawings

[0020] Figure 1A It is a schematic diagram of an implementation environment of the solution provided by the embodiment of the present application;

[0021] Figure 1BA schematic diagram of a product performance related to an embodiment of this application;

[0022] Figure 2 A schematic diagram of a system architecture related to an embodiment of this application;

[0023] Figure 3 A schematic flowchart of a data management method provided by an embodiment of this application;

[0024] Figure 4 A schematic flowchart of another data management method provided by an embodiment of this application;

[0025] Figure 5 A schematic flowchart of another data management method provided by an embodiment of this application;

[0026] Figure 6 A schematic diagram of a data management device provided by an embodiment of this application;

[0027] Figure 7 A schematic block diagram of an electronic device provided by an embodiment of this application. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0029] It should be understood that in the embodiments of this application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0030] In the description of this application, unless otherwise specified, "at least one" means one or more, and "a plurality" means two or more than two. In addition, "and / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the preceding and following associated objects. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0031] It should also be understood that the first, second, etc. descriptions in the embodiments of the present application are only for schematic and distinguishing description objects, without order, and do not represent special limitations on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.

[0032] It should also be understood that specific features, structures, or characteristics related to the embodiments in the specification are included in at least one embodiment of the present application. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.

[0033] In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0034] The present application may relate to the field of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.

[0035] First, the relevant knowledge related to the present application will be described:

[0036] Data platform: A data management and R & D platform used by data developers. The data platform can provide functions for accessing and utilizing data, helping organizations efficiently manage and analyze massive data to promote data-driven decision-making and business optimization. Its core functions include but are not limited to data collection, data storage, data processing, data analysis, and visualization, etc. The data platform plays an important role in multiple fields such as the e-commerce field, the financial field, the healthcare field, and the transportation field.

[0037] Application platform: A data visualization display platform used by data analysts or operators, such as drawing data line charts, bar charts, etc. Generally, the data platform is more inclined to be used by data developers, and the threshold for data analysts or operators is relatively high. Therefore, data analysts transfer data from the data platform to the application platform to lower the data usage threshold. Exemplarily, the application platform can copy data from other data platforms as its data source, or directly interact with other data platforms to obtain data.

[0038] Optionally, the data platform can integrate the functions of the application platform, that is, it can not only perform data development but also perform data display.

[0039] Data sharing: It refers to the process in which data is transferred from one data platform (or application platform) to another data platform (or application platform) for use, that is, the data interaction process between platforms. The more frequent the data sharing, the higher the usage frequency of the data.

[0040] Basic metadata: That is, the descriptive information of data, which is used to describe the basic attributes and characteristics of data, and can provide key information such as the structure, definition, source, storage location, access rights, data lineage, data quality, and integrity of data. It helps the data platform understand, manage, and use data, and is an important part of the data platform for data development and management.

[0041] Data lineage: Also known as data pedigree, data origin, data genealogy, etc., it refers to a natural association relationship similar to human blood relationship formed among data during the entire life cycle of data, from its generation, processing, processing, fusion, transfer to final extinction. That is to say, data lineage is used to describe the source and destination relationship between data, that is, where the data comes from and where it goes. Among them, the same data can have multiple sources, and can also be generated by processing multiple data. Data lineage, as important information describing the association and transfer between data, belongs to a part of the basic metadata.

[0042] Optionally, data lineage can be obtained by parsing and extracting key content based on the programming code in the data platform.

[0043] Parent data: In the data lineage, the dataset or data source that generates or provides the original data is called the parent data. The new data obtained by processing, transforming, or generating the parent data is called the child data. Through the data lineage, the source and generation process of the parent data can be traced, which is of great significance for understanding the source of data, ensuring data quality and security.

[0044] Optionally, the same child data can have multiple parent data, that is, multiple data sources or datasets can jointly generate a child data.

[0045] Figure 1A It is a schematic diagram of the solution implementation environment provided by an exemplary embodiment of the present application. The solution implementation environment may include multiple nodes 101, and each node 101 may deploy a data platform for data development and management. Among them, the data platform can provide full-link data development capabilities including data integration, data development, and task operation and maintenance, as well as data governance and operation capabilities such as data maps, data quality, and data security, to help enterprises achieve cost reduction and efficiency improvement and maximize data value during the data construction and application process. Optionally, an application platform can also be integrated in the data platform.

[0046] In some embodiments, data sharing can be carried out between the data platforms of different nodes 101. For example, data can be transferred from one data platform to another for use. The usage methods after data sharing include, but are not limited to, obtaining analysis metrics, item portraits, target object portraits, etc. based on the shared data, providing important data support for business operation and management decisions.

[0047] Among them, each data platform can correspond to a data source, and the data of different data sources can be feature summarized, that is, the data from different sources are unifiedly associated and merged to construct a new data feature table.

[0048] Exemplarily, the node 101 can be a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can also become a node of the blockchain.

[0049] The server can be one or more. When there are multiple servers, at least two servers are used to provide different services, and / or at least two servers are used to provide the same service. For example, the same service is provided in a load balancing manner. The embodiments of the present application do not limit this.

[0050] It should be understood that Figure 1A This is only an exemplary illustration and does not specifically limit the application scenarios of the embodiments of the present application. For example, Figure 1A exemplarily shows four nodes 101 Figure 1A It can also include other numbers of nodes. The nodes can also include application platforms, or the application scenarios can also include terminal devices, etc. The present application does not limit this.

[0051] In the embodiments of the present application, in each node, the data platform can evaluate the application value score of the data assets to indicate the importance of the data assets. For data assets with higher importance, their impact on the overall data is greater, and the daily maintenance requirement is higher.

[0052] In related solutions, the manifestation form of the important level of data assets is obtained from the business through manual research. Subsequently, professionals divide a large range of important data assets and provide them to the business side for manual annotation. After obtaining the results, data processing and maintenance are carried out. This implementation process mainly consumes manpower and cannot accurately display the importance degree of data assets.

[0053] To solve the above technical problems, the embodiments of the present application provide a data management method, apparatus, device, and storage medium, which can improve the accuracy of data asset value assessment.

[0054] Specifically, the embodiments of the present application can obtain at least one data feature of the first data, estimate the value score according to the at least one data feature, obtain at least one first value score corresponding to the first data. If the first value score is not lower than a preset value, the parent data or child data corresponding to the first data is obtained through the data lineage corresponding to the first data to obtain the second data, and the value score corresponding to the second data is updated according to the first value score to obtain at least one second value score corresponding to the second data.

[0055] In the embodiments of the present application, when the first value score of the first data is not lower than the preset value, the parent data or child data corresponding to the first data is obtained through the data lineage of the first data to obtain the second data, and the value score of the second data is updated according to the first value score of the first data. It can be realized that when the first data is a high-value data asset, the value score of the parent asset or child asset corresponding to the first data, that is, the second data, is updated according to the first value score of the first data to make up for the value score of the parent asset or child asset of the first data, and the importance of the data asset is transmitted to the upper-level parent asset or the lower-level child asset, that is, the cross-layer transmission of the data asset value score is realized, so that the data importance assessment is more complete, fair and reasonable, thereby improving the accuracy of data asset value assessment.

[0056] In addition, the embodiments of the present application can automatically identify the importance of data assets without manual annotation of the scope of important assets, thereby reducing the time for dividing important data assets.

[0057] The embodiments of the present application can be applied in complex data application scenarios, can intuitively and accurately evaluate the importance of each data asset, and save a large amount of repetitive work of manual screening, investigation and confirmation. The embodiments of the present application have been verified in relevant data platforms, and can cover more than 70% of the known important data assets. At the same time, the embodiments of the present application can also extract more important data assets that have not been discovered, and point out the direction for data maintenance and data quality assurance, such as maintaining and optimizing important data assets and allocating higher priorities.

[0058] When developers need to understand the importance of data assets in an organization, architecture or product, they can obtain the application value score of the data assets through the embodiments of the present application. Among them, the higher the score of the application value score, the higher the importance of the data asset, and thus the greater the impact of the data asset on the overall data, the higher the daily maintenance requirements, and the higher the requirements for data quality assurance.

[0059] Exemplarily,Figure 1B Shows a schematic diagram of a product performance in the application of application value score in data quality management applications. As Figure 1B shown, in the data governance page of data assets, the governance analysis obtains the application value score, storage health score, and computing health score as 77, 80, and 98 respectively. Then, based on the application value score, storage health score, and computing health score, the big data governance index is comprehensively obtained as 84.27 points. Optionally, the resource usage (such as storage resources) can also be displayed on the page as 20.32 for the storage of the data warehouse (DW), with a 6% month-on-month increase, 0.09 for the storage of the data bank (DBANK), with a 1% month-on-month increase, and the number of offline tables in the entity data is 689, with an 8% month-on-month increase.

[0060] In this Figure 1B shown application scenario, the application value score can realize the digital display of data assets. Specifically, for non-professional users, the ability to understand the value of data assets is relatively weak, and it is not convenient to communicate professional knowledge about data assets. By displaying the importance of data assets through the application value score, the internal characteristics of data assets can be abstracted into a more intuitive representation, such as user recognition, data usage frequency, data coverage area, etc. The level of the data application value score can directly reflect the quality of data value.

[0061] Therefore, by obtaining the application value score of data assets, the importance of data assets can be more intuitively displayed, indicating the direction for data maintenance, data quality assurance, data development, etc. in daily life.

[0062] In addition, the application value score of data assets also has extensive applications in fields such as enterprise cost management and model training. For example, in the enterprise cost management scenario, for hundreds of thousands of data assets of an enterprise, the storage and computing costs consumed for asset updates every day are extremely high, and it is difficult for managers to judge the value of data over time. By evaluating the value score of data assets through the solution of the embodiment of the present application, the governance of data assets with low value scores (i.e., useless assets) can be realized, such as specifying low-cost data processing strategies such as deletion, freezing, compression, or merging, effectively reducing data management costs. Another example is in the model training scenario. The model needs to use a large amount of high-quality sample data during training, such as object portraits, content recommendation data, etc. The value of the sample data has a relatively large impact on the final effect of the model. By scoring the value of the sample data through the application value score system provided by the present application, high-quality data samples can be quickly extracted according to the data scores and added to the model training, saving a large amount of data screening work.

[0063] Figure 2A schematic diagram of a system architecture involved in an embodiment of the present application. The system architecture sequentially includes modules such as basic metadata 210, application value indicators 220, core computing 230, application value score calculation 240, and application platform 250 from bottom to top.

[0064] Among them, the basic metadata 210 can obtain the basic metadata of data assets in a data platform (such as data platform a, data platform b, etc.) and obtain the data lineage relationships therein. Optionally, the data lineage relationships can include table lineage relationships and task lineage relationships. Among them, the table lineage relationship refers to the flow and transmission process of data between different data tables, and the task lineage relationship refers to the flow and conversion relationship of data during specific tasks of data processing.

[0065] The application value indicator 220 module obtains data feature in each dimension by processing and parsing the basic metadata. Optionally, the application value indicator 220 module can obtain data features in three major dimensions, namely user recognition, popularity application, and shared value feature dimensions. Among them, user recognition refers to features such as users' likes and collections of data, popularity application refers to features such as how many times the data is used and how many people use it within a certain period, and shared value refers to features such as which organizations and people use the data, whether the influence breaks through its own product and is recognized and used by other data products, and the monetary value of the data.

[0066] Optionally, in the embodiment of the present application, since each data feature is used to calculate the application value score of a data asset, the data feature can also be referred to as an application value indicator or a feature indicator. That is, in the embodiment of the present application, data feature, application value indicator, feature indicator, etc. represent the same or similar meanings and can be replaced with each other without special instructions.

[0067] The core computing 230 module is used to estimate the value score of each feature indicator. Optionally, the core computing 230 module can estimate the value score of each feature indicator through modules such as a neural network model, indicator weighting, and automatic grading and scoring. Here, by adding a neural network model as a supplementary value score estimation algorithm, it can help the solution be more intelligently applied to more complex scenarios.

[0068] The application value score 240 module can perform cross-layer transmission according to the value scores of each feature indicator, transmit the importance of the data asset to the upper-layer parent asset, and perform weighted summation of the value scores of each feature indicator of each data asset to obtain the final value score of each data asset. Optionally, the application value score 240 module also performs hierarchical weighting according to the attribution information of the data asset to obtain the value score of an organization or an individual (i.e., the creator), which is used to represent the importance of the organization or the creator in the data field.

[0069] Optionally, the application platform 250 may display the application value score of the data asset. Optionally, the application platform 250 may also display the application value score of an organization or an individual.

[0070] It should be noted that in the embodiments of the present application, the application value score, value score, value obtained score, etc. represent the same or similar meanings, and can be replaced with each other without special instructions.

[0071] Next, the technical solutions of the embodiments of the present application will be described in detail through some embodiments. These several embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0072] Figure 3 FIG. is a schematic flowchart of a data management method 300 provided by an embodiment of the present application. The method 300 may be executed by any electronic device having data processing capabilities, such as a server. As an example, the method 300 may be executed by Figure 1A each node 101 in, or the data platform deployed in each node 101, which is not limited. As Figure 3 shown, the method may include steps S310 to S340.

[0073] S310, obtain at least one data feature of the first data.

[0074] Exemplarily, the first data, i.e., the data asset, may be data content in various business fields, including but not limited to the video field, news field, e-commerce field, financial field, etc.

[0075] Optionally, the first data may include but not limited to structured data, semi-structured data, unstructured data, etc. Among them, structured data is data with a clear data model and structure, usually stored in a database, data warehouse or data lake in the form of a table. Semi-structured data is a data type between structured and unstructured data, with a certain structure but not completely regular, and is widely used in Internet and big data applications, such as Extensible Markup Language (XML), JavaScript Object Notation (JSON), etc. Unstructured data is data that cannot be represented by traditional row and column methods (such as a table), and usually exists in the form of text, pictures, audio or video.

[0076] Optionally, in the embodiments of the present application, at least one data feature of the first data may be obtained according to the basic metadata of the first data.

[0077] Specifically, the basic metadata of the first data is the descriptive information used to describe the first data. Exemplarily, the basic metadata of the first data may include the daily active user count (DAU), page view count (PV), unique visitor count (UV), like count, favorite count, usage platform, data performance after sharing, usage method after sharing, and information such as who uses the data after sharing, etc. Optionally, the basic metadata of the first data may further include the data lineage of the first data, such as the parent data or child data of the first data.

[0078] Exemplarily, the usage platform of the first data may be, for example, a portrait platform, a data warehouse platform, etc., without limitation. Among them, the portrait platform is built based on object tag data and serves the business, facilitating business operators to create portraits or classify single objects or groups of objects for personalized recommendation and refined operation in the business. The data warehouse platform aggregates structured data from different sources for comparison and analysis in the field of business intelligence, enabling cross-business and cross-system data integration and providing unified data support for management analysis and business decision-making.

[0079] One implementable way, refer to Figure 2 , the basic metadata 210 module can obtain the basic metadata of the input first data, and the application value index 220 module processes and analyzes the basic metadata of the first data to obtain at least one data feature of the first data.

[0080] Optionally, the at least one data feature includes at least one of a data non-sharing type feature and a data sharing type feature; among them, the data non-sharing type feature is the feature of the data in the data product it generates, and the data sharing type feature is the feature of the data after it is shared to another data product.

[0081] Among them, the data product may include a data platform. Specifically, the process of data generation and use may occur in different data products, or the data may be transferred among different data products during use. The process of data being transferred from one data product to another for use is data sharing.

[0082] The data non-sharing type feature is the feature of the data in the data product it generates, or the feature of the data before it is shared to other data products. Exemplarily, the data non-sharing type feature includes at least one of the page view count PV, unique visitor count UV, daily active user count DAU, favorite count, and like count. Among them, the favorite count, like count, page view count PV, etc. can be Figure 2 examples of user recognition dimension features inFigure 2 Example of medium heat application dimension features.

[0083] The data sharing - related features include at least one of information on data sharing situation, data sharing usage type, data sharing service role information, historical transaction proportion value information, and substitutability value information.

[0084] Among them, the information on data sharing situation can be the actual data performance after data sharing, such as the number of people using it per day, the number of access times per day, etc. The data sharing usage type can be the usage method after data sharing, such as obtaining analysis indicators, item portraits, object portraits, etc. based on the shared data. The data sharing usage type can provide important data support for business operation and management decision - making. The shared service role information can be the object groups that the data sharing will be used by, such as management personnel, financial personnel, advertising placement personnel, etc.

[0085] Among them, the historical transaction proportion value information can include the proportion of the data asset in the total historical income of the organization to which it belongs. The total historical income can include the total monthly income, the total annual income, etc., without limitation. The substitutability value information can include the total number of data assets similar to the data asset attributes and data asset meanings. In some embodiments, the historical transaction proportion value information and the substitutability value information can be referred to as monetary value information.

[0086] Specifically, the application products of data assets can obtain economic benefits through providing information services, account sales, product sales, etc. Data assets also have monetary attributes. By calculating the value score of data assets according to their monetary attributes, the value score can reflect the single value of data assets, the contribution degree to the enterprise, the irreplaceability in the field, etc. Thus, the value score is no longer a single digital display, but a more persuasive measurement system.

[0087] Therefore, in the embodiments of the present application, by obtaining the data sharing - related features of data, including data sharing situation, data sharing usage type, data sharing service role information, monetary value information, etc., the value of data assets is reflected not only in the data sharing usage process but also in the data sharing results.

[0088] In some embodiments, when classifying the characteristic attributes of the first data into various categories, such as user recognition, access heat, access object group, data sharing value, etc., if new data features are mined, the data features can be added to each category according to the data feature type and fused with the existing data features.

[0089] In some possible implementation manners, the basic metadata of different data sources of the first data can be uniformly associated and merged to construct a data feature table, which includes at least one of the above-mentioned data features. Specifically, different data features may come from different data platforms. When using data features, it is necessary to integrate data features from different sources, that is, perform feature integration, so as to perform pre-processing for estimating the value scores of different data features of the first data subsequently.

[0090] S320, estimate value scores according to at least one data feature to obtain at least one first value score corresponding to the first data.

[0091] Exemplarily, for each data feature in at least one data feature of the first data, a first value score corresponding to each data feature can be obtained, or value scores can be estimated for a certain type of data features in at least one data feature to obtain the first value score of this type of data features.

[0092] For example, for each data feature in the data non-sharing type features of the first data, such as each feature index like PV, UV, DAU, number of collections, number of likes, etc., a first value score can be calculated respectively, that is, each data feature corresponds to a separate first value score.

[0093] For another example, for the data sharing type features of the first data, such as data sharing situation information, data sharing usage type, and data sharing service role information, etc., a first value score is calculated jointly, that is, the data sharing type features of the first data correspond to a first value score, that is, the data sharing type features provide an overall value score.

[0094] In a possible implementation manner, the first value score corresponding to each data feature can be determined according to the position of the value corresponding to each data feature in the value distribution corresponding to each data feature.

[0095] Exemplarily, when calculating the first value score for each data feature of the first data respectively, the first value score corresponding to each data feature can be determined according to the position of the value (i.e., the feature value) corresponding to each type of feature data in the value distribution corresponding to each type of feature data. In this way, hierarchical processing of data features can be realized, that is, the higher the data feature value, the higher the first value score corresponding to the data feature, so as to ensure that each level of feature value of the data feature has different scores.

[0096] Optionally, the value distribution corresponding to each type of data feature can be statistically obtained, and the value distribution varies according to the different categories or businesses to which the data belongs.

[0097] For example, when the value corresponding to a data feature is at the median or mean in the value distribution corresponding to the data feature, it is determined that the first value score corresponding to the data feature is 50 points; when the position of the value corresponding to the data feature in the value distribution of the data feature is relatively forward or backward with respect to the median or mean, the first value score corresponding to the data feature is appropriately increased or decreased according to the forward or backward distance.

[0098] For another example, it can be configured that when the value corresponding to a data feature exceeds a specific value, the corresponding first value score is a specific score value. For example, when pv exceeds 10, the corresponding first value score is 10 points, and when pv exceeds 60, the corresponding first value score is 20 points, etc.

[0099] Optionally, according to the position of the value corresponding to each non - shared data feature, such as PV, UV, DAU, number of collections, number of likes, etc., in the value distribution corresponding to each data feature, the first value score corresponding to each data feature can be determined. That is to say, for non - shared data features, hierarchical processing can be performed, such that the higher the feature value corresponding to this type of feature, the higher the corresponding value score.

[0100] Exemplarily, referring to Figure 2 , the automatic grading and scoring unit in the core calculation 230 module can automatically grade and score each data feature according to the value corresponding to each data feature, and obtain the first value score corresponding to each data feature.

[0101] In a possible implementation manner, a data feature can be input into a vector converter to map the data feature into a first vector representation, and the first vector representation is input into a neural network model to output the first value score corresponding to the data feature. Among them, the vector converter and the neural network model are trained according to data with value score labels.

[0102] Optionally, the above - mentioned neural network model includes a Multilayer Perceptron Classifier.

[0103] Exemplarily, when calculating the first value score according to the data features of the first data, a neural network model can be used to learn the importance degree of each data feature and output the corresponding first value score. Specifically, each data feature can be input into a vector converter to convert the data feature into a vector format suitable for processing by the neural network model, for example, converted into specific coordinate values in a coordinate system. After the coordinate values are input into the neural network model, the model calculates the importance degree of the data features for the coordinate values, and finally maps the features to the corresponding categories. Among them, each category corresponds to a type of score, and a type of score corresponds to a type of importance degree category).

[0104] As a specific example, a vector can be selected as the vector converter. The number of input layers configured in the model is 6, corresponding to 6 feature inputs; the number of hidden layers is 6 or 7, which is used for non-linear transformation of data; the number of output layers is 6, corresponding to the final 6 types of scores (such as 0, 10, 30, 60, 80, 100 respectively).

[0105] Exemplarily, data with known importance levels can be labeled to obtain data with value score labels, and then the vector converter and the neural network model can be trained based on the data with value score labels to obtain the trained vector converter and neural network model. Exemplarily, the data with known importance levels includes, but is not limited to, important data, general data, low-value data, zero-value data, etc. For data with each importance level, corresponding value score labels can be respectively labeled. For example, zero-value data corresponds to a value score of 0, and important data corresponds to a value score of 100, and so on. Optionally, the importance level categories of the data can correspond to the number of output layers of the neural network model.

[0106] During the model training process, for data with each importance level, at least one data feature corresponding to the data can be obtained, and the at least one data feature is input into the vector converter to map the data feature into a vector representation. Then, the vector representation is input into the neural network model, and the value score output by the model is constrained to approach the value score label corresponding to the data, so as to update the parameters of the vector converter and the neural network model to obtain the trained vector converter and neural network model.

[0107] In some implementable ways, the maximum number of iterations can be configured to be 100 times when training the model, and the step size of the gradient descent algorithm can be selected as the L-BFGS (Limited-memory Broyden–Fletcher–Goldfarb–Shanno) optimization algorithm, and the embodiments of the present application do not limit this.

[0108] It should be understood that the importance levels of the same data may be different in different business fields. Therefore, the vector converter and the neural network model corresponding to the business can be trained based on the data with known importance levels in a certain business field, so as to realize the value score estimation of the data in this business field.

[0109] In some embodiments, at least one data feature can be input into the vector converter to map each data feature into a first vector representation, and then the at least one vector representation is jointly input into the neural network model, and the model outputs an overall first value score corresponding to the at least one data feature. Exemplarily, the at least one data feature can be a certain type of data feature, such as a data sharing type feature.

[0110] It can be understood that in the embodiments of the present application, by introducing a neural network model algorithm to estimate the value score of data features, it is possible to not only estimate the value score as a whole with at least one data feature to obtain an overall value score value, but also make the solution of the present application more intelligently applied to more complex application scenarios.

[0111] Exemplarily, referring to Figure 2 , the neural network model unit in the core calculation 230 module can input each data feature and output a comprehensive first value score by learning the importance degree of each feature.

[0112] S330, when the first value score is not lower than the preset value, obtain the parent data or child data corresponding to the first data according to the data lineage of the first data to obtain the second data.

[0113] Specifically, if the first value score is not lower than the preset value, it means that the first data is a data asset with a high value score. For a data asset with a high value score, in the embodiments of the present application, the corresponding parent data (i.e., parent class asset) or child data (subclass asset) is obtained according to its corresponding data lineage to obtain the second data.

[0114] As an example, the preset value can be configured to 70 points, that is, a data asset with a score higher than or equal to 70 points is a high-value score data asset.

[0115] As an example, for the first table in the database, a new second table can be generated by processing the first table through a programming development language. Based on the parsing of the programming development language, the source and destination of the table can be obtained, that is, the data lineage. For example, the second table is derived from the first table. Therefore, the first table is the source table of the second table, and the second table is the destination table of the first table. Therefore, in the data lineage, the first table is the parent data or parent class asset of the second table, and the second table is the child data or subclass asset of the first table.

[0116] It should be noted that the second data can be the parent data or child data corresponding to the first data, or can include both the parent data corresponding to the first data and the child data of the first data. The embodiments of the present application do not make any limitations in this regard.

[0117] S340, update the value score corresponding to the second data according to the first value score to obtain at least one second value score of the second data.

[0118] Specifically, when the first data is a high-value data asset, the value score of the parent asset or child asset corresponding to the first data, that is, the second data, can be updated according to the relatively high first value score of the first data, so as to make up for the value score of the parent asset or child asset of the first data, realize the transfer of the importance of the data asset to the upper-level parent asset or the lower-level child asset, realize the cross-layer transfer of the value score of the data asset, make the evaluation of data importance more perfect, fair and reasonable, and thus improve the accuracy of the evaluation of the value score of the data asset.

[0119] Exemplarily, referring to Figure 2 , the application value score calculation module 240 can update the corresponding value score of the second data according to the first value score to obtain the second value score of the second data, that is, realize the cross-layer transfer of the value score of the data asset.

[0120] Optionally, when the value score evaluation is performed according to a certain data feature #1 of the first data and the obtained first value score is not lower than the preset value, the value score of the second data obtained based on the feature of the same type as the data feature #1 can be updated according to the first value score to obtain the corresponding second value score, so as to make the evaluation of the importance of the data feature #1 more perfect, fair and reasonable, and improve the accuracy of the evaluation of the value score of the data feature #1.

[0121] Exemplarily, the data feature #1 includes but is not limited to PV, UV, DAU, the number of collections, the number of likes, data sharing-related features, etc.

[0122] Therefore, in the embodiment of the present application, when the first value score of the first data is not lower than the preset value, the parent data or child data corresponding to the first data is obtained according to the data lineage of the first data to obtain the second data, and the value score of the second data is updated according to the first value score of the first data, so that when the first data is a high-value data asset, the value score of the parent asset or child asset corresponding to the first data, that is, the second data, can be updated according to the first value score of the first data, so as to make up for the value score of the parent asset or child asset of the first data, realize the transfer of the importance of the data asset to the upper-level parent asset or the lower-level child asset, that is, realize the cross-layer transfer of the value score of the data asset, make the evaluation of data importance more perfect, fair and reasonable, and thus improve the accuracy of the evaluation of the value score of the data asset.

[0123] In addition, the embodiment of the present application can automatically identify the importance of data assets without manually marking the scope of important assets, thereby reducing the time for dividing important data assets.

[0124] In some embodiments, step S340 may be specifically implemented as obtaining an additional score for the value score of the second data according to the first value score, and obtaining the second value score of the second data based on the value score of the second data and the additional score.

[0125] Specifically, an additional score for the corresponding value score of the second data may be obtained according to the first value score of the first data, and the additional score is then added to the corresponding value score of the second data to update the corresponding value score of the second data, so as to make up for the value score of the second data to obtain the second value score of the second data, and to realize the transfer of the importance of the data asset to the upper-level parent asset or to the lower-level child asset.

[0126] Specifically, an additional score for the value score of the second data may be obtained according to the first value score and the attenuation coefficient corresponding to the first value score.

[0127] Specifically, in order to reflect that the first value score of the first data is only one factor affecting the corresponding value score of the second data, based on this, the first value score of the first data may be multiplied by an appropriate attenuation coefficient to reflect the influence degree of the first value score of the first data on the corresponding value score of the second data, so that the second value score of the second data is more perfect and reasonable.

[0128] In one implementation manner, the first value score may be multiplied by the attenuation coefficient to obtain the additional score.

[0129] Optionally, the attenuation coefficient value corresponding to the preset value when the attenuation is close to 0 at the maximum transfer depth may also be determined according to the relationship among the value score of the data, the data transfer depth, and the attenuation coefficient, and then the attenuation coefficient value may be determined as the attenuation coefficient corresponding to the first value score.

[0130] Specifically, the relationship among the value score of the data, the data transfer depth, and the attenuation coefficient is used to represent the attenuation situation of the value score of the data with the change of the data transfer depth. Exemplarily, in the case of one-way traceability transfer, the data value score may exponentially decay with the change of the data transfer depth, for example, it may be expressed as the following formula (1):

[0131] Score=Score0*μ n (1)

[0132] Wherein, Score represents the attenuated value score, Score0 represents the initial value score, μ represents the attenuation coefficient, and n represents the data transfer depth.

[0133] In one implementation, according to the data warehouse layering rules, it is mainly divided into the original data (Operational Data Store, ODS), the data detail layer (Data Warehouse Detail, DWD), the data service layer (Data Warehouse Service, DWS), and the data application layer (Application Data Store, ADS). Since the new data assets generated after being processed in the data warehouse usually enter another layer, the maximum depth of data transfer is 2 to the 4th power, that is, 16 layers. It can be understood that too much transfer depth affects program performance and the additional score is approximately equal to 0, while too low transfer depth cannot cover the entire data link.

[0134] It should be noted that the data warehouse is usually divided into 4 layers (ODS, DWD, DWS, ADS), and the transfer depth is 2 4 = 16 times is to cover the data flow requirements in extremely complex scenarios. At the same time, 2 is selected as the base for calculation because in a standard data warehouse, an asset usually generates another asset after being processed (a multiple relationship of 2). Through exponential decay, it can maximize the guarantee that data reaches the deepest layer while avoiding excessive fractional supplementation.

[0135] In addition, by setting the threshold corresponding to the data value score (i.e., the above preset value), high-value data assets can be guaranteed. High-value data assets can be used as the prerequisite for cross-layer transfer of the value score. Exemplarily, in the value score system provided in the embodiments of the present application, high-value assets can be determined according to the 80 / 20 distribution principle of data quality. For example, when applying 2 category features based on user recognition and popularity, the maximum value score reaches 90 points. At this time, according to the 80 / 20 distribution principle of quality, data with a value score of more than 70 points will become high-value assets, exceeding most data assets.

[0136] Combining the relationship between the value score of data, the data transfer depth, and the attenuation coefficient, for the prerequisite of cross-layer transfer of the value score (such as the value score reaching 70 points and above), it is expected that the additional score is close to 0 after at most the maximum data transfer depth times. At this time, the corresponding attenuation coefficient μ is 0.6. Exemplarily, for a value score of Score0 = 70, when the attenuation coefficient μ is 0.6, the additional score after 16 transfers is 0.02, which is close to 0. Therefore, the embodiments of the present application can not only ensure that the parent data asset gets score supplementation, but also avoid excessive score supplementation, which may affect its own value score too much.

[0137] Optionally, the attenuation coefficient may also be determined based on at least one of the blood relationship distance coefficient between the first data and the second data, the blood relationship depth coefficient between the first data and the second data, the time decay factor of the first data, the usage frequency coefficient of the first data, and the weight of the first data.

[0138] Specifically, the blood relationship distance coefficient is used to characterize the closeness of the blood relationship between the first data and the second data. For example, when the first data is a direct sub - data of the second data, the blood relationship distance coefficient is 1.0; when the second data is an indirect sub - data of the second data, the blood relationship distance coefficient is 0.5. The blood relationship depth coefficient is used to characterize the depth of the blood relationship between the first data and the second data, such as sub - data, grand - data, etc. For example, when the first data is a sub - data of the second data, the blood relationship depth coefficient is 0.8; when the first data is a grand - data of the second data, the blood relationship depth coefficient is 0.5. The time decay factor of the first data characterizes the influence of the update time of the first data on the evaluation score of the first data. For example, when the first data is new data (such as the update time is 1 day away from the current time), the time decay factor value can be 0.8; when the first data is old data (such as the update time is 1 year away from the current time), the time decay factor value can be 0.3. The usage frequency coefficient of the first data is used to characterize the usage frequency of the first data. For example, when the usage frequency of the first data is high, the usage frequency coefficient is 1.2; when the usage frequency of the first data is low, the usage frequency is 0.8. The weight of the first data is used to indicate the importance of the first data to the second data. For example, when the first data is important to the second data, the weight can be 1.0; when the first data is not important to the second data, the weight can be 0.7.

[0139] Exemplarily, the basic attenuation coefficient (such as 0.6 above) can be multiplied by at least one of the blood relationship distance coefficient, the blood relationship depth coefficient, the time decay factor, the usage frequency coefficient of the first data, and the weight as the final attenuation coefficient.

[0140] Exemplarily, data comparison can be performed based on different attenuation coefficients, and the attenuation coefficient value can be determined according to whether the final score meets the ideal score. Optionally, the value of the attenuation coefficient is between 0 and 1, for example, it can take the value of 0.6.

[0141] In some embodiments, referring to Figure 4 , method 300 may further include the following steps S350 and S360.

[0142] S350, when the second value score is not lower than the preset value, obtain the parent data or sub - data corresponding to the second data according to the data blood relationship corresponding to the second data to obtain the third data.

[0143] Specifically, the process of obtaining the third data by acquiring the parent data or child data corresponding to the second data is similar to the process of obtaining the third data by acquiring the parent data or child data corresponding to the first data. Step S350 can refer to Figure 3 the relevant description in step S330 in

[0144] S360. Update the value score corresponding to the third data according to the second value score to obtain at least one third value score of the third data.

[0145] Specifically, the process of obtaining the third value score according to the second value score is similar to the process of obtaining the second value score according to the first value score. Step S360 can refer to Figure 3 the relevant description in step S340 in

[0146] That is to say, after updating the value score of the parent data or child data according to the first value score of the first data, it is also possible to identify again whether the parent data or child data (i.e., the second data) meets the criteria of high-value score data assets, and continue to transfer across layers upward or downward, so that the importance of high-value score data assets can continue to be transferred across layers, further improving the data importance evaluation process.

[0147] It can be understood that the processes of steps S350 and S360 can be repeated. That is, as long as the value score of the parent data or child data meets the criteria of high-value score data assets, then the value score of the parent data or child data can continue to be transferred across layers upward or downward to make up for the value score of its corresponding parent data or child data. In this repeated process, only the value score of the parent data can be transferred across layers upward, or only the value score of the child data can be transferred across layers downward, or both the value score of the parent data can be transferred across layers upward and the value score of the child data can be transferred across layers downward.

[0148] In some embodiments, continue to refer to Figure 4 , method 300 may further include the following step S370.

[0149] S370. In the case where the second value score is lower than the preset value, determine that the second value score is not used to update the value score of the parent data or child data corresponding to the second data. At this time, do not acquire the parent data or child data corresponding to the second data.

[0150] Specifically, when it is recognized that the parent data or child data (i.e., the second data) does not meet the criteria of high-value score data assets, that is, the parent data or child data is not a high-value score data asset, it can be determined that the importance of the parent data or child data will no longer be transferred across layers.

[0151] Optionally, the preset value corresponding to the second value score can also be dynamically adjusted according to factors such as the data quality of the second data and the historical scoring trend. For example, when the data quality score of the second data is relatively high, a lower preset value can be set, so that the second value score of the second data can participate in the value score update of its corresponding parent data or child data. Another example is that when the historical scoring trend of the second data in the recent month shows an upward trend, a lower preset value can be set, so that the second value score of the second data can also participate in the value score update of its corresponding parent data or child data.

[0152] A possible implementation is to repeatedly execute steps S350 and S360 until the value score of the parent data or child data does not meet the criteria for high-value score data assets.

[0153] Exemplarily, referring to Figure 2 , the application value score 240 module can execute the process of steps S350 to S370 to achieve cross-layer transfer of high value scores of data assets.

[0154] Therefore, after updating the value score of the parent data or child data according to the first value score of the first data in the embodiment of the present application, it is also possible to re-identify whether the parent data or child data (i.e., the second data) meets the criteria for high-value score data assets, and continue to transfer cross-layer upward or downward, so that the importance of high-value score data assets can continue to be transferred cross-layer until the value score of the parent data or child data no longer meets the criteria for high-value score data assets, further improving the data importance evaluation process and obtaining a more fair, reasonable, and accurate value score.

[0155] In some embodiments, continuing to refer to Figure 4 , when the number of at least one second value score is at least two, method 300 may further include the following step S380:

[0156] S380, perform a weighted sum of at least two second value scores to obtain the total value score of the second data.

[0157] Exemplarily, when the value scores of the related data of the second data meet the condition of not being used to update the value score of the second data, a weighted sum of at least two second value scores of the second data can be performed to obtain the total value score of the second data.

[0158] Specifically, when the value scores of the related data of the second data, such as parent data, child data, sibling data, grandchild data, etc., all satisfy that they are not used to update the value score of the second data, for example, for the related data of the second data, the processes of repeatedly executing steps S350 and S360 are performed, and the value score of the second data has been updated. At this time, the total value score of the second data can be obtained according to at least one updated second value score of the second data.

[0159] Exemplarily, each second value score corresponds to one or more data features, indicating the importance degree of the corresponding data features in business applications. Therefore, when the number of second value scores is at least two, the at least two second value scores can be weighted and summed to obtain the total value score of the second data, so as to characterize the overall importance degree of the second data in business applications.

[0160] Exemplarily, for each data feature such as PV, UV, DAU, number of collections, number of likes, etc., a second value score can be respectively corresponding, and for data sharing type features, a second value score can be corresponding. For each second value score, it can be respectively multiplied by the corresponding weight coefficient, and then added to calculate the total value score of the second data.

[0161] One implementable way is that the weight of the value score corresponding to the data feature can be determined according to the importance degree of the data feature or whether it is a common feature. Exemplarily, the weight coefficients corresponding to common data features such as PV, UV, DAU, number of collections, number of likes, etc., can be 1, and the weight coefficient corresponding to data sharing type features can be 0.7.

[0162] Exemplarily, continue to refer to Figure 2 , the index weighting unit in the core calculation 230 module can execute the process of step S380 to implement the weighted sum of the value scores of different feature indexes.

[0163] It should be noted that in this application, after the cross-layer transfer of the value scores of all data assets is completed, the final value scores of each data asset can be weighted and summed respectively to obtain the total value scores of each data asset. Exemplarily, in Figure 4 , after performing steps S350 to S370 on all data to complete the cross-layer transfer of the value scores of all data assets, step S380 can be executed, that is, for each data asset, its corresponding at least two value scores are weighted and summed to obtain the total value score of each data asset.

[0164] It should be understood that in the embodiments of this application, the calculation of the total value score of the second data is used as an example for description. The calculation process of the total value scores of other data assets is similar to the calculation process of the total value score of the second data, and can refer to the relevant descriptions in the above text.

[0165] In some embodiments, with continued reference to Figure 4 , method 300 may further include the following step S390:

[0166] S390. When the total value score of the second data exceeds the maximum value of the preset value score range, determine that the total value score of the second data is the maximum value of the value score range.

[0167] Exemplarily, the maximum value of the preset value score range may be 100 points. When the total value score of the second data obtained according to step S380 exceeds 100 points, it may be determined that the total value score of the second data is still 100 points.

[0168] It should be noted that when the relevant solution uses data portraits and usage features as the scoring criteria for the level of application value scores, it heavily depends on the integrity of the basic data in the application, which may lead to the disadvantage of not being able to obtain high scores after lacking some unimportant data support. However, in the embodiments of the present application, by setting the maximum value of the value score range, when the total value score of the data exceeds the maximum value of the value score range, the total value score of the data still takes this maximum value, so that the lack of some unimportant data that affects the total value score has little or no impact on the total value score, thus solving the problem that the application value score heavily depends on the integrity of the basic data.

[0169] For example, the embodiments of the present application can configure higher weight values for important or frequently used data features. For example, if a data has access popularity after being used, then this data can obtain the value score corresponding to the access popularity feature index. For example, the score of this value score can reach 90 points. If the basic metadata is complete, then the total value score is relatively high, such as 140 points. Since 140 points exceeds the preset full score of 100 points, the total value score is still recorded as 100 points. This makes some special data (such as other data indicators besides access popularity) have little impact on the total value score. That is to say, even without the support of such special data, the total value score can still reach a high score. Therefore, the disadvantage of not being able to obtain high scores after lacking the support of special data can be solved.

[0170] In some embodiments, after classifying data features into different categories, when a new feature is mined, it can be added to the existing data features according to the feature category, fused with the existing features and re - allocated weights to evaluate the value score. Optionally, when the user finds that the value score does not match the actual situation, data can be supplemented into the basic metadata, or the weights of the data features can be re - allocated to calculate a more reasonable value score.

[0171] Exemplarily, the total value score of the data ranges from 1 to 100. The higher the score, the higher the corresponding application value. For example, data assets with a score below 20 are low-value data assets, those with a score between 20 and 80 are general assets, and those with a score above 80 are important assets.

[0172] Therefore, in the embodiments of the present application, by managing the basic metadata, adjusting the weights of data features, or adjusting the distribution of feature values, the user can make the value score estimation scheme adapt to different business scenarios, continuously strengthen the versatility in different fields, and meet the data value score evaluations of different forms, businesses, and standards.

[0173] In some embodiments, the total value score of at least one data may also be obtained. According to the attribution information of the at least one data, at least one data belonging to the target object is obtained, and then based on the total value score of the at least one data belonging to the target object, the data asset score of the target object is obtained.

[0174] Exemplarily, the total value score of all data assets can be obtained according to the method provided in the embodiments of the present application. After that, according to the attribution information of each data asset, at least one data belonging to the target object is obtained, and then based on the total value score of all data assets belonging to the target object, the data asset score of the target object is calculated. The target object may include an organization or an individual, without limitation.

[0175] Exemplarily, three application scenarios can be obtained through the attribution information of the data assets, which are as follows:

[0176] 1. Score each organization (business department, center, or group) to represent the importance of the organization in the data field;

[0177] 2. Score each data asset to represent the importance of the data asset in the data field;

[0178] 3. Score each object (such as the creator) to represent the importance of the object in the data field.

[0179] One implementable way is to sum up the total value scores of all data assets belonging to the same organization to obtain the value total score of the organization. Optionally, the average value obtained by dividing the value total score of the organization by the number of data assets belonging to the organization can be the average value score of the organization. Optionally, this average value score can be the final score of the application value.

[0180] The total value score of all data assets attributed to the same object can be summed to obtain the total value score of the object. Optionally, the average value score of the object can be obtained by dividing the total value score of the object by the number of data assets attributed to the object. Optionally, the average value score can be the final score of the application value.

[0181] Exemplarily, continue to refer to Figure 2 , the application value score 240 module can determine the application value scores of organizations and individuals according to the total value score of each data asset and the attribution information of the data asset.

[0182] Optionally, in the embodiments of the present application, the data asset score of the target object can also be obtained according to the total value score of at least one data attributed to the target object and the score interval coefficient.

[0183] Specifically, when calculating the data asset score of the target object according to the total value score of at least one data attributed to the target object, a certain coefficient, such as the score interval coefficient, can also be multiplied by the total value score of each data asset according to the weighted summation algorithm, and then averaged to obtain the data asset score of the target object. Exemplarily, the data asset score of the target object is as follows:

[0184] Data asset score = ∑(score of each asset * score interval coefficient) / total number of assets

[0185] In this way, the concept of data governance is incorporated to encourage users to optimize low-value assets into high-value assets through governance means. Exemplarily, the corresponding relationship between the score interval coefficient and the score interval (score interval -> interval coefficient): [0,10] -> 0.36, (10,20] -> 1.025, (20,40] -> 1.34, (40,60] -> 1.92, (60,80] -> 2.87, (80,100] -> 4.33. If a user performs governance and the value score of a certain asset changes from 10 points to 40 points, only 3.6 points will be provided for the asset owner when it is at 10 points, but after this asset is improved by 40 points, 52 points will be provided for the owner.

[0186] Therefore, in the embodiments of the present application, by introducing a coefficient product (i.e., a score interval coefficient) for each score interval where the data asset is located, high-value data assets will be multiplied by multiples ranging from 1 to 4.33 (for example, a data asset with an original score of 90 will provide a score of 90 * 4.33 in the final calculation), and low-value data assets will be multiplied by multiples ranging from 0.3 to 1 (for example, a data asset with an original score of 9 will provide a score of 9 * 0.36 in the final calculation). Under this algorithm model, the more high-value data assets there are, the higher the total score of the subject to which the data assets belong, and the more low-value data assets there are, the lower the total score of the subject to which the data assets belong. This is more conducive to promoting the willingness of the data asset owner to improve data governance and transforming their low-quality assets into high-quality assets.

[0187] One implementable way is to obtain a judgment matrix, where the element a in the judgment matrix ij represents the importance difference between the score interval corresponding to the j-th column and the score interval corresponding to the i-th row, and then, according to the elements in each row of the judgment matrix, determine the score interval coefficient of the score interval corresponding to each row.

[0188] Exemplarily, the score interval coefficient can be obtained according to the Analytic Hierarchy Process (AHP), the Technique for Order Preference by Similarity to an Ideal Solution (TOPSIS), etc. Specifically, as shown in Table 1, a two-dimensional judgment matrix can be established, with each row corresponding to each score interval and each column also corresponding to each score interval. The score intervals are, for example, [0, 10), [10, 20), [20, 40), [40, 60), [60, 80), [80, 100]. At the same time, set 5 values (1 - 5) to represent the importance level. Compare the factors in each level pairwise and use simple numbers to represent the importance difference between them. Among them, the importance level can be obtained through user input.

[0189] For example, in Table 1, a 51 is the element in the 5th row and 1st column. The score interval corresponding to the row is [60, 80), and the score interval corresponding to the column is [0, 10). The corresponding importance level is 5, indicating that the importance of the same asset after being promoted from the score interval [0, 10) to the score interval [60, 80) reaches the maximum value of 5, thereby indicating the improvement of asset quality. Conversely, a 15 is the element in the 1st row and 5th column. The score interval corresponding to the row is [0, 10), and the score interval corresponding to the column is [60, 80). The corresponding importance level is 1 / 5 (i.e., 0.2). Here, 5 corresponds to the previously set value, indicating that the importance of the same asset after being reduced from the score interval [60, 80) to the score interval [0, 10) reaches the minimum value of 0.2, thereby indicating the regression of asset quality. By analogy, the values corresponding to all elements in Table 1 can be obtained.

[0190] Then, for the elements corresponding to each score interval in the judgment matrix, determine the score interval coefficients for the score intervals corresponding to each row. For example, the average value of the respective elements corresponding to each row can be taken to obtain the score interval coefficient corresponding to each row. For example, for the score interval corresponding to the first row being [0, 10), the element values are 1, 0.25, 0.25, 0.25, 0.2, 0.2 respectively, and the final score interval coefficient obtained after taking the average is 0.358333. Similarly, the score interval coefficients corresponding to [10, 20), [20, 40), [40, 60), [60, 80), [80, 100] can be obtained as 1.025, 1.338889, 1.922222, 2.86667, 4.333333, etc. For the convenience of calculation, each score interval coefficient can be approximately equal to 0.36, 1.025, 1.34, 1.92, 2.87, 4.33.

[0191] Table 1

[0192] Weight value Score range 0-10 10-20 20-40 40-60 60-80 80-100 0.358333 0-10 1 0.25 0.25 0.25 0.2 0.2 1.025 10-20 4 1 0.5 0.25 0.2 0.2 1.338889 20-40 4 2 1 0.5 0.33333 0.2 1.922222 40-60 4 4 2 1 0.33333 0.2 2.866667 60-80 5 5 4 3 1 0.2 4.333333 80-100 5 5 5 5 5 1

[0193] Figure 5 is a schematic flowchart of another data management method 500 provided by an embodiment of the present application. This method 500 can be executed by any electronic device with data processing capabilities, such as a server. As an example, method 500 can be executed by Figure 1A each node 101 therein, or the data platform deployed in each node 101, without limitation.

[0194] It should be understood that Figure 5 shows the steps or operations of data management method 500, but these steps or operations are only examples. Embodiments of the present application can also perform other operations or Figure 5 transformations of each operation therein. In addition, Figure 5 each step therein can be executed in a different order from that Figure 5 presented, and it is possible that not all the operations Figure 5 presented are to be executed.

[0195] As Figure 5 shown, this method 500 can include steps S501 to S512.

[0196] S501, input basic metadata.

[0197] Specifically, the basic metadata of at least one data asset can be input. Optionally, the data asset can include Figure 3 the first data and the second data therein. Specifically, the basic metadata can refer to the relevant descriptions in the above text.

[0198] S502, obtain at least one data non - sharing class feature.

[0199] Specifically, at least one data non - sharing class feature of each data asset can be determined according to the basic metadata of at least one input data asset. Exemplarily, as Figure 5 shown, the data non - sharing class features can include PV, UV, number of collections, number of likes, etc., without limitation. Specifically, the data non - sharing class features can refer to the relevant descriptions in Figure 3 .

[0200] S503. Obtain at least one data sharing class feature.

[0201] Specifically, at least one data sharing class feature of each data asset can be determined according to the basic metadata of at least one input data asset. Exemplarily, as Figure 5 shown, the data sharing class features can include data sharing information, data sharing type, shared service role, historical transaction proportion value, substitutability value, etc., without limitation. Specifically, the data sharing class features can refer to the relevant descriptions in Figure 3 .

[0202] Optionally, the data non - sharing class features and the data sharing class features can be stored in the same location.

[0203] S504. Grade and score.

[0204] Specifically, for the data non - sharing class features obtained in step S502, such as PV, UV, number of collections, number of likes, etc., they can be graded and scored according to the magnitude of the numerical values (feature values) corresponding to the data features. For example, the larger the numerical value (or feature value), the higher the score. Exemplarily, taking the data asset as the first data, the scoring result can correspond to Figure 3 the example of at least one first value score corresponding to the first data in Figure 3 . Specifically, the process of grading and scoring can refer to the relevant descriptions in step S320 in

[0205] S505. Model scoring.

[0206] Specifically, for the data sharing class features obtained in step S503, such as data sharing information, data sharing type, shared service role, historical transaction proportion value, substitutability value, etc., the data features can be input into a vector converter to obtain a vector representation, and then the vector representation is input into a neural network model, and the model outputs the scoring result corresponding to the overall data sharing class features. Exemplarily, taking the data asset as the first data, the scoring result can correspond to Figure 3 the example of at least one first value score corresponding to the first data in Figure 3 . Specifically, the process of model scoring can refer to the relevant descriptions in step S320 in

[0207] S506, Score statistics.

[0208] Exemplarily, the scoring results of step S504 and the scoring results in step S505 can be summarized in a scoring table.

[0209] It should be noted that steps S502 to S505 can correspond to the feature summarization layer. Specifically, in the process of feature integration in steps S502 to S505, different data features may come from different data platforms. At this time, these data features from different sources need to be integrated together to prepare for subsequent value scoring using data features. It can be understood that aggregating the features of data from different sources on the same data asset is a data connection process.

[0210] S507, Rule hit.

[0211] Exemplarily, for the score statistics result in step S506, cross-layer transfer rule hit matching can be performed. For example, when the value score corresponding to a certain data feature is not lower than (or greater than or equal to) a preset value, the rule is hit.

[0212] S508, Cross-layer transfer of value scores.

[0213] Specifically, perform cross-layer transfer of value scores for the value scores with rule hits. Specifically, the parent data or child data of the data asset can be obtained according to the data lineage of the data asset corresponding to the value score, and then the value score of the parent data or child data is updated according to the value score, obtaining the value score of the parent data or child data. For example, an additional score can be obtained by multiplying the value score by a preset decay coefficient, and then the additional score is added to the value score of the parent data or child data to obtain the updated value score of the parent data or child data.

[0214] Optionally, it can also be determined again whether the parent data or child data meets the high value score standard, and the value score is continuously transferred cross-layer upward, repeating this process until the parent asset or child asset no longer meets the high value asset standard.

[0215] S509, Total score of data assets.

[0216] Specifically, the feature scores and model scores of each data asset can be summarized. For example, they are multiplied by the corresponding weight coefficients 1 and 0.7 respectively, and then added together to obtain the total score of each data asset.

[0217] Optionally, this embodiment of the present application allows the total score of the data asset to exceed 100 points. When it exceeds 100 points, the expression form of the total score of the data asset is still 100 points, so as to meet the value score evaluation of different forms, services, and standards.

[0218] S510, Obtain the organization score.

[0219] Exemplarily, the total score can be obtained by weighted summation of the value scores of all data assets under the same organization name. Optionally, the total score can also be divided by the number of data assets under the organization name to obtain an average value as the final score of the organization.

[0220] S511, Obtain the asset score.

[0221] S512, Obtain the personal score.

[0222] Exemplarily, the total score can be obtained by weighted summation of the value scores of all data assets under the same individual (such as the creator) name. Optionally, the total score can also be divided by the number of data assets under the individual name to obtain an average value as the final score of the organization.

[0223] Therefore, the embodiments of the present application can automatically identify the importance of data assets and reduce the time for manually classifying the importance of data assets. In the design of value score estimation, in order to ensure the integrity and accuracy of the value score estimation function of data assets, the concept of cross-layer transmission is created, which enables data to provide data indicators in the data warehouse after being processed multiple times, such as product DAU, Region of Interest (ROI), which are more concerned, and the parent assets or sub-assets in the data processing process play a connecting role and also receive corresponding attention as an indispensable part, thereby realizing the upward or downward transmission of the importance of data assets and ensuring fairness, accuracy, and irreplaceability in the process of data output.

[0224] The preferred embodiments of the present application have been described in detail above with reference to the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all belong to the protection scope of the present application. For example, in the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present application will not separately describe various possible combination methods. Again, any combination can be made between various different embodiments of the present application as long as it does not violate the idea of the present application, and it should also be regarded as the content disclosed by the present application.

[0225] It should also be understood that in the various method embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0226] The method provided by the embodiments of the present application has been described above. Next, the device provided by the embodiments of the present application will be described.

[0227] Figure 6 It is a schematic diagram of a data management device 10 provided by an embodiment of the present application. As Figure 6 shown, the device 10 includes: a first acquisition unit 11, a scoring unit 12, a second acquisition unit 13, and an update unit 14.

[0228] The first acquisition unit 11 is configured to acquire at least one data feature of the first data;

[0229] The scoring unit 12 is configured to estimate a value score based on the at least one data feature to obtain at least one first value score corresponding to the first data;

[0230] The second acquisition unit 13 is configured to, when the first value score is not lower than a preset value, acquire the parent data or child data corresponding to the first data according to the data lineage corresponding to the first data to obtain second data;

[0231] The update unit 14 is configured to update the value score corresponding to the second data according to the first value score to obtain at least one second value score of the second data.

[0232] In some embodiments, the update unit 14 is configured to update the value score corresponding to the second data according to the first value score to obtain at least one second value score of the second data, including:

[0233] Obtain an additional score of the value score of the second data according to the first value score and the attenuation coefficient corresponding to the first value score;

[0234] Obtain the second value score of the second data according to the value score of the second data and the additional score.

[0235] In some embodiments, the update unit 14 is further configured to:

[0236] Determine the attenuation coefficient value corresponding to the preset value when the maximum transmission depth decays to be close to 0 according to the relationship between the value score of the data, the data transmission depth, and the attenuation coefficient;

[0237] Determine the attenuation coefficient value as the attenuation coefficient corresponding to the first value score.

[0238] In some embodiments, the update unit 14 is further configured to:

[0239] In the case where the second value score is lower than the preset value, it is determined that the second value score is not used to update the value scores of the parent data or child data corresponding to the second data.

[0240] In some embodiments, the at least one data feature includes at least one of a data non - shared type feature and a data shared type feature;

[0241] Among them, the data non - shared type feature is the feature of the data in the data product it generates, and the data shared type feature is the feature after the data is shared to another data product.

[0242] In some embodiments, the data non - shared type feature includes at least one of page view volume, number of independent visitors, number of daily active users, number of favorites, and number of likes.

[0243] In some embodiments, the data shared type feature includes at least one of data sharing situation information, data sharing usage type, data sharing service role information, historical transaction proportion value information, and substitutability value information.

[0244] In some embodiments, the scoring unit 12 is used to estimate a value score according to the at least one data feature, and obtain at least one first value score corresponding to the first data, including:

[0245] Determine the first value score corresponding to each data feature according to the position of the value corresponding to each data feature in the value distribution corresponding to each data feature.

[0246] In some embodiments, the scoring unit 12 is used to estimate a value score according to the at least one data feature, and obtain at least one first value score corresponding to the first data, including:

[0247] Input the data feature into a vector converter to map the data feature into a first vector representation;

[0248] Input the first vector representation into a neural network model to output the first value score corresponding to the data feature; wherein, the vector converter and the neural network model are trained according to data with value score labels.

[0249] In some embodiments, the neural network model includes a multi - layer perceptron classifier.

[0250] In some embodiments, the apparatus 10 further includes a summing unit, which is used to perform a weighted sum of at least two second value scores of the second data to obtain a total value score of the second data when the value scores of the related data of the second data meet the condition of not being used to update the value score of the second data.

[0251] In some embodiments, the summing unit is further configured to:

[0252] When the total value score of the second data exceeds the maximum value of the preset value score range, determine that the total value score of the second data is the maximum value of the value score range.

[0253] In some embodiments, a third obtaining unit is further included, configured to:

[0254] Obtain the total value score of at least one data;

[0255] Obtain at least one data belonging to the target object according to the attribution information of the at least one data;

[0256] Obtain the data asset score of the target object according to the total value scores of at least one data belonging to the target object.

[0257] In some embodiments, the third obtaining unit is further configured to:

[0258] Obtain a judgment matrix, where the element a in the judgment matrix ij represents the importance difference between the score interval corresponding to the j-th column and the score interval corresponding to the i-th row;

[0259] Determine the score interval coefficient of the score interval corresponding to each row according to the elements of each row in the judgment matrix;

[0260] Wherein, the obtaining the data asset score of the target object according to the total value scores of at least one data belonging to the target object includes:

[0261] Obtain the data asset score of the target object according to the total value scores of at least one data belonging to the target object and the score interval coefficient.

[0262] In some embodiments, the first obtaining unit 10 is configured to obtain at least one data feature of the first data, including:

[0263] Obtain at least one data feature of the first data according to the basic metadata of the first data.

[0264] In the embodiment of the present application, when the first value score of the first data is not lower than a preset value, the parent data or child data corresponding to the first data is obtained according to the data lineage of the first data to obtain the second data, and the value score of the second data is updated according to the first value score of the first data. When the first data is a high-value data asset, the value score of the parent asset or child asset corresponding to the first data, that is, the second data, can be updated according to the first value score of the first data, so as to make up for the value score of the parent asset or child asset of the first data, and realize the transfer of the importance of the data asset to the upper-level parent asset or the lower-level child asset, that is, realize the cross-layer transfer of the value score of the data asset, making the evaluation of the data importance more perfect, fair and reasonable, thereby improving the accuracy of the evaluation of the value score of the data asset.

[0265] In addition, the embodiment of the present application can automatically identify the importance of data assets without manual annotation of the scope of important assets, thereby reducing the time for dividing important data assets.

[0266] It should be understood that the apparatus embodiment and the method embodiment can correspond to each other, and similar descriptions can refer to the method embodiment. To avoid repetition, it will not be elaborated here. Specifically, Figure 6 The data management apparatus 10 shown can execute the corresponding processes in the above method embodiments, and the foregoing and other operations and / or functions of each module in the apparatus 10 respectively correspond to the corresponding processes of the above respective methods 300. For the sake of brevity, it will not be elaborated here.

[0267] The apparatus of the embodiment of the present application has been described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional module can be implemented in the form of hardware, or in the form of instructions in software, or in the form of a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiment of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0268] Figure 7 is a schematic block diagram of an electronic device 30 provided by the embodiment of the present application.

[0269] As Figure 7 shown, the electronic device 30 may include:

[0270] A memory 31 and a processor 32. The memory 31 is used to store a computer program and transmit the program code to the processor 32. In other words, the processor 32 can call and run the computer program from the memory 31 to implement the method in the embodiments of the present application.

[0271] For example, the processor 32 can be used to execute the data management method in the above method embodiment according to the instructions in the computer program, including:

[0272] Obtain at least one data feature of the first data;

[0273] Estimate a value score according to the at least one data feature to obtain at least one first value score corresponding to the first data;

[0274] When the first value score is not lower than a preset value, obtain the parent data or child data corresponding to the first data according to the data lineage corresponding to the first data to obtain second data;

[0275] Update the value score corresponding to the second data according to the first value score to obtain at least one second value score corresponding to the second data.

[0276] In some embodiments of the present application, the processor 32 may include but is not limited to:

[0277] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0278] In some embodiments of the present application, the memory 31 includes but is not limited to:

[0279] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0280] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 31 and executed by the processor 32 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0281] As Figure 7 shown, the electronic device 30 may further include:

[0282] A transceiver 33, which may be connected to the processor 32 or the memory 31.

[0283] Among them, the processor 32 can control the transceiver 33 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 33 may include a transmitter and a receiver. The transceiver 33 may further include an antenna, and the number of antennas may be one or more.

[0284] It should be understood that the various components in the electronic device are connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.

[0285] The present application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments. Or rather, the embodiments of the present application also provide a computer program product containing instructions. When the instructions are executed by a computer, the computer executes the methods in the above method embodiments.

[0286] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).

[0287] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0288] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.

[0289] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0290] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A data management method, characterized in that: include: Acquire at least one data feature of the first data; Estimating the value score according to the at least one data feature to obtain at least one first value score corresponding to the first data; When the first value score is not lower than a preset value, obtaining the parent data or the child data corresponding to the first data according to the data lineage relationship corresponding to the first data to obtain the second data; The value score corresponding to the second data is updated according to the first value score to obtain at least one second value score for the second data.

2. The method according to claim 1, characterized in that The updating of the value score corresponding to the second data according to the first value score to obtain at least one second value score of the second data includes: Obtaining additional points for the value score of the second data according to the first value score and the attenuation coefficient corresponding to the first value score; A second value score of the second data is obtained according to the value score of the second data and the additional score.

3. The method according to claim 2, characterized in that Also includes: According to the relationship between the value score of the data, the data transmission depth and the attenuation coefficient, determining the attenuation coefficient value corresponding to the preset value when the maximum transmission depth is attenuated to be close to 0; The attenuation coefficient value is determined as the attenuation coefficient corresponding to the first value score.

4. The method according to claim 1, characterized in that: Also includes: When the second value score is lower than the preset value, it is determined that the second value score is not used to update the value score of the parent data or the child data corresponding to the second data.

5. The method according to claim 1, characterized in that The at least one data feature includes at least one of a data non-sharing feature and a data sharing feature; The non-data sharing feature is a feature of the data in the data product generated by the data, and the data sharing feature is a feature of the data after it is shared with another data product.

6. The method according to claim 5, characterized in that The non-data sharing characteristics include at least one of page views, number of independent visitors, number of daily active users, number of collections and number of likes; the data sharing characteristics include at least one of data sharing situation information, data sharing usage type, data sharing service role information, historical transaction ratio value information, and substitutability value information.

7. The method according to any one of claims 1 to 6, characterized in that: The estimating the value score according to the at least one data feature to obtain at least one first value score corresponding to the first data includes: The first value score corresponding to each of the data features is determined according to the position of the numerical value corresponding to each of the data features in the distribution of the numerical values ​​corresponding to each of the data features.

8. The method according to any one of claims 1 to 6, characterized in that: The estimating the value score according to the at least one data feature to obtain at least one first value score corresponding to the first data includes: Inputting the data features into a vector converter to map the data features into a first vector representation; The first vector representation is input into a neural network model, and the first value score corresponding to the data feature is output; wherein the vector converter and the neural network model are trained based on data with value score labels.

9. The method according to any one of claims 1 to 6, characterized in that: When the number of the at least one second value score is at least two, the method further comprises: When the value score of the data related to the second data satisfies the requirement not to update the value score of the second data, a weighted sum is taken for at least two second value scores of the second data to obtain a total value score of the second data.

10. The method according to claim 9, characterized in that Also includes: If the total value score of the second data exceeds the maximum value of a preset value score range, the total value score of the second data is determined to be the maximum value of the value score range.

11. The method according to claim 9, characterized in that Also includes: Obtain the total value score of at least one data; Acquire at least one data belonging to the target object according to the attribution information of the at least one data; A data asset score of the target object is obtained according to the total value score of at least one data belonging to the target object.

12. The method according to claim 11, characterized in that Also includes: Get the judgment matrix, the element a in the judgment matrix ij Indicates the importance difference between the score interval corresponding to the jth column and the score interval corresponding to the ith row; Determine the score interval coefficient of the score interval corresponding to each row according to the elements of each row in the judgment matrix; Wherein, obtaining the data asset score of the target object according to the total value score of at least one data belonging to the target object includes: A data asset score of the target object is obtained according to the total value score of at least one data belonging to the target object and the score interval coefficient.

13. The method according to any one of claims 1 to 6, characterized in that: The acquiring at least one data feature of the first data includes: At least one data feature of the first data is acquired according to basic metadata of the first data.

14. A data management device, characterized in that: include: A first acquisition unit, configured to acquire at least one data feature of first data; A scoring unit, configured to estimate a value score according to the at least one data feature, and obtain at least one first value score corresponding to the first data; A second acquisition unit is used to acquire the parent data or child data corresponding to the first data according to the data blood relationship corresponding to the first data to obtain the second data when the first value score is not lower than a preset value; An updating unit is used to update the value score corresponding to the second data according to the first value score to obtain at least one second value score of the second data.

15. An electronic device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 13.

16. A computer storage medium, characterized in that: Used to store a computer program, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1 to 13.

17. A computer program product, characterized in that The method comprises a computer program code, and when the computer program code is executed by an electronic device, the electronic device executes the method according to any one of claims 1 to 13.

Citation Information

Cited By

  • Matrix type data layering intelligent management system

    CN121070987A