Data processing method and device and electronic equipment

By establishing a blood relationship map in the data recommendation system and collecting and quantifying the usage and cost data of data resources, the problem of uncontrollable resource costs is solved, resource optimization and cost management are made transparent, and the return on investment of the business is improved.

CN120687622APending Publication Date: 2025-09-23NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510828606.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In data recommendation systems, resource cost management is opaque, resulting in uncontrollable resource usage costs, and traditional methods lack effective cost attribution and resource optimization methods.

Method used

By establishing a blood relationship map, we collect usage data and cost data of data resources in various business scenarios, and quantify them to obtain the usage cost corresponding to each business scenario.

Benefits of technology

It achieves a detailed understanding of the resource usage costs of each business scenario, optimizes resource allocation, reduces resource waste, and improves the return on investment of the business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687622A_ABST
    Figure CN120687622A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device and electronic equipment, and the method comprises the steps: building a blood relationship graph according to a business scene contained in a target data system and data resources associated with the business scene; wherein the blood relationship graph is used for indicating a corresponding relationship between each business scene contained in the target data system and the data resource; based on the blood relationship graph, collecting use data and cost data of the data resources by each business scene; and quantifying the use data and the cost data to obtain the use cost corresponding to each service scene. According to the mode, the use data and the cost data of each service for the data resources are collected through the blood relationship graph, and the use cost of each service for the data resources is obtained through quantitative calculation of the use data and the cost data, so that a user can know the use cost corresponding to each service in detail, and the user experience is improved. And resource optimization is carried out on each service based on the use cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data management, and in particular to a data processing method, device, and electronic device. Background Art

[0002] Data recommendation systems utilize a variety of big data and server resources. During the growth phase, a minimum resource allocation is typically requested initially, with investment increased later based on business growth. Capacity expansion and contraction are also implemented based on the monitoring of various data components. However, as business volume increases, resource costs rise. From a manager's perspective, the resource usage costs incurred by a large number of online businesses are completely unknown. Historical data storage and computing usage accumulates, making it impossible to accurately clean up the storage of many offline businesses. Blindly deleting data or stopping operations can impact online businesses, ultimately leading to uncontrollable resource costs. Summary of the Invention

[0003] The present disclosure aims to provide a data processing method, apparatus, and electronic device to accurately calculate the usage cost corresponding to each service, so as to make the cost of each service controllable.

[0004] In a first aspect, the present disclosure provides a data processing method, which includes: establishing a blood relationship map based on the business scenarios contained in the target data system and the data resources associated with the business scenarios; wherein the blood relationship map is used to indicate: the correspondence between the various business scenarios and data resources contained in the target data system; based on the blood relationship map, collecting usage data and cost data of the data resources for each business scenario; quantifying the usage data and cost data to obtain the usage cost corresponding to each business scenario.

[0005] In the second aspect, the present disclosure provides a data processing device, which includes: a relationship establishment module, which is used to establish a blood relationship map based on the business scenarios contained in the target data system and the data resources associated with the business scenarios; wherein the blood relationship map is used to indicate: the correspondence between the various business scenarios and data resources contained in the target data system; a data collection module, which is used to collect usage data and cost data of data resources for each business scenario based on the blood relationship map; a data quantification module, which is used to quantify the usage data and cost data to obtain the usage cost corresponding to each business scenario.

[0006] In a third aspect, the present disclosure provides an electronic device, which includes a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned data processing method.

[0007] In a fourth aspect, the present disclosure provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned data processing method.

[0008] The embodiments of the present disclosure bring the following beneficial effects:

[0009] The present disclosure provides a data processing method, device, and electronic device. First, a blood relationship map is established based on the business scenarios and data resources associated with the business scenarios contained in the target data system. The blood relationship map is used to indicate the corresponding relationship between the various business scenarios and data resources contained in the target data system. Then, based on the blood relationship map, the usage data and cost data of the data resources for each business scenario are collected. The usage data and cost data are then quantified to obtain the usage cost corresponding to each business scenario. This method collects the usage data and cost data of the data resources for each business through the blood relationship map, and obtains the usage cost of the data resources for each business through quantitative calculation of the usage data and cost data. This allows users to understand the usage cost corresponding to each business in detail and optimize resources for each business based on the usage cost.

[0010] Other features and advantages of the present disclosure will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by practicing the above-mentioned technology of the present disclosure.

[0011] In order to make the above-mentioned objects, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are specifically listed below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0013] Figure 1 A flowchart of a data processing method provided in an embodiment of the present disclosure;

[0014] Figure 2 A schematic diagram of a blood relationship map provided in an embodiment of the present disclosure;

[0015] Figure 3 A flowchart of full-link resource intelligent diagnosis and management provided by an embodiment of the present disclosure;

[0016] Figure 4 A schematic structural diagram of a data processing device provided in an embodiment of the present disclosure;

[0017] Figure 5 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all of them. Generally, the components of the embodiments of the present disclosure described and shown in the drawings herein can be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure as claimed, but merely represents selected embodiments of the present disclosure. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without creative effort shall fall within the scope of protection of the present disclosure.

[0020] With the continuous development of the gaming industry, the design of various in-game systems has become increasingly complex, leading to information overload for players. Therefore, personalized design scenarios have emerged in games. Currently, games have integrated personalized recommendation services across a wide range of topics, such as online shopping predictions, online shopping rankings, friend recommendations, gameplay recommendations, dynamic difficulty adjustment, real-time gift packs, matchmaking hosting, and input association. Personalized recommendation services are advanced business intelligence services built on massive data mining. Essentially, they are information filtering and sorting systems that provide users with personalized information services and decision support by predicting their item preferences.

[0021] Currently, the interactive entertainment recommendation system powers over a thousand businesses across over a hundred games, with tens of billions of requests per month. For such a highly concurrent and large-scale service system, during the growth phase, the team primarily focuses on project results and problem-solving. Algorithm strategy and engineering development personnel also focus on technical implementation solutions, easily overlooking cost factors. To meet higher technical metrics such as QPS (Queries Per Second) and ELC (Engineering Logical Complexity), or to meet greater data computing and data storage requirements, the most direct and effective approach is to invest more resources. However, from a global perspective, the lack of effective resource cost management ultimately leads to diminishing marginal returns—a mismatch between technological advancements and resource growth. In some cases, resource expansion may even outpace technological advancements, significantly reducing the project's return on investment (ROI).

[0022] A complete recommendation service involves the development of offline and real-time features, models, and algorithmic strategies. This involves the use of a variety of big data and server resources, including Spark, Flink, Kafka, Redis, ES, HBase, Hive, and GPUs. During the growth phase, a minimum resource allocation is typically requested initially, with investment increased based on business growth. Capacity expansion and contraction are also implemented based on the monitoring of various data components. As business volume increases, resource costs rise. However, from a manager's perspective, the resource usage costs of a large number of online businesses are completely unknown. Historical data storage and computing usage continue to accumulate, making it impossible to accurately clean up the storage of many offline businesses. Blindly deleting data or stopping operations can impact online businesses, ultimately leading to high monthly resource costs with no clear path to action.

[0023] The above existing technical solutions face technical challenges in multi-dimensional resource management: First, the system needs to integrate eight types of computing and storage data resources, such as Redis, Es, Hbase, and Kafka, which may have inherent defects such as inconsistent resource measurement standards and coarse cost attribution granularity; second, traditional management methods rely on manual experience for resource diagnosis and lack an evaluation system based on business value linkage, making it difficult to identify redundant resources in a timely manner; third, existing technologies only rely on basic utilization indicators to determine the health of resource usage, which can easily cause resource optimization decisions to deviate from business goals.

[0024] Based on the above problems, embodiments of the present invention provide a data processing method, apparatus, and electronic device. This technology can be applied to scenarios of business usage cost calculation and data asset management in data systems.

[0025] In order to facilitate understanding of the embodiments of the present disclosure, a data processing method disclosed in the embodiments of the present disclosure is first described in detail. Figure 1 As shown, the method includes the following specific steps:

[0026] Step S102: establishing a blood relationship map based on the business scenarios and data resources associated with the business scenarios included in the target data system; wherein the blood relationship map is used to indicate: the corresponding relationship between each business scenario and data resources included in the target data system.

[0027] The specific system corresponding to the target data system can be determined based on R&D requirements. For example, the target data system can be a data recommendation system or a data storage system. The target data system can include one or more business scenarios. Each business scenario will apply for the use of at least one data resource for operation. The data resources used by each business scenario can be the same or different.

[0028] The target data system's resource usage primarily involves data computation and storage for algorithmic strategies. Data resources can include Spark clusters, Flink clusters, Kafka clusters, Redis clusters, HBase clusters, Elk clusters, Hive, GPUs (Graphics Processing Units), and other big data and server resources. Spark clusters are used for offline computing, Flink and Kafka clusters for operational computing, Redis, HBase, and Elk clusters for feature data storage, Hive for log data storage, and GPUs for training and prediction of deep learning models or large models.

[0029] In the specific implementation, through unified upstream constraints and tool integration for the use of various data resources, that is, through top-down requirements and platform construction capabilities, all data resources are unified and integrated into the platform, thereby collecting the blood relationship between each business scenario and each data resource, establishing a fine-grained mapping relationship from business scenarios to data resources, and obtaining a blood relationship map.

[0030] Step S104: Based on the blood relationship graph, collect usage data and cost data of data resources in various business scenarios.

[0031] Because the lineage relationship map is used to indicate the data resources used in each business scenario, the usage and cost data of each data resource can be used to determine the usage and cost data of each business scenario. The cost data indicates the total storage capacity or total capacity of the data resource, while the usage data indicates the amount of data storage or data indexes used by the data resource.

[0032] In specific implementations, different data resources correspond to different usage data and cost data. For example, the usage data for Redis clusters, HBase clusters, and Elk clusters is data storage capacity, and the corresponding cost data is total capacity. Cost data for GPUs can include hosting fees and storage fees, while GPU usage data can include GPU usage time or space usage.

[0033] Step S106: quantify the usage data and cost data to obtain the usage cost corresponding to each business scenario.

[0034] After collecting the usage data and cost data corresponding to the data resources, the cost calculation or evaluation method for each data resource is different, and needs to be collected and calculated separately. Then, the usage data, cost data, etc. of each type of data resources in each business scenario are quantified, so as to obtain the usage cost corresponding to each business scenario based on the quantified usage data and cost data.

[0035] An embodiment of the present invention provides a data processing method, which collects usage data and cost data of data resources for each business through a blood relationship map, and obtains the usage cost of data resources for each business through quantitative calculation of the usage data and cost data, so that users can understand the usage cost corresponding to each business in detail and optimize resources for each business based on the usage cost.

[0036] The following examples are used to describe how to establish a blood relationship map.

[0037] Specifically, the specific process of establishing a blood relationship map based on the business scenarios and data resources associated with the business scenarios contained in the target data system may include: determining the data resources applied for by each business scenario contained in the target data system, and the data processing units used for the operation of each data resource; for each business scenario, establishing a correspondence between the data resources corresponding to the business scenario and the data processing units used for the operation of the data resources to obtain the blood relationship corresponding to the business scenario; integrating the blood relationships corresponding to each business scenario to obtain a blood relationship map.

[0038] In a specific implementation, the data resources requested by each business scenario are automatically allocated by the system or pre-set by the user. The data processing units used to operate each data resource can also be dynamically allocated or pre-set. Based on the data resources requested by each business scenario and the data processing units used to operate each data resource, a lineage relationship can be established for each business scenario. This lineage relationship indicates the corresponding relationship between the data resources corresponding to each business scenario and the data processing units used to operate that data resource.

[0039] Different data resources use different data processing units, which may include but are not limited to: HDFS file path, Kafka Topic partition, HBase table Region, Elk index, etc.

[0040] In an optional embodiment, a three-layer lineage model can be constructed based on a graph database. This three-layer lineage model is also known as a lineage relationship graph. This lineage relationship graph includes a business layer, a logical layer, and a physical layer. The business layer is used to store the various business scenarios contained in the target data system, the logical layer is used to store the data resources associated with each business scenario, and the physical layer is used to store the data processing units used to operate each data resource.

[0041] In its specific implementation, the graph database is a data management system designed with nodes and edges as its basic storage units and efficient storage and querying of graph data as its design principle. The business layer in the aforementioned lineage relationship graph uses specific business scenarios as its vertices. The logical layer associates each vertex of the business layer with a corresponding data resource. Such data resources may include, but are not limited to, algorithm model jobs, offline real-time feature jobs, Redis clusters, Hbase clusters, Elk clusters, etc. The physical layer maps the computation or storage usage generated by each data resource during operation to a specific data processing unit. This allows for the collection of lineage mapping and quantitative usage data, as well as the cost data of each data resource instance.

[0042] like Figure 2 FIG. 1 is a schematic diagram of a blood relationship map provided by an embodiment of the present invention. Figure 2 The business layer is used to store the business scenarios included in the recommendation system, including social recommendation scenarios, matching hosting scenarios, and gift package recommendation scenarios; the logical layer includes at least Redis cluster, Hbase cluster, Elk cluster, etc.; the physical layer includes at least HBase table Region, Elk index Index and key mode.

[0043] In practical applications, the relationship strength between the business layer, logical layer, and physical layer can be calculated through a dynamic weight algorithm to determine the lineage of business scenarios and data resources, thereby achieving two-way traceability.

[0044] The above approach establishes a fine-grained relationship map of physical resources from business units to specific key levels by implementing unified upstream constraints and tool integration for the use of various data resources.

[0045] The following examples are used to quantify the data.

[0046] Specifically, the above-mentioned specific process of collecting usage data and cost data of data resources for each business scenario based on the blood relationship map may include: determining the data resources corresponding to each business scenario based on the blood relationship map, and collecting the cost data of the data resources corresponding to each business scenario; collecting the operation data corresponding to the data processing unit used when the data resources corresponding to the business scenario are running, and determining the usage data of the data resources corresponding to the business scenario based on the operation data; determining the usage data of the data resources corresponding to each business scenario based on the blood relationship map and the usage data of the data resources corresponding to the business scenario.

[0047] In specific implementations, a distributed proxy architecture can be employed. A lightweight metadata collection agent is deployed on each compute node or minimum storage unit of the data resources corresponding to each business scenario to collect usage data corresponding to each data processing unit in real time. For example, runtime data such as Spark SQL execution plans, Flink job DAG topologies, and Kafka topic subscription relationships can be captured (this runtime information is also known as usage data). A regular expression rule engine is used to perform semantic parsing on static metadata such as Hive tables, Redis keys, HBase tables, and Elasticsearch indexes. The resulting data includes usage data.

[0048] Based on the above description, the specific process of quantifying usage data and cost data to obtain the usage costs corresponding to each business scenario may include: for each data resource, based on the data quantification logic corresponding to the current data resource and the usage data and cost data of the current data resource for each business scenario, determining the usage cost of the current data resource for each business scenario; wherein different data resources are respectively configured with corresponding data quantification logic; and integrating the usage costs of each business scenario for each data resource to obtain the usage costs corresponding to each business scenario.

[0049] In specific implementations, different data resources correspond to different data quantification logic, which is the method for calculating usage costs based on usage data and cost data. Specifically, this disclosure includes a multi-dimensional resource cost quantification system that can quantify usage data and cost data for different data resources.

[0050] In one specific embodiment, the data quantification logic for the Redis cluster is as follows: a small number of nodes are sampled to scan specific keys, and then classified into pattern keys based on the relationship graph and regular expressions, thereby classifying the keys into specific business scenarios. For example, in a game friend recommendation business with a business ID of 155326, if a job writes data with the pattern 155326_feature_${role_id}, then the specific keys 155326_feature_111 and 155326_feature_222 will be included in the resource storage of the game friend recommendation business, representing the cost of using the Redis cluster for the game friend recommendation business.

[0051] In another specific embodiment, the data quantization logic corresponding to the GPU device can be implemented by the following formula, where the service fee is the cost of using the GPU device in the service scenario:

[0052] Business expenses = (hosting fees + storage fees) * business share + 500;

[0053] Business ratio = business (number of GPU cards * duration) / total (number of GPU cards * duration);

[0054] In another specific embodiment, a time series database (TSDB) can be built through spatiotemporal dimension aggregation technology to store minute-level fine-grained indicators, and a sliding window aggregator can be developed to implement data quantification logic. In other words, real-time data can be counted by 5-minute windows through Flink, and daily or weekly statistics can be used to support report data.

[0055] In another specific embodiment, the data quantification logic can also be viewed through the system cost quantification model. The system cost quantification model can be simply understood as combining various data resources and business scenarios more closely through blood relationship maps and quantification. By making more reasonable use of the underlying data resources, the costs of various specific businesses of the target data system can be reduced, making its ROI higher and thus achieving higher economic value. For example, when the target data system is a game recommendation system, a recommendation resource utility equivalent (RUE) indicator is developed to convert GPU inference time into the marginal cost of gift package purchase conversion rate; a feedback value model for gameplay recommendation resources is constructed (such as the player retention value brought by optimizing the waiting time of the matching hosting system); and a dynamic game algorithm for real-time / offline resources is designed (using reinforcement learning to balance the resource competition between Flink real-time feature updates and Hive offline training). This method can improve the resource guarantee rate of high-value business units and reduce the waste of edge business resources. This is a unique capability in the gaming field that traditional Internet data governance solutions cannot achieve.

[0056] The following embodiments are used to describe data diagnosis and governance methods.

[0057] Specifically, after quantifying the usage data and cost data to obtain the usage cost corresponding to each business scenario, a health diagnosis of the data resources is performed based on the usage cost corresponding to each business scenario to obtain a health diagnosis result of the data resources.

[0058] In specific implementation, health diagnosis of various data resources is performed based on quantitative data. Health diagnosis plans can be customized by studying the characteristics of each type of data resource, and millisecond-level indicator normalization across heterogeneous resources can be achieved through adaptive collection agents. A standardized resource model that includes resource health and value coefficients is created, etc., so as to calculate the health score of each product and business's use of various data resources. This health score is also the above-mentioned health diagnosis result.

[0059] For example, the Redis health score calculation rule is: the physical layer's pattern key identifier is manageable (service offline, no read or write in the past month, whether there is an expiration date, etc.), and the storage ratio of the pattern key to be managed is calculated. If the full score is 100, 100*the storage ratio to be managed is the health score.

[0060] In practical applications, the health diagnosis methods disclosed herein include offline job health diagnosis and health game diagnosis of multimodal systems. Offline job health diagnosis includes: platforming data resource-related configurations through a low-code job development platform. After the user submits a job, the resource usage adaptive system analyzes and uses reasonable resources to monitor data during the job execution, including the job's memory usage, data distribution at each stage, and details of computational time. The monitoring data is analyzed to calculate the requested CPU Cores utilization, memory utilization, data skew, and other optimizable conditions, and then update the data in the resource adaptive system.

[0061] The game-based health diagnosis of this multimodal system includes developing a business health adversarial assessment network (for example, diagnosing resource competition between the GMV goal of mall ranking and the retention goal of gameplay recommendations), constructing a cross-temporal health status map (dynamically balancing sudden QPS increases for real-time gift packages with SLA guarantees for offline feature operations), and designing a game scenario-aware anomaly propagation tree (for example, analyzing the butterfly effect of Redis cache failure on friend recommendation latency). Compared to traditional single-dimensional detection solutions, this approach accelerates fault location in complex recommendation scenarios by eight times, even with an average of tens of billions of requests per month.

[0062] In practical applications, game-based diagnosis of the health of multimodal systems can be understood as: Diagnosing the health of various data resources requires more than simply measuring resource storage size and computational complexity. It must also be closely integrated with the actual business scenarios. For example, some business logic requires highly wasteful use of resources to achieve optimal results. While storage and computational complexity are high, this business's resource utilization is already industry-leading. Alternatively, others in this business could not achieve the same results using fewer resources. Therefore, even significant storage redundancy cannot be considered a sign of low health. Therefore, it is necessary to explore diverse methods for measuring health from a business implementation perspective, based on various business scenarios within the recommendation system, to achieve more accurate health diagnosis.

[0063] Furthermore, data resources can be elastically scaled and / or cleaned based on usage data corresponding to the data resources and health diagnosis results of the data resources.

[0064] In specific implementation, we can study the technical characteristics of each type of data resource, develop a multi-dimensional evaluation model that integrates the Necessity of Life Index (NEI) and cost-effectiveness ratio (CER), use algorithms to achieve predictive identification of zombie resources, and ultimately build an automated and intelligent governance architecture. Through quantitative health diagnosis results, data resources can be elastically scaled and cleaned with one click to improve resource utilization.

[0065] In an optional embodiment, the specific process of elastically scaling and / or data cleaning of data resources based on the usage data corresponding to the data resources and the health diagnosis results of the data resources may include: inputting the usage data corresponding to the data resources and the health diagnosis results of the data resources into a pre-prepared three-dimensional evaluation model, and outputting the identification results for the data resources through the three-dimensional evaluation model; wherein the three-dimensional evaluation model is used to perform quantitative analysis of the resource activity, citation value and business value of the data resources; and elastically scaling and / or data cleaning of the data resources based on the identification results.

[0066] In its implementation, a three-dimensional assessment model for the survival necessity index is constructed, conducting quantitative analysis across the three dimensions of resource activity, citation value, and business value. Specifically, each dimension is assigned a specific weight and rules. For example, citation value is weighted at 60, resource activity at 30, and business value at 10. The quantiles for each of these three indicators are then calculated, and the final score is derived through a weighted calculation. If the score is below 30%, the data resource is marked as a data resource for governance.

[0067] In a specific embodiment, the resource activity dimension is calculated using an exponential time decay function f(t) = e^(-0.001t), where t is the number of days without access. When t>30 days, the contribution value of the time factor drops to less than 5% of the initial value, significantly reducing the survival score of long-term idle resources. The reference value dimension ReferenceScore = Σ(δ^n), where δ is the decay factor and n is the number of consecutive days without references. When a data resource has no new references for 30 consecutive days, the reference value score drops to 5.9% of the initial value, effectively identifying isolated resources that have not been referenced for a long time. The business value dimension constructs a business topology network based on the knowledge graph and uses an improved PageRank algorithm to calculate the relevance score:

[0068] PR(p_i)=(1-d) / N+d*Σ(PR(p_j) / L(p_j))

[0069] Among them, d is the damping coefficient, and d is the business weight factor introduced to give weight bonus to resources on the key business chain.

[0070] like Figure 3 The figure shows a flow chart of full-link intelligent diagnosis and governance of resources provided by an embodiment of the present invention. First, a blood relationship must be established between business and data resources. By implementing unified upstream constraints and tool closures on the use of various data resources, a fine-grained blood relationship map of physical resources from business units to specific key levels must be established. Secondly, data such as resource usage and cost must be collected. The cost calculation or evaluation method for each data component is different, and they need to be collected and calculated separately. Finally, the usage, cost and other data of each product and business for each type of resource must be quantified. Then, a health diagnosis is performed on each type of resource based on the quantified data. By studying the characteristics of each type of data component, a health diagnosis plan is customized, and a health score is calculated for the use of each type of resource by each product and business. Finally, by studying the characteristics of various data components, automated diagnosis and governance tools are developed to diagnose and govern resources.

[0071] This paper proposes a data asset governance solution and implementation specifications for target data systems, including lineage building, usage quantification, health diagnosis, and problem management. This solution addresses the technical challenges of multidimensional resource management (integrating eight data resources, including Redis, Es, HBase, and Kafka), unifies resource measurement standards, and provides coarse-grained cost attribution. It replaces traditional resource management methods (which rely on manual experience) for resource diagnosis and establishes an evaluation system based on business value linkage to promptly identify redundant resources. It also establishes a dynamic evaluation model that incorporates dimensions such as data timeliness and business relevance to calculate resource health, integrating resource optimization decisions with business objectives. Furthermore, this method can significantly reduce project costs, eliminate resource waste, and thus improve business ROI.

[0072] Corresponding to the above method embodiment, the embodiment of the present invention provides a data processing device, such as Figure 4 As shown, the device includes

[0073] The relationship establishment module 40 is used to establish a blood relationship map based on the business scenarios contained in the target data system and the data resources associated with the business scenarios; wherein the blood relationship map is used to indicate: the corresponding relationship between each business scenario and data resources contained in the target data system.

[0074] The data collection module 41 is used to collect usage data and cost data of data resources in various business scenarios based on the blood relationship map.

[0075] The data quantification module 42 is used to quantify the usage data and cost data to obtain the usage cost corresponding to each business scenario.

[0076] The above-mentioned data processing device collects the usage data and cost data of data resources for each business through the blood relationship map, and obtains the usage cost of data resources for each business through quantitative calculation of the usage data and cost data, so that users can understand the usage cost corresponding to each business in detail and optimize the resources of each business based on the usage cost.

[0077] Furthermore, the above-mentioned relationship establishment module 40 is used to: determine the data resources applied for by each business scenario contained in the target data system, and the data processing units used for the operation of each data resource; for each business scenario, establish a corresponding relationship between the data resources corresponding to the business scenario and the data processing units used for the operation of the data resources to obtain the blood relationship corresponding to the business scenario; integrate the blood relationship corresponding to each business scenario to obtain a blood relationship map.

[0078] Furthermore, the above-mentioned blood relationship map includes a business layer, a logical layer and a physical layer; wherein, the business layer is used to store the various business scenarios contained in the target data system, the logical layer is used to store the data resources associated with each business scenario, and the physical layer is used to store the data processing units used for the operation of each data resource.

[0079] Furthermore, the above-mentioned data collection module 41 is used to: determine the data resources corresponding to each business scenario based on the blood relationship map, and collect the cost data of the data resources corresponding to each business scenario; collect the operation data corresponding to the data processing unit used when the data resources corresponding to the business scenario are running, and determine the usage data of the data resources corresponding to the business scenario based on the operation data; determine the usage data of the data resources corresponding to each business scenario based on the blood relationship map and the usage data of the data resources corresponding to the business scenario.

[0080] Furthermore, the above-mentioned data quantification module 42 is used to: for each data resource, determine the usage cost of the current data resource for each business scenario based on the data quantification logic corresponding to the current data resource and the usage data and cost data of the current data resource for each business scenario; wherein different data resources are respectively configured with corresponding data quantification logic; integrate the usage cost of each data resource for each business scenario to obtain the usage cost corresponding to each business scenario.

[0081] Furthermore, the above-mentioned device also includes a health diagnosis module, which is used to: after quantifying the usage data and cost data to obtain the usage cost corresponding to each business scenario, perform health diagnosis on the data resources based on the usage cost corresponding to each business scenario to obtain the health diagnosis result of the data resources.

[0082] Furthermore, the above-mentioned device also includes a resource governance module, which is used to: elastically scale and / or clean up data resources based on usage data corresponding to the data resources and health diagnosis results of the data resources.

[0083] Furthermore, the above-mentioned resource governance module is also used to: input the usage data corresponding to the data resources and the health diagnosis results of the data resources into a pre-prepared three-dimensional evaluation model, and output the identification results for the data resources through the three-dimensional evaluation model; wherein the three-dimensional evaluation model is used to perform quantitative analysis of the resource activity, citation value and business value of the data resources; and perform elastic scaling and / or data cleaning processing on the data resources based on the identification results.

[0084] The data processing device provided in the embodiment of the present disclosure has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding content in the aforementioned method embodiment.

[0085] The present disclosure also provides an electronic device, such as Figure 5 As shown, the electronic device includes a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned data processing method.

[0086] Specifically, the above-mentioned data processing method includes: establishing a blood relationship map based on the business scenarios contained in the target data system and the data resources associated with the business scenarios; wherein the blood relationship map is used to indicate: the correspondence between the various business scenarios and data resources contained in the target data system; based on the blood relationship map, collecting the usage data and cost data of the data resources for each business scenario; quantifying the usage data and cost data to obtain the usage cost corresponding to each business scenario.

[0087] The above data processing method collects the usage data and cost data of data resources for each business through the blood relationship map, and obtains the usage cost of data resources for each business through quantitative calculation of the usage data and cost data, so that users can understand the usage cost corresponding to each business in detail and optimize the resources of each business based on the usage cost.

[0088] In an optional embodiment, the above-mentioned step of establishing a blood relationship map based on the business scenarios contained in the target data system and the data resources associated with the business scenarios includes: determining the data resources applied for by each business scenario contained in the target data system, and the data processing units used for the operation of each data resource; for each business scenario, establishing a correspondence between the data resources corresponding to the business scenario and the data processing units used for the operation of the data resources to obtain the blood relationship corresponding to the business scenario; integrating the blood relationships corresponding to each business scenario to obtain a blood relationship map.

[0089] In an optional embodiment, the above-mentioned blood relationship map includes a business layer, a logical layer and a physical layer; wherein the business layer is used to store the various business scenarios contained in the target data system, the logical layer is used to store the data resources associated with each business scenario, and the physical layer is used to store the data processing units used for the operation of each data resource.

[0090] In an optional embodiment, the above-mentioned step of collecting usage data and cost data of data resources for each business scenario based on the blood relationship map includes: determining the data resources corresponding to each business scenario based on the blood relationship map, and collecting the cost data of the data resources corresponding to each business scenario; collecting the operation data corresponding to the data processing unit used when the data resources corresponding to the business scenario are running, and determining the usage data of the data resources corresponding to the business scenario based on the operation data; determining the usage data of the data resources corresponding to each business scenario based on the blood relationship map and the usage data of the data resources corresponding to the business scenario.

[0091] In an optional embodiment, the above-mentioned step of quantifying the usage data and cost data to obtain the usage cost corresponding to each business scenario includes: for each data resource, based on the data quantification logic corresponding to the current data resource and the usage data and cost data of the current data resource for each business scenario, determining the usage cost of the current data resource for each business scenario; wherein different data resources are respectively configured with corresponding data quantification logic; and integrating the usage cost of each data resource for each business scenario to obtain the usage cost corresponding to each business scenario.

[0092] In an optional embodiment, after the step of quantifying the usage data and cost data to obtain the usage costs corresponding to each business scenario, the above method also includes: performing a health diagnosis on the data resources based on the usage costs corresponding to each business scenario to obtain the health diagnosis results of the data resources.

[0093] In an optional embodiment, the above method further includes: elastically scaling and / or data cleaning of the data resources based on usage data corresponding to the data resources and health diagnosis results of the data resources.

[0094] In an optional embodiment, the above-mentioned steps of elastically scaling and / or data cleaning of data resources based on the usage data corresponding to the data resources and the health diagnosis results of the data resources include: inputting the usage data corresponding to the data resources and the health diagnosis results of the data resources into a pre-prepared three-dimensional evaluation model, and outputting the identification results for the data resources through the three-dimensional evaluation model; wherein the three-dimensional evaluation model is used to perform quantitative analysis of resource activity, citation value and business value of the data resources; and elastically scaling and / or data cleaning of the data resources based on the identification results.

[0095] Furthermore, Figure 5 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 101 , the communication interface 103 and the memory 100 are connected via the bus 102 .

[0096] The memory 100 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The system network element and at least one other network element are connected via Bluetooth through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0097] The processor 101 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 101 or by instructions in the form of software. The above-mentioned processor 101 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as a random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or register. The storage medium is located in the memory 100, and the processor 101 reads the information in the memory 100 and, in conjunction with its hardware, completes the steps of the method of the aforementioned embodiment.

[0098] An embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned data processing method.

[0099] Specifically, the above-mentioned data processing method includes: establishing a blood relationship map based on the business scenarios contained in the target data system and the data resources associated with the business scenarios; wherein the blood relationship map is used to indicate: the correspondence between the various business scenarios and data resources contained in the target data system; based on the blood relationship map, collecting the usage data and cost data of the data resources for each business scenario; quantifying the usage data and cost data to obtain the usage cost corresponding to each business scenario.

[0100] The above data processing method collects the usage data and cost data of data resources for each business through the blood relationship map, and obtains the usage cost of data resources for each business through quantitative calculation of the usage data and cost data, so that users can understand the usage cost corresponding to each business in detail and optimize the resources of each business based on the usage cost.

[0101] In an optional embodiment, the above-mentioned step of establishing a blood relationship map based on the business scenarios contained in the target data system and the data resources associated with the business scenarios includes: determining the data resources applied for by each business scenario contained in the target data system, and the data processing units used for the operation of each data resource; for each business scenario, establishing a correspondence between the data resources corresponding to the business scenario and the data processing units used for the operation of the data resources to obtain the blood relationship corresponding to the business scenario; integrating the blood relationships corresponding to each business scenario to obtain a blood relationship map.

[0102] In an optional embodiment, the above-mentioned blood relationship map includes a business layer, a logical layer and a physical layer; wherein the business layer is used to store the various business scenarios contained in the target data system, the logical layer is used to store the data resources associated with each business scenario, and the physical layer is used to store the data processing units used for the operation of each data resource.

[0103] In an optional embodiment, the above-mentioned step of collecting usage data and cost data of data resources for each business scenario based on the blood relationship map includes: determining the data resources corresponding to each business scenario based on the blood relationship map, and collecting the cost data of the data resources corresponding to each business scenario; collecting the operation data corresponding to the data processing unit used when the data resources corresponding to the business scenario are running, and determining the usage data of the data resources corresponding to the business scenario based on the operation data; determining the usage data of the data resources corresponding to each business scenario based on the blood relationship map and the usage data of the data resources corresponding to the business scenario.

[0104] In an optional embodiment, the above-mentioned step of quantifying the usage data and cost data to obtain the usage cost corresponding to each business scenario includes: for each data resource, based on the data quantification logic corresponding to the current data resource and the usage data and cost data of the current data resource for each business scenario, determining the usage cost of the current data resource for each business scenario; wherein different data resources are respectively configured with corresponding data quantification logic; and integrating the usage cost of each data resource for each business scenario to obtain the usage cost corresponding to each business scenario.

[0105] In an optional embodiment, after the step of quantifying the usage data and cost data to obtain the usage costs corresponding to each business scenario, the above method also includes: performing a health diagnosis on the data resources based on the usage costs corresponding to each business scenario to obtain the health diagnosis results of the data resources.

[0106] In an optional embodiment, the above method further includes: elastically scaling and / or data cleaning of the data resources based on usage data corresponding to the data resources and health diagnosis results of the data resources.

[0107] In an optional embodiment, the above-mentioned steps of elastically scaling and / or data cleaning of data resources based on the usage data corresponding to the data resources and the health diagnosis results of the data resources include: inputting the usage data corresponding to the data resources and the health diagnosis results of the data resources into a pre-prepared three-dimensional evaluation model, and outputting the identification results for the data resources through the three-dimensional evaluation model; wherein the three-dimensional evaluation model is used to perform quantitative analysis of resource activity, citation value and business value of the data resources; and elastically scaling and / or data cleaning of the data resources based on the identification results.

[0108] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a terminal device, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0109] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate the description of this disclosure and simplify the description. They do not indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0110] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure shall be subject to the scope of protection of the claims.

Claims

1. A data processing method, characterized in that: The method comprises: Establish a blood relationship map based on the business scenarios included in the target data system and the data resources associated with the business scenarios; wherein the blood relationship map is used to indicate: the corresponding relationship between each business scenario included in the target data system and the data resources; Based on the blood relationship graph, collect usage data and cost data of the data resources for each business scenario; The usage data and the cost data are quantified to obtain usage costs corresponding to each of the business scenarios.

2. The method according to claim 1, characterized in that The step of establishing a blood relationship map based on the business scenarios included in the target data system and the data resources associated with the business scenarios includes: Determine the data resources required for each business scenario in the target data system, as well as the data processing units used to operate each data resource; For each of the business scenarios, a correspondence is established between the data resources corresponding to the business scenario and the data processing units used to run the data resources, thereby obtaining a blood relationship corresponding to the business scenario; The blood relationship corresponding to each of the business scenarios is integrated to obtain the blood relationship map.

3. The method according to claim 2, characterized in that The blood relationship map includes a business layer, a logical layer and a physical layer; wherein the business layer is used to store the various business scenarios contained in the target data system, the logical layer is used to store the data resources associated with each business scenario, and the physical layer is used to store the data processing units used for the operation of each data resource.

4. The method according to claim 1, wherein The step of collecting usage data and cost data of the data resources for each business scenario based on the blood relationship map includes: Based on the blood relationship map, determine the data resources corresponding to each business scenario, and collect cost data of the data resources corresponding to each business scenario; Collecting operation data corresponding to a data processing unit used when the data resource corresponding to the business scenario is running, and determining usage data of the data resource corresponding to the business scenario based on the operation data; Based on the blood relationship map and the usage data of the data resources corresponding to the business scenarios, the usage data of the data resources corresponding to each business scenario is determined.

5. The method according to claim 1, wherein The step of quantifying the usage data and the cost data to obtain the usage cost corresponding to each business scenario includes: For each data resource, based on the data quantification logic corresponding to the current data resource and the usage data and cost data of the current data resource for each business scenario, determine the usage cost of the current data resource for each business scenario; wherein different data resources are respectively configured with corresponding data quantification logic; The usage costs of the data resources in each of the business scenarios are integrated to obtain the usage costs corresponding to the business scenarios.

6. The method according to claim 1, wherein After the step of quantifying the usage data and the cost data to obtain the usage cost corresponding to each business scenario, the method further includes: Based on the usage costs corresponding to each of the business scenarios, a health diagnosis is performed on the data resources to obtain a health diagnosis result of the data resources.

7. The method according to claim 6, characterized in that The method further comprises: Based on the usage data corresponding to the data resource and the health diagnosis result of the data resource, the data resource is elastically scaled and / or data cleaned.

8. The method according to claim 7, characterized in that The step of elastically scaling and / or cleaning the data resource based on the usage data corresponding to the data resource and the health diagnosis result of the data resource includes: Inputting usage data corresponding to the data resource and health diagnosis results of the data resource into a pre-defined three-dimensional assessment model, and outputting identification results for the data resource through the three-dimensional assessment model; wherein the three-dimensional assessment model is used to quantitatively analyze the resource activity, citation value, and business value of the data resource; The data resources are elastically scaled and / or data cleaned based on the recognition results.

9. A data processing device, characterized in that: The device comprises: A relationship establishment module is used to establish a blood relationship map based on the business scenarios included in the target data system and the data resources associated with the business scenarios; wherein the blood relationship map is used to indicate: the corresponding relationship between each business scenario included in the target data system and the data resources; A data collection module, configured to collect usage data and cost data of the data resources for each business scenario based on the blood relationship map; The data quantification module is used to quantify the usage data and the cost data to obtain the usage cost corresponding to each business scenario.

10. An electronic device, characterized in that: The electronic device includes a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the data processing method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the data processing method according to any one of claims 1 to 8.