Data resource processing method, equipment and medium
By acquiring the priority values of physical, risk, and scenario dimensions of data resources, the classification labels of data resources are automatically determined, which solves the problem of resource waste and underutilization caused by manual labeling errors, and realizes the rational classification and utilization of data resources.
Patent Information
- Application Number
- CN202511517062.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-23
AI Technical Summary
In existing technologies, the initial labeling of data resources based on human experience may be incorrect, leading to the problem that high-value data resources are not fully utilized or low-value data resources are wasted.
By acquiring the priority values of the physical, risk, and scenario dimensions of the target data resources, and combining multiple dimensional features, the classification labels are automatically and objectively determined, and warning information is issued when the labels are inconsistent.
It enables automatic and objective acquisition of classification labels for data resources, avoiding resource waste and underutilization caused by incorrect initial labels.
Smart Images

Figure CN120974247A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric digital data processing, in particular to a data resource processing method, device and medium. BACKGROUND
[0002] In traditional data asset management, classification of data resources usually relies on personal experience of data administrators or business experts, for example, manually labeling data resources as high value and low value according to personal experience. Such manual label management method has the following problems: lack of verification mechanism for initial labels, if the initial label is set incorrectly, for example, the initial label of high-value data resources is mistakenly set as low value, then the high-value data resources will not be deeply or comprehensively utilized; or the initial label of low-value data resources is set as high value, then the low-value data resources may be allocated more resources (such as exclusive storage quota and high-power servers, etc.), resulting in resource waste. SUMMARY
[0003] The present application aims to provide a data resource processing method, device and medium to solve the problem that the initial label manually labeled on data resources according to personal experience in the prior art may be incorrect.
[0004] According to a first aspect of the present application, a data resource processing method is provided, the method comprising the following steps: obtaining a physical dimension priority value of the target data resource according to a preset physical characteristic parameter of the target data resource; the preset physical characteristic parameter comprises at least one of the following parameters: data storage size, number of records.
[0005] obtaining a risk dimension priority value of the target data resource according to a preset risk characteristic parameter of the target data resource; the preset risk characteristic parameter comprises at least one of the following parameters: data encryption state, desensitization level.
[0006] obtaining a scene dimension priority value of the target data resource according to a preset scene characteristic parameter of the target data resource; the preset scene characteristic parameter of the target data resource is obtained according to a scene type of the target data resource.
[0007] determining a measurement priority value of the target data resource according to the physical dimension priority value, the risk dimension priority value and the scene dimension priority value of the target data resource.
[0008] classifying the target data resource according to the measurement priority value of the target data resource and a preset measurement threshold value, obtaining a classification label of the target data resource, and issuing a preset warning information if the classification label of the target data resource is inconsistent with an initial label of the target data resource.
[0009] Furthermore, the process of obtaining the preset scene feature parameters of the target data resources includes: Obtain the field names and field types of the target data resource; the field types include quantitative numerical types and non-quantitative numerical types.
[0010] Construct field features of the target data resource; the field features of the target data resource include the name of each field included in the target data resource and the preset features corresponding to each field; wherein, when the type of a field is a quantitative numerical type, the preset features corresponding to the field include at least one of the following features: mean, variance, and value range; when the type of a field is a non-quantitative numerical type, the preset features corresponding to the field include at least one of the following features: number of unique values, percentage of unique values, and non-empty rate.
[0011] The scenario type of the target data resource is obtained based on the field characteristics of the target data resource and the trained scenario reasoning model.
[0012] If the first list includes mapping relationships corresponding to scene types of the target data resource, then the scene feature parameters in the first list that have mapping relationships with the scene types of the target data resource are determined as preset scene feature parameters of the target data resource; the first list includes several mapping relationships between scene types and scene feature parameters.
[0013] Furthermore, obtaining the scene dimension priority value of the target data resource based on the preset scene feature parameters of the target data resource includes: The judgment threshold corresponding to the preset scene feature parameters of the scene type for acquiring target data resources.
[0014] The comparison result between the parameter values corresponding to the preset scene feature parameters of the target data resource and the judgment threshold is obtained.
[0015] The scene dimension priority value of the target data resource is determined based on the comparison results and the second list; the second list includes the mapping relationship between the comparison results and the scene dimension priority values corresponding to the preset scene feature parameters of the scene type of the target data resource.
[0016] Furthermore, the measurement priority value of the target data resource is determined based on its physical dimension priority value, risk dimension priority value, and scenario dimension priority value, including: Obtain several baseline feature vectors corresponding to a specified scene type; the specified scene type is the scene type of the target data resource.
[0017] Obtain the similarity between the feature vector of the target data resource and each benchmark feature vector corresponding to the specified scenario type; the feature vector of the target data resource includes the parameter value corresponding to at least one of the following parameters: the preset physical feature parameter, the preset risk feature parameter, and the preset scenario feature parameter of the target data resource.
[0018] The weight combination corresponding to the benchmark feature vector with the maximum similarity is determined as the target weight combination; any weight combination includes the weight corresponding to the physical dimension, the weight corresponding to the risk dimension, and the weight corresponding to the scenario dimension.
[0019] The measurement priority value of the target data resource is obtained based on the weights corresponding to the physical dimensions in the target weight combination, the priority value of the physical dimensions of the target data resource, the weights corresponding to the risk dimensions in the target weight combination, the priority value of the risk dimensions, and the weights corresponding to the scenario dimensions in the target weight combination and the priority value of the scenario dimensions.
[0020] Furthermore, the process of obtaining the baseline feature vector corresponding to the specified scene type includes: Obtain sample data resources corresponding to a specified scenario type.
[0021] Cluster the sample data resources corresponding to the specified scene type based on the feature vector of each sample data resource corresponding to the specified scene type to obtain several sample data resource clusters.
[0022] For any sample data resource cluster, the central feature vector of the feature vectors corresponding to the sample data resources included in the sample data resource cluster is determined as the reference feature vector corresponding to the sample data resource cluster.
[0023] Furthermore, the process of obtaining the weight combination corresponding to the benchmark feature vector of the sample data resource cluster includes: fitting the sample data resources and corresponding measurement priority values included in the sample data resource cluster, and determining the weight combination consisting of the weights corresponding to the physical dimension, the risk dimension, and the scenario dimension obtained from the fitting as the weight combination corresponding to the benchmark feature vector of the sample data resource cluster.
[0024] Furthermore, the data resource measurement processing method also includes: if the classification label of the target data resource is consistent with the initial label of the target data resource, then no preset warning information is issued.
[0025] Furthermore, classifying the target data resources based on the measurement priority value and the preset measurement threshold includes: obtaining the comparison result of the measurement priority value and the preset measurement threshold of the target data resources, and determining the classification label of the target data resources based on the comparison result and the third list; the third list includes the mapping relationship between the comparison result and the classification label.
[0026] According to a second aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described data resource processing method.
[0027] According to a third aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for processing data resources.
[0028] Compared with the prior art, the present invention has at least the following beneficial effects: This invention determines the measurement priority value of target data resources based on their physical dimension priority value, risk dimension priority value, and scenario dimension priority value. It then classifies the target data resources according to their measurement priority value and a preset measurement threshold. Thus, this invention can automatically and objectively acquire classification labels for target data resources by combining multiple dimensions of their characteristics. Furthermore, if the classification label of a target data resource is inconsistent with its initial label, a preset warning message is issued to avoid problems such as incomplete or wasted data resources due to mismatch between the initial label and the target data resource. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 A flowchart illustrating the data resource processing method provided in Embodiment 1 of the present invention; Figure 2 A flowchart illustrating the process of obtaining preset scene feature parameters of the target data resource provided in Embodiment 1 of the present invention; Figure 3 A flowchart illustrating the steps for obtaining the scene dimension priority value of a target data resource according to Embodiment 1 of the present invention; Figure 4 A flowchart illustrating the steps for determining the metering priority value of a target data resource, as provided in Embodiment 1 of the present invention; Figure 5 This is a flowchart illustrating the process of obtaining the baseline feature vector corresponding to a specified scene type, as provided in Embodiment 1 of the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Example 1: According to this embodiment, as Figure 1 As shown, a method for processing data resources is provided, the method comprising the following steps: S100, obtain the physical dimension priority value of the target data resource according to the preset physical characteristic parameters of the target data resource; the preset physical characteristic parameters include at least one of the following parameters: data storage size, number of records.
[0033] In this embodiment, the target data resource is the data resource of the classification label to be obtained, and the data resource is a collection of data.
[0034] In this embodiment, the preset physical characteristic parameters are predefined parameters that describe the physical attributes of the data resource and reflect its basic characteristics. For example, the preset physical characteristic parameters include data storage size and number of records, where data storage size is the storage space occupied by the data resource, and the number of records is the number of data entries contained in the data resource (for example, when the target data resource is a database table, the number of records refers to the number of data rows it contains).
[0035] In one specific implementation, if the number of preset physical feature parameters is greater than or equal to 2, obtaining the physical dimension priority value of the target data resource based on the preset physical feature parameters includes: obtaining the parameter value corresponding to each preset physical feature parameter of the target data resource; normalizing the parameter value corresponding to each preset physical feature parameter; multiplying each normalization result by the corresponding preset weight; and summing the products to obtain the physical dimension priority value of the target data resource. Optionally, the preset weight corresponding to each preset physical feature parameter is an empirical value. If the number of preset physical feature parameters is 1, the parameter value corresponding to the preset physical feature parameter will be normalized, and the normalization result will be determined as the physical dimension priority value.
[0036] In this embodiment, the larger the volume of the target data resource, the higher the priority value of the corresponding physical dimension, which usually means that the information value contained in the target data resource is higher.
[0037] S200, obtain the risk dimension priority value of the target data resource according to the preset risk characteristic parameters of the target data resource; the preset risk characteristic parameters include at least one of the following parameters: data encryption status, desensitization level.
[0038] In this embodiment, the preset risk characteristic parameters are predefined parameters used to measure the risk of data resources. For example, the preset risk characteristic parameters include data encryption status and desensitization level. Data encryption status refers to whether the data has been encrypted; optionally, data encryption status includes encrypted and unencrypted. Desensitization level refers to the degree of desensitization of sensitive data. For example, sensitive data includes mobile phone numbers and ID card numbers, and desensitization levels include no desensitization, partial desensitization, and complete desensitization. As a specific implementation, different data encryption statuses correspond to different encryption scores, and different desensitization levels correspond to different desensitization scores. For example, when the data encryption status is encrypted, the encryption score is 1; when the data encryption status is unencrypted, the encryption score is 0. When the desensitization level is no desensitization, the desensitization score is 0; when the desensitization level is partially desensitized, the desensitization score is 0.5; and when the desensitization level is completely desensitized, the desensitization score is 1. Multiply the encryption score by its corresponding encryption weight to obtain the encryption priority value; multiply the de-identification score by its corresponding de-identification weight to obtain the de-identification priority value; the sum of the encryption priority value and the de-identification priority value is determined as the risk dimension priority value. Optionally, the de-identification weight and encryption weight are empirical values, for example, both the de-identification weight and the encryption weight are 0.5.
[0039] In this embodiment, the higher the risk of the target data resource, the lower the priority value of the corresponding risk dimension, which usually means that the quality of the target data resource is lower.
[0040] S300, obtain the scene dimension priority value of the target data resource according to the preset scene feature parameters of the target data resource; the preset scene feature parameters of the target data resource are obtained according to the scene type of the target data resource.
[0041] As a specific implementation method, such as Figure 2 As shown, the process of obtaining the preset scene feature parameters of the target data resource includes: S301, Obtain the field names and field types of the target data resource; the field types include quantitative numerical types and non-quantitative numerical types.
[0042] In this embodiment, the field name refers to the identifier of the smallest data unit in the data resource (such as transaction ID, customer ID, transaction amount, transaction time, etc.). The field type refers to the category of the field data, where quantitative value types are types that can be quantified and support mathematical operations, such as transaction amount; non-quantitative value types are types that are not quantitative value types. As a specific implementation method, the fields of the target data resource can be obtained through existing data parsing tools.
[0043] S302, Construct the field features of the target data resource; the field features of the target data resource include the name of each field included in the target data resource and the preset features corresponding to each field; wherein, when the type of a field is a quantitative numerical type, the preset features corresponding to the field include at least one of the following features: mean, variance, and value range; when the type of a field is a non-quantitative numerical type, the preset features corresponding to the field include at least one of the following features: number of unique values, percentage of unique values, and non-empty rate.
[0044] In this embodiment, field features are quantitative descriptions of field attributes. The number of unique values refers to the number of distinct values, the percentage of unique values refers to the ratio of the number of unique values to the total number of records, and the non-null rate is the ratio of the number of non-null values to the total number of records.
[0045] In this embodiment, the field characteristics of the target data resource can reflect the data patterns and business characteristics of the target data resource, and the scenario type of the target data resource can be determined based on the field characteristics of the target data resource.
[0046] S303, obtain the scene type of the target data resource based on the field characteristics of the target data resource and the trained scene reasoning model.
[0047] In this embodiment, the scenario type is the application scenario of the data resource, which reflects the core purpose of the data resource. The scenario inference model is a multi-class machine learning classification model (such as a random forest or neural network model), with field features as input and scenario type as output. The training data consists of sample data labeled with scenario types. As a specific implementation, the training phase uses sample data resources of different scenario types. The input for any sample data resource is the field feature corresponding to that sample data resource, and the output for any sample data resource is the scenario type label labeled with that sample data resource. Thus, by utilizing the powerful pattern recognition capability of the machine learning classification model, the scenario type to which an unknown data resource most likely belongs can be automatically inferred, achieving automated scenario type identification. As a specific implementation, the multi-class machine learning classification model can use BERT or its variants (RoBERTa, DeBERTa), or gradient boosting trees (such as LightGBM, XGBoost) can be selected. If a BERT-like model is chosen, this type of model is pre-trained on massive amounts of text and possesses powerful language semantic understanding capabilities, effectively capturing the semantic information of field names. For the numerical information of fields, preprocessing is required first through normalization (eliminating differences in magnitude) and linear mapping (converting single-dimensional numerical values into numerical embedding vectors of the same dimension as the BERT text embeddings). Then, the text embedding vectors and numerical embedding vectors are concatenated and fused along the feature dimension, enabling the model to capture both semantic and numerical features simultaneously. The structure and training process of multi-class machine learning classification models are existing technologies and will not be elaborated here.
[0048] In this embodiment, the field features of the target data resource are input into a trained scenario inference model. The scenario type with the highest probability in the output of the trained scenario inference model is the scenario type of the target data resource. For example, the field features of the target data resource are: Field 1 is named [customer_age], with an average value of 35.2 and a variance of 64.0; Field 2 is named [customer_region], with 150 unique values and a unique value ratio of 0.75; Field 3 is named [transaction_amount], with an average value of 255.5 and a variance of 100.0. The scenario type with the highest probability in the output of the trained scenario inference model is the customer profile. Here, customer_age represents the customer's age, customer_region represents the customer's region, and transaction_amount represents the transaction amount.
[0049] S304, if the first list includes mapping relationships corresponding to scene types of the target data resource, then the scene feature parameters in the first list that have mapping relationships with the scene types of the target data resource are determined as preset scene feature parameters of the target data resource; the first list includes mapping relationships between several scene types and scene feature parameters.
[0050] In this embodiment, the core requirements differ across scenarios. A first list allows for the targeted extraction of key parameters for each scenario, enabling further evaluation of data resources based on these extracted parameters. The first list is a pre-established mapping table between scenario types and scenario feature parameters, used to record core metrics of interest for different scenario types. For example, when the scenario type is a customer profiling scenario, preset scenario feature parameters include customer attribute completeness, etc.
[0051] In this embodiment, if the first list does not include the mapping relationship corresponding to the scene type of the target data resource, a preset prompt message is output so that the scene feature parameters corresponding to the scene type of the target data resource can be given manually.
[0052] As a specific implementation method, such as Figure 3 As shown, obtaining the scene dimension priority value of the target data resource based on the preset scene feature parameters of the target data resource includes: S310, Obtain the judgment threshold corresponding to the preset scene feature parameters corresponding to the scene type of the target data resource.
[0053] In this embodiment, the judgment threshold is a pre-set benchmark value for measuring scene feature parameters. Optionally, the judgment thresholds corresponding to different preset scene feature parameters for different scene types are empirical values.
[0054] S320, obtain the comparison result between the parameter values corresponding to the preset scene feature parameters of the target data resource and the judgment threshold.
[0055] In this embodiment, the comparison result can be obtained by comparing the parameter value corresponding to the preset scene feature parameter of the target data resource with the judgment threshold; optionally, the comparison result includes being higher than the judgment threshold, lower than the judgment threshold, etc.
[0056] S330, determine the scene dimension priority value of the target data resource based on the comparison result and the second list; the second list includes the mapping relationship between the comparison result and the scene dimension priority value corresponding to the preset scene feature parameters of the scene type of the target data resource.
[0057] In this embodiment, the second list is a pre-established mapping table of scene types, scene feature parameters, comparison results, scores, and parameter weights. This table records the scores and weights of different comparison results for different scene feature parameters of different scene types. First, the score corresponding to each scene feature parameter of the target data resource is obtained through the comparison results of each scene feature parameter included in the scene type of the target data resource. Then, by accumulating the scores and corresponding parameter weights of all scene feature parameters included in the scene type of the target data resource, the scene dimension priority value of the target data resource can be obtained.
[0058] Therefore, this embodiment can obtain the scenario dimension priority value of data resources of different scenario types, so that the obtained dimension priority value is related to the scenario type, thereby improving the flexibility and matching of obtaining the scenario dimension priority value.
[0059] In this embodiment, the higher the priority value of the scene dimension of the target data resource, the higher the quality or value of the target data resource in its scene.
[0060] S400 determines the metering priority value of the target data resource based on the physical dimension priority value, risk dimension priority value, and scenario dimension priority value of the target data resource.
[0061] This embodiment assesses data from three core dimensions: physical, risk, and scenario, covering multiple aspects of data value assessment and improving the rationality of the measurement priority value of the acquired target data resources. As a specific implementation method, such as... Figure 4 As shown, S400 includes: S410, obtain several baseline feature vectors corresponding to a specified scene type; the specified scene type is the scene type of the target data resource.
[0062] As a specific implementation method, the process of obtaining the baseline feature vector corresponding to a specified scene type includes, for example:Figure 5 As shown: S411, obtain sample data resources corresponding to the specified scene type.
[0063] In this embodiment, sample data resources corresponding to a specified scenario type are obtained by collecting a batch of historical data resources that have been manually confirmed and whose metering priority values are known.
[0064] S412, cluster the sample data resources corresponding to the specified scene type according to the feature vector of each sample data resource corresponding to the specified scene type to obtain several sample data resource clusters.
[0065] In this embodiment, the feature vector corresponding to any sample data resource of a specified scenario type includes the parameter value of at least one of the following parameters: preset physical feature parameter of the target data resource, preset risk feature parameter, and preset scenario feature parameter.
[0066] In this embodiment, the feature vector corresponding to any sample data resource of a specified scene type and the feature vector of the target data resource are composed of parameter values corresponding to the same parameters, and the parameter values corresponding to the same position in the feature vector corresponding to any sample data resource of a specified scene type and the feature vector of the target data resource represent the same type of parameter.
[0067] It should be understood that sample data resources within the same cluster obtained by clustering are sample data resources with relatively similar features, while sample data resources within different clusters are sample data resources with significantly different features. Optionally, the clustering uses the k-means clustering method, where k is an empirical value or determined by the elbow rule; an N-dimensional space is constructed, where N is the dimension of the feature vectors. Thus, the feature vector corresponding to any sample data corresponds to a point in the space, and the distance between the two points corresponding to the feature vectors of any two sample data is determined as the distance between the feature vectors corresponding to the two sample data. Those skilled in the art will know that the process of k-means clustering is prior art and will not be described in detail here.
[0068] S413, For any sample data resource cluster, the central feature vector of the feature vectors corresponding to the sample data resources included in the sample data resource cluster is determined as the reference feature vector corresponding to the sample data resource cluster.
[0069] It should be understood that the central eigenvector of the eigenvectors corresponding to the eigenvectors ...
[0070] As a specific implementation, for any sample data resource cluster, the process of obtaining the weight combination corresponding to the benchmark feature vector of the sample data resource cluster includes: fitting the sample data resources and corresponding measurement priority values included in the sample data resource cluster, and determining the weight combination consisting of the weights corresponding to the physical dimension, the risk dimension, and the scenario dimension obtained by fitting as the weight combination corresponding to the benchmark feature vector of the sample data resource cluster.
[0071] As a specific implementation method, a linear regression model is adopted, with the measurement priority value being Y, Y = Wp × P + Wc × C + Ws × S, where Wp, Wc, and Ws are the weights corresponding to the physical dimension, risk dimension, and scenario dimension, respectively, and P, C, and S are the priority values for the physical dimension, risk dimension, and scenario dimension, respectively. The priority values for the physical dimension, risk dimension, and scenario dimension, along with the labeled measurement priority values, corresponding to the sample data resources included in any sample data resource cluster, are input into Y = Wp × P + Wc × C + Ws × S. The least squares method is used to solve for the weights corresponding to the physical dimension, risk dimension, and scenario dimension, and the resulting weight combination is determined as the weight combination corresponding to that sample data resource cluster.
[0072] In this embodiment, considering that data resources may have different focuses under the same scenario type, that is, data resources under the same scenario are applicable to different measurement priority value subtypes; this embodiment finds these subtypes through clustering and learns an optimal weight combination for each subtype, so that the obtained measurement priority value can be more accurate.
[0073] S420, obtain the similarity between the feature vector of the target data resource and each benchmark feature vector corresponding to the specified scene type; the feature vector of the target data resource includes the parameter value corresponding to at least one of the following parameters: the preset physical feature parameter of the target data resource, the preset risk feature parameter, and the preset scene feature parameter.
[0074] In this embodiment, the feature vector of the target data resource has the same dimension as each baseline feature vector corresponding to the specified scene type; optionally, the similarity is cosine similarity.
[0075] S430, the weight combination corresponding to the benchmark feature vector corresponding to the maximum similarity is determined as the target weight combination; any weight combination includes the weight corresponding to the physical dimension, the weight corresponding to the risk dimension, and the weight corresponding to the scene dimension.
[0076] In this embodiment, the weight combination corresponding to the benchmark feature vector with the maximum similarity is determined as the target weight combination that can best match the target data resource.
[0077] S440, obtain the measurement priority value of the target data resource based on the weight corresponding to the physical dimension in the target weight combination, the priority value of the physical dimension of the target data resource, the weight corresponding to the risk dimension in the target weight combination, the priority value of the risk dimension, the weight corresponding to the scenario dimension in the target weight combination, and the priority value of the scenario dimension.
[0078] In this embodiment, a weighted summation formula is used to obtain the metering priority value of the target data resource, which will not be elaborated here.
[0079] Therefore, this embodiment determines which sample cluster the target data resource belongs to by comparing similarity, and then assigns the most matching weight combination to the target data resource to ensure that the weight combination is adapted to the characteristics of the target data resource, thereby improving the accuracy of the final obtained measurement priority value.
[0080] S500: Classify the target data resources according to the metering priority value and the preset metering threshold, and obtain the classification label of the target data resources. If the classification label of the target data resources is inconsistent with the initial label of the target data resources, a preset warning message is issued.
[0081] In this embodiment, the measurement threshold is the critical value for classifying data resources; the initial label of the target data resource is a label that is pre-marked on the target data resource before processing it.
[0082] In this embodiment, the higher the metering priority value of the target data resource, the higher the usable value of the target data resource.
[0083] In this embodiment, classifying target data resources based on their measurement priority value and a preset measurement threshold includes: obtaining a comparison result between the target data resource's measurement priority value and the preset measurement threshold; and determining the classification label of the target data resource based on the comparison result and a third list. The third list includes a mapping relationship between the comparison result and the classification label. As a specific implementation, the classification labels include high-value labels, medium-value labels, and low-value labels. The preset measurement thresholds include a first measurement threshold and a second measurement threshold. If the measurement priority value of the target data resource is greater than the first measurement threshold, the target data resource is classified as a high-value label; if the measurement priority value of the target data resource is less than the second measurement threshold, the target data resource is classified as a low-value label; otherwise, the target data resource is classified as a medium-value label. The first measurement threshold is greater than the second measurement threshold.
[0084] In this embodiment, the data resource measurement processing method further includes: if the classification label of the target data resource is consistent with the initial label of the target data resource, then no preset warning information is issued.
[0085] This embodiment can detect errors in the initial label by comparing the classification label with the initial label, triggering an alert in time and avoiding confusion in processing priorities due to label errors.
[0086] This embodiment determines the measurement priority value of the target data resource based on its physical dimension priority value, risk dimension priority value, and scenario dimension priority value. The target data resource is then classified according to its measurement priority value and a preset measurement threshold. Thus, this embodiment can automatically and objectively acquire the classification labels of the target data resource by combining its multiple dimensions of characteristics. Furthermore, if the classification label of the target data resource is inconsistent with its initial label, a preset warning message is issued to avoid problems such as incomplete and unutilized data resources or resource waste caused by the mismatch between the initial label and the target data resource.
[0087] Example 2: This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the following steps: The physical dimension priority value of the target data resource is obtained based on the preset physical characteristic parameters of the target data resource; the preset physical characteristic parameters include at least one of the following parameters: data storage size, number of records.
[0088] The risk dimension priority value of the target data resource is obtained based on the preset risk characteristic parameters of the target data resource; the preset risk characteristic parameters include at least one of the following parameters: data encryption status and desensitization level.
[0089] The scene dimension priority value of the target data resource is obtained based on the preset scene feature parameters of the target data resource; the preset scene feature parameters of the target data resource are obtained based on the scene type of the target data resource.
[0090] The measurement priority of the target data resources is determined based on the physical dimension priority, risk dimension priority, and scenario dimension priority of the target data resources.
[0091] The target data resources are classified according to their measurement priority value and preset measurement threshold, and the classification labels of the target data resources are obtained. If the classification labels of the target data resources are inconsistent with the initial labels of the target data resources, a preset warning message is issued.
[0092] Example 3: This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it performs the following steps: The physical dimension priority value of the target data resource is obtained based on the preset physical characteristic parameters of the target data resource; the preset physical characteristic parameters include at least one of the following parameters: data storage size, number of records.
[0093] The risk dimension priority value of the target data resource is obtained based on the preset risk characteristic parameters of the target data resource; the preset risk characteristic parameters include at least one of the following parameters: data encryption status and desensitization level.
[0094] The scene dimension priority value of the target data resource is obtained based on the preset scene feature parameters of the target data resource; the preset scene feature parameters of the target data resource are obtained based on the scene type of the target data resource.
[0095] The measurement priority of the target data resources is determined based on the physical dimension priority, risk dimension priority, and scenario dimension priority of the target data resources.
[0096] The target data resources are classified according to their measurement priority value and preset measurement threshold, and the classification labels of the target data resources are obtained. If the classification labels of the target data resources are inconsistent with the initial labels of the target data resources, a preset warning message is issued.
[0097] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0098] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of the invention. It should also be understood that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.
Claims
1. A method for processing data resources, characterized in that, The method includes the following steps: The physical dimension priority value of the target data resource is obtained based on the preset physical characteristic parameters of the target data resource; the preset physical characteristic parameters include at least one of the following parameters: data storage size, number of records; The risk dimension priority value of the target data resource is obtained based on the preset risk characteristic parameters of the target data resource; the preset risk characteristic parameters include at least one of the following parameters: data encryption status, de-identification level; The scene dimension priority value of the target data resource is obtained based on the preset scene feature parameters of the target data resource; the preset scene feature parameters of the target data resource are obtained according to the scene type of the target data resource. The measurement priority value of the target data resource is determined based on the physical dimension priority value, risk dimension priority value, and scenario dimension priority value of the target data resource; The target data resources are classified according to their measurement priority value and preset measurement threshold, and the classification labels of the target data resources are obtained. If the classification labels of the target data resources are inconsistent with the initial labels of the target data resources, a preset warning message is issued.
2. The data resource processing method according to claim 1, characterized in that, The process of obtaining the preset scene feature parameters of the target data resources includes: Obtain the field names and field types of the target data resource; the field types include quantitative and non-quantitative numerical types. Construct field features of the target data resource; the field features of the target data resource include the name of each field included in the target data resource and the preset features corresponding to each field; wherein, when the type of a field is a quantitative numerical type, the preset features corresponding to the field include at least one of the following features: mean, variance, and value range; when the type of a field is a non-quantitative numerical type, the preset features corresponding to the field include at least one of the following features: number of unique values, percentage of unique values, and non-empty rate; The scenario type of the target data resource is obtained based on the field characteristics of the target data resource and the trained scenario reasoning model; If the first list includes mapping relationships corresponding to scene types of the target data resource, then the scene feature parameters in the first list that have mapping relationships with the scene types of the target data resource are determined as preset scene feature parameters of the target data resource; the first list includes several mapping relationships between scene types and scene feature parameters.
3. The data resource processing method according to claim 1, characterized in that, The priority values for obtaining the scene dimension of the target data resource based on the preset scene feature parameters include: The judgment threshold corresponding to the preset scene feature parameters corresponding to the scene type of the target data resource; The comparison result between the parameter values corresponding to the preset scene feature parameters of the target data resource and the judgment threshold is obtained; The scene dimension priority value of the target data resource is determined based on the comparison results and the second list; the second list includes the mapping relationship between the comparison results and the scene dimension priority values corresponding to the preset scene feature parameters of the scene type of the target data resource.
4. The data resource processing method according to claim 1, characterized in that, The measurement priority of target data resources is determined based on their physical dimension priority, risk dimension priority, and scenario dimension priority, including: Obtain several baseline feature vectors corresponding to a specified scene type; the specified scene type is the scene type of the target data resource; Obtain the similarity between the feature vector of the target data resource and each benchmark feature vector corresponding to the specified scene type; the feature vector of the target data resource includes the parameter value corresponding to at least one of the following parameters: the preset physical feature parameter, the preset risk feature parameter, and the preset scene feature parameter of the target data resource; The weight combination corresponding to the benchmark feature vector with the maximum similarity is determined as the target weight combination; any weight combination includes the weights corresponding to the physical dimension, the weights corresponding to the risk dimension, and the weights corresponding to the scenario dimension. The measurement priority value of the target data resource is obtained based on the weights corresponding to the physical dimensions in the target weight combination, the priority value of the physical dimensions of the target data resource, the weights corresponding to the risk dimensions in the target weight combination, the priority value of the risk dimensions, and the weights corresponding to the scenario dimensions in the target weight combination and the priority value of the scenario dimensions.
5. The data resource processing method according to claim 4, characterized in that, The process of obtaining the baseline feature vector corresponding to a specified scene type includes: Obtain sample data resources corresponding to a specified scenario type; Based on the feature vector corresponding to each sample data resource of the specified scenario type, the sample data resources of the specified scenario type are clustered to obtain several sample data resource clusters. For any sample data resource cluster, the central feature vector of the feature vectors corresponding to the sample data resources included in the sample data resource cluster is determined as the reference feature vector corresponding to the sample data resource cluster.
6. The data resource processing method according to claim 5, characterized in that, For any sample data resource cluster, the process of obtaining the weight combination corresponding to the benchmark feature vector of the sample data resource cluster includes: fitting the sample data resources and corresponding measurement priority values included in the sample data resource cluster, and determining the weight combination consisting of the weights corresponding to the physical dimension, the risk dimension, and the scenario dimension obtained by fitting as the weight combination corresponding to the benchmark feature vector of the sample data resource cluster.
7. The data resource processing method according to claim 1, characterized in that, The data resource measurement processing method further includes: if the classification label of the target data resource is consistent with the initial label of the target data resource, then no preset warning information is issued.
8. The data resource processing method according to claim 1, characterized in that, Classifying target data resources based on their measurement priority value and a preset measurement threshold includes: obtaining a comparison result between the measurement priority value and the preset measurement threshold of the target data resources; and determining the classification label of the target data resources based on the comparison result and a third list; the third list includes the mapping relationship between the comparison result and the classification label.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data resource processing method as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the data resource processing method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Customer information storage management method and system based on block chain
CN117055818A
Data classification and grading processing method and system based on machine learning
CN117216668A
Sea area development suitability evaluation method and system
CN119359157A
Small sample constraint-oriented network data asset security classification method and system
CN120336963A
Human resource archive filing management system
CN120743852A