An enterprise data management method based on data analysis

By constructing an enterprise data management methodology, collecting and processing multi-dimensional data, extracting core features, establishing a data value-risk linkage assessment model, and dynamically updating parameters, the problems of low data quality, one-sided value assessment, and lack of risk warning have been solved. This enables accurate identification of high-value data and risk warning, adapts to enterprise changes, and improves the accuracy of data management and decision support.

CN122134128APending Publication Date: 2026-06-02HUAXU ZHENGXIN INTELLECTUAL PROPERTY SERVICE CANGZHOU CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAXU ZHENGXIN INTELLECTUAL PROPERTY SERVICE CANGZHOU CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing enterprise data management methods suffer from problems such as low data quality, one-sided value assessment, lack of risk warning, and poor model adaptability, leading to decreased data management accuracy and decision-making bias.

Method used

By collecting multi-dimensional data from both inside and outside the enterprise, and after preprocessing, extracting business relevance, data value density, and risk warning characteristics, a data value-risk linkage assessment model is constructed. The model parameters are dynamically updated in conjunction with user feedback to achieve high-value data identification and risk warning.

Benefits of technology

It improves data accuracy and consistency, accurately identifies high-value data, adapts to changes in enterprise business, avoids decision-making biases, and provides secure data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134128A_ABST
    Figure CN122134128A_ABST
Patent Text Reader

Abstract

This invention relates to a data analysis-based enterprise data management method, comprising the following steps: collecting multi-dimensional data from both internal and external sources of the enterprise and preprocessing the data; extracting three core features from the preprocessed effective data; constructing a data value-risk linkage assessment model based on the core features, calculating the comprehensive value score and risk level of the enterprise data; sorting the data in descending order based on the comprehensive value score, outputting high-value data for enterprise decision support, and triggering corresponding early warning mechanisms according to the risk level; and dynamically updating the parameters of the data value-risk linkage assessment model based on feedback data from enterprise users regarding the effectiveness of data application, achieving adaptive optimization of the model. In this invention, by integrating three types of features—business relevance, value density, and risk warning—and introducing a time decay factor, a multi-dimensional value assessment model can be constructed to accurately identify high-value data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise data processing and management technology, specifically to an enterprise data management method based on data analysis. Background Technology

[0002] With the deepening of digital transformation, enterprise data is characterized by multi-source nature, heterogeneity, and massive growth, encompassing internal business data, financial data, operational data, and external industry data, policy data, and market data. This data is the core basis for enterprise strategic decision-making, business optimization, and risk control; however, current enterprise data management methods have the following shortcomings:

[0003] Data collection lacks a systematic approach, failing to effectively integrate internal and external multi-dimensional data. Furthermore, invalid data (such as null values ​​and duplicate records) and noisy data (such as abnormal fluctuations and non-business-related data) are mixed in, resulting in low data quality. Data value assessment relies on a single dimension, depending solely on business relevance, ignoring key factors such as data timeliness and risk impact, making it impossible to accurately identify high-value data. The lack of a dynamic adaptation mechanism and fixed model parameters make it difficult to cope with changes in business operations and data distribution migration, leading to a decline in management accuracy over time. Finally, the absence of a data risk and value linkage assessment system prevents simultaneous early warning of potential risks (such as decision-making biases caused by data anomalies) during data management. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] This invention provides a data analysis-based enterprise data management method that solves the problems of low data quality, one-sided value assessment, lack of risk warning, and poor model adaptability in existing enterprise data management.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the present invention provides the following technical solution: an enterprise data management method based on data analysis, comprising the following steps:

[0008] Step 1: Collect multi-dimensional data from both inside and outside the enterprise, and preprocess the data;

[0009] Step 2: Extract three types of core features from the preprocessed effective data. The core features include business relevance features, data value density features, and risk warning features.

[0010] Step 3: Based on the core features, construct a data value-risk linkage assessment model to calculate the comprehensive value score and risk level of enterprise data;

[0011] Step 4: Sort the data in descending order based on the comprehensive value score, output high-value data for enterprise decision support, and trigger corresponding early warning mechanisms according to the risk level;

[0012] Step 5: Based on feedback data from enterprise users regarding the effectiveness of data application, dynamically update the parameters of the data value-risk linkage assessment model to achieve adaptive optimization of the model.

[0013] Preferably, the multi-dimensional data in step 1 includes internal data, external data, and auxiliary data. The internal data includes business data, financial data, and operational data. The external data includes industry data, policy data, and market data. The auxiliary data includes enterprise size tags, industry type tags, and data collection timestamps. The data preprocessing in step 1 includes invalid data removal, noise data filtering, and data standardization. The invalid data removal includes null value removal, duplicate record deduplication, and non-business data filtering. The noise data filtering includes rule-based filtering and statistical filtering. The data standardization includes text encoding unification, numerical normalization, and unstructured text structuring transformation.

[0014] In a further preferred embodiment, the business relevance feature described in step 2 is preferred. The calculation formula is as follows:

[0015] ;

[0016] In the formula, The number of keywords matching the data text with the industry's core business thesaurus. This represents the total number of keywords in the data text.

[0017] In a further preferred embodiment, the data value density feature described in step 2 is... The calculation formula is as follows:

[0018] ;

[0019] In the formula, , as well as For data integrity Data timeliness and data scarcity The feature weights, and .

[0020] In a further preferred embodiment, the risk warning feature described in step 2 is... Risk levels are classified and quantified based on the data anomaly volatility coefficient D. The calculation formula for the data anomaly volatility coefficient D is as follows:

[0021] ;

[0022] In the formula, For the target data value, This is the average of similar data over the past three months.

[0023] In a further preferred embodiment, the data value-risk linkage assessment model in step 3 includes a comprehensive value scoring model, the calculation formula of which is as follows:

[0024] ;

[0025] In the formula, To score the overall value, For the first Weight parameters of class features, For feature quantization values, The time decay factor, For bias terms;

[0026] The time decay factor The calculation formula is as follows:

[0027] ;

[0028] In the formula, For the current time, For data collection time, This is the attenuation coefficient.

[0029] In a further preferred embodiment, the data value-risk linkage assessment model also includes a risk level calculation model, the calculation formula of which is as follows:

[0030] ;

[0031] In the formula, Risk level, and This is the risk threshold.

[0032] In a further preferred embodiment, the model parameter update in step 5 employs gradient descent, and the update objects include feature weights. Bias terms The update formulas for the feature weights and bias terms are as follows:

[0033] ;

[0034] ;

[0035] In the formula, For the number of iterations, For learning rate, For loss function, , This is the partial derivative of the loss function.

[0036] (III) Beneficial Effects

[0037] Compared with existing technologies, this invention provides an enterprise data management method based on data analysis, which has the following beneficial effects:

[0038] In this invention, multi-source data integration covers data across the entire enterprise scenario, and through dual screening including invalid data removal and noise data filtering, combined with standardized processing, the accuracy and consistency of the data can be effectively improved.

[0039] In this invention, by integrating three types of features—business relevance, value density, and risk warning—and introducing a time decay factor, a multi-dimensional value assessment model can be constructed to accurately identify high-value data.

[0040] In this invention, by establishing a linkage assessment mechanism between data value and risk, potential risks of high-value data can be warned simultaneously during the data management process, thus avoiding decision-making bias.

[0041] In this invention, a dynamic update mechanism based on user feedback is used to continuously optimize model parameters and risk thresholds, which can adapt to changes in enterprise business and data distribution migration, ensuring long-term management accuracy.

[0042] In this invention, by ranking the comprehensive value scores and outputting high-value data, and combining this with risk level-differentiated early warnings, accurate and secure data support can be provided for enterprise decision-making. Attached Figure Description

[0043] Figure 1 This is a flowchart illustrating a data analytics-based enterprise data management approach implemented according to the scheme. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] Please see Figure 1 A data analysis-based enterprise data management method includes the following steps:

[0046] Step 1: Collect multi-dimensional data from both inside and outside the enterprise, and preprocess the data;

[0047] Step 2: Extract three types of core features from the preprocessed effective data. The core features include business relevance features, data value density features, and risk warning features.

[0048] Step 3: Construct a data value-risk linkage assessment model based on core features to calculate the comprehensive value score and risk level of enterprise data;

[0049] Step 4: Sort the data in descending order based on the comprehensive value score, output high-value data for enterprise decision support, and trigger corresponding early warning mechanisms according to the risk level;

[0050] Step 5: Based on feedback data from enterprise users regarding the effectiveness of data application, dynamically update the parameters of the data value-risk linkage assessment model to achieve adaptive optimization of the model.

[0051] In this embodiment, internal data includes business data (order volume, production capacity data, customer retention rate), financial data (revenue, cost, profit margin), and operational data (process response time, equipment failure rate); external data includes industry data (market share, industry growth rate), policy data (tax policies, regulatory requirements), and market data (competitor prices, consumer preferences); auxiliary data includes enterprise size tags, industry type tags, and data collection timestamps.

[0052] In this embodiment, data preprocessing includes invalid data removal, noisy data filtering, and data standardization. Invalid data removal includes removing null values, completely duplicate records (based on a combination of "data type + collection time + core fields" for deduplication), and non-business-related data (such as external crawler data unrelated to the company's main business). Noisy data filtering includes using rule-based filtering (setting a reasonable threshold range for numerical data and removing outliers exceeding the range) and statistical filtering (based on a normal distribution and removing extreme values ​​exceeding 99.7% of the data distribution range). Data standardization can: uniformly encode text data into UTF-8 format, removing garbled characters and meaningless special symbols; map numerical data to the [0,1] interval through Min-Max normalization; and use the BERT model to perform word segmentation, entity recognition, and intent classification on unstructured text data (such as policy documents and customer feedback) to transform it into structured features.

[0053] In this embodiment, the extraction process of the three types of core features in step 2 is as follows:

[0054] For business relevance characteristics First, a core business thesaurus (including keywords, synonyms, and near-synonyms) is constructed for the company's industry. Then, the target data text is matched against the core business thesaurus, and the ratio of the number of successfully matched keywords to the total number of keywords in the data text is calculated using the following formula:

[0055] ;

[0056] In the formula, The number of keywords matching the data text with the industry's core business thesaurus. This represents the total number of keywords in the data text. The value range is [0,1]. The closer the value is to 1, the stronger the correlation between the data and the enterprise's business.

[0057] For data value density characteristics First, define the set of value influencing factors, including data integrity ( Values ​​range from 0 to 1, calculated in reverse based on the field's missing rate; data timeliness ( (Values ​​range from 0 to 1) Data scarcity ( (Values ​​range from 0 to 1, quantified based on industry data availability); then, the value density feature is calculated through weighted summation, as shown in the following formula: ;

[0058] In the formula, , as well as For data integrity Data timeliness and data scarcity The feature weights, and The initial values ​​for the three are set based on industry experience.

[0059] For risk warning characteristics First, calculate the data anomaly fluctuation coefficient D, and then quantify it based on the deviation rate between the target data and the average of similar data over the past three months. The formula for calculating the data anomaly fluctuation coefficient D is as follows:

[0060] ;

[0061] In the formula, For the target data value, This is the average of similar data over the past three months.

[0062] In this embodiment, the data value-risk linkage assessment model may include a comprehensive value scoring model and a risk level calculation model.

[0063] The calculation formula for the comprehensive value scoring model is as follows:

[0064] ;

[0065] In the formula, For the first The comprehensive value score of each data point ranges from [0, +∞), with a higher value indicating higher data value. These correspond to business relevance feature F1, data value density feature F2, and risk warning feature F3, respectively (risk features have a negative impact and are adjusted by weighting symbols). For the first Weight parameters of class features, >0、 >0、 <0, and satisfy |w1|+|w2|+|w3|=1, with the initial value determined through industry sample training; For the first The first data item Class feature quantization value; This is a bias term with a value range of [0, 0.2]. It is used to correct the baseline value of the feature weighted sum and avoid missing high-value data with low feature values. The time decay factor is used to mitigate the impact of outdated data, and its calculation formula is as follows:

[0066] ;

[0067] In the formula, For the current time, For data collection time, This is the attenuation coefficient, with a value range of [0.01, 0.1]. (Industry data) You can take 0.03, real-time business data. You can take 0.08.

[0068] The calculation formula for the risk level calculation model is as follows:

[0069] ;

[0070] In the formula, For the first The risk level of each data point and This is a risk threshold, which can be adjusted according to the company's risk tolerance. This is a value-risk linkage factor used to reflect the risk amplification effect of high-value data.

[0071] In this embodiment, the model parameters are updated using gradient descent, and the updated objects include feature weights. Bias terms The update formulas for feature weights and bias terms are as follows:

[0072] ;

[0073] ;

[0074] In the formula, This represents the number of iterations. The learning rate, with a value in the range [0.001, 0.01], is used to control the update step size; For loss function, , Loss functions right The partial derivatives of .

[0075] loss function The calculation formula is as follows:

[0076] ;

[0077] In the formula, The number of user feedback samples; The first one labeled by the user The actual value score of each data point (based on the effectiveness of data application, with a value of 0-10). The overall value score predicted by the model (normalized to the [0,10] interval); The first one labeled by the user The actual risk level of each data point (low risk = 1, medium risk = 2, high risk = 3); The quantified value of the risk level predicted by the model; Risk item weights are used to balance the loss weights in value assessment and risk assessment.

[0078] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. The terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, unless otherwise explicitly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. Moreover, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0079] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A data analysis-based enterprise data management method, characterized in that, Includes the following steps: Step 1: Collect multi-dimensional data from both inside and outside the enterprise, and preprocess the data; Step 2: Extract three types of core features from the preprocessed effective data. The core features include business relevance features, data value density features, and risk warning features. Step 3: Based on the core features, construct a data value-risk linkage assessment model to calculate the comprehensive value score and risk level of enterprise data; Step 4: Sort the data in descending order based on the comprehensive value score, output high-value data for enterprise decision support, and trigger corresponding early warning mechanisms according to the risk level; Step 5: Based on feedback data from enterprise users regarding the effectiveness of data application, dynamically update the parameters of the data value-risk linkage assessment model to achieve adaptive optimization of the model.

2. The enterprise data management method based on data analysis according to claim 1, characterized in that: The multi-dimensional data mentioned in step 1 includes internal data, external data, and auxiliary data. The internal data includes business data, financial data, and operational data. The external data includes industry data, policy data, and market data. The auxiliary data includes enterprise size tags, industry type tags, and data collection timestamps.

3. The enterprise data management method based on data analysis according to claim 2, characterized in that: The data preprocessing in step 1 includes invalid data removal, noisy data filtering, and data standardization. Invalid data removal includes null value removal, duplicate record deduplication, and non-business data filtering. Noisy data filtering includes rule-based filtering and statistical filtering. Data standardization includes text encoding unification, numerical normalization, and unstructured text structuring transformation.

4. A data analysis-based enterprise data management method according to any one of claims 1-3, characterized in that: The business relevance features mentioned in step 2 The calculation formula is as follows: ; In the formula, The number of keywords matching the data text with the industry's core business thesaurus. This represents the total number of keywords in the data text.

5. The enterprise data management method based on data analysis according to claim 4, characterized in that: The data value density feature described in step 2 The calculation formula is as follows: ; In the formula, , as well as For data integrity Data timeliness and data scarcity The feature weights, and .

6. The enterprise data management method based on data analysis according to claim 5, characterized in that: Risk warning features described in step 2 Risk levels are classified and quantified based on the data anomaly volatility coefficient D. The calculation formula for the data anomaly volatility coefficient D is as follows: ; In the formula, For the target data value, This is the average of similar data over the past three months.

7. The enterprise data management method based on data analysis according to claim 6, characterized in that: The data value-risk linkage assessment model mentioned in step 3 includes a comprehensive value scoring model, the calculation formula of which is as follows: ; In the formula, To score the overall value, For the first Weight parameters of class features, For feature quantization values, The time decay factor, For bias terms; The time decay factor The calculation formula is as follows: ; In the formula, For the number of iterations, For the current time, For data collection time, This is the attenuation coefficient.

8. The enterprise data management method based on data analysis according to claim 7, characterized in that: The data value-risk linkage assessment model also includes a risk level calculation model, the calculation formula of which is as follows: ; In the formula, Risk level, and This is the risk threshold.

9. The enterprise data management method based on data analysis according to claim 8, characterized in that: The model parameter update in step 5 uses gradient descent, and the update objects include feature weights. Bias terms The update formulas for the feature weights and bias terms are as follows: ; ; In the formula, For learning rate, For loss function, , This is the partial derivative of the loss function.