DMS data management system
By introducing a DMS data management system with sparseness evaluation, dynamic Bayesian estimation and data dependency detection, the problems of insufficient data accuracy evaluation and inefficient acquisition efficiency are solved, and efficient and reliable data acquisition and processing are achieved.
Patent Information
- Application Number
- CN202411569180.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-11-05
AI Technical Summary
When existing data management systems process complex data sets, there are problems such as insufficient data accuracy evaluation, low data collection efficiency and limited data set optimization methods, resulting in wrong judgments and low data quality during data processing.
The sparseness evaluation module, uncertainty evaluation module and optimization acquisition strategy module are introduced to optimize the data acquisition strategy through sparseness evaluation, dynamic Bayesian estimation and data dependency detection to realize hierarchical data acquisition.
It improves the accuracy of data accuracy evaluation, optimizes the data acquisition process, and enhances data quality control. It is especially suitable for the processing of dynamic and large-scale data sets, reducing the cost of redundant data generation and processing time.
Smart Images

Figure CN119474844B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management systems, in particular to a DMS data management system. Background Art
[0002] With the rapid development of big data technology, data collection, management, and analysis are becoming increasingly important. However, existing data management systems often have the following problems when processing complex data sets:
[0003] Insufficient data accuracy assessment: Existing data management systems lack effective evaluation models and are unable to accurately measure the overall accuracy and uncertainty of data sets. This can easily lead to misjudgments during data processing and affect the reliability of data decisions.
[0004] Inefficient data collection: Traditional data collection methods do not effectively optimize and layer data sources and features, resulting in low data collection and analysis efficiency. This is especially true when processing large or dynamic data sets, which often makes it difficult to meet real-time requirements.
[0005] Limited means of dataset optimization: Existing systems often lack effective optimization strategies when dealing with dataset complexity and are unable to perform targeted optimization collection based on the characteristics and uncertainties of different data sources, resulting in high dataset redundancy and low data quality. Summary of the Invention
[0006] Based on the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a DMS data management system to solve the above-mentioned technical problems.
[0007] To achieve the above objectives, the present invention provides the following technical solutions: DMS data management system, comprising:
[0008] Sparsity evaluation module: used to perform sparsity evaluation on the comprehensive data set in the DMS data management system and output data sparsity evaluation results;
[0009] Uncertainty assessment module: used to perform dynamic Bayesian estimation and uncertainty quantification on the comprehensive data set in the DMS data management system, and output uncertainty quantification assessment results;
[0010] An acquisition strategy optimization module is configured to optimize the acquisition strategy according to the data sparsity evaluation result and the uncertainty quantification evaluation result;
[0011] Data layered collection module: used to detect the data dependency of the collected data, score the data source through a preset data quality assessment algorithm, and perform data layered collection in descending order of the data source score.
[0012] The present invention is further configured such that the sparsity evaluation module includes a sampling sparsity evaluation unit, a sampling divergence evaluation unit, and a data sparsity evaluation result unit;
[0013] Sampling sparsity evaluation unit: used to apply low sampling rate detection algorithm to comprehensive data sets to evaluate the sampling sparsity of data dimensions;
[0014] Sampling divergence evaluation unit: for evaluating the sampling divergence of data dimensions whose sampling sparsity values are lower than a preset sparsity threshold according to a preset sampling divergence analysis;
[0015] Data sparsity evaluation result unit: used to output the sampling sparsity and the sampling divergence as a data sparsity evaluation result.
[0016] The present invention is further configured such that the calculation logic of the low sampling rate detection algorithm is: ,in, For the Sampling sparsity of data dimensions, is the length of the time series, is a regulating factor used to control the sensitivity of the nonlinear function. For the The data dimension is The data value of a time series sample point, For the The data dimension is The center value of the time series sample points;
[0017] The calculation logic of the preset sampling divergence analysis is: ,in, For the The sampling divergence of the data dimension, is the actual distribution probability, representing the The data dimension is The observation frequency of a time series sample point, is the theoretical expected probability, representing the ideal data distribution, A small amount to avoid singularities in the logarithmic function.
[0018] The present invention is further configured such that the uncertainty assessment module includes a prior distribution construction module, a Bayesian update module and an uncertainty quantification module;
[0019] Prior distribution construction module, used to construct prior distribution for comprehensive data sets in the DMS data management system according to data dimensions;
[0020] A Bayesian update module is used to apply a dynamic Bayesian update algorithm based on the prior distribution to adjust the parameter estimates of the data dimension in real time and construct a posterior distribution;
[0021] The uncertainty quantification module is used to calculate the uncertainty of the data dimension according to the posterior distribution.
[0022] The present invention is further configured such that the construction logic of the prior distribution is: ,in, For the The prior distribution of the data dimensions, is a normalization constant used to ensure that the sum of the prior distribution is 1, For the The time series sample space of data dimensions, For the The data dimension is The data value of the time series sample point and the Parameters of data dimensions Energy function of
[0023] The construction logic of the posterior distribution is: ,in, For the The posterior distribution of the data dimensions, For the The likelihood function of the data dimension is expressed as When it is observed The probability of For the The parameter space of data dimensions;
[0024] The calculation logic of the uncertainty of the data dimension is: ,in, For the The uncertainty of each data dimension.
[0025] The present invention is further configured to optimize the acquisition strategy, including:
[0026] Perform a weighted summation of sampling sparsity and uncertainty to obtain a collection priority score, and perform collection in descending order of the collection priority score;
[0027] The collection volume of data dimensions is dynamically adjusted based on the collection priority score. The calculation logic of the collection volume is as follows: ,in, For the The amount of data collected in each dimension, For the Priority scoring for collection of each data dimension, The total amount of available collection resources.
[0028] The present invention is further configured such that the data layered acquisition module includes a dependency detection unit, a data source scoring unit, and a data layered acquisition unit;
[0029] A dependency detection unit, configured to construct a data dependency matrix based on the collected data;
[0030] A data source scoring unit, configured to score the data source using a preset data quality assessment algorithm according to the data dependency matrix;
[0031] The data layer collection unit is used to collect data in layers in descending order of the scores of the data sources.
[0032] The present invention is further configured such that the construction logic of the data dependency matrix is: ,in, is the data dependency matrix, For the The data source is in The value of the data dimension, For the The data source is in The value of the data dimension, To remove the The number of data sources for each data source, For the The adoption space of data dimensions, Small values introduced to avoid singularities of the logarithmic function;
[0033] The logic for scoring data sources is: ,in, For the The ratings of the data sources, is the length of the time series, is a regulating factor used to control the impact of dependency on the score.
[0034] The present invention is further configured such that the logic of layered collection is: ,in, For the The data dimension is The amount of data collected from each data source, For the The amount of data collected in each dimension.
[0035] The present invention provides a DMS data management system, comprising: a sparsity assessment module for performing sparsity assessment on a comprehensive data set in the DMS data management system and outputting a data sparsity assessment result; an uncertainty assessment module for performing dynamic Bayesian estimation and uncertainty quantification on the comprehensive data set in the DMS data management system and outputting an uncertainty quantification assessment result; an optimization acquisition strategy module for optimizing the acquisition strategy based on the data sparsity assessment result and the uncertainty quantification assessment result; and a data hierarchical acquisition module for performing data dependency detection on the collected data, scoring the data source using a preset data quality assessment algorithm, and performing data hierarchical acquisition in descending order of the data source score. The beneficial effects produced include:
[0036] 1. Improve the accuracy of data accuracy assessment: This invention introduces an accuracy assessment module and an uncertainty assessment module, which can dynamically evaluate comprehensive data sets, avoiding the errors caused by the single assessment method in traditional methods. By quantifying data uncertainty, it can provide more accurate assessment results, thereby improving the credibility and reliability of the data.
[0037] 2. Optimize the data collection process: By optimizing the collection strategy module, the present invention can perform hierarchical collection and processing of data sets from different sources based on data accuracy and uncertainty assessment results. This approach not only reduces the generation of redundant data but also significantly improves data collection efficiency. It is particularly suitable for processing dynamic, large-scale data sets, greatly reducing the time cost of collection and processing.
[0038] 3. Enhanced data quality control: The present invention provides a mechanism that combines hierarchical data collection and data quality assessment, which can automatically monitor and evaluate data quality, and adjust data collection strategies in real time according to fluctuations in data quality, ensuring high-quality data input and output, thereby providing a reliable data foundation for the data analysis process.
[0039] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. In the drawings:
[0041] Figure 1 This is a schematic structural diagram of a DMS data management system according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following describes the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art will readily appreciate the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the various details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.
[0043] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0044] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.
[0045] DMS data management system, such as Figure 1 Shown, including:
[0046] Sparsity evaluation module: used to perform sparsity evaluation on the comprehensive data set in the DMS data management system and output data sparsity evaluation results;
[0047] Uncertainty assessment module: used to perform dynamic Bayesian estimation and uncertainty quantification on the comprehensive data set in the DMS data management system, and output uncertainty quantification assessment results;
[0048] An acquisition strategy optimization module is configured to optimize the acquisition strategy according to the data sparsity evaluation result and the uncertainty quantification evaluation result;
[0049] Data layered collection module: used to detect the data dependency of the collected data, score the data source through a preset data quality assessment algorithm, and perform data layered collection in descending order of the data source score.
[0050] Specifically, the sparsity assessment module is used to perform sparsity assessment on the comprehensive data set in the DMS data management system. Its main task is to identify sparse areas in the data set, that is, areas with low sampling frequency or insufficient samples. The assessment results are used for subsequent strategy optimization. The present invention is further configured such that the sparsity assessment module includes a sampling sparsity assessment unit, a sampling divergence assessment unit, and a data sparsity assessment result unit.
[0051] Sampling sparsity evaluation unit: used to apply a low sampling rate detection algorithm to the comprehensive data set to evaluate the sampling sparsity of the data dimension; the present invention is further configured such that the calculation logic of the low sampling rate detection algorithm is: ,in, For the Sampling sparsity of data dimensions, is the length of the time series, is a regulating factor used to control the sensitivity of the nonlinear function. For the The data dimension is The data value of a time series sample point, For the The data dimension is The data center value of each time series sample point; specifically, the above calculation logic is used to evaluate the sampling sparsity of the data dimension, and quantify the sparsity of the sampling by detecting the deviation between the sample points of different time series and the center value. The higher the sparsity, the less sampling of the data dimension, and the data may not be complete. The calculation of sampling sparsity is achieved through nonlinear functions, which can accurately reflect the distribution of data at different time points. The sparsity of each data dimension is determined by the difference between the observed value in the time series and the center value of the dimension. The greater the difference, the higher the sparsity may be. Through nonlinear functions , can effectively amplify the impact of sparse sampling areas, quantify the sparsity of sampling, and further, adjust the factor It is used to control the sensitivity of nonlinear function to data differences, with a value range of 0.1 to 10. The larger the value, the more sensitive the function is to differences in the data.
[0052] Sampling divergence evaluation unit: for evaluating the sampling divergence of data dimensions whose sampling sparsity values are lower than a preset sparsity threshold according to a preset sampling divergence analysis; the calculation logic of the preset sampling divergence analysis is: ,in, For the The sampling divergence of the data dimension, is the actual distribution probability, representing the The data dimension is The observation frequency of a time series sample point, is the theoretical expected probability, representing the ideal data distribution, To avoid the slight amount of singularity of the logarithmic function. Specifically, the above calculation logic calculates the sampling divergence of each data dimension by comparing the difference between the actual sampling distribution and the expected distribution. Sampling divergence is used to evaluate whether the data of a certain dimension deviates from the expected sampling distribution. By accumulating the logarithmic difference between the actual distribution probability and the theoretical distribution probability at each time series point, the degree of deviation of the data is quantified; sampling divergence evaluation is used to quantify the difference between the actual sampling and theoretical distribution of the data dimension. By comparing the actual sampling distribution and expected distribution , evaluate the degree of deviation of the data dimension from the ideal distribution at different time points; further, Take a very small value, including 10 −6 , to avoid the denominator being zero when calculating the logarithmic function. There is no restriction on the specific value here. By performing logarithmic processing and divergence calculation on the data distribution, sample points with significant deviations can be effectively amplified, helping the system identify abnormal sampling points or data noise, thereby further improving data quality.
[0053] Data sparsity assessment result unit: used to output the sampling sparsity and sampling divergence as a data sparsity assessment result. Specifically, the output here can be either a weighted sum of the sampling sparsity and sampling divergence, or output as a data pair. By combining sampling sparsity and sampling divergence, this unit can comprehensively assess data sparsity issues, identifying both insufficient data frequency and imbalanced data distribution, thereby comprehensively measuring data quality;
[0054] The present invention is further configured such that the uncertainty assessment module includes a prior distribution construction module, a Bayesian update module and an uncertainty quantification module;
[0055] The prior distribution construction module is used to construct a prior distribution for the comprehensive data set in the DMS data management system according to the data dimension; the present invention is further configured such that the construction logic of the prior distribution is: ,in, For the The prior distribution of the data dimensions, is a normalization constant used to ensure that the sum of the prior distribution is 1, For the The time series sample space of data dimensions, For the The data dimension is The data value of the time series sample point and the Parameters of data dimensions Energy function; Specifically, the above calculation logic is used to construct a prior distribution for the comprehensive data set in the DMS data management system and determine the initial probability distribution of each data dimension. In dynamic Bayesian estimation, the prior distribution is the basis of the entire reasoning process and represents the system's initial assumptions about the data model parameters when there is no observed data. The energy function is used to define the relationship between data and model parameters, and the normalization constant is combined to ensure that the sum of the probability distribution is 1, ultimately generating a prior distribution for each data dimension; further, the parameters is the model parameter, indicating the A set of prior parameters for each data dimension. According to the specific application scenario, The combination of parameters including mean, variance, regression coefficient, etc. has a value range determined by the distribution characteristics of the data and is not specifically restricted here. The prior distribution reflects the system's initial assumptions about the parameters of each data dimension, helps guide the system on how to infer parameters in the absence of data, and provides a basis for subsequent dynamic adjustments.
[0056] The Bayesian update module is used to apply the dynamic Bayesian update algorithm based on the prior distribution, adjust the parameter estimates of the data dimension in real time, and construct the posterior distribution; the construction logic of the posterior distribution is: ,in, For the The posterior distribution of the data dimensions, For the The likelihood function of the data dimension is expressed as When it is observed The probability of For the The parameter space of each data dimension; specifically, the above calculation logic is based on the principle of Bayesian reasoning, combining the prior distribution with the observed data, dynamically adjusting the parameter estimation of each data dimension, and finally constructing the posterior distribution. This module updates the system's inference of the data in real time through Bayes' theorem, especially when new data arrives, it can adjust the previous assumptions and estimates based on the latest observations. The posterior distribution represents the system's correction of the uncertainty of the model parameters after combining the observed data; the likelihood function depends on the relationship between the observed data and the parameters. Its value range is 0 to 1, indicating that given the parameters Under, the observed data The probability of a prediction error. Larger values indicate a better fit between the parameters and the observed data. Bayesian updating dynamically adjusts model parameters based on real-time observations. As new data arrives, the system continuously optimizes parameter estimates, making the model more accurate and flexible. The Bayesian updating module provides the system with a mechanism to cope with data uncertainty and noise. Even when data is sparse or incomplete, the system can still make reasonable inferences based on prior knowledge, reducing model uncertainty.
[0057] The uncertainty quantification module is used to calculate the uncertainty of the data dimension based on the posterior distribution; the calculation logic of the uncertainty of the data dimension is: ,in, For the The uncertainty of each data dimension; specifically, the above calculation logic is based on the Bayesian posterior distribution to quantify the uncertainty of the data dimension. The core of this module is to use entropy as a metric to evaluate the degree of uncertainty in parameter estimation. The entropy calculation can give the uncertainty quantification result of each data dimension. The larger the entropy value, the better the model is at determining the parameters. The higher the uncertainty of the estimate, the smaller the entropy value, which means the more certain the model is about the parameter.
[0058] The present invention is further configured to optimize the acquisition strategy, including:
[0059] The sampling sparsity and uncertainty are weighted and summed to obtain the collection priority score, and collection is performed in descending order according to the collection priority score. Specifically, through the weighted collection priority score, the system can dynamically adjust the order of data collection, give priority to collecting dimensions with high scores, and help the system collect the most important data with limited collection resources, thereby improving the coverage and completeness of the overall data.
[0060] The collection volume of data dimensions is dynamically adjusted based on the collection priority score. The calculation logic of the collection volume is as follows: ,in, For the The amount of data collected in each dimension, For the Priority scoring for collection of each data dimension, is the total amount of available collection resources. Specifically, the above logic is used to dynamically adjust the amount of data collected for each dimension based on the collection priority score of each data dimension. By weighting the available resources according to the priority score, the system can better allocate limited collection resources and prioritize the collection of data dimensions with high scores. The ultimate goal is to optimize the data collection strategy and ensure that resources are reasonably allocated to the dimensions that need them most; based on the collection priority score of each data dimension, , the available collection resources Assigned to each data dimension , and rationally allocates resources based on the scoring weights, ensuring that collection resources are prioritized for high-scoring dimensions. By dynamically allocating resources based on collection priority scores, the system can allocate limited collection resources to the dimensions most in need, maximizing data quality. The system considers both the sparsity of data sampling and the uncertainty of the model's understanding of data dimensions, rationally allocating collection resources through weighted scoring to ensure comprehensive data coverage and model stability.
[0061] The present invention is further configured such that the data layered acquisition module includes a dependency detection unit, a data source scoring unit, and a data layered acquisition unit;
[0062] The dependency detection unit is used to construct a data dependency matrix based on the collected data. The present invention is further configured such that the construction logic of the data dependency matrix is: ,in, is the data dependency matrix, For the The data source is in The value of the data dimension, For the The data source is in The value of the data dimension, To remove the The number of data sources for each data source, For the The adoption space of data dimensions, To avoid the tiny values introduced by the singularity of the logarithmic function; specifically, the above calculation logic constructs a data dependency matrix to evaluate the mutual dependence of different data sources in a specific data dimension. By comparing the collection values of multiple data sources on the same data dimension, the matrix quantifies the degree of dependence between data sources. The higher the dependency, the stronger the dependence of a data source on other data sources in that dimension, thus reflecting the independence or redundancy of the data; through the dependency matrix, the system can identify which data sources are dependent on other data sources and which data sources are more independent. This is of great significance for optimizing data collection and reducing redundant data; data sources with higher dependence can reduce the collection frequency, thereby reducing system resource consumption. On the contrary, independent data sources with lower dependence can be retained first or the collection frequency can be increased to ensure data diversity and independence.
[0063] The data source scoring unit is used to score the data source using a preset data quality assessment algorithm according to the data dependency matrix; the logic for scoring the data source is: ,in, For the The ratings of the data sources, is the length of the time series, is an adjustment factor used to control the impact of dependency on the score. Specifically, the above calculation logic scores the quality of each data source based on the constructed data dependency matrix. The score quantifies the independence, integrity, and reliability of the data source based on data dependency. The purpose of the score is to give higher scores to data sources with strong independence and high data quality, and lower scores to data sources with strong dependence by considering the degree of dependence of the data source on different time series. The scoring results are used to guide subsequent collection strategies and data processing optimization. The adjustment factor Used to adjust the impact of dependency on the score. This makes the score more sensitive to changes in dependency, with a value range of 0.1 to 10. By scoring the dependency of a data source, the system can clearly identify the independence and reliability of the data source. Data sources with strong independence and low dependency receive higher scores, which helps guide data collection strategies.
[0064] The data layered collection unit is used to collect data in layers according to the descending order of the scores of the data sources. The present invention is further configured such that the logic of the layered collection is: ,in, For the The data dimension is The amount of data collected from each data source, For the Specifically, the above calculation logic is based on the score of each data source. , and collect data from each data source in descending order. The system will sort the total amount of collected resources according to the score. Allocate resources proportionally to different data sources. Data sources with higher scores will be allocated more resources and prioritized for collection. The purpose of this logic is to ensure that the system prioritizes the collection of data sources with strong independence and high data quality, thereby improving the reliability and validity of the overall data. By allocating resources based on scores, the system can prioritize the collection of high-quality, independent data sources, ensuring the most effective use of resources and improving the quality of overall data collection. It automatically reduces data collection from low-scoring data sources, reduces the collection of redundant data, optimizes the efficiency of data storage and processing, and realizes dynamic hierarchical collection of data sources. The above mechanism is flexible and dynamic, and can adjust the collection strategy in real time based on the score, providing an efficient resource allocation solution for complex data management systems.
[0065] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0066] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0067] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0068] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0069] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0070] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0071] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0072] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0073] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0074] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0075] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. DMS data management system, characterized by: include: Sparsity evaluation module: used to perform sparsity evaluation on the comprehensive data set in the DMS data management system and output data sparsity evaluation results; Uncertainty assessment module: used to perform dynamic Bayesian estimation and uncertainty quantification on the comprehensive data set in the DMS data management system, and output uncertainty quantification assessment results; An acquisition strategy optimization module is configured to optimize the acquisition strategy according to the data sparsity evaluation result and the uncertainty quantification evaluation result; Data layered collection module: used to detect the data dependency of the collected data, score the data source using a preset data quality assessment algorithm, and perform data layered collection in descending order of the data source score; The data layered acquisition module includes a dependency detection unit, a data source scoring unit, and a data layered acquisition unit; the dependency detection unit is used to construct a data dependency matrix based on the collected data; A data source scoring unit, configured to score the data source using a preset data quality assessment algorithm according to the data dependency matrix; A data stratification collection unit, used to collect data in stratified order according to the scores of the data sources in descending order; The construction logic of the data dependency matrix is: ,in, is the data dependency matrix, For the The data source is in The value of the data dimension, For the The data source is in The value of the data dimension, To remove the The number of data sources for each data source, For the The adoption space of data dimensions, Small values introduced to avoid singularities of the logarithmic function; The logic for scoring data sources is: ,in, For the The ratings of the data sources, is the length of the time series, is a regulating factor used to control the impact of dependency on the score.
2. The DMS data management system according to claim 1, characterized in that: The sparsity evaluation module includes a sampling sparsity evaluation unit, a sampling divergence evaluation unit and a data sparsity evaluation result unit; Sampling sparsity evaluation unit: used to apply low sampling rate detection algorithm to comprehensive data sets to evaluate the sampling sparsity of data dimensions; Sampling divergence evaluation unit: for evaluating the sampling divergence of data dimensions whose sampling sparsity values are lower than a preset sparsity threshold according to a preset sampling divergence analysis; Data sparsity evaluation result unit: used to output the sampling sparsity and the sampling divergence as a data sparsity evaluation result.
3. The DMS data management system according to claim 2, characterized in that: The calculation logic of the low sampling rate detection algorithm is: ,in, For the Sampling sparsity of data dimensions, is the length of the time series, is a regulating factor used to control the sensitivity of the nonlinear function. For the The data dimension is The data value of a time series sample point, For the The data dimension is The center value of the time series sample points; The calculation logic of the preset sampling divergence analysis is: ,in, For the The sampling divergence of the data dimension, is the actual distribution probability, representing the The data dimension is The observation frequency of a time series sample point, is the theoretical expected probability, representing the ideal data distribution, A small amount to avoid singularities in the logarithmic function.
4. The DMS data management system according to claim 2, characterized in that: The uncertainty assessment module includes a prior distribution construction module, a Bayesian update module and an uncertainty quantification module; Prior distribution construction module, used to construct prior distribution for comprehensive data sets in the DMS data management system according to data dimensions; A Bayesian update module is used to apply a dynamic Bayesian update algorithm based on the prior distribution to adjust the parameter estimates of the data dimension in real time and construct a posterior distribution; The uncertainty quantification module is used to calculate the uncertainty of the data dimension according to the posterior distribution.
5. The DMS data management system according to claim 4, characterized in that: The construction logic of the prior distribution is: ,in, For the The prior distribution of the data dimensions, is a normalization constant used to ensure that the sum of the prior distribution is 1, For the The time series sample space of data dimensions, For the The data dimension is The data value of the time series sample point and the Parameters of data dimensions Energy function of The construction logic of the posterior distribution is: ,in, For the The posterior distribution of the data dimensions, For the The likelihood function of the data dimension is expressed as When it is observed The probability of For the The parameter space of data dimensions; The calculation logic of the uncertainty of the data dimension is: ,in, For the The uncertainty of each data dimension.
6. The DMS data management system according to claim 4, characterized in that: Optimize collection strategies, including: Perform a weighted summation of sampling sparsity and uncertainty to obtain a collection priority score, and perform collection in descending order of the collection priority score; The collection volume of data dimensions is dynamically adjusted based on the collection priority score. The calculation logic of the collection volume is as follows: ,in, For the The amount of data collected in each dimension, For the The collection priority score of each data dimension, The total amount of available collection resources.
7. The DMS data management system according to claim 1, characterized in that: The logic of layered collection is: ,in, For the The data dimension is The amount of data collected from each data source, For the The amount of data collected in each dimension.
Citation Information
Patent Citations
Graded forecasting method based on user liveness
CN102982466A
E-commerce recommendation method based on dynamic interest group identifier and generative adversarial network
CN112231583A