Method and system for unstructured data quality optimization based on mathematical model

By constructing a data quality-related feature set, calculating dynamic weights and nonlinear hierarchies, and combining it with an LSTM model for time-varying optimization, the problems of insufficient dynamism and low resource allocation efficiency in unstructured data quality assessment and optimization are solved, achieving efficient and accurate data management.

CN120596799BActive Publication Date: 2026-02-17GUANGDONG NANFANG NEWSPAPER MEDIA GRP NEW MEDIA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510740258.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2025-05-30
Filing Date
2025-06-05
Publication Date
2026-02-17
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing technologies for unstructured data quality assessment and optimization suffer from insufficient dynamism, static feature weights, limited optimization strategies, and low resource allocation efficiency. They are ill-suited to adapting to the dynamic changes and diversity of data streams, resulting in delayed assessment results and low resource utilization efficiency.

Method used

A mathematical model-based approach is adopted to construct a data quality-related feature set, calculate dynamic weights and nonlinear hierarchies, combine LSTM models for time-varying optimization, generate data quality optimization strategies, design resource allocation priorities, and conduct comprehensive evaluation to improve data quality.

Benefits of technology

It enables dynamic assessment and optimization of unstructured data quality, improves data management efficiency, adapts to the diversity and real-time nature of data sources, enhances assessment accuracy and resource utilization, and provides multi-dimensional and dynamic assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596799B_ABST
    Figure CN120596799B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of unstructured data quality optimization method and system based on mathematical model, including receiving digital asset unstructured data, obtain data source diversity and inflow rate number, construct data quality related feature set;By feature extraction to feature set, the dynamic weight of unstructured data and nonlinear classification are calculated, data quality score and grade are obtained;Based on quality score and grade combination LSTM model carries out time-varying optimization, generates data quality optimization strategy.The present application realizes the dynamic evaluation and optimization of unstructured data quality by mathematical model, significantly improves digital asset management efficiency.Integrated data source diversity and real-time, accurately construct feature set;Dynamic weight and nonlinear classification improve the evaluation accuracy;Time-varying optimization and resource allocation priority design improve resource utilization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of digital asset management, and particularly relates to a non-structured data quality optimization method and system based on a mathematical model. BACKGROUND

[0002] With the rapid development of digital economy, digital assets such as blockchain data, Internet of Things data, and social media data are increasingly important in various industries. Non-structured data, due to its diversity and high generation rate, has become the core of digital asset management. However, existing technologies have significant shortcomings in non-structured data quality evaluation and optimization, limiting the efficient mining and application of data value.

[0003] Existing non-structured data quality evaluation methods mainly rely on static indicators such as completeness, accuracy, and consistency, lacking comprehensive quantification of data source diversity and real-time performance. These methods usually use simple statistical analysis or rule-based evaluation, which cannot adapt to the dynamic changes of data flow, resulting in lagging evaluation results that are difficult to reflect the real-time state of data quality. In addition, existing optimization strategies mainly focus on data cleaning or preprocessing, ignoring dynamic resource allocation and time-varying optimization based on quality evaluation, making it difficult to achieve continuous quality improvement.

[0004] In terms of weight distribution, existing technologies usually use fixed weights or subjectively set weight models, failing to fully consider the dynamic correlation and non-linear grading needs between features. This leads to a lack of refinement in quality evaluation results, making it difficult to guide precise optimization strategies. At the same time, existing methods lack priority design based on quality scores and change trends in resource allocation, resulting in low resource utilization efficiency. In addition, existing technologies have less comprehensive evaluation of optimization effects, focusing only on single indicators, and failing to provide multi-dimensional and dynamic evaluation results.

[0005] Therefore, existing technologies have problems such as lack of dynamicity, static feature weights, single optimization strategies, and low resource allocation efficiency in non-structured data quality evaluation and optimization. There is an urgent need for an evaluation and optimization method that can comprehensively consider data source diversity, real-time performance, and dynamic changes to improve the quality management level of non-structured data of digital assets and meet the complex needs of blockchain, Internet of Things, and social media scenarios. SUMMARY

[0006] The purpose of the present application is to provide a non-structured data quality optimization method and system based on a mathematical model to solve the problems of lack of dynamicity, static feature weights, single optimization strategies, and low resource allocation efficiency in existing technologies.

[0007] To achieve one of the above-mentioned purposes, an embodiment of the present application provides a non-structured data quality optimization method based on a mathematical model, which comprises,

[0008] receiving the digital asset unstructured data, obtaining data source diversity and inflow rate number, and constructing a data quality related feature set;

[0009] calculating the dynamic weight and non-linear classification of the unstructured data by feature extraction on the feature set, to obtain data quality scores and grades;

[0010] based on the quality score and grade, combining an LSTM model for time-varying optimization to generate a data quality optimization strategy.

[0011] As a further improvement of an embodiment of the application, the method further comprises that the constructing a data quality related feature set comprises,

[0012] extracting the variability driven features of the digital unstructured data by principal component analysis to construct the data quality related feature set;

[0013] the variability driven features are expressed as:

[0014]

[0015] wherein subscript j is the jth feature in the data quality related feature set, is the initial weight of the feature, is the coefficient of variation of the feature, is the information entropy of the feature, is the heterogeneity-real-time index, expressed as:

[0016]

[0017] wherein, is the proportion of the ith data source in the digital asset unstructured data, v is the data flow rate, p is the rate sensitivity coefficient, and N is the total data amount of the digital asset unstructured data.

[0018] As a further improvement of an embodiment of the application, the method further comprises that the "calculating the dynamic weight and non-linear classification of the unstructured data by feature extraction on the feature set" comprises,

[0019] the calculation formula of the dynamic weight is:

[0020]

[0021] wherein, is the feature extraction score, is the feature-quality coupling matrix, expressed as:

[0022]

[0023] wherein, is a quality indicator selected from one or more of integrity, accuracy and consistency, wherein integrity represents the completeness of data, accuracy represents the compliance of data with true value, and consistency represents the logical consistency between data; η is a decay coefficient, is a feature correlation coefficient of the quality indicator .

[0024] The non-linear grading is represented as:

[0025]

[0026] wherein, is a quality score, is a grading sensitivity, , is the mean of the quality score . is a dynamic grading coefficient.

[0027] As a further improvement of an embodiment of the present application, the method further comprises that the time-varying optimization based on the quality score and the grade combined LSTM model comprises,

[0028] by minimizing the prediction error of the digital asset quality indicator , an optimization strategy for improving the quality of unstructured data is generated, and an LSTM model is used to take the quality score Q and the dynamic weight as input features to predict the quality indicator;

[0029] The optimization strategy includes adjusting the data collection frequency, optimizing the feature extraction algorithm, or dynamically allocating computing resources to improve the level of the quality indicator;

[0030] The optimization objective of the time-varying optimization is to minimize the prediction error of the digital asset quality indicator to generate an optimization strategy for improving the quality of unstructured data, and the optimization objective is represented as:

[0031]

[0032] wherein, is an optimization weight, is the actual value of the quality indicator, is the predicted value of the LSTM model, is a regularization coefficient, and θ is an optimization parameter.

[0033] As a further improvement of an embodiment of the present application, the method further comprises resource allocation based on the data quality optimization strategy;

[0034] By designing a resource allocation priority, the resource is preferentially allocated to the data source with a data quality lower than a preset threshold, and the resource allocation priority is is expressed as:

[0035]

[0036] wherein, is a quality reference threshold, indicating a standardized target value of a digital asset quality score, is a time-varying value of the quality score Q, is a resource allocation adjustment constant, used to control the sensitivity of the priority .

[0037] As a further improvement of an embodiment of the present application, the method further comprises comprehensive evaluation of the optimization effect according to the allocation result;

[0038] A feedback gain is used as an index of the optimization effect of the data quality optimization strategy; the feedback gain is expressed as:

[0039]

[0040] wherein, is the optimized digital asset quality, expressed as:

[0041]

[0042] wherein, is a quality index increment, and t is a time variable.

[0043] As a further improvement of an embodiment of the present application, the method further comprises quantifying the overall effect of the unstructured data optimization strategy of the digital asset by using a comprehensive evaluation score, and the comprehensive evaluation score The calculation formula of the comprehensive evaluation score is:

[0044]

[0045] A final report is generated based on the comprehensive evaluation score, which is used to guide subsequent data quality management.

[0046] To achieve one of the above-mentioned purposes, an embodiment of the present application further provides an unstructured data quality optimization system based on a mathematical model, which comprises a receiving module, a quality evaluation module, a time-varying optimization module and a resource allocation module;

[0047] The receiving module is used to receive unstructured data of a digital asset, obtain data source diversity and inflow rate, and construct a data quality related feature set;

[0048] The quality evaluation module is used for calculating the dynamic weight and nonlinear grading of the unstructured data by feature extraction on the feature set, to obtain the data quality score and grade;

[0049] The time-varying optimization module is used for time-varying optimization based on the quality score and grade combined with the LSTM model, to generate the data quality optimization strategy.

[0050] To achieve one of the above-mentioned purposes, an embodiment of the present application further provides an electronic device comprising a memory and a processor, characterized in that the memory stores a computer program executable on the processor, and the processor executes the program to implement the steps of the above-mentioned mathematical model-based unstructured data quality optimization method.

[0051] To achieve one of the above-mentioned purposes, an embodiment of the present application further provides a storage medium storing a computer program, characterized in that the computer program is executed by a processor to implement the steps of the above-mentioned mathematical model-based unstructured data quality optimization method.

[0052] Compared with the prior art, the present application provides a mathematical model-based unstructured data quality optimization method and system, which realizes dynamic evaluation and optimization of unstructured data quality through mathematical models, significantly improving the efficiency of digital asset management. The diversity and real-time nature of the data source are comprehensively considered to accurately construct a feature set; dynamic weight and nonlinear grading improve evaluation accuracy; time-varying optimization and resource allocation priority design improve resource utilization; comprehensive evaluation quantifies optimization effect, ensuring continuous quality improvement. It is suitable for scenarios such as blockchain, Internet of Things, and social media, effectively solving the problems of insufficient dynamicity, static weight, and single optimization in the prior art, and providing an efficient and accurate quality management solution for complex data environments. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 is the overall flowchart of the mathematical model-based unstructured data quality optimization method described in the present application.

[0054] Figure 2 is the architecture schematic diagram of the mathematical model-based unstructured data quality optimization system described in the present application. DETAILED DESCRIPTION

[0055] The present application will be described in detail below with reference to the specific embodiments shown in the accompanying drawings. However, these embodiments do not limit the present application, and any changes in structure, method, or function made by those of ordinary skill in the art based on these embodiments are included within the scope of the present application.

[0056] Embodiments of the present application are described below in detail, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary only, for the purpose of explanation, and are not to be taken as limiting of the present application.

[0057] In the first embodiment of the present application, the present application provides a non-structured data quality optimization method based on a mathematical model, as shown in Figure 1 The method comprises,

[0058] S1: receiving digital asset non-structured data, obtaining data source diversity and inflow rate number, and constructing a data quality related feature set;

[0059] S2: calculating the dynamic weight and non-linear classification of non-structured data by feature extraction on the feature set, to obtain data quality score and grade;

[0060] S3: combining the quality score and grade with the LSTM model for time-varying optimization to generate a data quality optimization strategy.

[0061] In one specific embodiment of the present application, the data quality related feature set is constructed, specifically,

[0062] The variability driven features of the digital non-structured data are extracted by principal component analysis to construct the data quality related feature set;

[0063] The variability driven features are expressed as:

[0064]

[0065] wherein subscript j is the jth feature in the data quality related feature set, is the initial weight of the feature, is the coefficient of variation of the feature, is the information entropy of the feature, is the heterogeneity-real-time index, expressed as:

[0066]

[0067] wherein, is the proportion of the ith data source in the digital asset non-structured data, v is the data flow rate, p is the rate sensitivity coefficient, and N is the total data amount of the digital asset non-structured data.

[0068] It should be noted that the unstructured data described in the present application includes but is not limited to text, image, video and log data, and is widely used in blockchain data management, Internet of Things data processing and social media data analysis scenarios. Before constructing the data quality related feature set, the received digital asset unstructured data needs to be preprocessed. The preprocessing includes data cleaning (removing redundancy, missing or abnormal values), format standardization and normalization processing to ensure data consistency and computability. For example, for text data, remove invalid characters and duplicate content; for image data, adjust the resolution and color space; for video data, extract key frames or metadata; for log data, parse timestamps and event labels. Preprocessing ensures the quality and consistency of the input data, providing a reliable foundation for feature extraction.

[0069] Further, the core of feature set construction is to extract the variability driven features of digital unstructured data through principal component analysis (PCA). PCA maps the original data to a low-dimensional space through linear transformation, retains the main variability information of the data, and reduces the computational complexity. The specific steps include: data matrix construction: convert unstructured data into feature vector form, such as text word vector, image pixel matrix or video frame feature. Covariance matrix calculation: analyze the correlation between feature vectors to generate a covariance matrix. Eigenvalue decomposition: extract the principal eigenvectors of the covariance matrix to determine the main variability direction. Feature selection: select the top k principal components (k is determined by the cumulative contribution rate, usually ≥ 85%), forming the variability driven feature set. In addition to PCA, other feature extraction methods can be supplemented according to the data type, such as term frequency-inverse document frequency (TF-IDF) for text data or convolutional neural network (CNN) feature extraction for image data, to enhance the diversity and adaptability of the features.

[0070] Further, the constructed feature set quantifies the variability of unstructured data, providing input for subsequent dynamic weight calculation, nonlinear grading and time-varying optimization. The feature set supports multiple data types (text, image, video, log) and adapts to scenarios such as blockchain (transaction record analysis), Internet of Things (sensor data fusion) and social media (user behavior analysis). Feature extraction algorithms (such as PCA) are combined with mathematical models to ensure the comprehensiveness and accuracy of the feature set.

[0071] In one embodiment of the present application, by performing feature extraction on the feature set, the dynamic weight and nonlinear grading of unstructured data are calculated, specifically,

[0072] The calculation formula of the dynamic weight is:

[0073]

[0074] wherein, a score for feature extraction, a feature-quality coupling matrix, denoted as:

[0075]

[0076] wherein, a quality indicator, selected from one or more of completeness, accuracy and consistency, wherein completeness represents the completeness of data, accuracy represents the compliance of data with true values, and consistency represents the logical consistency among data; η is a decay coefficient, a feature correlation coefficient with the quality indicator

[0077] the non-linear grading , denoted as:

[0078]

[0079] wherein, a quality score, a grading sensitivity, , a mean value of the quality score a dynamic grading coefficient.

[0080] Further, obtained by normalized decomposition of the variability-driven feature , and the specific decomposition process is as follows:

[0081] feature vector construction: the feature vector is denoted as a weighted sum of the feature set, wherein each feature is composed of , and j is the feature index;

[0082] principal component analysis decomposition: PCA is applied to the feature set, the covariance matrix of the feature vector is calculated, the principal component is extracted, and the score of each feature is generated, wherein i corresponds to the i-th feature, consistent with j;

[0083] standardization: the is normalized to the interval [0, 1], and the formula is , to ensure the comprehensive correlation with ;

[0084] boundary processing: when approaches zero, a zero-division constant is set to ensure the stability of calculation.

[0085] It should be noted that the dynamic weight ​​The calculation of the dynamic weight is used to quantify the contribution of the feature set to the data quality evaluation, fully considering the dynamic correlation between the features and the quality indicators. The calculation process is as follows The correlation between the features and the quality indicators is introduced, so that the weight can dynamically reflect the real-time changes of the data quality. For example, in the blockchain transaction data, if the accuracy is highly related to a certain feature , the weight of the feature will be increased accordingly.

[0086] Further, the nonlinear grading is used to map the data quality score to the grading interval to realize the grading evaluation of the data quality. The nonlinear grading adopts the Sigmoid function, which realizes the nonlinear mapping of the quality score through and , and ensures that the grading result can reflect the subtle differences in quality. For example, in the Internet of Things data processing, low-quality data may be mapped to a lower level, and high-quality data may be mapped to a higher level, facilitating the development of subsequent optimization strategies.

[0087] In one specific embodiment of the present application, the time-varying optimization is performed based on the combination of the quality score and the grade LSTM model, specifically,

[0088] By minimizing the prediction error of the digital asset quality indicator , an optimization strategy for improving the quality of unstructured data is generated, and the LSTM model is used to predict the quality indicator with the quality score Q and the dynamic weight as input features;

[0089] The optimization strategy includes adjusting the data collection frequency, optimizing the feature extraction algorithm, or dynamically allocating computing resources to improve the level of the quality indicator;

[0090] The optimization objective of the time-varying optimization is to minimize the prediction error of the digital asset quality indicator to generate an optimization strategy for improving the quality of unstructured data, and the optimization objective is represented as:

[0091]

[0092] wherein, is the optimization weight, is the actual value of the quality indicator, is the predicted value of the LSTM model, is the regularization coefficient, and θ is the optimization parameter.

[0093] ​It should be noted that the core objective of time-varying optimization is to generate dynamic optimization strategies by minimizing the prediction error of digital asset quality indicators, continuously improving the quality of unstructured data. The optimization process utilizes LSTM models to capture the temporal dependence of data quality, combining quality scores and dynamic weights to dynamically adjust data processing flows (such as data collection frequency, computational resource allocation).

[0094] Further, the optimization weight is obtained by linear transformation, specifically:

[0095]

[0096] where is the weight adjustment coefficient, determined according to the importance of application scenarios (such as blockchain, Internet of Things) or quality indicators (integrity, accuracy, consistency), this transformation enhances or weakens the contribution of the optimization objective to ensure reflects the relative importance of each quality indicator in time-varying optimization.

[0097] Further, the LSTM model is used to predict the quality indicator , with the quality score Q and dynamic weight as input features. The construction and application of the LSTM model includes the following steps:

[0098] Input feature preparation: the quality score Q and dynamic weight are combined into a time series input, normalized to the [0, 1] interval to improve model stability.

[0099] Model structure design: the LSTM model includes an input layer, multiple LSTM units (usually 2-3 layers, with 64-256 units), a fully connected layer, and an output layer. The tanh activation function is used to handle non-linear relationships, and the Dropout layer (dropout rate such as 0.2) is used to prevent overfitting.

[0100] Training process: the model is trained using historical data (including Q, and ), the optimizer is selected as Adam, and the loss function is mean squared error. During training, the batch size (such as 32) and the number of iterations (such as 100) are adjusted according to the data size.

[0101] Prediction output: the LSTM model predicts the quality indicator at the future time step based on the current and historical input, which is used for the calculation of the optimization objective.

[0102] Further, the dynamic regularization coefficient ​According to the digital asset quality decay factor adaptive adjustment. The quality decay factor is calculated by analyzing the trend of data quality change over time (such as the integrity decline rate), and the formula is:

[0103]

[0104] wherein, is the initial regularization coefficient, is the adjustment coefficient, which controls the influence of the decay factor, is the time-varying change of the quality score, reflecting the change of the quality over time. When is larger, decreases, allowing the model to be more flexible to fit the data; when is smaller, increases, enhancing regularization to improve model stability.

[0105] Further, the optimization strategy is generated based on the results of , specifically including:

[0106] Data collection frequency adjustment: when the predicted quality is lower than the threshold, increase the data collection frequency to improve the integrity.

[0107] Resource allocation optimization: preferentially allocate computing or storage resources to low-quality data sources, based on the subsequent resource allocation priority formula.

[0108] Preprocessing improvement: adjust data cleaning or feature extraction algorithms according to the prediction error, such as enhancing outlier detection.

[0109] In one specific embodiment of the present application, resource allocation is performed based on the data quality optimization strategy, specifically,

[0110] By designing resource allocation priority, resources are preferentially allocated to data sources with quality lower than the preset threshold, and the resource allocation priority is expressed as:

[0111]

[0112] wherein, is the quality reference threshold, representing the standardized target value of the digital asset quality score, is the time-varying value of the quality score Q, is the resource allocation adjustment constant, used to control the sensitivity of the priority .

[0113] It should be noted that the goal of resource allocation is to dynamically schedule limited computing, storage or bandwidth resources according to data quality optimization strategies and data quality trends, to prioritize improving data sources with quality below the preset threshold, thereby improving the overall quality level of unstructured data. The resource allocation priority is quantified by a mathematical formula to ensure the accuracy and efficiency of resource allocation. Formula By ensuring that data sources with lower quality have higher priority, resources are allocated to improve their quality.

[0114] Further, the calculation process of resource allocation priority includes the following steps:

[0115] Quality difference calculation: Calculate the difference between the current quality score and the reference threshold to reflect the quality gap of the data source. The larger the difference, the lower the quality of the data source, and the more resources need to be allocated.

[0116] Weight adjustment: Use optimization weights to weight the quality difference according to the importance of the data source or quality indicators, ensuring that resource allocation reflects feature priorities. For example, in an Internet of Things scenario, the accuracy of sensor data may have a higher .

[0117] Normalization: Normalize the priority by the denominator to avoid excessively large or small priority values and maintain the stability of the allocation. The choice of needs to be adjusted according to the total amount of resources and the number of data sources.

[0118] Boundary processing: When approaches zero or negative, set to a non-zero constant to ensure calculation stability; when is too small, set a lower limit to avoid the risk of division by zero.

[0119] Further, resource allocation is based on priority guidance, which includes the following dynamic scheduling of resource types:

[0120] Computing resources: Prioritize allocating CPU / GPU computing power to low-quality data sources for enhanced data cleaning, feature extraction or model training. For example, for blockchain transaction log data, increase computing power to improve data consistency.

[0121] Storage resources: Prioritize allocating storage space to low-quality data sources to ensure high availability of critical data. For example, for social media image data, allocate more storage to support high-resolution data retention.

[0122] Bandwidth resource: Prioritize network bandwidth allocation to low-quality data sources to speed up data collection or transmission. For example, increase bandwidth for Internet of Things sensor data to improve real-time performance.

[0123] In one embodiment of the present application, the optimization effect is comprehensively evaluated according to the allocation result, specifically,

[0124] The feedback gain is used as an index of the optimization effect of the data quality optimization strategy; the feedback gain is expressed as:

[0125]

[0126] wherein, is the quality of the optimized digital assets, expressed as:

[0127]

[0128] wherein, is the quality index increment, and t is the time variable.

[0129] It should be noted that the goal of comprehensive evaluation is to quantify the effect of data quality optimization strategy through feedback gain, to evaluate the improvement of unstructured data quality after resource allocation, and to generate a final report to guide subsequent data management. The feedback gain As a core evaluation index, the difference between the quality scores before and after optimization, the resource allocation priority and the feature-quality coupling relationship are comprehensively considered. The feedback gain The difference between the quality scores before and after optimization is compared , combined with the resource allocation priority and the feature-quality coupling, to quantify the effect of the optimization strategy. A higher indicates that the optimization effect is significant and the resource allocation is efficient.

[0130] In one embodiment of the present application, the overall effect of the unstructured data optimization strategy of the digital assets is quantified by the comprehensive evaluation score;

[0131] Specifically, the comprehensive evaluation score is calculated by the formula:

[0132]

[0133] Based on the comprehensive evaluation score, a final report is generated to guide subsequent data quality management.

[0134] It should be noted that the goal of analyzing the data quality improvement effect is to quantify the overall effect of the unstructured data optimization strategy of the digital assets through the comprehensive evaluation score The overall improvement effect of the quantization optimization strategy on the quality of unstructured data is evaluated by comprehensively considering the optimized quality score, feedback gain, nonlinear grading, and feature-quality coupling relationship, thereby providing a scientific basis for data quality management. By integrating the optimized quality , feedback gain , nonlinear grading , and feature-quality coupling , the effect of the optimization strategy is comprehensively evaluated, and the larger the value is, the more significant the optimization effect is.

[0135] Further, based on the comprehensive evaluation score , a final report is generated to guide subsequent data quality management. The report content includes:

[0136] Optimization effect summary: according to , the overall effect of the optimization strategy is evaluated. If is higher than the preset threshold, it means that the optimization is significant; if it is lower than the threshold, it needs to be analyzed.

[0137] Quality indicator analysis: list the quality scores Q and Q' before and after optimization, feedback gain , and nonlinear grading , analyze the improvement of each quality indicator (such as completeness, accuracy).

[0138] Resource allocation evaluation: combined with the resource allocation priority , evaluate the resource utilization efficiency, and identify whether the resource scheduling strategy needs to be adjusted.

[0139] Improvement suggestions: according to and the quality indicator increment , suggestions for optimizing data collection frequency, feature extraction algorithm, or resource allocation are proposed. For example, in the Internet of Things scenario, if the completeness improvement is insufficient, the sensor data collection frequency can be increased.

[0140] The final report is presented in a visual form (such as charts, tables) for easy understanding and decision-making by users.

[0141] In the second embodiment of the present application, a mathematical model-based unstructured data quality optimization system is provided, which includes a receiving module 1, a quality evaluation module 2, a time-varying optimization module 3, and a resource allocation module 4;

[0142] The receiving module 1 is used to receive digital asset unstructured data, obtain data source diversity and inflow rate, and construct a data quality-related feature set.

[0143] The quality evaluation module 2 is used for calculating the dynamic weight and nonlinear grading of the unstructured data by feature extraction on the feature set, so as to obtain the data quality score and grade.

[0144] The time-varying optimization module 3 is used for time-varying optimization based on the quality score and grade combined with the LSTM model, so as to generate the data quality optimization strategy.

[0145] In the third embodiment of the present application, the present application provides an electronic device comprising a memory and a processor, characterized in that the memory stores a computer program capable of running on the processor, and the processor implements the steps in the above-mentioned mathematical model-based unstructured data quality optimization method when executing the program.

[0146] In the fourth embodiment of the present application, the present application provides a storage medium storing a computer program, characterized in that the computer program implements the steps in the above-mentioned mathematical model-based unstructured data quality optimization method when executed by a processor.

[0147] In summary, the present application provides a mathematical model-based unstructured data quality optimization method and system, which realizes dynamic evaluation and optimization of unstructured data quality through mathematical models, significantly improving the efficiency of digital asset management. The diversity and real-time nature of the data sources are comprehensively considered to accurately construct a feature set; dynamic weight and nonlinear grading improve evaluation accuracy; time-varying optimization and resource allocation priority design improve resource utilization; comprehensive evaluation quantifies optimization effects to ensure continuous quality improvement. The present application is suitable for scenarios such as blockchain, Internet of Things, and social media, effectively solving the problems of insufficient dynamicity, static weight, and single optimization in the prior art, and providing an efficient and accurate quality management solution for complex data environments.

[0148] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described modules can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0149] The modules described as separate components can or can not be physically separated, and the components shown as modules can or can not be physical modules, i.e., they can be located in one place or distributed to multiple network modules. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment.

[0150] In addition, the functional modules in each embodiment of the present application can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated module can be realized in the form of hardware or hardware plus software function module.

[0151] The integrated module realized in the form of the software function module can be stored in a computer readable storage medium. The software function module is stored in a storage medium and includes a plurality of instructions for enabling a computer system (which can be a personal computer, a server, or a network system, etc.) or a processor to execute part of the steps of the method described in the various embodiments of the present application. The storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media capable of storing program codes.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for non-structured data quality optimization based on mathematical model, characterized in that: Comprising, receiving unstructured data of digital assets, obtaining data source diversity and inflow rate, and constructing a data quality related feature set; extracting features from the feature set, calculating dynamic weights and non-linear classification of unstructured data, and obtaining data quality scores and levels; based on the quality score and level, combining the LSTM model for time-varying optimization to generate data quality optimization strategy; the time-varying optimization based on the quality score and level combined with the LSTM model includes, By minimizing the prediction error of the digital asset quality indicator An optimization strategy for improving the quality of unstructured data is generated, and an LSTM model is used with quality score Q and dynamic weight as input features for quality indicator prediction; the optimization strategy includes adjusting data collection frequency, optimizing feature extraction algorithm or dynamically allocating computing resources to improve the level of the quality indicator; the optimization target of the time-varying optimization is to minimize the prediction error of the digital asset quality indicator to generate an optimization strategy to improve the quality of unstructured data, and the optimization target is represented as: ; wherein, for optimizing the weights, for the actual value of the quality indicator, for the predicted value of the LSTM model, is a regularization coefficient, and θ is an optimization parameter; based on the data quality optimization strategy, resource allocation includes, By designing a resource allocation priority, the resource is preferentially allocated to the data source with data quality lower than the preset threshold, and the resource allocation priority is represented as: ; wherein, is a quality reference threshold, representing a standardized target value for the quality score of the digital asset, is a time-varying value for the quality score Q, is a resource allocation adjustment constant, used to control the sensitivity of the priority .

2. The method for non-structured data quality optimization based on mathematical model according to claim 1, characterized in that: the construction of the data quality related feature set includes, extracting variability driven features of the unstructured data of the digital assets by principal component analysis to construct the data quality related feature set; The variability driving feature is represented as: ; wherein subscript j is the jth feature in the data quality related feature set, is the initial weight of the feature, is the coefficient of variation of the feature, is the information entropy of the feature, is the heterogeneity-real-time index, expressed as: ; wherein, is the proportion of the i-th type of data source in the unstructured data of the digital asset, v is the data flow rate, p is the rate sensitivity coefficient, and N is the total data volume of the unstructured data of the digital asset.

3. The method of claim 2, wherein: the "extracting features from the feature set, calculating dynamic weights and non-linear classification of unstructured data" includes, The dynamic weight The calculation formula is: ; wherein, is a feature extraction score, is a feature-quality coupling matrix, represented as: ; wherein, is a quality indicator selected from one or more of integrity, accuracy and consistency, wherein integrity represents the completeness of the data, accuracy represents the agreement of the data with the true value, and consistency represents the logical consistency between the data; η is a decay coefficient, is a feature correlation coefficient with the quality indicator . said non-linear hierarchy is represented as: ; wherein, is the quality score, is the grading sensitivity, , is the quality score is the mean, is the dynamic grading coefficient.

4. The method of claim 3, wherein: also includes comprehensive evaluation of optimization effect according to the allocation result; adopting a feedback gain as an index of optimization effect of the data quality optimization strategy; the feedback gain is represented as: ; wherein, is the optimized digital asset quality, expressed as: ; wherein, is the quality index increment, t is the time variable.

5. The method for non-structured data quality optimization based on mathematical model according to claim 4, characterized in that: also includes, Quantify the overall effectiveness of the unstructured data optimization strategy for digital assets using a composite score, the composite score is calculated as: ; generating a final report based on the comprehensive evaluation score to guide subsequent data quality management.

6. A mathematical model based unstructured data quality optimization system applied to the mathematical model based unstructured data quality optimization method as claimed in claim 1, characterized in that: comprising a receiving module, a quality evaluation module, a time-varying optimization module and a resource allocation module; the receiving module is used to receive unstructured data of digital assets, obtain data source diversity and inflow rate, and construct a data quality related feature set; the quality evaluation module is used to extract features from the feature set, calculate dynamic weights and non-linear classification of unstructured data, and obtain data quality scores and levels; the time-varying optimization module is used to combine the LSTM model based on the quality score and level for time-varying optimization to generate data quality optimization strategy.

7. An electronic device comprising a memory and a processor, characterized in that: the memory stores a computer program executable on the processor, and the processor executes the program to realize the steps in the unstructured data quality optimization method based on mathematical model according to any one of claims 1-5.

8. A storage medium, storing a computer program, characterized in that: the computer program is executed by the processor to realize the steps in the unstructured data quality optimization method based on mathematical model according to any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-source heterogeneous quality information fusion processing method and system for power grid main equipment

    CN115099338A

  • Air quality forecasting system based on LSTM-MEA-SVR

    CN116796291A