Mathematical model-based unstructured data quality optimization method and system

By constructing a data quality-related feature set, calculating dynamic weights and nonlinear grading, and combining the LSTM model for time-varying optimization, the problems of insufficient dynamics and low resource allocation efficiency in unstructured data quality assessment and optimization are solved, achieving efficient and accurate data management.

CN120596799AActive Publication Date: 2025-09-05GUANGDONG NANFANG NEWSPAPER MEDIA GRP NEW MEDIA CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510740258.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-05-30
Filing Date
2025-06-05
Publication Date
2025-09-05
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing technologies in unstructured data quality assessment and optimization have problems such as insufficient dynamics, static feature weights, single optimization strategies, and low resource allocation efficiency. They are difficult to adapt to dynamic changes in data flows, resulting in delayed assessment results and inefficient resource utilization.

Method used

A mathematical model-based approach is used to construct a data quality-related feature set, calculate dynamic weights and nonlinear grading, combine the LSTM model for time-varying optimization, generate optimization strategies, design resource allocation priorities, and conduct a comprehensive evaluation.

Benefits of technology

It realizes dynamic evaluation and optimization of unstructured data quality, improves data management efficiency, adapts to the diversity and real-time nature of data sources, improves evaluation accuracy and resource utilization, and provides multi-dimensional and dynamic evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596799A_ABST
    Figure CN120596799A_ABST
Patent Text Reader

Abstract

The invention relates to an unstructured data quality optimization method and system based on a mathematical model, and the method comprises the steps: receiving the unstructured data of digital assets, obtaining the diversity and inflow rate of a data source, and constructing a data quality related feature set; performing feature extraction on the feature set, and calculating the dynamic weight and nonlinear grading of the unstructured data to obtain a data quality score and grade; and performing time-varying optimization based on the quality score and the grade in combination with an LSTM model to generate a data quality optimization strategy. According to the method, the dynamic evaluation and optimization of the unstructured data quality are realized through the mathematical model, and the digital asset management efficiency is remarkably improved. Integrating data source diversity and real-time performance, and accurately constructing a feature set; the evaluation accuracy is improved through dynamic weight and nonlinear grading; the resource utilization rate is improved through time-varying optimization and resource allocation priority design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital asset management, and in particular relates to a method and system for optimizing the quality of unstructured data based on a mathematical model. Background Art

[0002] With the rapid development of the digital economy, digital assets such as blockchain data, IoT data, and social media data are becoming increasingly important across various industries. Unstructured data, due to its diversity and rapid generation rate, has become central to digital asset management. However, existing technologies for assessing and optimizing unstructured data quality have significant shortcomings, limiting the efficient mining and application of data value.

[0003] Existing unstructured data quality assessment methods primarily rely on static metrics such as completeness, accuracy, and consistency, lacking comprehensive quantification of the diversity and real-time nature of data sources. These methods typically employ simple statistical analysis or rule-based assessments, which are unable to adapt to the dynamic nature of data streams. This results in delayed assessment results and a failure to reflect the real-time status of data quality. Furthermore, existing optimization strategies often focus on data cleaning or preprocessing, neglecting dynamic resource allocation and time-varying optimization based on quality assessment, making it difficult to achieve continuous quality improvement.

[0004] When it comes to weight allocation, existing technologies typically use fixed weights or subjectively defined weight models, failing to fully consider the dynamic correlations between features and the need for nonlinear grading. This results in a lack of refinement in quality assessment results, making it difficult to guide precise optimization strategies. Furthermore, existing methods lack prioritization based on quality scores and changing trends in resource allocation, resulting in inefficient resource utilization. Furthermore, existing technologies rarely provide comprehensive evaluations of optimization effectiveness, focusing solely on a single metric and failing to provide multi-dimensional, dynamic evaluation results.

[0005] Therefore, existing technologies in the quality assessment and optimization of unstructured data have problems such as insufficient dynamics, static feature weights, single optimization strategy, and low resource allocation efficiency. There is an urgent need for an assessment and optimization method that can comprehensively consider the diversity, real-time and dynamic changes of data sources to improve the quality management level of digital asset unstructured data and meet the complex needs of scenarios such as blockchain, Internet of Things and social media. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for optimizing the quality of unstructured data based on a mathematical model, so as to solve the problems of the existing technology such as insufficient dynamics, static feature weights, single optimization strategy, and low resource allocation efficiency.

[0007] To achieve one of the above-mentioned objectives, an embodiment of the present invention provides a method for optimizing the quality of unstructured data based on a mathematical model, the method comprising:

[0008] Receive unstructured digital asset data, obtain data source diversity and inflow rate data, and build a feature set related to data quality;

[0009] By extracting features from feature sets, calculating the dynamic weight and nonlinear classification of unstructured data, we can obtain data quality scores and grades.

[0010] Based on the quality score and grade combined with the LSTM model, time-varying optimization is performed to generate a data quality optimization strategy.

[0011] As a further improvement of an embodiment of the present invention, the method further includes: constructing a data quality related feature set includes:

[0012] Extracting variability-driving features of the digital unstructured data through principal component analysis to construct the data quality-related feature set;

[0013] The variability driving characteristics Expressed as:

[0014]

[0015] Among them, subscript j is the jth feature in the data quality related feature set, is the initial weight of the feature, is the characteristic coefficient of variation, is the feature information entropy, is the heterogeneity-real-time index, expressed as:

[0016]

[0017] in, is the proportion of the i-th data source in the unstructured data of digital assets, v is the data flow rate, ρ is the rate sensitivity coefficient, and N is the total amount of unstructured data of digital assets.

[0018] As a further improvement of an embodiment of the present invention, the method further includes that the “calculating the dynamic weight and nonlinear classification of unstructured data by extracting features from the feature set” includes:

[0019] The dynamic weight The calculation formula is:

[0020]

[0021] in, is the feature extraction score, is the characteristic-mass coupling matrix, expressed as:

[0022]

[0023] in, is a quality indicator, selected from one or more of completeness, accuracy and consistency, where completeness indicates the completeness of the data, accuracy indicates the degree of conformity between the data and the true value, and consistency indicates the logical consistency between the data; η is the attenuation coefficient, Features and quality indicators Correlation coefficient of

[0024] The nonlinear grading , expressed as:

[0025]

[0026] in, is the quality score, For graded sensitivity, , Quality Score The mean of is the dynamic classification coefficient.

[0027] As a further improvement of an embodiment of the present invention, the method further includes that the time-varying optimization based on the quality score and grade combined with the LSTM model includes:

[0028] By minimizing the digital asset quality index The prediction error is used to generate an optimization strategy for improving the quality of unstructured data, and the LSTM model is used to calculate the quality score Q and dynamic weights. Used as input features for quality indicator prediction;

[0029] The optimization strategy includes adjusting the data collection frequency, optimizing the feature extraction algorithm or dynamically allocating computing resources to improve the level of the quality indicator;

[0030] The optimization goal of the time-varying optimization is to minimize the prediction error of the digital asset quality indicator to generate an optimization strategy for improving the quality of unstructured data. The optimization goal is expressed as:

[0031]

[0032] in, To optimize the weights, is the actual value of the quality indicator, is the predicted value of the LSTM model, is the regularization coefficient, and θ is the optimization parameter.

[0033] As a further improvement of an embodiment of the present invention, the method further includes allocating resources based on a data quality optimization strategy;

[0034] By designing resource allocation priorities, resources are guided to be allocated preferentially to data sources whose data quality is lower than a preset threshold. Expressed as:

[0035]

[0036] in, is the quality reference threshold, which represents the standardized target value of the digital asset quality score. is the time-varying value of the quality score Q, Allocate resource adjustments to control priorities sensitivity.

[0037] As a further improvement of an embodiment of the present invention, the method further includes performing a comprehensive evaluation of the optimization effect based on the allocation result;

[0038] Feedback gain is used as an indicator of the optimization effect of the data quality optimization strategy; the feedback gain Expressed as:

[0039]

[0040] in, is the quality of digital assets after optimization, expressed as:

[0041]

[0042] in, is the quality index increment, and t is the time variable.

[0043] As a further improvement of an embodiment of the present invention, the method further includes quantifying the overall effect of the digital asset unstructured data optimization strategy using a comprehensive evaluation score. The calculation formula is:

[0044]

[0045] A final report is generated based on the comprehensive evaluation score to guide subsequent data quality management.

[0046] To achieve one of the above-mentioned objects of the invention, an embodiment of the present invention further provides an unstructured data quality optimization system based on a mathematical model, the system comprising a receiving module, a quality assessment module, a time-varying optimization module and a resource allocation module;

[0047] The receiving module is used to receive unstructured digital asset data, obtain data source diversity and inflow rate data, and construct a data quality-related feature set;

[0048] The quality assessment module is used to extract features from the feature set, calculate the dynamic weight and nonlinear classification of unstructured data, and obtain data quality scores and grades;

[0049] The time-varying optimization module is used to perform time-varying optimization based on the quality score and grade combined with the LSTM model to generate a data quality optimization strategy.

[0050] In order to achieve one of the above-mentioned purposes of the invention, an embodiment of the present invention also provides an electronic device, including a memory and a processor, characterized in that the memory stores a computer program that can be run on the processor, and when the program is executed on the processor, the steps in the above-mentioned unstructured data quality optimization method based on mathematical models are implemented.

[0051] In order to achieve one of the above-mentioned purposes of the invention, an embodiment of the present invention further provides a storage medium, which stores a computer program, and is characterized in that when the computer program is executed by a processor, it implements the steps in the above-mentioned unstructured data quality optimization method based on the mathematical model.

[0052] Compared with existing technologies, the present invention provides a mathematical model-based unstructured data quality optimization method and system. This method uses mathematical models to dynamically evaluate and optimize unstructured data quality, significantly improving digital asset management efficiency. It integrates the diversity and real-time nature of data sources to accurately construct feature sets; dynamic weighting and nonlinear grading improve evaluation accuracy; time-varying optimization and resource allocation priority design improve resource utilization; and comprehensive evaluation quantifies optimization effects to ensure continuous quality improvement. Applicable to scenarios such as blockchain, the Internet of Things, and social media, this method effectively addresses the problems of existing technologies, such as insufficient dynamism, static weighting, and single optimization, providing efficient and accurate quality management solutions for complex data environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is an overall flow chart of the unstructured data quality optimization method based on mathematical model described in the present invention.

[0054] Figure 2 It is a schematic diagram of the architecture of the unstructured data quality optimization system based on mathematical model described in the present invention. DETAILED DESCRIPTION

[0055] The present invention will be described in detail below with reference to the specific embodiments shown in the accompanying drawings. However, these embodiments do not limit the present invention, and any structural, methodological, or functional changes made by those skilled in the art based on these embodiments are all within the scope of protection of the present invention.

[0056] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and are not to be construed as limiting the present invention.

[0057] In the first embodiment of the present invention, the present invention provides a method for optimizing the quality of unstructured data based on a mathematical model, such as Figure 1 As shown, the method includes,

[0058] S1: Receive unstructured digital asset data, obtain data source diversity and inflow rate data, and construct a data quality-related feature set;

[0059] S2: By extracting features from the feature set, the dynamic weight and nonlinear classification of unstructured data are calculated to obtain the data quality score and grade;

[0060] S3: Based on the quality score and grade combined with the LSTM model, time-varying optimization is performed to generate a data quality optimization strategy.

[0061] In a specific embodiment of the present invention, a data quality related feature set is constructed, specifically,

[0062] Extracting variability-driving features of the digital unstructured data through principal component analysis to construct the data quality-related feature set;

[0063] The variability driving characteristics Expressed as:

[0064]

[0065] Among them, subscript j is the jth feature in the data quality related feature set, is the initial weight of the feature, is the characteristic coefficient of variation, is the feature information entropy, is the heterogeneity-real-time index, expressed as:

[0066]

[0067] in, is the proportion of the i-th data source in the unstructured data of digital assets, v is the data flow rate, ρ is the rate sensitivity coefficient, and N is the total amount of unstructured data of digital assets.

[0068] It should be noted that the unstructured data described in this invention includes, but is not limited to, text, images, videos, and log data, and is widely used in scenarios such as blockchain data management, IoT data processing, and social media data analysis. Before constructing a data quality-related feature set, the received unstructured digital asset data must be preprocessed. This preprocessing includes data cleaning (removing redundancies, missing values, or outliers), format standardization, and normalization to ensure data consistency and computability. For example, for text data, invalid characters and duplicate content are removed; for image data, the resolution and color space are adjusted; for video data, keyframes or metadata are extracted; and for log data, timestamps and event tags are parsed. Preprocessing ensures the quality and consistency of input data, providing a reliable foundation for feature extraction.

[0069] Furthermore, the core of feature set construction lies in extracting the variability-driving features of digital unstructured data through principal component analysis (PCA). PCA maps the raw data into a low-dimensional space through linear transformation, retaining the data's main variation information and reducing computational complexity. Specific steps include: Data matrix construction: Converting unstructured data into feature vector form, such as word vectors for text, pixel matrices for images, or frame features for videos. Covariance matrix calculation: Analyzing the correlation between feature vectors to generate a covariance matrix. Eigenvalue decomposition: Extracting the principal eigenvectors of the covariance matrix to determine the main direction of variation. Feature selection: Selecting the top k principal components (k is determined by the cumulative contribution rate, typically ≥85%) to form a variability-driving feature set. In addition to PCA, other feature extraction methods can be supplemented based on the data type, such as term frequency-inverse document frequency (TF-IDF) for text data or convolutional neural network (CNN) feature extraction for image data, to enhance feature diversity and adaptability.

[0070] Furthermore, the constructed feature set quantifies the variability of unstructured data, providing input for subsequent dynamic weight calculation, nonlinear grading, and time-varying optimization. The feature set supports a variety of data types (text, images, videos, and logs) and is suitable for scenarios such as blockchain (transaction record analysis), the Internet of Things (sensor data fusion), and social media (user behavior analysis). Feature extraction algorithms (such as PCA) are combined with mathematical models to ensure the comprehensiveness and accuracy of the feature set.

[0071] In a specific embodiment of the present invention, the dynamic weight and nonlinear classification of unstructured data are calculated by extracting features from the feature set, specifically,

[0072] The dynamic weight The calculation formula is:

[0073]

[0074] in, is the feature extraction score, is the characteristic-mass coupling matrix, expressed as:

[0075]

[0076] in, is a quality indicator, selected from one or more of completeness, accuracy and consistency, where completeness indicates the completeness of the data, accuracy indicates the degree of conformity between the data and the true value, and consistency indicates the logical consistency between the data; η is the attenuation coefficient, Features and quality indicators Correlation coefficient of

[0077] The nonlinear grading , expressed as:

[0078]

[0079] in, is the quality score, For graded sensitivity, , Quality Score The mean of is the dynamic classification coefficient.

[0080] Further, By driving the variability The standardized decomposition is obtained, and the specific decomposition process is as follows:

[0081] Feature vector construction: It is expressed as the weighted sum of the feature set, where each feature is represented by Composition, j is the feature index;

[0082] Principal Component Analysis Decomposition: Apply PCA to the feature set and calculate the eigenvectors The covariance matrix of , extracts the principal components, and generates the scores of each feature , where i corresponds to the i-th feature, which is consistent with j;

[0083] Standardization: Normalization is only in the [0,1] interval, the formula is , ensure that comprehensive associations;

[0084] Boundary processing: When When approaching zero, set the zero division prevention constant to ensure Computationally stable.

[0085] It should be noted that dynamic weight The calculation of is used to quantify the contribution of feature sets to data quality assessment, taking into full account the dynamic correlation between features and quality indicators. Dynamic weight The calculation process is through The introduction of the correlation between features and quality indicators enables the weights to dynamically reflect the real-time changes in data quality. For example, in blockchain transaction data, if the accuracy With a certain feature Highly correlated, the weight of this feature will increase accordingly.

[0086] Furthermore, nonlinear grading It is used to map the data quality score to the grading interval to achieve hierarchical evaluation of data quality. Nonlinear grading uses Sigmoid function. and Implementing nonlinear mapping of quality scores ensures that the grading results can reflect subtle differences in quality. For example, in IoT data processing, low-quality data may be mapped to a lower level, while high-quality data may be mapped to a higher level, facilitating the formulation of subsequent optimization strategies.

[0087] In a specific embodiment of the present invention, time-varying optimization is performed based on quality score and grade combined with LSTM model, specifically,

[0088] By minimizing the digital asset quality index The prediction error is used to generate an optimization strategy for improving the quality of unstructured data, and the LSTM model is used to calculate the quality score Q and dynamic weights. Used as input features for quality indicator prediction;

[0089] The optimization strategy includes adjusting the data collection frequency, optimizing the feature extraction algorithm or dynamically allocating computing resources to improve the level of the quality indicator;

[0090] The optimization goal of the time-varying optimization is to minimize the prediction error of the digital asset quality indicator to generate an optimization strategy for improving the quality of unstructured data. The optimization goal is expressed as:

[0091]

[0092] in, To optimize the weights, is the actual value of the quality indicator, is the predicted value of the LSTM model, is the regularization coefficient, and θ is the optimization parameter.

[0093] It should be noted that the core goal of time-varying optimization is to generate dynamic optimization strategies to continuously improve the quality of unstructured data by minimizing the prediction error of digital asset quality indicators. The optimization process utilizes an LSTM model to capture the temporal dependence of data quality. This combines quality scores with dynamic weights to dynamically adjust data processing procedures (such as data collection frequency and computing resource allocation).

[0094] Further, optimize the weight By dynamic weight Perform linear transformation to obtain:

[0095]

[0096] in, is the weight adjustment coefficient, which is determined by the importance of the application scenario (such as blockchain, Internet of Things) or quality indicators (completeness, accuracy, consistency). This transformation strengthens or weakens In the optimization goal Contribution to ensure Reflects the relative importance of each quality indicator in time-varying optimization.

[0097] Furthermore, the LSTM model is used to predict quality indicators , with quality score Q and dynamic weight As input features. The construction and application of the LSTM model includes the following steps:

[0098] Input feature preparation: quality score Q and dynamic weight The time series input is normalized to the interval [0,1] to improve the model stability.

[0099] Model Architecture: The LSTM model consists of an input layer, multiple LSTM units (typically 2-3 layers with 64-256 units), a fully connected layer, and an output layer. A tanh activation function is used to handle nonlinear relationships, and a dropout layer (with a dropout rate of, for example, 0.2) is used to prevent overfitting.

[0100] Training process: Use historical data (including Q, and ) to train the model, using Adam as the optimizer and mean squared error as the loss function. During training, adjust the batch size (e.g., 32) and number of iterations (e.g., 100) based on the data size.

[0101] Prediction output: The LSTM model predicts quality indicators for future time steps based on current and historical inputs. , used to calculate the optimization target.

[0102] Furthermore, the dynamic regularization coefficient Adaptively adjust based on the digital asset quality attenuation factor. The quality attenuation factor is calculated by analyzing the changing trend of data quality over time (such as the integrity degradation rate). The formula is:

[0103]

[0104] in, is the initial regularization coefficient, To adjust the coefficient, control the influence of the attenuation factor, is the time-varying variation of the quality score, reflecting the change of quality over time. When it is larger, Decreasing allows the model to fit the data more flexibly; when When smaller, Increase to enhance regularization to improve model stability.

[0105] Furthermore, the optimization strategy is based on The results are generated, including:

[0106] Data collection frequency adjustment: when the prediction quality When below the threshold, data collection frequency is increased to improve completeness.

[0107] Resource allocation optimization: Prioritize computing or storage resources for low-quality data sources based on subsequent resource allocation priority formulas.

[0108] Preprocessing improvements: Adjust data cleaning or feature extraction algorithms based on prediction errors, such as enhancing outlier detection.

[0109] In a specific embodiment of the present invention, resource allocation is performed based on a data quality optimization strategy, specifically,

[0110] By designing resource allocation priorities, resources are guided to be allocated preferentially to data sources whose data quality is lower than a preset threshold. Expressed as:

[0111]

[0112] in, is the quality reference threshold, which represents the standardized target value of the digital asset quality score. is the time-varying value of the quality score Q, Allocate resource adjustments to control priorities sensitivity.

[0113] It should be noted that the goal of resource allocation is to dynamically schedule limited computing, storage, or bandwidth resources based on data quality optimization strategies and data quality change trends, giving priority to improving data sources with quality below a preset threshold, thereby improving the overall quality level of unstructured data. Resource allocation priorities are quantified through mathematical formulas to ensure the accuracy and efficiency of resource allocation. Formula Prioritize resources to improve the quality of lower quality data sources by ensuring they receive higher priority.

[0114] Furthermore, resource allocation priority The calculation process includes the following steps:

[0115] Quality difference calculation: calculate the current quality score With reference threshold The difference between the two reflects the quality gap of the data source. The larger the difference, the lower the quality of the data source, and resources need to be allocated first.

[0116] Weight adjustment: using optimized weights ,weighting the quality difference based on the importance of the data source or ,quality metric, ensuring that resource allocation reflects the feature priority.,For example, in an IoT scenario, the accuracy of sensor data may have a ,higher priority. .

[0117] Normalization: by denominator Normalized priority , avoid priority values ​​that are too large or too small and maintain the stability of the allocation. The choice needs to be adjusted according to the total amount of resources and the number of data sources.

[0118] Boundary processing: When When close to zero or negative value, set is a non-zero constant to ensure computational stability; when If the value is too small, set a lower limit to avoid the risk of division by zero.

[0119] Furthermore, resource allocation is based on priority Guidance, including dynamic scheduling of the following resource types:

[0120] Computing resources: Prioritize CPU / GPU computing power to lower-quality data sources for enhanced data cleaning, feature extraction, or model training. For example, for blockchain transaction log data, increasing computing power can improve data consistency.

[0121] Storage resources: Prioritize storage space for lower-quality data sources to ensure high availability of critical data. For example, for social media image data, more storage can be allocated to support high-resolution data retention.

[0122] Bandwidth resources: Prioritize allocating network bandwidth to low-quality data sources to speed up data collection or transmission. For example, for IoT sensor data, increasing bandwidth can improve real-time performance.

[0123] In a specific embodiment of the present invention, a comprehensive evaluation of the optimization effect is performed based on the allocation results, specifically,

[0124] Feedback gain is used as an indicator of the optimization effect of the data quality optimization strategy; the feedback gain Expressed as:

[0125]

[0126] in, is the quality of digital assets after optimization, expressed as:

[0127]

[0128] in, is the quality index increment, and t is the time variable.

[0129] It should be noted that the goal of the comprehensive evaluation is to quantify the effectiveness of the data quality optimization strategy through the feedback gain indicator, evaluate the degree of improvement in the quality of unstructured data after resource allocation, and generate a final report to guide subsequent data management. As the core evaluation indicator, it comprehensively considers the difference in quality scores before and after optimization, resource allocation priority, and feature-quality coupling relationship. By comparing the quality score difference before and after optimization , and combined with resource allocation priority and feature-quality coupling to quantify the effect of the optimization strategy. It indicates that the optimization effect is significant and resource allocation is efficient.

[0130] In one embodiment of the present invention, a comprehensive evaluation score is used to quantify the overall effectiveness of the digital asset unstructured data optimization strategy;

[0131] Specifically, the comprehensive evaluation score The calculation formula is:

[0132]

[0133] A final report is generated based on the comprehensive evaluation score to guide subsequent data quality management.

[0134] It should be noted that the goal of analyzing the effect of data quality improvement is to use comprehensive evaluation scores The overall improvement effect of the optimization strategy on the quality of unstructured data is quantified, and the optimized quality score, feedback gain, nonlinear classification and feature-quality coupling relationship are comprehensively considered to provide a scientific basis for data quality management. Optimized quality through integration , feedback gain , nonlinear grading and feature-mass coupling , comprehensively evaluate the effect of the optimization strategy, and the larger the value, the more significant the optimization effect.

[0135] Furthermore, based on the comprehensive evaluation score , generate a final report to guide subsequent data quality management. The report includes:

[0136] Optimization effect summary: According to value, evaluate the overall effect of the optimization strategy. If it is higher than the preset threshold, it means that the optimization is significant; if it is lower than the threshold, it is necessary to analyze the deficiencies.

[0137] Quality indicator analysis: List the quality scores Q and Q' before and after optimization, feedback gain and nonlinear grading , analyze the improvement of various quality indicators (such as completeness and accuracy).

[0138] Resource Allocation Assessment: Combined with Resource Allocation Priorities , evaluate resource utilization efficiency and identify whether resource scheduling strategies need to be adjusted.

[0139] Improvement suggestions: According to and quality indicator increments ,proposes suggestions for optimizing data collection frequency, feature extraction algorithms, or resource allocation.,For example, in an IoT scenario, if the integrity improvement is insufficient,,the sensor data collection frequency can be increased.

[0140] The final report is presented in visual form (such as charts and tables) to facilitate user understanding and decision-making.

[0141] In a second embodiment of the present invention, the present invention provides an unstructured data quality optimization system based on a mathematical model, the system comprising a receiving module 1, a quality assessment module 2, a time-varying optimization module 3 and a resource allocation module 4;

[0142] The receiving module 1 is used to receive unstructured digital asset data, obtain data source diversity and inflow rate data, and construct a data quality-related feature set;

[0143] The quality assessment module 2 is used to extract features from the feature set, calculate the dynamic weight and nonlinear classification of unstructured data, and obtain the data quality score and grade;

[0144] The time-varying optimization module 3 is used to perform time-varying optimization based on the quality score and grade combined with the LSTM model to generate a data quality optimization strategy.

[0145] In a third embodiment of the present invention, the present invention provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the program is executed on the processor, the steps in the above-mentioned method for optimizing the quality of unstructured data based on a mathematical model are implemented.

[0146] In a fourth embodiment of the present invention, the present invention provides a storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the above-mentioned method for optimizing the quality of unstructured data based on a mathematical model.

[0147] In summary, the present invention provides a mathematical model-based unstructured data quality optimization method and system, which achieves dynamic evaluation and optimization of unstructured data quality through mathematical models, significantly improving the efficiency of digital asset management. It integrates the diversity and real-time nature of data sources to accurately construct feature sets; dynamic weights and nonlinear grading improve evaluation accuracy; time-varying optimization and resource allocation priority design improve resource utilization; and comprehensive evaluation quantifies optimization effects to ensure continuous quality improvement. Applicable to scenarios such as blockchain, the Internet of Things, and social media, it effectively addresses the problems of existing technologies such as insufficient dynamism, static weights, and single optimization, providing efficient and accurate quality management solutions for complex data environments.

[0148] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the modules described above can refer to the corresponding process in the aforementioned method implementation, and will not be repeated here.

[0149] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of these modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0150] In addition, the functional modules in each embodiment of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0151] The above-mentioned integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules are stored in a storage medium and include a number of instructions for causing a computer system (which may be a personal computer, server, or network system, etc.) or a processor to execute some of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A method for optimizing the quality of unstructured data based on a mathematical model, characterized by: include, Receive unstructured digital asset data, obtain data source diversity and inflow rate data, and build a feature set related to data quality; By extracting features from feature sets, calculating the dynamic weight and nonlinear classification of unstructured data, we can obtain data quality scores and grades. Based on the quality score and grade combined with the LSTM model, time-varying optimization is performed to generate a data quality optimization strategy.

2. The unstructured data quality optimization method based on a mathematical model according to claim 1, characterized in that: The construction of the data quality related feature set includes: Extracting variability-driving features of the digital unstructured data through principal component analysis to construct the data quality-related feature set; The variability driving characteristics Expressed as: ; Among them, subscript j is the jth feature in the data quality related feature set, is the initial weight of the feature, is the characteristic coefficient of variation, is the feature information entropy, is the heterogeneity-real-time index, expressed as: ; in, is the proportion of the i-th type of data source in the unstructured data of digital assets, v is the data flow rate, ρ is the rate sensitivity coefficient, and N is the total amount of unstructured data of digital assets.

3. The method for optimizing the quality of unstructured data based on a mathematical model according to claim 2, characterized in that: The "calculating dynamic weights and nonlinear grading of unstructured data by extracting features from feature sets" includes: The dynamic weight The calculation formula is: ; in, is the feature extraction score, is the characteristic-mass coupling matrix, expressed as: ; in, is a quality indicator, selected from one or more of completeness, accuracy and consistency, where completeness indicates the completeness of the data, accuracy indicates the degree of conformity between the data and the true value, and consistency indicates the logical consistency between the data; η is the attenuation coefficient, Characterized by and quality indicators Correlation coefficient of The nonlinear grading , expressed as: ; in, is the quality score, For graded sensitivity, , Quality Score The mean of is the dynamic classification coefficient.

4. The method for optimizing the quality of unstructured data based on a mathematical model according to claim 3, characterized in that: The time-varying optimization based on quality score and grade combined with LSTM model includes: By minimizing the digital asset quality index The prediction error is used to generate an optimization strategy for improving the quality of unstructured data, and the LSTM model is used to calculate the quality score Q and dynamic weights. Used as input features for quality indicator prediction; The optimization strategy includes adjusting the data collection frequency, optimizing the feature extraction algorithm or dynamically allocating computing resources to improve the level of the quality indicator; The optimization goal of the time-varying optimization is to minimize the prediction error of the digital asset quality indicator to generate an optimization strategy for improving the quality of unstructured data. The optimization goal is expressed as: ; in, To optimize the weights, is the actual value of the quality indicator, is the predicted value of the LSTM model, is the regularization coefficient and θ is the optimization parameter.

5. The method for optimizing the quality of unstructured data based on a mathematical model according to claim 4, characterized in that: It also includes resource allocation based on data quality optimization strategies; By designing resource allocation priorities, resources are guided to be allocated preferentially to data sources whose data quality is lower than a preset threshold. Expressed as: ; in, is the quality reference threshold, which represents the standardized target value of the digital asset quality score. is the time-varying value of the quality score Q, Allocate resource adjustments to control priorities sensitivity.

6. The method for optimizing the quality of unstructured data based on a mathematical model according to claim 5, characterized in that: It also includes a comprehensive evaluation of the optimization effect based on the allocation results; Feedback gain is used as an indicator of the optimization effect of the data quality optimization strategy; the feedback gain Expressed as: ; in, is the quality of digital assets after optimization, expressed as: ; in, is the quality index increment, and t is the time variable.

7. The method for optimizing the quality of unstructured data based on a mathematical model according to claim 6, characterized in that: Also includes, A comprehensive evaluation score is used to quantify the overall effect of the digital asset unstructured data optimization strategy. The calculation formula is: ; A final report is generated based on the comprehensive evaluation score to guide subsequent data quality management.

8. A mathematical model-based unstructured data quality optimization system, characterized by: It includes receiving module, quality assessment module, time-varying optimization module and resource allocation module; The receiving module is used to receive unstructured digital asset data, obtain data source diversity and inflow rate data, and construct a data quality-related feature set; The quality assessment module is used to extract features from the feature set, calculate the dynamic weight and nonlinear classification of unstructured data, and obtain data quality scores and grades; The time-varying optimization module is used to perform time-varying optimization based on the quality score and grade combined with the LSTM model to generate a data quality optimization strategy.

9. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program that can be run on the processor, and when the program is executed on the processor, the steps of the unstructured data quality optimization method based on a mathematical model as described in any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the unstructured data quality optimization method based on a mathematical model as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-source heterogeneous quality information fusion processing method and system for power grid main equipment

    CN115099338A

  • Air quality forecasting system based on LSTM-MEA-SVR

    CN116796291A

  • Network resource real-time scheduling method based on multi-objective optimization algorithm

    CN118250167A

  • Big data intelligent decision analysis method and system based on machine learning

    CN119338510A

  • Apparatus and method for optimizing resource of system based on self-adaptive for natwork application

    KR1020160116434A