A management system for business data

By extracting feature vectors and quality assessments from multi-source heterogeneous business data, an adaptive storage strategy matrix is ​​generated, which solves the problem of storage strategy mismatch in existing technologies and achieves efficient and accurate business data storage management.

CN121996172BActive Publication Date: 2026-08-25BEIJING TIANTUO LIXING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610220862.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-24
Publication Date
2026-08-25
Estimated Expiration
2046-02-24

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately capture the correlation between the business value density and structural characteristics of multi-source heterogeneous business data, resulting in severe homogenization of data cluster partitioning, mismatched storage strategies, resource waste, or access performance bottlenecks, and difficulty in adapting to the storage needs of various types of business data clusters.

Method used

The data acquisition and analysis module extracts the basic business feature vector set of multi-source heterogeneous business data. The intelligent quality perception classification module performs feature extraction, classification and quality diagnosis to generate a business deep coding feature vector set and quality confidence assessment value. Combined with the storage-driven evaluation module and the policy fusion inference module, an adaptive storage policy matrix is ​​generated to achieve precise storage management.

Benefits of technology

It achieves accurate capture and correlation matching of data characteristics and quality status, optimizes storage resource allocation, balances storage costs and access performance, improves storage management efficiency and accuracy, and avoids resource waste and performance bottlenecks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996172B_ABST
    Figure CN121996172B_ABST
Patent Text Reader

Abstract

The application discloses a kind of for the management system of business data, it is related to data management technical field.The management system for the business data of the application, comprising: data acquisition analysis module, for obtaining multi-source heterogeneous business data, and extracting business basic feature vector set;Intelligent quality perception classification module, for the business basic feature vector set input into pre-trained quality perception classification network, generate quality confidence evaluation value;Storage drive evaluation module, for extracting storage drive evaluation value;Strategy fusion inference module, for the fusion inference processing of above-mentioned evaluation value, obtain adaptive storage strategy matrix, the adaptive storage strategy matrix in the adaptive storage execution module of the application is to each type of business data cluster Take corresponding storage management measures, so that data characteristics and quality state are clear and controllable, to avoid the mismatch of storage strategy caused by feature and quality matching inaccuracy, and then improve management efficiency and precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, specifically to a management system for business data. Background Technology

[0002] Business data management is a technical system within an enterprise's IT architecture specifically responsible for the persistent storage, organization, maintenance, lifecycle management, and access support of various business data generated during operations. Its core task is to achieve efficient, economical, and reliable utilization of storage resources while meeting business continuity and compliance requirements, which is related to the enterprise's operating costs, data asset value, and responsiveness.

[0003] As enterprises deepen their digital transformation, business data is showing an explosive growth and increasing complexity. This data is typically characterized by multiple sources and heterogeneity: multiple sources mean that data is generated and stored in various independent systems such as customer relationship management, enterprise resource planning, supply chain management, and log platforms; heterogeneity is manifested in the diversity of data structure, such as well-organized relational database tables, semi-structured documents, unstructured text logs and multimedia files, as well as differences in business attributes, such as different value densities, access patterns, update frequencies and compliance requirements.

[0004] Based on the above findings, the limitations of existing technologies include at least the following issues: Existing technologies struggle to accurately capture the correlation between the business value density and structural characteristics of data, leading to severe homogenization in data cluster partitioning. This makes it difficult to provide accurate basis for differentiated storage strategies and easily overlooks the impact of data quality on storage strategy adaptation. Furthermore, the lack of diagnosis and quantitative assessment of data cluster quality status can easily result in mismatches such as insufficient high-quality data storage resources and excessive resource consumption by low-quality data. Additionally, the lack of integration between storage-driven requirements and quality confidence assessment values ​​in storage strategy matching makes it difficult to balance the mutually constraining objectives of storage cost, access performance, and data reliability. This can easily lead to wasted storage resources or bottlenecks in business access performance, making it difficult to adapt to the storage needs of various types of business data clusters, thereby reducing the accuracy and efficiency of business data storage management. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a management system for business data, which solves the problems of inaccurate matching of data characteristics and quality, and inefficient management due to mismatched storage strategies.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a management system for business data, comprising: a data acquisition and analysis module, used to acquire multi-source heterogeneous business data and extract the business basic feature vector set of the multi-source heterogeneous business data; The intelligent quality-aware classification module is used to input the business basic feature vector set into a pre-trained quality-aware classification network. The quality-aware classification network includes a feature extraction sub-network, a classification sub-network, and a quality detection sub-network. The feature extraction sub-network performs feature transformation and deep mining processing on the business basic feature vector set to generate a business deep-encoded feature vector set corresponding to multi-source heterogeneous business data. The classification sub-network performs collaborative classification processing on the business deep-encoded feature vector set to extract a two-dimensional classification label set for the multi-source heterogeneous business data and classify the multi-source heterogeneous business data into several types of business data clusters. The quality detection sub-network performs quality diagnosis processing on the business data clusters to obtain a quality confidence assessment value corresponding to each type of business data cluster. The storage-driven evaluation module is used to extract the storage-driven evaluation value corresponding to each type of business data cluster based on each type of business data cluster; the policy fusion inference module is used to perform fusion inference processing on the storage-driven evaluation value and the quality confidence evaluation value to obtain the adaptive storage policy matrix corresponding to each type of business data cluster; the adaptive storage execution module is used to take corresponding storage management measures for each type of business data cluster based on the adaptive storage policy matrix.

[0007] Further, the specific steps for extracting the basic business feature vector are as follows: preprocessing the multi-source heterogeneous business data; performing feature construction processing based on the preprocessed multi-source heterogeneous business data to generate an original feature set; standardizing the original feature set and concatenating it into a basic business feature vector set in a preset order.

[0008] Further, the feature extraction sub-network includes a feature splitting layer, an embedding layer, a feature fusion layer, and a deep encoding layer. The specific steps for generating the business deep encoding feature vector set are as follows: In the feature splitting layer, the business basic feature vector set is separated to generate a business discrete classification feature vector set and a business numerical feature set corresponding to multi-source heterogeneous business data; In the embedding layer, the business discrete classification feature vector set is densely encoded to generate a discrete feature embedding vector set corresponding to multi-source heterogeneous business data; In the feature fusion layer, the discrete feature embedding vector set and the business continuous numerical feature vector set are concatenated and transformed to generate a fused enhanced feature vector set corresponding to multi-source heterogeneous business data; In the deep encoding layer, the fused enhanced feature vector set is deep-mined and mapped to output the business deep encoding feature vector set.

[0009] Further, the specific steps for generating the fusion-enhanced feature vector set are as follows: performing dimensional concatenation processing on the discrete feature embedding vector set and the continuous numerical feature vector set of the business to generate a business fusion feature vector set corresponding to multi-source heterogeneous business data; performing transformation enhancement processing on the business fusion feature vector set to generate the fusion-enhanced feature vector set.

[0010] Furthermore, the classification sub-network includes a feature dimensionality reduction layer, a label allocation layer, and a classification output layer. The specific steps for generating several types of business data clusters are as follows: In the feature dimensionality reduction layer, the business deep-encoded feature vector set is subjected to dimensionality compression processing to generate a two-dimensional feature vector set corresponding to multi-source heterogeneous business data; In the label allocation layer, the two-dimensional feature vector set is subjected to discretized label allocation processing to generate a two-dimensional classification label set corresponding to multi-source heterogeneous business data; In the classification output layer, the multi-source heterogeneous business data is classified based on the two-dimensional classification label set to output several types of business data clusters.

[0011] Further, the specific steps for generating the two-dimensional classification label set are as follows: normalizing the two-dimensional feature vector set; performing attribution determination processing based on the normalized two-dimensional feature vector set and combined with a preset business classification granularity to generate a spatial attribution mapping set corresponding to multi-source heterogeneous business data; and performing label allocation processing on the two-dimensional feature vector set based on the spatial attribution mapping set to generate the two-dimensional classification label set.

[0012] Furthermore, the quality detection sub-network includes a type-aware diagnostic layer, a quality adaptation layer, and a confidence output layer. The specific steps for obtaining the quality confidence assessment value corresponding to each type of business data cluster are as follows: In the type-aware diagnostic layer, matching processing is performed on each type of business data cluster to generate a diagnostic configuration vector corresponding to the corresponding business data cluster; in the quality adaptation layer, the diagnostic configuration vector is subjected to targeted extraction processing to generate a quality adaptation set corresponding to the corresponding business data cluster; in the confidence output layer, the quality adaptation set is fused to output the quality confidence assessment value.

[0013] Further, the specific steps for extracting the storage driver evaluation value are as follows: Based on each type of business data cluster, extract the storage requirement evaluation set corresponding to each type of business data cluster; perform comprehensive processing on the storage requirement evaluation set to generate the storage driver evaluation value corresponding to each type of business data cluster.

[0014] Further, the specific steps for obtaining the adaptive storage strategy matrix are as follows: Construct a strategy library containing multiple candidate storage strategies, and generate a corresponding strategy feature vector for each candidate storage strategy; extract the classification label corresponding to each type of business data cluster based on each type of business data cluster; construct a storage strategy requirement vector corresponding to each type of business data cluster based on the quality confidence assessment value, the storage-driven assessment value, and the classification label; filter the candidate storage strategies in the strategy library based on the classification label and the quality confidence assessment value to obtain a subset of candidate strategies corresponding to each type of business data cluster; perform strategy optimization processing based on the storage strategy requirement vector and the subset of candidate strategies to construct the adaptive storage strategy matrix.

[0015] Further, the specific steps of the strategy optimization process are as follows: analyze the matching distance between the storage strategy demand vector and the strategy feature vector of each candidate storage strategy in the candidate strategy subset, and sort the candidate storage strategy subset based on the matching distance; based on a multi-objective optimization method, select the Pareto optimal non-dominated strategy solution set from the sorted candidate strategy subset; and determine the preferred storage strategy and alternative storage strategy corresponding to each type of business data cluster from the non-dominated strategy solution set based on preset rules, so as to generate the adaptive storage strategy matrix.

[0016] The present invention has the following beneficial effects: (1) The management system for business data constructs a full-link data processing mechanism through an intelligent quality perception classification module to achieve accurate capture and correlation matching of data features and quality status. The feature extraction sub-network first separates discrete classification features and continuous numerical features through a feature splitting layer. The embedding layer transforms discrete features into dense embedding vectors. Then, through the fusion layer splicing transformation and the deep coding layer deep mining, business deep coding features that can accurately represent the data value density and structural features are generated. The classification sub-network reduces the dimensionality of deep features and accurately classifies multi-source heterogeneous data into several business data clusters through label allocation and classification output, avoiding feature recognition bias. At the same time, the quality detection sub-network extracts multi-dimensional quality indicators and merges them to generate quality confidence evaluation values ​​through type-aware diagnosis matching of exclusive diagnostic configuration vectors, thereby realizing the quantitative representation of the quality status of each data cluster. This makes the data features and quality status clear and controllable, so as to avoid storage strategy mismatch caused by inaccurate feature and quality matching, thereby improving management efficiency and accuracy.

[0017] (2) The management system for business data extracts a storage demand assessment set for each type of data cluster through the storage-driven assessment module, and generates a storage-driven assessment value that accurately reflects the intensity of storage demand. The strategy fusion reasoning module deeply integrates the assessment value with the quality confidence assessment value and the classification label to construct a comprehensive storage strategy demand vector. On this basis, the system first selects a subset of suitable candidate strategies through the classification label and the quality confidence assessment value, and then selects the preferred and alternative strategies by calculating the matching distance ranking and Pareto optimal non-dominated solution set. Finally, it generates an adaptive storage strategy matrix, thereby effectively balancing the mutual constraints between storage cost, access performance and data reliability, and thus improving the adaptability of storage management measures.

[0018] (3) The management system for business data classifies multi-source heterogeneous data into several business data clusters with distinct characteristics, and then generates an adaptive storage strategy matrix containing preferred and alternative strategies for each data cluster. This avoids the chaotic state of scattered and untargeted storage measures. At the same time, based on accurate classification and quality assessment, high-quality data can obtain sufficient storage resources to ensure business access needs, while low-quality data is matched with a suitable simplified storage strategy, which will not over-occupy resources and realize on-demand allocation of storage resources to complete the full-link processing from data access to storage execution. This reduces manual management costs and ensures the orderliness and rationality of storage management, thereby significantly improving the overall efficiency of business data storage management.

[0019] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0020] Figure 1 This is a block diagram of a management system for business data according to the present invention.

[0021] Figure 2 This is a flowchart illustrating the specific steps involved in generating the business depth-coded feature vector set in a management system for business data according to the present invention.

[0022] Figure 3 This is a flowchart illustrating the specific steps involved in obtaining the adaptive storage strategy matrix in a management system for business data according to the present invention. Detailed Implementation

[0023] Please see Figure 1This invention provides a technical solution: a management system for business data, comprising: a data acquisition and analysis module for acquiring multi-source heterogeneous business data and extracting a set of basic business feature vectors from the multi-source heterogeneous business data; and an intelligent quality perception classification module for inputting the set of basic business feature vectors into a pre-trained quality perception classification network, wherein the quality perception classification network includes a feature extraction sub-network, a classification sub-network, and a quality detection sub-network. The feature extraction subnetwork performs feature transformation and deep mining on the basic business feature vector set to generate a business deep coding feature vector set corresponding to multi-source heterogeneous business data; the classification subnetwork performs collaborative classification on the business deep coding feature vector set to extract a two-dimensional classification label set of multi-source heterogeneous business data and classifies the multi-source heterogeneous business data into several types of business data clusters; the quality detection subnetwork performs quality diagnosis on the business data clusters to obtain the quality confidence assessment value corresponding to each type of business data cluster. The storage-driven evaluation module is used to extract the storage-driven evaluation value corresponding to each type of business data cluster based on each type of business data cluster; the policy fusion inference module is used to perform fusion inference processing on the storage-driven evaluation value and the quality confidence evaluation value to obtain the adaptive storage policy matrix corresponding to each type of business data cluster. The adaptive storage execution module is used to take corresponding storage management measures for each type of business data cluster based on the adaptive storage policy matrix. Specifically, it reads the adaptive storage policy matrix and, for each data cluster defined by a two-dimensional classification label... The system identifies the business data cluster type, parses its corresponding preferred storage policy configuration, and maps the abstract parameters in the policy configuration (such as storage media policy, persistence policy, and lifecycle policy) to a set of specific operation instructions that can be executed by the underlying storage system.

[0024] Specifically, multi-source heterogeneous business data refers to the raw data of multiple business data records generated by the target enterprise during its operation, including but not limited to the following data categories consisting of multiple data records: for example, customer relationship data consisting of customer profile records, communication records, service request records, etc.; transaction process data consisting of order records, payment records, invoice records, etc. The specific steps for extracting the basic business feature vector are as follows: Preprocessing of multi-source heterogeneous business data involves: performing data cleaning, format standardization, and source marking operations on each data record to generate a standardized data copy for feature extraction. This operation is performed only on the copy and does not change the storage content of the original business data. Data cleaning: Missing value handling: Traverse each data field. If the value of the field is empty or null, perform the following judgment and operation: Forward filling: If and only if the data record has an adjacent predecessor record that is strictly sorted by timestamp, and the value of the same field in the predecessor record is not empty, fill the current missing value with the value of the predecessor record. Mean fill: If and only if the field is numeric and does not belong to the time series context, calculate the arithmetic mean of all non-null values ​​of the field in the current data batch and fill it with this mean; Marker retention: For data that does not meet any of the above filling conditions, its null value is retained, and in the feature vector generated for the data, the feature dimension corresponding to the field is set to a preset special value that represents the missing value (such as -999). Outlier Handling: For numeric fields, perform the following operations: Identification: Based on the historical data distribution of the field, calculate its first quartile (Q1) and third quartile (Q3), and define its reasonable value range as [Q1-1.5*IQR, Q3+1.5*IQR], where IQR (interquartile range) = Q3-Q1. Values ​​outside this range are identified as outliers. All identified outliers are corrected to the median of all non-outliers in the current data batch. Format standardization: All date and time strings are identified using a pre-built format parsing library (such as strptime) to recognize their original format (e.g., YYYY-MM-DD HH:MM:SS) and uniformly converted to Unix timestamps (integers) in milliseconds; the character encoding of all text data is detected and converted to UTF-8 encoding format; for currency amount fields, their currency symbols (e.g., ¥, $, €) and units are identified and uniformly converted to the specified base currency (e.g., ¥) according to a pre-built exchange rate table; Source tag: Attach a structured source tag to each data record. The tag should contain at least: a globally unique business data source ID, and the date and batch in which the data was captured. Based on the preprocessed multi-source heterogeneous business data, feature construction is performed to generate an original feature set, which specifically includes: structural type features: parsing the physical storage carrier and logical organization of the data and mapping them to predefined type codes; specifically: if the data originates from a relational database table and has a fixed field pattern, the code is 1; if the data is a self-describing JSON or XML document, the code is 2; if the data is a plain text log or document without a fixed structure, the code is 3; if the data is a binary file such as an image, audio, or video, the code is 4. Data size characteristics: Record the physical storage size in bytes of the data record and calculate its logarithm to base 10 as the first sub-characteristic; record the total number of logical records in the data table or file to which it belongs and calculate its logarithm to base 10 as the second sub-characteristic. Time attribute features: Extract business timestamps (such as order creation time, log generation time) from data records, calculate the difference between the timestamp and the current system time, and convert it into a floating-point number in days; Data source characteristics: The globally unique business data source ID (i.e., which data record it corresponds to) is directly recorded in the source marking step as a classification identifier; Content statistical features: Sampling statistical analysis is performed on the core fields of the data records. For numeric fields (such as amount and quantity), the arithmetic mean and standard deviation of all numeric field values ​​in the record are calculated as two sub-features. For text fields (such as name and description), the total number of different characters in all text fields in the record is counted as a sub-feature. The original feature set is formed based on the above features. The original feature set is standardized and concatenated into a business basic feature vector set in a preset order. Specifically, the two logarithmic sub-features of the data scale feature, the floating-point values ​​of the time attribute feature, and the mean and standard deviation of the content statistics feature, a total of five continuous numerical features, are standardized using Z-score. For the integer encoding of the structural type feature (1, 2, 3, 4) and the ID of the data source feature, one-hot encoding is used to convert them into a sparse binary vector with only the corresponding position set to 1 and the rest set to 0. The features that have undergone Z-score standardization are concatenated with two classification feature vectors that have undergone one-hot encoding in a fixed order: structural type features, data size features (logarithmic size), data size features (logarithmic number of records), time attribute features, data source features, content statistics features, content statistics features, and content statistics features. This process generates a basic business feature vector that corresponds one-to-one with each data record and forms the basic feature vector set.

[0025] In this implementation plan, data is preprocessed by cleaning, standardizing, and labeling to fill in missing values, correct outliers, and unify time, encoding, and currency formats without altering the original data, ensuring data purity and usability. Then, multi-dimensional features such as structure type and data scale are extracted in a targeted manner to match the business data attributes. Finally, a feature vector set is generated through standardization and ordered concatenation, providing accurate and unified input basis for subsequent feature extraction and classification evaluation of the network.

[0026] Specifically, such as Figure 2As shown, the feature extraction sub-network includes a feature splitting layer, an embedding layer, a feature fusion layer, and a deep encoding layer. The specific steps for generating the business deep encoding feature vector set are as follows: In the feature splitting layer, the business basic feature vector set is separated to generate the business discrete classification feature vector set and business numerical feature set corresponding to multi-source heterogeneous business data. Specifically, each feature dimension in the business basic feature vector corresponding to each data record is identified as a discrete classification dimension or a continuous numerical dimension. The business discrete classification feature vector set includes the dimensions of structural type features and data source features corresponding to each data record; The business continuous numerical feature set includes the dimensions of data scale features, time attribute features, and content statistical features corresponding to each data record; In the embedding layer, dense encoding is performed on the discrete classification feature vector set of the business data to generate a discrete feature embedding vector set corresponding to the multi-source heterogeneous business data, specifically as follows: For each data record corresponding to a discrete business classification feature, perform two parallel embedding lookup operations: Structure type feature embedding: Using the structure type feature encoding value of the record as an index, the corresponding row vector is found in a trainable first embedding matrix to obtain the structure type embedding vector. The number of rows in the first embedding matrix is ​​equal to the total number of structure types, and the number of columns is the preset structure type embedding dimension. During the training process, this matrix learns the semantic relationship between different business data structure types in terms of data organization, storage characteristics and processing logic. Data source feature embedding: Using the data source feature ID of the record as an index, the corresponding row vector is found in a trainable second embedding matrix to obtain the data source embedding vector. The number of rows in the second embedding matrix is ​​equal to the total number of unique data sources, and the number of columns is the preset data source embedding dimension. During the training process, this matrix learns the semantic features of different data sources in terms of data quality, update mode and business value. The structure type embedding vector and the data source embedding vector are concatenated along the feature dimension to form an intermediate fusion vector for the data record. This intermediate fusion vector is then input into a trainable fully connected layer for linear projection. This fully connected layer has a preset unified output dimension. By performing a linear transformation on the input vector and adding a bias, a non-linear activation function (such as ReLU) is applied to the result of the linear transformation to generate the final discrete feature embedding vector for the data record. This process learns the deep interaction between the structure type semantics and the data source semantics, for example, specific types of business data often originate from certain specific records, and encodes this cross-dimensional business association into a unified, dense vector representation. All discrete feature embedding vectors are organized according to the corresponding order of the original data records to form a discrete feature embedding vector set. In the feature fusion layer, the discrete feature embedding vector set and the continuous numerical feature vector set of the business are concatenated and transformed to generate the fusion enhanced feature vector set corresponding to the multi-source heterogeneous business data. In the deep encoding layer, the fused enhanced feature vector set undergoes deep mining and mapping processing to output a business deep encoded feature vector set. Specifically, this layer consists of one or more sequentially stacked deep encoding blocks. By simulating complex global dependencies between features and performing high-order nonlinear transformations, each deep encoding block contains a multi-head business association attention sublayer (which allows the model to learn information from different representation subspaces simultaneously and calculate the interaction weights between all feature dimensions within the same data record) and a feedforward neural network sublayer (used for further nonlinear transformations and feature mapping). Residual connections and layer normalization techniques are employed to ensure training stability. Taking the fused enhanced feature vector corresponding to a certain data record as an example: Let the first The input to each coded block is the fused enhanced feature vector. ,Will The input business-related attention sublayer generates multiple sets of query, key, and value representations through trainable linear projection, examining features from different business perspectives. Its core computation involves obtaining attention weights, which reflect the importance of a certain feature dimension (e.g., data source) to another feature dimension (e.g., update time) from a specific business perspective. The computation of a single attention head can be formally represented as: ; in, , , Depend on Through different trainable weight matrices (Query the projected weight matrix). (Key projection weight matrix) (The value projection weight matrix) is projected to obtain, Given the dimension of the key vector, a multi-head mechanism performs multiple such calculations in parallel to capture diverse business association patterns, ultimately culminating in a trainable weight matrix. Project the image to obtain the output representation. The original business feature information is preserved through residual connections and then stably distributed after layer normalization. ; Will (Intermediate feature vector) Input business pattern feedforward sublayer. This sublayer performs independent nonlinear transformations on the features at each location, aiming to combine and enhance the captured business associations to form a higher-level pattern. Its calculation is defined as: ; in, , , , These are trainable parameters specific to the feedforward sublayer of this coding block. Let be a mathematical function, representing the first... The operation performed by the dedicated feedforward neural network sublayer in each encoding block allows this transformation to learn complex rules such as storage priority patterns (abstract features to be learned) if the data is frequently accessed (numerical features) and comes from the core transaction system (discrete semantics). The output of the encoding block is then obtained through residual connections and layer normalization. Repeating these steps, the output vector of the last encoding block is the business depth encoding feature vector obtained after deep mining of the current data record. Each dimension of this business depth encoding feature vector corresponds to a fusion-derived feature of the original specific features, such as: Structure and Source Fusion Features: Derived from the embedding vector fusion of structure type features (encoding 1-4) and data source features (source ID), representing the structure type preference corresponding to data from a specific source (such as core trading systems outputting mostly regular data). The fusion of scale and time characteristics is derived from the fusion of data scale characteristics and time attribute characteristics, representing the relationship between data volume and timeliness (e.g., large-volume data with strong timeliness requires high-performance storage). Content and value integration characteristics: derived from the integration of content statistical characteristics and data source characteristics, representing the correlation between data content complexity and business value (e.g., records with high average transaction amounts have high business value density).

[0027] The specific steps for generating the fused and enhanced feature vector set are as follows: The discrete feature embedding vector set and the continuous numerical feature vector set of business are concatenated in a dimension to generate a business fusion feature vector set corresponding to multi-source heterogeneous business data. Specifically, for each data record, the two vectors are concatenated end to end in the feature dimension direction to generate the business fusion feature vector of the data record. This concatenation operation is performed in parallel on all data records to form a business fusion feature vector set. The business fusion feature vector set is transformed and enhanced to generate a fusion enhanced feature vector set, specifically based on: One fully connected layer ( ≥2, in this implementation example Take 3), perform multi-layer nonlinear transformation, deeply explore the complex relationship between discrete semantic features and continuous numerical features, and extract a higher level of comprehensive business representation. The input business fusion feature vector (denoted as) The input is a first fully connected layer, which has a trainable weight matrix. and bias vector After performing a linear transformation, this layer undergoes a nonlinear mapping using the ReLU activation function to obtain the first layer's output vector. It initially integrates features of different types and introduces nonlinearity; The first layer output vector Input the second fully connected layer, which has a trainable weight matrix. and bias vector Similarly, a linear transformation and ReLU activation are performed to obtain the second layer output vector: Learn more complex cross-feature interactions; The second output vector The input to the third fully connected layer undergoes a linear transformation, but without using a non-linear activation function, to obtain a fused enhanced feature vector (corresponding to each data record). This vector encodes the complex non-linear relationships between different feature dimensions (such as data from a specific source often having a certain scale pattern) and extracts a high-order pattern that represents the overall business status and value of the data.

[0028] In this implementation scheme, the semantic association between structure type and data source is learned through the embedding layer, and then the two types of features are spliced ​​together through the fusion layer. Finally, the global dependency relationship is mined through the deep coding layer. The generated deep coding features integrate multi-dimensional attributes such as structure, scale, and time, which can accurately represent the business association pattern of the data. This provides a reliable feature basis that meets business needs for the accurate clustering of the subsequent classification sub-network and the effective evaluation of the quality detection sub-network.

[0029] Specifically, the classification sub-network includes a feature dimensionality reduction layer, a label assignment layer, and a classification output layer. The specific steps for generating several types of business data clusters are as follows: In the feature dimensionality reduction layer, the feature vector set of deep business encoding is compressed to generate a two-dimensional feature vector set corresponding to multi-source heterogeneous business data, specifically as follows: Taking the business depth encoded feature vector corresponding to a certain data record as an example, this layer uses a trainable weight matrix. With bias vector Perform a linear transformation and output the two-dimensional vector corresponding to the data record. ,in, As the first component of the vector, it is designed to be a continuous measure representing the density of business value. As the second component, it is designed as a continuous measure characterizing structural features; therefore, This represents the projection coordinates of the data in the two-dimensional space of business value and structural characteristics—the horizontal coordinate. Reflecting the value level, the vertical axis... Both reflect the structure type and together characterize the position of data in the business semantic space; The first column vector learns how to... The training objective is to amplify features strongly correlated with value judgments (e.g., scale and time fusion features: small difference between business timestamps and current time, and large data scale, reflecting active and valuable data; content and value fusion features: high mean of numerical fields and semantics of high-value sources, reflecting high-value transactional data; structure and source fusion features: regular structure encoding and core system source, reflecting regular data of core business, etc.) and suppress features that are irrelevant to value or mainly related to structure, thereby consolidating high-dimensional value information into a single value. On the coordinates; The second column vector learns how to focus and combine in a similar way. Features strongly correlated with structure type (e.g., structure and source fusion features: unstructured encoding 3 / 4 and non-core system sources, reflecting loose / fluid structure data; scale and time fusion features: logarithmically large record size and binary file structure encoding 4, reflecting large multimedia file data, etc.) are projected onto... Coordinates, the magnitude of which directly represents the position of the data in the structural type spectrum; This learning is guided by a classification task loss function, ensuring that the final learned mapping guarantees, to the greatest extent possible, that data with similar business value and structural type in the original high-dimensional feature space have similar points on the two-dimensional plane. They are also close to each other; Perform this operation on all data records to obtain the coordinates of each record on the value-structure two-dimensional plane, forming a two-dimensional feature vector set; In the label allocation layer, the two-dimensional feature vector set is discretized and labeled to generate a two-dimensional classification label set corresponding to multi-source heterogeneous business data; In the classification output layer, multi-source heterogeneous business data is classified based on a two-dimensional classification label set, outputting several types of business data clusters, specifically: For each unique two-dimensional category label that appears Create a logical container, iterate through all data records, and assign the records to the corresponding containers based on their labels; Each non-empty container is defined as a business data cluster, which is defined by its label. A unique identifier, containing all original business data records with that tag.

[0030] The specific steps for generating a two-dimensional classification label set are as follows: Normalization of the two-dimensional feature vector set is specifically performed by: statistically analyzing the values ​​of all vectors in the two-dimensional feature vector set. The minimum and maximum values ​​of the axis, and in The minimum and maximum values ​​of the axes, for each two-dimensional eigenvector in the set. Perform min-max normalization; Based on the normalized two-dimensional feature vector set, and combined with the preset business classification granularity (including the number of rows) With column number The attribution determination process is performed to generate a spatial attribution mapping set corresponding to multi-source heterogeneous business data. Specifically, it is based on: and The normalized two-dimensional space (i.e., the unit square [0, 1] × [0, 1]) is uniformly divided into... OK A rectangular grid of columns, with the row index of the grid. The numbers from top to bottom are 1, 2, ... Column index The numbers from left to right are 1, 2, ... Each grid cell is indexed by its row. and column indexes Uniquely identified, denoted as a unit Its covered coordinate range is: For the right and top boundaries, when the coordinates are equal to 1, they belong to the last unit (i.e., or ); Each two-dimensional feature vector after normalization Determine the grid cell to which it belongs and calculate its row index. and column indexes The formula is: ,in This indicates rounding up, specifically when... hour, ;when hour, Therefore, a grid cell identifier is assigned to each two-dimensional feature vector. The correspondence between all two-dimensional feature vectors and their grid cell identifiers is recorded to form a spatial attribution mapping set. This mapping set records which grid cell each data record (by its two-dimensional feature vector index) is assigned to. Label assignment is performed on a two-dimensional feature vector set based on a spatial attribution mapping set to generate a two-dimensional classification label set. Specifically, according to the mapping relationship recorded in the spatial attribution mapping set, a two-dimensional classification label is generated for each data record. This label is directly derived from the row and column index of its corresponding grid cell. This indicates that the label format is... Among them, row index Corresponding to the discrete classification of data in the structural type spectrum dimension (from regular to fluid), column index This corresponds to the discrete classification of data in the dimension of business value density (from low to high). Iterate through all data records and assign their corresponding two-dimensional classification labels in the order they appear in the original data. Extract the data and arrange them in the same order to form a two-dimensional classification label set, where each label indicates the business data type to which the record belongs.

[0031] In this implementation plan, the dimensionality reduction layer focuses on the data value density and structural characteristics, gathers key information to generate two-dimensional vectors, and allows similar data to be adjacent in the two-dimensional plane. After normalization and grid partitioning, two-dimensional labels are accurately assigned, and business data clusters are formed by clustering according to labels. The generated clusters are uniquely identified by labels, with clear characteristics, which can effectively distinguish different types of data, avoid classification homogenization, and also allow the storage strategy to accurately meet the needs of different data clusters, avoiding strategy mismatch problems.

[0032] Specifically, the quality inspection sub-network includes a type-aware diagnostic layer, a quality adaptation layer, and a confidence output layer. The specific steps to obtain the quality confidence assessment value corresponding to each type of business data cluster are as follows: In the type-aware diagnostic layer, matching processing is performed on each type of business data cluster to generate a diagnostic configuration vector corresponding to the business data cluster, specifically as follows: During the training phase, a diagnostic strategy mapping relationship is pre-established based on business management objectives. This relationship explicitly defines different two-dimensional classification labels. The corresponding quality assessment focus and calculation requirements; Obtain the two-dimensional classification labels of this business data cluster. ; Based on predefined diagnostic strategy mapping relationships (i.e., during system deployment or training phases, a diagnostic strategy mapping table is pre-established, and the construction of this table is based on: expert business knowledge: defining different business data types, defined by their two-dimensional classification labels). Characterize the key quality dimensions to focus on; Historical data analysis: Analyze the patterns of various quality problems in historical data to verify and calibrate the above mapping relationship, and determine specific calculation parameters (such as thresholds and rule sets) for each dimension. Online query and generation: At runtime, for a given cluster of business data: Read its two-dimensional classification label ,by (For indexing, query the predefined diagnostic strategy mapping table), for tags Generate or match a structured diagnostic configuration vector, which is a data structure containing specific technical parameters, and explicitly includes at least: The set of quality dimensions to be activated: clearly specify the quality dimensions that need to be calculated for this type of business data cluster (e.g., calculate several or all of the following dimensions: business compliance, integrity, anomaly, and business effectiveness). Specific calculation rules and parameters for each dimension: For each quality dimension that needs to be activated, provide its specific calculation method or parameters. For example, for the business compliance dimension: check whether the data records violate the predefined business rule set, such as: status transition rule: when the order status is shipped, the logistics tracking number field cannot be empty; numerical logic rule: the invoice amount must be equal to the sum of the corresponding order amount; association existence rule: the customer ID must exist in the customer master table. Traverse each record in the cluster, apply all relevant rules for verification, and count the number of records that violate any rule; output metric: business rule violation rate = (number of records that violate the rule / total number of records in the cluster); Complete Dimension: This dimension tracks the degree of missing key fields in the records. It includes fields that must be checked for this type of data, such as [customer name, order amount, creation time]. Missing fields not listed in the table are not counted. It iterates through each record in the cluster, checking if the value of each field in its key field list is NULL or equal to a preset missing value (e.g., -999). Output metric: Missing Rate = (Total number of missing field values ​​ / (Number of key fields × Total number of records in the cluster)). Anomaly Dimensions: Numerical fields: use standard deviation thresholding, with parameters including: standard deviation multiple (e.g., 3σ) or quantile method, with parameters: upper quantile (e.g., 0.99) and lower quantile (e.g., 0.01); Textual fields (for fluid / loose structures): rare value detection can be used, with parameters including: lowest frequency threshold (e.g., <0.1% of total occurrences); Based on historical data or current batch data, calculate the normal value range for each specified field, traverse records within the cluster, and count the number of times its value falls outside the normal range. Output metric: Anomaly density = (total number of detected anomalies / total number of field values ​​examined); Business effectiveness dimension: Timestamp field, specifying the business time field used to calculate effectiveness (such as order creation time, log generation time); Linear decay model (suitable for scenarios where business value decays evenly within a defined period and drops sharply after expiration, such as contract data, promotional activity data): The parameter is the validity period (days, such as 90 days), effectiveness score = max(0, 1-(current time-business timestamp) / validity period); Exponential decay model (suitable for scenarios where business value continuously and rapidly decays over time, but never completely reaches zero, such as behavioral logs and market data): The parameter is the decay coefficient (λ, e.g., 0.01), and the effectiveness score = exp(-λ × (current time - business timestamp)); For each record in the cluster, calculate the time difference between its business timestamp and the current system time (which can be converted to days), and substitute it into the decay model to calculate the score of a single record; Output metric: average effectiveness score of the cluster = (sum of effectiveness scores of all records / total number of records in the cluster); In the quality adaptation layer, the diagnostic configuration vector is extracted in a targeted manner to generate the quality adaptation set corresponding to the business data cluster. Specifically, for each type of business data cluster, the set of quality dimensions to be activated corresponding to that type is read from the diagnostic configuration vector. For dimensions marked as inactive, their output scores will be set to the default value (such as 0). For each activated quality dimension, its predefined calculation function is called, and the specific calculation rules and parameters specified for that dimension in the diagnostic configuration vector are passed in. The calculation is performed on all records of the current business data cluster to generate the business rule violation rate, missing rate, anomaly density, and cluster average effective score corresponding to each type of business data cluster. The data is then standardized and its specific values ​​are mapped between 0 and 1 (except for the default value) to form a quality adaptation set. In the confidence output layer, the quality fit set is fused to output the quality confidence assessment value. Specifically, based on historical data during training, the weight coefficients corresponding to the business rule violation rate, missing rate, anomaly density, and cluster average effective score for various types are calculated (i.e., historical business data is classified to form two-dimensional labels). The data clusters were manually labeled with an overall quality score for each cluster. For each historical data cluster, the raw indicators in four quality dimensions (compliance, completeness, anomalies, and effectiveness) were uniformly calculated and standardized into a scoring vector. Data clusters with the same labels were then assigned a score. All historical clusters are grouped together. For each group of data, linear regression is used to minimize the error between the predicted and labeled values, resulting in an optimal set of initial weights that best fit the overall quality of that data type. The obtained weights are normalized (to a sum of 1) and used as the final weight coefficient vector for that type. This vector is then used in conjunction with the quality fit set to generate the quality confidence assessment value for each type of business data cluster. The calculation formula is as follows: ; in, This is a quality confidence assessment value. For the rate of violation of business rules, The missing rate, It is an abnormal density. The average effective score for the cluster. , , , The weighting coefficients are, in order, the business rule violation rate, the missing rate, the anomaly density, and the average effective score of the cluster.

[0033] The pre-training steps of the quality-aware classification network are as follows: A pre-training dataset is constructed by collecting heterogeneous data from multiple sources across all business types of the target enterprise. Data cleaning, format standardization, source labeling, feature construction, and standardization are completed according to a predetermined process to generate a set of basic business feature vectors for pre-training. Business experts then label each data point with two-dimensional manual labels indicating value density and structural type. After clustering by label, quality confidence values ​​(0-1 intervals) are assigned to each cluster. Simultaneously, diagnostic configuration vectors are preset for each type of label, and a diagnostic configuration vector library is constructed. Subsequently, network parameters are initialized. In the feature extraction subnetwork, the embedding matrices for structure type and data source are initialized as normally distributed random matrices, the parameters of the fully connected layer are initialized as normally distributed, and the bias is set to 0. The projection matrix of the attention sublayer of the deep coding layer and the parameters of the feedforward network are initialized as normally distributed. The weight matrix of the feature dimensionality reduction layer of the classification subnetwork is initialized as normally distributed, and the bias is set to 0. The weight coefficients of the confidence output layer of the quality detection subnetwork are initialized as uniformly distributed. The remaining modules without training parameters load preset rules.

[0034] Multi-task joint training is implemented. In the first stage, the feature extraction sub-network is pre-trained. The input pre-trained feature vector set outputs deep encoded features. The parameters are optimized using contrastive loss to narrow the distance between features with the same label and widen the distance between features with different labels. After multiple rounds of iteration, the parameters with the best clustering accuracy are retained. In the second stage, the parameters of the pre-trained feature extraction sub-network are loaded, and the entire network is jointly trained: the input feature vector is inferred through the entire chain to output classification labels, data clusters and quality confidence values. A joint loss function of cross-entropy loss (measuring classification bias) and mean squared error loss (measuring quality assessment bias) is used. The parameters are updated through AdamW optimizer, gradient clipping is set to prevent explosion, and an early stopping strategy is used to terminate training.

[0035] Finally, the model is validated and saved. The training set and validation set are divided in a 7:3 ratio. The classification accuracy (target ≥90%) and the mean absolute error of quality assessment (target ≤0.05) are used as evaluation indicators. The optimal full network parameters are saved for validation. New business data can be collected quarterly for small-batch fine-tuning to update the model parameters to adapt to new business scenarios and ensure that the network continues to adapt to business needs.

[0036] In this implementation plan, dedicated diagnostic configurations are matched according to data cluster labels, quality dimensions are activated in a targeted manner, and multi-dimensional indicators such as compliance and completeness are accurately calculated. These are then integrated to generate quality confidence assessment values. Combined with full-process pre-training to optimize network parameters, the quality assessment is ensured to be accurate and reliable, avoiding missed detections and misjudgments. This clearly quantifies the quality status of each type of data cluster, allowing subsequent storage strategies to accurately adapt to the needs of different quality data, thereby avoiding insufficient high-quality data resources and excessive resource consumption by low-quality data.

[0037] Specifically, the steps for extracting storage-driven evaluation values ​​are as follows: Based on each type of business data cluster, extract the storage demand evaluation set (including storage resource consumption value and content redundancy value) corresponding to each type of business data cluster. Specifically, for each data record in the business data cluster, take the mean of the first sub-feature and the second sub-feature in the data scale feature, and take the mean of all such results for the business data cluster, and perform maximum and minimum normalization to finally obtain the storage resource consumption value. The higher the value, the greater the pressure on storage space occupied by the data cluster. For each data record in the business data cluster, extract the statistical value of the total number of different characters in the text field from its content statistical features, calculate the arithmetic mean of the above statistical values ​​for all records in the cluster, normalize this mean (map to the 0-1 range), and finally obtain the content redundancy value. The lower the value (closer to 0), the more limited the character set of the data content and the higher the possibility of pattern repetition, that is, the higher the internal redundancy. The storage requirement assessment set is comprehensively processed to generate storage driver assessment values ​​corresponding to each type of business data cluster. Specifically, this involves reading the data from that type of business data cluster. and ,according to Map its value to the [0, 1] interval to obtain the business value density value of this type of business data cluster; A baseline weight (e.g., 0.5 for both) is preset for storage resource consumption and content redundancy. This baseline weight is then adjusted based on the business value density of the business data cluster for that type of business data. ,in, To correct the baseline weight of the processed storage resource consumption value, As the baseline weight for storage resource consumption value, This represents the business value density value. To adjust the coefficient (which can be 0.2 in this example), the baseline weight for the corrected content redundancy value is 1- ; Based on the corrected storage resource consumption value Content redundancy value The baseline weights are analyzed in conjunction with storage resource consumption and content redundancy values, i.e. This allows us to obtain the storage driver evaluation value corresponding to each type of business data cluster. .

[0038] In this implementation plan, based on the scale characteristics and content statistical characteristics of data clusters, the storage resource consumption value and content redundancy value are accurately extracted, and the data's pressure on storage space and inherent redundancy are clearly quantified. At the same time, combined with the business value density to correct the weight, storage-driven evaluation values ​​are generated through collaborative analysis, so that the evaluation results are more in line with the actual storage needs of different data clusters, providing accurate demand references for the subsequent fusion reasoning of storage strategies, and helping to make storage management more reasonable.

[0039] Specifically, such as Figure 3 As shown, the specific steps to obtain the adaptive storage strategy matrix are as follows: Construct a strategy library containing multiple candidate storage strategies, and generate a corresponding strategy feature vector for each candidate storage strategy. Specifically, the strategy library is a set of predefined storage configuration schemes available for the system to choose from, and each candidate storage strategy... Defined by specific configuration parameters in three dimensions: Storage media type (e.g., NVMe SSD, SATA SSD, SAS HDD, SATA HDD, archive tape). Data persistence solutions (e.g., local multiple replicas, cross-availability zone synchronous replication, erasure coding, and single replicas combined with regular backups); Lifecycle management rules (e.g., permanent retention, scheduled migration to cold storage, automatic deletion upon expiration); To perform quantization matching, for each strategy Generate a three-dimensional policy feature vector. ,in: The value ∈ [0, 1] represents the cost-effectiveness score, with a lower value indicating a higher unit storage cost; ∈[0,1] represents the performance rating, with higher values ​​indicating better expected access performance (such as IOPS and latency); ∈[0,1] represents the management complexity score, and the higher the value, the higher the operational complexity of the strategy; These scores are calculated by applying a predefined scoring model to the above configuration parameters. For example, NVMe SSD media will be assigned a higher score. and lower Cross-availability zone replication solutions will significantly improve... ; The predefined scoring model is implemented as follows: Establish a basic scoring table: For each configuration parameter, establish a basic scoring mapping table; Storage Media Baseline Rating: A base cost score is assigned to each media type based on market average price and performance benchmarks. and performance base score ,For example:

[0040] Basic scoring for persistence solutions: A basic score for management complexity is assigned to each solution based on its implementation complexity and resource consumption. and reliability gain ,For example:

[0041] Lifecycle rule basic score: A management complexity adjustment score is set based on the degree of automation and strategy complexity. For example, the rules for permanent retention are simple. =0.1; the timed migration rules are complex. =0.4; Synthetic strategy feature vector: The final feature vector is calculated through weighted synthesis and normalization. , , ; Cost-benefit score The main cost is influenced by the cost of the storage medium. Additionally, high-performance persistence solutions may incur extra costs. The calculation formula is as follows: ; in, , Weighting coefficients (e.g.) =0.8, =0.2), This indicates the cost penalty associated with a high-reliability solution; Performance rating It is mainly determined by the performance of the storage medium. ; Management complexity score It combines the complexity of persistence schemes with the management overhead of lifecycle rules; ; in, , Weighting coefficients (e.g.) =0.7, =0.3); After the strategy library is built, the three scoring dimensions of all strategies are subjected to min-max normalization to ensure that they strictly fall within the interval [0, 1] while maintaining their original relative relationships, thus obtaining... , , ; Based on each type of business data cluster, a classification label corresponding to each type of business data cluster is extracted. Specifically, for each categorized business data cluster, its type is identified by a unique two-dimensional classification label. ; Based on the quality confidence assessment value, storage-driven assessment value, and classification labels, a storage strategy requirement vector is constructed for each type of business data cluster, specifically as follows: For the label is The business data clusters, whose quality confidence assessment values ​​are known. and storage driver evaluation value Calculate a three-dimensional storage strategy requirement vector. To quantify the cluster's demand for storage resources, the calculation formula is as follows: First, the input parameters are normalized to make them , , , Both are in the interval [0, 1]; ; ; ; in, , , , , , The weighting coefficients are adjustable, and the sum of the coefficients in each equation is 1. The design logic is as follows: Cost requirements : The lower, The lower the price, the greater the need for cost control. The higher the willingness (to accept low-cost strategies); Performance requirements : and The higher the level, the greater the demand for access performance. The higher; Management needs The more complex the structure ( (The larger) and the quality confidence assessment value The higher the level, the greater the need for refined management (such as high availability and strong consistency). The higher; Finally, Perform normalization processing; in, , , , , , The steps to obtain it are as follows: Collect a set of historical business data, which has been categorized and has real storage access logs and cost records; Calculate the normalized features for each historical business data cluster: , , , ; Based on the storage logs, label the actual storage cost efficiency (denoted as ) for each historical data cluster. ), actual access performance requirements (denoted as Actual management input intensity (referred to as) These labeled values ​​are obtained by statistically analyzing the unit storage cost, average access frequency, and number of maintenance interventions for the cluster data, and then normalized to the [0, 1] interval. Fitting weight coefficients: ; ; ; Cost demand weight ( , ): By solving the linear regression problem The least squares method was used to fit the result. and The initial values ​​are then normalized to make their sum equal to 1; Performance requirement weight ( , ): By solving The results were obtained through fitting. Management demand weight ( , ): By solving The results were obtained through fitting. Based on classification labels and quality confidence assessment values, candidate storage strategies in the strategy library are filtered to obtain a subset of candidate strategies corresponding to each type of business data cluster, specifically: According to the label It searches for the allowed range of policy feature values ​​in a predefined mapping table, for example, if ≤0.25 (normalized data) and If the value is ≥0.7 (high value), then only retain [the remaining values]. A strategy with a performance factor of ≥0.6 (high performance); if If the value is ≥0.75 (unstructured data), then certain specific strategies that only apply to relational databases will be excluded. Quality constraint filtering: based on quality confidence assessment value Apply quality threshold rules, for example: like A value <0.3 indicates very poor data quality; all such data should be excluded. >0.7 (high management complexity) and A strategy with a cost of <0.3 (extremely high) avoids over-investing in low-quality data; like A value >0.8 indicates high data quality and must be retained. A strategy with a value of ≥0.5 (providing a certain level of persistence); Through the above filtering, a subset of candidate strategies corresponding to the current business data cluster is obtained. ; Based on the storage strategy requirement vector and the candidate strategy subset, strategy optimization is performed to construct an adaptive storage strategy matrix.

[0042] The specific steps of the strategy optimization process are as follows: Analyze the matching distance between the storage strategy demand vector and the strategy feature vector of each candidate storage strategy in the candidate strategy subset, and sort the candidate storage strategy subset based on the matching distance. For a subset of candidate strategies Each candidate storage strategy Calculate its strategy feature vector Storage strategy requirements vector of the current business data cluster Euclidean distance between As the matching distance: ; The smaller the value, the better the strategy. The higher the degree of matching with the needs of the current data cluster, the better. For all ∈ Sort in ascending order; Based on a multi-objective optimization method, a Pareto-optimal set of non-dominated policy solutions is selected from the ranked subset of candidate policies. Specifically, each candidate policy is considered as a solution to be optimized on three objectives: minimizing the cost-benefit difference. Minimize performance level differences Minimize differences in management complexity ; Using the fast non-dominated sorting algorithm to screen the Pareto front: for two strategies and ,like Not inferior to in all three objectives (Right now ≤ , ≤ , ≤ ), and is strictly superior to at least one objective. Then it is called Dominate ; Traverse the subset of candidate strategies to find all strategies that are not dominated by any other strategy. These strategies constitute the first Pareto front, i.e., the set of solutions for non-dominated strategies. The strategies in this set achieve the best trade-off between cost, performance, and management complexity, and are difficult to improve simultaneously. Based on preset rules, the preferred storage strategy and (at least one) alternative storage strategy corresponding to each type of business data cluster are determined from the solution set of non-dominated strategies to generate an adaptive storage strategy matrix, which is as follows: First, define the ideal point. It represents the theoretically optimal solution (but in reality, no single strategy may be able to achieve it). Then, calculate the non-dominated solution set. Each candidate strategy To the ideal point Euclidean distance (Other proximity measures may also be used): ; Determine the preferred storage strategy: Select Minimal Strategy This is the preferred storage strategy for this type of business data cluster because it is the solution that comes closest to the ideal point under multiple objective trade-offs. Determine alternative storage strategies: Select One or two strategies with values ​​second only to the preferred strategy can be selected as alternative storage strategies. Alternatively, strategies that perform exceptionally well on a single objective (such as...) can also be chosen. or (Minimum) but Suboptimal strategies are added to the alternative list to address potential changes in future business needs. Construct an adaptive storage strategy matrix, which is a logical table, where: Rows: Two-dimensional category labels for each type For indexing, there is a corresponding business data cluster type; Column / Content: Each row records the data cluster corresponding to this type: Preferred strategy ID and its complete configuration parameters (media, persistence, lifecycle); List of alternative strategy IDs; Policy change trigger threshold, for example, when the cluster A drop of more than 20% or When the value increases by more than 30%, a policy reassessment is triggered.

[0043] In this implementation plan, a quantitative strategy library is first constructed, and a three-dimensional feature vector is generated for each strategy. Then, a requirement vector is generated by combining the classification label, quality and driving evaluation value of the data cluster. The optimal strategy is selected through a two-layer screening of labels and quality and Pareto optimization. Finally, a matrix containing the preferred strategy, alternative strategies and change thresholds is generated. This balances the needs of cost, performance and management dimensions, and can cope with business changes. It avoids the adaptation shortcomings of traditional static strategies and makes the storage solution accurately meet the actual needs of various data clusters.

[0044] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0045] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A management system for business data, characterized in that, include: The data acquisition and analysis module is used to acquire multi-source heterogeneous business data and extract the business basic feature vector set of the multi-source heterogeneous business data; The intelligent quality perception classification module is used to input the business basic feature vector set into a pre-trained quality perception classification network, which includes a feature extraction sub-network, a classification sub-network, and a quality detection sub-network. The feature extraction sub-network performs feature transformation and deep mining processing on the business basic feature vector set to generate a business deep coding feature vector set corresponding to multi-source heterogeneous business data. The classification sub-network performs collaborative classification processing on the service deep coding feature vector set, extracts the two-dimensional classification label set of the multi-source heterogeneous service data, and classifies the multi-source heterogeneous service data into several types of service data clusters. The quality inspection sub-network performs quality diagnosis processing on the service data clusters to obtain the quality confidence assessment value corresponding to each type of service data cluster. The storage driver evaluation module is used to extract the storage driver evaluation value corresponding to each type of business data cluster based on each type of business data cluster. The strategy fusion inference module is used to perform fusion inference processing on the storage driver evaluation value and the quality confidence evaluation value to obtain an adaptive storage strategy matrix corresponding to each type of business data cluster. The adaptive storage execution module is used to take corresponding storage management measures for each type of business data cluster based on the adaptive storage strategy matrix.

2. The management system for business data according to claim 1, characterized in that, The specific steps for extracting the basic feature vector of the business are as follows: Preprocess multi-source heterogeneous business data; Based on the preprocessed multi-source heterogeneous business data, feature construction processing is performed to generate an original feature set; The original feature set is standardized and then concatenated into a business basic feature vector set in a preset order.

3. The management system for business data according to claim 1, characterized in that, The feature extraction subnetwork includes a feature splitting layer, an embedding layer, a feature fusion layer, and a deep encoding layer. The specific steps for generating the business deep encoded feature vector set are as follows: In the feature splitting layer, the business basic feature vector set is separated to generate a business discrete classification feature vector set and a business numerical feature set corresponding to multi-source heterogeneous business data. In the embedding layer, the discrete classification feature vector set of the business is densely encoded to generate a discrete feature embedding vector set corresponding to multi-source heterogeneous business data. In the feature fusion layer, the discrete feature embedding vector set and the continuous numerical feature vector set of the business are concatenated and transformed to generate a fusion enhanced feature vector set corresponding to multi-source heterogeneous business data. In the deep coding layer, the fusion enhancement feature vector set is subjected to deep mining and mapping processing to output the business deep coding feature vector set.

4. The management system for business data according to claim 3, characterized in that, The specific steps for generating the fused enhanced feature vector set are as follows: The discrete feature embedding vector set and the continuous numerical feature vector set of the business are concatenated dimensionally to generate a business fusion feature vector set corresponding to multi-source heterogeneous business data. The business fusion feature vector set is transformed and enhanced to generate the fusion enhanced feature vector set.

5. The management system for business data according to claim 1, characterized in that, The classification sub-network includes a feature dimensionality reduction layer, a label assignment layer, and a classification output layer. The specific steps for generating several types of business data clusters are as follows: In the feature dimensionality reduction layer, the business depth encoding feature vector set is subjected to dimensionality compression processing to generate a two-dimensional feature vector set corresponding to multi-source heterogeneous business data; In the label allocation layer, the two-dimensional feature vector set is discretized and labeled to generate a two-dimensional classification label set corresponding to multi-source heterogeneous business data. In the classification output layer, multi-source heterogeneous business data are classified based on the two-dimensional classification label set, and several types of business data clusters are output.

6. The management system for business data according to claim 5, characterized in that, The specific steps for generating the two-dimensional classification label set are as follows: The two-dimensional feature vector set is normalized. Based on the normalized two-dimensional feature vector set, and combined with the preset business classification granularity, the spatial attribution mapping set corresponding to multi-source heterogeneous business data is generated. The two-dimensional feature vector set is labeled based on the spatial attribution mapping set to generate the two-dimensional classification label set.

7. The management system for business data according to claim 1, characterized in that, The quality detection sub-network includes a type-aware diagnostic layer, a quality adaptation layer, and a confidence output layer. The specific steps for obtaining the quality confidence assessment value corresponding to each type of business data cluster are as follows: In the type-aware diagnostic layer, each type of business data cluster is matched to generate a diagnostic configuration vector corresponding to the business data cluster. In the quality adaptation layer, the diagnostic configuration vector is extracted in a targeted manner to generate a quality adaptation set corresponding to the business data cluster. In the confidence output layer, the quality fit set is fused to output the quality confidence assessment value.

8. The management system for business data according to claim 1, characterized in that, The specific steps for extracting the storage driver evaluation value are as follows: Based on each type of business data cluster, extract the storage requirement assessment set corresponding to each type of business data cluster; The storage requirement assessment set is processed to generate storage-driven assessment values ​​for each type of business data cluster.

9. The management system for business data according to claim 1, characterized in that, The specific steps to obtain the adaptive storage strategy matrix are as follows: Construct a policy library containing multiple candidate storage policies, and generate a corresponding policy feature vector for each candidate storage policy; Based on each type of business data cluster, extract the corresponding category label for each type of business data cluster; Based on the quality confidence assessment value, the storage driver assessment value, and the classification label, construct a storage strategy requirement vector corresponding to each type of business data cluster; Based on the classification labels and the quality confidence assessment values, candidate storage strategies in the strategy library are filtered to obtain a subset of candidate strategies corresponding to each type of business data cluster; Based on the storage strategy requirement vector and the candidate strategy subset, strategy optimization is performed to construct the adaptive storage strategy matrix.

10. The management system for business data according to claim 9, characterized in that, The specific steps of the strategy optimization process are as follows: Analyze the matching distance between the storage strategy demand vector and the strategy feature vector of each candidate storage strategy in the candidate strategy subset, and sort the candidate storage strategy subset based on the matching distance; Based on a multi-objective optimization method, Pareto optimal non-dominated policy solutions are selected from the sorted subset of candidate policies. Based on preset rules, the preferred storage strategy and alternative storage strategy corresponding to each type of business data cluster are determined from the non-dominated strategy solution set to generate the adaptive storage strategy matrix.

Citation Information

Patent Citations

  • Data asset quality scoring method and system based on multi-dimensional features

    CN120125105A

  • Apparatus and method for assessing data quality for text analysis

    KR102019207B1