An intelligent data management system
Through the data acquisition, decoupling analysis, and map generation modules of the intelligent data management system, the problems of inaccurate field identification and low risk identification accuracy of multi-source heterogeneous data have been solved, realizing intelligent support for high-quality data analysis and risk assessment.
Patent Information
- Application Number
- CN202511250895.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-03
AI Technical Summary
In existing technologies, the collection, processing and fusion of multi-source heterogeneous data suffer from problems such as inaccurate field identification, unclear semantics, and loose fusion structure, making it difficult to meet the needs of high-quality data analysis. Furthermore, the lack of systematic modeling of field behavior disturbances, structural dependencies and risk propagation paths leads to low accuracy in risk identification results.
An intelligent data management system is adopted, including a data acquisition module, a decoupling analysis module, a graph generation module, and a behavior analysis module. Data is collected through a multi-source access controller, and mutual information indicators are calculated using field discretization and kernel density estimation. An attribute fusion graph structure is constructed to generate a risk control cognitive graph. Combined with action execution logs, field behavior is analyzed to construct a comprehensive quality score.
It improves the accuracy of semantic recognition and structural fusion of data fields, enhances the granularity and precision of data risk assessment, realizes the linkage between intelligent assessment and compression processing of data quality, and meets the needs of high-quality data analysis.
Smart Images

Figure CN120763566B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent data management, in particular to an intelligent data management system. BACKGROUND
[0002] With the development of big data, artificial intelligence and financial technology, the value of data in enterprise operation, risk control and regulatory compliance is increasingly prominent. Data management systems gradually evolve from traditional data storage and calling to intelligent data recognition, analysis and evaluation. Especially in scenarios involving high-dimensional structured data, the correlation between fields, behavior stability and risk sensitivity become important basis for evaluating data quality and risk level.
[0003] In the prior art, the collection, processing and fusion of multi-source heterogeneous data still face problems such as inaccurate field recognition, unclear semantics and loose fusion structure, which are difficult to meet the needs of high-quality data analysis. At the same time, in the field of data risk control, there is a lack of systematic modeling methods for field behavior disturbance, structure dependence and risk propagation path, resulting in low accuracy of risk identification results. In addition, data quality evaluation usually relies on missing rate or consistency rules, which is difficult to fully reflect the comprehensive quality level of fields in the structure, semantics and risk dimensions.
[0004] Therefore, there is an urgent need for a data management system that can support multi-dimensional evaluation, has structure perception ability and risk reasoning ability, to realize high-quality modeling of complex data and intelligent risk control support. SUMMARY
[0005] Based on the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide an intelligent data management system to solve the above technical problems.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical scheme: an intelligent data management system, comprising: a data acquisition module, a decoupling analysis module, a graph generation module, a behavior analysis module, and a quality evaluation module;
[0007] The data acquisition module: uses a multi-source access controller to acquire data information, pre-processes the data information to obtain a data set, and extracts field information based on the data set;
[0008] The decoupling analysis module: calculates mutual information indicators based on field information through field discretization and kernel density estimation, and calculates information bonding degree combined with disturbance change rate after field fusion;
[0009] The graph generation module: constructs an attribute fusion graph structure based on field information, calculates node embedding vectors through a graph convolution network, constructs risk mapping skewness combined with path disturbance directionality features in the fusion graph, and generates a risk control cognitive graph;
[0010] The behavior analysis module: collects action execution logs, extracts field information behavior sequences in a historical time window, calculates field disturbance response degree, and constructs field behavior interference factors;
[0011] The quality evaluation module: based on field information, extracts key field missing proportion and information entropy to calculate information completeness index, and combines information cementation degree, risk mapping skewness and field behavior interference factor to construct comprehensive quality score.
[0012] The application further provides that the data acquisition module comprises:
[0013] The multi-source access controller collects multi-modal data information, and the data information includes transaction flow data, user behavior trajectory data, KYC information and document data.
[0014] The rule-driven structured data cleaning technology is used for pre-processing the data information, and the pre-processing includes field renaming, field type unification, missing value completion, abnormal value correction and time format standardization, and a data set is generated.
[0015] The field recognition algorithm based on machine learning is applied to the data set to extract fields, and field information is obtained by integration.
[0016] The application further provides that the decoupling analysis module comprises:
[0017] Based on the field information, the plurality of fields in the data set are discretized, the probability distribution of the field value is estimated by using the kernel density estimation method, and based on the joint probability distribution and the marginal distribution between the fields after discretization, the mutual information index representing the structural dependence relationship between the field pairs is calculated.
[0018] The application further provides that, based on the field information, a field fusion operation is performed on any two fields, and based on the disturbance change rate after field fusion, the information cementation degree representing the field structure collaboration degree is generated in combination with the mutual information index.
[0019] The application further provides that the graph generation module comprises:
[0020] Based on the field information, an attribute fusion graph structure is generated by using a graph structure construction tool, the nodes in the attribute fusion graph structure represent field entities, and the edges represent the fusion relationship between the fields.
[0021] Based on the attribute fusion graph structure, the graph convolution network is used for feature aggregation and representation learning of each node, and the node embedding vector representing the field structure and semantic association is generated.
[0022] The application is further configured to utilize a graph structure analysis method to construct a multi-hop path set in an attribute fusion graph structure, and utilize a path disturbance awareness mechanism to extract path disturbance directionality features.
[0023] Based on the path disturbance directionality features, the change trend of the multi-hop path in the disturbance propagation process is analyzed, and a risk mapping skewness index is constructed.
[0024] The application is further configured to assign risk labels to the nodes in the attribute fusion graph structure according to the risk mapping skewness index by using a threshold segmentation method, wherein the risk labels include low risk, medium risk and high risk.
[0025] All nodes with a high risk label are screened, and the context adjacent nodes and paths of the nodes are extracted to construct a risk control cognitive graph.
[0026] The application is further configured that the behavior analysis module comprises:
[0027] The action execution log saved by the collection system comprises fields, operation events and timestamps;
[0028] Based on the field information, the field behavior sequence with a clear time sequence in the action execution log within a preset historical time window is extracted;
[0029] Based on the operation events and timestamps of the field behavior sequence, the field behavior interference factor for representing the stability and intervention sensitivity of the field is constructed by calculating the change trend of the features of the field before and after the operation events.
[0030] The application is further configured that the quality evaluation module comprises:
[0031] Based on the field information, the missing proportion of the corresponding records of the preset key field in the data set is extracted, and the information entropy index calculated based on the field value distribution is combined to generate the information completeness index for measuring the information coverage rate and content richness of the field.
[0032] The application is further configured to weight and fuse the information completeness index, the information cementation degree, the risk mapping skewness and the field behavior interference factor to construct a comprehensive quality score for representing the field structure rationality, semantic adhesion, behavior stability and risk correlation degree.
[0033] Based on the quality score, a threshold segmentation strategy is used to divide the quality level of each data in the data set, and the data records with a comprehensive quality score higher than a preset threshold are identified as high-quality data.
[0034] According to the requirements of the supervision agreement, the high-quality data is applied to generate a compressed summary report by using a data compression template.
[0035] The application provides a kind of intelligent data management system, the system is by: data acquisition module: utilize multi-source access controller to collect data information, data information is preprocessed to obtain data set, field information is extracted based on data set;Decoupling analysis module: according to field information, mutual information index is calculated by field discretization and kernel density estimation, and information cementation degree is calculated by combining the disturbance change rate after field fusion;Atlas generation module: attribute fusion graph structure is constructed according to field information, node embedding vector is calculated by graph convolution network, risk mapping skewness is constructed by combining path disturbance directionality features in fusion graph, and risk control cognitive atlas is generated;Behavior analysis module: action execution log is collected, behavior sequence of field information in historical time window is extracted, field disturbance response degree is calculated, and field behavior interference factor is constructed;Quality evaluation module: information completeness index is calculated based on field information extraction key field missing ratio and information entropy, and comprehensive quality score is constructed by combining information cementation degree, risk mapping skewness and field behavior interference factor, the beneficial effects include:
[0036] Improve the accuracy of data field semantic recognition and structure fusion: by introducing field behavior interference factor and information cementation degree, the semantic coupling relationship and structure dependence characteristics between fields in multi-source data are deeply modeled, and the quality and stability of field mapping and fusion are significantly improved.
[0037] Enhance the granularity and accuracy of data risk assessment: combined with risk mapping skewness index, the dynamic response relationship between fields and high-risk labels is constructed, so that the system can identify potential risk paths and field abnormal disturbance, and improve the risk discovery ability and the agility of data governance.
[0038] Realize the linkage of intelligent evaluation and compression processing of data quality: the quality level of structured data records is divided by fusing information completeness index and comprehensive quality score mechanism, and compression summary report is automatically generated based on high-quality data, which takes into account data utilization efficiency and regulatory compliance requirements.
[0039] The above description is only a summary of the technical solutions of the application, in order to more clearly understand the technical means of the application, which can be implemented according to the content of the specification, and in order to make the above and other purposes, features and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creating laborious work. In the drawings:
[0041] Figure 1 FIG. 1 shows a schematic diagram of a structure of an intelligent data management system according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0042] Other advantages and effects of the present application can be easily understood by those skilled in the art from the description of the present application. The present application can also be implemented or applied by other different specific embodiments, and the details in the description can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, but not for limiting the protection scope of the present application.
[0043] It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present application, and only the components related to the present application are shown in the diagrams, but not the number, shape and size of the components when actually implemented. The shapes, number and proportions of the components when actually implemented can be arbitrarily changed, and the layout pattern of the components can be more complex.
[0044] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present application, however, it is obvious for those skilled in the art that the embodiments of the present application can be implemented without these specific details, and in other embodiments, the known structures and devices are shown in the form of block diagrams instead of details, to avoid making the embodiments of the present application difficult to understand.
[0045] Embodiment:
[0046] An intelligent data management system, as shown in FIG. 1, comprises: Figure 1
[0047] Data acquisition module: data information is acquired by using a multi-source access controller, and a data set is obtained by preprocessing the data information, and field information is extracted based on the data set;
[0048] Decoupling analysis module: mutual information indicators are calculated by field discretization and kernel density estimation according to the field information, and information cementation is calculated by combining the disturbance change rate after field fusion;
[0049] Graph generation module: an attribute fusion graph structure is constructed according to the field information, node embedding vectors are calculated by a graph convolution network, risk mapping skewness is constructed by combining the path disturbance directionality features in the fusion graph, and a risk control cognitive graph is generated;
[0050] Behavior analysis module: action execution logs are collected, behavior sequences of the field information in a historical time window are extracted, field disturbance response degrees are calculated, and field behavior interference factors are constructed;
[0051] The quality evaluation module: based on field information extraction key field missing proportion and information entropy calculation information completeness index, combined with information cementation degree, risk mapping skewness and field behavior interference factor to build comprehensive quality score.
[0052] The application is further provided with the data acquisition module comprising:
[0053] Through the multi-source access controller, multi-modal data information is collected, and the data information includes transaction flow data, user behavior trajectory data, KYC information and document data.
[0054] Based on the rule-driven structured data cleaning technology, the data information is preprocessed, and the preprocessing includes field renaming, field type unification, missing value completion, abnormal value correction and time format standardization, and a data set is generated.
[0055] The field extraction is applied to the data set by a machine learning-based field recognition algorithm, and the field information is integrated. Specifically, the system first accesses data sources from different platforms, formats and protocols through a multi-source access controller, including user uploaded documents, log servers, business databases, etc. Specifically, KYC information is obtained by uploading user identification or form documents, and key information fields are extracted by OCR recognition technology; user behavior track data is uploaded in real time by a log server, recording user operation paths and behavior events; transaction flow data is synchronized from a business database at regular intervals; document data such as contracts or audit forms is converted into data records after structured extraction; all original data is uniformly accessed to the system and format compatibility and data merging are completed to form a preliminary original data information set. Subsequently, the system performs rule-driven structured data cleaning on the original data, mainly including the following five processing operations: field renaming: uniformly naming fields with different sources but the same semantics, for example, "cust_name", "user_name" and "name" are uniformly named as "user name"; field type unification: standardizing the data type of the field, such as converting the date "2021 / 12 / 12" to the standard date format "2021-12-12", and converting the string "123" representing the numerical value to the integer 123; missing value completion: according to the field rule library or the default value strategy set, the missing items are completed, for example, if the transaction time is missing, the record creation time is used to fill in or the set date is filled in to ensure that the data is not empty while distinguishing the missing data; abnormal value correction: identifying and repairing the data with obvious logical abnormalities, for example, identifying the absolute value of the transaction amount with a negative value or marking it as abnormal to prevent misleading subsequent analysis; time format standardization: unify the different time representation formats in multiple sources to the international standard time format: such as ISO8601, to ensure the consistency and parsability of the time field. After cleaning and preprocessing, the system forms a data set with consistent structure and clear semantics. Next, a machine learning-based field recognition algorithm is applied to the data set to automatically extract key field information. This stage of the algorithm identifies fields with clear semantics such as "transaction amount", "user identifier", "behavior event" through feature matching, context analysis and pattern recognition, and labels and integrates these fields. Finally, a unified field information set is constructed to provide a standardized input data structure for subsequent processing steps.
[0056] The application further provides that the decoupling analysis module comprises:
[0057] Based on field information, a plurality of fields in the data set are discretized, and a kernel density estimation method is used to estimate the probability distribution of the field values; based on the joint probability distribution and the marginal distribution between the discretized fields, mutual information indicators representing the structural dependency relationship between the field pairs are calculated. Specifically, the mutual information indicator is a statistical quantity that measures the strength of the association between two variables, which is used to evaluate whether the information of one field depends on another field, and can reveal whether there is a structural or semantic connection between the fields, and can reveal the structural coupling relationship between the fields; it is used to identify redundant fields, isolated fields or potential semantic dependencies, and is one of the calculation bases of information cohesion degree. The higher the mutual information, the more information shared between the two fields, and the stronger the correlation. Discretization is to convert continuous fields into discrete values or buckets through equal frequency binning or equal interval binning, which is used to improve the robustness of mutual information calculation and provide a more stable estimation basis for the kernel density estimation method. The kernel density estimation method is a non-parametric method used to estimate the probability distribution of the field, which is often used to replace the histogram to provide a smooth probability density function estimation, and is suitable for continuous variables, and is used to estimate the probability distribution part in mutual information. Mutual information indicator calculation logic: wherein, is the mutual information indicator; is a field, representing field information Two fields in field information is a set of key field information automatically extracted by a machine learning-based field recognition algorithm; is a probability distribution function; and are field values, i.e. and are specific values of and , for example: if the field “ =“region” has field values: {“Beijing”, “Shanghai”, “Guangzhou”}”, and the field “ =“occupation” has field values: {“doctor”, “teacher”, “programmer”}”, then
[0058] The application is further configured to perform a field fusion operation on any two fields based on field information, generate information cementation degree for representing field structure cooperation degree based on disturbance change rate after field fusion and mutual information index. Specifically, the information cementation degree represents structural consistency and semantic adhesion degree between field pairs in a field set, reflects consistency, stability and structural coupling strength after data field fusion, is used for measuring rationality and structure quality of field fusion, is an important dimension of data quality evaluation, can assist in identifying structure isolation, semantic dispersion or redundant fields, is one of input indexes of comprehensive quality score. Information cementation degree calculation logic: wherein, is information cementation degree; is field information, specifically, the ith data of field information, is the size of the field set; is field fusion, used for merging two fields into a new composite field, simulating possible interaction between fields through splicing; is disturbance change rate, the change amplitude of the fusion field is obtained by calculating the gradient change of the value after field fusion in the data dimension, and is used for reflecting the sensitivity after field coupling; is the two norm of disturbance response, used for measuring the stability of the field fusion in different data records; is a small constant, used for avoiding the problem of denominator explosion caused by zero or minimum value of gradient.
[0059] The application is further configured that the atlas generation module comprises:
[0060] generate an attribute fusion graph structure based on field information using a graph structure construction tool, the nodes in the attribute fusion graph structure represent field entities, and the edges represent the fusion relationship between fields;
[0061] Based on the attribute fusion graph structure, a graph convolution network is used to aggregate and represent learning of each node to generate node embedding vectors representing field structure and semantic association. Specifically, first, an attribute fusion graph is constructed using a graph structure construction tool, including NetworkX, Neo4j, DGL or PyG. Each node in the attribute fusion graph corresponds to a specific field entity, such as "name", "region", "occupation", and each edge represents a certain fusion relationship between fields, such as semantic similarity, co-occurrence frequency, hierarchical structure, etc. Then, according to the connection relationship and field information between the nodes in the attribute fusion graph, each field node is assigned an initial feature vector, which is generated by the BERT, Word2Vec, FastText node initial feature generation tool to generate text embedding of the field name. Subsequently, based on the attribute fusion graph and the initial feature vector, a graph neural network model such as GCN, GAT or GraphSAGE is introduced to perform feature aggregation and representation learning. Through multiple rounds of message passing mechanism, the graph neural network integrates the feature information of adjacent nodes to dynamically update the representation of each field node, and finally generates node embedding vectors that capture the field structure characteristics and semantic association.
[0062] The application further provides that a multi-hop path set is constructed in the attribute fusion graph structure using a graph structure analysis method, and a path disturbance directionality feature is extracted using a path disturbance perception mechanism.
[0063] Based on the path perturbation directionality feature, the change trend of multi-hop path in the perturbation propagation process is analyzed, and the risk mapping skewness index is constructed. Specifically, the risk mapping skewness index is used to measure the deviation degree of the potential risk of a node in the attribute fusion graph after the multi-hop path propagation, which is a comprehensive parameter for measuring the influence of path perturbation propagation on the node state. The construction process of the risk mapping skewness index is as follows: first, for each node in the graph, the node embedding vector is generated by using the graph neural network model. In order to improve the stability, multiple vectors are obtained by using multiple sampling methods, and the average value of these vectors is calculated to represent the typical embedding feature of the node. Then, all nodes considered as "normal" are extracted from the historical samples, and the center point of their embedding vectors is calculated as the supervision center vector. The difference between the average embedding vector of the node and the supervision center vector is compared to obtain the structural deviation degree as the first evaluation. Then, starting from the current node, the path set within all fixed hops is constructed, which records the direction and path of information propagation between nodes. For each path, the change direction of the embedding value of each node in the path is analyzed, and the directional deviation index of the path perturbation is calculated, and the result is normalized by a standardization function. The perturbation values of all paths are averaged to represent the average perturbation trend experienced by the node in the information propagation process, which is the second evaluation. Finally, the structural deviation degree and the path perturbation trend are combined by weighting to generate the final risk mapping skewness index, which is used to measure the potential risk level of the node. The calculation logic of the risk mapping skewness index is as follows: wherein, is the risk mapping skewness index; is the node embedding vector; is the attribute fusion graph structure; is the node embedding vector mean, which is the average value of all node embedding vectors in the attribute fusion graph structure corresponding to the i-th data; is the supervision center vector, which is obtained by calculating the center point of the embedding vectors of all nodes considered as "normal" from the historical samples; is the norm type, which is set to 2 here, representing the Euclidean norm; is the structural deviation, which is the structural deviation of the data structure from the supervision center, the larger the distance, the more the data structure deviates from the standard, and the higher the risk; is the multi-hop path set, which is the possible multi-hop path between fields starting from i, and the multi-hop path set records all related paths; is the k-th path in the multi-hop path set, which is a path chain composed of multiple field nodes, and the path reflects the logical connection relationship between fields; is the sigmoid function, which is used to convert the path deviation directionality value to a standard risk score between 0 and 1; For path deviation directionality, by The calculation is obtained, For the angle between the embedding vectors of two nodes in the jth path segment, the node embedding vectors of the two nodes on the path are used to evaluate whether the path direction deviates by cosine similarity or angle calculation, thereby constructing the path deviation directionality index.
[0064] The application further provides that, according to the risk mapping skewness index, a threshold segmentation method is used to assign risk labels to the nodes in the attribute fusion graph structure, the risk labels including: low risk, medium risk and high risk.
[0065] All nodes with a high risk label are screened, and the context adjacent nodes and paths of the nodes are extracted to construct a risk control cognitive graph. Specifically, the risk control cognitive graph is a local subgraph extracted from the original graph structure based on the risk assessment result, and contains high-risk nodes and their adjacent context information, i.e. adjacent nodes and paths. After calculating the risk mapping skewness index of all nodes, a threshold segmentation method is used to classify the nodes; by presetting two thresholds, the entire score interval is divided into three risk level regions, corresponding to low risk, medium risk and high risk respectively. The thresholds are set by quantile points of data distribution, historical experience or optimization objective function. All nodes evaluated as high risk are screened, and their structural context is expanded outward, extracting their directly connected adjacent nodes and the paths they participate in. This process not only contains the connection relationship of the graph structure, but also contains the attribute coordination mode embodied in the path; finally, the high-risk nodes and their adjacent structures are unified and sorted to construct an independent subgraph, i.e. a risk control cognitive graph, which shows the upstream and downstream logic of the high-risk nodes, the associated paths and the potential influencing factors, facilitating precise intervention and tracking analysis by supervisors.
[0066] The application further provides that the behavior analysis module comprises:
[0067] The action execution log saved by the collection system comprises: fields, operation events and timestamps;
[0068] Based on the field information, the field behavior sequence with a clear time sequence in the action execution log within a preset historical time window is extracted;
[0069] Based on the operation event and timestamp of the field behavior sequence, a field behavior interference factor for representing the stability and intervention sensitivity of the field is constructed by calculating the change trend of the characteristics of the field before and after the operation event. Specifically, the field behavior interference factor is an important parameter index for measuring the "sensitivity" and "stability" change trend of a certain field in the system operation process, which can be understood as: whether a certain field is frequently intervened, such as frequently modified, approved, and whether these intervention behaviors have a significant influence on the running, judgment or decision of the system; the higher the value of this factor, the more frequently the field is operated in the time window, or the operation has a great influence on the behavior result; and the lower the value, the more stable the field is, or the sensitivity to the operation result is lower. Field behavior interference factor calculation logic: wherein, is the field behavior interference factor; is a preset historical time window length, which is used for evaluating the total time step of the behavior field interference relationship, and can also be understood as the maximum time span in the time window, and the default is 30 days; is the operation event, which is the operation event at the time point t collected in the action execution log; is the field value, which specifically represents the value of the s-th field in the i-th data, wherein i is the data object number, and s is the index of the field value on the data object; is the behavior field sensitivity gradient, and the operation event is the sensitivity change rate of the field to the behavior, that is, the response degree of the field change to the behavior triggering, which is observed by model fitting, such as LSTM or attention-based time sequence model, to observe the gradient influence of the field change on the behavior triggering; is the time step number, which represents which record in the time step is currently being processed.
[0070] The application further provides that the quality evaluation module comprises:
[0071] Based on field information, the missing proportion of the preset key field corresponding to the record in the data set is extracted, and the information entropy index calculated by combining the field value distribution is generated to generate the information completeness index for measuring the field information coverage and content richness. Specifically, the information completeness index is a comprehensive evaluation index for measuring whether the information in a data object is complete, whether the fields are complete, and whether the content is rich. It is mainly used to quantitatively score the structural completeness of a single data record or data object in data quality management, data asset evaluation, data cleaning pre-analysis and other scenarios. By scanning each record against the preset key field list, the number of null fields is counted to obtain the missing proportion in combination with the total number of fields; based on field information, firstly, the frequency of all possible values of each field appearing in the sample data is counted; then the frequency is normalized to convert it into a probability distribution; then, using the probability distribution, the information entropy value is calculated by using existing calculation tools or functions. Information entropy is used to reflect the diversity and uncertainty of field information, and the larger the value is, the richer the information is. In the prior art, the calculation method of information entropy is mature and widely used, and common implementations include scientific calculation libraries such as the entropy function in the SciPy library of Python, related functions in the machine learning library, and entropy calculation modules in statistical analysis software such as R language and Matlab, which can be directly called to realize fast and accurate calculation. Finally, the calculated information entropy value and the missing proportion are weighted and fused according to the preset weight to obtain a comprehensive information completeness index, which is used to quantify the structural completeness of the data object. The weight is 0.3 by default because the information completeness index considers two indicators: the field missing proportion reflects the completeness of the data, and the information entropy reflects the diversity and richness of the data content. The missing proportion directly affects the basic data quality, while the information entropy reflects the detailed value of the field content. The weight setting of 0.3 means that the information entropy has a moderate influence on the information completeness index, but it will not excessively cover the core role of the missing proportion, so as to maintain the balance and reasonableness of the overall evaluation.
[0072] The application further provides that the information completeness index, the information cementation degree, the risk mapping skewness and the field behavior interference factor are weighted and fused to construct a comprehensive quality score for representing the field structure rationality, the semantic adhesion, the behavior stability and the risk correlation degree;
[0073] Based on the quality score, a threshold segmentation strategy is used to divide the quality level of each data in the data set, and the data record with a comprehensive quality score higher than a preset threshold is identified as high-quality data.
[0074] According to the requirements of the regulatory agreement, the data compression template is applied to generate a compressed summary report for high-quality data. Specifically, the comprehensive quality score is a comprehensive index for evaluating the overall quality level of a single data record, which integrates the rationality of data structure, the adhesion of field semantics, the stability of field behavior, and the risk relevance of multiple dimensions of information, and can help the system to distinguish between high-quality data and low-quality data, and is mainly used in data quality management, risk assessment, data screening, compliance audit, and compressed summary generation, etc. Records with high comprehensive quality scores indicate that their field structure is reasonable and close, the semantic expression is clear, the historical behavior interference is less, the risk information is complete and relevant, and are suitable for subsequent analysis, modeling and supervision; Records with low scores may have problems such as structural disorder, semantic ambiguity, frequent manual modification, or abnormal risks. The system first calculates the information adhesion degree of each record in the data set based on field information, which involves analyzing the statistical dependence between fields and the stability after fusion. At the same time, the information completeness index of the record is calculated to measure the missing of key fields and the information richness of field content. The system collects the historical operation log corresponding to the record to calculate the field behavior interference factor, which reflects the operation sensitivity and stability of the field. Based on the constructed risk cognition graph, the risk mapping skewness of the record field is calculated to evaluate its potential risk level. Combined with the above indicators, the completeness index is adjusted by using the preset weight to form a weighted positive indicator combination. The positive indicators include information adhesion degree and information completeness index, and the negative impact indicators include field behavior interference factor and risk mapping skewness index, as well as a small constant to prevent the denominator from being 0 and ensure stable calculation. Finally, the scoring results are classified using a predefined threshold to filter out high-quality data. For the filtered high-quality data, a compressed summary report is automatically generated according to the pre-set regulatory agreement and data compression template, which is used for data transmission and storage.
[0075] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any skilled person in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An intelligent data management system, characterized in that, include: Data acquisition module: Collects data information using the multi-source access controller, preprocesses the data information to obtain a dataset, and extracts field information based on the dataset; Decoupling Analysis Module: Based on field information, calculates mutual information index through field discretization and kernel density estimation, and calculates information cohesion degree by combining the perturbation change rate after field fusion. This includes: performing field fusion operation on any two fields based on field information, and generating information cohesion degree to characterize the synergy of field structure based on the perturbation change rate after field fusion and the mutual information index. The graph generation module constructs an attribute fusion graph structure based on field information, calculates node embedding vectors through a graph convolutional network, and constructs a risk mapping skewness by combining the directional features of path perturbation in the fusion graph to generate a risk control cognitive graph. This includes: constructing a multi-hop path set in the attribute fusion graph structure using graph structure analysis methods; extracting the directional features of path perturbation using a path perturbation perception mechanism; and analyzing the changing trends of multi-hop paths during perturbation propagation based on the directional features of path perturbation to construct a risk mapping skewness index. Behavior Analysis Module: Collects action execution logs, extracts behavior sequences of field information within a historical time window, calculates the degree of field disturbance response, and constructs field behavior interference factors. This includes: collecting action execution logs stored in the system, where the action execution logs include fields, operation events, and timestamps; extracting field behavior sequences with a clear temporal order from the action execution logs within a preset historical time window based on the field information; and constructing field behavior interference factors to characterize field stability and intervention sensitivity by calculating the changing trends of field characteristics before and after the operation events based on the operation events and timestamps of the field behavior sequences. Quality assessment module: Based on the proportion of missing key fields extracted from field information and information entropy, the information completeness index is calculated, and a comprehensive quality score is constructed by combining information cohesion, risk mapping skewness and field behavior interference factors.
2. The intelligent data management system according to claim 1, characterized in that, The data acquisition module includes: Multimodal data information is collected through a multi-source access controller, including: transaction flow data, user behavior trajectory data, KYC information, and document data; The data is preprocessed using rule-driven structured data cleaning technology. The preprocessing includes: field renaming, field type unification, missing value completion, outlier correction, and time format standardization to generate a dataset. Fields are extracted from the dataset using a machine learning-based field recognition algorithm, and the information is then integrated to obtain the field information.
3. The intelligent data management system according to claim 1, characterized in that, The decoupling analysis module includes: Based on field information, multiple fields in the dataset are discretized, and the probability distribution of field values is estimated using the kernel density estimation method. Based on the joint probability distribution and marginal distribution of the discretized fields, the mutual information index representing the structural dependency between field pairs is calculated.
4. The intelligent data management system according to claim 1, characterized in that, The map generation module includes: Based on the field information, a graph structure building tool is used to generate an attribute fusion graph structure, in which nodes represent field entities and edges represent fusion relationships between fields. Based on the attribute fusion graph structure, a graph convolutional network is used to perform feature aggregation and representation learning on each node, generating node embedding vectors that represent field structure and semantic correlation.
5. The intelligent data management system according to claim 4, characterized in that, Based on the risk mapping skewness index, a threshold segmentation method is used to assign risk labels to nodes in the attribute fusion graph structure. The risk labels include: low risk, medium risk, and high risk. Filter all nodes with high-risk labels, extract the context adjacent nodes and paths of the nodes, and construct a risk control cognitive graph.
6. The intelligent data management system according to claim 1, characterized in that, The quality assessment module includes: Based on field information, the missing proportion of corresponding records of preset key fields in the dataset is extracted. Combined with the information entropy index calculated by the field value distribution, an information completeness index is generated to measure the field information coverage and content richness.
7. The intelligent data management system according to claim 6, characterized in that, The information completeness index, information cohesion, risk mapping skewness, and field behavior interference factor are weighted and integrated to construct a comprehensive quality score that characterizes the rationality of field structure, semantic cohesion, behavioral stability, and risk correlation. Based on the quality score, a threshold segmentation strategy is used to classify the quality level of each data in the dataset, and data records with a comprehensive quality score higher than the preset threshold are identified as high-quality data. In accordance with regulatory requirements, a compressed summary report is generated for high-quality data using a data compression template.
Citation Information
Patent Citations
E-commerce platform security authentication system and method
CN120342766A
Domain name information processing and displaying method based on multi-modal data fusion
CN120342997A