Multi-dimensional data integrated work file tracing method and system
By constructing a multi-dimensional data integration method for tracing work archives, and utilizing multi-dimensional correlation strength, temporal evolution patterns, and security sensitivity characteristics, this method solves the problems of insufficient data integration and one-sided path evaluation in existing technologies. It achieves high-precision and reliable tracing path generation and optimization, and improves the interpretability and adaptability of tracing results.
Patent Information
- Application Number
- CN202511676169.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-02-06
AI Technical Summary
Existing methods for tracing work records suffer from insufficient data integration, inadequate feature extraction, and biased path evaluation, resulting in low confidence and poor interpretability of tracing results, making it difficult to meet the high-precision tracing requirements in complex business scenarios.
An ontology-based semantic fusion model is constructed to extract features of multidimensional association strength, temporal evolution pattern and comprehensive security sensitivity. A multidimensional tracing model and path comprehensive weight evaluation mechanism are adopted, combined with improved spatiotemporal coupling multidimensional visualization technology, to realize path generation, evaluation and optimization.
It significantly improves the accuracy and reliability of traceability paths, enhances adaptability and robustness in complex scenarios, optimizes path generation and logical consistency, supports the adaptive and intelligent traceability process, and provides multi-dimensional visualization analysis.
Smart Images

Figure CN121480980A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent data management and analysis, in particular to a multi-dimensional data integrated work record tracing method and system. BACKGROUND
[0002] With the in-depth application of information technology in the field of human resource management, organization archives management, work record data presents the characteristics of multi-source heterogeneity, large volume and dynamic evolution. These data are widely distributed in relational databases, unstructured documents, business system logs and external data sources, making it increasingly complex to accurately trace specific work trajectories or associated relationships. Under this background, how to efficiently and reliably reconstruct the complete link of work records from massive and heterogeneous data has become a key technical challenge to improve the quality of organizational decision-making and risk management capability.
[0003] At present, the existing work record tracing methods have obvious limitations. First, at the data level, existing methods often rely on a few types of data sources, lack effective integration and semantic fusion of multi-source heterogeneous data, resulting in weak data foundation and one-sided information. Second, at the feature extraction level, existing technologies usually only focus on simple time sequence patterns, and fail to effectively integrate multi-dimensional features such as correlation strength, dynamic behavior patterns and security risk attributes, resulting in insufficient feature expression. Third, at the path generation and evaluation level, existing solutions mostly use a single path search algorithm, and the evaluation standard is one-sided, lacking comprehensive quantitative evaluation of the internal logic consistency, stability and business compliance of the path, resulting in low confidence and poor interpretability of the tracing results, making it difficult to meet the high-precision tracing needs in complex business scenarios.
[0004] The present application aims to overcome the above-mentioned defects and provides a multi-dimensional data integrated work record tracing method and system. By constructing an ontology-based semantic fusion model, the unified representation and integration of multi-source heterogeneous work record data are realized, providing a comprehensive and consistent data foundation for tracing. By designing feature extraction models that respectively quantify multi-dimensional correlation strength, time sequence evolution pattern and comprehensive security sensitivity, the limitation of insufficient feature dimensionality is overcome. By constructing a multi-dimensional tracing model that integrates correlation, time sequence and security features, and introducing path comprehensive weight, confidence evaluation and multi-level condition judgment mechanism, intelligent generation, accurate evaluation and adaptive optimization of the tracing path are realized, significantly improving the reliability, robustness and business fit of the tracing results. SUMMARY
[0005] In view of the defects in the prior art, the present application provides a multi-dimensional data integrated work record tracing method and system.
[0006] In a first aspect, the present application provides a multi-dimensional data integrated work archive tracing method, comprising the following steps: obtaining work archive data; based on the work archive data, extracting first archive features, second archive features and third archive features; according to the first archive features, the second archive features and the third archive features, obtaining a tracing path and a confidence score; using the tracing path and the confidence score, judging a first level condition, a second level condition and a third level condition to obtain a judgment result; according to the judgment result, combining an improved space-time coupling multi-dimensional visualization technology to start a tracing optimization process. According to the tracing optimization process, the work archive is traced.
[0007] Optionally, based on the work archive data, the first archive features, the second archive features and the third archive features are extracted, which comprises: based on the work archive data, constructing a first archive feature extraction model; through the first archive feature extraction model, extracting first archive features, the first archive features comprising multi-dimensional correlation feature values; based on the work archive data, constructing a second archive feature extraction model; through the second archive feature extraction model, extracting second archive features, the second archive features comprising time series context abnormal feature values; based on the work archive data, constructing a third archive feature extraction model; through the third archive feature extraction model, extracting third archive features, the third archive features comprising comprehensive sensitivity feature values.
[0008] Optionally, the first archive feature extraction model satisfies the following relationship: , wherein, is a multi-dimensional correlation feature value between any two archive entities i and j, is a co-occurrence frequency of entity i and entity j in all data, is a co-occurrence frequency of entity i and entity k in all data, is a neighbor entity set of entity i, is a neighbor entity set of entity j, is a mutual information smoothing factor, , is a network embedding vector of entity i and j, is a preset weight parameter, is a number of common neighbors of entity i and j, is a point mutual information of entity i and j; the second archive feature extraction model satisfies the following relationship: , wherein, is a time series context abnormal feature value, is an operation a content vectorization representation, W is a time window size for calculating a short-term content mean, is a set of all possible operation types, is an operation type a weight, is the probability of occurrence of operation type in the current sliding window, is the probability of occurrence of operation type in the long-term history window, t is the time index of the current operation in the time sequence, k is a loop variable used in the summation operation; the third archive feature extraction model satisfies the following relationship: , wherein, is the comprehensive sensitivity feature value of the archive a in the time window T, D is the number of dimensions of static sensitivity, is the static sensitivity score of the archive a in the dth dimension, is the combined weight of the dth dimension, is the information entropy of the access sequence of the archive a in the time window T, , is the minimum and maximum information entropy of all archives in the time window T, is the neighbor set of the archive a in the implicit association network, is the comprehensive sensitivity of the neighbor archive b in the previous time window , is the shortest path distance between the archives a and b in the implicit association network, is the potential field attenuation coefficient.
[0009] Optionally, according to the first archive feature, the second archive feature and the third archive feature, the traceability path and the confidence score are obtained, which comprises: according to the first archive feature, the second archive feature and the third archive feature, a traceability confidence evaluation model is constructed, the traceability confidence evaluation model fuses multi-dimensional information, and the multi-dimensional information includes association strength, network topology structure and time sequence consistency information; the candidate traceability path is determined by using the traceability confidence evaluation model; the traceability confidence evaluation model is established based on the candidate traceability path; and the confidence score is obtained through the traceability confidence evaluation model.
[0010] Optionally, the first level condition, the second level condition and the third level condition are judged by using the traceability path and the confidence score, and the judgment result is obtained, which comprises: the path consistency measurement function is established by using the traceability path; the path quality comprehensive index function is established according to the confidence score and the path consistency measurement function; the path quality comprehensive index is obtained according to the path quality comprehensive index function. A dynamic acceptable domain model is established; a dynamic judgment threshold is obtained based on the dynamic acceptable domain model; based on the path quality comprehensive index and the dynamic judgment threshold, the first-level conditions, the second-level conditions, and the third-level conditions are judged to obtain the judgment results.
[0011] Optionally, based on the comprehensive path quality index and the dynamic judgment threshold, the first-level conditions, second-level conditions, and third-level conditions are judged to obtain the judgment result, including: based on the comprehensive path quality index and the dynamic judgment threshold, the first-level condition judgment is performed to determine whether the following conditions are met simultaneously: , , in, This is a comprehensive index of path quality. The threshold is dynamically determined, and C is the path confidence score. Based on the confidence threshold, , These are the minimum and maximum theoretical values of the consistency measure, respectively. Construct a path stability index, which satisfies the following relationship: , Where S is the path stability index. For the edge The associated feature values, For operation Abnormal index, This is the stability tolerance coefficient; The path stability index is used to perform a second-level conditional judgment to determine whether the following conditions are met: , in, As a stability benchmark, The coefficient of variation function, Let P represent the overall sensitivity of file a, and P be the current assessment path. This represents the maximum permissible value for the sensitivity coefficient of variation. Construct a business logic compliance function, which satisfies the following relationship: , Where B represents the business logic compliance, and n represents the number of path nodes. For business semantic similarity function, Let i be the i-th node in the path. For functions that conform to the type, Let j be the type of the j-th node. The set of allowed node types; Using the aforementioned business logic compliance, a third-level condition judgment is performed to determine whether the following conditions are met: , in, As a benchmark for business compliance, For similarity coefficient function, The set of neighbors of the starting node. The set of neighbors of the terminating node; the judgment result is obtained through the first-level condition judgment, the second-level condition judgment and the third-level condition judgment, and the judgment result includes the set of valid traceability paths, the set of optimized paths and the identifier of the type of failure.
[0012] Optionally, based on the judgment result, the traceability optimization process is initiated as follows: if the first condition judgment, the second condition judgment, and the third condition judgment all pass, the traceability result is determined to be valid, and a detailed traceability report is output; if any one of the first condition judgment, the second condition judgment, and the third condition judgment fails, the traceability result is determined to be invalid, and the traceability optimization process is initiated.
[0013] Optionally, the traceability optimization process includes: using an improved spatiotemporal coupled multidimensional visualization technology to jointly project the traceability path in the time dimension, relationship dimension, and risk dimension to obtain a multidimensional presentation model; constructing an adaptive optimization engine for the traceability path based on the multidimensional presentation model; establishing a dynamic evolution model of traceability credibility based on the adaptive optimization engine; and establishing a traceability pattern knowledge base based on the dynamic evolution model.
[0014] Optionally, based on the multi-dimensional presentation model, the adaptive optimization engine for tracing paths includes: For paths that fail the first-level condition judgment, the feature enhancement optimization mode is activated; for paths that fail the second-level condition judgment, the structure enhancement optimization mode is activated; and for paths that fail the third-level condition judgment, the semantic alignment optimization mode is activated.
[0015] Secondly, the present invention provides a multi-dimensional data integration work file tracing system. The system uses the aforementioned multi-dimensional data integration work file tracing method, comprising: an acquisition module for acquiring work file data; an analysis module for extracting a first file feature, a second file feature, and a third file feature based on the work file data; obtaining a tracing path and a confidence score based on the first file feature, the second file feature, and the third file feature; using the tracing path and the confidence score to judge first-level conditions, second-level conditions, and third-level conditions, and obtaining a judgment result; initiating a tracing optimization process based on the judgment result and combined with improved spatiotemporal coupling multi-dimensional visualization technology; and an output module for realizing the tracing of work files according to the tracing optimization process.
[0016] Compared with the prior art, the beneficial effects of the present invention include: 1. Improve the accuracy and reliability of tracing paths. By integrating multidimensional association strength, temporal evolution patterns, and dynamic permission sensitivity, a multidimensional tracing model is constructed, overcoming the one-sidedness caused by insufficient dimensions in existing technologies. By introducing a path confidence assessment model and combining association strength distribution, temporal anomalies, and sensitivity fluctuations for comprehensive quantification, the accuracy of identifying real and effective paths is significantly improved.
[0017] 2. Enhance adaptability and robustness in complex scenarios. By employing methods such as correlation feature extraction based on mutual information and network embedding, temporal anomaly detection through semantic and structural fusion, and sensitivity assessment of topological potential field, we can effectively address complex situations such as sparse data, hidden anomalies, and cascading risks, and maintain stable performance even in environments with incomplete data or dynamic evolution.
[0018] 3. Optimize path generation and logical consistency. The improved path generation algorithm integrates association strength, network topology, and temporal consistency, and incorporates a multi-level condition judgment mechanism, including quality screening, stability verification, and business logic verification, to ensure that the output path conforms to business logic in terms of structure, time, and semantics, avoiding interference from invalid or unreasonable paths.
[0019] 4. Achieve adaptive and intelligent traceability processes. A dynamic acceptable domain model, path optimization engine, and online confidence update mechanism are designed to dynamically adjust evaluation criteria and path quality based on task characteristics and user feedback. This supports system self-optimization and knowledge accumulation, significantly improving traceability efficiency and long-term applicability.
[0020] 5. Supports multi-dimensional visualization and interactive analysis. Through spatiotemporal coupling visualization technology, the path is presented in an integrated manner across the dimensions of association, time, and risk, and cross-view interactive linkage is provided to help users intuitively understand complex traceability relationships and improve decision support capabilities. Attached Figure Description
[0021] Figure 1 This is a flowchart of a multi-dimensional data integration working file tracing method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a multi-dimensional data integration work file traceability system according to an embodiment of the present invention. Detailed Implementation
[0022] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.
[0023] Throughout this specification, references to "an embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, the phrases "in an embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0024] Please see Figure 1 The present invention provides a method for tracing work records through multi-dimensional data integration, the method comprising the following steps: S1. Obtain work file data.
[0025] In one embodiment, raw working file data is first extracted in parallel from multiple heterogeneous data sources. These data sources include, but are not limited to, relational databases, unstructured document libraries, business system logs, and external data obtained by web crawlers, ensuring the comprehensiveness of data collection. During the extraction process, a data source identifier and an initial timestamp are attached to each data record to maintain data traceability.
[0026] Furthermore, the extracted raw work file data undergoes data cleaning and alignment. The raw work file data includes structured data, unstructured data, and semi-structured data. The structured data includes database tables, the unstructured data includes text reports, images, and videos, and the semi-structured data includes XML format logs.
[0027] Specifically, for structured data, null value imputation, outlier correction, and format standardization are performed; for unstructured data, key information is extracted and vectorized using natural language processing and computer vision techniques, such as converting text reports into word vectors and identifying and feature-encoding seals or signatures in images; for semi-structured data, format parsing and key field mapping are performed.
[0028] Furthermore, a data fusion model is constructed to map the cleaned and aligned work file data to the same feature space.
[0029] Specifically, an ontology-based semantic fusion method is used to identify, disambiguate, and associate the same entity from different data sources, and generate a unified digital archive object containing multiple features for each entity, thereby completing the data fusion preprocessing and obtaining preprocessed working archive data.
[0030] S2. Based on the work file data, extract the first file feature, the second file feature, and the third file feature.
[0031] In one embodiment, the working file data generated in step S1 is first invoked to provide a consistent data foundation for multi-dimensional feature extraction; based on the different requirements of the tracing task for correlation, temporal behavior and security attributes, the first, second and third features are extracted respectively.
[0032] Furthermore, a first archival feature is extracted, which quantifies the multidimensional association strength between archival entities. To achieve this, the present invention constructs a composite association measurement algorithm, which, based on existing co-occurrence analysis, introduces contribution weighting based on information entropy and structural similarity compensation based on network embedding.
[0033] Specifically, a first archival feature extraction model is constructed, which satisfies the following relationship: , in, Let i be the multidimensional association feature value between any two file entities i and j. Let i be the co-occurrence frequency of entity i and entity j in all data. The co-occurrence frequency of entity i and entity k in all data is obtained by counting the number of times the two entities appear together in all documents, logs or related records. Let i be the set of neighboring entities. The set of neighboring entities of entity j is obtained by finding all entities that have direct edges to entity i or j in the initial association network that has been constructed. is the mutual information smoothing factor, which is used to control the sensitivity of the mutual information term and is adjusted through cross-validation; and are the network embedding vectors of entities i and j, which are obtained by training the entity relationship graph into the graph embedding model; is the preset weight parameter, and its value range is [0, 1], which is used to adjust the proportion of the two similarity measures in the compensation term; is the number of common neighbors of entities i and j, which is obtained by calculating the size of the intersection of the neighbor sets of the two entities; is the pointwise mutual information between entities i and j; it is calculated through the following expression: , where and are the probability estimates of the entities appearing alone, is the probability estimate of the entities co-occurring.
[0034] Furthermore, the first archive feature extraction model is used to extract the first archive features.
[0035] Compared with the prior art, the improvement of the first archive feature extraction model lies in: 1. Introducing the mutual information weight, which can filter out high-frequency but uninformative co-occurrences, such as stop words like "of", "is", etc., and emphasizes statistically significant associations.
[0036] 2. Introducing the structural similarity compensation, which combines the global network embedding information and the local common neighbor information. Even if two entities have few direct co-occurrences, as long as their roles in the network structure are similar, they can be assigned a high association feature value, significantly improving the association discovery ability in sparse data scenarios.
[0037] Furthermore, the second archive features are extracted. The second archive features reflect the temporal features of the dynamic evolution pattern of the archive operation sequence and are used for the temporal evolution model. The present invention designs a temporal anomaly-sensitive feature that combines the mutation of the operation content and the violation of the context pattern.
[0038] Specifically, a second archive feature extraction model is constructed. The second archive feature extraction model satisfies the following relational expression: , where is the temporal context anomaly feature value, is the content vector representation of the operation , which is obtained by inputting text information such as the operation description and operation object into the pre-trained model and taking the vector output of the label; W is the time window size for calculating the short-term content mean, which is determined according to the typical cycle of the operation sequence in the business scenario; The set of all possible operation types is obtained by deduplicating and encoding all operation types in the historical operation log; For operation type The weight is obtained by normalizing the product of the frequency of occurrence of this type of operation in historical abnormal data and the information entropy, and is used to distinguish the importance of different operation types. For the operation type within the current sliding window The probability of occurrence is obtained by statistically analyzing and normalizing the frequency of various operations within the current window; For operation types in the long history window The probability of occurrence is obtained by statistically analyzing the frequency of various operations over a relatively long period and normalizing it; t is the time index of the current operation in the time sequence, and k is the loop variable used in the summation operation.
[0039] Furthermore, second archival features are extracted using the second archival feature extraction model.
[0040] Compared to existing technologies, the improvements of the second archival feature extraction model are as follows: 1. Introducing a short-term content drift metric based on deep semantics can capture subtle but anomalous semantic changes in the content being operated on, rather than just anomalous time intervals.
[0041] 2. By introducing a weighted operation type distribution deviation measure, detection can be performed from the perspective of macroscopic sequence patterns. This can uncover hidden anomalies where individual operations appear normal, but the combination patterns deviate from historical norms, thus achieving the fusion of microscopic and macroscopic anomaly features.
[0042] Furthermore, a third set of archival features is extracted, which assesses the comprehensive sensitivity of an archive in a dynamic access control environment. This sensitivity integrates its static attributes, dynamic access behavior patterns, and topological risks within the implicit relational network. This invention abandons existing static weighting or simple frequency statistics methods and proposes a dynamic sensitivity quantification method based on information entropy and topological potential.
[0043] Specifically, a third-archive feature extraction model is constructed, which satisfies the following expression: , in, Let D be the comprehensive sensitivity feature value of file a within the time window T, and let D be the number of dimensions of static sensitivity, which is determined according to the predefined sensitivity evaluation dimensions. The static sensitivity score of archive a on the dth dimension is obtained by automatically scoring the archive metadata through a rule-based system or machine learning model, and then calibrating it through expert review. The combined weight of the d-th dimension is obtained by normalizing the relative importance of each dimension to the overall sensitivity after calculating it using the analytic hierarchy process. The information entropy of the access sequence of file a within time window T is obtained by calculating the joint information entropy of the operation type, visitor role, time distribution and other features of all access events within the window, and is used to quantify the unpredictability of access behavior. , To find the minimum and maximum information entropy of all files within a time window T, calculate the information entropy of all files in the system by traversing the system. The value is obtained by taking the extreme value; The set of neighbors of file a in the implicit association network is constructed by setting a threshold based on the multidimensional association feature values obtained from the first file feature extraction model. For neighbor file b in the previous time window The overall sensitivity is calculated by recursively applying this formula, with an initial value... Determined by its geometric mean static sensitivity; The shortest path distance between files a and b in the implicit association network is calculated using a graph search algorithm in the implicit association network. is the potential field attenuation coefficient, used to control the attenuation rate of sensitivity propagation on the network, and is obtained by optimization on the validation set through grid search.
[0044] Furthermore, the third archive features are extracted using the aforementioned third archive feature extraction model.
[0045] Compared to existing technologies, the improvements of the third archive feature extraction model are as follows: 1. Geometric mean is used instead of the linear weighting method commonly used in existing technologies. This non-linear fusion can better reflect that the low sensitivity of any dimension will significantly suppress the overall score, thus more accurately portraying the inherent vulnerability of the archives and overcoming the risk masking problem that may exist in linear methods.
[0046] 2. Information entropy of access behavior is introduced as a dynamic penalty term; unlike existing technologies that only focus on access frequency, this model assesses risk by quantifying the uncertainty and randomness of access patterns; the more unpredictable an access sequence is, the greater its potential risk, thus upgrading the risk assessment dimension from access quantity to access quality.
[0047] 3. A recursive topological potential field model is proposed to simulate the propagation of sensitivity in interconnected networks. It surpasses existing methods that isolate nodes for evaluation or use PageRank to measure global importance. It can dynamically and locally capture network phenomena of risk propagation and diffusion on specific network paths, and has significant advantages in identifying potential cascading risks caused by interconnected relationships.
[0048] In this embodiment, during the tracing of the leak incident involving the internal project "Alpha" of a certain company, relevant work file data was obtained. The extracted first file feature quantifies the association strength between personnel and documents: for example, employee A and document X co-occurred 15 times, their neighbor set overlap was 0.6, and their network embedding vector similarity was 0.8, resulting in a first file feature value of 0.72; employee B and document Y co-occurred only 3 times, but their structural similarity compensation was high, resulting in a feature value of 0.65; the second file feature reflects the anomaly of the operation sequence: within 72 hours before the leak incident, an access operation sequence of employee C to document Z was detected, with a short-term content drift metric of 0.85, an operation type distribution deviation of 0.7, and a temporal anomaly feature value of 0.78, significantly higher than the historical average of 0.3. The third archival feature assessment evaluated the dynamic sensitivity of the archives: the static attribute sensitivity score of document X was 8.5 (out of 10), its recent access behavior information entropy was 2.1, and the sensitivity of its neighboring document W in the previous period in the association network was 7.2. The overall sensitivity of document X in this period was calculated to be 7.9, which belongs to the high-risk level.
[0049] S3. Based on the first archive feature, the second archive feature, and the third archive feature, obtain the tracing path and confidence score.
[0050] In one embodiment, firstly, the first, second, and third archival features extracted in step S2 are integrated to construct a multi-dimensional tracing model capable of comprehensively evaluating the reliability of tracing paths. This model aims to overcome the problems of single tracing dimensions and one-sided evaluation standards in existing technologies. The multi-dimensional tracing model includes a tracing path generation model and a tracing confidence assessment model.
[0051] Specifically, a tracing path generation model is constructed. Based on the first archival feature, this model searches for all possible connection paths in the archival association graph. An improved k-shortest path algorithm is employed, which not only considers the number of hops but also introduces path weights based on the first archival feature, ensuring that the generated candidate paths have high confidence in both structural connectivity and semantic relevance. It is important to note that existing k-shortest path algorithms are a class of algorithms used to find the top k shortest paths in a graph; that is, given a starting point and an ending point, they sort the paths in ascending order by path length or weight and return the top k paths among all possible paths.
[0052] The expression for the traceability path generation model is as follows: ; in, The comprehensive weight score of candidate path P is used to sort and filter all candidate paths. The higher the score, the higher the priority of the path. P is a candidate trace path. It is generated by the improved k-shortest path algorithm and is a sequence of nodes and edges. Let be a directed edge in path P; obtained by traversing adjacent pairs of nodes on path P. For the edge The dynamic importance weight; calculated by determining the two entities connected by the edge. and The geometric mean of the PageRank values is used to obtain the edge type, combined with the prior importance of the edge. For entities and The first archival feature between them is calculated by the first archival feature extraction model in step S2; The power exponent of the association feature is obtained by optimizing on the validation set through grid search, and is used to control the nonlinear effect of association strength in the weights; The length of path P is obtained by directly counting the number of nodes in the path: The maximum allowed path length is preset based on business needs and the trade-offs in tracing efficiency: The path length penalty coefficient is learned from historical data through a machine learning model and is used to adjust the degree of suppression of excessively long paths. It is a smoothing constant, a very small positive number set to prevent the denominator from being zero, and is determined by empirical value; Here, the variance function is used to calculate the degree centrality of nodes along the path: For nodes Degree centrality; by statistically analyzing nodes in the archival association graph. The total number of connected edges is obtained as follows: The maximum degree centrality of all nodes in the graph is obtained by traversing the degree of all nodes in the graph and taking the maximum value. This is the set of timestamps corresponding to all edges on path P, obtained by extracting the timestamps of operations or events associated with each edge; The time interval between consecutive events on the path is calculated. The difference between adjacent timestamps is obtained; The average time interval for the path is calculated by... The arithmetic mean is obtained; The standard deviation of the path time interval is calculated by all The standard deviation is obtained, where t is time.
[0053] Compared to existing technologies, the improvements of the aforementioned tracing path generation model are as follows: 1. By integrating information from three dimensions—association strength, network topology, and temporal consistency—into a unified weighting function, a set of candidate paths with richer semantics and better alignment with business logic is generated.
[0054] 2. A penalty term based on time interval variability was introduced. It can identify and suppress paths with chaotic and illogical timelines, such as those where the time sequence of events on a path is reversed or the intervals are extremely uneven.
[0055] 3. By simultaneously introducing static correlation feature values and dynamic importance weights based on the global network structure in the numerator, and penalizing paths with unstable topological structures by calculating the variance of the path node degree centrality ratio in the denominator, this design that integrates macroscopic network structure and microscopic correlation features overcomes the one-sidedness of existing edge weights.
[0056] Furthermore, a traceability confidence assessment model is constructed. This model is the core of the multi-dimensional traceability model. Its function is to calculate a quantitative confidence score for each candidate traceability path by integrating the first, second, and third archival features. The confidence score is used to measure the probability that the traceability path is a real and valid path.
[0057] Specifically, the traceability confidence assessment model satisfies the following relationship: , in, The confidence score is given to the candidate tracing path P. The higher the score, the higher the confidence of the path. P is a candidate tracing path, which consists of a series of consecutive archive entity nodes. The number of entity nodes contained in path P is obtained by directly counting the number of nodes on the path; Let P be a directed edge connecting entities. and ; For entities With entity The first archival feature between them is calculated by the first archival feature extraction model in step S2; For the border The set of all operations occurring within the corresponding time interval or logical relationship is filtered out by querying the time series database, focusing on the entity. and All operation records that occurred during the association's validity period are obtained; For operation The anomaly index is obtained by normalizing the temporal context anomaly feature value output by the second archive feature extraction model in step S2 using the Sigmoid function; This is the path-level anomaly tolerance coefficient, which is optimized on the validation set through grid search and is used to adjust the impact of time-series anomalies on the overall confidence score. The third archival feature of archival entity a within the time window T is calculated using the third archival feature extraction model in step S2. The preset global maximum comprehensive sensitivity threshold is set according to business security requirements and is used to normalize the average sensitivity on the path. is the Gini coefficient function, used to measure the degree of inequality of the association characteristics of each edge on the path. It is obtained by calculating the Gini coefficient of these characteristics. The larger the value, the more uneven the distribution of association strength. Let P be the union of the sets of operations corresponding to all edges on path P; The maximum anomaly index among all operations on path P is determined by traversing... All in the set And obtain the maximum value; The coefficient of variation function is used to measure the overall sensitivity of each node on the path. The degree of dispersion of these sensitivity values is obtained by calculating the ratio of the standard deviation to the mean of these sensitivity values; This is the sensitivity dispersion normalization factor, set according to the distribution of the coefficient of variation in historical path data, used to balance the magnitude of the coefficient of variation term.
[0058] Compared to existing technologies, the improvements of the traceability confidence assessment model are as follows: 1. By introducing a path-level nonlinear fusion mechanism, the geometric average of the key factors of all edges on the path is calculated, which significantly reduces the negative impact of weak links in the path on the overall evaluation results and better reflects the overall robustness of the path.
[0059] 2. By introducing three penalty terms—uneven distribution of association strength, worst-case time series anomaly, and sensitivity volatility—the system systematically identifies and penalizes paths that have acceptable average association but poor internal stability, contain high-risk nodes, or have inconsistent security attributes.
[0060] 3. By organically integrating static relationships, dynamic temporal behaviors, and security status, a unified quantitative assessment across time and space and across attributes is achieved, which is more in line with the complexity of real-world traceability scenarios.
[0061] Furthermore, using the aforementioned traceability confidence assessment model, all candidate traceability paths are calculated, and the top k paths with the highest confidence scores are selected as preliminary traceability results, outputting their path sequences and corresponding confidence scores.
[0062] In this embodiment, possible leakage paths are searched in the personnel-document association graph. The main candidate paths generated include: path P1 (employee A → document X → external IP address), with a length of 2, a mean association feature value of 0.70, a time interval variation coefficient of 0.2, a node degree centrality variance of 0.3, and a comprehensive weight score of 8.5; path P2 (employee B → document Y → employee C → document Z → cloud storage account), with a length of 4, a mean association feature value of 0.55, a time interval variation coefficient of 0.8, a node degree centrality variance of 0.7, and a comprehensive weight score of 5.2. Calculated using the traceability confidence assessment model, the confidence score of path P1 is 0.88, and the confidence score of path P2 is 0.62. Finally, the top 3 paths with the highest confidence scores are selected as the preliminary traceability results, with path P1 ranking first.
[0063] S4. Using the tracing path and the confidence score, make judgments on the first condition, the second condition and the third condition to obtain the judgment result.
[0064] In one embodiment, a preliminary tracing result is first received from the output of step S3. The preliminary tracing result includes a set of candidate tracing paths and their corresponding confidence scores, i.e., confidence scores. Furthermore, a path consistency metric function is constructed. This consistency metric function aims to quantify the coherence and rationality of the internal logic of a path. The expression of the path consistency metric function is as follows: ; in, This is the consistency metric for path P; P is a candidate traceability path, derived from the output of step S3. The number of entity nodes contained in path P is obtained by directly counting the number of nodes on the path; It is a directed edge in path P, obtained by traversing adjacent node pairs on path P; For nodes and nodes The Jaccard similarity coefficient of the neighbor set is calculated by... This is used to measure the closeness of adjacent nodes on a path within the network structure. For nodes The representative timestamp is obtained by extracting the median of the timestamps of the major events or operations associated with that node; The scaling factor for timestamp differences is obtained by calculating the standard deviation of timestamp differences for all possible edges and is used to normalize the impact of timestamp differences. The semantic similarity between node a and the overall topic m of path P is obtained by calculating the cosine similarity between the text description vector of node a and the average vector of the text description vectors of all nodes on path P: This represents the mean semantic topic similarity of all nodes on path P; is the entropy of all edge types on path P, obtained by calculating the information entropy of the distribution of these edge types, and is used to measure the diversity of relation types in the path; z is an edge in the tracing path P; The maximum value of the edge type entropy among all candidate paths is obtained by iterating through the edge type entropy of all candidate paths and taking the maximum value.
[0065] Compared to existing technologies, the improvement of the path consistency metric function lies in: 1. By integrating structural consistency, temporal consistency, and semantic consistency, a more comprehensive framework for evaluating the rationality of path logic is provided, effectively solving the problem that existing technologies do not fully address the path attributes.
[0066] 2. A penalty term based on semantic similarity variance is introduced, which can effectively identify and suppress drift paths that, although locally related, deviate from the core theme in terms of overall semantics. This is crucial for ensuring the focus of tracing results in terms of business meaning.
[0067] 3. By introducing edge type entropy as a penalty factor, the priority of paths with overly complex relation types and lacking dominant logic can be reduced, making the tracing results more interpretable and consistent in terms of relation types.
[0068] Furthermore, a multi-dimensional joint verification mechanism that goes beyond conventional threshold judgment is constructed. By introducing a comprehensive path quality index and a dynamic acceptable domain model, intelligent judgment of tracing results is achieved.
[0069] Specifically, firstly, a comprehensive path quality index function is constructed. This function comprehensively considers the confidence, complexity, and consistency of a path to form a unified evaluation index, the expression of which is: , in, The overall quality index of path P; The confidence score for path P is derived from the output of step S3. The length of path P is obtained by counting the number of nodes in the path; This is the path length normalization factor, calculated as the median of all candidate path lengths. The consistency measure for path P is calculated using the function defined in step S4; The weighting parameters are determined by combining entropy weighting with expert scoring. This is the set of candidate paths, derived from the output of step S3. The variance function is used to calculate the degree of dispersion of the relative confidence values, where i and j are path indices.
[0070] Furthermore, a dynamic acceptable domain model is established. This model dynamically adjusts the judgment criteria based on the characteristics of the current tracing task and the distribution of historical data. Its expression is: ; in, The threshold is determined dynamically. The set of quality indices for historically valid paths is extracted from the historical tracing results database; M is the median function, reflecting the typical level of the quality index; D is the mean absolute deviation, measuring the dispersion of the quality index. represents the current set of candidate paths, derived from the output of step S3; E is the information entropy function, which evaluates the diversity of the current path set. The maximum path set entropy value is obtained through theoretical calculation. , The adjustment coefficients are obtained by optimizing them on the validation set through machine learning.
[0071] Furthermore, a multi-level condition judgment process is designed.
[0072] Specifically, in order to achieve path quality screening, a first-level condition judgment is performed to determine the comprehensive path quality index. Has the dynamic threshold been reached? ,Right now Simultaneously, the path confidence C is required to satisfy the following relationship: , in, The base confidence threshold is determined by optimizing model performance on the validation set. , These are the minimum and maximum theoretical values of the consistency measure, respectively; obtained through theoretical analysis combined with historical data statistics.
[0073] Further, the first-level condition judgment is performed, namely path stability verification.
[0074] Specifically, a path stability index S is constructed, wherein the path stability index S satisfies the following relationship: , The stability index must meet the following conditions: , in, For the edge The associated feature values are derived from the output of the first archive feature extraction model in step S2; For operation The anomaly index comes from the output of the second archive feature extraction model in step S2; The stability tolerance coefficient is obtained by optimization on the validation set through grid search; Set as a stability benchmark based on business needs; The coefficient of variation is used to calculate the degree of dispersion of the sensitivity values; The overall sensitivity of file a is derived from the output of the third file feature extraction model in step S2; The maximum allowable value for the sensitivity variation coefficient is set according to business security requirements; P is the current evaluation path, derived from the output of step S3.
[0075] Further, a third level of condition judgment is performed, namely, a business logic consistency check.
[0076] Specifically, construct the business logic compliance function: , The business logic compliance must satisfy the following relationship: , Where B represents the business logic compliance, and n represents the number of path nodes, which is obtained by counting the path nodes; This is the business semantic similarity function, obtained by calculating the cosine similarity of the node's business description vector; Let be the i-th node in the path, which comes from the path sequence; For the type conformance function, check whether the node type is within the allowed range; Extract the type of the j-th node from its attributes; The set of allowed node types is set according to business rules; The benchmark for business compliance is determined through expert evaluation. The Jaccard similarity coefficient function is used to calculate the similarity between two sets. The set of neighbors of the starting node is extracted from the association graph; The set of neighbors of the terminating node is extracted from the association graph.
[0077] Furthermore, if a candidate path passes all three levels of judgment simultaneously, the tracing result is deemed valid, and a detailed tracing report is output. If any judgment fails, the tracing optimization process is initiated, and corresponding optimization strategies are adopted based on the specific type of judgment that failed.
[0078] The advantage of this determination method is that: 1. By introducing a comprehensive quality index, multiple evaluation dimensions are integrated in a non-linear manner, avoiding the limitations of insufficient dimensional judgment and solving the misjudgment problem caused by independent judgment of each condition in the existing technology.
[0079] 2. By establishing a dynamic acceptable domain, the threshold is adaptively adjusted based on historical data and current path set characteristics, overcoming the problem of poor applicability of the threshold in different scenarios of existing methods and improving the flexibility of judgment.
[0080] 3. By designing a multi-level progressive verification system, from basic quality screening to stability verification and then to business logic testing, a progressively in-depth verification system has been formed to ensure that only truly high-quality paths can pass all checkpoints.
[0081] 4. By combining the ratio of the weakest to the strongest correlation in the path with the maximum anomaly index, unstable paths that have good average indicators but obvious weak links are effectively identified.
[0082] In this embodiment, the preliminary tracing results are evaluated using a three-level judgment process. The overall quality index of path P1 is 0.85, higher than the dynamic threshold of 0.75; its basic confidence level of 0.88 meets the requirements; its stability index is 0.8 (the ratio of the weakest association to the strongest association is 0.7, and the maximum anomaly index is 0.3), higher than the baseline value of 0.6; its business logic compliance is 0.9 (node type compliance is 1.0, and the similarity between the start and end nodes is 0.8), higher than the required value of 0.7; path P1 successfully passes all three levels of judgment. The overall quality index of path P2 is 0.68, lower than the dynamic threshold; its stability index is 0.4 (the ratio of the weakest association to the strongest association is 0.3, and the maximum anomaly index is 0.7), lower than the baseline value; its business logic compliance is 0.5, which does not meet the requirements; therefore, path P2 is determined to require optimization.
[0083] S5. Based on the judgment result, and combined with the improved spatiotemporal coupling multidimensional visualization technology, initiate the traceability optimization process.
[0084] In one embodiment, the judgment result from step S4 is first received, which includes a set of valid traceability paths, a set of optimized paths, and a corresponding failure judgment type identifier; based on the judgment result, a multi-mode traceability output and adaptive optimization mechanism is executed to realize full-link intelligent traceability of work files.
[0085] Furthermore, a multi-dimensional presentation model for traceability results is constructed. This model overcomes the limitations of insufficient linear path display in existing technologies by employing an improved spatiotemporal coupling multi-dimensional visualization technology to jointly project the traceability path in the time dimension, relationship dimension, and risk dimension.
[0086] Specifically, the multi-dimensional presentation model satisfies the following relationship: , in, A multidimensional visual representation of path P; It is a timestamp sequence of the path, obtained by extracting the timestamps of the associated events of each node in the path; The spatial projection function maps path nodes to spatial coordinates of the association graph based on the features of the first archive, and is obtained by combining the force-directed algorithm with multidimensional scaling analysis. This is a time projection function that maps event timestamps to a non-uniform scale on the timeline, compressing and expanding the scale according to the importance of the event. For the risk projection function, the third file feature values are mapped to visual codes of color depth and icon size; This is a cross-dimensional connection function that ensures the selection and state synchronization of the same entity in different dimensional views. Operators that represent the overlay of multi-dimensional views.
[0087] Compared to existing technologies, the improvement of the spatiotemporally coupled multidimensional visualization technology lies in its joint projection of the tracing path across time, relationships, and risk dimensions, thereby achieving a more three-dimensional and richer visualization. Specifically, this multidimensional presentation model maps the path as a superposition of spatial projection, temporal projection, and risk projection, and embeds a cross-dimensional connection mechanism to ensure interactive linkage between views. Spatial projection constructs a graph coordinate system based on the relationships between nodes, utilizing force-guided layout and multidimensional scaling analysis. Temporal projection non-uniformly scales the timeline based on event importance to highlight key event nodes. Risk projection intuitively encodes risk levels through visual variables such as color and size. Furthermore, the cross-dimensional connection function ensures that user operations in different views can be synchronized in real time, enhancing multi-view collaborative analysis capabilities and significantly improving the expressiveness of the tracing path and user interpretation efficiency.
[0088] Furthermore, an adaptive optimization engine for tracing paths is constructed. For paths that fail the judgment in S4, corresponding optimization strategies are activated based on the specific type of failure.
[0089] Specifically, the adaptive optimization engine includes three optimization modes: For paths that fail the first-level condition judgment (path quality screening), activate the feature enhancement optimization mode: ; in, The optimized path; It is the set of neighborhood paths of a path, obtained by performing local replacement, edge addition or deletion operations on path nodes in the association graph; The comprehensive quality index of the candidate path is calculated using the comprehensive path quality index function in step S4. The similarity between the original path and the candidate path is obtained by calculating the weighted sum of the node overlap rate and the edge sequence edit distance of the two paths; The similarity weight coefficients are dynamically adjusted through reinforcement learning in the interactive feedback. P represents the candidate path, and P represents the original path.
[0090] For paths that fail the second-level condition judgment (path stability verification), initiate the structure reinforcement optimization mode: , in, The optimized path, For the associated feature values in path P The weakest edge is obtained by iterating through the first file of all edges in the path and taking the minimum value. For bridge nodes, by finding simultaneous with and It has highly correlated feature values, that is and The nodes are obtained, The threshold for associated features is determined by statistically analyzing the upper quartiles of all edge associated feature values.
[0091] For paths that fail the third-level condition check (business logic consistency check), initiate semantic alignment optimization mode: ; in, The optimized path, For semantic alignment functions; It serves as a business constraint rule library, extracted from domain ontology and business specification documents; Map is a node mapping function that maps node pairs in the original path to equivalent node pairs that conform to business logic, achieved by calculating the matching degree between node semantic description and business constraints; Let i be a directed edge formed by two adjacent nodes in path P, where i is the node index. It serves as a context or semantic knowledge base used in the mapping process, containing conceptual relationships and attributes from the domain ontology, to help determine whether two nodes are equivalent in business logic.
[0092] Furthermore, a dynamic evolution model for traceability credibility is constructed. This model establishes an online update mechanism for confidence scores based on user feedback on traceability results, thereby enabling the system to self-optimize.
[0093] Specifically, the dynamic evolution model satisfies the following relationship: ; in, The path confidence for the (n+1)th iteration; The decay factor for historical confidence is determined by the time decay function; It is a direct feedback factor, obtained through users' explicit rating of the path; As an indirect feedback factor, it is calculated through implicit feedback signals such as user subsequent query behavior, dwell time, and path expansion depth; To determine the fusion weight, user experience metrics are optimized through traffic splitting tests.
[0094] Furthermore, a mechanism for the accumulation and transfer of traceability knowledge is constructed. Validated high-quality traceability paths and their characteristic patterns are stored in a traceability pattern library to support the rapid initiation of traceability tasks and pattern reuse.
[0095] Specifically, the knowledge accumulation mechanism adopts the following approach: ; in, For tracing pattern knowledge base; For path The abstract pattern description is obtained by extracting key feature templates of the path; The vectorized representation of the path is obtained by inputting the features of the path nodes and edges into a graph neural network; The metadata for the path includes generation time, usage frequency, and applicable scenario tag information, where N is the total number of paths.
[0096] Compared to existing technologies, this method has the advantage of multi-dimensional visualization and interactive innovation. It uses spatiotemporally coupled multi-dimensional projection technology to achieve an integrated presentation of the traceability path in the associated space, time axis, and risk dimension. Combined with a cross-dimensional connection mechanism, it ensures consistent interaction between multiple views, significantly improving user experience and traceability efficiency. The adaptive optimization engine designs differentiated optimization strategies for different failure judgment types, forming a complete optimization loop through feature enhancement, structural reinforcement, and semantic alignment, fundamentally improving path quality. The dynamic evolutionary learning mechanism establishes an online confidence update model based on user feedback, integrating direct and indirect feedback signals to achieve system self-evolution and adapt to changes in the business environment. The knowledge accumulation and transfer architecture builds a traceability pattern library, enabling the accumulation and reuse of high-quality traceability experience, improving system startup speed and traceability quality, and forming continuously value-added knowledge assets.
[0097] In this embodiment, an optimization process is initiated for the failed path P2. In feature enhancement optimization, a new path P2' is generated by replacing the node "Employee C" with "Employee D" (which has a high correlation with both document Y and document Z) in the association graph, and its overall quality index is improved to 0.78. In structural reinforcement optimization, a bridge node "Project Server" is added to the weakest edge "Employee B → Document Y" in P2 to form a reinforced path P2''. In semantic alignment optimization, the node "Cloud Storage Account" that does not conform to business logic is mapped to the compliant "Authorized External Collaboration Platform". The optimized path P2''' has an overall quality index of 0.82, passing all judgment conditions. At the same time, the visualization results of path P1 are projected in three dimensions on the timeline, association graph, and risk heatmap to support user interactive exploration.
[0098] S6. Based on the traceability optimization process, traceability of work files is achieved.
[0099] In one embodiment, the work files were traced according to the traceability optimization process in step S5.
[0100] In this embodiment, through the above process, two effective tracing paths were ultimately determined: the primary path P1 shows that employee A directly leaked the document X to an external IP, with a confidence level of 0.88; the secondary path P2''' shows that employee B, via the project server, had employee D operate document Z to an authorized external platform, which may pose a compliance risk, with a confidence level of 0.82. Based on these tracing results, the security team conducted an in-depth investigation into employee A and confirmed the violation; at the same time, a compliance review was conducted on the external collaboration process involved in path P2''', and control measures were improved; all tracing paths, judgment results, and optimization processes were recorded in the knowledge base for rapid matching and tracing of similar events in the future.
[0101] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of a multi-dimensional data integration work file traceability system according to an embodiment of the present invention. The system includes an input device, a processor, an output device, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions. The processor is configured to call the program instructions. The system uses a multi-dimensional data integration work file traceability method. The input device includes an acquisition module, the processor includes an analysis module, and the output device includes an output module. The acquisition module is used to acquire work file data; The analysis module is used to extract first, second, and third file features based on the work file data; obtain a tracing path and confidence score based on the first, second, and third file features; use the tracing path and confidence score to judge the first, second, and third level conditions and obtain the judgment result; and initiate the tracing optimization process based on the judgment result and the improved spatiotemporal coupling multidimensional visualization technology. The output module is used to enable the traceability of work files according to the traceability optimization process.
[0102] In summary, this invention extracts three types of archival features—quantitative correlation strength, temporal anomaly patterns, and dynamic security sensitivity—by collecting multi-source heterogeneous data in parallel and constructing a fusion model. Based on this, a multi-dimensional traceability model integrating correlation networks, temporal behavior, and security attributes is designed to generate candidate paths and calculate confidence scores. A three-level progressive intelligent judgment mechanism is implemented by introducing a comprehensive path quality index and a dynamic acceptable domain model. For failed paths, three optimization modes—feature enhancement, structural reinforcement, and semantic alignment—are employed, combined with spatiotemporal coupling visualization technology and a dynamic confidence evolution mechanism, ultimately achieving precise, adaptive, and intelligent end-to-end traceability of working archives.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A method for tracing work records using multi-dimensional data integration, characterized in that, The method includes the following steps: Obtain work file data; Based on the work file data, extract the first file feature, the second file feature, and the third file feature; Based on the first archival feature, the second archival feature, and the third archival feature, obtain the tracing path and confidence score; Using the tracing path and the confidence score, the first-level conditions, the second-level conditions, and the third-level conditions are judged to obtain the judgment results; Based on the judgment results, and combined with the improved spatiotemporal coupling multidimensional visualization technology, the traceability optimization process is initiated; According to the aforementioned traceability optimization process, the traceability of work files can be achieved.
2. The method for tracing work archives through multi-dimensional data integration according to claim 1, characterized in that, Based on the aforementioned work file data, the extraction of the first file feature, the second file feature, and the third file feature includes: Based on the aforementioned work file data, a first file feature extraction model is constructed; The first archive feature is extracted using the first archive feature extraction model. The first archive feature includes multidimensional correlation feature values. Based on the aforementioned work file data, a second file feature extraction model is constructed; The second archive feature is extracted using the second archive feature extraction model. The second archive feature includes temporal context anomaly feature values. Based on the aforementioned work file data, a third file feature extraction model is constructed; The third archive feature is extracted using the third archive feature extraction model. The third archive feature includes a comprehensive sensitivity feature value.
3. The method for tracing work archives through multi-dimensional data integration according to claim 2, characterized in that, The first archive feature extraction model satisfies the following relationship: , in, Let i be the multidimensional association feature value between any two file entities i and j. Let i be the co-occurrence frequency of entity i and entity j in all data. Let i be the co-occurrence frequency of entity i and entity k in all data. Let i be the set of neighboring entities. Let j be the set of neighboring entities. Mutual information smoothing factor, , Let i be the network embedding vectors of entities i and j. The preset weight parameters, Let be the number of common neighbors of entities i and j. The mutual information between entities i and j; the second archive feature extraction model satisfies the following relationship: , in, These are time-series context anomaly feature values. For operation The content is vectorized, and W is the size of the time window used to calculate the short-term content mean. For the set of all possible operation types, For operation type The weight, For the operation type within the current sliding window The probability of its occurrence, For operation types in the long history window The probability of occurrence is given by t, where t is the time index of the current operation in the time sequence, and k is the loop variable used in the summation operation; the third archive feature extraction model satisfies the following relationship: , in, Let be the comprehensive sensitivity feature value of file 'a' within the time window T, and D be the number of dimensions of static sensitivity. Score the static sensitivity of file a on the d-th dimension. Let d be the combined weight of the d-th dimension. Let the information entropy be the sequence of accesses to file a within time window T. , The minimum and maximum information entropy of all files within the time window T. Let a be the set of neighbors of file a in the implicit association network. For neighbor file b in the previous time window Overall sensitivity Let be the shortest path distance between files a and b in the implicit association network. is the potential field attenuation coefficient.
4. The method for tracing work archives through multi-dimensional data integration according to claim 1, characterized in that, Based on the first archival feature, the second archival feature, and the third archival feature, the tracing path and confidence score are obtained, including: Based on the first archive feature, the second archive feature, and the third archive feature, a traceability confidence assessment model is constructed. The traceability confidence assessment model integrates multi-dimensional information, including information on association strength, network topology, and temporal consistency. Candidate tracing paths are determined using the aforementioned tracing confidence assessment model; Based on the candidate tracing paths, a tracing confidence assessment model is established; The confidence score is obtained through the aforementioned traceability confidence assessment model.
5. The method for tracing work archives through multi-dimensional data integration according to claim 1, characterized in that, Using the tracing path and the confidence score, the first-level conditions, second-level conditions, and third-level conditions are judged, and the judgment results include: Using the aforementioned tracing path, a path consistency measurement function is established; Based on the confidence score and the path consistency measurement function, a comprehensive path quality index function is established; based on the comprehensive path quality index function, the comprehensive path quality index is obtained. Establish a dynamic acceptable domain model; based on the dynamic acceptable domain model, obtain a dynamic judgment threshold; Based on the comprehensive path quality index and the dynamic judgment threshold, the first-level conditions, the second-level conditions, and the third-level conditions are judged to obtain the judgment results.
6. The method for tracing work archives through multi-dimensional data integration according to claim 5, characterized in that, Based on the comprehensive path quality index and the dynamic judgment threshold, the first-level conditions, second-level conditions, and third-level conditions are judged, and the judgment results include: Based on the comprehensive path quality index and the dynamic judgment threshold, a first-level condition judgment is performed to determine whether the following conditions are met simultaneously: , , in, This is a comprehensive index of path quality. The threshold is dynamically determined, and C is the path confidence score. Based on the confidence threshold, , These are the minimum and maximum theoretical values of the consistency measure, respectively. Construct a path stability index, which satisfies the following relationship: , Where S is the path stability index. For the edge The associated feature values, For operation Abnormal index, This is the stability tolerance coefficient; The path stability index is used to perform a second-level conditional judgment to determine whether the following conditions are met: , in, As a stability benchmark, The coefficient of variation function, Let P represent the overall sensitivity of file a, and P be the current assessment path. This represents the maximum permissible value for the sensitivity coefficient of variation. Construct a business logic compliance function, which satisfies the following relationship: , Where B represents the business logic compliance, and n represents the number of path nodes. For business semantic similarity function, Let i be the i-th node in the path. For functions that conform to the type, Let j be the type of the j-th node. The set of allowed node types; Using the aforementioned business logic compliance, a third-level condition judgment is performed to determine whether the following conditions are met: , in, As a benchmark for business compliance, For similarity coefficient function, The set of neighbors of the starting node. The set of neighbors of the terminating node; The judgment results are obtained through the first-level condition judgment, the second-level condition judgment, and the third-level condition judgment. The judgment results include a set of valid traceability paths, a set of optimized paths, and an identifier of the type of judgment failure.
7. The method for tracing work archives through multi-dimensional data integration according to claim 1, characterized in that, Based on the judgment result, the traceability optimization process is initiated as follows: If the first, second, and third conditions are all met, the tracing result is deemed valid, and a detailed tracing report is output. If any one of the first, second, or third conditional checks fails, the traceability result is deemed invalid, and the traceability optimization process is initiated.
8. The method for tracing work archives through multi-dimensional data integration according to claim 7, characterized in that, The traceability optimization process includes: By utilizing the improved spatiotemporal coupling multidimensional visualization technology, the tracing path is jointly projected in the time dimension, relationship dimension, and risk dimension to obtain a multidimensional presentation model; Based on the aforementioned multi-dimensional presentation model, an adaptive optimization engine for tracing paths is constructed. Based on the aforementioned adaptive optimization engine, a dynamic evolution model for traceability credibility is established; A traceability pattern knowledge base is established based on the aforementioned dynamic evolution model.
9. The method for tracing work archives through multi-dimensional data integration according to claim 8, characterized in that, Based on the aforementioned multi-dimensional presentation model, the adaptive optimization engine for tracing paths includes: For paths that fail the first-level condition judgment, activate the feature enhancement optimization mode; For paths that fail the second-level condition judgment, activate the structure reinforcement optimization mode; For paths that fail the third-level condition judgment, initiate the semantic alignment optimization mode.
10. A multi-dimensional data integration work file traceability system, wherein the system uses the multi-dimensional data integration work file traceability method according to any one of claims 1 to 9, characterized in that, The system includes: The acquisition module is used to acquire work file data; The analysis module is used to extract first, second, and third file features based on the work file data; obtain a tracing path and confidence score based on the first, second, and third file features; use the tracing path and confidence score to judge the first, second, and third level conditions and obtain the judgment result; and initiate the tracing optimization process based on the judgment result and the improved spatiotemporal coupling multidimensional visualization technology. The output module is used to enable the traceability of work files according to the traceability optimization process.
Citation Information
Patent Citations
Archive management system and method based on data analysis
CN119830308A
Abnormal behavior identification method and system based on multi-dimensional data analysis
CN120012004A
Whole-course tracing method and system for group meal food
CN120806992A
Electronic archive credible evidence storage and intelligent tracing management system and method based on block chain
CN120850360A
File tracing method and device based on sensitive information, equipment and storage medium
CN120851015A