A method and system for constructing an archive record traceability model based on big data analysis
By constructing a traceability model for archival records and utilizing big data analytics, efficient traceability and accurate tracking of archival records are achieved, solving the problems of low efficiency and insufficient security in existing technologies and improving the efficiency and accuracy of archival management.
Patent Information
- Application Number
- CN202510444331.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Existing methods for tracing archival records rely on manual operations and simple databases, resulting in low efficiency, insufficient security, and difficulty in adapting to the needs of archival management in a big data environment. Furthermore, they cannot effectively identify abnormal operations or tampering.
A big data analytics-based archival record tracing model is constructed, comprising an archival tracing database, an identity verification layer, a data verification layer, and an evaluation and prediction layer. Through tracing coding, key verification, data classification, time-series causal correlation calculation, and model optimization, efficient tracing and accurate tracking of archival records are achieved.
It improves the efficiency and accuracy of archival record tracing, enhances data preprocessing capabilities and model adaptability, supports the terminal tracing needs of different archival record systems, and has universal applicability.
Smart Images

Figure CN119988317B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of archive tracing, in particular to a method and system for constructing an archive record tracing model based on big data analysis. BACKGROUND
[0002] With the rapid development of information technology, the number and types of archive records are increasing, and the security, integrity and traceability of archive information are becoming increasingly important. Effectively tracing and managing archive records to ensure the security and integrity of archive information has become a pressing problem.
[0003] Existing archive record tracing methods usually rely on manual operation and simple database queries, and their data storage lacks effective classification and partitioning mechanisms. When dealing with large-scale data, there are problems such as low efficiency, insufficient security and limited tracing ability, and they are easily affected by human factors, resulting in inaccurate tracing results. In addition, when faced with complex archive relationships and operations, traditional methods cannot effectively identify abnormal operations or tampering in archive records, lack effective prediction and evaluation mechanisms, and are difficult to adapt to the needs of archive management in a big data environment. The present application proposes a method and system for constructing an archive record tracing model based on big data analysis, aiming to realize regular storage, privacy encryption, data verification, efficient tracing and causal prediction of archive records through a multi-level technical architecture, and proposes a method for constructing an archive record tracing model based on big data analysis to overcome the shortcomings of existing archive record tracing methods, realize effective temporal and causal correlation of operation records and archive data, and improve the efficiency and accuracy of archive record tracing. SUMMARY
[0004] The present application aims to provide a method and system for constructing an archive record tracing model based on big data analysis.
[0005] To achieve the above-mentioned purpose, the present application is implemented according to the following technical solutions:
[0006] The present application comprises the following steps:
[0007] Constructing an archive tracing database: obtaining archive chain archive information, tracing and encoding the archive information, and storing the archive information in a data classification and partitioning manner;
[0008] Constructing an identity verification layer: setting a tracing key, verifying the identity of the operator according to the tracing key, and reading information or performing archive record operations;
[0009] Constructing a data verification layer: extracting target archive information to obtain target archive features, determining archive information integrity and nodes to be traced according to the target archive features;
[0010] constructing an evaluation prediction layer: determining a time sequence causal correlation degree vector according to the corresponding archive feature and the historical operation vector of the node to be traced, and constructing an archive record prediction model, inputting target archive information into the archive record prediction model to output an archive record prediction result; the archive record prediction result includes an archive operation vector and a causal vector;
[0011] optimizing an archive record tracing model: optimizing the archive record tracing model according to the deviation of the causal vector, and tracing the target archive according to the optimized archive record tracing model.
[0012] Further, the method for constructing the archive tracing database comprises:
[0013] obtaining the entire archive chain archive information, extracting the starting time stamp corresponding to the starting record operation, extracting the starting record personnel corresponding to the starting record operation, and tracing and encoding the archive information according to the starting time stamp, the starting record personnel and the total record number;
[0014] dividing the archive information into archive metadata and archive entity data, classifying the archive entity data to obtain external input entity data, archive receiving entity data and actual record entity data, storing the archive entity data in zones according to the entity data classification result, generating new archive metadata according to the entity data classification result and the zoned storage condition, classifying all archive metadata to obtain description metadata, management metadata and access metadata, and storing the metadata in zones according to the metadata classification result.
[0015] Further, the method for constructing the identity verification layer comprises the following steps:
[0016] setting a tracing key, performing identity verification and obtaining tracing permission according to the tracing key, and updating the archive information according to the specific operation and the corresponding document tracing code; the tracing key includes a primary tracing key and a secondary tracing key; the tracing permission includes a document access permission and a document record permission;
[0017] inputting a primary tracing key to obtain a document access permission, performing document viewing and generating an access record, and updating the access metadata by finding the corresponding archive information storage location through the tracing code; the primary tracing key is composed of an access code, a permission code and an identity code, the access code is used to verify the access permission, the permission code is used to determine the document access range of the tracing database, and the identity code is used to generate the access identity;
[0018] Or input the secondary traceability key to obtain the document recording permission, record the document, and find the corresponding archive information storage location to update the archive information through the traceability code; the secondary traceability key is composed of a record code, a permission code, and an identity code, the record code is used to verify the recording permission, the permission code is used to determine the document recording range of the traceability database, and the identity code is used to generate the record identity.
[0019] Further, the method for constructing the data verification layer comprises:
[0020] Input the primary traceability key to obtain the document access permission, find the corresponding management metadata according to the target archive traceability code, extract the target archive entity data storage location in the management metadata, and extract the target archive entity data according to the target archive entity data storage location;
[0021] Process the target archive entity data to obtain the target archive features, and divide the target archive features into target archive format features and target archive content features using hierarchical clustering; the target archive format features are obtained by processing the finally recorded target archive entity data; the target archive content features include external input content features, archive receiving content features, and actual recording content features;
[0022] Compare the target archive format features with the standard archive format features, determine the target archive integrity, and update the description metadata;
[0023] Define each archive recording operation on the archive chain as an archive recording node, calculate the similarity of the actual recording content feature vectors between adjacent archive recording nodes, and determine the to-be-traced node according to the similarity calculation result.
[0024] Further, the method for determining the time sequence causal correlation degree vector comprises:
[0025] Determine the data size and time operation parameters of each archive recording node on the archive chain according to the description metadata, determine the sliding window size and step according to the data size and time operation parameters, calculate the cross-correlation of each window using the sliding window method to obtain the time sequence correlation degree of the to-be-traced node;
[0026] Determine the recording vector of the to-be-traced node according to the operation record in the management metadata, calculate the deviation of the external input content feature vector and the archive receiving content feature vector to obtain the operation deviation vector of the to-be-traced node, perform feature fusion on the recording vector and the archive receiving content feature vector of the to-be-traced node to obtain a first action vector, perform feature cross on the first action vector and the operation deviation vector of the to-be-traced node to obtain a second action vector, calculate the content feature deviation between the to-be-traced node and the previous archive recording node to obtain a content feature change vector, and calculate the correlation degree between the content feature change vector and the second action vector of the to-be-traced node to obtain the causal correlation degree of the to-be-traced node.
[0027] Define traceability nodes i Content feature change vector , For tracing nodes i No. j Each content change feature element n Define the source node based on the number of feature elements. i Second action vector , For tracing nodes i No. j The expression for calculating the causal correlation degree of the second action element and the node to be traced is:
[0028] ,
[0029] ,
[0030] ,
[0031] in For tracing nodes i Content feature change vector With the second action vector The degree of causal relationship For relevance weight, For vectors and The cosine Pearson correlation coefficient, For vectors and Pearson correlation coefficient, For vectors and mutual information, To be the maximum mutual information, for and The joint probability density, for Marginal probability distribution, for Marginal probability distribution, For tracing nodes i No. j Operational deviation For tracing nodes i Mean of operational deviation;
[0032] The temporal correlation degree and causal correlation degree of the target file's traceable nodes are combined according to the time series to obtain the temporal correlation degree vector and the causal correlation degree vector. The temporal correlation degree vector and the causal correlation degree vector of the target file are then fused to obtain the temporal causal correlation degree vector.
[0033] Further, the method of outputting the archive record prediction result comprises:
[0034] The content matrix is composed of the target archive content features, the record matrix is composed of the record vectors of the nodes to be traced, the actual record content feature deviation of the initial and ending record nodes of the target archive is calculated to obtain the beginning and ending feature deviation vector, and the content matrix, the record matrix, the beginning and ending feature deviation vector, and the time sequence causal correlation degree vector of the target archive are combined to form an archive record tracing set;
[0035] An archive record prediction model is constructed according to the archive record tracing set, the archive record prediction model comprises a random forest algorithm, an elastic network regression algorithm, and an Adam optimizer; the archive record prediction model is used to predict the influence of the document record operation of each tracing node on the overall document content, and extract the operation record of the corresponding tracing node;
[0036] The random forest algorithm divides the archive record tracing set into training data and test data according to a ratio of 7:3;
[0037] The elastic network regression algorithm learns the relationship between the content matrix, the record matrix, the time sequence causal correlation degree, and the overall deviation vector, and performs regression analysis to predict the influence of each tracing node on the overall deviation, and outputs an operation causal vector;
[0038] The Adam optimizer is used to adaptively adjust the network learning rate;
[0039] The target archive information is input into the archive record prediction model to output an archive record prediction result; the archive record prediction result comprises an archive operation vector and a causal vector.
[0040] Further, the method of optimizing the archive record tracing model comprises:
[0041] The deviation between the predicted value and the actual value of the causal vector is calculated, the target function is determined according to the deviation of the causal vector and the archive operation vector, and the expression is:
[0042] ,
[0043] Wherein is the target function of the optimization of the archive record tracing model, is the predicted value of the i th element of the causal vector, i is the actual value of the i th element of the causal vector, is the dimension of the causal vector, i is the i th element of the archive operation vector, m is the mean value of the elements of the archive operation vector, is the i th element of the archive operation vector, i is the mean value of the elements of the archive operation vector, is the i th element of the archive operation vector, The file operation vector element variance is The hyperparameter of the explanatory term weight is The first k Parameter of the model, K The number of model parameters;
[0044] The search population is mapped by chaos, and the carrier operation is optimized chaos variable:
[0045] ,
[0046] ,
[0047] Wherein The first t +1th chaos variable, The first t th chaos variable, T The first t th chaos variable, The current search number, , The first i th chaos variable sequence of the first j th optimization variable after carrier operation, The upper limit of the first i th optimization variable, The lower limit of the first i th optimization variable, The first i th chaos variable corresponding to the first j th optimization variable;
[0048] The population fitness is calculated, and the explosion operator is introduced according to the population fitness to dynamically adjust the particle swarm:
[0049] ,
[0050] ,
[0051] Wherein The number of sparks produced by the firework i Explosion, The amplitude of the firework i Explosion, h The total number of sparks, The maximum explosion amplitude, , The maximum and minimum fitness of the population respectively, g The number of individuals currently participating in the calculation in the population, The fitness of the current position of the firework particle, A non-zero natural number;
[0052] Perform crossover and mutation operations on the population to update the particle velocity and position:
[0053] ,
[0054] ,
[0055] ,
[0056] in For the first i The causal bias particle in the first t Speed at +1 iteration d As a dimension, For the first i The causal bias particle in the first t Speed at the next iteration For the first t Causal weights at the next iteration , , A random number in the range (0,1). , Let be the acceleration constant. For the first i The causal bias particle in the first t Position at the next iteration For the first i The causal bias particle in the first t The historical best position at the next iteration. For the first t The historical optimal positions of all causal bias particles at the next iteration. , Divided into the first t The upper and lower bounds of the positions of all causal biased particles in the next iteration. For the first i The causal bias particle in the first t The position at +1 iteration, , Causal weights at the beginning and end;
[0057] The iteration continues until the objective function is minimized or the maximum number of iterations is reached, at which point the iteration stops.
[0058] The target archives are traced and recorded based on the optimized archive record tracing model.
[0059] Secondly, a record tracing system based on big data analysis includes:
[0060] An archive traceability database module is configured to obtain archive chain archive information, traceably encode the archive information, and data classify and partition store the archive information.
[0061] An identity verification module is configured to verify the identity of an operator according to a traceable key to read information or perform an archive record operation.
[0062] A data verification module is configured to extract target archive management metadata, extract target archive entity data according to the target archive management metadata, process the target archive entity data to obtain target archive features, and determine target archive integrity and determine a node to be traced according to the target archive features.
[0063] An evaluation prediction module is configured to determine a time sequence causal correlation degree vector according to corresponding archive features and a historical operation vector of the node to be traced, construct an archive record prediction model, input target archive information into the archive record prediction model, and output an archive record prediction result.
[0064] A model optimization module is configured to optimize the archive record traceability model according to a causal vector deviation.
[0065] An intelligent management module is configured to store, view, and manage the archive information and the archive record prediction result, generate a materialized view according to a timeline, the archive operation vector, and the causal vector, and show archive record history to a user.
[0066] The present application has the following beneficial effects:
[0067] Compared with the prior art, the present application has the following technical effects:
[0068] The present application can improve data preprocessing capability and enhance model adaptability in archive record traceability, thereby improving the efficiency and precision of archive record traceability, optimizing archive record traceability technology, greatly saving resources, improving work efficiency, realizing traceability of archive records, providing support for effective traceability and precise management of technical archives in various industries, and being suitable for terminal traceability requirements of archive records of different archive record traceability systems and different users, and having certain universality. BRIEF DESCRIPTION OF DRAWINGS
[0069] Figure 1 The present application is a step flowchart of a construction method of an archive record traceability model based on big data analysis. DETAILED DESCRIPTION
[0070] The application will be further described in the following specific examples, the illustrative examples of the application and the description used to explain the application, but not as a limitation of the application.
[0071] The application discloses a method and system for constructing an archive record tracing model based on big data analysis, which comprises the following steps:
[0072] As shown in Figure 1 In the embodiment, the following steps are included:
[0073] Constructing an archive tracing database: obtaining archive chain archive information, tracing coding the archive information, data classification and partition storage of the archive information;
[0074] Constructing an identity verification layer: setting a tracing key, verifying the identity of the operator according to the tracing key to read information or perform archive record operation;
[0075] Constructing a data verification layer: extracting target archive information to obtain target archive features, determining archive information integrity and a node to be traced according to the target archive features;
[0076] Constructing an evaluation and prediction layer: determining a time sequence causal correlation degree vector according to the archive features corresponding to the node to be traced and a historical operation vector, constructing an archive record prediction model, inputting the target archive information into the archive record prediction model, and outputting an archive record prediction result; the archive record prediction result comprises an archive operation vector and a causal vector;
[0077] Optimizing the archive record tracing model: optimizing the archive record tracing model according to the deviation of the causal vector, and tracing the target archive according to the optimized archive record tracing model.
[0078] In the embodiment, the method for constructing the archive tracing database comprises:
[0079] Obtaining the entire archive chain archive information, extracting a starting time stamp corresponding to a starting record operation, extracting a starting record personnel corresponding to the starting record operation, and tracing coding the archive information according to the starting time stamp, the starting record personnel and the total record number;
[0080] The archive information is divided into archive metadata and archive entity data, the archive entity data is classified to obtain external input entity data, archive receiving entity data and actual record entity data, the archive entity data is stored in zones according to the classification result of the entity data, and new archive metadata is generated according to the classification result of the entity data and the storage condition in the zones, the metadata classification is performed on all archive metadata to obtain description metadata, management metadata and access metadata, and the metadata is stored in zones according to the classification result of the metadata; the description metadata includes the operation time of each record node of the archive, the recorder and the size of the entity data; the management metadata includes the operation record of each operation point and the storage location of the entity data;
[0081] In the actual evaluation, taking the work archive management of a production line for one production task as an example, the management information including the operators and operation time is obtained: OP001-2024-06-01-08: 15, OP002-2024-06-01-10: 30, OP003-2024-06-01-13: 45, OP004-2024-06-02-08: 00, OP005-2024-06-02-11: 00, OP001-2024-06-03-14: 30, OP006-2024-06-04-09: 15, OP007-2024-06-05-16: 00, OP002-2024-06-06-07: 45, OP008-2024-06-07-10: 00, OP009-2024-06-08-13: 30;
[0082] Therefore, the traceability code is determined as OP001-OP009-2024-06-01-08: 15-2024-06-08-13: 30, and the archive metadata and the archive entity data are stored in zones, the external input entity data is the adjustment parameter input by the operator in the system, the archive receiving entity data is the adjustment parameter actually executed by the system, and the actual record entity data includes the production efficiency of the adjusted machine, the transmission speed of the production line, the number of workers, the final product yield, the sample quality of the finished product, the defective rate and the feedback of the workers.
[0083] In the embodiment, the method for constructing the identity verification layer includes the following steps:
[0084] A traceability key is set, identity verification is performed according to the traceability key and traceability permission is obtained, and archive information is updated according to specific operations and corresponding document traceability codes; the traceability key includes a first traceability key and a second traceability key; the traceability permission includes a document access permission and a document record permission;
[0085] Input the first-level traceability key to obtain document access permission, view the document, generate access records, find the corresponding archive information storage location through the traceability code, and update the access metadata; the first-level traceability key is composed of an access code, a permission code, and an identity code, the access code is used to verify access permission, the permission code is used to determine the document access range of the traceability database, and the identity code is used to generate an access identity; the access records include access identity, access time, and access file traceability code;
[0086] Or input the second-level traceability key to obtain document record permission, perform document record, and find the corresponding archive information storage location through the traceability code to update the archive information; the second-level traceability key is composed of a record code, a permission code, and an identity code, the record code is used to verify record permission, the permission code is used to determine the document record range of the traceability database, and the identity code is used to generate a record identity; the updated archive information includes traceability code, archive metadata, and archive entity data, the original archive entity data is retained, external input entity data, archive received entity data, and final record entity data are stored in a partition, description metadata is updated according to the record identity, record time, and final entity data size, management metadata is updated according to the specific operation of the document record, and the starting timestamp, starting operator, and total record number are obtained according to the updated description metadata and management metadata, and the traceability code is updated again;
[0087] In actual evaluation, taking the production line that continues to operate without completing the current production task as an example, the operator OP004 adjusts the transmission speed of the production line according to the worker work intensity feedback on June 9, 2024 at 8:00, inputs the traceability key Record1998OP004, verifies the record permission through the record code Record, determines the current recordable document range through the permission code 1998, generates a record identity through the identity code OP004, and updates the corresponding archive information according to the traceability code, and updates the traceability code to OP001-OP004-2024-06-01-08:15-2024-06-09-08:00;
[0088] Or taking the traceability of the target document as an example, the visitor inputs the traceability key Query0218QP001, verifies the access permission through the access code Query, determines the current accessible document range through the permission code 0218, generates an access identity through the identity code QP001, and generates access records according to the specific access time and the traceability code of the accessed document.
[0089] In this embodiment, the method for constructing the data verification layer includes:
[0090] The input first-level traceability key obtains a document access right, finds corresponding management metadata according to a target archive traceability code, extracts a target archive entity data storage location in the management metadata, and extracts target archive entity data according to the target archive entity data storage location;
[0091] The target archive entity data is processed to obtain target archive features; the method for processing the target archive entity data includes calculating the text importance of each word by performing text classification on document data using TF-IDF to obtain text features, performing speech recognition and classification on audio data using a deep neural network combined with a mel-frequency cepstrum coefficient to obtain audio features, extracting edge / color / gray / local information of image data using an image processing method to obtain image features, and extracting video image features of video data using a pre-trained convolutional neural network; the image processing method includes linear discriminant analysis, Canny edge detection, SIFT feature description, color histogram, and gray level co-occurrence matrix; the target archive features are divided into target archive format features and target archive content features using hierarchical clustering; the target archive format features are obtained by processing the target archive entity data of the final record; and the target archive content features include external input content features, archive receiving content features, and actual record content features;
[0092] The target archive format features are compared with standard archive format features to determine the integrity of the target archive and update the description metadata;
[0093] Each archive record operation on the archive chain is defined as an archive record node, the similarity of actual record content feature vectors between adjacent archive record nodes is calculated, and when the similarity of actual record content feature vectors between two adjacent archive record nodes is less than a similarity threshold, it is determined that the archive record operation of the later-occurring archive record node is abnormal, and the archive record node is set as a to-be-traced node;
[0094] In actual evaluation, the work information of a production task of a production line is recorded and traced to extract target archive format features, the target archive format features are compared with standard archive format features to determine the data integrity as 0.99, which is greater than the integrity threshold 0.95, and it is determined that the target archive data is complete;
[0095] The similarity of actual record content feature vectors between two adjacent archive record nodes is calculated as 0.92, 0.89, 0.93, 0.96, 0.90, 0.95, 0.97, 0.87, 0.94, and 0.98 in time sequence, adjacent archive record nodes with a content feature vector similarity less than a similarity threshold 0.93 are screened, and the later-occurring archive record node in the four groups of archive record nodes is determined as a to-be-traced node.
[0096] In this embodiment, the method for determining the temporal causal correlation vector includes:
[0097] Based on the metadata description, determine the data size and time operation parameters of each record node in the archive chain. Based on the data size and time operation parameters, determine the sliding window size and step size. Use the sliding window method to calculate the cross-correlation of each window to obtain the temporal correlation degree of the node to be traced.
[0098] Taking the tracing of this archival record as an example, the tracing node and the preceding and following record nodes are determined as a sliding window. The window size is set according to the size of the archival information corresponding to the three record nodes, and the sliding window step size is set according to the time interval between two tracing nodes. The corresponding step sizes are 3h15min, 48h45min, and 65h15min.
[0099] The record vector of the node to be traced is determined based on the operation records in the management metadata. The deviation between the external input content feature vector and the archive received content feature vector is calculated to obtain the operation deviation vector of the node to be traced. The record vector of the node to be traced and the archive received content feature vector are fused to obtain the first action vector. The first action vector of the node to be traced and the operation deviation vector are cross-referenced to obtain the second action vector. The deviation between the actual recorded content feature of the node to be traced and the previous archive record node is calculated to obtain the content feature change vector. The correlation between the content feature change vector of the node to be traced and the second action vector is calculated to obtain the causal correlation degree of the node to be traced.
[0100] Define traceability nodes i Content feature change vector , For tracing nodes i No. j Each content change feature element n Define the source node based on the number of feature elements. i Second action vector , For tracing nodes i No. j The expression for calculating the causal correlation degree of the second action element and the node to be traced is:
[0101] ,
[0102] ,
[0103] ,
[0104] in For tracing nodes i Content feature change vector With the second action vector The degree of causal relationship For relevance weight, For vectors and The cosine Pearson correlation coefficient, For vectors and Pearson correlation coefficient, For vectors and mutual information, To be the maximum mutual information, for and The joint probability density, for Marginal probability distribution, for Marginal probability distribution, For tracing nodes i No. j Operational deviation For tracing nodes i Mean of operational deviation;
[0105] The temporal correlation degree and causal correlation degree of the target file's traceable nodes are combined according to the time series to obtain the temporal correlation degree vector and the causal correlation degree vector. The temporal correlation degree vector and the causal correlation degree vector of the target file are then fused to obtain the temporal causal correlation degree vector.
[0106] In the actual assessment, the time-series causal correlation vector obtained is (0.45, 0.85, 0.80, 0.35, 0.30, 0.88, 0.50, 0.25, 0.90, 0.40, 0.20).
[0107] In this embodiment, the method for outputting the predicted results of the archive records includes:
[0108] The content matrix is composed of the content features of the target archive, and the record matrix is composed of the record vectors of the nodes to be traced. The deviation of the actual record content features of the initial and final record nodes of the target archive is calculated to obtain the beginning and end feature deviation vector. The content matrix, record matrix, beginning and end feature deviation vector and temporal causal correlation vector of the target archive are combined to form the archive record tracing set.
[0109] An archival record prediction model is constructed based on the archival record traceability set. The archival record prediction model includes the random forest algorithm, the elastic network regression algorithm, and the Adam optimizer. The archival record prediction model is used to predict the impact of document record operations at each traceability node on the overall document content and to extract the operation records of the corresponding traceability node.
[0110] The random forest algorithm divides the archival record source set into training data and test data in a 7:3 ratio;
[0111] The elastic network regression algorithm learns the relationship between the content matrix, record matrix, temporal causal correlation degree and overall deviation vector, and performs regression analysis to predict the impact of each source node on the overall deviation, outputting the operational causal vector;
[0112] The Adam optimizer is used to adaptively adjust the network learning rate.
[0113] The target archive information is input into the archive record prediction model, which outputs the archive record prediction result; the archive record prediction result includes archive operation vector and causal vector.
[0114] In this embodiment, the method for optimizing the archival record tracing model includes:
[0115] Calculate the deviation between the predicted and actual values of the causal vector. Based on the causal vector deviation and the file operation vector, determine the objective function, expressed as:
[0116] ,
[0117] in The objective function for optimizing the archival record traceability model. For the causal vector, the first i Predicted values for each element For the causal vector, the first i The actual value of each element. m As a causal vector dimension, For the file operation vector number i One element, The mean of the elements in the file operation vector. The variance of the vector elements in the file operation. For the hyperparameters of the explanatory term weights, For the model's first k One parameter, K This refers to the number of model parameters;
[0118] The search population is subjected to chaotic mapping, and carrier operations are performed to optimize the chaotic variables:
[0119] ,
[0120] ,
[0121] in For the first t +1 chaotic variable, For the first t A chaotic variable, T A chaotic variable, t This represents the current number of searches. For chaotic random numbers, , For the first i The first optimization variable is the... j The optimized variable after carrier operation in a sequence of chaotic variables. For the first i The upper bound of the optimization variables. For the first i The lower bound of each optimization variable. For the first i The corresponding optimization variable is the first j One chaotic variable;
[0122] Calculate the population fitness, and then introduce an explosion operator to dynamically adjust the particle swarm based on the population fitness.
[0123] ,
[0124] ,
[0125] in For fireworks i The number of sparks produced by the explosion For fireworks i The magnitude of the explosion h For the total number of sparks, For the maximum explosion amplitude, , These represent the maximum and minimum fitness of the population, respectively. g This represents the number of individuals currently participating in the calculation within the population. The fitness of the current position of the firework particle. It is a non-zero natural number;
[0126] Perform crossover and mutation operations on the population to update the velocity and position of the particles:
[0127] ,
[0128] ,
[0129] ,
[0130] in For the first i The causal bias particle in the first t Speed at +1 iteration d As a dimension, For the first i The causal bias particle in the first t Speed at the next iteration For the first t Causal weights at the next iteration , 、 is a random number of (0, 1), 、 is an acceleration constant, is the position of the th causal bias particle at the th iteration, i is the position of the th causal bias particle at the th iteration, t is the history optimal position of the th causal bias particle at the th iteration, is the history optimal position of the th causal bias particle at the th iteration, i is the history optimal position of all causal bias particles at the th iteration, t is the history optimal position of all causal bias particles at the th iteration, is the upper and lower bounds of the position of all causal bias particles at the th iteration, t is the upper and lower bounds of the position of all causal bias particles at the th iteration, 、 is the position of the th causal bias particle at the th iteration, t 、 is the position of the th causal bias particle at the th iteration, i 、 t is the causal weight at the beginning and end; 、 is the causal weight at the beginning and end;
[0131] iterating until the target function is minimized or the maximum number of iterations is reached;
[0132] According to the optimized archive record tracing model, the target archive is recorded and traced;
[0133] In the actual evaluation, the tracing code of the target archive is obtained, the target archive is recorded and traced according to the optimized archive record tracing model, and the corresponding causal vector is (0.357, 0.786, 0.714, 0.268, 0.225, 0.816, 0.408, 0.163, 0.833, 0.3265, 0.1020). The change of the final entity data caused by the archive record operation corresponding to the tracing point is: 1. OP002-2024-06-01-10:30 record point adjustment machine production rate causes production efficiency to increase from 90% to 92%, and the defect rate to decrease from 2% to 1.8%, 2. OP003-2024-06-01-13:45 record point adjusts the temperature of the input content from 180°C to 182°C, and the defect rate further decreases to 1.5%, 3. OP001-2024-06-03-14:30 record point optimizes the final product, causing the sampling quality of the product to improve to the superior level, and the defect rate to decrease to 1%, 4. OP002-2024-06-06-07:45 record point increases the transmission speed of the production line to 25 m / min, and the number of production workers increases by 0.2 person / m.
[0134] In a second aspect, an archive record tracing system based on big data analysis includes:
[0135] An archive traceability database module is configured to obtain archive chain archive information, traceably encode the archive information, and data classify and partition store the archive information;
[0136] An identity verification module is configured to verify the identity of an operator according to a traceable key to read information or perform an archive record operation;
[0137] A data verification module is configured to extract target archive management metadata, extract target archive entity data according to the target archive management metadata, process the target archive entity data to obtain target archive features, determine target archive integrity, and determine a node to be traced according to the target archive features;
[0138] An evaluation prediction module is configured to determine a time sequence causal correlation degree vector according to archive features corresponding to the node to be traced and a historical operation vector, construct an archive record prediction model, input target archive information into the archive record prediction model, and output an archive record prediction result;
[0139] A model optimization module is configured to optimize the archive record traceability model according to a causal vector deviation;
[0140] An intelligent management module is configured to store, view, and manage the archive information and the archive record prediction result, generate a materialized view according to a timeline, the archive operation vector, and the causal vector, and show archive record history to a user.
[0141] The above merely describes preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for constructing an archival record traceability model based on big data analysis, characterized in that, The method comprises the following steps: S1, constructing an archive traceability database: obtaining archive chain archive information, traceability coding the archive information, data classification and partition storage of the archive information; S2, constructing an identity verification layer: setting a traceability key, verifying the identity of the operator according to the traceability key to read information or perform archive record operations; S3, constructing a data verification layer: extracting target archive information to obtain target archive features, determining archive information integrity and a node to be traced according to the target archive features; S4, constructing an evaluation prediction layer: determining a time sequence causal correlation degree vector according to the archive features corresponding to the node to be traced and a historical operation vector, and constructing an archive record prediction model, inputting the target archive information into the archive record prediction model to output an archive record prediction result; the archive record prediction result comprises an archive operation vector and a causal vector; S5, optimizing the archive record traceability model: optimizing the archive record traceability model according to the deviation of the causal vector, and performing record traceability on the target archive according to the optimized archive record traceability model; The method for determining the time sequence causal correlation degree vector comprises: According to the data size and time operation parameters of each archive record node on the archive chain determined according to the description metadata, the size of the sliding window and the step are determined according to the data size and time operation parameters, the cross-correlation of each window is calculated by using the sliding window method to obtain the time sequence correlation degree of the node to be traced; According to the operation record in the management metadata, the record vector of the node to be traced is determined, the deviation between the external input content feature vector and the archive received content feature vector is calculated to obtain the operation deviation vector of the node to be traced, the record vector of the node to be traced and the archive received content feature vector are fused to obtain a first action vector, the first action vector of the node to be traced and the operation deviation vector are cross-featured to obtain a second action vector, the content feature change vector is obtained by calculating the actual record content feature deviation between the node to be traced and the previous archive record node, and the causal correlation degree of the node to be traced is obtained by calculating the correlation degree between the content feature change vector of the node to be traced and the second action vector. Definition of the traceability node i Content feature change vector , For the traceability node i The first j Content change feature element, n The number of feature elements, definition of the traceability node i Second action vector , For the traceability node i The first j Second action element, the traceability node causal correlation degree calculation expression is: , , , in For tracing nodes i Content feature change vector With the second action vector The degree of causal relationship For relevance weight, For vectors and The cosine Pearson correlation coefficient, For vectors and Pearson correlation coefficient, For vectors and mutual information, To be the maximum mutual information, for and The joint probability density, for Marginal probability distribution, for Marginal probability distribution, For tracing nodes i No. j Operational deviation For tracing nodes i Mean of operational deviation; According to the time sequence combination of the target archive, the time sequence correlation degree and the causal correlation degree of the node to be traced are obtained to obtain the time sequence correlation degree vector and the causal correlation degree vector, and the time sequence correlation degree vector and the causal correlation degree vector of the target archive are fused to obtain the time sequence causal correlation degree vector.
2. The method as claimed in claim 1, wherein, The method for constructing the archive traceability database comprises: Obtaining the entire archive chain archive information, extracting the starting timestamp corresponding to the starting record operation, extracting the starting record personnel corresponding to the starting record operation, and traceability coding the archive information according to the starting timestamp, the starting record personnel and the total record number; The file information is divided into file metadata and file entity data, the file entity data is classified to obtain external input entity data, file receiving entity data and actual record entity data, the file entity data is stored in zones according to the classification result of the entity data, and new file metadata is generated according to the classification result of the entity data and the storage condition in the zones, the metadata classification is performed on all the file metadata to obtain description metadata, management metadata and access metadata, and the metadata is stored in zones according to the classification result of the metadata.
3. The method as claimed in claim 1, wherein, The method for constructing the identity authentication layer comprises the following steps: The traceability key is set, the identity is authenticated according to the traceability key, the traceability permission is obtained, and the file information is updated according to the specific operation and the corresponding document traceability code; the traceability key comprises a first traceability key and a second traceability key; the traceability permission comprises a document access permission and a document record permission; The first traceability key is input to obtain the document access permission, the document is viewed, and the access record is generated; the corresponding file information storage location is found through the traceability code, and the access metadata is updated; the first traceability key is composed of an access code, a permission code and an identity code; the access code is used for verifying the access permission, the permission code is used for determining the document access range of the traceability database, and the identity code is used for generating the access identity; Or the second traceability key is input to obtain the document record permission, the document is recorded, and the corresponding file information storage location is found through the traceability code to update the file information; the second traceability key is composed of a record code, a permission code and an identity code; the record code is used for verifying the record permission, the permission code is used for determining the document record range of the traceability database, and the identity code is used for generating the record identity.
4. The method as claimed in claim 1, wherein, The method for constructing the data verification layer comprises: The first traceability key is input to obtain the document access permission, the corresponding management metadata of the target file is found according to the target file traceability code, the storage location of the target file entity data in the management metadata is extracted, and the target file entity data is extracted according to the storage location of the target file entity data; The target file entity data is processed to obtain the target file feature, and the target file feature is divided into a target file format feature and a target file content feature by adopting hierarchical clustering; the target file format feature is obtained by processing the finally recorded target file entity data; the target file content feature comprises an external input content feature, a file receiving content feature and an actual record content feature; The target file format feature is compared with a standard file format feature, the target file integrity is determined, and the description metadata is updated; An archive record node is defined for each archive record operation on the file chain, the similarity of the actual record content feature vectors between adjacent archive record nodes is calculated, and the node to be traced is determined according to the similarity calculation result.
5. The method as claimed in claim 1, wherein, The method for outputting the archive record prediction result comprises: The target file content feature is used to form a content matrix, the record vector of the node to be traced is used to form a record matrix, the actual record content feature deviation of the initial and end record nodes of the target file is calculated to obtain a start-end feature deviation vector, and the content matrix, the record matrix, the start-end feature deviation vector and the time sequence causal correlation degree vector of the target file are used to form an archive record traceability set; The archive record prediction model is used to predict the influence of the document record operation of each traceability node on the overall document content, and extract the operation record of the corresponding traceability node; The random forest algorithm divides the archive record traceability set into training data and test data according to 7:3; The elastic network regression algorithm learns the relationship between the content matrix, the record matrix, the time sequence causal correlation degree and the overall bias vector, and performs regression analysis to predict the influence of each traceability node on the overall bias, and outputs the operation causal vector; The Adam optimizer is used to adaptively adjust the network learning rate; The target archive information is input into the archive record prediction model to output the archive record prediction result; The archive record prediction result includes the archive operation vector and the causal vector.
6. The method as claimed in claim 1, wherein, The method for optimizing the archive record traceability model comprises: Calculate the deviation of the causal vector prediction value and the actual value, and determine the target function according to the causal vector deviation and the archive operation vector, the expression is: , in The objective function for optimizing the archival record traceability model. For the causal vector, the first i Predicted values for each element For the causal vector, the first i The actual value of each element. m As a causal vector dimension, For the file operation vector number i One element, The mean of the elements in the file operation vector. The variance of the vector elements in the file operation. For the hyperparameters of the explanatory term weights, For the model's first k One parameter, K This refers to the number of model parameters; Chaotic mapping is performed on the search population, and carrier operation is performed to optimize chaotic variables: , , wherein is the t +1th chaotic variable, is the t th chaotic variable, T th chaotic variable, t is the current search number, is the chaotic random number, , is the i th chaotic variable sequence of the j th optimization variable after carrier operation, is the upper limit of the i th optimization variable, is the lower limit of the i th optimization variable, is the i th chaotic variable corresponding to the j th optimization variable; Calculate the fitness of the population, and introduce the explosion operator to dynamically adjust the particle swarm according to the fitness of the population: , , wherein is the firework i is the number of sparks produced by the explosion, is the firework i is the amplitude of the explosion, h is the total number of sparks, is the maximum amplitude of the explosion, , are the maximum and minimum fitness of the population, respectively, g is the number of individuals currently participating in the calculation in the population, is the fitness of the current position of the firework particle, is a non-zero natural number; Perform cross mutation operation on the population, and update the speed and position of the particles: , , , wherein is the number of causal bias particles, i is the velocity of the t th causal bias particle at the d th iteration, is the dimensionality, i is the velocity of the t th causal bias particle at the th iteration, t is the causal weight at the th iteration, , is a random number between (0,1), , is an acceleration constant, is the position of the i th causal bias particle at the t th iteration, is the history optimal position of the i th causal bias particle at the t th iteration, is the history optimal position of all causal bias particles at the t th iteration, , is the upper and lower bound of all causal bias particles positions at the t th iteration, is the position of the i th causal bias particle at the t +1th iteration, , is the causal weight at the beginning and end. Iterate continuously until the target function is minimized or the maximum number of iterations is reached; According to the optimized archive record traceability model, record traceability is performed on the target archive.
7. A big data analytics based archival record provenance system to perform the method of any one of claims 1-6, characterized in that, It includes: An archive traceability database module is used to obtain archive chain archive information, traceability code the archive information, and store the archive information in data classification and partition; An identity verification module is used to verify the identity of the operator according to the traceability key to read information or perform archive record operation; A data verification module is used to extract target archive management metadata, extract target archive entity data according to the target archive management metadata, process the target archive entity data to obtain target archive features, and determine the target archive integrity and the to-be-traced node according to the target archive features; An evaluation prediction module is used to determine the time sequence causal correlation degree vector according to the archive features corresponding to the to-be-traced node and the historical operation vector, and construct an archive record prediction model, input target archive information into the archive record prediction model to output archive record prediction results; A model optimization module is used to optimize the archive record traceability model according to the causal vector deviation; An intelligent management module is used to store, view and manage the archive information and the archive record prediction results, generate a materialized view according to the timeline, the archive operation vector and the causal vector, and show the archive record history to the user.
Citation Information
Patent Citations
Tracing method and system for fault root cause of Internet of Things equipment, and storage medium
CN117743843A
Data tracing method and system based on big data and block chain multi-dimensional features
CN119557607A